Feedback sources¶
People's feedback is what evalr's judges learn from and are measured against. It is recorded by the application people use: artifactr and reflexr record typed feedback on threads, turns, runs and chains, and your own application may keep ratings in a table. A feedback source turns one type of that feedback into evalr examples, each an input with the verdict a person gave it.
FeedbackSource is one of evalr's ports: a small protocol that the libraries implement in their [evals] extras, so evalr never imports them (ADR-0006).
The examples on this page use the Thread and Helpfulness types from Getting started.
The port¶
A feedback source has an input_type, a verdict_type (the feedback type) and an examples() method that yields Examples asynchronously. It promises three things, which check_feedback_source checks:
- Every example has a verdict: the feedback itself.
- Ids are stable, derived from the feedback, so splits and stores treat an example the same way every time it is collected.
- Iterating again yields the same examples, until more feedback arrives.
collect gathers a source's examples into a dataset, ready to split, save and train on:
from evalr import Example, collect
from evalr.memory import InMemoryFeedbackSource
source = InMemoryFeedbackSource(
[
Example[Thread, Helpfulness](
id="fb_0192",
input=Thread(request="I was charged twice", reply="I refunded the duplicate charge."),
verdict=Helpfulness(rating=5, resolved=True, reason="Fixed at once"),
),
Example[Thread, Helpfulness](
id="fb_0193",
input=Thread(request="My order is late", reply="Orders arrive within a week."),
verdict=Helpfulness(rating=2, resolved=False, reason="It ignored my order"),
),
],
input_type=Thread,
verdict_type=Helpfulness,
)
feedback = await collect("helpfulness", source, description="People's ratings of support replies")
train, validate = feedback.split(0.2)
InMemoryFeedbackSource yields a fixed list, which suits tests and prototypes.
reflexr¶
reflexr's [evals] extra ships reflexr.evals.LogFeedbackSource, a feedback source over a workspace's log. Each piece of one feedback type becomes an example: the verdict is the feedback, and the input is built by your application from what the feedback is about, since only it knows what its evaluators judge.
Given a reflexr workspace in which people give TriageQuality feedback (a reflexr Feedback type) on the runs of a triage rule:
from pydantic import BaseModel
from evalr import collect
from reflexr.evals import FeedbackContext, LogFeedbackSource
class TriageRun(BaseModel):
"""What a triage judge reads: the service, and what the run did about it."""
service: str
emitted: int
def triage_run(context: FeedbackContext[TriageQuality]) -> TriageRun:
assert context.run is not None # feedback on a run comes with the run's record
return TriageRun(
service=str(context.run.run.scope["service"]), emitted=len(context.run.emitted)
)
source = LogFeedbackSource(
workspace,
feedback_type=TriageQuality, # the verdict type
input_type=TriageRun,
input=triage_run,
targets={"run"},
)
triage = await collect("triage-quality", source)
- The input builder receives a
FeedbackContext: thefeedback_givenenvelope (who gave it, and when), the feedback, and what it is about. Feedback on a run or a firing comes with aRunRecordof the run, the events that made its rule fire, and the events it emitted; feedback on a causal chain comes with the chain's envelopes. The builder may be async, to read more of the log. - Ids are the feedback's event ids, and an example of feedback on a run carries the trace of the run's latest attempt, so datasets link back to the traces they came from.
- Metadata records the tenant, the workspace, the kind of target, who gave the feedback and its position in the log.
targets=limits the source to feedback on some kinds of target (run,firing,chain); every kind by default.
The same extra runs evalr evaluators as rules (EvaluatorAction), replays examples against a candidate agent, graph or model as an experiment task (replay_task), and computes rule-level measures from the log. See reflexr's architecture and its ADR-0020.
artifactr¶
artifactr records typed feedback from people and evaluators on artifact versions, threads, turns and messages (artifactr's evaluation guide). Its [evals] extra, planned in artifactr's RFC-0002, will provide the same pieces: datasets from the log, with each piece of feedback's target as context, experiment tasks that replay a turn, online evaluators and the end-to-end measures. Until it ships, an artifactr application can write a source of its own that reads the feedback_given events in a workspace's log, the way the next section reads a table of ratings.
A source of your own¶
Anything that yields examples with verdicts is a feedback source. Here, ratings an application keeps in its own database:
from collections.abc import AsyncIterator
from dataclasses import dataclass
@dataclass
class Rating:
"""A row of the application's ratings table."""
id: int
request: str
reply: str
stars: int
solved: bool
comment: str | None
class RatingsSource:
"""People's ratings of replies, as examples."""
input_type = Thread
verdict_type = Helpfulness
def __init__(self, rows: list[Rating]) -> None:
self.rows = rows # in an application, a query
async def examples(self) -> AsyncIterator[Example[Thread, Helpfulness]]:
for row in self.rows:
yield Example[Thread, Helpfulness](
id=f"rating-{row.id}", # stable: derived from the rating
input=Thread(request=row.request, reply=row.reply),
verdict=Helpfulness(rating=row.stars, resolved=row.solved, reason=row.comment),
)
ratings = RatingsSource([Rating(7, "Cancel my plan", "Done.", 4, True, None)])
from_ratings = await collect("ratings", ratings)
Before relying on a source, check it against the port's contract with check_feedback_source, as reflexr does for LogFeedbackSource (Testing with the contracts).