Skip to content

Changelog

All notable changes to this project are documented in this file.

The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.

[Unreleased]

Features

  • core: Own the score mapping and ports the libraries share (#32)
  • online: Run evaluators on live traffic, sampled and within a budget (#21)
  • measures: Define task completion, drop-off and rewrites generically (#20)
  • hf: Publish, pin and import datasets on the Hugging Face Hub (#19)
  • langfuse: Add the Langfuse score sink and experiment tracker (#18)
  • langfuse: Add the Langfuse dataset store (#16)
  • decision: Calibrate decision thresholds on a dataset (#15)
  • decision: Add DecisionEvaluator on a decision-only view (#14)
  • dspy: Save and load versioned judges (#13)
  • dspy: Train judges with GEPA (#12)
  • dspy: Add DspyJudge with signatures derived from its types (#11)
  • core: Add the optimizer port, fallback composition and measurement (#10)
  • core: Add agreement and calibration metrics (#9)
  • core: Add scores, the score sink and experiment tracker ports and their adapters (#8)
  • core: Add the dataset store and feedback source ports and their adapters (#7)
  • core: Add examples, datasets, deterministic splits and formatters (#5)
  • core: Add verdicts, field kinds, the evaluator protocol and function evaluators (#4)

Bug fixes

  • Read people's feedback in rates, type experiment verdicts, and drop difflib's autojunk (#33) (breaking)
  • Build generic results through their parametrized alias on Python 3.12 (#30)
  • langfuse: Match the dataset store and its fake to a real Langfuse (#29)
  • core: Never shorten a list that windowing alone can fit (#25)

Documentation

  • Regenerate the changelog, and have the site regenerate it on every build (#28)
  • Render lists on the site as GitHub does, and gate scripts/ like the library (#27)
  • Publish the site at evalr.alexnodeland.com from main (#24)
  • Add the documentation site and brand (#23)
  • Describe v0.1 in the README and refresh the changelog (#22)
  • adr: Build evalr as ports and adapters (#6)
  • Add the evalr design: RFC-0001 and ADRs (#1)

Refactoring

  • One set of Langfuse score adapters, scores that check their own values, and one docs build (#34) (breaking)

Miscellaneous

  • Lay the foundation for the evalr library (#3)
  • Initial commit