Skip to content
#

ai-evals

Here are 76 public repositories matching this topic...

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.

  • Updated Aug 5, 2026
  • Python

Core engine behind Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations

  • Updated Jul 29, 2026
  • JavaScript
agent-workspace-architecture

Agent workspace architecture — the reference implementation of an agent-ready memory layer, demonstrated end-to-end in Claude Code: roles library, typed memory, hooks, scheduled agents, self-audits, loop selection, measurement-gated self-improvement. Interactive tour, fork-ready samples.

  • Updated Aug 3, 2026
  • Python

Eval-first AI agent that triages property maintenance emails. The real work is the eval system around it: trace-driven error analysis, code graders and validated LLM-as-judge (TPR/TNR), component and end-to-end evals, a failure taxonomy, and a CI regression gate. LangGraph, FastAPI, Langfuse.

  • Updated Jun 7, 2026
  • Python

Frontend for Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations

  • Updated Jul 29, 2026
  • TypeScript

Improve this page

Add a description, image, and links to the ai-evals topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the ai-evals topic, visit your repo's landing page and select "manage topics."

Learn more