Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
-
Updated
Aug 5, 2026 - Python
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
Operational doctrine for practical AI systems design.
Benchmarking Open-Ended Inference Optimization by AI Agents
pytest for LLM apps: record API calls once, replay them forever, and test meaning without flaky live runs.
Evaluation Infrastructure for AI Agents
Core engine behind Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations
Learn to evaluate AI products for production — 21 hands-on lessons on evals, metrics, fairness, agents, red teaming, and release decisions for working PMs.
Collection of frameworks and tools for AI evalations, including tool-use, agentic AI, MCP, and multimodal
Governed multi-LLM Responsible AI control plane for prior-authorization decision support, with PHI-safe audit trails, HITL review, evidence packets, and deterministic governance evals.
Agent workspace architecture — the reference implementation of an agent-ready memory layer, demonstrated end-to-end in Claude Code: roles library, typed memory, hooks, scheduled agents, self-audits, loop selection, measurement-gated self-improvement. Interactive tour, fork-ready samples.
Eval-first AI agent that triages property maintenance emails. The real work is the eval system around it: trace-driven error analysis, code graders and validated LLM-as-judge (TPR/TNR), component and end-to-end evals, a failure taxonomy, and a CI regression gate. LangGraph, FastAPI, Langfuse.
Benchmark LLM jailbreak resilience across providers with standardized tests, adversarial mode, rich analytics, and a clean Web UI.
Public research on LLM internals, Jacobian lenses, SAE steering, nonlinear dynamics, evaluation, and inspectable AI systems.
Frontend for Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations
Offline, auditable benchmark for one-shot LLM market decisions.
Open-source toolkit for assessing whether an AI workflow is ready for production: governance, RAG quality, evals, observability, human review, cost, risk, and business value.
Playbook for PMs shipping AI products with PRDs, evals, HITL, launch gates, cost, and observability.
AI agents built spaCy model recommenders, then a fresh AI peer-review panel rejected all three
Add a description, image, and links to the ai-evals topic page so that developers can more easily learn about it.
To associate your repository with the ai-evals topic, visit your repo's landing page and select "manage topics."