Evolutionary multi-agent runtime that breeds, evaluates, and improves autonomous agents across reproducible epochs to converge on optimization of a goal.
-
Updated
Aug 9, 2026 - Python
Evolutionary multi-agent runtime that breeds, evaluates, and improves autonomous agents across reproducible epochs to converge on optimization of a goal.
Core engine behind Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations
Benchmark, observe, and govern tool-using AI agents before production.
Frontend for Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations
A benchmark measuring AI agent reliability on paired Amharic and English tasks in simulated Ethiopian workflows.
A Multiagent System for Adversarial Evaluation and Safety Checking of Other Computer Agents.
Turn agent telemetry into eval jobs — outcome, quality, and spend — correlated with business outcomes per agent run.
Python SDK for Calibrate's API
Run AI agent evaluations in your CI/CD pipeline using Calibrate and catch regressions
Trustra Agent Eval Harness Open-core evaluation and tracing harness for AI agents. Run test datasets against an agent, score results, capture production traces, and build a tamper-evident hash-chained record of what an agent actually did.
Open-source agent assurance: turn a PRD and agent URL into frozen test suites with evidence-backed PASS/FAIL/UNVERIFIABLE verdicts. Web console, REST API, SQLite. Domain packs for regulated teams (fintech, insurance, health).
To associate your repository with the agent-evaluation-tools topic, visit your repo's landing page and select "manage topics."