Lynn GitHub 镜像仓 · Primary repository: https://github.com/MerkyorLynn/Lynn · Downloads: https://download.merkyorlynn.com/download.html
-
Updated
Sep 11, 2026 - TypeScript
Lynn GitHub 镜像仓 · Primary repository: https://github.com/MerkyorLynn/Lynn · Downloads: https://download.merkyorlynn.com/download.html
Meta-harness optimization loop wired onto Islo sandboxes. POC: 0/5→5/5 in four proposer steps. Built on islo.dev.
Observe an agent run and GroundEval drafts the policy and diagram for you, no hand-written policy required, then scores what it checked, what it skipped, and what it wasn't allowed to touch.
Scenario Testing for AI Agents
Live, open-source benchmark for comparing AI coding agents on real GitHub issues
AI-operated company. Building agent-friend: universal tool adapter for AI agents. @tool → OpenAI, Claude, Gemini, MCP. Live 24/7 on Twitch.
Agent 测试转型手册:资深自动化测试工程师的 AI QE 学习路径 + 练手代码
Framework-agnostic evaluation harness for Go — test your MCP servers and AI agents with scored, CI-ready checks.
Project page for Meta-harness on Islo (POC). https://zozo123.github.io/meta-harness-on-islo-page/
Binary-criteria evaluation harness for Claude skills with planned extension to plugins, agents, and MCP servers. Score every change yes/no across 7 layers — package integrity, trigger quality, functional quality, regression protection, baseline value, model variance, rollout safety. Never gradients.
A reasoning benchmark runner for comparing LLMs as OpenClaw agents use them. 52 prompts, 3 eval sets, 11 traps, LLM-as-judge, tier-based leaderboard.
Transcript-first evaluation tool for comparing coding-agent sessions across Codex, Claude Code, and Pi.
Vendor-neutral research umbrella for measuring AI plugin, agent, and MCP server quality across CLI runtimes (Claude Code, Gemini CLI, Copilot CLI, Codex CLI).
A curated list of benchmarks, harnesses, leaderboards, and tools for evaluating AI coding agents.
检测 AI Agent 代理指标与真实业务结果背离的审计工具,面向增长与运营团队
PandaProbe harness turns agent failures into fixes
A durable, long-running agent that improves and evaluates other agents.
开源通用 AI Agent 真实任务评测 · 同 Prompt、客观开奖、评分细则全公开 | Open-source evaluation of general-purpose AI Agents on real-world tasks with verifiable outcomes — by PingWest / 硅星人
Field guide to building production-grade evaluation infrastructure for LLM agent systems. Based on real production work at Airbnb (BPI Virtual Analyst), Shell (NLP), and contributions to LangChain, LiveKit, and Ragas.
Turn agent telemetry into eval jobs — outcome, quality, and spend — correlated with business outcomes per agent run.
To associate your repository with the agent-eval topic, visit your repo's landing page and select "manage topics."