Skip to content
#

reliability-testing

Here are 19 public repositories matching this topic...

End-to-end reliability testing for multi-step AI agent workflows. Chaos testing, cascade failure detection, E2E benchmarks. What DeepEval/Langfuse/LangSmith can't test. Works with LangChain, CrewAI, AutoGen, OpenAI. npm install agent-reliability

  • Updated Apr 17, 2026
  • TypeScript

Conducted in Python, this analysis examines the internal consistency (using Kuder-Richardson Formula 20 [KR-20]), inter-rater reliability (using Intraclass Correlation Coefficient [ICC] and Cohen's Kappa), item difficulty and discrimination indices, and the relationship between multiple-choice and essay sections for mixed-format achievement tests.

  • Updated Jun 19, 2026
  • Jupyter Notebook

Portable, content-addressed reliability evidence for LLM systems. Capture how a model behaves under perturbation; preserve, verify, and diff the evidence across model changes.

  • Updated Jun 12, 2026
  • Python

Add this topic to your repo

To associate your repository with the reliability-testing topic, visit your repo's landing page and select "manage topics."

Learn more