Skip to content
@QuesmaOrg

Quesma

Making AI agents production-ready through independent evaluation and training.

Pinned Loading

  1. CompileBench CompileBench Public

    Benchmark of LLMs on real open-source projects against dependency hell, legacy toolchains, and complex build systems.

    Astro 59 8

  2. otel-bench otel-bench Public

    OpenTelemetry Benchmark - can AI trace your failed login?

    Shell 21 4

Repositories

Showing 10 of 25 repositories
  • quesma-shipper Public

    Collects AI coding-agent sessions from developer machines, scrubs secrets, encrypts with age, and uploads to your organisation's storage.

    QuesmaOrg/quesma-shipper's past year of commit activity
    Go 6 Apache-2.0 0 2 3 Updated Sep 11, 2026
  • awesome-ai-tokenomics Public

    A curated list on AI token economics: what tokens cost, where they get wasted, and how to cut the bill. Tools, benchmarks, papers, and copy-paste configs for the token economy of LLMs and coding agents.

    QuesmaOrg/awesome-ai-tokenomics's past year of commit activity
    Python 174 CC0-1.0 21 0 1 Updated Sep 8, 2026
  • shipper-protocol Public

    The wire contract between Quesma Shipper and its control plane.

    QuesmaOrg/shipper-protocol's past year of commit activity
    Go 4 Apache-2.0 0 0 1 Updated Sep 4, 2026
  • terminal-bench-science Public Forked from harbor-framework/terminal-bench-science

    Terminal Bench for Science

    QuesmaOrg/terminal-bench-science's past year of commit activity
    Shell 1 Apache-2.0 330 0 0 Updated Jul 25, 2026
  • otel-bench Public

    OpenTelemetry Benchmark - can AI trace your failed login?

    QuesmaOrg/otel-bench's past year of commit activity
    Shell 21 Apache-2.0 4 0 4 Updated Jul 14, 2026
  • BinaryAudit Public

    An open-source benchmark for evaluating AI agents' ability to find backdoors hidden in compiled binaries.

    QuesmaOrg/BinaryAudit's past year of commit activity
    Shell 99 6 1 1 Updated Jul 14, 2026
  • QuesmaOrg/trival-prompt-bench's past year of commit activity
    HTML 1 0 0 1 Updated Jul 14, 2026
  • CompileBench Public

    Benchmark of LLMs on real open-source projects against dependency hell, legacy toolchains, and complex build systems.

    QuesmaOrg/CompileBench's past year of commit activity
    Astro 59 MIT 8 2 3 Updated Jul 14, 2026
  • gradient-engineer Public

    Gradient Engineer: 60‑Second Linux Analysis (Nix + LLM)

    QuesmaOrg/gradient-engineer's past year of commit activity
    Go 18 MIT 3 0 1 Updated Jul 3, 2026
  • terminal-bench-3 Public Forked from harbor-framework/terminal-bench

    For of TB3.0

    QuesmaOrg/terminal-bench-3's past year of commit activity
    Python 0 491 0 0 Updated Jun 17, 2026