Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,10 @@ All notable changes to `@tangle-network/agent-eval` and its sibling `agent-eval-

## [Unreleased]

---

## [0.150.2] — 2026-08-20

### Added

- `economics.spend` on the supervisor-run report and `spendUsd` on the rollup (#660): both total-spend measurements as named fields with their denominators, never one bare number. `spend.journalDerived` (journal `metered` + `settled` rows) answers execution accounting — what execution observably consumed; `spend.closeRecord` (loops `state.json` `result.spentUsd`, or Runtime `result.json` `spentTotal.usd`, which the analyzer now reads) answers billing — what the store recorded as settled at close. Each run-level measurement carries its record count; each rollup measurement carries `runs`, its own denominator, because the two sums cover different run sets (measured in discovery-lab: 4.762B input tokens journal-derived over 918 runs vs 4.780B close-record over 868 runs, ~0.4% apart). Divergence is a signal to read, not an error. `totalUsd` stays as the collapsed compatibility field; its docstring names the pick order and points to `spend`.
Expand Down
2 changes: 1 addition & 1 deletion clients/python/pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ build-backend = "hatchling.build"

[project]
name = "agent-eval-rpc"
version = "0.150.1"
version = "0.150.2"
description = "Python RPC client, official optimizer bridge, and DSPy metric adapter for @tangle-network/agent-eval."
readme = "README.md"
requires-python = ">=3.10"
Expand Down
2 changes: 1 addition & 1 deletion clients/python/src/agent_eval_rpc/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -53,7 +53,7 @@
try:
__version__ = version("agent-eval-rpc")
except PackageNotFoundError:
__version__ = "0.150.1"
__version__ = "0.150.2"

__all__ = [
"Client",
Expand Down
2 changes: 1 addition & 1 deletion clients/python/uv.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "@tangle-network/agent-eval",
"version": "0.150.1",
"version": "0.150.2",
"description": "Evaluate and improve AI agents from runs, traces, judges, and feedback. Compare candidates, cluster failures, measure lift, and gate releases.",
"homepage": "https://github.com/tangle-network/agent-eval#readme",
"repository": {
Expand Down
2 changes: 1 addition & 1 deletion src/analyst/benchmark-implementation.ts
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@ export const ANALYST_BENCHMARK_DEPENDENCY_LOCK_FILES = Object.freeze([
])

export const ANALYST_BENCHMARK_DEPENDENCY_LOCK_SHA256 =
'5c5e71bc3a4fd59416a0b56ac94808fac4b633de1d402824a07ebcc219a94cf2'
'61ad75e662e240404f523d7dad66e18145df9c35b1ce8630468970fef2a00162'

/** The published benchmark evidence was produced at this package version, by
* the retired one-shot direct runner, before trace analysts moved to the
Expand Down
Loading