Skip to content

A model call's cost is the provider's reported charge; the catalog price is estimated_cost - #253

Merged
deepfates merged 3 commits into
mainfrom
claude/imp-reported-cost
Sep 28, 2026
Merged

deepfates merged 3 commits into
mainfrom
claude/imp-reported-cost

Conversation

@deepfates

@deepfates deepfates commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Fixes #252.

What changed

  • Imp.Core.LMResponse.cost (and :cost on the :model_response event) is the charge the provider reported, or nil when it reported none.
  • OpenRouter with the caller's own key ("is_byok": true): OpenRouter's "cost" covers only its fee, and the provider bills the key separately. OpenRouter reports that upstream charge as cost_details.upstream_inference_cost. cost is the fee plus the upstream charge, or nil if the upstream figure is missing. Understating spend is the dangerous direction for a cap, so a missing figure means unknown, not zero.
  • The new estimated_cost field (and :estimated_cost on the event) holds ReqLLM's catalog price. billing stays where it was: the breakdown behind that estimate. Its docs now say so instead of calling it the provider's breakdown.
  • Docs for LMResponse and Imp.LM describe both figures and what nil means. There is a CHANGELOG entry under Unreleased, a one-line row in decisions.md, and a regenerated priv/public_api.json for the new struct field.

Where the numbers come from (read from ReqLLM 1.24.0 and checked live)

  • ReqLLM's usage step (Step.Usage.handle/1 → Usage.Cost.merge/3) prices the tokens from the catalog. It writes that price under atom keys: :cost (the breakdown map) and :total_cost.
  • For OpenAI-format responses, Provider.Defaults.parse_openai_usage/2 keeps every usage field it doesn't interpret under its original string key. OpenRouter's charge therefore arrives as usage["cost"], next to "cost_details" and "is_byok".
  • v0.6.0 read atom :cost first, so the estimate always won whenever the catalog had a price.
  • I made one live request (openrouter:openai/gpt-4.1-nano, max_tokens 5, no openrouter_usage option) on v0.6.0 code. Usage carried "cost" => 2.2e-6 and :total_cost => 2.0e-6, and Imp reported cost: 2.0e-6. This settled two things:
    • OpenRouter reports its charge without being asked, so Imp does not add usage: %{include: true}. A comment in reported_cost/2 says why.
    • Part of the gap is ReqLLM's rounding: Billing.component_cost/2 rounds each line item to 1e-6 USD. For small calls the estimate can be off by tens of percent even when catalog prices are correct.
  • Streaming is not covered by this PR. ReqLLM never prices a streamed call, so its estimated_cost is nil, and the docs say so. The streamed response's usage metadata is lost before it reaches Imp.Core.response/1, so it has no cost either. That bug is fixed in a separate PR.
  • Other providers: every catalog-priced provider that reports no charge now gets cost: nil plus an estimate. That includes Anthropic, OpenAI and Google called directly, Groq, xAI and others.

Every use of cost in lib, and what I decided

  • lib/imp/core.ex: response/1 and reported_cost/2 changed as described above. reported_cost/2 still falls back to top-level metadata :cost, which is how a client other than ReqLLM reports a charge (a test covers it).
  • lib/imp/lm.ex: the :model_response event gains :estimated_cost. :billing is unchanged.
  • lib/imp/trajectory.ex: ATIF model_observation copies the event metadata whole, so it carries estimated_cost with no code change (a test covers it).
  • lib/imp/optimizer/budget.ex:488 and lib/mix/tasks/imp.benchmark.gepa_campaign.ex:536: these read ReqLLM's [:req_llm, :token_usage] telemetry (total_cost/cost), which is ReqLLM's catalog estimate. They don't read LMResponse, so this change doesn't touch them. Optimizer budgets therefore still count estimates. I left that alone: it is a prospective limit with its own pricing, and changing it is a separate decision.
  • lib/imp/optimizer/trajectory.ex:624 / trajectory_types.ex: a GEPA trajectory's own usage struct. It is not built from LMResponse.cost, so no change.
  • lib/imp/optimizer/playbook*.ex (cost_usd) and lib/mix/tasks/imp.benchmark.* ("cost", "estimated_cost", token_cost): these are their own accounting from configured prices or benchmark results. They don't read LMResponse, so no change.
  • max_reflection_cost (GEPA) reads ReflectionStrategy.observable_cost, which is unrelated. No change.

Tests

test/model_response_cost_test.exs has 10 tests. Main ones:

  • "an OpenRouter response keeps OpenRouter's charge as cost and ReqLLM's price as the estimate". This goes through the real path: an OpenRouter-shaped body from a local HTTP server, ReqLLM decoding, and ReqLLM's usage step pricing the tokens from inline pricing components. cost == 0.0025 (reported) and estimated_cost ≈ 0.0019 (catalog).
  • "a priced call reports the provider's charge as cost and the catalog price apart". Same check at the event level.
  • "a call ReqLLM priced but the provider did not report has no cost, only an estimate". cost: nil, estimated_cost: 0.001858.
  • "an OpenRouter call on the caller's own key costs the fee plus the upstream charge" and "... with no upstream charge has no cost" (no cost_details, an empty map, or a nil upstream figure).
  • Also covered: nothing priced (both nil, no :billing), ATIF carries both, string/Decimal parsing for both fields, a non-ReqLLM client's metadata :cost, and unreadable figures read as nil.

Falsification: I reverted lib/ to v0.6.0 (git diff lib > patch; git checkout lib) and ran the file: 10 tests, 8 failures. Every test named above failed. On the OpenRouter test, v0.6.0 gave the estimate as cost. Then I restored the patch. For the BYOK clause, I removed only that clause: its two tests failed (12 tests, 2 failures), and then I restored it.

Checks run: mix format --check-formatted, mix compile --warnings-as-errors, mix docs --warnings-as-errors, mix docs.check, mix quality.check, mix dialyzer.check, and mix check. The final full run, after the review fixes: 59 doctests, 9 properties, 3543 tests, 0 failures, 13 skipped (221 excluded).

Two earlier full runs failed, for reasons outside this change:

  • PublicAPIManifestTest failed because the manifest needed regenerating for the new field. That is fixed in this PR.
  • Imp.TimedOutConnectionTest "the call after a timed-out call is answered on a new connection" failed once on a loaded machine. Both calls had a 200 ms timeout, but only the first needs to time out. The separate commit "TimedOutConnectionTest: only the first call has to time out" gives the second call 5 s. refute first == second still catches a second call that lands on the stuck connection. That follows from how the test is written; I did not rerun it on mint 1.11 to prove it.

Version: 0.7.0, I think

A struct field is additive, but the value of cost changes for existing callers. Calls to Anthropic, OpenAI or Google directly went from a number to nil, and OpenRouter calls change from the estimate to the charge. decisions.md says a release that "changes what a caller receives or can rely on bumps the minor version". So the CHANGELOG entry is marked "Breaking:" with a migration note. The case for 0.6.1 is that the docs always promised the reported charge, so this only makes the code match its contract. But a host summing cost for direct-provider calls would quietly stop counting, and that is the kind of change the ruling exists to flag.

Unsure / for the owner

  • Dwell: Dwell.Spend already counts nil as an unpriced call, and its docs say a cap never binds for a provider that reports nothing. OpenRouter residents get real charges after upgrading. A resident on a direct provider would stop counting toward its cap unless Dwell chooses to fall back on estimated_cost.
  • billing keeps its name even though it now clearly means the estimate's breakdown. Renaming it would break callers for a cosmetic gain, so I didn't.

…ice is estimated_cost

LMResponse.cost read ReqLLM's usage :cost first, which is ReqLLM's catalog
estimate, so OpenRouter's own charge (the string-keyed "cost" in its usage
object) was overridden. cost is now the reported charge or nil, and the
estimate is the new estimated_cost field. Fixes #252.
…or a streamed call

With the caller's own key, OpenRouter's "cost" is its fee only; the
upstream charge is cost_details.upstream_inference_cost. cost is their sum,
or nil when the upstream figure is missing.
Both calls had a 200 ms timeout, so the second failed on a loaded machine.
It now has 5 s; a second call that lands on the stuck connection still
fails the refute on the connection it came in on.
@deepfates
deepfates merged commit 4696e50 into main Sep 28, 2026
10 checks passed
@deepfates
deepfates deleted the claude/imp-reported-cost branch September 28, 2026 22:19
deepfates added a commit that referenced this pull request Sep 28, 2026
ReqLLM's two-element model tuple is {provider, keyword}, naming the model
in :id or :model; {provider, binary} is not a shape ReqLLM or Imp accepts.
CHANGELOG: #253's breaking cost change moves under Changed.
@deepfates deepfates mentioned this pull request Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

LMResponse.cost is ReqLLM's catalog estimate, not the provider's reported charge

1 participant