A model call's cost is the provider's reported charge; the catalog price is estimated_cost - #253
Merged
Merged
Conversation
…ice is estimated_cost LMResponse.cost read ReqLLM's usage :cost first, which is ReqLLM's catalog estimate, so OpenRouter's own charge (the string-keyed "cost" in its usage object) was overridden. cost is now the reported charge or nil, and the estimate is the new estimated_cost field. Fixes #252.
…or a streamed call With the caller's own key, OpenRouter's "cost" is its fee only; the upstream charge is cost_details.upstream_inference_cost. cost is their sum, or nil when the upstream figure is missing.
Both calls had a 200 ms timeout, so the second failed on a loaded machine. It now has 5 s; a second call that lands on the stuck connection still fails the refute on the connection it came in on.
deepfates
added a commit
that referenced
this pull request
Sep 28, 2026
ReqLLM's two-element model tuple is {provider, keyword}, naming the model
in :id or :model; {provider, binary} is not a shape ReqLLM or Imp accepts.
CHANGELOG: #253's breaking cost change moves under Changed.
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #252.
What changed
Imp.Core.LMResponse.cost(and:coston the:model_responseevent) is the charge the provider reported, ornilwhen it reported none."is_byok": true): OpenRouter's"cost"covers only its fee, and the provider bills the key separately. OpenRouter reports that upstream charge ascost_details.upstream_inference_cost.costis the fee plus the upstream charge, ornilif the upstream figure is missing. Understating spend is the dangerous direction for a cap, so a missing figure means unknown, not zero.estimated_costfield (and:estimated_coston the event) holds ReqLLM's catalog price.billingstays where it was: the breakdown behind that estimate. Its docs now say so instead of calling it the provider's breakdown.LMResponseandImp.LMdescribe both figures and whatnilmeans. There is a CHANGELOG entry under Unreleased, a one-line row indecisions.md, and a regeneratedpriv/public_api.jsonfor the new struct field.Where the numbers come from (read from ReqLLM 1.24.0 and checked live)
Step.Usage.handle/1→Usage.Cost.merge/3) prices the tokens from the catalog. It writes that price under atom keys::cost(the breakdown map) and:total_cost.Provider.Defaults.parse_openai_usage/2keeps every usage field it doesn't interpret under its original string key. OpenRouter's charge therefore arrives asusage["cost"], next to"cost_details"and"is_byok".:costfirst, so the estimate always won whenever the catalog had a price.openrouter:openai/gpt-4.1-nano, max_tokens 5, noopenrouter_usageoption) on v0.6.0 code. Usage carried"cost" => 2.2e-6and:total_cost => 2.0e-6, and Imp reportedcost: 2.0e-6. This settled two things:usage: %{include: true}. A comment inreported_cost/2says why.Billing.component_cost/2rounds each line item to 1e-6 USD. For small calls the estimate can be off by tens of percent even when catalog prices are correct.estimated_costisnil, and the docs say so. The streamed response's usage metadata is lost before it reachesImp.Core.response/1, so it has no cost either. That bug is fixed in a separate PR.cost: nilplus an estimate. That includes Anthropic, OpenAI and Google called directly, Groq, xAI and others.Every use of cost in lib, and what I decided
lib/imp/core.ex:response/1andreported_cost/2changed as described above.reported_cost/2still falls back to top-level metadata:cost, which is how a client other than ReqLLM reports a charge (a test covers it).lib/imp/lm.ex: the:model_responseevent gains:estimated_cost.:billingis unchanged.lib/imp/trajectory.ex: ATIFmodel_observationcopies the event metadata whole, so it carriesestimated_costwith no code change (a test covers it).lib/imp/optimizer/budget.ex:488andlib/mix/tasks/imp.benchmark.gepa_campaign.ex:536: these read ReqLLM's[:req_llm, :token_usage]telemetry (total_cost/cost), which is ReqLLM's catalog estimate. They don't readLMResponse, so this change doesn't touch them. Optimizer budgets therefore still count estimates. I left that alone: it is a prospective limit with its own pricing, and changing it is a separate decision.lib/imp/optimizer/trajectory.ex:624/trajectory_types.ex: a GEPA trajectory's own usage struct. It is not built fromLMResponse.cost, so no change.lib/imp/optimizer/playbook*.ex(cost_usd) andlib/mix/tasks/imp.benchmark.*("cost","estimated_cost",token_cost): these are their own accounting from configured prices or benchmark results. They don't readLMResponse, so no change.max_reflection_cost(GEPA) readsReflectionStrategy.observable_cost, which is unrelated. No change.Tests
test/model_response_cost_test.exshas 10 tests. Main ones:cost == 0.0025(reported) andestimated_cost ≈ 0.0019(catalog).cost: nil,estimated_cost: 0.001858.cost_details, an empty map, or a nil upstream figure).nil, no:billing), ATIF carries both, string/Decimal parsing for both fields, a non-ReqLLM client's metadata:cost, and unreadable figures read asnil.Falsification: I reverted
lib/to v0.6.0 (git diff lib > patch; git checkout lib) and ran the file: 10 tests, 8 failures. Every test named above failed. On the OpenRouter test, v0.6.0 gave the estimate ascost. Then I restored the patch. For the BYOK clause, I removed only that clause: its two tests failed (12 tests, 2 failures), and then I restored it.Checks run:
mix format --check-formatted,mix compile --warnings-as-errors,mix docs --warnings-as-errors,mix docs.check,mix quality.check,mix dialyzer.check, andmix check. The final full run, after the review fixes:59 doctests, 9 properties, 3543 tests, 0 failures, 13 skipped (221 excluded).Two earlier full runs failed, for reasons outside this change:
PublicAPIManifestTestfailed because the manifest needed regenerating for the new field. That is fixed in this PR.Imp.TimedOutConnectionTest"the call after a timed-out call is answered on a new connection" failed once on a loaded machine. Both calls had a 200 ms timeout, but only the first needs to time out. The separate commit "TimedOutConnectionTest: only the first call has to time out" gives the second call 5 s.refute first == secondstill catches a second call that lands on the stuck connection. That follows from how the test is written; I did not rerun it on mint 1.11 to prove it.Version: 0.7.0, I think
A struct field is additive, but the value of
costchanges for existing callers. Calls to Anthropic, OpenAI or Google directly went from a number tonil, and OpenRouter calls change from the estimate to the charge.decisions.mdsays a release that "changes what a caller receives or can rely on bumps the minor version". So the CHANGELOG entry is marked "Breaking:" with a migration note. The case for 0.6.1 is that the docs always promised the reported charge, so this only makes the code match its contract. But a host summingcostfor direct-provider calls would quietly stop counting, and that is the kind of change the ruling exists to flag.Unsure / for the owner
Dwell.Spendalready countsnilas an unpriced call, and its docs say a cap never binds for a provider that reports nothing. OpenRouter residents get real charges after upgrading. A resident on a direct provider would stop counting toward its cap unless Dwell chooses to fall back onestimated_cost.billingkeeps its name even though it now clearly means the estimate's breakdown. Renaming it would break callers for a cosmetic gain, so I didn't.