Streamed ReqLLM calls record the usage and cost the provider reported - #255
Merged
Merged
Conversation
…and cost
ReqLLM's stream is a Stream.resource, which reports its end after a
suspension as {:halted, _}. The client ended such a stream without the
terminal event, so every streamed call was recorded with no usage and no
cost. Both ends now close with one terminal event; a stream whose metadata
carries a provider error or finishes :error or :cancelled ends as an
Imp.LMError instead of a completion, and a failure event carries the
metadata that arrived before it. The streamed path also accepts an inline
model spec.
# Conflicts: # CHANGELOG.md
Merge ReqLLM's metadata handle into the terminal event, so an incomplete
stream (no finish, no [DONE]) ends as {:stream_finished, :incomplete} and
usage only the handle reported is kept. A failed stream returns its partial
response to Imp.LM.record/3, which records its usage and cost and returns
the reason alone. Imp.Clients.ReqLLM.model_identity/1 names provider and
model id for every accepted model shape, replacing Execution's copy.
# Conflicts: # .dialyzer_ignore.exs
ReqLLM's two-element model tuple is {provider, keyword}, naming the model
in :id or :model; {provider, binary} is not a shape ReqLLM or Imp accepts.
CHANGELOG: #253's breaking cost change moves under Changed.
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
Every provider stream from
Imp.Clients.ReqLLM.stream/3now ends with exactly one terminal event. A stream that did not complete is a failure, and a failed stream records the spend the provider reported.StreamResponse.streamis aStream.resource. Resumed after it has run out, it reports{:halted, _}, not{:done, _}. The client dropped that case with a bare{:halt, state}, so a real stream never emitted thedone: trueevent with usage, model and finish reason. Every real streamed call was recorded with no usage and no cost. The client's reducer only suspends and is only resumed with{:cont, _}, so neither result can mean that a consumer stopped. Both now go throughfinish_stream/1.finish_stream/1awaitsStreamResponse.metadata_handle(bounded at 5 s; an exit or raise adds nothing). It takes the handle's finish reason, and its usage when no chunk reported any. The handle reports:incompletewhen the body ended with no termination event (stream_server.exextract_final_metadata/1).:error,:cancelledor:incomplete, ends in{:error, %Imp.LMError{}}withretryable: true, because the stream had started. It never ends as a completion.ReqLLM.StreamResponse.events/1ends the first three as:error/:cancelledevents. It carries:incompleteas the reason of a:finishevent, which Imp does not count as a completion. Before this change, the partial text of such a stream was collected as the answer.Imp.Streaming.Executionkeeps that metadata and returns{:error, reason, partial}toImp.LM.record/3, wherepartialis anImp.Core.LMResponsebuilt from it.record/3putsusage,cost,estimated_cost(andbilling) on the failed:model_responseand counts the usage inImp.Usage. The caller still receives{:error, reason}. The reason for this change: a host sumscostas money spent, and a stream that failed after it was billed must not count as free.Imp.Clients.ReqLLM.model_identity/1(@doc false) returns the provider (as a string) and the model id for every model shape the client accepts: a string, a{provider, opts}tuple (ReqLLM's own two-element shape, naming the model in:idor:model), a{provider, model, opts}tuple, or a map with atom or string keys. The old{provider, binary}clause is dropped: neither ReqLLM nor Imp accepts that shape. The client'sresponse_metadataandExecution.normalize_metadata/2both use it. Before, a streamed call with an inline map spec or a{provider, opts}spec failed withString.Chars not implemented. A streamed call now records the model the provider reported, as a non-streamed call does, and otherwise the configured model id. The fallback was the whole spec string, so theImp.Usagekey for a stream that reports no model changes fromopenai/openai:gpt-testtoopenai/gpt-test. One test was updated for this.relayed_error/1and the stream path share oneprovider_error/1mapping.Imp.LMErrormoduledoc now says that an error sent inside a stream arrives without its provider code, because ReqLLM's decoder keeps only"message". Such an error therefore has statusniland is retryable.metadata_handle: self()now start a realMetadataHandle(tests andMix.Tasks.Imp.Benchmark.Trace.StreamReqLLM).origin/main(including A model call's cost is the provider's reported charge; the catalog price is estimated_cost #253 and Accept every reasoning effort ReqLLM accepts, including max #254) is merged in; the CHANGELOG keeps both lists under Unreleased, with A model call's cost is the provider's reported charge; the catalog price is estimated_cost #253's breaking cost change under Changed and the fixes under Fixed.How it was tested
The tests are in
test/req_llm_stream_end_test.exs. Every stream there is a realStream.resource: ReqLLM decoding OpenRouter-shaped SSE fromImp.Test.LocalHTTP, or aStream.resourcestub with a real metadata handle.Imp.Run,usage["cost"]and:cost)Imp.Run)Imp.LM.record/3withImp.Usage.track){:openai, id: ...}and a string-keyed map; the identity test also covers{:openai, id: ...}and{:openai, model: ...}, and both failed before the{provider, opts}clause was added)Falsification: I reverted each change, ran the tests, and restored it.
lib/reverted, all four failed, and with only the{:halted, _}clause reverted, all four failed too.:incompletemapping: the incomplete test fails.model_identitywithout the string-keyed and 2-tuple clauses: the identity test and the tuple/string-keyed test fail.to_stringnaming: 3 fail, among them the tuple/string-keyed test and the run-record test.lm.exnot counting the partial: therecord/3usage test fails.mix checkpassed locally:59 doctests, 9 properties, 3555 tests, 0 failures, 13 skipped (221 excluded).mix dialyzer.checkpassed locally:Total errors: 142, Skipped: 142, Unnecessary Skips: 0, with no unused filters. #254 removed theresolve_model/1pattern_match_coventry. This branch moves theprovider_meta || %{}guard_failpin to line 1917; its finding and reason are unchanged.Unsure / open
:model_responsedoes not record the partial text asoutput, only the error, usage and cost.