From 3891eb430442afa55d2da7619c8a0f5463715650 Mon Sep 17 00:00:00 2001 From: deepfates Date: Mon, 28 Sep 2026 16:52:19 -0700 Subject: [PATCH 1/3] Release 0.7.0 @version 0.7.0, CHANGELOG Unreleased becomes 0.7.0, RELEASE_NOTES for 0.7.0, and {:imp, "~> 0.7"} and v0.7.0 links in the README, docs, livebooks and deployment example. The streamed-call model change and the incomplete-stream error are marked Breaking with their migrations. --- CHANGELOG.md | 49 +++--- README.md | 6 +- RELEASE_NOTES.md | 155 +++++++----------- docs/coming-from-dspy.md | 4 +- docs/diving-deeper/modules-and-composition.md | 2 +- docs/diving-deeper/saving-and-artifacts.md | 2 +- .../running-in-your-application.md | 2 +- docs/getting-started/setting-up.md | 4 +- docs/getting-started/where-to-go-next.md | 2 +- docs/production.md | 8 +- examples/deployment/README.md | 2 +- examples/deployment/mix.exs | 2 +- livebooks/01_real_lm_front_door.livemd | 2 +- livebooks/02_without_a_provider.livemd | 2 +- livebooks/03_evaluate_and_optimize.livemd | 2 +- livebooks/04_tools_agents_mcp_rlm.livemd | 2 +- livebooks/05_operating_imp.livemd | 4 +- mix.exs | 2 +- priv/public_api.json | 2 +- test/package_contract_test.exs | 2 +- 20 files changed, 118 insertions(+), 138 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 6116804e..7418559e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,7 +2,7 @@ User-visible changes to Imp are recorded here. -## Unreleased +## 0.7.0 — 2026-09-28 ### Changed @@ -16,13 +16,34 @@ User-visible changes to Imp are recorded here. provider key reports OpenRouter's fee plus the upstream charge, or `nil` when the upstream charge is missing. A call to any catalog-priced provider that reports no charge (Anthropic, OpenAI, Google, Groq, xAI and others) - now has a `nil` cost where it had the estimate. The estimate is the new `estimated_cost` field on - `Imp.Core.LMResponse` and `:estimated_cost` on the `:model_response` event; - `billing` is the breakdown behind that estimate, as it always was, and its - docs now say so. Migration: a host that sums `cost` treats `nil` as an - unknown charge, not a free one; a host that wants the old number for calls - with no reported charge reads `estimated_cost` for them, knowing it is an - estimate. + now has a `nil` cost where it had the estimate. The estimate is the new + `estimated_cost` field on `Imp.Core.LMResponse` and `:estimated_cost` on + the `:model_response` event, `nil` for a streamed call; `billing` is the + breakdown behind that estimate, as it always was, and its docs now say so. + Migration: a host that sums `cost` treats `nil` as an unknown charge, not a + free one; a host that wants the old number for calls with no reported + charge reads `estimated_cost` for them, knowing it is an estimate. +- Breaking: a stream from `Imp.Clients.ReqLLM` that did not complete ends in + `{:error, %Imp.LMError{}}`, not in a completion of whatever text arrived + first: one that carries a provider error, finishes with reason `:error` or + `:cancelled`, or is incomplete, its body ending with no finish and no + `[DONE]` (reason `{:stream_finished, :incomplete}`). The error event + carries the usage and other metadata that arrived before it. Migration: a + caller of `Imp.stream/3` with `provider_stream: true` handles that error as + any failed call; the request reached the provider, so it may have been + charged. +- Breaking: a call streamed through `Imp.Clients.ReqLLM` records the model + the provider reported, as a non-streamed call does, and otherwise the + configured model id (`gpt-test`, where it recorded the whole + `openai:gpt-test`), in the `req_llm` metadata of its `:model_response` + event's `:response` and in its `Imp.Usage` key (`"openai/gpt-test"`). + A call to a client built from a string-keyed spec map or a + `{provider, opts}` or `{provider, model, opts}` tuple, streamed or not, + records its provider, where it recorded none, so its `Imp.Usage` key gains + the `"provider/"` prefix. Migration: a host that matched the + provider-prefixed model of a streamed call matches the model id, or the + model the provider reported, and one that reads `Imp.Usage` by key for such + a client uses the prefixed key. ### Fixed @@ -34,23 +55,13 @@ User-visible changes to Imp are recorded here. with exactly one terminal event, `done: true` with the provider's usage (including its `"cost"`), model and finish reason, taking the finish reason, and usage no chunk reported, from ReqLLM's metadata handle. -- A stream that did not complete ends in `{:error, %Imp.LMError{}}`, not in - a completion of whatever text arrived first: one that carries a provider - error, finishes with reason `:error` or `:cancelled`, or is incomplete, - its body ending with no finish and no `[DONE]` (reason - `{:stream_finished, :incomplete}`). The error event carries the usage and - other metadata that arrived before it. - A streamed call that fails after the provider reported usage records that usage and cost on its failed `:model_response` event and in `Imp.Usage`, since the provider may have charged for it. The caller still receives `{:error, reason}`. - A streamed call to a client built from an inline spec map, with atom or string keys, or a `{provider, opts}` tuple no longer fails with - `{:lm_stream_failed, "protocol String.Chars not implemented ..."}`. A - streamed call records the model the provider reported, as a non-streamed - call does, and otherwise the configured model id (`gpt-test`, where it - recorded the whole `openai:gpt-test`), so its `Imp.Usage` key changes the - same way. + `{:lm_stream_failed, "protocol String.Chars not implemented ..."}`. - `:reasoning_effort` accepts `max`, when an LM is built, on a call and in a saved program. Imp's accepted efforts are read from ReqLLM's own `reasoning_effort` option, so they are every effort ReqLLM accepts, on every diff --git a/README.md b/README.md index e7f6208a..e8c72ea6 100644 --- a/README.md +++ b/README.md @@ -145,7 +145,7 @@ text or JSON you can score, such as an agent's tool descriptions. ## Install ```elixir -{:imp, "~> 0.6"} +{:imp, "~> 0.7"} ``` Imp needs Elixir 1.19 or later and a C and C++ compiler, for the native code @@ -154,7 +154,7 @@ access, because erlexec's build fetches rebar3 plugins. It reaches models throug [ReqLLM](https://hex.pm/packages/req_llm), so any provider ReqLLM supports works. -Imp 0.6 is experimental. Its API may still change, and its optimizers need +Imp 0.7 is experimental. Its API may still change, and its optimizers need large-scale benchmarking. Bug reports and pull requests are welcome. ## Learn @@ -162,7 +162,7 @@ large-scale benchmarking. Bug reports and pull requests are welcome. - [Getting started](docs/getting-started/index.md) builds one program step by step, from the first call to a supervised server, with real scores. - [Coming from DSPy](docs/coming-from-dspy.md) maps DSPy's names to Imp's. -- [Tutorials](https://github.com/deepfates/imp/tree/v0.6.0/livebooks) are Livebook notebooks +- [Tutorials](https://github.com/deepfates/imp/tree/v0.7.0/livebooks) are Livebook notebooks you can run offline or with a key. - The [cheatsheet](docs/cheatsheet.cheatmd) has the common calls on one page. diff --git a/RELEASE_NOTES.md b/RELEASE_NOTES.md index 2d170314..9dae4295 100644 --- a/RELEASE_NOTES.md +++ b/RELEASE_NOTES.md @@ -1,26 +1,24 @@ -# Imp v0.6.0 +# Imp v0.7.0 Imp is a framework for typed, optimizable language-model programs on the BEAM. Declare a task as named inputs and outputs, call it like any other Elixir program, measure it on examples, compile it with an optimizer, and run the selected program under OTP. -This release fixes bugs found after 0.5.0: failures that were reported as -success, credentials that reached reports and checkpoints, requests that could -be sent twice, and checkpoints that could not be resumed. It is `0.6.0` rather -than a patch because several of those fixes change what a caller receives: -`Imp.collect/3` returns a prediction; a ReActV2 turn whose model fails returns -a `StepError`; examples without declared inputs are refused; a reply that -answers no output is a parse error; `ReqLLMBatch` no longer resends a request -that may have run; an MCP 503 or 529 is `:refused`; `Imp.Datasets.csv/3` -refuses some files 0.5.0 loaded; an `Imp.react` signature may not have a -`tools` field; and GEPA and `Imp.Observability` report errors where they -reported `:ok` or `:failed`. +This release makes the money a model call reports mean one thing, fixes +streamed calls through `Imp.Clients.ReqLLM`, and fixes reasoning efforts. It +is `0.7.0` rather than `0.6.1` because three of those fixes change what a +caller receives: a call's `cost` is `nil` where the provider reported no +charge, where it was ReqLLM's catalog estimate; a stream that did not +complete returns `{:error, %Imp.LMError{}}`, where it returned the text that +had arrived; and a streamed call's recorded model is the model id or the +model the provider reported, where it was the whole `"provider:model"` +string. ## Install ```elixir -{:imp, "~> 0.6"} +{:imp, "~> 0.7"} ``` Every dependency comes from Hex. Use a path dependency only while developing @@ -47,91 +45,62 @@ an application has. ## Headline changes -- An agent's step prompt asks for one format, as DSPy's does. With an LM that - calls tools natively, the step no longer also asks for a `tool_calls` field; - an LM that cannot is shown each tool's description and arguments in the - prompt and writes its calls in `tool_calls`. -- Failures are reported instead of passing as success. A ReActV2 turn whose - model fails returns `Imp.Predict.ReActV2.StepError` with the history as far - as it got; a GEPA run that continued past failed proposals reports - `:with_errors`; the Chat, JSON and XML adapters report a reply that answers - no output as a parse error; and evaluation and optimizers refuse examples - that never declared their inputs, whose labels were passed to the program - as inputs. -- Redaction runs before a term is converted, in reports, results, GRPO - checkpoints, session records and saved programs, so a client, retriever or - OAuth store in them is no longer written with its secrets. The SIMBA, - MIPROv2, InferRules, random-search and GEPA checkpoints redact their failure - reasons. Redaction still replaces the whole string, as in 0.5.0, and - recognizes more credential shapes. -- Renderers (`:system_renderer`, `:output_renderer`) shape the JSON fallback - and every JSON or XML request, so a fallback sends the request the host - shaped. -- `Imp.Clients.ReqLLMBatch` never sends again a request that may have run, and - waits for the provider's `retry-after` before a retry. -- A streamed call is recorded in its run like any other, and a reply with text - and tool calls keeps all of them, streamed or not. -- GEPA checkpoints resume: from a pending proposal batch, from a program that - is an agent, and in a fresh VM. +- `cost` on `Imp.Core.LMResponse` and on the `:model_response` event is the + charge the provider reported, and `nil` when it reported none. OpenRouter + reports its charge unasked, and Imp now reads it; a call made through + OpenRouter with the caller's own provider key reports OpenRouter's fee plus + the upstream charge. A provider whose response carries no charge (the + Anthropic, OpenAI and Google APIs called directly among them) gives a `nil` + `cost`. ReqLLM's catalog price is the new `estimated_cost`. +- A call streamed through `Imp.Clients.ReqLLM` (`Imp.stream/3` with + `provider_stream: true`) records the usage and cost the provider reported; + in 0.6.0 every such call was recorded with no usage and no cost. A stream + that stops with a provider error, with finish reason `:error` or + `:cancelled`, or with a body that ends before the provider finished, is an + error, and the usage that arrived before it is still recorded. A streamed + call to a client built from a spec map or a `{provider, opts}` tuple no + longer fails. +- `:reasoning_effort` accepts every effort ReqLLM accepts, `max` among them, + read from ReqLLM's own option. An effort given as a string, as every effort + loaded from a saved program is, reaches ReqLLM as its atom; in 0.6.0 it + failed every call on most providers, OpenRouter (without + `openrouter_reasoning_wire: :nested`), Anthropic, Google and Groq among + them. -## Upgrading from 0.5 +## Upgrading from 0.6 -1. Change the dependency to `{:imp, "~> 0.6"}`, run `mix deps.get`, and - commit `mix.lock`. `nimble_csv` and `mint` are new direct dependencies. -2. `Imp.collect/3` returns `{:ok, prediction}` or `{:error, reason}`; read - fields with `Imp.get(prediction, :answer)` instead of matching a string. -3. Where you checked `Imp.Prediction.complete?/1` after a ReActV2 model - failure, match `{:error, %Imp.Predict.ReActV2.StepError{reason: reason, - history: history}}` and store `history` as you store a finished turn's. A - step refused by an `Imp.OperationalSafetyError` ends the turn this way at - once; where you call the program yourself, match - `{:error, %Imp.Predict.ReActV2.StepError{reason: - %Imp.OperationalSafetyError{}}}`. `Imp.Evaluate` and the optimizers - already raise it. -4. Call `Imp.with_inputs/2` on every example you evaluate or optimize on, and - build evaluation rows with `Imp.example/1 |> Imp.with_inputs(...)` rather - than plain maps or pair lists. Give each field of an example or prediction - once, under one spelling. To keep a signature's instructions, pass it to - `Imp.Signature.new/2` without instructions. -5. A field called `tools` in an `Imp.react` task signature is refused, and a - saved program with one no longer loads. Rebuild the program with the field - renamed and save it; to keep a saved program's optimized instructions and - demos, rename the field in the saved file, or optimize again. -6. Match `%Imp.AdapterParseError{kind: :missing_fields}` where you relied on a - prediction of defaults, or on ProgramOfThought's `:missing_program`, for a - reply that answered no output. -7. Compare `Imp.Clients.ReqLLM` clients by `model`, not by the whole struct, - which now carries `:tool_calling`; `%Imp.Predict.ReActV2{}` likewise - carries `:tool_order`. -8. Treat a `ReqLLMBatch` request that ends `:ambiguous` as possibly run, and - check it with the provider before sending it again. A 0.5.0 checkpoint's - `:transient_failure` requests become `:ambiguous` on resume. -9. An MCP call answered with 503 or 529 is `:refused`, not `:unknown`. -10. Quote CSV fields that hold a quote (`"12"" pipe",1`) and remove spaces - between a comma and a quoted field; `Imp.Datasets.csv/3` refuses both. -11. Match `Imp.Observability.Status`'s new `:succeeded_with_errors`, which an - optimizer report with errors gives where it gave `:failed`. Read GEPA's - `report.errors` rather than expecting `status: :ok`: a run that went on - past failed proposals reports `:with_errors`. Review `:beam_native` - stoppers built on `consecutive_outcome/2`, which now counts an iteration - that raised as `:proposal_error` instead of `:none`. -12. Keep 0.5.0 away from files 0.6.0 writes: it cannot read a trajectory with - atom keys (in a GEPA or Playbook checkpoint, or saved with `Imp.dump/1`), - an optimizer report that holds an `Imp.History`, or a `ReqLLMBatch` - checkpoint. -13. A ReActV2 turn that reaches `max_iters` with text beside tool calls it - did not run now answers with that text, where it answered `nil`; the calls - are still listed as unexecuted. A host that publishes every non-empty - answer should decide whether to publish such a turn's text. -14. Re-evaluate saved agents on held-out data: the step prompt and tool roster - order changed, so they send different prompt text. -15. A `:system_renderer` or `:output_renderer` now also shapes the JSON - fallback and JSON and XML requests. Build on `opts[:default_system]` and - `opts[:default_outputs]` rather than ignoring the options, or the fallback - sends a Chat-shaped prompt. +1. Change the dependency to `{:imp, "~> 0.7"}`, run `mix deps.get`, and + commit `mix.lock`. No dependency was added or removed. +2. Treat a `nil` `cost` as an unknown charge, not a free call. A host that + wants the 0.6.0 number for a call with no reported charge reads + `estimated_cost` (or `:estimated_cost` on the event), knowing it is an + estimate; it is `nil` for a streamed call. +3. Where you call `Imp.stream/3` with `provider_stream: true`, handle + `{:error, %Imp.LMError{}}` from a stream that did not complete as a failed + call. The request reached the provider, so it may have been charged. +4. Where you match the model recorded for a streamed call (the `req_llm` + metadata of its `:model_response` event, or its `Imp.Usage` key), expect + the model the provider reported, or the configured model id without its + provider prefix: `gpt-test` for `"openai:gpt-test"`, and the `Imp.Usage` + key `"openai/gpt-test"`. A client built from a string-keyed spec map or a + tuple now records its provider, streamed or not, so its `Imp.Usage` key + is `"provider/model"` where it was the model alone. +5. Expect spend totals built on `Imp.Usage` or on `:model_response` events to + grow: streamed calls now carry their usage and cost, and a streamed call + that fails after the provider reported usage records it on its failed + `:model_response` event and in `Imp.Usage`. ## Known limits +- The usage map (`:usage` on the `:model_response` event, `Imp.Usage`, + `Imp.Prediction.get_lm_usage/1`) is ReqLLM's, unchanged. For a + non-streamed call its `:cost` and `:total_cost` are ReqLLM's catalog + estimate, not a charge, and OpenRouter's charge is its `"cost"`. Read money + from `cost` and `estimated_cost`. +- An error the provider sends inside a stream has `status` `nil` and + `retryable: true`, because ReqLLM's stream decoder keeps only its message; + the same error in a non-streamed response may carry a status that says not + to retry. - `mint` 1.11.0 leaves an HTTP/1 connection open after a receive timeout, and Finch 0.23.0, the newest release, returns it to its pool with the unanswered request still on it. A later request the pool gives that diff --git a/docs/coming-from-dspy.md b/docs/coming-from-dspy.md index e60dae08..456d6837 100644 --- a/docs/coming-from-dspy.md +++ b/docs/coming-from-dspy.md @@ -12,7 +12,7 @@ brought in. ## The five-minute version ```elixir -# pip install dspy -> {:imp, "~> 0.6"} +# pip install dspy -> {:imp, "~> 0.7"} # lm = dspy.LM("openai/...") -> lm = Imp.req_llm("openai:gpt-5.4-mini", api_key: ...) # dspy.Predict("q -> a") -> program = Imp.predict("q -> a", lm: lm) # program(q="...") -> {:ok, pred} = Imp.call(program, %{q: "..."}) @@ -87,7 +87,7 @@ error row too. Each model request has a receive timeout, and has no time limit of its own (an MCP tool call does, 30 seconds by default), so to bound an agent, start it with `Imp.start_run/3` and cancel it when you choose. The -[deployment example](https://github.com/deepfates/imp/blob/v0.6.0/examples/deployment/README.md) +[deployment example](https://github.com/deepfates/imp/blob/v0.7.0/examples/deployment/README.md) is a complete OTP application. **Optimizers see what you name.** DSPy finds predictors by walking a module's diff --git a/docs/diving-deeper/modules-and-composition.md b/docs/diving-deeper/modules-and-composition.md index f7dfabaa..83c9695a 100644 --- a/docs/diving-deeper/modules-and-composition.md +++ b/docs/diving-deeper/modules-and-composition.md @@ -149,7 +149,7 @@ A custom module is your code, so it is not saved as a whole. What an optimizer chose for it is data: `Imp.ProgramParameters.values/1` reads it, `Imp.ProgramParameters.apply_values/2` puts it on a freshly built program, and `Imp.Optimizer.Artifact` writes it to a checksummed file. The -[deployment example](https://github.com/deepfates/imp/blob/v0.6.0/examples/deployment/README.md) +[deployment example](https://github.com/deepfates/imp/blob/v0.7.0/examples/deployment/README.md) does this in a supervised application. ### Built-in variants diff --git a/docs/diving-deeper/saving-and-artifacts.md b/docs/diving-deeper/saving-and-artifacts.md index fa60ff26..4a865c53 100644 --- a/docs/diving-deeper/saving-and-artifacts.md +++ b/docs/diving-deeper/saving-and-artifacts.md @@ -21,7 +21,7 @@ Save the whole program when it is made of Imp's modules: `Imp.predict/2`, when the program is your own struct implementing `Imp.Module`, or when you want the program's code to live in your release and only its tuned text to cross the persistence boundary. The -[deployment example](https://github.com/deepfates/imp/tree/v0.6.0/examples/deployment) +[deployment example](https://github.com/deepfates/imp/tree/v0.7.0/examples/deployment) does the second. ### 2. JSON with a checksum diff --git a/docs/getting-started/running-in-your-application.md b/docs/getting-started/running-in-your-application.md index 9991337e..44b6c125 100644 --- a/docs/getting-started/running-in-your-application.md +++ b/docs/getting-started/running-in-your-application.md @@ -121,7 +121,7 @@ Changing models is then a configuration change, and rotating a key is a restart. The -[deployment example](https://github.com/deepfates/imp/blob/v0.6.0/examples/deployment/README.md) +[deployment example](https://github.com/deepfates/imp/blob/v0.7.0/examples/deployment/README.md) is the complete version of this page: a two-stage program like `TicketTriage`, learned parameters loaded and reloaded as an `Imp.Optimizer.Artifact`, crashed and timed-out calls contained, and the whole thing restarted in a fresh OS diff --git a/docs/getting-started/setting-up.md b/docs/getting-started/setting-up.md index f4bbe249..71186599 100644 --- a/docs/getting-started/setting-up.md +++ b/docs/getting-started/setting-up.md @@ -5,7 +5,7 @@ Add Imp to a Mix project: ~~~elixir def deps do [ - {:imp, "~> 0.6"} + {:imp, "~> 0.7"} ] end ~~~ @@ -13,7 +13,7 @@ end or, for a script or a Livebook notebook: ~~~elixir -Mix.install([{:imp, "~> 0.6"}]) +Mix.install([{:imp, "~> 0.7"}]) ~~~ Imp needs Elixir 1.19 and a C and C++ compiler, for the native code in two diff --git a/docs/getting-started/where-to-go-next.md b/docs/getting-started/where-to-go-next.md index 4ee7c144..fedef501 100644 --- a/docs/getting-started/where-to-go-next.md +++ b/docs/getting-started/where-to-go-next.md @@ -51,7 +51,7 @@ The module documentation is the reference for every function. [Running Imp in production](../production.md) covers what the last page began: supervision, provider failures, runs you can observe and cancel, and telemetry. The -[deployment example](https://github.com/deepfates/imp/blob/v0.6.0/examples/deployment/README.md) +[deployment example](https://github.com/deepfates/imp/blob/v0.7.0/examples/deployment/README.md) is a complete application to copy from. ## Try it in a notebook diff --git a/docs/production.md b/docs/production.md index 7e52a8db..b69242ee 100644 --- a/docs/production.md +++ b/docs/production.md @@ -5,16 +5,16 @@ where its credentials come from, how long a call may take and what it may cost. This page covers those decisions in the order you meet them: starting Imp, credentials, concurrency, timeouts, cost and caching, persistence and telemetry. The -[deployment example](https://github.com/deepfates/imp/tree/v0.6.0/examples/deployment) +[deployment example](https://github.com/deepfates/imp/tree/v0.7.0/examples/deployment) is a complete OTP application that puts them together. ## Imp in your supervision tree -Imp is an OTP application, and adding `{:imp, "~> 0.6"}` to your +Imp is an OTP application, and adding `{:imp, "~> 0.7"}` to your dependencies starts it with yours. It supervises its settings, its response cache and the task pools that bound its work. It opens no listener and starts no subprocess, and the MCP and ACP runtimes start only when you use -them. In a script, `Mix.install([{:imp, "~> 0.6"}])` starts it too. +them. In a script, `Mix.install([{:imp, "~> 0.7"}])` starts it too. ## Credentials at runtime @@ -252,7 +252,7 @@ standard error. ## The reference application -[`examples/deployment`](https://github.com/deepfates/imp/tree/v0.6.0/examples/deployment) +[`examples/deployment`](https://github.com/deepfates/imp/tree/v0.7.0/examples/deployment) is a small OTP application that loads a checksummed artifact at startup, reads its key from the environment, serves calls from bounded supervised tasks, answers overload and timeouts with errors instead of blocking, and reloads diff --git a/examples/deployment/README.md b/examples/deployment/README.md index 8b2080ae..c7b130f9 100644 --- a/examples/deployment/README.md +++ b/examples/deployment/README.md @@ -150,4 +150,4 @@ Use a whole-program artifact for a supported portable Imp shape. Use its selected predictor parameters should cross the persistence boundary. During source development set `IMP_PATH` to the Imp checkout. Other -applications use the Hex dependency in `mix.exs`, `{:imp, "~> 0.6"}`. +applications use the Hex dependency in `mix.exs`, `{:imp, "~> 0.7"}`. diff --git a/examples/deployment/mix.exs b/examples/deployment/mix.exs index 0fa677c8..47a086b7 100644 --- a/examples/deployment/mix.exs +++ b/examples/deployment/mix.exs @@ -11,7 +11,7 @@ defmodule ImpDeployment.MixProject do defp deps do case System.get_env("IMP_PATH") do - nil -> [{:imp, "~> 0.6"}] + nil -> [{:imp, "~> 0.7"}] path -> [{:imp, path: path}] end end diff --git a/livebooks/01_real_lm_front_door.livemd b/livebooks/01_real_lm_front_door.livemd index 90ffc1b4..56293a97 100644 --- a/livebooks/01_real_lm_front_door.livemd +++ b/livebooks/01_real_lm_front_door.livemd @@ -7,7 +7,7 @@ imp_path = System.get_env("IMP_PATH") lockfile = imp_path && Path.join(imp_path, "mix.lock") cond do - is_nil(imp_path) -> Mix.install([{:imp, "~> 0.6"}]) + is_nil(imp_path) -> Mix.install([{:imp, "~> 0.7"}]) File.regular?(lockfile) -> Mix.install([{:imp, path: imp_path}], lockfile: lockfile) true -> Mix.install([{:imp, path: imp_path}]) end diff --git a/livebooks/02_without_a_provider.livemd b/livebooks/02_without_a_provider.livemd index c8343e8b..d18c9766 100644 --- a/livebooks/02_without_a_provider.livemd +++ b/livebooks/02_without_a_provider.livemd @@ -7,7 +7,7 @@ imp_path = System.get_env("IMP_PATH") lockfile = imp_path && Path.join(imp_path, "mix.lock") cond do - is_nil(imp_path) -> Mix.install([{:imp, "~> 0.6"}]) + is_nil(imp_path) -> Mix.install([{:imp, "~> 0.7"}]) File.regular?(lockfile) -> Mix.install([{:imp, path: imp_path}], lockfile: lockfile) true -> Mix.install([{:imp, path: imp_path}]) end diff --git a/livebooks/03_evaluate_and_optimize.livemd b/livebooks/03_evaluate_and_optimize.livemd index 9d63ec3c..0cf1b308 100644 --- a/livebooks/03_evaluate_and_optimize.livemd +++ b/livebooks/03_evaluate_and_optimize.livemd @@ -7,7 +7,7 @@ imp_path = System.get_env("IMP_PATH") lockfile = imp_path && Path.join(imp_path, "mix.lock") cond do - is_nil(imp_path) -> Mix.install([{:imp, "~> 0.6"}]) + is_nil(imp_path) -> Mix.install([{:imp, "~> 0.7"}]) File.regular?(lockfile) -> Mix.install([{:imp, path: imp_path}], lockfile: lockfile) true -> Mix.install([{:imp, path: imp_path}]) end diff --git a/livebooks/04_tools_agents_mcp_rlm.livemd b/livebooks/04_tools_agents_mcp_rlm.livemd index 581fbb64..3afba394 100644 --- a/livebooks/04_tools_agents_mcp_rlm.livemd +++ b/livebooks/04_tools_agents_mcp_rlm.livemd @@ -7,7 +7,7 @@ imp_path = System.get_env("IMP_PATH") lockfile = imp_path && Path.join(imp_path, "mix.lock") cond do - is_nil(imp_path) -> Mix.install([{:imp, "~> 0.6"}]) + is_nil(imp_path) -> Mix.install([{:imp, "~> 0.7"}]) File.regular?(lockfile) -> Mix.install([{:imp, path: imp_path}], lockfile: lockfile) true -> Mix.install([{:imp, path: imp_path}]) end diff --git a/livebooks/05_operating_imp.livemd b/livebooks/05_operating_imp.livemd index 720406b9..666d4fee 100644 --- a/livebooks/05_operating_imp.livemd +++ b/livebooks/05_operating_imp.livemd @@ -7,7 +7,7 @@ imp_path = System.get_env("IMP_PATH") lockfile = imp_path && Path.join(imp_path, "mix.lock") cond do - is_nil(imp_path) -> Mix.install([{:imp, "~> 0.6"}]) + is_nil(imp_path) -> Mix.install([{:imp, "~> 0.7"}]) File.regular?(lockfile) -> Mix.install([{:imp, path: imp_path}], lockfile: lockfile) true -> Mix.install([{:imp, path: imp_path}]) end @@ -103,7 +103,7 @@ Enum.each(busy, &Task.await/1) refused ``` -The [deployment example](https://github.com/deepfates/imp/tree/v0.6.0/examples/deployment) +The [deployment example](https://github.com/deepfates/imp/tree/v0.7.0/examples/deployment) wraps the same pattern in a `ProgramServer` that also reloads parameters while serving. diff --git a/mix.exs b/mix.exs index fb186dd7..4ab03b71 100644 --- a/mix.exs +++ b/mix.exs @@ -1,7 +1,7 @@ defmodule Imp.MixProject do use Mix.Project - @version "0.6.0" + @version "0.7.0" def project do [ diff --git a/priv/public_api.json b/priv/public_api.json index 9dd3b0ac..35ddbeaa 100644 --- a/priv/public_api.json +++ b/priv/public_api.json @@ -10338,7 +10338,7 @@ "types": [] } ], - "package_version": "0.6.0", + "package_version": "0.7.0", "schema_version": 3, "scope": "Curated Imp API compiled from package-shipped sources. Exports come from non-hidden documentation entries; callbacks, types, and struct fields are category- and policy-gated." } diff --git a/test/package_contract_test.exs b/test/package_contract_test.exs index 2461566c..6c7223c3 100644 --- a/test/package_contract_test.exs +++ b/test/package_contract_test.exs @@ -114,7 +114,7 @@ defmodule PackageContractTest do version = Mix.Project.config()[:version] dependency = hex_dependency(version) - assert version == "0.6.0" + assert version == "0.7.0" assert File.read!("RELEASE_NOTES.md") =~ "# Imp v#{version}" assert File.read!("CHANGELOG.md") =~ "## #{version}" assert File.read!("examples/deployment/mix.exs") =~ dependency From 4af753af6b2b4fc821c68688b6b42230bc6a3b6f Mon Sep 17 00:00:00 2001 From: deepfates Date: Mon, 28 Sep 2026 20:09:38 -0700 Subject: [PATCH 2/3] Release notes: apply the fact-check corrections A streamed call records the configured model id; Imp.Usage does not see streamed calls, so failed-stream spend is on the :model_response event only (new Known limit); OpenRouter's charge was read in 0.6.0 only for uncatalogued models; how ReqLLM's providers take an effort, here and in the Imp.Clients.ReqLLM moduledoc; both tuple shapes; using the filler for history is Imp's; the reloaded note's tool_calls is a plain map; and the CHANGELOG note entry says how to convert an old note. --- CHANGELOG.md | 35 +++++++++++++++++++------------- RELEASE_NOTES.md | 41 ++++++++++++++++++++++---------------- lib/imp/clients/req_llm.ex | 16 ++++++++------- 3 files changed, 54 insertions(+), 38 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index cf806b39..2ff790a8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -32,17 +32,18 @@ User-visible changes to Imp are recorded here. caller of `Imp.stream/3` with `provider_stream: true` handles that error as any failed call; the request reached the provider, so it may have been charged. -- Breaking: a call streamed through `Imp.Clients.ReqLLM` records the model - the provider reported, as a non-streamed call does, and otherwise the - configured model id (`gpt-test`, where it recorded the whole - `openai:gpt-test`), in the `req_llm` metadata of its `:model_response` - event's `:response` and in its `Imp.Usage` key (`"openai/gpt-test"`). +- Breaking: a call streamed through `Imp.Clients.ReqLLM` records the + configured model id without its provider prefix (`gpt-test`, where it + recorded the whole `openai:gpt-test`) in the `req_llm` metadata of its + `:model_response` event's `:response`, and its `Imp.Usage` entry is keyed + `"openai/gpt-test"`. In 0.6.0 a real streamed call had no `Imp.Usage` + entry at all. A call to a client built from a string-keyed spec map or a `{provider, opts}` or `{provider, model, opts}` tuple, streamed or not, records its provider, where it recorded none, so its `Imp.Usage` key gains the `"provider/"` prefix. Migration: a host that matched the - provider-prefixed model of a streamed call matches the model id, or the - model the provider reported, and one that reads `Imp.Usage` by key for such + provider-prefixed model of a streamed call matches the model id, and one + that reads `Imp.Usage` by key for such a client uses the prefixed key. - Breaking: `Imp.Predict.ReActV2`'s `:last_request_note` reaches the model as a user message with no assistant reply after it, as documented, in the @@ -50,7 +51,8 @@ User-visible changes to Imp are recorded here. stored as a history entry with inputs and no outputs. The chat adapter renders that as a finished exchange, so the note was followed by an assistant message the model never gave, each field reading "Not supplied - for this conversation history message." That filler is Imp's; DSPy 3.2.1 + for this conversation history message." Using that filler for history is + Imp's; DSPy 3.2.1 renders a missing history output as `None`. The note, and the entry for inputs no step spent that comes before it when the first step failed, now carry `tool_calls: %Imp.Adapter.Types.ToolCalls{tool_calls: []}` and @@ -58,7 +60,9 @@ User-visible changes to Imp are recorded here. host that recognises the note in `metadata.history` by its shape (only the first input's key) matches it by that key with an empty `tool_calls` list, or by its text. A note recorded by an earlier version keeps the old - shape and still renders with the filler. `History` entries a host writes + shape and still renders with the filler; to render it as this release + does, add `tool_calls: %Imp.Adapter.Types.ToolCalls{tool_calls: []}` and + `tool_call_results: []` to that entry. `History` entries a host writes keep Imp's rendering. - In written tool mode (an LM that cannot call tools natively, where earlier steps replay as text), a stored step that recorded no call and no other @@ -78,17 +82,20 @@ User-visible changes to Imp are recorded here. (including its `"cost"`), model and finish reason, taking the finish reason, and usage no chunk reported, from ReqLLM's metadata handle. - A streamed call that fails after the provider reported usage records that - usage and cost on its failed `:model_response` event and in - `Imp.Usage`, since the provider may have charged for it. The caller still + usage and cost on its failed `:model_response` event, since the provider + may have charged for it. The caller still receives `{:error, reason}`. - A streamed call to a client built from an inline spec map, with atom or - string keys, or a `{provider, opts}` tuple no longer fails with + string keys, or a `{provider, opts}` or `{provider, model, opts}` tuple no + longer fails with `{:lm_stream_failed, "protocol String.Chars not implemented ..."}`. - `:reasoning_effort` accepts `max`, when an LM is built, on a call and in a saved program. Imp's accepted efforts are read from ReqLLM's own `reasoning_effort` option, so they are every effort ReqLLM accepts, on every - provider; ReqLLM's provider maps or clamps it to what that provider's API - takes. OpenRouter receives `"max"` in either wire field. + provider. ReqLLM's provider translates it where it has a translation + (Anthropic and Google turn it into a thinking budget); OpenAI, OpenRouter, + Groq and xAI receive the effort as written, and it is the provider's to + accept. OpenRouter receives `"max"` in either wire field. - A string effort such as `"high"` reaches ReqLLM as its atom. ReqLLM checks the effort against its atom list before any provider sees it, so in 0.6.0 a string effort, including every effort loaded from a saved program, failed diff --git a/RELEASE_NOTES.md b/RELEASE_NOTES.md index 5cf35d09..c9e4065f 100644 --- a/RELEASE_NOTES.md +++ b/RELEASE_NOTES.md @@ -12,8 +12,8 @@ never gave. It is `0.7.0` rather than `0.6.1` because four of those fixes change what a caller receives: a call's `cost` is `nil` where the provider reported no charge, where it was ReqLLM's catalog estimate; a stream that did not complete returns `{:error, %Imp.LMError{}}`, where it returned the text that -had arrived; and a streamed call's recorded model is the model id or the -model the provider reported, where it was the whole `"provider:model"` +had arrived; a streamed call's recorded model is the configured model id +without its provider prefix, where it was the whole `"provider:model"` string; and the history entry of ReActV2's `:last_request_note` carries empty `tool_calls` and `tool_call_results`, where it held the note alone. @@ -49,7 +49,9 @@ an application has. - `cost` on `Imp.Core.LMResponse` and on the `:model_response` event is the charge the provider reported, and `nil` when it reported none. OpenRouter - reports its charge unasked, and Imp now reads it; a call made through + reports its charge unasked, and Imp now reads it for every model; 0.6.0 + read it only for models ReqLLM's catalog does not price and reported the + catalog estimate for the rest. A call made through OpenRouter with the caller's own provider key reports OpenRouter's fee plus the upstream charge. A provider whose response carries no charge (the Anthropic, OpenAI and Google APIs called directly among them) gives a `nil` @@ -59,8 +61,9 @@ an application has. in 0.6.0 every such call was recorded with no usage and no cost. A stream that stops with a provider error, with finish reason `:error` or `:cancelled`, or with a body that ends before the provider finished, is an - error, and the usage that arrived before it is still recorded. A streamed - call to a client built from a spec map or a `{provider, opts}` tuple no + error, and the usage that arrived before it is still recorded on the + failed `:model_response` event. A streamed call to a client built from a + spec map or a `{provider, opts}` or `{provider, model, opts}` tuple no longer fails. - `:reasoning_effort` accepts every effort ReqLLM accepts, `max` among them, read from ReqLLM's own option. An effort given as a string, as every effort @@ -72,7 +75,8 @@ an application has. message with no assistant message after it, in the last request and whenever the returned history is passed back. It was followed by an assistant message the model never gave, each field reading "Not supplied - for this conversation history message." That text is Imp's own; DSPy 3.2.1 + for this conversation history message." Using that text for history is + Imp's; DSPy 3.2.1 renders a missing history output as `None`. The note's history entry, and the entry for inputs no step spent that comes before it when the first step failed, now carry `tool_calls: %Imp.Adapter.Types.ToolCalls{tool_calls: @@ -94,16 +98,15 @@ an application has. `{:error, %Imp.LMError{}}` from a stream that did not complete as a failed call. The request reached the provider, so it may have been charged. 4. Where you match the model recorded for a streamed call (the `req_llm` - metadata of its `:model_response` event, or its `Imp.Usage` key), expect - the model the provider reported, or the configured model id without its - provider prefix: `gpt-test` for `"openai:gpt-test"`, and the `Imp.Usage` - key `"openai/gpt-test"`. A client built from a string-keyed spec map or a - tuple now records its provider, streamed or not, so its `Imp.Usage` key - is `"provider/model"` where it was the model alone. -5. Expect spend totals built on `Imp.Usage` or on `:model_response` events to - grow: streamed calls now carry their usage and cost, and a streamed call - that fails after the provider reported usage records it on its failed - `:model_response` event and in `Imp.Usage`. + metadata of its `:model_response` event), expect the configured model id + without its provider prefix: `gpt-test` for `"openai:gpt-test"`. A client + built from a string-keyed spec map or a tuple now records its provider, so + the `Imp.Usage` key of its calls is `"provider/model"` where it was the + model alone. +5. Expect spend totals built on `:model_response` events to grow: streamed + calls now carry their usage and cost, and a streamed call that fails after + the provider reported usage records it on its failed `:model_response` + event. 6. If you store the histories ReActV2 returns (`metadata.history`) and recognise the last-request note by its exact shape (only the first input's key), accept the new `tool_calls` and `tool_call_results` fields, or match @@ -113,10 +116,14 @@ an application has. `tool_calls: %Imp.Adapter.Types.ToolCalls{tool_calls: []}` and `tool_call_results: []` to that entry; for a history saved with `Imp.History.dump/1`, load it with `Imp.History.load!/1`, add them, and - dump it again. + dump it again. In a history reloaded with `Imp.History.load!/1` the field + is the plain map `%{tool_calls: []}`. ## Known limits +- `Imp.Usage` does not see streamed calls: they run in a separate process, + so usage tracked around `Imp.collect` or a streamed `Imp.call` is empty. + Read usage from the run's `:model_response` events. - The usage map (`:usage` on the `:model_response` event, `Imp.Usage`, `Imp.Prediction.get_lm_usage/1`) is ReqLLM's, unchanged. For a non-streamed call its `:cost` and `:total_cost` are ReqLLM's catalog diff --git a/lib/imp/clients/req_llm.ex b/lib/imp/clients/req_llm.ex index ca575744..baf09666 100644 --- a/lib/imp/clients/req_llm.ex +++ b/lib/imp/clients/req_llm.ex @@ -35,13 +35,15 @@ defmodule Imp.Clients.ReqLLM do event, and the provider stream is cancelled. `:reasoning_effort` is the one reasoning option, on the client or on a call. - It takes any value of ReqLLM's own `reasoning_effort` option, such as `high`, - `xhigh`, `max` or `default`, as an atom or a string, on every provider; - ReqLLM's provider maps or clamps it to what that provider's API takes. A - call naming `nil` spends no reasoning on that call whatever the client is - configured with. Native reasoning fields (`Imp.Predict`) set the same - option, so a client configured with an effort and a program that asks for - one never disagree. + It takes any value of ReqLLM's own `reasoning_effort` option, such as + `high`, `xhigh`, `max` or `default`, as an atom or a string, on every + provider. ReqLLM's provider translates it where it has a translation + (Anthropic and Google turn it into a thinking budget); OpenAI, OpenRouter, + Groq and xAI receive the effort as written, and it is the provider's to + accept. A call naming `nil` spends no reasoning on that call whatever the + client is configured with. Native reasoning fields (`Imp.Predict`) set the + same option, so a client configured with an effort and a program that asks + for one never disagree. OpenRouter accepts the effort in two wire fields, and its endpoint catalog says which one an endpoint supports: ReqLLM's top-level `reasoning_effort` From 71bb26aa41548b0dd0c2b22eb96f02cb3006ab38 Mon Sep 17 00:00:00 2001 From: deepfates Date: Mon, 28 Sep 2026 20:22:01 -0700 Subject: [PATCH 3/3] Dialyzer: re-pin the provider_meta guard after the moduledoc reflow --- .dialyzer_ignore.exs | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/.dialyzer_ignore.exs b/.dialyzer_ignore.exs index a30b4699..ff2d9f74 100644 --- a/.dialyzer_ignore.exs +++ b/.dialyzer_ignore.exs @@ -55,7 +55,7 @@ # defensive guard: ReqLLM.Response types provider_meta as map() with a %{} # default, but the struct does not enforce it (a caller can build one with # nil), and ReqLLM's own OpenTelemetry attributes guard it with is_map/1. - {"lib/imp/clients/req_llm.ex", :guard_fail, 1917}, + {"lib/imp/clients/req_llm.ex", :guard_fail, 1919}, # defensive error clause on an always-ok internal call {"lib/imp/clients/training.ex", :pattern_match, {1215, 13}}, # defensive error clause on an always-ok internal call