From 3d7350edd961391815f9081958ef5c102910c375 Mon Sep 17 00:00:00 2001 From: deepfates Date: Mon, 28 Sep 2026 06:38:03 -0700 Subject: [PATCH 1/8] Release 0.6.0 (draft: version pending owner) Turn Unreleased into 0.6.0: merge the changelog's sections, put every breaking change under Changed with its migration, and replace the release notes with the 0.6.0 document. Bump the version, the install lines, the tag-pinned links and the API manifest's package version. --- CHANGELOG.md | 706 +++++++++--------- README.md | 9 +- RELEASE_NOTES.md | 318 +++----- docs/coming-from-dspy.md | 4 +- docs/diving-deeper/modules-and-composition.md | 2 +- docs/diving-deeper/saving-and-artifacts.md | 2 +- .../running-in-your-application.md | 2 +- docs/getting-started/setting-up.md | 4 +- docs/getting-started/where-to-go-next.md | 2 +- docs/production.md | 8 +- examples/deployment/README.md | 2 +- examples/deployment/mix.exs | 2 +- livebooks/01_real_lm_front_door.livemd | 2 +- livebooks/02_without_a_provider.livemd | 2 +- livebooks/03_evaluate_and_optimize.livemd | 2 +- livebooks/04_tools_agents_mcp_rlm.livemd | 2 +- livebooks/05_operating_imp.livemd | 4 +- mix.exs | 2 +- priv/public_api.json | 2 +- test/package_contract_test.exs | 2 +- 20 files changed, 470 insertions(+), 609 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 442e9cdd..5d971b00 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,7 +2,7 @@ User-visible changes to Imp are recorded here. -## Unreleased +## 0.6.0 — 2026-09-28 ### Security @@ -19,64 +19,228 @@ User-visible changes to Imp are recorded here. or Azure `sig`. It passed these unchanged into run events, traces, trajectories and saved programs. As before, the whole string is replaced, and every shape it caught before is still caught. -- `Imp.ExternalCommand` redacts captured output with the same rules: output - that holds a credential is `"[REDACTED]"`, so the credentials printed beside - a recognized one (`env | grep AWS`, a credentials file) go with it. It used - two patterns of its own (`sk-` and `Bearer`) before the shared rules. The - optimizer's pricing URL check uses the same rules instead of its own. - A `Bearer` token of 16 or more characters with a digit in it is found wherever it ends: mid-line (`Bearer https://…`) or before a newline - in a multi-line string (`"Authorization: Bearer \nmachine …"`). - Both were missed. A word without a digit after `Bearer`, followed by more - text, is still read as prose and left alone. -- A string of near misses for `Bearer` or `Basic` in one letter case - (`BEARER x` lines) is redacted in linear time; it took quadratic time. + (`"Authorization: Bearer \nmachine …"`). Both were missed. A word + without a digit after `Bearer`, followed by more text, is still read as + prose and left alone. +- Near misses for `Bearer` or `Basic` in one letter case (`BEARER x` lines) + are redacted in linear time; they took quadratic time. +- `Imp.ExternalCommand` redacts captured output with the same rules, in + place of its own `sk-` and `Bearer` patterns: output that holds a + credential is `"[REDACTED]"` whole, so the credentials printed beside a + recognized one (`env | grep AWS`, a credentials file) go with it. The + optimizer's pricing URL check uses the same rules instead of its own. +- Writers redact a term before converting it: optimizer reports + (`Imp.Optimizer.Report.dump/1`, `json_safe/1`, `json_projection/1`), + experiment results, `Imp.Evaluate.Result.save_as_json/2` and + `save_as_csv/2`, BetterTogether's bootstrap diagnostics, optimizer artifact + candidates, GRPO session checkpoints, Optimize Anything results, ACP session + records and saved programs (a Predict's demos and metadata, KNN examples, + memory retriever documents). They converted first, so a client, retriever + or MCP OAuth store in them was written with its header values, URL query + secrets or store key, and a secret in a map key that is a tuple, a list or + a struct was written as it was. `Imp.Redaction.drop_credentials/1` redacts + such structs too. The SIMBA, MIPROv2, InferRules and random-search + checkpoints redact their failure reasons; the instructions, demos and scores + a resumed run continues from are kept as they are. Output that holds no + secret is unchanged. +- `Imp.Redaction.redact/2` hides a connection struct's header values and the + query, fragment and user info of its URLs, as `inspect/1` already did: + `Imp.Clients.ReqLLM`, `Imp.Retrievers.HTTP`, `Imp.Tracking.MLflow`, + `Imp.Tracking.WandB`, `Imp.Optimize.Anything.Config.Tracking`, + `ExMCP.Client` and `ExMCP.Transport.HTTP`. A retriever in a tool result kept + an `X-Subscription-Token` header, a cookie and a `?key=` URL in + `render_inspection/2`, run-event JSON and ATIF. +- `Imp.Redaction.redact/2` redacts the secrets in `Imp.MCP.OAuth.Store` (the + derived key), `Imp.MCP.OAuth.Pending` (the authorization URL and `state`) + and `Imp.MCP.OAuth.Flow` (the authorization URL, PKCE transaction and + registered client), keeping each struct and its other fields. They were + walked as plain maps, so the store's key reached any redacted output that + held a store. `code_verifier` is a credential name, so a PKCE transaction + map is redacted on its own too. +- A map key that is a string shaped like a credential is redacted. + `Imp.Optimizer.Trajectory` names a key it refuses (one that is not an atom + or a string) by its type; it printed the key, credential included. +- The `ReqLLM.Error.API.Request` inside an `Imp.LMError` from + `Imp.Clients.ReqLLM` carries at most one response header, `retry-after`. + Any other header ReqLLM left on the error is removed, so cookies, account + identifiers and request ids no longer reach logs, checkpoints or run events + through `inspect/1` of the error. +- An ACP session store redacts the history and transcript it writes, except + the provider's `reasoning_content` and `reasoning_details`. A resumed + session therefore replays a user message or tool result that held a + credential as `"[REDACTED]"`, both to the client and in the model's + history. The session's `_meta` is stored as sent; Imp never reads it. +- A two-element list with a name first is a key and its value only when it is + an element of a list, as JSON writes config, headers and tool results + (`[["api_key", key], ["model", "gpt"]]`); held directly as a field or map + value, it is data. Saving and `Imp.Redaction.drop_credentials/1` read every + such list as a pair, so a saved program whose example had an input named + like a credential (`with_inputs([:api_key, :question])`), or whose metadata + held such a list, did not load back as it was; `Imp.Redaction.redact/1` + rewrote an Avatar tool schema's `"required" => ["api_key", "query"]`; and + `Imp.Optimizer.Parameter` refused input keys `["api_key", "question"]`. One + divergence from 0.5.0 follows: a flat `["api_key", key]` held directly as a + value, under a key that is not itself a credential name, is not redacted, + since it has the shape of a list of two names. ### Changed - Breaking: `Imp.Observability.Status` has a new state, - `:succeeded_with_errors`, and `Imp.Observability.status/1` gives it for - every optimizer report whose `errors` is not empty, from any optimizer, - where it gave `:failed`: an optimizer that returned a report returned a - program. Code that matches every `Status` state adds the new one; code that - treated `:failed` as "the report has errors" matches - `:succeeded_with_errors`. -- Breaking: `Imp.collect/3` now returns `{:ok, prediction}` or - `{:error, reason}`, as `Imp.call/2` does, instead of a string. The string - joined the values of every output field with no separator, so - `question -> reasoning, answer` collected as `"Because.Paris"`. Code that - matched a string reads the field from the prediction instead: + `:succeeded_with_errors`, and `Imp.Observability.status/1` gives it, where + it gave `:failed`, for every optimizer report whose `errors` is not empty: + an optimizer that returned a report returned a program. Migration: code that + matches every `Status` state adds the new one; code that treated `:failed` + as "the report has errors" matches `:succeeded_with_errors`. +- Breaking: `Imp.collect/3` returns `{:ok, prediction}` or + `{:error, reason}`, as `Imp.call/2` does, instead of a string that joined + every output field's value with no separator (`question -> reasoning, + answer` collected as `"Because.Paris"`). Migration: `{:ok, prediction} = Imp.collect(program, inputs)`, then `Imp.get(prediction, :answer)`. -- Breaking: an `Imp.Predict.ReActV2` turn that could not get a model response - returns `{:error, %Imp.Predict.ReActV2.StepError{}}`, where it returned +- Breaking: an `Imp.Predict.ReActV2` turn whose last request (made after a + step failed or the turn was interrupted) also fails returns + `{:error, %Imp.Predict.ReActV2.StepError{}}`, where it returned `{:ok, prediction}` with no outputs and `termination_reason: :incomplete`. - That is a turn whose last request, after a step failed or the turn was - interrupted, failed too: the LM returned an `Imp.LMError`, its client raised - (`{:lm_failed, client, exception}`), or a renderer raised - (`{:adapter_format_failed, adapter, exception}`). The error's `:reason` is - that request's error unchanged, which `Imp.Errors.retryable?/1` and - `Imp.Errors.context_window_exceeded?/1` read through the struct, and its - `:history` is the turn's `Imp.History` as far as it got, the value the - incomplete prediction carried in its `:history` metadata, with every tool - call that ran and its result. A step that fails and whose last request is - answered still ends `:last_text`, `:forced_submit` or `:extracted`, and - `:incomplete` is kept for a turn whose last request was answered without a - valid answer, or that ran out of time or context window. `Imp.ACP` saves the - history of such a turn to its session before failing the turn. Migration: - match `{:error, %Imp.Predict.ReActV2.StepError{reason: reason, history: - history}}` where you checked `Imp.Prediction.complete?/1` after a model - failure, and store `history` as you store a finished turn's. - `Imp.Errors.retryable?/1` on it says whether the last model request may be - sent again, not the turn: its tools have run. Inspecting it shows the reason - and the history's size, not the history. + `:reason` is that request's error unchanged: an `Imp.LMError`, + `{:lm_failed, client, exception}` or + `{:adapter_format_failed, adapter, exception}`, which + `Imp.Errors.retryable?/1` and `context_window_exceeded?/1` read through the + struct. `retryable?/1` says whether the last model request may be sent + again, not the turn: its tools have run. `:history` is the turn's + `Imp.History` as far as it got, with every tool call that ran and its + result; inspecting the error shows the history's size, not the history. + A turn whose last request is answered still ends `:last_text`, + `:forced_submit` or `:extracted`, and `:incomplete` is kept for one answered + without a valid answer or out of time or context window. `Imp.ACP` saves + the history to its session before failing the turn. Migration: where you + checked `Imp.Prediction.complete?/1` after a model failure, match + `{:error, %Imp.Predict.ReActV2.StepError{reason: reason, history: + history}}` and store `history` as you store a finished turn's. - Breaking: an `Imp.Predict.ReActV2` step refused by an `Imp.OperationalSafetyError` (a route, cost, transport or budget guard) ends - the turn at once with that error as the `StepError`'s `:reason`. It made the - turn's last request instead, and an answer to that request ended the turn - `:last_text` or `:forced_submit`, passing the guard by. Migration: none for - a caller that already treats safety errors as fatal; `Imp.Evaluate` and the - optimizers find the guard inside the `StepError` and raise it. + the turn at once with a `StepError` whose `:reason` is that error. The turn + made its last request instead, and an answer to it passed the guard by. + `Imp.Evaluate` and the optimizers find the guard inside the `StepError` and + raise it. Migration: none for a caller that treats safety errors as fatal. +- Breaking: an `Imp.react` task signature cannot have a field named `tools`, + which is now the step's tool list, and loading a program saved with one + raises the same `ArgumentError`. Migration: rename the field and save the + program again. +- Breaking: `Imp.Clients.ReqLLM` has a `:tool_calling` field: whether the + model calls tools natively, read from the registry once when the client is + built or loaded, and kept with the model it was read for. A client whose + model is swapped looks the new model up. A client built by `Imp.req_llm/2`, + or loaded, therefore no longer equals a bare `%Imp.Clients.ReqLLM{}` for the + same model. Migration: compare clients by `model`, not by the whole struct. +- Breaking: in `Imp.Adapter.Chat`, `Imp.Adapter.JSON` and `Imp.Adapter.XML`, + a reply that answers none of the requested outputs (no + `[[ ## field ## ]]` section, no output key, no output tag) is an + `Imp.AdapterParseError` of kind `:missing_fields` naming every output, so + the JSON fallback or a retry runs. When every output was optional or + defaulted, as a ReActV2 step's are, such a reply parsed as defaults and + `nil`s, so an XML agent given prose finished with `answer: nil`. DSPy 3.3.1 + fills defaults there; Imp does not. For a signature that names an output in + `metadata[:text_field]`, Chat and XML read prose as that field (so a + ReActV2 step's prose under XML is its `next_thought`), and a blank + completion, the one exemption, as a step that said nothing. A reply that + writes fields in some format (a JSON object with an output's key, a + `[[ ## field ## ]]` line, an output's tag) is not prose: Chat and XML parse + it or report it, so a step that spelled out a tool call as JSON runs that + tool instead of answering with the JSON text. A JSON `{}` for a ReActV2 + step is `:missing_fields`, where it was a step with `nil` fields. + `Imp.Predict.ProgramOfThought`, whose outputs are all optional, sends prose + to the JSON fallback instead of regenerating with a missing-program error. + Migration: match `%Imp.AdapterParseError{kind: :missing_fields}` where you + relied on a prediction of defaults, or on ProgramOfThought's + `:missing_program`, for a reply that answered nothing. +- Breaking: `Imp.Clients.ReqLLMBatch` never sends a request again when it may + already have run. A dispatch that timed out, crashed, threw or exited was + retried as transient, so one request could run and be billed several times; + it is now `:ambiguous` and final, and a dispatcher may return + `{:ambiguous, reason}` itself. `req_llm_dispatcher/2` makes every call with + `max_retries: 0` (ReqLLM's own retry step could send a request four times in + one batch attempt); retries only a request that never reached the provider + (connection refused, or no pooled connection free, as Req reports it) or + that the provider answered with 408, 425, 429, 503 or 529; treats any other + 4xx as terminal; and treats any other 5xx, and a timeout or closed + connection with no response, as `:ambiguous`, where a 5xx was retried. On + resume, a checkpoint written by 0.5.0 is rewritten at schema version 2 and + its `:transient_failure` requests become `:ambiguous`, since 0.5.0 recorded + timeouts and dispatcher crashes that way. Migration: check an `:ambiguous` + request with the provider before sending it again. +- Breaking: an MCP call the server answers with HTTP 503 or 529 is + `:refused`, like a 429, where it was `:unknown`: RFC 9110 defines 503 as the + server being unable to handle the request, and providers answer overload + with 503 or 529. MCP and language-model calls read a status the same way. + Migration: code that treated a 503 or 529 as possibly run matches + `:refused` for them. +- Breaking: `Imp.Datasets.csv/3` parses RFC 4180 CSV with NimbleCSV, a new + dependency (`nimble_csv ~> 1.3`); it raised `FunctionClauseError` on any + quoted field and could not read a quoted line break. A line break may be + CRLF, LF or a bare CR, and a leading byte order mark is dropped, where it was + read into the first column's name. A blank line is skipped, but a `""` line + is a row with one empty value, so a file with more than one column refuses + it (`invalid CSV row at :: expected N fields, got 1`). A + malformed file raises `Imp.Datasets.Error` naming the line its record starts + on, with that record's first 200 characters as `record`. Some files 0.5.0 + loaded are refused: + - a quote inside an unquoted field, such as an inch mark (`12" pipe,1`): + `invalid CSV at :: a quote opened on this line is never + closed`, or, when the line holds two (`12" by 3" board,1`), + `invalid CSV at :: unexpected escape character " in "..."`. + - a space between a comma and a quoted field (`x, "y"`): + `invalid CSV at :: unexpected escape character " in "..."`. + Migration: quote the whole field and double the quotes inside it + (`"12"" pipe",1`), or remove the space. +- Breaking: `Imp.Example.inputs/1` and `labels/1` raise `ArgumentError` when + the example never declared its inputs, as DSPy raises `ValueError`. They + returned every field as inputs, labels included, so a program was given the + answer and scored on it. `Imp.evaluate/4`, `Imp.Evaluate.run/2`, + `Imp.Experiment.Data.new/1` and every optimizer that runs a program on + examples check each dataset before any model call, and the error names the + function, the dataset and the row. Migration: call `Imp.with_inputs/2` on + every example you evaluate or optimize on. +- Breaking: `Imp.Evaluate` and `Imp.Experiment.Data` refuse a row that is a + plain map or a field pair list, which cannot declare its inputs; Evaluate + turned it into an example whose labels reached the program. Migration: + build the row with `Imp.example/1 |> Imp.with_inputs(...)`. +- Breaking: `Imp.Example.new/1` and `Imp.Prediction.new/2` raise when a field + is given twice, as an atom and a string or as a repeated key; one value was + silently dropped. Migration: give each field once, under one spelling. +- Breaking: `Imp.Signature.new/2` (and `Imp.signature/2`) applies new + instructions to an existing signature; it returned the signature unchanged. + Migration: to keep a signature's instructions, pass it without + instructions. +- Breaking: a GEPA run that continued past failed proposals reports them, so + it reports `status: :with_errors` where it reported `:ok` with no errors. + Each failure is an entry in `report.errors` with its `iteration`, its + `diagnostics` and its `candidate_id` (the rejected candidate in + `candidates`, or `nil` when the failure left none there); + `metadata.failed_proposals` counts them. With `raise_on_exception: false` + that is any failure: a reflection call, reflection strategy, evaluation or + validation that raised, threw or exited, or an iteration that did. With the + default, `true`, it is the failures GEPA already recorded and went on from: + a reflection that returned no usable instruction, a failed reflective + dataset in a parallel slot, and a reflection interrupted before a resume. + The program returned is still the best candidate found, the baseline when + every proposal failed, as DSPy's GEPA returns it. A slot cancelled because + a sibling failed first is rejected, not counted as failed. Migration: code + that checked `status == :ok` reads `report.errors`. +- Breaking: under GEPA's `execution_profile: :beam_native` with + `raise_on_exception: false`, an iteration that raised is recorded as a + rejection with a `{:proposal_error, reason}` reason, so + `Stopper.consecutive_outcome/2` counts it as `:proposal_error`, as the + default profile already did, where it counted `:none`. When evaluating the + first two proposals raises and the third solves the task, + `consecutive_outcome(:proposal_error, 2)` now stops after two iterations + with the baseline, where the run went on and returned the solving + candidate, and `consecutive_outcome(:none, 2)` no longer stops it after the + first. An iteration that threw or exited, which crashed the run, is recorded + the same way. + Migration: review stoppers built on `consecutive_outcome(:none, _)` or + `(:proposal_error, _)` for `:beam_native` runs. - `Imp.Clients.ReqLLM` marks `context_window_exceeded` on more providers' length refusals, streamed or not: Anthropic's "prompt is too long: N tokens > M maximum" and "input length and `max_tokens` exceed context limit", @@ -89,76 +253,21 @@ User-visible changes to Imp are recorded here. non-streamed `context_length_exceeded` was marked, so the others failed the call instead of letting `Imp.Predict.ReActV2` leave out older episodes or end the turn `:incomplete`. +- Imp 0.5.0 cannot read some files this release writes: + - a trajectory with a map that has an atom key, which GEPA checkpoints, + Playbook checkpoints and `Imp.dump/1` write through + `Imp.Optimizer.Trajectory`; + - an optimizer report (`Imp.Optimizer.Report.encode_term/1`, `dump/1`) + that holds an `Imp.History`, now written under a `"history"` tag where it + was a plain map; + - an `Imp.Clients.ReqLLMBatch` checkpoint, now at schema version 2. ### Fixed -- Optimizer reports (`Imp.Optimizer.Report.dump/1`, `json_safe/1`, - `json_projection/1`), experiment results, `Imp.Evaluate.Result.save_as_json/2` - and `save_as_csv/2`, BetterTogether's bootstrap diagnostics, optimizer - artifact candidates, GRPO session checkpoints, Optimize Anything results, - ACP session records and saved programs (a Predict's demos and metadata, KNN - examples, memory retriever documents) redact a term before converting it. - They converted first, so a client, retriever or MCP OAuth store in them was - written with its header values, URL query secrets or store key, and a - secret in a map key that is a tuple, a list or a struct was written as it - was. The SIMBA, MIPROv2, InferRules and random-search checkpoints redact - their failure reasons the same way; the instructions, demos and scores a - resumed run continues from are kept. `Imp.Redaction.drop_credentials/1` - redacts such structs too. Output that holds no secret is unchanged. -- An ACP session store redacts the history and transcript it writes, except - the provider's `reasoning_content` and `reasoning_details`. A resumed session - therefore replays a user message or tool result that held a credential as - `"[REDACTED]"`, both to the client and in the model's history. The session's - `_meta` is stored as sent; Imp never reads it. -- A saved program whose example has an input named like a credential - (`with_inputs([:api_key, :question])`), or whose metadata holds such a list, - loads back as it was. Saving and `Imp.Redaction.drop_credentials/1` read - every two-element list whose first item was a name as a key and its value, - and replaced or dropped the second item. A two-element list with a name - first is now a pair when it is an element of a list, as JSON writes config, - headers and tool results (`[["api_key", key], ["model", "gpt"]]`); held - directly as a field or map value, it is data. This also stops - `Imp.Redaction.redact/1` from rewriting an Avatar tool schema's - `"required" => ["api_key", "query"]` and `Imp.Optimizer.Parameter` from - refusing input keys `["api_key", "question"]`. One divergence from 0.5.0 - follows: a flat `["api_key", key]` held directly as a value, under a key - that is not itself a credential name, is not redacted, since it has the - shape of a list of two names. -- A map key that is a string shaped like a credential is redacted. - `Imp.Optimizer.Trajectory` names a key it refuses (one that is not an atom - or a string) by its type; it printed the key, credential included. -- `Imp.Optimizer.Report.load!/1` and `Imp.Optimizer.GRPO.Checkpoint.load!/1` - read the `"[REDACTED]"` atom marker in a VM that has not loaded - `Imp.Redaction`, and a MIPROv2 checkpoint resumes in a fresh VM: its loader - loads the modules whose atoms it decodes. It failed with "not an already - existing atom". -- A reply that answers none of the requested outputs is a parse error in - `Imp.Adapter.Chat`, `Imp.Adapter.JSON` and `Imp.Adapter.XML`: an - `Imp.AdapterParseError` of kind `:missing_fields` naming every output, so - the JSON fallback or a retry runs. For Chat that is a completion with no - `[[ ## field ## ]]` section for any output, for JSON an object with none of - the output keys, and for XML a reply with none of the output tags (prose, a - JSON object, the other adapters' markers). Such a reply parsed as a - prediction of defaults and `nil`s when every output was optional or - defaulted, as a ReActV2 step's are, so an XML agent given prose finished - with `answer: nil`. DSPy 3.3.1 fills defaults there; Imp does not. - For a signature that names an output in `metadata[:text_field]`, Chat and - XML read prose as that field and a blank completion as a step that said - nothing; the blank completion is the only reply exempt from the rule. So a - ReActV2 step's prose under XML is its `next_thought`. A reply that writes - the fields in some format (a JSON object with an output's key, a - `[[ ## field ## ]]` line, an output's tag) is not prose: Chat and XML parse - it in their own format or report it, and the JSON fallback reads it, so a - step that spelled out a tool call as JSON runs that tool instead of ending - with the JSON text as its answer. A JSON `{}` for a ReActV2 step, under - `Imp.Adapter.JSON` or in the fallback, is now `:missing_fields` where it was - a step with `nil` fields. `Imp.Predict.ProgramOfThought`, whose outputs are - all optional, now sends a prose reply to the JSON fallback instead of - regenerating with a missing-program error. - `Imp.inspect_history/2` renders any history. A turn holding a term JSON has no encoding for, such as the `{:error, {:unknown_tool, name}}` result a ReActV2 history keeps for a call to a tool that does not exist, raised - `Protocol.UndefinedError`. Turns now render through the same conversion + `Protocol.UndefinedError`. Turns render through the conversion `Imp.Observability.render_inspection/2` uses: tuples become lists, structs become maps, and pids, references and functions become their `inspect/1` text. Redaction runs first, as before. @@ -171,109 +280,62 @@ User-visible changes to Imp are recorded here. `Imp.Trajectory.to_atif/2`, which returned the raw bytes, so encoding their result raised. Output for values without such binaries is unchanged, and a request's `tools_hash` is identical. -- `Imp.Redaction.redact/2` redacts the secrets in `Imp.MCP.OAuth.Store` (the - derived key), `Imp.MCP.OAuth.Pending` (the authorization URL and `state`) - and `Imp.MCP.OAuth.Flow` (the authorization URL, PKCE transaction and - registered client), keeping each struct and its other fields. They were - walked as plain maps, whose field names are not credential names, so the - store's key reached any redacted output that held a store. `code_verifier` - is a credential name, so a PKCE transaction map is redacted on its own too. -- `Imp.Redaction.redact/2` hides a connection struct's header values and the - query, fragment and user info of its URLs, as `inspect/1` already did: - `Imp.Clients.ReqLLM`, `Imp.Retrievers.HTTP`, `Imp.Tracking.MLflow`, - `Imp.Tracking.WandB`, `Imp.Optimize.Anything.Config.Tracking`, - `ExMCP.Client` and `ExMCP.Transport.HTTP`. A retriever in a tool result kept - an `X-Subscription-Token` header, a cookie and a `?key=` URL in - `render_inspection/2`, run-event JSON and ATIF. - A call streamed with `Imp.stream(program, inputs, provider_stream: true)` is recorded in its run like any other model call: a `:model_request` event with the request's `:purpose`, and a `:model_response` event with the usage and cost the provider reported. It recorded neither, so a streamed turn left no model record, no cost and no ATIF model step. -- A model turn that says something and calls tools keeps what it said. Through +- A model turn that says something and calls tools keeps both. Through `Imp.req_llm/2` the text was dropped whenever the reply had tool calls, so a ReActV2 step's `next_thought` was empty; streamed with `provider_stream: true`, the text was kept but the tool calls were lost, all of them when text - arrived and all but the last otherwise. Both now return the text and every - tool call, and the adapter reads the text as it reads a text reply, into - `next_thought` for ReActV2, as DSPy does. So when a ReActV2 run reaches - `max_iters` and its last reply has text beside tool calls it did not run, - that text is now the answer, where the answer was `nil`; the calls are still - listed as unexecuted. The text also appears in history turns and in the ATIF - model step. -- An Avatar tool ends with its caller. Its task kept running after the - caller was killed, after `Imp.Run.cancel/3` and after the run's owner died; - it now ends when the caller does. It still runs unlinked, so a crash is an - observation, and takes no place in the task pool. It sees the caller's - `Imp.context/2` settings, run context and deadline, which it did not, and - parallel work it starts runs on the caller's place in the pool when the - caller has one. A tool that timed out, or whose task exited, reads as - `:unknown` in `Imp.Tool.outcome/1` rather than `:result`, since it may have - acted. - Inside a run, Avatar records each tool call as `:tool_call` and - `:tool_result` events with `metadata.outcome`, and arguments that fail the - tool's schema are refused before the tool starts. -- An RLM call made outside a run no longer leaves its model or tool call - running, holding a place in the task pool, when the calling process is - killed. An RLM call no longer leaves the pool place of one of its own tasks - recorded in the calling process, which made parallel work that process - started afterwards (`Imp.Predict.Parallel.map/3`, `Imp.Evaluate.run/2`) run - one item at a time. -- The package ships the TRL worker that `Imp.Clients.TRLTrainer` starts by - default: `priv/trl_worker/worker.py`, its `pyproject.toml` and `uv.lock`, and - the default contract. In 0.5.0 the defaults named files the package did not - contain, so GRPO training from Hex needed a source checkout. The default - contract now pins transformers 5.10.1, the version the lockfile installs; it - named 5.5.0, which the worker refused at startup. -- An `Imp.react` step asks for one thing, whichever adapter formats it. With - an LM that calls tools natively, the tools are sent natively and the step's - prompt no longer also describes `tool_calls` as a field to write, in the - Chat, JSON and XML adapters, their demos and the JSON fallback: the step - output that native calls fill is left out before any adapter formats, as - DSPy does. A model that followed that description wrote a JSON object, and - for a signature with one text output the object became the answer. An LM - whose client says it cannot call tools (a ReqLLM model the registry lists - without tool calling, `Imp.Clients.TRLLM`) is sent no tools and asked to - write its calls in `tool_calls`, whose description shows the shape of a - call with an example. A `tools` input lists every tool, `submit` included, - in DSPy's words: its name, its description, and its arguments as JSON (the - schema's properties, its `required` list and its `$defs`; the properties in - the schema's order when it keeps one, a `Jason.OrderedObject`, and by name - for an ordinary map). Its earlier steps are replayed as text rather than as tool blocks - a provider without declared tools may refuse. A task signature can no - longer have a field named `tools`, and loading a program saved with one is - refused with the same error. `Imp.Clients.ReqLLM` reads whether the model - calls tools from the registry once, when the client is built, and keeps it - with the model it was read for (`:tool_calling`); a client whose model is - swapped looks the new model up. So a client built by `Imp.req_llm/2`, or - loaded, no longer equals a bare `%Imp.Clients.ReqLLM{}` for the same model. The guidance says where a text answer - goes, the same in every format: in `next_thought`, with `tool_calls` left - empty when the model writes its calls. It no longer asks for plain text - beside a structure that asks for fields; a plain-text reply to a Chat step - is still read as the answer. A stored turn that carries only - the task's answer is replayed as that answer, not as step fields marked - "Not supplied"; it is replayed by the adapter itself, so a host - `:output_renderer` is not consulted for it. `Imp.LM.Budgeted` and BootstrapFewShot's rollout LM answer - for the LM they wrap, for this and for reasoning and response-format - support. -- `Imp.react` sends its tool roster in the order the tools were declared, then - `submit`, and a saved agent keeps that order. It was the order of the tool - names' atoms, which can differ between processes and changed the prompt a - provider caches. + arrived and all but the last otherwise. The adapter reads the text as it + reads a text reply, into `next_thought` for ReActV2, as DSPy does. So when a + ReActV2 run reaches `max_iters` and its last reply has text beside tool + calls it did not run, that text is now the answer, where the answer was + `nil`; the calls are still listed as unexecuted. The text also appears in + history turns and in the ATIF model step. +- An `Imp.react` step asks for one thing, whichever adapter formats it, as + DSPy's does. With an LM that calls tools natively, the tools are sent + natively and the prompt (Chat, JSON or XML, their demos and the JSON + fallback) no longer also describes `tool_calls` as a field to write; a model + that followed that description wrote a JSON object, which for a signature + with one text output became the answer. An LM whose client says it cannot + call tools (a ReqLLM model the registry lists without tool calling, + `Imp.Clients.TRLLM`) is sent no tools. Its step has a `tools` input listing + every tool, `submit` included, in DSPy's words: name, description, and + arguments as JSON (the schema's properties, in its order when it is a + `Jason.OrderedObject` and by name otherwise, its `required` list and its + `$defs`). It writes its calls in `tool_calls`, whose description shows a + call's shape with an example, and its earlier steps are replayed as text, + not as tool blocks a provider without declared tools may refuse. The + guidance says the same in every format: a text answer goes in + `next_thought`, with `tool_calls` left empty when the model writes its + calls. It no longer asks for plain text beside a structure that asks for + fields; a plain-text reply to a Chat step is still read as the answer. A + stored turn that carries only the task's answer is replayed as that answer, + not as step fields marked "Not supplied", by the adapter itself, so a host + `:output_renderer` is not consulted for it. `Imp.LM.Budgeted` and + BootstrapFewShot's rollout LM answer for the LM they wrap, for tool calling + as for reasoning and response-format support. +- `Imp.react` sends its tool roster in declared order, then `submit`, and a + saved agent keeps that order. It was the order of the tool names' atoms, + which can differ between processes and changed the prompt a provider + caches. - The JSON fallback after an unparseable Chat or XML reply sends the same request in JSON: the program's `adapter_opts` renderers (`:system_renderer`, - `:output_renderer`), an agent loop's guidance and the demos go with it. - Before, it sent the stock JSON prompt, so a host that shapes its prompt with - renderers got a different prompt on every fallback, and an `Imp.react` step - lost its tool guidance. `Imp.Adapter.JSON` and `Imp.Adapter.XML` now honor + `:output_renderer`), an agent loop's guidance and the demos go with it. It + sent the stock JSON prompt, so a host that shapes its prompt with renderers + got a different prompt on every fallback, and an `Imp.react` step lost its + tool guidance. `Imp.Adapter.JSON` and `Imp.Adapter.XML` honor `:system_renderer`, `:output_renderer` and `:guidance` whenever they are used, and end a request whose inputs are all in the history with the output - requirements as a user message of their own instead of appending them to the - last tool result. A renderer receives the formatting adapter's own + requirements as a user message of their own instead of appending them to + the last tool result. A renderer receives the formatting adapter's own rendering in its options, as `:default_system` and, for an `:output_renderer` that takes the options as a fourth argument, `:default_outputs`, so one that builds on the default builds on the format - of the request and a fallback never asks for two formats. + of the request. - The `[:imp, :adapter, :parse, :json_fallback]` event names the adapter whose reply failed as `:adapter` (it always said `Imp.Adapter.Chat`, also for XML) and the adapter that retried as `:fallback_adapter`. @@ -287,162 +349,79 @@ User-visible changes to Imp are recorded here. an Elixir string cannot hold, makes the parse an `Imp.AdapterParseError`. This reaches the Chat, JSON and XML adapters and GEPA's instruction proposal. -- `Imp.Clients.ReqLLMBatch` no longer sends a request again when it may - already have run. A dispatch that timed out, crashed, threw or exited was - retried as transient, so one request could run and be billed several - times; it is now `:ambiguous` and final, as an uncommitted dispatch already - was on resume. A dispatcher may return `{:ambiguous, reason}` itself. - `req_llm_dispatcher/2`: - - makes every call with `max_retries: 0`. ReqLLM's own retry step resent a - timed-out or refused request up to three more times inside one batch - attempt, so a single attempt could send a request four times. - - retries only a request that never reached the provider (connection - refused, or no pooled connection free, as Req reports it) or that the - provider answered with 408, 425, 429, 503 or 529. - - treats any other 4xx as terminal. - - treats a 500, 502, 504 or any other 5xx but 503 and 529, and a timeout - or closed connection with no response, as `:ambiguous`. A 5xx was - retried. - Before a retry the batch waits: for the provider's `retry-after` (seconds - or an HTTP date) when it sent one, otherwise with exponential backoff and - jitter, capped by the new `:max_retry_wait` option (default 60 s). A +- Before a retry, `Imp.Clients.ReqLLMBatch` waits for the provider's + `retry-after` (seconds, or an IMF-fixdate, RFC 850 or asctime date) when it + sent one, otherwise with exponential backoff and jitter capped by the new + `:max_retry_wait` option (default 60 s). `Imp.Clients.ReqLLM` keeps + `retry-after` on its errors for this; ReqLLM's own decoding dropped it. A waiting request takes no dispatch slot from the others. When the provider asks for longer than `:max_retry_wait`, or the wait would pass the `Imp.Deadline` in force, the request is not retried in that run: it stays `:transient_failure`, the summary is not `complete?`, and `resume/3` retries it. The checkpoint keeps the time each such request may be sent - again (`not_before`, UTC), and `resume/3` waits for it under the same - rules or stops the retry again without sending. `retry-after` may be an - IMF-fixdate, an RFC 850 date or an asctime date. A process that traps - exits and is stopped by its parent during the wait exits at once with the - parent's reason; a dispatch wave in progress is still waited for, up to - `:timeout`, and a process started with bare `spawn` has no parent to - listen for and sleeps uninterrupted. -- The `ReqLLM.Error.API.Request` inside an `Imp.LMError` from - `Imp.Clients.ReqLLM` carries at most one response header, `retry-after`, - so the caller can wait before retrying (ReqLLM's own decoding dropped it). - Any other header ReqLLM left on the error is removed, so cookies, account - identifiers and request ids no longer reach logs, checkpoints or run - events through `inspect/1` of the error. - A checkpoint written by 0.5.0 is rewritten at schema version 2 on resume, - and its `:transient_failure` requests become `:ambiguous`, since 0.5.0 - recorded timeouts and dispatcher crashes that way. -- An MCP call the server answers with HTTP 503 or 529 is `:refused`, like a - 429, where it was `:unknown`: RFC 9110 defines 503 as the server being - unable to handle the request, and providers answer overload with 503 or - 529. MCP and language-model calls read a status the - same way. -- `Imp.Datasets.csv/3` reads quoted fields. It raised `FunctionClauseError` - on any quoted field and could not read a quoted line break. It now parses - RFC 4180 CSV with NimbleCSV, a new dependency (`nimble_csv ~> 1.3`). A - line break may be CRLF, LF or a bare CR (as Excel for Mac writes), and a - leading byte order mark is dropped; it was read into the first column's - name. A blank line is skipped; a record holding only `""` is a row with - an empty value, so in a file with more than one column a `""` line is now - refused as `invalid CSV row at :: expected N fields, got 1`, - where 0.5.0 raised `FunctionClauseError` on it as on any quoted field. A malformed file raises `Imp.Datasets.Error` naming the line its - record starts on, with the start of that record (at most 200 characters) - as `record`. Some files 0.5.0 loaded are now refused: - - a quote inside an unquoted field, such as an inch mark (`12" pipe,1`): - `invalid CSV at :: a quote opened on this line is never - closed`, or, when the line holds two (`12" by 3" board,1`), - `invalid CSV at :: unexpected escape character " in "..."`. - - a space between a comma and a quoted field (`x, "y"`): - `invalid CSV at :: unexpected escape character " in "..."`. - Quote the whole field and double the quotes inside it (`"12"" pipe",1`), - or remove the space. + again (`not_before`, UTC), and `resume/3` waits for it under the same rules + or stops the retry again without sending. A process that traps exits and is + stopped by its parent during the wait exits at once with the parent's + reason; a dispatch wave in progress is still waited for, up to `:timeout`, + and a process started with bare `spawn` has no parent to listen for and + sleeps uninterrupted. +- An Avatar tool ends with its caller. Its task kept running after the caller + was killed, after `Imp.Run.cancel/3` and after the run's owner died. It + still runs unlinked, so a crash is an observation, and takes no place in the + task pool. It sees the caller's `Imp.context/2` settings, run context and + deadline, which it did not, and parallel work it starts runs on the + caller's place in the pool when the caller has one. A tool that timed out, + or whose task exited, reads as `:unknown` in `Imp.Tool.outcome/1` rather + than `:result`, since it may have acted. Inside a run, Avatar records each + tool call as `:tool_call` and `:tool_result` events with + `metadata.outcome`, and arguments that fail the tool's schema are refused + before the tool starts. +- An RLM call made outside a run no longer leaves its model or tool call + running, holding a place in the task pool, when the calling process is + killed. An RLM call no longer leaves the pool place of one of its own tasks + recorded in the calling process, which made parallel work that process + started afterwards (`Imp.Predict.Parallel.map/3`, `Imp.Evaluate.run/2`) run + one item at a time. +- The package ships the TRL worker that `Imp.Clients.TRLTrainer` starts by + default: `priv/trl_worker/worker.py`, its `pyproject.toml` and `uv.lock`, and + the default contract. In 0.5.0 the defaults named files the package did not + contain, so GRPO training from Hex needed a source checkout. The default + contract pins transformers 5.10.1, the version the lockfile installs; it + named 5.5.0, which the worker refused at startup. +- `Imp.Optimizer.Report.load!/1` and `Imp.Optimizer.GRPO.Checkpoint.load!/1` + read the `"[REDACTED]"` atom marker in a VM that has not loaded + `Imp.Redaction`, and a MIPROv2 checkpoint resumes in a fresh VM: its loader + loads the modules whose atoms it decodes. They failed with "not an already + existing atom". - `Imp.Evaluate.Result.save_as_csv/2` writes with the same CSV module that `Imp.Datasets.csv/3` reads with, so what it writes reads back unchanged. The bytes it writes are the same as before. - An instruction an optimizer sets on `Imp.Predict.ProgramOfThought` or `Imp.Predict.CodeAct` reaches the extraction step, which kept the old - instructions when GEPA, MIPROv2, COPRO, SIMBA or InferRules set it. + instructions when GEPA, MIPROv2, COPRO, SIMBA or InferRules set it, and such + a program saves and loads: `Imp.Saving.load!/1` raised "saved + ProgramOfThought planner instructions must match task instructions". `Imp.Optimizer.InstructionSearch` sets instructions the same way. -- A ProgramOfThought or CodeAct whose instruction an optimizer set saves and - loads. `Imp.Saving.load!/1` raised "saved ProgramOfThought planner - instructions must match task instructions" for it. - -### Examples and datasets - -Every change here is breaking for code that relied on the old behaviour. - -- `Imp.Example.inputs/1` and `labels/1` raise `ArgumentError` when the example - never declared its inputs, as DSPy raises `ValueError`. They returned every - field as inputs, labels included, so a program was given the answer and - scored on it. `Imp.evaluate/4`, `Imp.Evaluate.run/2`, - `Imp.Experiment.Data.new/1` and every optimizer that runs a program on - examples check each dataset before any model call, and the error names the - function, the dataset and the row. Migration: call `Imp.with_inputs/2` on - every example you evaluate or optimize on. -- `Imp.Evaluate` and `Imp.Experiment.Data` refuse a row that is a plain map or - a field pair list, which cannot declare its inputs; Evaluate turned it into - an example whose labels reached the program. Migration: build the row with - `Imp.example/1 |> Imp.with_inputs(...)`. -- `Imp.Example.new/1` and `Imp.Prediction.new/2` raise when a field is given - twice, as an atom and a string or as a repeated key; one value was silently - dropped. Migration: give each field once, under one spelling. -- `Imp.Signature.new/2` (and `Imp.signature/2`) applies new instructions to an - existing signature; it returned the signature unchanged. Migration: to keep - a signature's instructions, pass it without instructions. - -### Documentation - -- The install instructions ask for a C and a C++ compiler: jaxon builds native - code from C and erlexec from C++. They asked for a C++ compiler only. They - also say the first compile needs network access, for erlexec's rebar3 - plugins. -- The package's Changelog link opens the changelog on HexDocs, and the 0.4.0 - entry links the benchmark pages as they were at v0.4.0, not on `main`. -- Each Livebook's setup cell only installs Imp: from Hex, or from the - checkout `IMP_PATH` names. It no longer searches for a source checkout. -- The MCP example on the Tools and agents page defines its server and its - imported tools, where it used an undefined `imported`, and closes the - server when the call fails. +- `examples/deployment/agent_optimization.exs` writes the optimizer's + rejections and history with `Imp.Optimizer.Report.json_safe/1`. It passed + them to `Jason.encode!/1`, which raised on a failure reason such as + `{:incomplete_evaluation, 1}`, so a finished run lost its record. ### GEPA - A GEPA report names real failures only: a row whose program call or metric failed, a row whose metric returned a value `Imp.Metrics` cannot read, and - a proposal error. Before, it treated the metric's feedback as a failure, so - a run whose metric returned feedback reported `errors`, - `status: :with_errors` and candidates named "Program call failed: …" when - nothing had failed. -- A GEPA report no longer crashes when a metric throws or exits: the - failure is shown as `{:throw, reason}` or `{:exit, reason}`. Before, - building the report raised `Protocol.UndefinedError`. Every diagnostic in - the report has its credential values redacted. -- A GEPA candidate rejected because its proposal failed is named "Proposal - failed: …", not "Program call failed: …". -- A GEPA run that continued past failed proposals reports them. Each - failure is an entry in `report.errors` with its `iteration`, its - `diagnostics` and its `candidate_id` (the rejected candidate in - `candidates`, or `nil` when the failure left none there); - `metadata.failed_proposals` counts them, and `status` is `:with_errors`. - With `raise_on_exception: false` that is any failure: a reflection call, - reflection strategy, evaluation or validation that raised, threw or - exited, or an iteration that did. With the default, `true`, it is the - failures GEPA already recorded and went on from: a reflection that returned - no usable instruction, a failed reflective dataset in a parallel slot, and - a reflection interrupted before a resume, so such a run now reports - `:with_errors` where it reported `:ok`. The program returned is still the - best candidate found, the baseline when every proposal failed, as DSPy's - GEPA returns it. A slot cancelled because a sibling failed first is - rejected, not counted as failed. Before, such a run reported `status: :ok` - and no errors. -- Under `execution_profile: :beam_native` with `raise_on_exception: false`, - an iteration that raised is recorded as a rejection with a - `{:proposal_error, reason}` reason, so `Stopper.consecutive_outcome/2` - counts it as a `:proposal_error` outcome, as the default profile already - did; before, its outcome was `:none`. For example, when evaluating the - first two proposals raises and the third proposal solves the task, - `consecutive_outcome(:proposal_error, 2)` now stops the run after those two - iterations with the baseline, where before the run went on to - `:max_iterations` and returned the solving candidate; and - `consecutive_outcome(:none, 2)`, which stopped that run after its first - iteration, no longer does. An iteration that threw or exited is recorded the - same way; before, it crashed the run. + a proposal error. It treated the metric's feedback as a failure, so a run + whose metric returned feedback reported `errors`, `status: :with_errors` + and candidates named "Program call failed: …" when nothing had failed. A + candidate rejected because its proposal failed is named "Proposal failed: + …". +- A GEPA report no longer crashes when a metric throws or exits: the failure + is shown as `{:throw, reason}` or `{:exit, reason}`, where building the + report raised `Protocol.UndefinedError`. Every diagnostic in the report has + its credential values redacted. - GEPA ends the run on an `Imp.OperationalSafetyError` (a budget, cost, - route or transport guard) whatever `raise_on_exception` says. Before, with + route or transport guard) whatever `raise_on_exception` says. With `raise_on_exception: false`, it recorded the refusal as a failed proposal and went on spending. - GEPA redacts a failure reason when it records it in a rejection, the @@ -459,38 +438,33 @@ Every change here is breaking for code that relied on the old behaviour. `:kill`) as a failed proposal when a module selector, a reflection strategy, or the adapter's evaluation or reflective dataset raises it in that process: the exit goes on up, as it must from a process that traps - exits and turns its owner's shutdown into an exit. Before, it was recorded - and the run went on. + exits and turns its owner's shutdown into an exit. - `Imp.Optimizer.Trajectory.dump/1` and `load!/1` round-trip a trajectory: a map with an atom key is written as its entries, each key tagged as an atom, so a prediction's metadata, a trace step (`%{predictor: :main}`) and metric metadata such as Optimize Anything's `objective_scores` load with the keys they had. Every map key was written as a string, so a trajectory whose prediction had metadata failed to load (`Imp.Prediction.new/2: invalid map - in :metadata … got: "trace"`). GEPA checkpoints, Playbook checkpoints and - `Imp.dump/1` use this codec; resuming GEPA from a checkpoint whose pending - proposal batch held a result from an `Imp.predict` program raised that - error. Loading never creates an atom: a key written as an atom the loading - VM does not have (a dynamic metric-metadata key, say) loads as its name, a - string. A prediction's string metadata keys become their atoms when those - exist, so `Imp.Prediction.get_lm_usage/1`, `complete?/1` and readers of - `:trace` find them. A map that holds a key as both an atom and a string is - refused on load, as on dump. A trajectory written by 0.5.0 still loads, including one + in :metadata … got: "trace"`), and resuming GEPA from a checkpoint whose + pending proposal batch held a result from an `Imp.predict` program raised + that error. Loading never creates an atom: a key written as an atom the + loading VM does not have loads as its name, a string. A prediction's string + metadata keys become their atoms when those exist, so + `Imp.Prediction.get_lm_usage/1`, `complete?/1` and readers of `:trace` find + them. A map that holds a key as both an atom and a string is refused on + load, as on dump. A trajectory written by 0.5.0 still loads, including one whose prediction has metadata; its other map keys load as the strings they - were written as. Imp 0.5.0 cannot read a trajectory written with atom keys. + were written as. - GEPA checkpoints an agent. A trajectory, and the report codec GEPA uses for a pending batch's reflective dataset and a result's outputs, write an `Imp.History` with `Imp.History.dump/1` and read it with `Imp.History.load!/1`; the trajectory codec also carries - `Imp.Adapter.Types.ToolCalls` and `ToolCallResults`. Before, the first - checkpoint that held an `Imp.react` program's trajectories raised + `Imp.Adapter.Types.ToolCalls` and `ToolCallResults`. The first checkpoint + that held an `Imp.react` program's trajectories raised `Trajectory.DecodeError: trajectory contains an unsupported struct: Imp.History`, under either profile. - An optimizer report (`Imp.Optimizer.Report.encode_term/1`, `dump/1`) that - holds an `Imp.History` now writes it under a `"history"` tag, where it wrote - a plain map; Imp 0.5.0 and earlier builds cannot read such a report. - A GEPA checkpoint resumes in a fresh VM: the loader loads the GEPA modules - whose atoms a checkpoint holds before decoding it. Before, + whose atoms a checkpoint holds before decoding it. `Imp.Optimizer.GEPA.compile_with_report/5` given a checkpoint in a VM that had not yet run GEPA raised `not an already existing atom` on names like `:cache_hits` and `:no_strict_improvement`. @@ -507,10 +481,20 @@ Every change here is breaking for code that relied on the old behaviour. - `Imp.Optimizer.GEPA.compile_with_report/5` no longer raises `KeyError` when it resumes from a checkpoint taken during a full validation; the interrupted validation's rejected candidate is in the report. -- `examples/deployment/agent_optimization.exs` writes the optimizer's - rejections and history with `Imp.Optimizer.Report.json_safe/1`. It passed - them to `Jason.encode!/1`, which raised on a failure reason such as - `{:incomplete_evaluation, 1}`, so a finished run lost its record. + +### Documentation + +- The install instructions ask for a C and a C++ compiler: jaxon builds native + code from C and erlexec from C++. They asked for a C++ compiler only. They + also say the first compile needs network access, for erlexec's rebar3 + plugins. +- The package's Changelog link opens the changelog on HexDocs, and the 0.4.0 + entry links the benchmark pages as they were at v0.4.0, not on `main`. +- Each Livebook's setup cell only installs Imp: from Hex, or from the + checkout `IMP_PATH` names. It no longer searches for a source checkout. +- The MCP example on the Tools and agents page defines its server and its + imported tools, where it used an undefined `imported`, and closes the + server when the call fails. ## 0.5.0 — 2026-09-26 diff --git a/README.md b/README.md index 4a195484..e7f6208a 100644 --- a/README.md +++ b/README.md @@ -145,7 +145,7 @@ text or JSON you can score, such as an agent's tool descriptions. ## Install ```elixir -{:imp, "~> 0.5"} +{:imp, "~> 0.6"} ``` Imp needs Elixir 1.19 or later and a C and C++ compiler, for the native code @@ -154,16 +154,15 @@ access, because erlexec's build fetches rebar3 plugins. It reaches models throug [ReqLLM](https://hex.pm/packages/req_llm), so any provider ReqLLM supports works. -Imp 0.5 is experimental and is its first release on Hex. Its API may still -change, and its optimizers need large-scale benchmarking. Bug reports and -pull requests are welcome. +Imp 0.6 is experimental. Its API may still change, and its optimizers need +large-scale benchmarking. Bug reports and pull requests are welcome. ## Learn - [Getting started](docs/getting-started/index.md) builds one program step by step, from the first call to a supervised server, with real scores. - [Coming from DSPy](docs/coming-from-dspy.md) maps DSPy's names to Imp's. -- [Tutorials](https://github.com/deepfates/imp/tree/v0.5.0/livebooks) are Livebook notebooks +- [Tutorials](https://github.com/deepfates/imp/tree/v0.6.0/livebooks) are Livebook notebooks you can run offline or with a key. - The [cheatsheet](docs/cheatsheet.cheatmd) has the common calls on one page. diff --git a/RELEASE_NOTES.md b/RELEASE_NOTES.md index 67309ba2..be0594e9 100644 --- a/RELEASE_NOTES.md +++ b/RELEASE_NOTES.md @@ -1,19 +1,26 @@ -# Imp v0.5.0 +# Imp v0.6.0 Imp is a framework for typed, optimizable language-model programs on the BEAM. Declare a task as named inputs and outputs, call it like any other Elixir program, measure it on examples, compile it with an optimizer, and run the selected program under OTP. -This release puts Imp on Hex. It is `0.5.0` rather than a patch because the -install line changes, an OTP release that uses the protocol adapters lists one -more application, and ReActV2 and `Imp.MCP.OAuth` change shapes a program may -depend on. +This release fixes bugs found after 0.5.0: failures that were reported as +success, credentials that reached reports and checkpoints, requests that could +be sent twice, and checkpoints that could not be resumed. It is `0.6.0` rather +than a patch because several of those fixes change what a caller receives: +`Imp.collect/3` returns a prediction; a ReActV2 turn whose model fails returns +a `StepError`; examples without declared inputs are refused; a reply that +answers no output is a parse error; `ReqLLMBatch` no longer resends a request +that may have run; an MCP 503 or 529 is `:refused`; `Imp.Datasets.csv/3` +refuses some files 0.5.0 loaded; an `Imp.react` signature may not have a +`tools` field; and GEPA and `Imp.Observability` report errors where they +reported `:ok` or `:failed`. ## Install ```elixir -{:imp, "~> 0.5"} +{:imp, "~> 0.6"} ``` Every dependency comes from Hex. Use a path dependency only while developing @@ -35,224 +42,95 @@ tree calls, and no cowlib release fixes it yet. ## Headline changes -- GEPA works on agents. Optimizing an `Imp.react` agent, the reflection model - reads the whole run (tool calls, tool results, the final answer) and the - agent's tools, and GEPA rewrites the agent's instruction. By default - `Imp.Optimizer.GEPA` behaves as DSPy's GEPA does; Imp's own search is - `execution_profile: :beam_native`. -- An `Imp.Deadline` reaches the work Imp starts for you: `Imp.parallel/3`, - evaluation rows, optimizer workers and runs inherit the caller's deadline, - and `Imp.start_run/3` takes `deadline:`. -- Imp depends on ExMCP 1.5 from Hex, unpatched. What Imp needed from the - `deepfates/ex_mcp` fork now lives in Imp: stdio MCP servers that end with - their connection, children included; a clean `PATH` for them inside a - release; trust for authorized remote servers; the connection options public - servers need; and the browser OAuth flow. -- `ReActV2` offers `submit` only to a signature that needs one. A task with - exactly one text output ends its turn on a step that answers in text, and an - interrupted turn makes one last request whose text is the answer instead of - failing. -- `ReActV2` gains `finish_on` for tools whose call is the answer. -- `:model_request` events record the whole request, and tool definitions are - emitted once per run as `:tools_sent`. -- An MCP tool call that got no answer says whether it was refused, had its - credential refused, was never sent, or may have run (`Imp.MCP.CallFailure`, - `Imp.Tool.outcome/1`), an error result that declares its outcome is read as - declared, and a failed tool call reaches the model as plain text. -- A ReAct prediction's fields are its outputs; how the turn ended is metadata, - in one vocabulary, with `Imp.Prediction.complete?/1`. -- `Imp.Run` and `Imp.ACP` refuse options they do not know, and - `Imp.Run.Event.kinds/0` lists every event kind. -- A host names its own run pool and limit (`Imp.Run.start/3`'s `:admission`), - and a failing run event sink is reported to the run's owner. -- `Imp.MCP.connect/2` takes `pool_size:`, so several calls to one HTTP server - run at once, and an HTTP call can take as long as its `:timeout` allows. -- A run no longer outlives its control process, and a cancellation that never - returns no longer holds a run. +- An agent's step prompt asks for one format, as DSPy's does. With an LM that + calls tools natively, the step no longer also asks for a `tool_calls` field; + an LM that cannot call tools natively is sent no tools and is told them, with + their descriptions and arguments, in a `tools` input. +- Failures are reported instead of passing as success. A ReActV2 turn whose + model fails returns `Imp.Predict.ReActV2.StepError` with the history as far + as it got; a GEPA run that continued past failed proposals reports + `:with_errors`; the Chat, JSON and XML adapters report a reply that answers + no output as a parse error; and evaluation and optimizers refuse examples + that never declared their inputs, which gave the program the answer. +- Redaction runs before a term is converted, in reports, results, + checkpoints, session records and saved programs, so a client, retriever or + OAuth store in them is no longer written with its secrets. Redaction still + replaces the whole string, as in 0.5.0, and recognizes more credential + shapes. +- `Imp.Clients.ReqLLMBatch` never sends again a request that may have run, and + waits for the provider's `retry-after` before a retry. +- A streamed call is recorded in its run like any other, and a reply with text + and tool calls keeps all of them, streamed or not. +- GEPA checkpoints resume: from a pending proposal batch, from a program that + is an agent, and in a fresh VM. +- The lock file takes `mint` 1.11.0, which fixes three advisories + (EEF-CVE-2026-91043, EEF-CVE-2026-92103, EEF-CVE-2026-94194). -## Breaking changes from v0.4.0 +## Upgrading from 0.5 -- Replace `{:imp, github: "deepfates/imp", tag: "v0.4.0"}` with - `{:imp, "~> 0.5"}`. `EX_MCP_PATH` is no longer read. -- A release that uses `Imp.MCP` or `Imp.ACP` adds `erlexec: :load` beside - `ex_mcp: :load`. -- Trust for an authorized remote MCP server is VM-wide. While a connection to - it is open, its exact origin (`scheme://host:port`) is in ExMCP's - `trusted_origins`, so any ExMCP client in the same VM may send credential - headers to that origin without consent. In 0.4.0 the trust belonged to the - one connection. No other origin is trusted, the origin is removed when the - last connection to it closes, and origins the host configured are left - alone. A host that runs other ExMCP clients it does not trust with those - origins should know this. -- `Imp.Optimizer.GEPA` defaults to DSPy's GEPA, `execution_profile: - :gepa_v0_1_4_merge`: merge on, no evaluation cache, perfect minibatches - skipped, the pinned RNG, and `:generations` turned into a metric budget when - `:max_metric_calls` is not given. Options the DSPy profiles fix (ComBee, - `:feedback_fn`, `:module_selector`, `:candidate_selection_strategy`, - `:proposal_concurrency`, `:reflection_strategy`, the frontier, sampling, - selection, evaluation and acceptance policies, `:max_reflection_calls`, and - `reflection_record_mode: :beam_native`) raise unless `execution_profile: - :beam_native` is given, which is the 0.4.0 behaviour. Resuming a checkpoint - written by a 0.4.0-default run raises under the new default; resume it with - `execution_profile: :beam_native`. -- `Imp.MCP.OAuth.begin/3` no longer takes `:flow`; a pre-registered client is - `client_registration: {:pre_registered, client_id, client_secret}` with - `client_issuer:` naming the authorization server it belongs to. A server - with no OAuth metadata at all is refused instead of given guessed endpoints. -- An `Imp.Tool` named with a string keeps the string, and tools imported from - an MCP server are named by the server's string. Code that compared an - imported tool's `name` to an atom compares it to the string. -- An MCP tool call that got no answer returns - `{:error, %Imp.MCP.CallFailure{}}` instead of - `{:mcp_tool_call_failed, server, reason}` or - `{:mcp_connection_unavailable, server, reason}`. A call that reaches its - `:timeout` is `%Imp.MCP.CallFailure{outcome: :unknown, reason: :timeout}`, - answered at the timeout while the request runs on; a call to an HTTP server - whose connections all stay busy until the timeout is `:not_sent` with - `reason: :no_idle_connection`. -- `"type" => "sse"` is MCP's deprecated HTTP+SSE transport, and its `"url"` - is the event stream's. In 0.4.0 it was Streamable HTTP with a standing GET - stream; a Streamable HTTP server is now `"type" => "http"`. An `sse` - descriptor with `"headers"` or `"auth"`, or with a query string in its URL, - is refused before anything is dialed (`:mcp_sse_credentials_refused`, - `:mcp_sse_url_refused`): the whole import under the default - `on_failure: :refuse`, only that server under `on_failure: :drop`. -- When a run's control process ends while the run is still going, the task is - killed after its registered cancellations are called; its monitor reports - `:killed`. -- A run's owner can receive `{:imp_run_event_sink_failed, run_id, details}` - and `{:imp_run_event_undelivered, run_id, event}`; an owner with a strict - `handle_info/2` needs clauses for them. -- `Imp.Run.start/3`, `Imp.ACP.start_link/1`, `Imp.ACP.run/1` and - `Imp.ACP.Local.start_link/1` raise `ArgumentError` for an option they do - not know. A transport's own options for `Imp.ACP` go in - `:transport_options`, and `:capabilities` is spelled `:agent_capabilities`. -- `Imp.predict/2`, `Imp.chain_of_thought/2` and `Imp.configure/1` raise - `ArgumentError` for an option or setting they do not know. Request options - such as `:temperature` go under `config:`; a setting of your own goes - through `Imp.context/2`. -- `:max_errors` and `:retriever` are no longer settings, and - `Imp.configure/1` and `Imp.context/2` refuse them. Pass `:max_errors` to - BootstrapFewShot, RandomSearch or COPRO (10 when not given) and a retriever - to the program. -- ReActV2 emits no `:final` event; `:run_finished` carries the prediction. - `Imp.Trajectory.to_atif/2`'s `extra.outcome` is `extra.terminal_event`, and - a tool result's `extra.outcome` is the recorded `Imp.Tool.outcome/1` - instead of `"returned"` or `"error"`. -- A ReActV2 or ReAct prediction's fields are its outputs only: `history`, - `termination_reason`, `termination_cause`, `termination_error`, - `finished_by_tool`, `unexecuted_tool_calls` and `context_projection` are in - `prediction.metadata`. `termination_reason` says how the turn ended, and a - turn without an answer is `:incomplete` with `termination_cause` saying why; - typed extraction is `:extracted` (no `completion_mode`), and - `Imp.Predict.ReAct` spells `:parse_failure` as `:parse_error` and `:direct` - as `:answered`. Use `Imp.Prediction.complete?/1` to ask whether a turn - answered. -- For a signature with one `:string` output, `ReActV2` offers no `submit` - tool, and a step answered in text with no tool call ends the turn. -- Errors have one shape per tag, with the reason as a term. A failed - `Imp.Clients.ReqLLM` request is `%Imp.LMError{}` (with `status`, - `retryable` and `context_window_exceeded`; `Imp.ContextWindowExceededError` - is gone), and a completion that cannot be parsed is - `%Imp.AdapterParseError{kind: ...}`, which `Imp.Predict` returns - directly instead of `%{reason: {:error, _}, trace: _}`. A raise inside a - client, program, tool, tool policy, retriever, optimizer or ACP callback - keeps the exception struct where 0.4.0 kept its message. - `{:tool_denied, tool}` is `{:tool_denied, tool, :tool_policy}`, and a run's - `:authorize` refusal is `{:tool_denied, tool, reason}`; `Refine` and - `Assertions` return `{:error, reason}`; - `Imp.optimize!` raises `Imp.Error` for a failed optimization. The CHANGELOG - lists every tag that changed. -- `Imp.Example` and `Imp.Prediction` keep string keys as strings. Code that - read a field of data loaded from JSON with `map.field` or `map[:field]` - reads it with `Imp.Example.get/2` or by its string key. -- `Imp.MCP.Client`, `Imp.MCP.HTTPClient`, `Imp.MCP.StreamableHTTPClient`, - `Imp.MCP.StdioClient`, `Imp.MCP.Catalog`, `Imp.MCP.import_tools`, - `Imp.ACP.MCP` and `Imp.Core.ToolCall`/`ToolResult` are gone. - `Imp.MCP.connect/2` imports tools; `Imp.ACP.ToolKind.derive_all/1` gives an - import's ACP tool kinds. -- A saved program holds no HTTP header, credential or not. An LM with custom - headers (a routing header such as `x-tenant` included) sends requests - without them after loading until it is rebound with `Imp.with_lm/2` or a - scoped `Imp.context/2`. -- `Imp.save!` refuses an LM whose `base_url` has a query, fragment or user - info. -- Prompts name types in plain words instead of Python annotations - (`one of: atlas, harbor` where 0.4.0 wrote `Literal['atlas', 'harbor']`), - values take their JSON spelling (`null`, `true`, `false`), and the - structured-output schema is named `outputs`. A non-string answer for a - string field is kept as its JSON text (`"true"`, not `"True"`), and a - `null` answer is no value rather than the string `"None"`. Fields, order - and parsing are unchanged, but a saved optimized program now sends - different prompt text. -- An `Imp.Telemetry` span's `[:exception]` event carries `:kind`, `:reason` - and `:stacktrace`, as `:telemetry.span/3` does, instead of `:error` as - text. -- An optimizer's `compile/N` is no longer documented where `Imp.optimize` or - `Imp.train` runs the optimizer; call those. -- `Imp.load!/1` reading a file is `Imp.read!/1`; `Imp.load/1` returns - `{:ok, program}` and `Imp.load!/1` takes the dumped map. -- `Imp.react/3` builds ReActV2 and `Imp.react_v2` is gone. - `Imp.Predict.ReAct`'s `mode: :dspy_3_2_1` is `mode: :dspy`. -- `Imp.Predict.Predict` is `Imp.Predict`. -- `Imp.Retrievers.KNN` is deleted; `Imp.Retrieve.Memory` retrieves by token - overlap. -- An LM is a struct or module whose `generate/3` takes it first. The - `%{module:, opts:}` map and a bare function are refused; so is a module - that defines only `generate/2`. A retriever module's `retrieve/3` takes - itself first. -- `Imp.MCP.CallFailure` has `server_name`, `tool_name` and `index`; an - `unavailable` entry has `server_name`; `:authorize` returns `:allow` or - `{:deny, reason}` and its context names the `descriptor`; - `:credentials` is `:credential_store`. -- A `:tool_policy` function returns `:allow` or `{:deny, reason}`, and a - refused call is `{:tool_denied, name, reason}`. -- `max_concurrency` is `num_threads` on evaluation, parallel, search, batch and - optimizer options. `Refine`'s `max_attempts` is `n`; `RLM` takes - `max_iterations` only. `Imp.Optimizer.RandomSearch` and `BootstrapRS` are - `Imp.Optimizer.BootstrapFewShotWithRandomSearch`. +1. Change the dependency to `{:imp, "~> 0.6"}`, run `mix deps.get`, and commit + `mix.lock`. `nimble_csv` is a new dependency. +2. `Imp.collect/3` returns `{:ok, prediction}` or `{:error, reason}`; read + fields with `Imp.get(prediction, :answer)` instead of matching a string. +3. Where you checked `Imp.Prediction.complete?/1` after a ReActV2 model + failure, match `{:error, %Imp.Predict.ReActV2.StepError{reason: reason, + history: history}}` and store `history` as you store a finished turn's. A + step refused by an `Imp.OperationalSafetyError` also ends the turn this + way, at once. +4. Call `Imp.with_inputs/2` on every example you evaluate or optimize on, and + build evaluation rows with `Imp.example/1 |> Imp.with_inputs(...)` rather + than plain maps or pair lists. Give each field of an example or prediction + once, under one spelling. To keep a signature's instructions, pass it to + `Imp.Signature.new/2` without instructions. +5. Rename a field called `tools` in an `Imp.react` task signature, and save + the program again; a saved program with one no longer loads. +6. Match `%Imp.AdapterParseError{kind: :missing_fields}` where you relied on a + prediction of defaults, or on ProgramOfThought's `:missing_program`, for a + reply that answered no output. +7. Compare `Imp.Clients.ReqLLM` clients by `model`, not by the whole struct, + which now carries `:tool_calling`. +8. Treat a `ReqLLMBatch` request that ends `:ambiguous` as possibly run, and + check it with the provider before sending it again. A 0.5.0 checkpoint's + `:transient_failure` requests become `:ambiguous` on resume. +9. An MCP call answered with 503 or 529 is `:refused`, not `:unknown`. +10. Quote CSV fields that hold a quote (`"12"" pipe",1`) and remove spaces + between a comma and a quoted field; `Imp.Datasets.csv/3` refuses both. +11. Match `Imp.Observability.Status`'s new `:succeeded_with_errors`, which an + optimizer report with errors gives where it gave `:failed`. Read GEPA's + `report.errors` rather than expecting `status: :ok`: a run that went on + past failed proposals reports `:with_errors`. Review `:beam_native` + stoppers built on `consecutive_outcome/2`, which now counts an iteration + that raised as `:proposal_error` instead of `:none`. +12. Keep 0.5.0 away from files 0.6.0 writes: it cannot read a trajectory with + atom keys (GEPA and Playbook checkpoints, `Imp.dump/1`), an optimizer + report that holds an `Imp.History`, or a `ReqLLMBatch` checkpoint. +13. Re-evaluate saved agents on held-out data: the step prompt and tool roster + order changed, so they send different prompt text. -## Upgrade path +## Known limits -1. Change the dependency line, run `mix deps.get`, and commit `mix.lock`. -2. Add `erlexec: :load` to any release that lists `ex_mcp: :load`. -3. Replace `OAuth.begin/3`'s `:flow` with `:client_registration` if you used it. -4. Match MCP call failures on `%Imp.MCP.CallFailure{outcome: ...}` (a 401 is - `:auth_refused`), compare imported tool names as strings, and give run - owners clauses for `:imp_run_event_sink_failed` and - `:imp_run_event_undelivered`. -5. Read a ReAct or ReActV2 prediction's `history` and `termination_*` from - `prediction.metadata`, and match `termination_reason` against the new - values. -6. Change `"type" => "sse"` descriptors for Streamable HTTP servers to - `"http"`. A server that needs credentials is reached over Streamable HTTP; - an `sse` descriptor takes none. -7. Match LM failures on `%Imp.LMError{}` (or ask `Imp.Errors.retryable?/1` - and `Imp.Errors.context_window_exceeded?/1`), parse failures on - `%Imp.AdapterParseError{kind: ...}`, and exception reasons on the struct - rather than its text. -8. Rename `Imp.react_v2` to `Imp.react`, `Imp.Predict.Predict` to - `Imp.Predict`, and `Imp.load!(path)` to `Imp.read!(path)`; give custom LMs - and retriever modules the `generate/3` and `retrieve/3` that take the - client first, and wrap an LM function in a struct that implements - `Imp.LM`. -9. Rename `max_concurrency:` to `num_threads:` where you configure - evaluation, optimizers or parallel calls; answer `:authorize` and - `:tool_policy` with `:allow` or `{:deny, reason}`; match a refused tool - call as `{:tool_denied, name, reason}`; read `server_name` and `tool_name` - from MCP failures and absences. Call `Imp.Signature.load!/1`, - `Imp.History.load!/1`, `Imp.Optimizer.Report.load!/1` and - `Imp.Clients.TrainingJob.load!/2` where you called `load`, and - `Imp.Clients.TrainingJob.read!/2` where you read a checkpoint file, and - the datasets' `read!` where you called their `load(path)`. -10. Rebind the LM of any loaded program that relies on custom headers, and - move a `base_url` query, fragment or user info into configuration the - host supplies at load time. -11. Update telemetry handlers for `[:exception]` to read `:kind`, `:reason` - and `:stacktrace`. -12. Re-evaluate saved optimized programs on held-out data, since their - prompt text changed, and run your application smoke test against the new - release; a one-text-output ReActV2 program now ends turns differently. +- A flat `["api_key", key]` held directly as a value, under a key that is not + itself a credential name, is not redacted, where 0.5.0 redacted it. A + two-element list with a name first is read as a key and value only as an + element of a list, so that lists of names such as + `with_inputs([:api_key, :question])` survive saving. +- A GEPA checkpoint redacts its failure reasons but otherwise holds the resume + state as it is: candidates, the evaluation cache, proposed instructions and + reflection data. Fast-Slow and Playbook checkpoints likewise hold their + prompts, rollouts and playbooks. Treat checkpoint files as sensitive. +- A ReActV2 step whose prose quotes its own field names (a JSON object with + `next_thought` or `tool_calls` keys, a `[[ ## field ## ]]` line, a + `` tag) is not read as prose and goes to the JSON fallback, + which usually costs one more call. If the model repeats the same reply, the + step ends `:incomplete`. +- When a streamed ReqLLM request fails, ReqLLM's own warning log prints the + error with its response headers. Imp strips them from the error it returns, + but cannot change that log line. +- In command output, a `Bearer` token of 12 to 15 characters, or one with no + digit, is shown when more text or a newline follows it. 0.5.0's command + pipeline hid any such token of 12 or more characters. +- A GEPA checkpoint taken between preparing and starting a full validation + resumes by rejecting that candidate, as before. The [CHANGELOG](CHANGELOG.md) records every user-visible change in this release. Generated module documentation is the complete API reference. Start diff --git a/docs/coming-from-dspy.md b/docs/coming-from-dspy.md index 2989213e..e60dae08 100644 --- a/docs/coming-from-dspy.md +++ b/docs/coming-from-dspy.md @@ -12,7 +12,7 @@ brought in. ## The five-minute version ```elixir -# pip install dspy -> {:imp, "~> 0.5"} +# pip install dspy -> {:imp, "~> 0.6"} # lm = dspy.LM("openai/...") -> lm = Imp.req_llm("openai:gpt-5.4-mini", api_key: ...) # dspy.Predict("q -> a") -> program = Imp.predict("q -> a", lm: lm) # program(q="...") -> {:ok, pred} = Imp.call(program, %{q: "..."}) @@ -87,7 +87,7 @@ error row too. Each model request has a receive timeout, and has no time limit of its own (an MCP tool call does, 30 seconds by default), so to bound an agent, start it with `Imp.start_run/3` and cancel it when you choose. The -[deployment example](https://github.com/deepfates/imp/blob/v0.5.0/examples/deployment/README.md) +[deployment example](https://github.com/deepfates/imp/blob/v0.6.0/examples/deployment/README.md) is a complete OTP application. **Optimizers see what you name.** DSPy finds predictors by walking a module's diff --git a/docs/diving-deeper/modules-and-composition.md b/docs/diving-deeper/modules-and-composition.md index b01df57e..f7dfabaa 100644 --- a/docs/diving-deeper/modules-and-composition.md +++ b/docs/diving-deeper/modules-and-composition.md @@ -149,7 +149,7 @@ A custom module is your code, so it is not saved as a whole. What an optimizer chose for it is data: `Imp.ProgramParameters.values/1` reads it, `Imp.ProgramParameters.apply_values/2` puts it on a freshly built program, and `Imp.Optimizer.Artifact` writes it to a checksummed file. The -[deployment example](https://github.com/deepfates/imp/blob/v0.5.0/examples/deployment/README.md) +[deployment example](https://github.com/deepfates/imp/blob/v0.6.0/examples/deployment/README.md) does this in a supervised application. ### Built-in variants diff --git a/docs/diving-deeper/saving-and-artifacts.md b/docs/diving-deeper/saving-and-artifacts.md index 5d0ed3d4..fa60ff26 100644 --- a/docs/diving-deeper/saving-and-artifacts.md +++ b/docs/diving-deeper/saving-and-artifacts.md @@ -21,7 +21,7 @@ Save the whole program when it is made of Imp's modules: `Imp.predict/2`, when the program is your own struct implementing `Imp.Module`, or when you want the program's code to live in your release and only its tuned text to cross the persistence boundary. The -[deployment example](https://github.com/deepfates/imp/tree/v0.5.0/examples/deployment) +[deployment example](https://github.com/deepfates/imp/tree/v0.6.0/examples/deployment) does the second. ### 2. JSON with a checksum diff --git a/docs/getting-started/running-in-your-application.md b/docs/getting-started/running-in-your-application.md index 7bfeaaae..9991337e 100644 --- a/docs/getting-started/running-in-your-application.md +++ b/docs/getting-started/running-in-your-application.md @@ -121,7 +121,7 @@ Changing models is then a configuration change, and rotating a key is a restart. The -[deployment example](https://github.com/deepfates/imp/blob/v0.5.0/examples/deployment/README.md) +[deployment example](https://github.com/deepfates/imp/blob/v0.6.0/examples/deployment/README.md) is the complete version of this page: a two-stage program like `TicketTriage`, learned parameters loaded and reloaded as an `Imp.Optimizer.Artifact`, crashed and timed-out calls contained, and the whole thing restarted in a fresh OS diff --git a/docs/getting-started/setting-up.md b/docs/getting-started/setting-up.md index e4c0f683..f4bbe249 100644 --- a/docs/getting-started/setting-up.md +++ b/docs/getting-started/setting-up.md @@ -5,7 +5,7 @@ Add Imp to a Mix project: ~~~elixir def deps do [ - {:imp, "~> 0.5"} + {:imp, "~> 0.6"} ] end ~~~ @@ -13,7 +13,7 @@ end or, for a script or a Livebook notebook: ~~~elixir -Mix.install([{:imp, "~> 0.5"}]) +Mix.install([{:imp, "~> 0.6"}]) ~~~ Imp needs Elixir 1.19 and a C and C++ compiler, for the native code in two diff --git a/docs/getting-started/where-to-go-next.md b/docs/getting-started/where-to-go-next.md index 740cc223..4ee7c144 100644 --- a/docs/getting-started/where-to-go-next.md +++ b/docs/getting-started/where-to-go-next.md @@ -51,7 +51,7 @@ The module documentation is the reference for every function. [Running Imp in production](../production.md) covers what the last page began: supervision, provider failures, runs you can observe and cancel, and telemetry. The -[deployment example](https://github.com/deepfates/imp/blob/v0.5.0/examples/deployment/README.md) +[deployment example](https://github.com/deepfates/imp/blob/v0.6.0/examples/deployment/README.md) is a complete application to copy from. ## Try it in a notebook diff --git a/docs/production.md b/docs/production.md index c7ea798a..7e52a8db 100644 --- a/docs/production.md +++ b/docs/production.md @@ -5,16 +5,16 @@ where its credentials come from, how long a call may take and what it may cost. This page covers those decisions in the order you meet them: starting Imp, credentials, concurrency, timeouts, cost and caching, persistence and telemetry. The -[deployment example](https://github.com/deepfates/imp/tree/v0.5.0/examples/deployment) +[deployment example](https://github.com/deepfates/imp/tree/v0.6.0/examples/deployment) is a complete OTP application that puts them together. ## Imp in your supervision tree -Imp is an OTP application, and adding `{:imp, "~> 0.5"}` to your +Imp is an OTP application, and adding `{:imp, "~> 0.6"}` to your dependencies starts it with yours. It supervises its settings, its response cache and the task pools that bound its work. It opens no listener and starts no subprocess, and the MCP and ACP runtimes start only when you use -them. In a script, `Mix.install([{:imp, "~> 0.5"}])` starts it too. +them. In a script, `Mix.install([{:imp, "~> 0.6"}])` starts it too. ## Credentials at runtime @@ -252,7 +252,7 @@ standard error. ## The reference application -[`examples/deployment`](https://github.com/deepfates/imp/tree/v0.5.0/examples/deployment) +[`examples/deployment`](https://github.com/deepfates/imp/tree/v0.6.0/examples/deployment) is a small OTP application that loads a checksummed artifact at startup, reads its key from the environment, serves calls from bounded supervised tasks, answers overload and timeouts with errors instead of blocking, and reloads diff --git a/examples/deployment/README.md b/examples/deployment/README.md index a35debff..8b2080ae 100644 --- a/examples/deployment/README.md +++ b/examples/deployment/README.md @@ -150,4 +150,4 @@ Use a whole-program artifact for a supported portable Imp shape. Use its selected predictor parameters should cross the persistence boundary. During source development set `IMP_PATH` to the Imp checkout. Other -applications use the Hex dependency in `mix.exs`, `{:imp, "~> 0.5"}`. +applications use the Hex dependency in `mix.exs`, `{:imp, "~> 0.6"}`. diff --git a/examples/deployment/mix.exs b/examples/deployment/mix.exs index 9c50035e..0fa677c8 100644 --- a/examples/deployment/mix.exs +++ b/examples/deployment/mix.exs @@ -11,7 +11,7 @@ defmodule ImpDeployment.MixProject do defp deps do case System.get_env("IMP_PATH") do - nil -> [{:imp, "~> 0.5"}] + nil -> [{:imp, "~> 0.6"}] path -> [{:imp, path: path}] end end diff --git a/livebooks/01_real_lm_front_door.livemd b/livebooks/01_real_lm_front_door.livemd index 36798171..90ffc1b4 100644 --- a/livebooks/01_real_lm_front_door.livemd +++ b/livebooks/01_real_lm_front_door.livemd @@ -7,7 +7,7 @@ imp_path = System.get_env("IMP_PATH") lockfile = imp_path && Path.join(imp_path, "mix.lock") cond do - is_nil(imp_path) -> Mix.install([{:imp, "~> 0.5"}]) + is_nil(imp_path) -> Mix.install([{:imp, "~> 0.6"}]) File.regular?(lockfile) -> Mix.install([{:imp, path: imp_path}], lockfile: lockfile) true -> Mix.install([{:imp, path: imp_path}]) end diff --git a/livebooks/02_without_a_provider.livemd b/livebooks/02_without_a_provider.livemd index c799b977..c8343e8b 100644 --- a/livebooks/02_without_a_provider.livemd +++ b/livebooks/02_without_a_provider.livemd @@ -7,7 +7,7 @@ imp_path = System.get_env("IMP_PATH") lockfile = imp_path && Path.join(imp_path, "mix.lock") cond do - is_nil(imp_path) -> Mix.install([{:imp, "~> 0.5"}]) + is_nil(imp_path) -> Mix.install([{:imp, "~> 0.6"}]) File.regular?(lockfile) -> Mix.install([{:imp, path: imp_path}], lockfile: lockfile) true -> Mix.install([{:imp, path: imp_path}]) end diff --git a/livebooks/03_evaluate_and_optimize.livemd b/livebooks/03_evaluate_and_optimize.livemd index 5ee25947..9d63ec3c 100644 --- a/livebooks/03_evaluate_and_optimize.livemd +++ b/livebooks/03_evaluate_and_optimize.livemd @@ -7,7 +7,7 @@ imp_path = System.get_env("IMP_PATH") lockfile = imp_path && Path.join(imp_path, "mix.lock") cond do - is_nil(imp_path) -> Mix.install([{:imp, "~> 0.5"}]) + is_nil(imp_path) -> Mix.install([{:imp, "~> 0.6"}]) File.regular?(lockfile) -> Mix.install([{:imp, path: imp_path}], lockfile: lockfile) true -> Mix.install([{:imp, path: imp_path}]) end diff --git a/livebooks/04_tools_agents_mcp_rlm.livemd b/livebooks/04_tools_agents_mcp_rlm.livemd index 1016fa6f..581fbb64 100644 --- a/livebooks/04_tools_agents_mcp_rlm.livemd +++ b/livebooks/04_tools_agents_mcp_rlm.livemd @@ -7,7 +7,7 @@ imp_path = System.get_env("IMP_PATH") lockfile = imp_path && Path.join(imp_path, "mix.lock") cond do - is_nil(imp_path) -> Mix.install([{:imp, "~> 0.5"}]) + is_nil(imp_path) -> Mix.install([{:imp, "~> 0.6"}]) File.regular?(lockfile) -> Mix.install([{:imp, path: imp_path}], lockfile: lockfile) true -> Mix.install([{:imp, path: imp_path}]) end diff --git a/livebooks/05_operating_imp.livemd b/livebooks/05_operating_imp.livemd index 6caa0db8..720406b9 100644 --- a/livebooks/05_operating_imp.livemd +++ b/livebooks/05_operating_imp.livemd @@ -7,7 +7,7 @@ imp_path = System.get_env("IMP_PATH") lockfile = imp_path && Path.join(imp_path, "mix.lock") cond do - is_nil(imp_path) -> Mix.install([{:imp, "~> 0.5"}]) + is_nil(imp_path) -> Mix.install([{:imp, "~> 0.6"}]) File.regular?(lockfile) -> Mix.install([{:imp, path: imp_path}], lockfile: lockfile) true -> Mix.install([{:imp, path: imp_path}]) end @@ -103,7 +103,7 @@ Enum.each(busy, &Task.await/1) refused ``` -The [deployment example](https://github.com/deepfates/imp/tree/v0.5.0/examples/deployment) +The [deployment example](https://github.com/deepfates/imp/tree/v0.6.0/examples/deployment) wraps the same pattern in a `ProgramServer` that also reloads parameters while serving. diff --git a/mix.exs b/mix.exs index 5837ef2a..efb33ce9 100644 --- a/mix.exs +++ b/mix.exs @@ -1,7 +1,7 @@ defmodule Imp.MixProject do use Mix.Project - @version "0.5.0" + @version "0.6.0" def project do [ diff --git a/priv/public_api.json b/priv/public_api.json index 8de1b3dc..fd825129 100644 --- a/priv/public_api.json +++ b/priv/public_api.json @@ -10337,7 +10337,7 @@ "types": [] } ], - "package_version": "0.5.0", + "package_version": "0.6.0", "schema_version": 3, "scope": "Curated Imp API compiled from package-shipped sources. Exports come from non-hidden documentation entries; callbacks, types, and struct fields are category- and policy-gated." } diff --git a/test/package_contract_test.exs b/test/package_contract_test.exs index 19f9cfe3..4496ef27 100644 --- a/test/package_contract_test.exs +++ b/test/package_contract_test.exs @@ -110,7 +110,7 @@ defmodule PackageContractTest do version = Mix.Project.config()[:version] dependency = hex_dependency(version) - assert version == "0.5.0" + assert version == "0.6.0" assert File.read!("RELEASE_NOTES.md") =~ "# Imp v#{version}" assert File.read!("CHANGELOG.md") =~ "## #{version}" assert File.read!("examples/deployment/mix.exs") =~ dependency From a6aa389abc99bf3f3206d299e6748a2bd6d82daa Mon Sep 17 00:00:00 2001 From: deepfates Date: Mon, 28 Sep 2026 06:39:36 -0700 Subject: [PATCH 2/8] Release notes: plainer tool wording; the interrupted-turn answer change in Upgrading --- RELEASE_NOTES.md | 13 +++++++++---- 1 file changed, 9 insertions(+), 4 deletions(-) diff --git a/RELEASE_NOTES.md b/RELEASE_NOTES.md index be0594e9..78aed10b 100644 --- a/RELEASE_NOTES.md +++ b/RELEASE_NOTES.md @@ -44,14 +44,15 @@ tree calls, and no cowlib release fixes it yet. - An agent's step prompt asks for one format, as DSPy's does. With an LM that calls tools natively, the step no longer also asks for a `tool_calls` field; - an LM that cannot call tools natively is sent no tools and is told them, with - their descriptions and arguments, in a `tools` input. + an LM that cannot is shown each tool's description and arguments in the + prompt and writes its calls in `tool_calls`. - Failures are reported instead of passing as success. A ReActV2 turn whose model fails returns `Imp.Predict.ReActV2.StepError` with the history as far as it got; a GEPA run that continued past failed proposals reports `:with_errors`; the Chat, JSON and XML adapters report a reply that answers no output as a parse error; and evaluation and optimizers refuse examples - that never declared their inputs, which gave the program the answer. + that never declared their inputs, whose labels were passed to the program + as inputs. - Redaction runs before a term is converted, in reports, results, checkpoints, session records and saved programs, so a client, retriever or OAuth store in them is no longer written with its secrets. Redaction still @@ -104,7 +105,11 @@ tree calls, and no cowlib release fixes it yet. 12. Keep 0.5.0 away from files 0.6.0 writes: it cannot read a trajectory with atom keys (GEPA and Playbook checkpoints, `Imp.dump/1`), an optimizer report that holds an `Imp.History`, or a `ReqLLMBatch` checkpoint. -13. Re-evaluate saved agents on held-out data: the step prompt and tool roster +13. A ReActV2 turn that reaches `max_iters` with text beside tool calls it + did not run now answers with that text, where it answered `nil`; the calls + are still listed as unexecuted. A host that publishes every non-empty + answer should decide whether to publish such a turn's text. +14. Re-evaluate saved agents on held-out data: the step prompt and tool roster order changed, so they send different prompt text. ## Known limits From 61bc81df497bf65aaced606047c407f0705a451e Mon Sep 17 00:00:00 2001 From: deepfates Date: Mon, 28 Sep 2026 07:14:33 -0700 Subject: [PATCH 3/8] Release notes and changelog: fact-check corrections and omissions --- CHANGELOG.md | 88 ++++++++++++++++++++++++++++++++---------------- RELEASE_NOTES.md | 72 +++++++++++++++++++++++++-------------- 2 files changed, 106 insertions(+), 54 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 5d971b00..53e56752 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,16 +9,19 @@ User-visible changes to Imp are recorded here. - The lock file takes `mint` 1.11.0 (and `hpax` 1.1.0), which fixes EEF-CVE-2026-91043, EEF-CVE-2026-92103 and EEF-CVE-2026-94194 in the HTTP client Req uses. The example projects' lock files take the same versions. + Imp's lock file does not bind an application that depends on Imp, and Imp + sets no `mint` floor: run `mix deps.update mint hpax` in your own project. - `Imp.Redaction` redacts a string that holds a PEM private key of any type, a PGP private key block or a PuTTY key file; a Stripe (`sk_live_`, `rk_live_`), GitHub, GitLab (`glpat-`), Hugging Face (`hf_`), Slack - (`xox?-`), SendGrid (`SG.`), npm (`npm_`), PyPI (`pypi-`), Google OAuth + (`xox?-`), SendGrid (`SG.`), npm (`npm_`), PyPI (`pypi-AgE`), Google OAuth (`ya29.`) or Vault (`hvs.`) token, or a JSON Web Token; an AWS access key id longer than 20 characters; a URL with a password in its user info; or a signed URL's `X-Amz-Signature`, `X-Amz-Security-Token`, `X-Goog-Signature` or Azure `sig`. It passed these unchanged into run events, traces, trajectories and saved programs. As before, the whole string is replaced, - and every shape it caught before is still caught. + and every shape it caught before is still caught. It tries 22 patterns + where 0.5.0 tried 7, so it is slower on text that holds no credential. - A `Bearer` token of 16 or more characters with a digit in it is found wherever it ends: mid-line (`Bearer https://…`) or before a newline (`"Authorization: Bearer \nmachine …"`). Both were missed. A word @@ -29,8 +32,11 @@ User-visible changes to Imp are recorded here. - `Imp.ExternalCommand` redacts captured output with the same rules, in place of its own `sk-` and `Bearer` patterns: output that holds a credential is `"[REDACTED]"` whole, so the credentials printed beside a - recognized one (`env | grep AWS`, a credentials file) go with it. The - optimizer's pricing URL check uses the same rules instead of its own. + recognized one (`env | grep AWS`, a credentials file) go with it. A + `Bearer` token of 12 to 15 characters, or one with no digit, followed by + more text or a newline is shown, where 0.5.0's own `Bearer` pattern hid any + such token of 12 or more characters. The optimizer's pricing URL check uses + the same rules instead of its own. - Writers redact a term before converting it: optimizer reports (`Imp.Optimizer.Report.dump/1`, `json_safe/1`, `json_projection/1`), experiment results, `Imp.Evaluate.Result.save_as_json/2` and @@ -59,6 +65,8 @@ User-visible changes to Imp are recorded here. walked as plain maps, so the store's key reached any redacted output that held a store. `code_verifier` is a credential name, so a PKCE transaction map is redacted on its own too. +- An error about a devset, evaluation row, field entry or demo that is not + an example names its type and key names, never its values. - A map key that is a string shaped like a credential is redacted. `Imp.Optimizer.Trajectory` names a key it refuses (one that is not an atom or a string) by its type; it printed the key, credential included. @@ -81,9 +89,10 @@ User-visible changes to Imp are recorded here. held such a list, did not load back as it was; `Imp.Redaction.redact/1` rewrote an Avatar tool schema's `"required" => ["api_key", "query"]`; and `Imp.Optimizer.Parameter` refused input keys `["api_key", "question"]`. One - divergence from 0.5.0 follows: a flat `["api_key", key]` held directly as a - value, under a key that is not itself a credential name, is not redacted, - since it has the shape of a list of two names. + divergence from 0.5.0 follows: a flat two-element name-first list that is + not itself an element of a list (at the top level, in a tuple, or under a + key that is not a credential name), such as `["password", "hunter2"]`, is + not redacted, since it has the shape of a list of two names. ### Changed @@ -119,21 +128,26 @@ User-visible changes to Imp are recorded here. `{:error, %Imp.Predict.ReActV2.StepError{reason: reason, history: history}}` and store `history` as you store a finished turn's. - Breaking: an `Imp.Predict.ReActV2` step refused by an - `Imp.OperationalSafetyError` (a route, cost, transport or budget guard) ends - the turn at once with a `StepError` whose `:reason` is that error. The turn - made its last request instead, and an answer to it passed the guard by. - `Imp.Evaluate` and the optimizers find the guard inside the `StepError` and - raise it. Migration: none for a caller that treats safety errors as fatal. + `Imp.OperationalSafetyError` (a route, cost, transport, budget or + cancellation guard) ends the turn at once with a `StepError` whose `:reason` + is that error. The turn made its last request instead, and an answer to it + passed the guard by. `Imp.Evaluate` and the optimizers find the guard inside + the `StepError` and raise it. Migration: where you call the program + yourself, match `{:error, %Imp.Predict.ReActV2.StepError{reason: + %Imp.OperationalSafetyError{}}}`. - Breaking: an `Imp.react` task signature cannot have a field named `tools`, which is now the step's tool list, and loading a program saved with one - raises the same `ArgumentError`. Migration: rename the field and save the - program again. + raises the same `ArgumentError`. Migration: rebuild the program with the + field renamed and save it; to keep a saved program's optimized instructions + and demos, rename the field in the saved file, or optimize again. - Breaking: `Imp.Clients.ReqLLM` has a `:tool_calling` field: whether the model calls tools natively, read from the registry once when the client is built or loaded, and kept with the model it was read for. A client whose model is swapped looks the new model up. A client built by `Imp.req_llm/2`, or loaded, therefore no longer equals a bare `%Imp.Clients.ReqLLM{}` for the - same model. Migration: compare clients by `model`, not by the whole struct. + same model. Likewise `%Imp.Predict.ReActV2{}` has a `:tool_order` field, + the tools' declared order. Migration: compare clients by `model`, not by + the whole struct. - Breaking: in `Imp.Adapter.Chat`, `Imp.Adapter.JSON` and `Imp.Adapter.XML`, a reply that answers none of the requested outputs (no `[[ ## field ## ]]` section, no output key, no output tag) is an @@ -148,7 +162,10 @@ User-visible changes to Imp are recorded here. writes fields in some format (a JSON object with an output's key, a `[[ ## field ## ]]` line, an output's tag) is not prose: Chat and XML parse it or report it, so a step that spelled out a tool call as JSON runs that - tool instead of answering with the JSON text. A JSON `{}` for a ReActV2 + tool instead of answering with the JSON text. So a step whose prose quotes + its own field names goes to the JSON fallback, which usually costs one more + call; beside native tool calls such a thought becomes `nil`, and the calls + still run. A JSON `{}` for a ReActV2 step is `:missing_fields`, where it was a step with `nil` fields. `Imp.Predict.ProgramOfThought`, whose outputs are all optional, sends prose to the JSON fallback instead of regenerating with a missing-program error. @@ -168,14 +185,15 @@ User-visible changes to Imp are recorded here. connection with no response, as `:ambiguous`, where a 5xx was retried. On resume, a checkpoint written by 0.5.0 is rewritten at schema version 2 and its `:transient_failure` requests become `:ambiguous`, since 0.5.0 recorded - timeouts and dispatcher crashes that way. Migration: check an `:ambiguous` + timeouts and dispatcher crashes that way, and a `schema_migration` event + records each. Migration: check an `:ambiguous` request with the provider before sending it again. - Breaking: an MCP call the server answers with HTTP 503 or 529 is `:refused`, like a 429, where it was `:unknown`: RFC 9110 defines 503 as the server being unable to handle the request, and providers answer overload - with 503 or 529. MCP and language-model calls read a status the same way. - Migration: code that treated a 503 or 529 as possibly run matches - `:refused` for them. + with 503 or 529. MCP calls and `Imp.Clients.ReqLLMBatch` read a status the + same way. Migration: code that treated a 503 or 529 as possibly run + matches `:refused` for them. - Breaking: `Imp.Datasets.csv/3` parses RFC 4180 CSV with NimbleCSV, a new dependency (`nimble_csv ~> 1.3`); it raised `FunctionClauseError` on any quoted field and could not read a quoted line break. A line break may be @@ -200,8 +218,12 @@ User-visible changes to Imp are recorded here. answer and scored on it. `Imp.evaluate/4`, `Imp.Evaluate.run/2`, `Imp.Experiment.Data.new/1` and every optimizer that runs a program on examples check each dataset before any model call, and the error names the - function, the dataset and the row. Migration: call `Imp.with_inputs/2` on - every example you evaluate or optimize on. + function, the dataset and the row. To check it, `Imp.Evaluate`, MIPROv2, + `Imp.Optimizer.InstructionSearch` and `Imp.Optimizer.SignatureOptimizer` + read a lazy dataset once and hold it in memory, so a one-shot stream is no + longer found empty on a second pass (every InstructionSearch candidate + scored 0.0 on one). Migration: call `Imp.with_inputs/2` on every example + you evaluate or optimize on. - Breaking: `Imp.Evaluate` and `Imp.Experiment.Data` refuse a row that is a plain map or a field pair list, which cannot declare its inputs; Evaluate turned it into an example whose labels reached the program. Migration: @@ -284,13 +306,18 @@ User-visible changes to Imp are recorded here. is recorded in its run like any other model call: a `:model_request` event with the request's `:purpose`, and a `:model_response` event with the usage and cost the provider reported. It recorded neither, so a streamed turn left - no model record, no cost and no ATIF model step. + no model record, no cost and no ATIF model step. A `stream/3` that raises + before returning a stream is recorded too and returns + `{:error, {:lm_failed, lm, error}}`, as a raising `generate/3` does. - A model turn that says something and calls tools keeps both. Through `Imp.req_llm/2` the text was dropped whenever the reply had tool calls, so a ReActV2 step's `next_thought` was empty; streamed with `provider_stream: true`, the text was kept but the tool calls were lost, all of them when text - arrived and all but the last otherwise. The adapter reads the text as it - reads a text reply, into `next_thought` for ReActV2, as DSPy does. So when a + arrived and all but the last otherwise. The raw LM output for such a reply + is `%{text: text, tool_calls: calls}` (`:text` only when it is not blank), + unstreamed and streamed. Keeping the text is what DSPy does; the adapter + then reads it as it reads a text reply, into `next_thought` for ReActV2, + where DSPy's Chat adapter leaves the field empty. So when a ReActV2 run reaches `max_iters` and its last reply has text beside tool calls it did not run, that text is now the answer, where the answer was `nil`; the calls are still listed as unexecuted. The text also appears in @@ -320,7 +347,7 @@ User-visible changes to Imp are recorded here. as for reasoning and response-format support. - `Imp.react` sends its tool roster in declared order, then `submit`, and a saved agent keeps that order. It was the order of the tool names' atoms, - which can differ between processes and changed the prompt a provider + which can differ between runs of the VM and changed the prompt a provider caches. - The JSON fallback after an unparseable Chat or XML reply sends the same request in JSON: the program's `adapter_opts` renderers (`:system_renderer`, @@ -335,7 +362,9 @@ User-visible changes to Imp are recorded here. rendering in its options, as `:default_system` and, for an `:output_renderer` that takes the options as a fourth argument, `:default_outputs`, so one that builds on the default builds on the format - of the request. + of the request. A renderer now shapes the JSON fallback and every JSON or + XML request; one that ignores its options sends a Chat-shaped prompt in the + fallback. - The `[:imp, :adapter, :parse, :json_fallback]` event names the adapter whose reply failed as `:adapter` (it always said `Imp.Adapter.Chat`, also for XML) and the adapter that retried as `:fallback_adapter`. @@ -359,8 +388,9 @@ User-visible changes to Imp are recorded here. `Imp.Deadline` in force, the request is not retried in that run: it stays `:transient_failure`, the summary is not `complete?`, and `resume/3` retries it. The checkpoint keeps the time each such request may be sent - again (`not_before`, UTC), and `resume/3` waits for it under the same rules - or stops the retry again without sending. A process that traps exits and is + again (`not_before`, UTC) and records a `retry_stopped` event, and + `resume/3` waits for it under the same rules or stops the retry again + without sending. A process that traps exits and is stopped by its parent during the wait exits at once with the parent's reason; a dispatch wave in progress is still waited for, up to `:timeout`, and a process started with bare `spawn` has no parent to listen for and diff --git a/RELEASE_NOTES.md b/RELEASE_NOTES.md index 78aed10b..bb7f61d4 100644 --- a/RELEASE_NOTES.md +++ b/RELEASE_NOTES.md @@ -53,43 +53,56 @@ tree calls, and no cowlib release fixes it yet. no output as a parse error; and evaluation and optimizers refuse examples that never declared their inputs, whose labels were passed to the program as inputs. -- Redaction runs before a term is converted, in reports, results, +- Redaction runs before a term is converted, in reports, results, GRPO checkpoints, session records and saved programs, so a client, retriever or - OAuth store in them is no longer written with its secrets. Redaction still - replaces the whole string, as in 0.5.0, and recognizes more credential - shapes. + OAuth store in them is no longer written with its secrets. The SIMBA, + MIPROv2, InferRules, random-search and GEPA checkpoints redact their failure + reasons. Redaction still replaces the whole string, as in 0.5.0, and + recognizes more credential shapes. +- Renderers (`:system_renderer`, `:output_renderer`) shape the JSON fallback + and every JSON or XML request, so a fallback sends the request the host + shaped. - `Imp.Clients.ReqLLMBatch` never sends again a request that may have run, and waits for the provider's `retry-after` before a retry. - A streamed call is recorded in its run like any other, and a reply with text and tool calls keeps all of them, streamed or not. - GEPA checkpoints resume: from a pending proposal batch, from a program that is an agent, and in a fresh VM. -- The lock file takes `mint` 1.11.0, which fixes three advisories - (EEF-CVE-2026-91043, EEF-CVE-2026-92103, EEF-CVE-2026-94194). +- Imp's lock file takes `mint` 1.11.0, which fixes three advisories + (EEF-CVE-2026-91043, EEF-CVE-2026-92103, EEF-CVE-2026-94194). That lock + does not reach your application, and Imp sets no `mint` floor: update your + own lock (Upgrading, step 1). ## Upgrading from 0.5 -1. Change the dependency to `{:imp, "~> 0.6"}`, run `mix deps.get`, and commit - `mix.lock`. `nimble_csv` is a new dependency. +1. Change the dependency to `{:imp, "~> 0.6"}`, run `mix deps.get` and + `mix deps.update mint hpax`, and commit `mix.lock`. `nimble_csv` is a new + dependency. 2. `Imp.collect/3` returns `{:ok, prediction}` or `{:error, reason}`; read fields with `Imp.get(prediction, :answer)` instead of matching a string. 3. Where you checked `Imp.Prediction.complete?/1` after a ReActV2 model failure, match `{:error, %Imp.Predict.ReActV2.StepError{reason: reason, history: history}}` and store `history` as you store a finished turn's. A - step refused by an `Imp.OperationalSafetyError` also ends the turn this - way, at once. + step refused by an `Imp.OperationalSafetyError` ends the turn this way at + once; where you call the program yourself, match + `{:error, %Imp.Predict.ReActV2.StepError{reason: + %Imp.OperationalSafetyError{}}}`. `Imp.Evaluate` and the optimizers + already raise it. 4. Call `Imp.with_inputs/2` on every example you evaluate or optimize on, and build evaluation rows with `Imp.example/1 |> Imp.with_inputs(...)` rather than plain maps or pair lists. Give each field of an example or prediction once, under one spelling. To keep a signature's instructions, pass it to `Imp.Signature.new/2` without instructions. -5. Rename a field called `tools` in an `Imp.react` task signature, and save - the program again; a saved program with one no longer loads. +5. A field called `tools` in an `Imp.react` task signature is refused, and a + saved program with one no longer loads. Rebuild the program with the field + renamed and save it; to keep a saved program's optimized instructions and + demos, rename the field in the saved file, or optimize again. 6. Match `%Imp.AdapterParseError{kind: :missing_fields}` where you relied on a prediction of defaults, or on ProgramOfThought's `:missing_program`, for a reply that answered no output. 7. Compare `Imp.Clients.ReqLLM` clients by `model`, not by the whole struct, - which now carries `:tool_calling`. + which now carries `:tool_calling`; `%Imp.Predict.ReActV2{}` likewise + carries `:tool_order`. 8. Treat a `ReqLLMBatch` request that ends `:ambiguous` as possibly run, and check it with the provider before sending it again. A 0.5.0 checkpoint's `:transient_failure` requests become `:ambiguous` on resume. @@ -103,31 +116,40 @@ tree calls, and no cowlib release fixes it yet. stoppers built on `consecutive_outcome/2`, which now counts an iteration that raised as `:proposal_error` instead of `:none`. 12. Keep 0.5.0 away from files 0.6.0 writes: it cannot read a trajectory with - atom keys (GEPA and Playbook checkpoints, `Imp.dump/1`), an optimizer - report that holds an `Imp.History`, or a `ReqLLMBatch` checkpoint. + atom keys (in a GEPA or Playbook checkpoint, or saved with `Imp.dump/1`), + an optimizer report that holds an `Imp.History`, or a `ReqLLMBatch` + checkpoint. 13. A ReActV2 turn that reaches `max_iters` with text beside tool calls it did not run now answers with that text, where it answered `nil`; the calls are still listed as unexecuted. A host that publishes every non-empty answer should decide whether to publish such a turn's text. 14. Re-evaluate saved agents on held-out data: the step prompt and tool roster order changed, so they send different prompt text. +15. A `:system_renderer` or `:output_renderer` now also shapes the JSON + fallback and JSON and XML requests. Build on `opts[:default_system]` and + `opts[:default_outputs]` rather than ignoring the options, or the fallback + sends a Chat-shaped prompt. ## Known limits -- A flat `["api_key", key]` held directly as a value, under a key that is not - itself a credential name, is not redacted, where 0.5.0 redacted it. A - two-element list with a name first is read as a key and value only as an - element of a list, so that lists of names such as - `with_inputs([:api_key, :question])` survive saving. -- A GEPA checkpoint redacts its failure reasons but otherwise holds the resume - state as it is: candidates, the evaluation cache, proposed instructions and - reflection data. Fast-Slow and Playbook checkpoints likewise hold their - prompts, rollouts and playbooks. Treat checkpoint files as sensitive. +- A flat two-element name-first list that is not itself an element of a + list (at the top level, in a tuple, or under a key that is not a credential + name), such as `["password", "hunter2"]`, is not redacted, where 0.5.0 + redacted it. Such a list is read as a key and value only as an element of a + list, so that lists of names such as `with_inputs([:api_key, :question])` + survive saving. +- The GEPA, SIMBA, MIPROv2, InferRules and random-search checkpoints redact + their failure reasons but otherwise hold the resume state as it is: GEPA's + candidates, evaluation cache, proposed instructions and reflection data, and + the others' instructions, demos and scores. Fast-Slow and Playbook + checkpoints likewise hold their prompts, rollouts and playbooks. Treat + checkpoint files as sensitive. - A ReActV2 step whose prose quotes its own field names (a JSON object with `next_thought` or `tool_calls` keys, a `[[ ## field ## ]]` line, a `` tag) is not read as prose and goes to the JSON fallback, which usually costs one more call. If the model repeats the same reply, the - step ends `:incomplete`. + step ends `:incomplete`. Beside native tool calls, such a thought becomes + `nil`; the calls still run. - When a streamed ReqLLM request fails, ReqLLM's own warning log prints the error with its response headers. Imp strips them from the error it returns, but cannot change that log line. From d76506035b252b731d74d5db95a3c6d4ffbf2e4d Mon Sep 17 00:00:00 2001 From: deepfates Date: Mon, 28 Sep 2026 07:36:41 -0700 Subject: [PATCH 4/8] Changelog: signature fields named nil, true or false --- CHANGELOG.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 53e56752..da8e102a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -426,6 +426,10 @@ User-visible changes to Imp are recorded here. - `Imp.Evaluate.Result.save_as_csv/2` writes with the same CSV module that `Imp.Datasets.csv/3` reads with, so what it writes reads back unchanged. The bytes it writes are the same as before. +- A signature field written as `nil`, `true` or `false` keeps that name as + text, so `"nil: string -> a"` saves and loads with a field named `"nil"`. + The name became the atom `nil`, which saved as `""`. A field given the atom + `nil`, `true` or `false` as its name raises `ArgumentError`. - An instruction an optimizer sets on `Imp.Predict.ProgramOfThought` or `Imp.Predict.CodeAct` reaches the extraction step, which kept the old instructions when GEPA, MIPROv2, COPRO, SIMBA or InferRules set it, and such From d4ec28a00e336188d1ecb405f6ab76f2e6612a72 Mon Sep 17 00:00:00 2001 From: deepfates Date: Mon, 28 Sep 2026 09:28:32 -0700 Subject: [PATCH 5/8] Declare mint ~> 1.11 (Imp matches Mint's error structs); derive the version in the deployment test --- CHANGELOG.md | 11 ++++++----- RELEASE_NOTES.md | 13 ++++++------- mix.exs | 6 ++++++ test/deployment_reference_test.exs | 3 ++- 4 files changed, 20 insertions(+), 13 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index da8e102a..5d8f8233 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,11 +6,12 @@ User-visible changes to Imp are recorded here. ### Security -- The lock file takes `mint` 1.11.0 (and `hpax` 1.1.0), which fixes - EEF-CVE-2026-91043, EEF-CVE-2026-92103 and EEF-CVE-2026-94194 in the HTTP - client Req uses. The example projects' lock files take the same versions. - Imp's lock file does not bind an application that depends on Imp, and Imp - sets no `mint` floor: run `mix deps.update mint hpax` in your own project. +- Imp depends on `mint` `~> 1.11` directly, which fixes EEF-CVE-2026-91043, + EEF-CVE-2026-92103 and EEF-CVE-2026-94194 in the HTTP client Req uses. Imp + matches Mint's error structs to tell a request that was never sent from one + that may have run, and did not declare the dependency. `mix deps.get` in an + application that locked `mint` 1.10.1 moves it to 1.11.0. Imp's lock file + and the example projects' take `mint` 1.11.0 and `hpax` 1.1.0. - `Imp.Redaction` redacts a string that holds a PEM private key of any type, a PGP private key block or a PuTTY key file; a Stripe (`sk_live_`, `rk_live_`), GitHub, GitLab (`glpat-`), Hugging Face (`hf_`), Slack diff --git a/RELEASE_NOTES.md b/RELEASE_NOTES.md index bb7f61d4..7059f3ab 100644 --- a/RELEASE_NOTES.md +++ b/RELEASE_NOTES.md @@ -68,16 +68,15 @@ tree calls, and no cowlib release fixes it yet. and tool calls keeps all of them, streamed or not. - GEPA checkpoints resume: from a pending proposal batch, from a program that is an agent, and in a fresh VM. -- Imp's lock file takes `mint` 1.11.0, which fixes three advisories - (EEF-CVE-2026-91043, EEF-CVE-2026-92103, EEF-CVE-2026-94194). That lock - does not reach your application, and Imp sets no `mint` floor: update your - own lock (Upgrading, step 1). +- Imp requires `mint` 1.11, which fixes three advisories + (EEF-CVE-2026-91043, EEF-CVE-2026-92103, EEF-CVE-2026-94194). Imp uses + Mint's error structs directly and now declares the dependency, so + `mix deps.get` moves an application's locked `mint` to 1.11. ## Upgrading from 0.5 -1. Change the dependency to `{:imp, "~> 0.6"}`, run `mix deps.get` and - `mix deps.update mint hpax`, and commit `mix.lock`. `nimble_csv` is a new - dependency. +1. Change the dependency to `{:imp, "~> 0.6"}`, run `mix deps.get`, and + commit `mix.lock`. `nimble_csv` and `mint` are new direct dependencies. 2. `Imp.collect/3` returns `{:ok, prediction}` or `{:error, reason}`; read fields with `Imp.get(prediction, :answer)` instead of matching a string. 3. Where you checked `Imp.Prediction.complete?/1` after a ReActV2 model diff --git a/mix.exs b/mix.exs index efb33ce9..14c5b242 100644 --- a/mix.exs +++ b/mix.exs @@ -128,6 +128,12 @@ defmodule Imp.MixProject do {:jason, "~> 1.4"}, {:jaxon, "~> 2.0.8"}, {:jsv, "~> 0.21"}, + # Imp matches Mint's error structs (Mint.TransportError, Mint.HTTPError) + # to tell a request that was never sent from one that may have run + # (Imp.Clients.ReqLLM, Imp.MCP.CallFailure), so it depends on Mint + # directly. 1.11.0 is the first release without EEF-CVE-2026-91043, + # EEF-CVE-2026-92103 and EEF-CVE-2026-94194. + {:mint, "~> 1.11"}, {:nimble_csv, "~> 1.3"}, {:nimble_options, "~> 1.1"}, {:req, "~> 0.6"}, diff --git a/test/deployment_reference_test.exs b/test/deployment_reference_test.exs index 97db1e01..872d1a7f 100644 --- a/test/deployment_reference_test.exs +++ b/test/deployment_reference_test.exs @@ -455,7 +455,8 @@ defmodule DeploymentReferenceTest do readme = File.read!(Path.join(@example_root, "README.md")) assert mix_file =~ ~s(elixir: "~> 1.19") - assert mix_file =~ ~s({:imp, "~> 0.5"}) + [major, minor | _patch] = String.split(Mix.Project.config()[:version], ".") + assert mix_file =~ ~s({:imp, "~> #{major}.#{minor}"}) assert mix_file =~ "IMP_PATH" assert readme =~ "bounded supervised task" assert readme =~ "IMP_MODEL" From c8664191d033bbb7160ad07ac8360294edb0cf75 Mon Sep 17 00:00:00 2001 From: deepfates Date: Mon, 28 Sep 2026 09:32:14 -0700 Subject: [PATCH 6/8] decisions: Imp follows semantic versioning; 0.5.0 is published --- decisions.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/decisions.md b/decisions.md index a54f88d8..08d27a7c 100644 --- a/decisions.md +++ b/decisions.md @@ -8,6 +8,7 @@ not necessarily when it was made. | Date | Decision | Source and reason | Status | Retires when | | --- | --- | --- | --- | --- | +| 2026-09-28 | Imp follows semantic versioning. Before 1.0, a release that changes what a caller receives or can rely on bumps the minor version (`0.5` to `0.6`), and a release of fixes that change nothing a caller relies on bumps the patch; `{:imp, "~> 0.x"}` then never takes a breaking release unasked. | Owner, 2026-09-28. `CHANGELOG.md` marks each breaking change "Breaking:" with its migration, and `RELEASE_NOTES.md` names them. | In force. | Does not retire. | | 2026-09-26 | `Imp.Optimizer.GEPA` with no `:execution_profile` runs DSPy's GEPA (`:gepa_v0_1_4_merge`, or `:gepa_v0_1_4` with `use_merge: false`), and its reflection records are DSPy's; Imp's own search is `execution_profile: :beam_native`, chosen explicitly. | Owner, 2026-09-26. An upstream name promises upstream semantics; the published GEPA results (`research/RESULTS.md` R3, R5) used the pinned profile, as `:gepa_v0_1_4`; and no reason for the BEAM-native default was ever recorded. `lib/imp/optimizer/gepa.ex` moduledoc, `test/gepa_agent_reflection_test.exs`. | In force. | Does not retire. | | 2026-09-25 | Prompts and parsed values use neutral spellings: types in words (`string`, `integer`, `true or false`, `one of: a, b`, `list of strings`, `object`, declared once in `Imp.Adapter.FieldType`), values in JSON (`null`, `true`, `["x", "y"]`), in every adapter, ReAct, and the MIPROv2, GEPA and SIMBA prompts. A non-string answer for a string field is kept as its JSON text; a null is no value. Parity with DSPy means the same fields, order and constraints in the prompt, the same fields and types accepted, and the same errors raised, not DSPy's text or Python's `str()` spelling. | The owner's intention, 2026-09-25: Python spellings (`Literal[...]`, `True`, `None`, `repr(Example)`) read as foreign in an Elixir library and teach the model nothing the neutral words do not. `test/adapter_type_wording_test.exs`, `test/upstream_exam/adapters_test.exs`; the golden trace and the MIPROv2 proposer differentials compare prompts after putting DSPy's spellings into Imp's words (`Imp.DSPyWording`). | In force. | Does not retire. | | 2026-09-25 | A tool call's value alone never reads as `:refused`: `Imp.Tool.outcome/1` says `:refused` only for an `Imp.MCP.CallFailure` or an MCP error result that declares it in `structuredContent.outcome` (`"refused"`, `"auth_refused"`, `"unknown"`), and ReActV2 and RLM decide their own refusals where they refuse the call and record them as `metadata.outcome`. | `lib/imp/tool.ex` `outcome/1`, `test/mcp_call_outcome_test.exs`. A tool can return any term, including one shaped like Imp's refusals, and `:refused` tells a caller that repeating the call is safe; reading it from a term a tool that ran can produce would retry a write that landed. The declared outcome is how a server that knows more (Kite, for a write that may have been applied) says so. | In force. | Does not retire. | @@ -30,7 +31,7 @@ not necessarily when it was made. | 2026-09-13 | Absorption renames: `IMP_ACP_PATH` becomes `IMP_PATH` (copied workspace example), `IMP_ACP_ROOT` becomes `IMP_ROOT` (host application preset), launcher is `examples/workspace_agent/scripts/workspace-agent-acp`. Module names `Imp.ACP`, `Imp.ACP.Host` and the `deepfates.com/imp-acp` wire metadata namespace keep their names. | Retired `imp_acp` README. Saved executable and environment references are updated explicitly, not by rewriting contact history. | Amended 2026-09-25: `Imp.ACP.MCP` is gone before 0.5.0, the first Hex release; it copied `Imp.MCP.Import` to add one field, which `Imp.ACP.ToolKind.derive_all(import.annotations)` computes. | Does not retire. | | 2026-09-13 | The workspace agent's session store keeps the `imp_acp/workspace_agent/sessions` directory name. | `examples/workspace_agent/README.md`. Preserves saved sessions; it is a storage location, not a dependency on the retired package. | In force. | Saved sessions are migrated, or the owner accepts losing restart continuity for them. | | 2026-09-13 | Ordinary Imp boot starts no protocol listener, subprocess or remote connection and does not start the ExMCP application. Protocol entry points (`Imp.ACP.*`, non-empty `Imp.MCP.connect/2`) start it explicitly. Releases using them declare `applications: [ex_mcp: :load, erlexec: :load]`; erlexec starts on the first stdio MCP connection. | `mix.exs` dependency comment, `docs/production.md` "Protocol runtime in releases", commit `728f8c77`. Prediction and optimizer processes must not open listeners or acquire protocol boot output. | In force. | Does not retire while Imp is a library inside a host application. | -| 2026-09-23 | Imp is a Hex package, installed with `{:imp, "~> 0.5"}`, and every dependency is a Hex package. ExMCP is `~> 1.5`, unpatched. What Imp needed from the `deepfates/ex_mcp` fork lives in Imp, each piece with its reason and the condition that retires it beside it: `Imp.MCP.OwnedStdio` (server process groups, the child `PATH`), `Imp.MCP.Trust` (trust for authorized remote origins), the HTTP client options in `Imp.MCP.Connections`, and `Imp.MCP.OAuth.Flow`. Publishing a version is the owner's action. | Owner ruling, 2026-09-24: ExMCP from Hex, with what can live in our code moved there rather than carried as patches. `mix hex.build` and `mix package.check` build and consume the package; `test/mcp_stdio_lifecycle_test.exs`, `test/mcp_stdio_environment_test.exs`, `test/mcp_trust_test.exs`, `test/mcp_http_public_server_test.exs` and `test/mcp_oauth_test.exs` fail without each piece. | In force; `0.5.0` is built but not published. | Does not retire. | +| 2026-09-23 | Imp is a Hex package, installed with `{:imp, "~> 0.5"}`, and every dependency is a Hex package. ExMCP is `~> 1.5`, unpatched. What Imp needed from the `deepfates/ex_mcp` fork lives in Imp, each piece with its reason and the condition that retires it beside it: `Imp.MCP.OwnedStdio` (server process groups, the child `PATH`), `Imp.MCP.Trust` (trust for authorized remote origins), the HTTP client options in `Imp.MCP.Connections`, and `Imp.MCP.OAuth.Flow`. Publishing a version is the owner's action. | Owner ruling, 2026-09-24: ExMCP from Hex, with what can live in our code moved there rather than carried as patches. `mix hex.build` and `mix package.check` build and consume the package; `test/mcp_stdio_lifecycle_test.exs`, `test/mcp_stdio_environment_test.exs`, `test/mcp_trust_test.exs`, `test/mcp_http_public_server_test.exs` and `test/mcp_oauth_test.exs` fail without each piece. | In force; `0.5.0` is published on Hex. | Does not retire. | | 2026-09-24 | An `Imp.Tool` keeps the type of the name it was given, and a tool imported from an MCP server is named by the server's string. No string is turned into an atom. | `lib/imp/tool.ex` `new/4`, `test/tool_schema_runtime_test.exs`. Converting a string to an atom when one happened to be loaded made a name's type depend on what else ran in the VM, so the same server's tools came back as a mix of atoms and strings; lookups and tool policies already compare names by text. | In force. | Does not retire. | | 2026-09-13 | `mix.exs` resolves ExMCP three ways: a bundled `vendor/ex_mcp` if present, else `EX_MCP_PATH`, else the GitHub pin. | `mix.exs`. No reason was recorded beside the conditional. | Retired 2026-09-23: ExMCP comes from Hex, one way. | Retired. | | 2026-09-13 | Known interoperability limit, not a feature: when an HTTP MCP server selects protocol version `2025-03-26`, ExMCP's `notifications/initialized` can still carry the client's `2025-11-25` header, and a strict older server may reject the session. Compatibility with strict older HTTP servers is not claimed. | `docs/production.md`, commit `902a5546`. ExMCP 1.5.0's connection manager sends that notification before it settles the transport's version (`establish_legacy_protocol/4`). | Open defect, documented. | ExMCP settles the HTTP version before sending that notification and a test against a strict older server passes. | From feba21a0dcbb5f4a736377c0d53b968c464485ab Mon Sep 17 00:00:00 2001 From: deepfates Date: Mon, 28 Sep 2026 11:28:55 -0700 Subject: [PATCH 7/8] Hold mint 1.10.1: mint 1.11.0 with Finch 0.23.0 reuses timed-out HTTP/1 connections mint 1.11.0 leaves an HTTP/1 connection open after a receive timeout and Finch 0.23.0 returns it to its pool, so the next request on it waits behind the unanswered one and times out. Declare mint ~> 1.10, lock 1.10.1 and hpax 1.0.4 here and in both examples, ignore the three mint advisories with their reasons until a Finch release includes sneako/finch#397, add a regression test that fails on 1.11.0 and passes on 1.10.1, and say so in the CHANGELOG and release notes. --- .audit_ignore | 51 +++++++++++++++++++ CHANGELOG.md | 28 +++++++---- RELEASE_NOTES.md | 29 +++++++++-- examples/deployment/mix.lock | 4 +- examples/workspace_agent/mix.lock | 4 +- mix.exs | 9 ++-- mix.lock | 4 +- test/timed_out_connection_test.exs | 80 ++++++++++++++++++++++++++++++ 8 files changed, 185 insertions(+), 24 deletions(-) create mode 100644 test/timed_out_connection_test.exs diff --git a/.audit_ignore b/.audit_ignore index 5d772a43..baef8735 100644 --- a/.audit_ignore +++ b/.audit_ignore @@ -58,3 +58,54 @@ EEF-CVE-2026-43966 # RETIRE the moment cowlib publishes a release that validates cookie/1, and # move the lock to that release. EEF-CVE-2026-43969 + +# mint advisories. Verified 2026-09-28 against the ERLEF CNA records that +# hex.audit serves (https://api.osv.dev/v1/vulns/) and the mint 1.10.1 and +# 1.11.0 sources. +# +# Shared facts. mint 1.11.0 fixes all three, and mix.lock holds 1.10.1 anyway. +# mint 1.11.0 no longer closes an HTTP/1 connection on a receive timeout (the +# `{:error, %Mint.TransportError{reason: :timeout}}` clause of +# Mint.HTTP1.recv/3), and Finch 0.23.0, the newest release, returns any open +# connection to its pool (transfer_if_open in lib/finch/http1/pool.ex). The +# next request on that connection is written behind the unanswered one and +# times out, and so does a retry that lands there; if the server answers late, +# a later request on it raises CaseClauseError inside Finch. On 1.10.1 the +# retry opens a new connection. test/timed_out_connection_test.exs fails on +# 1.11.0 with Finch 0.23.0 and passes on 1.10.1. Req and ReqLLM speak HTTP/1 +# unless a caller configures otherwise (Req's default protocols are [:http1]; +# ReqLLM's @default_stream_pool_protocols is [:http1]), and Imp opens HTTP/2 +# only when a caller asks, through the benchmark parity task's +# --req-llm-pool-protocols. +# RETIRE all three together when a Finch release closes a connection that +# still has a request in flight, such as one that includes +# https://github.com/sneako/finch/pull/397 ("Close abandoned HTTP/1 +# connections after request errors", open and unreleased on 2026-09-28), or +# when a mint release closes the connection on a receive timeout again. Then +# move the lock to that Finch release and mint 1.11 in the same change, and +# the timed-out-connection test must still pass. + +# EEF-CVE-2026-91043 / CVE-2026-91043 / GHSA-9x8p-qrf4-jq7g (HIGH) - HPACK +# indexed cookie fields in a Mint HTTP/2 response bypass max_header_list_size +# and exhaust client memory. HTTP/2 only (Mint.HTTP2): not reachable through +# Imp's HTTP/1 default. Reachable in an application that configures HTTP/2 +# for Req or ReqLLM against a malicious server. Retire as above. +EEF-CVE-2026-91043 + +# EEF-CVE-2026-92103 / CVE-2026-92103 / GHSA-q95c-ccq6-j5j6 (MEDIUM) - the +# Mint HTTP/2 client buffers a frame up to 16 MiB before enforcing +# max_frame_size. HTTP/2 only (Mint.HTTP2.Frame): not reachable through Imp's +# HTTP/1 default. Retire as above. +EEF-CVE-2026-92103 + +# EEF-CVE-2026-94194 / CVE-2026-94194 / GHSA-gvrc-75rc-7gj9 (MEDIUM) - the +# Mint HTTP/1 client applies chunked framing when chunked is not the final +# transfer coding, and keeps an HTTP/1.0 connection open after a response with +# Transfer-Encoding. This one is HTTP/1 and is reachable: a malicious server +# behind an intermediary that follows RFC 9112 can desynchronize the +# intermediary and Mint on a pooled connection and poison the responses to +# later requests. It needs both the malicious server and such an intermediary +# between it and Imp. It is ignored because the fix comes only with mint +# 1.11.0, whose timeout behaviour breaks every HTTP/1 client of Finch 0.23.0, +# not because it is a false positive. Retire as above. +EEF-CVE-2026-94194 diff --git a/CHANGELOG.md b/CHANGELOG.md index 1034bed2..e3092014 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,12 +6,16 @@ User-visible changes to Imp are recorded here. ### Security -- Imp depends on `mint` `~> 1.11` directly, which fixes EEF-CVE-2026-91043, - EEF-CVE-2026-92103 and EEF-CVE-2026-94194 in the HTTP client Req uses. Imp - matches Mint's error structs to tell a request that was never sent from one - that may have run, and did not declare the dependency. `mix deps.get` in an - application that locked `mint` 1.10.1 moves it to 1.11.0. Imp's lock file - and the example projects' take `mint` 1.11.0 and `hpax` 1.1.0. +- Imp's lock file and the example projects' keep `mint` 1.10.1, which has + three advisories that `mint` 1.11.0 fixes: EEF-CVE-2026-91043 and + EEF-CVE-2026-92103 are in Mint's HTTP/2 client, and EEF-CVE-2026-94194 is + in HTTP/1 chunked framing and needs an intermediary between the client and + a malicious server. `mint` 1.11.0 leaves an HTTP/1 connection open after a + receive timeout, and Finch 0.23.0 reuses it, so requests after a timeout + fail (see Known limits in the release notes). `.audit_ignore` lists the + three advisories with the reason and the condition that removes them: a + Finch release that includes https://github.com/sneako/finch/pull/397, + which is open and unreleased. - `Imp.Redaction` redacts a string that holds a PEM private key of any type, a PGP private key block or a PuTTY key file; a Stripe (`sk_live_`, `rk_live_`), GitHub, GitLab (`glpat-`), Hugging Face (`hf_`), Slack @@ -290,11 +294,13 @@ User-visible changes to Imp are recorded here. dependency (`Imp.Core` reads a reported cost given as a `Decimal`). It declares `plug` and `plug_cowboy` with `runtime: false` for the demo MCP servers, which the package leaves out: Imp starts neither, and no Imp code - that ships uses them. It declares `mint` `~> 1.11` (see Security). Of these, - only `mint` moves an existing lock: an application that locked `mint` below - 1.11.0 moves to 1.11.0 on `mix deps.get`. The other requirements are ones - Req and ExMCP already set (ExMCP requires `decimal` `~> 3.0`), so an - application resolves no new package and keeps the versions it has. + that ships uses them. It declares `mint` `~> 1.10`, since it matches Mint's + error structs to tell a request that was never sent from one that may have + run. The other requirements are ones Req and ExMCP already set (ExMCP + requires `decimal` `~> 3.0`), so they move no lock. The `mint` requirement + moves only a lock below 1.10: `mix deps.get` then takes the newest release, + 1.11.0, unless the application asks for `{:mint, "~> 1.10.1"}` (see + Security). - Imp no longer depends on `jsv`, which nothing in Imp uses. ReqLLM still requires it. diff --git a/RELEASE_NOTES.md b/RELEASE_NOTES.md index 7059f3ab..2c6d38ec 100644 --- a/RELEASE_NOTES.md +++ b/RELEASE_NOTES.md @@ -40,6 +40,11 @@ containing CR or LF, and a fresh `mix deps.get` resolves Cowboy 2.19.0. The seco for an outgoing `Cookie` request header, which nothing in Imp's dependency tree calls, and no cowlib release fixes it yet. +Imp declares `mint` `~> 1.10`, and a fresh `mix deps.get` resolves `mint` +1.11.0. That release fixes three advisories but reuses HTTP/1 connections +that timed out; Imp's own lock holds 1.10.1. See Known limits for the choice +an application has. + ## Headline changes - An agent's step prompt asks for one format, as DSPy's does. With an LM that @@ -68,10 +73,6 @@ tree calls, and no cowlib release fixes it yet. and tool calls keeps all of them, streamed or not. - GEPA checkpoints resume: from a pending proposal batch, from a program that is an agent, and in a fresh VM. -- Imp requires `mint` 1.11, which fixes three advisories - (EEF-CVE-2026-91043, EEF-CVE-2026-92103, EEF-CVE-2026-94194). Imp uses - Mint's error structs directly and now declares the dependency, so - `mix deps.get` moves an application's locked `mint` to 1.11. ## Upgrading from 0.5 @@ -131,6 +132,26 @@ tree calls, and no cowlib release fixes it yet. ## Known limits +- `mint` 1.11.0 leaves an HTTP/1 connection open after a receive timeout, + and Finch 0.23.0, the newest release, returns it to its pool with the + unanswered request still on it. A later request the pool gives that + connection waits behind the unanswered one and times out, and so does a + retry that lands there, until the server answers the first request or + closes the connection. If the late answer arrives while another request is + waiting, Finch raises `CaseClauseError`. This affects every Req or Finch user on HTTP/1, the + default for Req and ReqLLM, whose server can time out; ReqLLM spreads a + host's requests over several connections, so there only the requests that + draw the stuck one fail. With `mint` 1.10.1 a timeout closes the + connection and the next request opens a new one, so Imp's lock holds + 1.10.1. It has three advisories that 1.11.0 fixes: EEF-CVE-2026-91043 + (high) and EEF-CVE-2026-92103 are in Mint's HTTP/2 client only, and + EEF-CVE-2026-94194 is in HTTP/1 chunked framing and needs a malicious + server behind an intermediary that reads the framing strictly. An + application chooses one: add `{:mint, "~> 1.10.1"}` to its dependencies to + lock 1.10.1 and keep those advisories, or take 1.11.0 and accept that a + connection that timed out is reused until the server closes it. Finch has + an open, unreleased fix (https://github.com/sneako/finch/pull/397); a Finch + release that includes it ends the choice. - A flat two-element name-first list that is not itself an element of a list (at the top level, in a tuple, or under a key that is not a credential name), such as `["password", "hunter2"]`, is not redacted, where 0.5.0 diff --git a/examples/deployment/mix.lock b/examples/deployment/mix.lock index fbd6f1ef..4db647e2 100644 --- a/examples/deployment/mix.lock +++ b/examples/deployment/mix.lock @@ -11,7 +11,7 @@ "ex_json_schema": {:hex, :ex_json_schema, "0.11.5", "ea45f3238be135949dbbbcc9e8eb4682d6e561b8a66374e75d85a2e1d2bd4107", [:mix], [{:decimal, "~> 3.0", [hex: :decimal, repo: "hexpm", optional: false]}], "hexpm", "61ed2a8f07bd115e7ab6d45c147642a8c73b962bc419fadbb248046b9d3d0f20"}, "ex_mcp": {:hex, :ex_mcp, "1.5.0", "bf0a6862b306d4ba76c29848db110b49d1cbe83b304c9d1ba2ea6b09a7f83b9a", [:mix], [{:castore, "~> 1.0", [hex: :castore, repo: "hexpm", optional: false]}, {:ex_json_schema, "~> 0.10", [hex: :ex_json_schema, repo: "hexpm", optional: false]}, {:fuse, "~> 2.4", [hex: :fuse, repo: "hexpm", optional: true]}, {:jason, "~> 1.4", [hex: :jason, repo: "hexpm", optional: false]}, {:jose, "~> 1.11", [hex: :jose, repo: "hexpm", optional: false]}, {:mint, "~> 1.6", [hex: :mint, repo: "hexpm", optional: false]}, {:mint_web_socket, "~> 1.0", [hex: :mint_web_socket, repo: "hexpm", optional: false]}, {:plug, "~> 1.16", [hex: :plug, repo: "hexpm", optional: false]}, {:plug_cowboy, "~> 2.7", [hex: :plug_cowboy, repo: "hexpm", optional: false]}, {:telemetry, "~> 1.2", [hex: :telemetry, repo: "hexpm", optional: false]}], "hexpm", "3cb3cd60e2275277f519ec6db02cf5e92b342d88c5ba0d31e31bca498b28e4fb"}, "finch": {:hex, :finch, "0.23.0", "e3f9287ac25a8832f848b144c2b57346aac65b205e2e0629a52adfe6507fd837", [:mix], [{:mime, "~> 1.0 or ~> 2.0", [hex: :mime, repo: "hexpm", optional: false]}, {:mint, "~> 1.8", [hex: :mint, repo: "hexpm", optional: false]}, {:nimble_options, "~> 0.4 or ~> 1.0", [hex: :nimble_options, repo: "hexpm", optional: false]}, {:nimble_pool, "~> 1.1", [hex: :nimble_pool, repo: "hexpm", optional: false]}, {:telemetry, "~> 0.4 or ~> 1.0", [hex: :telemetry, repo: "hexpm", optional: false]}], "hexpm", "80e58d3f936f57e3fdf404f83a3642897ae6d9fb642934e46da4d8fe761b99d5"}, - "hpax": {:hex, :hpax, "1.1.0", "782931867cc23217c68fb5f68fe1a11f5e7544c7fda82c8a7019a5df5a4a1cdf", [:mix], [], "hexpm", "0b8d0f05832f55571d65ac720f79bf8994138ffbb133209dc4685eae0ad456a8"}, + "hpax": {:hex, :hpax, "1.0.4", "777de5d433b0fbdc7c418159c8055910faa8047ffdb3d6b31098d2a46cd7685c", [:mix], [], "hexpm", "afc7cb142ebcc2d01ce7816190b98ce5dd49e799111b24249f3443d730f377ca"}, "idna": {:hex, :idna, "7.1.0", "1067a13043538129602d2f2ce6899d8713125c7d19734aa557ce2e3ea55bd4f1", [:rebar3], [], "hexpm", "6ae959a025bf36df61a8cab8508d9654891b5426a84c44d82deaffd6ddf8c71f"}, "jason": {:hex, :jason, "1.4.5", "2e3a008590b0b8d7388c20293e9dcc9cf3e5d642fd2a114e4cbbb52e595d940a", [:mix], [{:decimal, "~> 1.0 or ~> 2.0 or ~> 3.0", [hex: :decimal, repo: "hexpm", optional: true]}], "hexpm", "b0c823996102bcd0239b3c2444eb00409b72f6a140c1950bc8b457d836b30684"}, "jaxon": {:hex, :jaxon, "2.0.8", "00951a79d354260e28d7e36f956c3de94818124768a4b22e0fc55559d1b3bfe7", [:make, :mix], [{:elixir_make, "~> 0.4", [hex: :elixir_make, repo: "hexpm", optional: false]}], "hexpm", "74532853b1126609615ea98f0ceb5009e70465ca98027afbbd8ed314d887e82d"}, @@ -19,7 +19,7 @@ "jsv": {:hex, :jsv, "0.24.0", "71b521244b51e1849ac7cae986ce391ae13f6ffb781235ad046a4f7ebbb9472b", [:mix], [{:abnf_parsec, "~> 2.0", [hex: :abnf_parsec, repo: "hexpm", optional: false]}, {:decimal, "~> 2.0 or ~> 3.0", [hex: :decimal, repo: "hexpm", optional: true]}, {:idna, "~> 6.0 or ~> 7.0", [hex: :idna, repo: "hexpm", optional: false]}, {:jason, "~> 1.0", [hex: :jason, repo: "hexpm", optional: true]}, {:texture, ">= 1.2.1", [hex: :texture, repo: "hexpm", optional: false]}], "hexpm", "a9829510d25fe6e16a84600ba8d2f7c3da5234401b796aa87741393b1c85aa27"}, "llm_db": {:hex, :llm_db, "2026.9.5", "357a594559c194f65616787961badde1256d5078992639b771437ef3e039d1e5", [:mix], [{:dotenvy, "~> 1.1", [hex: :dotenvy, repo: "hexpm", optional: false]}, {:igniter, "~> 0.7", [hex: :igniter, repo: "hexpm", optional: true]}, {:jason, "~> 1.4", [hex: :jason, repo: "hexpm", optional: false]}, {:req, "~> 0.5", [hex: :req, repo: "hexpm", optional: false]}, {:toml, "~> 0.7", [hex: :toml, repo: "hexpm", optional: false]}, {:zoi, "~> 0.10", [hex: :zoi, repo: "hexpm", optional: false]}], "hexpm", "8631c05b435ebedade83421d3d8b72b051af2caee822d953a21384cad95a9714"}, "mime": {:hex, :mime, "2.0.7", "b8d739037be7cd402aee1ba0306edfdef982687ee7e9859bee6198c1e7e2f128", [:mix], [], "hexpm", "6171188e399ee16023ffc5b76ce445eb6d9672e2e241d2df6050f3c771e80ccd"}, - "mint": {:hex, :mint, "1.11.0", "a713551624815c0435237b93d90ea8b9b14254690c66f732d0ef8930f76ff1d9", [:mix], [{:castore, "~> 0.1.0 or ~> 1.0", [hex: :castore, repo: "hexpm", optional: true]}, {:hpax, "~> 1.1", [hex: :hpax, repo: "hexpm", optional: false]}], "hexpm", "c6279ba2d6aa3a383a1d4cfbe7b59f42e6efd400f58d8e2acfeac48a438693ab"}, + "mint": {:hex, :mint, "1.10.1", "c53e70867cf74017716884d8d33e0742b08b32e9cdb0031cbc69a429dc5555e3", [:mix], [{:castore, "~> 0.1.0 or ~> 1.0", [hex: :castore, repo: "hexpm", optional: true]}, {:hpax, "~> 0.1.1 or ~> 0.2.0 or ~> 1.0", [hex: :hpax, repo: "hexpm", optional: false]}], "hexpm", "0ba2a904605ed8406393444fb8b3356dc58eb59ee6c7fb94ac3f015e1be129e8"}, "mint_web_socket": {:hex, :mint_web_socket, "1.0.6", "5ffcf350df5b90f2d7a04adf877165228804993714592512374218d4679e325a", [:mix], [{:mint, ">= 1.4.1 and < 2.0.0-0", [hex: :mint, repo: "hexpm", optional: false]}], "hexpm", "0c360e9012413f1c115a63532601eb5d63731aab7010949178769760686c1698"}, "nimble_csv": {:hex, :nimble_csv, "1.3.0", "b7f998dc62b222bce9596e46f028c7a5af04cb5dde6df2ea197c583227c54971", [:mix], [], "hexpm", "41ccdc18f7c8f8bb06e84164fc51635321e80d5a3b450761c4997d620925d619"}, "nimble_options": {:hex, :nimble_options, "1.1.1", "e3a492d54d85fc3fd7c5baf411d9d2852922f66e69476317787a7b2bb000a61b", [:mix], [], "hexpm", "821b2470ca9442c4b6984882fe9bb0389371b8ddec4d45a9504f00a66f650b44"}, diff --git a/examples/workspace_agent/mix.lock b/examples/workspace_agent/mix.lock index d72ea040..5d382494 100644 --- a/examples/workspace_agent/mix.lock +++ b/examples/workspace_agent/mix.lock @@ -11,7 +11,7 @@ "ex_json_schema": {:hex, :ex_json_schema, "0.11.5", "ea45f3238be135949dbbbcc9e8eb4682d6e561b8a66374e75d85a2e1d2bd4107", [:mix], [{:decimal, "~> 3.0", [hex: :decimal, repo: "hexpm", optional: false]}], "hexpm", "61ed2a8f07bd115e7ab6d45c147642a8c73b962bc419fadbb248046b9d3d0f20"}, "ex_mcp": {:hex, :ex_mcp, "1.5.0", "bf0a6862b306d4ba76c29848db110b49d1cbe83b304c9d1ba2ea6b09a7f83b9a", [:mix], [{:castore, "~> 1.0", [hex: :castore, repo: "hexpm", optional: false]}, {:ex_json_schema, "~> 0.10", [hex: :ex_json_schema, repo: "hexpm", optional: false]}, {:fuse, "~> 2.4", [hex: :fuse, repo: "hexpm", optional: true]}, {:jason, "~> 1.4", [hex: :jason, repo: "hexpm", optional: false]}, {:jose, "~> 1.11", [hex: :jose, repo: "hexpm", optional: false]}, {:mint, "~> 1.6", [hex: :mint, repo: "hexpm", optional: false]}, {:mint_web_socket, "~> 1.0", [hex: :mint_web_socket, repo: "hexpm", optional: false]}, {:plug, "~> 1.16", [hex: :plug, repo: "hexpm", optional: false]}, {:plug_cowboy, "~> 2.7", [hex: :plug_cowboy, repo: "hexpm", optional: false]}, {:telemetry, "~> 1.2", [hex: :telemetry, repo: "hexpm", optional: false]}], "hexpm", "3cb3cd60e2275277f519ec6db02cf5e92b342d88c5ba0d31e31bca498b28e4fb"}, "finch": {:hex, :finch, "0.23.0", "e3f9287ac25a8832f848b144c2b57346aac65b205e2e0629a52adfe6507fd837", [:mix], [{:mime, "~> 1.0 or ~> 2.0", [hex: :mime, repo: "hexpm", optional: false]}, {:mint, "~> 1.8", [hex: :mint, repo: "hexpm", optional: false]}, {:nimble_options, "~> 0.4 or ~> 1.0", [hex: :nimble_options, repo: "hexpm", optional: false]}, {:nimble_pool, "~> 1.1", [hex: :nimble_pool, repo: "hexpm", optional: false]}, {:telemetry, "~> 0.4 or ~> 1.0", [hex: :telemetry, repo: "hexpm", optional: false]}], "hexpm", "80e58d3f936f57e3fdf404f83a3642897ae6d9fb642934e46da4d8fe761b99d5"}, - "hpax": {:hex, :hpax, "1.1.0", "782931867cc23217c68fb5f68fe1a11f5e7544c7fda82c8a7019a5df5a4a1cdf", [:mix], [], "hexpm", "0b8d0f05832f55571d65ac720f79bf8994138ffbb133209dc4685eae0ad456a8"}, + "hpax": {:hex, :hpax, "1.0.4", "777de5d433b0fbdc7c418159c8055910faa8047ffdb3d6b31098d2a46cd7685c", [:mix], [], "hexpm", "afc7cb142ebcc2d01ce7816190b98ce5dd49e799111b24249f3443d730f377ca"}, "idna": {:hex, :idna, "7.1.0", "1067a13043538129602d2f2ce6899d8713125c7d19734aa557ce2e3ea55bd4f1", [:rebar3], [], "hexpm", "6ae959a025bf36df61a8cab8508d9654891b5426a84c44d82deaffd6ddf8c71f"}, "jason": {:hex, :jason, "1.4.5", "2e3a008590b0b8d7388c20293e9dcc9cf3e5d642fd2a114e4cbbb52e595d940a", [:mix], [{:decimal, "~> 1.0 or ~> 2.0 or ~> 3.0", [hex: :decimal, repo: "hexpm", optional: true]}], "hexpm", "b0c823996102bcd0239b3c2444eb00409b72f6a140c1950bc8b457d836b30684"}, "jaxon": {:hex, :jaxon, "2.0.8", "00951a79d354260e28d7e36f956c3de94818124768a4b22e0fc55559d1b3bfe7", [:make, :mix], [{:elixir_make, "~> 0.4", [hex: :elixir_make, repo: "hexpm", optional: false]}], "hexpm", "74532853b1126609615ea98f0ceb5009e70465ca98027afbbd8ed314d887e82d"}, @@ -19,7 +19,7 @@ "jsv": {:hex, :jsv, "0.22.0", "3a2bb35dd7d1ca0034437bee8099214e40d391a5e4f058f5a0800aef91c4dfeb", [:mix], [{:abnf_parsec, "~> 2.0", [hex: :abnf_parsec, repo: "hexpm", optional: false]}, {:decimal, "~> 2.0 or ~> 3.0", [hex: :decimal, repo: "hexpm", optional: true]}, {:idna, "~> 6.0 or ~> 7.0", [hex: :idna, repo: "hexpm", optional: false]}, {:jason, "~> 1.0", [hex: :jason, repo: "hexpm", optional: true]}, {:texture, ">= 1.2.1", [hex: :texture, repo: "hexpm", optional: false]}], "hexpm", "79bae1f970413c86771051a8ea0bd553cc1e866d285270be539f1ef3ca044f3c"}, "llm_db": {:hex, :llm_db, "2026.7.5", "38e345e753b027f095e9eb01157de911c1e75d5ac7d760042f771cbdeabb08c1", [:mix], [{:dotenvy, "~> 1.1", [hex: :dotenvy, repo: "hexpm", optional: false]}, {:igniter, "~> 0.7", [hex: :igniter, repo: "hexpm", optional: true]}, {:jason, "~> 1.4", [hex: :jason, repo: "hexpm", optional: false]}, {:req, "~> 0.5", [hex: :req, repo: "hexpm", optional: false]}, {:toml, "~> 0.7", [hex: :toml, repo: "hexpm", optional: false]}, {:zoi, "~> 0.10", [hex: :zoi, repo: "hexpm", optional: false]}], "hexpm", "93ceaf3448b8cea11358854388e83480ed4e84e39adc0af8a3629935d3d5dd7c"}, "mime": {:hex, :mime, "2.0.7", "b8d739037be7cd402aee1ba0306edfdef982687ee7e9859bee6198c1e7e2f128", [:mix], [], "hexpm", "6171188e399ee16023ffc5b76ce445eb6d9672e2e241d2df6050f3c771e80ccd"}, - "mint": {:hex, :mint, "1.11.0", "a713551624815c0435237b93d90ea8b9b14254690c66f732d0ef8930f76ff1d9", [:mix], [{:castore, "~> 0.1.0 or ~> 1.0", [hex: :castore, repo: "hexpm", optional: true]}, {:hpax, "~> 1.1", [hex: :hpax, repo: "hexpm", optional: false]}], "hexpm", "c6279ba2d6aa3a383a1d4cfbe7b59f42e6efd400f58d8e2acfeac48a438693ab"}, + "mint": {:hex, :mint, "1.10.1", "c53e70867cf74017716884d8d33e0742b08b32e9cdb0031cbc69a429dc5555e3", [:mix], [{:castore, "~> 0.1.0 or ~> 1.0", [hex: :castore, repo: "hexpm", optional: true]}, {:hpax, "~> 0.1.1 or ~> 0.2.0 or ~> 1.0", [hex: :hpax, repo: "hexpm", optional: false]}], "hexpm", "0ba2a904605ed8406393444fb8b3356dc58eb59ee6c7fb94ac3f015e1be129e8"}, "mint_web_socket": {:hex, :mint_web_socket, "1.0.6", "5ffcf350df5b90f2d7a04adf877165228804993714592512374218d4679e325a", [:mix], [{:mint, ">= 1.4.1 and < 2.0.0-0", [hex: :mint, repo: "hexpm", optional: false]}], "hexpm", "0c360e9012413f1c115a63532601eb5d63731aab7010949178769760686c1698"}, "nimble_csv": {:hex, :nimble_csv, "1.3.0", "b7f998dc62b222bce9596e46f028c7a5af04cb5dde6df2ea197c583227c54971", [:mix], [], "hexpm", "41ccdc18f7c8f8bb06e84164fc51635321e80d5a3b450761c4997d620925d619"}, "nimble_options": {:hex, :nimble_options, "1.1.1", "e3a492d54d85fc3fd7c5baf411d9d2852922f66e69476317787a7b2bb000a61b", [:mix], [], "hexpm", "821b2470ca9442c4b6984882fe9bb0389371b8ddec4d45a9504f00a66f650b44"}, diff --git a/mix.exs b/mix.exs index c58f0c3d..ee3f2dd0 100644 --- a/mix.exs +++ b/mix.exs @@ -139,9 +139,12 @@ defmodule Imp.MixProject do # Imp matches Mint's error structs (Mint.TransportError, Mint.HTTPError) # to tell a request that was never sent from one that may have run # (Imp.Clients.ReqLLM, Imp.MCP.CallFailure), so it depends on Mint - # directly. 1.11.0 is the first release without EEF-CVE-2026-91043, - # EEF-CVE-2026-92103 and EEF-CVE-2026-94194. - {:mint, "~> 1.11"}, + # directly. The floor is not 1.11: mint 1.11.0 leaves an HTTP/1 + # connection open after a receive timeout and Finch 0.23.0 reuses it, so + # mix.lock holds 1.10.1 (test/timed_out_connection_test.exs, + # .audit_ignore) until a Finch release includes + # https://github.com/sneako/finch/pull/397. + {:mint, "~> 1.10"}, {:nimble_csv, "~> 1.3"}, {:nimble_options, "~> 1.1"}, # The demo MCP servers (Imp.ACP.DemoMCPHTTPPlug, DemoMCPOAuthPlug, the diff --git a/mix.lock b/mix.lock index f3c69ac5..a0c22d51 100644 --- a/mix.lock +++ b/mix.lock @@ -19,7 +19,7 @@ "ex_mcp": {:hex, :ex_mcp, "1.5.0", "bf0a6862b306d4ba76c29848db110b49d1cbe83b304c9d1ba2ea6b09a7f83b9a", [:mix], [{:castore, "~> 1.0", [hex: :castore, repo: "hexpm", optional: false]}, {:ex_json_schema, "~> 0.10", [hex: :ex_json_schema, repo: "hexpm", optional: false]}, {:fuse, "~> 2.4", [hex: :fuse, repo: "hexpm", optional: true]}, {:jason, "~> 1.4", [hex: :jason, repo: "hexpm", optional: false]}, {:jose, "~> 1.11", [hex: :jose, repo: "hexpm", optional: false]}, {:mint, "~> 1.6", [hex: :mint, repo: "hexpm", optional: false]}, {:mint_web_socket, "~> 1.0", [hex: :mint_web_socket, repo: "hexpm", optional: false]}, {:plug, "~> 1.16", [hex: :plug, repo: "hexpm", optional: false]}, {:plug_cowboy, "~> 2.7", [hex: :plug_cowboy, repo: "hexpm", optional: false]}, {:telemetry, "~> 1.2", [hex: :telemetry, repo: "hexpm", optional: false]}], "hexpm", "3cb3cd60e2275277f519ec6db02cf5e92b342d88c5ba0d31e31bca498b28e4fb"}, "file_system": {:hex, :file_system, "1.1.1", "31864f4685b0148f25bd3fbef2b1228457c0c89024ad67f7a81a3ffbc0bbad3a", [:mix], [], "hexpm", "7a15ff97dfe526aeefb090a7a9d3d03aa907e100e262a0f8f7746b78f8f87a5d"}, "finch": {:hex, :finch, "0.23.0", "e3f9287ac25a8832f848b144c2b57346aac65b205e2e0629a52adfe6507fd837", [:mix], [{:mime, "~> 1.0 or ~> 2.0", [hex: :mime, repo: "hexpm", optional: false]}, {:mint, "~> 1.8", [hex: :mint, repo: "hexpm", optional: false]}, {:nimble_options, "~> 0.4 or ~> 1.0", [hex: :nimble_options, repo: "hexpm", optional: false]}, {:nimble_pool, "~> 1.1", [hex: :nimble_pool, repo: "hexpm", optional: false]}, {:telemetry, "~> 0.4 or ~> 1.0", [hex: :telemetry, repo: "hexpm", optional: false]}], "hexpm", "80e58d3f936f57e3fdf404f83a3642897ae6d9fb642934e46da4d8fe761b99d5"}, - "hpax": {:hex, :hpax, "1.1.0", "782931867cc23217c68fb5f68fe1a11f5e7544c7fda82c8a7019a5df5a4a1cdf", [:mix], [], "hexpm", "0b8d0f05832f55571d65ac720f79bf8994138ffbb133209dc4685eae0ad456a8"}, + "hpax": {:hex, :hpax, "1.0.4", "777de5d433b0fbdc7c418159c8055910faa8047ffdb3d6b31098d2a46cd7685c", [:mix], [], "hexpm", "afc7cb142ebcc2d01ce7816190b98ce5dd49e799111b24249f3443d730f377ca"}, "idna": {:hex, :idna, "7.1.0", "1067a13043538129602d2f2ce6899d8713125c7d19734aa557ce2e3ea55bd4f1", [:rebar3], [], "hexpm", "6ae959a025bf36df61a8cab8508d9654891b5426a84c44d82deaffd6ddf8c71f"}, "jason": {:hex, :jason, "1.4.5", "2e3a008590b0b8d7388c20293e9dcc9cf3e5d642fd2a114e4cbbb52e595d940a", [:mix], [{:decimal, "~> 1.0 or ~> 2.0 or ~> 3.0", [hex: :decimal, repo: "hexpm", optional: true]}], "hexpm", "b0c823996102bcd0239b3c2444eb00409b72f6a140c1950bc8b457d836b30684"}, "jaxon": {:hex, :jaxon, "2.0.8", "00951a79d354260e28d7e36f956c3de94818124768a4b22e0fc55559d1b3bfe7", [:make, :mix], [{:elixir_make, "~> 0.4", [hex: :elixir_make, repo: "hexpm", optional: false]}], "hexpm", "74532853b1126609615ea98f0ceb5009e70465ca98027afbbd8ed314d887e82d"}, @@ -30,7 +30,7 @@ "makeup_elixir": {:hex, :makeup_elixir, "1.0.1", "e928a4f984e795e41e3abd27bfc09f51db16ab8ba1aebdba2b3a575437efafc2", [:mix], [{:makeup, "~> 1.0", [hex: :makeup, repo: "hexpm", optional: false]}, {:nimble_parsec, "~> 1.2.3 or ~> 1.3", [hex: :nimble_parsec, repo: "hexpm", optional: false]}], "hexpm", "7284900d412a3e5cfd97fdaed4f5ed389b8f2b4cb49efc0eb3bd10e2febf9507"}, "makeup_erlang": {:hex, :makeup_erlang, "1.1.0", "835f7e60792e08824cda445639555d7bf1bbbddb1b60b306e33cb6f6db24dc74", [:mix], [{:makeup, "~> 1.0", [hex: :makeup, repo: "hexpm", optional: false]}], "hexpm", "1cd6780fb1dd1a03979abaed0fe82712b0625118fd5257d3ebbf73f960c73c3c"}, "mime": {:hex, :mime, "2.0.7", "b8d739037be7cd402aee1ba0306edfdef982687ee7e9859bee6198c1e7e2f128", [:mix], [], "hexpm", "6171188e399ee16023ffc5b76ce445eb6d9672e2e241d2df6050f3c771e80ccd"}, - "mint": {:hex, :mint, "1.11.0", "a713551624815c0435237b93d90ea8b9b14254690c66f732d0ef8930f76ff1d9", [:mix], [{:castore, "~> 0.1.0 or ~> 1.0", [hex: :castore, repo: "hexpm", optional: true]}, {:hpax, "~> 1.1", [hex: :hpax, repo: "hexpm", optional: false]}], "hexpm", "c6279ba2d6aa3a383a1d4cfbe7b59f42e6efd400f58d8e2acfeac48a438693ab"}, + "mint": {:hex, :mint, "1.10.1", "c53e70867cf74017716884d8d33e0742b08b32e9cdb0031cbc69a429dc5555e3", [:mix], [{:castore, "~> 0.1.0 or ~> 1.0", [hex: :castore, repo: "hexpm", optional: true]}, {:hpax, "~> 0.1.1 or ~> 0.2.0 or ~> 1.0", [hex: :hpax, repo: "hexpm", optional: false]}], "hexpm", "0ba2a904605ed8406393444fb8b3356dc58eb59ee6c7fb94ac3f015e1be129e8"}, "mint_web_socket": {:hex, :mint_web_socket, "1.0.6", "5ffcf350df5b90f2d7a04adf877165228804993714592512374218d4679e325a", [:mix], [{:mint, ">= 1.4.1 and < 2.0.0-0", [hex: :mint, repo: "hexpm", optional: false]}], "hexpm", "0c360e9012413f1c115a63532601eb5d63731aab7010949178769760686c1698"}, "mix_audit": {:hex, :mix_audit, "2.1.5", "c0f77cee6b4ef9d97e37772359a187a166c7a1e0e08b50edf5bf6959dfe5a016", [:make, :mix], [{:jason, "~> 1.4", [hex: :jason, repo: "hexpm", optional: false]}, {:yaml_elixir, "~> 2.11", [hex: :yaml_elixir, repo: "hexpm", optional: false]}], "hexpm", "87f9298e21da32f697af535475860dc1d3617a010e0b418d2ec6142bc8b42d69"}, "mox": {:hex, :mox, "1.3.2", "f34ca4331b1cce3125609c6de674739fe5f3df0d68df25510d64b6a94b2246de", [:mix], [{:nimble_ownership, "~> 1.0", [hex: :nimble_ownership, repo: "hexpm", optional: false]}], "hexpm", "97918a185e727f3128a826f01827b5b8168e1b815bda19eb4192c3ffdb6a8054"}, diff --git a/test/timed_out_connection_test.exs b/test/timed_out_connection_test.exs new file mode 100644 index 00000000..1a8e4590 --- /dev/null +++ b/test/timed_out_connection_test.exs @@ -0,0 +1,80 @@ +defmodule Imp.TimedOutConnectionTest do + use ExUnit.Case, async: true + + # This test is why Imp's lock holds mint 1.10.1 rather than 1.11.0. + # + # mint 1.11.0 leaves an HTTP/1 connection open after a receive timeout, with + # the request still in flight, and Finch 0.23.0 returns any open connection to + # its pool. The next request that pool gives that connection is written behind + # the one that was never answered, so it times out too, and so does every + # retry that lands there. On mint 1.10.1 the timeout closes the connection and + # the next request opens a new one. + # + # The Finch fix is https://github.com/sneako/finch/pull/397 ("Close abandoned + # HTTP/1 connections after request errors"), which is open and in no release. + # When a Finch release includes it, the lock takes that release and mint 1.11 + # together, and this test passes on both. + + @completion %{ + "id" => "late", + "object" => "chat.completion", + "model" => "timeout-model", + "choices" => [ + %{ + "index" => 0, + "finish_reason" => "stop", + "message" => %{"role" => "assistant", "content" => "pong"} + } + ], + "usage" => %{"prompt_tokens" => 1, "completion_tokens" => 1, "total_tokens" => 2} + } + + test "the call after a timed-out call is answered on a new connection" do + {:ok, seen} = Agent.start_link(fn -> [] end) + + # Bandit serves each HTTP/1 connection from one process, so the handler's + # pid names the connection a request came in on. The first request is held + # past the client's timeout; every later one is answered at once. + base_url = + Imp.Test.LocalHTTP.start(fn _request -> + connection = self() + count = Agent.get_and_update(seen, &{length(&1) + 1, &1 ++ [connection]}) + + if count == 1 do + receive do + :release -> :ok + after + 2_000 -> :ok + end + end + + {200, @completion} + end) + + lm = + Imp.req_llm( + %{ + provider: :openai, + id: "timeout-model", + model: "timeout-model", + base_url: base_url <> "/v1" + }, + api_key: "local-test-key", + cache: false + ) + + # ReqLLM's Finch pool spreads a host's requests at random over several + # one-connection shards, so a request after the timeout meets the stuck + # connection only when it draws the same shard. Choosing the first shard + # for every request makes that meeting certain. + finch = [name: ReqLLM.Application.finch_name(), pool_strategy: &hd/1] + opts = [timeout: 200, max_retries: 0, req_http_options: [finch: finch]] + messages = [%{role: :user, content: "ping"}] + + assert {:error, %Imp.LMError{retryable: true}} = Imp.LM.generate(lm, messages, opts) + assert {:ok, _reply} = Imp.LM.generate(lm, messages, opts) + + assert [first, second] = Agent.get(seen, & &1) + refute first == second + end +end From 73212a327697e2c19612a623fd3ac8d420b6265f Mon Sep 17 00:00:00 2001 From: deepfates Date: Mon, 28 Sep 2026 11:37:46 -0700 Subject: [PATCH 8/8] mint ~> 1.8, Finch's own requirement, so declaring it moves no lock --- CHANGELOG.md | 9 +++------ RELEASE_NOTES.md | 2 +- mix.exs | 5 +++-- 3 files changed, 7 insertions(+), 9 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index e3092014..56ba9f5b 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -294,13 +294,10 @@ User-visible changes to Imp are recorded here. dependency (`Imp.Core` reads a reported cost given as a `Decimal`). It declares `plug` and `plug_cowboy` with `runtime: false` for the demo MCP servers, which the package leaves out: Imp starts neither, and no Imp code - that ships uses them. It declares `mint` `~> 1.10`, since it matches Mint's + that ships uses them. It declares `mint` `~> 1.8`, since it matches Mint's error structs to tell a request that was never sent from one that may have - run. The other requirements are ones Req and ExMCP already set (ExMCP - requires `decimal` `~> 3.0`), so they move no lock. The `mint` requirement - moves only a lock below 1.10: `mix deps.get` then takes the newest release, - 1.11.0, unless the application asks for `{:mint, "~> 1.10.1"}` (see - Security). + run. Every one of these requirements is one Finch, Req or ExMCP already set + (ExMCP requires `decimal` `~> 3.0`), so none moves an application's lock. - Imp no longer depends on `jsv`, which nothing in Imp uses. ReqLLM still requires it. diff --git a/RELEASE_NOTES.md b/RELEASE_NOTES.md index 2c6d38ec..2d170314 100644 --- a/RELEASE_NOTES.md +++ b/RELEASE_NOTES.md @@ -40,7 +40,7 @@ containing CR or LF, and a fresh `mix deps.get` resolves Cowboy 2.19.0. The seco for an outgoing `Cookie` request header, which nothing in Imp's dependency tree calls, and no cowlib release fixes it yet. -Imp declares `mint` `~> 1.10`, and a fresh `mix deps.get` resolves `mint` +Imp declares `mint` `~> 1.8`, as Finch does, and a fresh `mix deps.get` resolves `mint` 1.11.0. That release fixes three advisories but reuses HTTP/1 connections that timed out; Imp's own lock holds 1.10.1. See Known limits for the choice an application has. diff --git a/mix.exs b/mix.exs index ee3f2dd0..fb186dd7 100644 --- a/mix.exs +++ b/mix.exs @@ -139,12 +139,13 @@ defmodule Imp.MixProject do # Imp matches Mint's error structs (Mint.TransportError, Mint.HTTPError) # to tell a request that was never sent from one that may have run # (Imp.Clients.ReqLLM, Imp.MCP.CallFailure), so it depends on Mint - # directly. The floor is not 1.11: mint 1.11.0 leaves an HTTP/1 + # directly. The requirement is Finch's own, so declaring it moves no + # application's lock. It is not 1.11: mint 1.11.0 leaves an HTTP/1 # connection open after a receive timeout and Finch 0.23.0 reuses it, so # mix.lock holds 1.10.1 (test/timed_out_connection_test.exs, # .audit_ignore) until a Finch release includes # https://github.com/sneako/finch/pull/397. - {:mint, "~> 1.10"}, + {:mint, "~> 1.8"}, {:nimble_csv, "~> 1.3"}, {:nimble_options, "~> 1.1"}, # The demo MCP servers (Imp.ACP.DemoMCPHTTPPlug, DemoMCPOAuthPlug, the