Make OpenAI Responses token ceilings templated and surface truncation - #72150
Draft
Lee-W wants to merge 1 commit into
Draft
Make OpenAI Responses token ceilings templated and surface truncation#72150Lee-W wants to merge 1 commit into
Lee-W wants to merge 1 commit into
Conversation
Dags need a per-run token ceiling that can vary by environment, but max_output_tokens and max_tool_calls were only reachable through the non-templated response_kwargs dict. Promote both to first-class, templated operator arguments, coerce the rendered string to a positive int before the request goes out, and reject an invalid or non-positive value instead of silently running with no ceiling. Also reject the same key being set both as an argument and inside response_kwargs, so one silently wins. OpenAI's Responses API has no monetary cost limit to expose, so this is a token ceiling only -- for a cost cap, use apache-airflow-providers-common-ai instead. Hitting max_output_tokens does not fail the request; the response comes back incomplete with truncated output text, which the previous 'may be empty' warning did not describe. Surface incomplete_details.reason in the log message so operators can tell a truncation apart from other non-completed statuses.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Dags need a per-run token ceiling that can vary by environment, but max_output_tokens and max_tool_calls were only reachable through the non-templated response_kwargs dict. Promote both to first-class, templated operator arguments, coerce the rendered string to a positive int before the request goes out, and reject an invalid or non-positive value instead of silently running with no ceiling. Also reject the same key being set both as an argument and inside response_kwargs, so one silently wins.
OpenAI's Responses API has no monetary cost limit to expose, so this is a token ceiling only -- for a cost cap, use
apache-airflow-providers-common-ai instead. Hitting max_output_tokens does not fail the request; the response comes back incomplete with truncated output text, which the previous 'may be empty' warning did not describe. Surface
incomplete_details.reason in the log message so operators can tell a truncation apart from other non-completed statuses.
Was generative AI tooling used to co-author this PR?
{pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.