Skip to content

Add connection-driven provider failover for common.ai LLM calls - #72156

Draft
Lee-W wants to merge 1 commit into
apache:mainfrom
astronomer:worktree-common-ai-fallback-conn-ids
Draft

Add connection-driven provider failover for common.ai LLM calls#72156
Lee-W wants to merge 1 commit into
apache:mainfrom
astronomer:worktree-common-ai-fallback-conn-ids

Conversation

@Lee-W

@Lee-W Lee-W commented Aug 27, 2026

Copy link
Copy Markdown
Member

A task pointed at one llm_conn_id had no failover story: the only way to survive a provider outage was to build a pydantic-ai FallbackModel in Dag code, which moves credentials outside Airflow connections and leaves the llm_conn_id path uncovered entirely. An outage therefore meant task failure followed by retries into the same outage.

Carrying the chain on the connection keeps the Dag naming a single connection while whoever administers connections owns the failover topology, so changing a standby provider is a connection edit rather than a Dag deployment. Resolving each hop through the hook registered for its own conn_type is what lets one chain span vendors whose credentials live in different connection fields.

Chains are rejected rather than followed recursively, and model_id is not inherited by the fallbacks: a flat chain is the one a reader can verify off a single connection, and model_id names a model of the primary's provider.


Was generative AI tooling used to co-author this PR?
  • Yes (please specify the tool below)

  • Read the Pull Request Guidelines for more information. Note: commit author/co-author name and email in commits become permanently public when merged.
  • For fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
  • When adding dependency, check compliance with the ASF 3rd Party License Policy.
  • For significant user-facing changes create newsfragment: {pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.

A task pointed at one llm_conn_id had no failover story: the only way to
survive a provider outage was to build a pydantic-ai FallbackModel in Dag
code, which moves credentials outside Airflow connections and leaves the
llm_conn_id path uncovered entirely. An outage therefore meant task failure
followed by retries into the same outage.

Carrying the chain on the connection keeps the Dag naming a single
connection while whoever administers connections owns the failover
topology, so changing a standby provider is a connection edit rather than a
Dag deployment. Resolving each hop through the hook registered for its own
conn_type is what lets one chain span vendors whose credentials live in
different connection fields.

Chains are rejected rather than followed recursively, and model_id is not
inherited by the fallbacks: a flat chain is the one a reader can verify off
a single connection, and model_id names a model of the primary's provider.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant