Skip to content

Expose structured model cost/usage in DatasetCreationResults #956

Description

@jeremyjordan

Priority Level

Medium (Nice to have)

Is your feature request related to a problem? Please describe.

As we iterate on Data Designer pipelines, we want to track the associated cost per run. This is useful alongside other metrics (e.g. quality scores, % of discarded rows) as we perform experiments to improve the pipeline. We would like to see how changes to the pipeline either reduce costs, improve quality, or a combination of the two.

Describe the solution you'd like

Data Designer tracks model usage internally, but does not expose or persist it through DatasetCreationResults.

This makes it difficult for generation pipelines to report tokens and cost. Could Data Designer expose a structured usage summary on each result?

The summary should include:

  • Model alias and model name.
  • Input and output tokens.
  • Successful and failed logical request counts.
  • Provider-reported cost and currency, when available.

For composite workflows, usage should be available per stage and as a workflow total. Data Designer should persist it with the run artifacts so resumed runs can reconstruct cumulative usage without double counting.

Missing usage or cost should remain null or explicitly unavailable. Data Designer should not silently estimate cost from an unversioned pricing table.

A suitable API could be:

result = data_designer.create(...)

result.model_usage

This would let downstream systems report cost and usage without importing private engine APIs.

Describe alternatives you've considered

No response

Agent Investigation

We found an internal Data Designer observer:

from data_designer.engine.models.usage_events import subscribe_token_usage

The callback receives one event after each successful logical model call. Each event contains:

  • Model alias and model name.
  • Input tokens.
  • Output tokens.
  • Runtime correlation metadata.

However, the callback is not a good reporting contract:

  • It is process-global. A subscriber can receive events from unrelated concurrent Data Designer runs.
  • Filtering by model alias reduces contamination but cannot prevent it when runs share an alias.
  • It is an internal API. Data Designer can change it without preserving compatibility.
  • It reports successful calls only. It omits failed calls, provider attempts, retries, and corrections.
  • A successful response without provider usage emits zero tokens. Consumers cannot distinguish missing usage from a real zero-token response.
  • It does not include monetary cost.
  • It does not persist usage. Resumed runs cannot recover usage from earlier process invocations.
  • It does not directly identify the composite workflow stage. Stage attribution requires an additional wrapper around each DataDesigner.create() call.
  • DatasetCreationResults does not retain the underlying model registry or its richer usage statistics.

Our workaround subscribed during workflow.run(), filtered expected model aliases, and serialized our own concurrent generation calls. It still could not protect against unrelated Data Designer calls in the same process. It also had to report incomplete or unknown usage as null.

A public usage summary on DatasetCreationResults would be more accurate, stable, and natural. Persisting that summary would also make composite workflows and resumed runs reliable.

Additional context

No response

Checklist

  • I've reviewed existing issues and the documentation
  • This is a design proposal, not a "please build this" request

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requesttriagedIssue reviewed and approved by a maintainer

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions