Skip to content

feat(video): add mixed references with matching budget reservations - #155

Merged
VickyXAI merged 15 commits into
BlockRunAI:mainfrom
KillerQueen-Z:codex/seedance-capability-parity
Sep 29, 2026
Merged

VickyXAI merged 15 commits into
BlockRunAI:mainfrom
KillerQueen-Z:codex/seedance-capability-parity

Conversation

@KillerQueen-Z

@KillerQueen-Z KillerQueen-Z commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Add client parameters for mixed reference media, supported output controls, 2.5 first/last frames and 30 images, and last-frame results. These fields are forwarded through the existing Base, Solana and account transports; acceptance depends on the gateway capability and rollout status below. Include the gateway’s existing reference-media surcharge in budget reservations before signing, and apply URL/SSRF checks to the new media inputs.

Model-specific combinations and unsupported controls are rejected before payment. Existing authoritative-quote checks, paid-request accounting and polling deadlines remain intact.

Validation

Full npm test: 1,321 passed. Focused schema/video regression suite: 48 passed. Re-measured the expanded tool schema and regenerated both context-cost cards; synchronized the README and overhead documentation (full: 13,456 tokens; media: 6,240). This resolves all three failing CI assertions. Earlier TypeScript checking and MCP Apps/server/declaration build passed. SSRF, quote rejection, deadline accounting and settlement tests use mocked transport.

Rollout and limits

This PR does not deploy or change R2V_ENABLED. Base #719 is merged. Solana #367 was closed after the product decision to offer reference media through api.blockrun.ai; Solana #374 (still open at this review) implements that refusal. Do not assume Solana capability parity or enable its reference-media switch. Account API #282 remains open. Publish only after the intended gateway capabilities are deployed; accepting a client parameter does not imply that every gateway supports it.

No new paid upstream renders were performed. New request/response behavior is validated against the captured public upstream parameter schema and mocked contract/payment tests; perform a small paid render smoke test before production enablement. Automatic duration, editing/extension, 2.5 reference video/audio and 1080p, callbacks and draft/flex remain held pending cost/output or lifecycle validation.

See SEEDANCE_CAPABILITIES.md for input examples and capability limits. No authentication, signed-quote format or settlement formula changes.

Related PRs

KillerQueen-Z and others added 8 commits September 23, 2026 13:29
… serves them

Review of BlockRunAI#155 against the three live gateways. Two defects sat on the money
path, and both only bite on the rail the feature actually runs on.

## The reserve was 2-3.2x short on the only rail that serves reference media

`(1 + videos + 0.3*audios)` multiplied the OUTPUT duration, so a clip was
priced as if it were as long as the render. The gateway bills per reference
SECOND (blockrun#730: +108,000 tokens for a 5s reference, +324,000 for a 15s
one — exactly 3x for 3x the length, no per-clip term), and since the caller
sends a URL and never a duration, every clip is quoted at the model's 15.2s
ceiling. seedance-2.0 at 4s with three videos and three audios reserved $4.45
against $14.31 owed.

That gap matters because reference media is served only by api.blockrun.ai,
which bills when the gateway ACCEPTS the job and issues no 402 — so this
estimate is the only pre-payment control there is. The ledger still self-heals
from `x-blockrun-cost-usd`, but the budget gate had already let the job
through and `confirmSpend` had already shown the human the wrong number.

The estimator now mirrors the gateway's arithmetic, reference term and
floored resolution factor included. The count validation also moves ABOVE the
model dispatch: it sat inside the Seedance branch, so grok and sora fell
through to their own tables and dropped the references from the reserve
silently — the fail-open the function's own throw-never-default rule exists
to prevent.

## Reference media was advertised on two rails that refuse it

Live probe 2026-09-26, both wallet gateways:

    {"error":"reference media is not available on this gateway",
     "message":"... served only by api.blockrun.ai ..."}

400, before any quote. The default install is a wallet rail, so following the
new example earned an unreadable `Unexpected status 400 (the endpoint did not
return a quote)` after spending DNS on every reference URL. Refused
client-side now, naming the rail that serves it.

## Also

- The eleven new guards returned through the catch-all's money classifier,
  so "output_format requires Seedance 2.5" arrived as "Video generation
  failed" with advice to try sora-2 — which supports none of these controls.
  They return formatError like their siblings, and each names the model, the
  limit and the offending field.
- The guards now run BEFORE the SSRF loop. 30 reference images on a model
  that rejects them outright used to buy 30 attacker-chosen DNS lookups,
  unbilled and invisible to the ledger, before the refusal. Hosts are also
  deduped, so one CDN is one lookup rather than thirty.
- The SSRF iterable is typed explicitly again: without it a future `.map`
  returning a short tuple compiles, yields `undefined`, and skips the check
  for a whole field.
- `last_frame_url` reaches the TEXT output on all three rails. It was in
  `structuredContent` alone, so a content-only client paid for a frame it
  could not see. `last_frame_backed_up` is guarded like its `backed_up`
  sibling instead of riding on the URL's presence.
- RealFace and first/last-frame refusals read their model lists off the sets.
  2.5 joined FIRST_LAST_FRAME_MODELS in this PR and the sentence did not.
- `reference_image_urls` gets the 2048-char cap its sibling clip URLs already
  had; `seed` gets a range; `safety_identifier`, `watermark` and `input_type`
  get the descriptions every other field has.

Tests: test/video-reference-media.test.ts pins the account rail end to end —
reserve against the gateway's own arithmetic at six probed combinations, the
duration-independence property the old term broke, wallet-rail refusal on
both chains, all 25 guard branches with their accepted twins, and SSRF across
the arrays including a private host behind a public first element and the
one-lookup-per-host rule. verify:prices says in writing that it cannot probe
this surcharge, since neither wallet gateway will quote it.

1338 pass, typecheck clean.
The new SEEDANCE_CAPABILITIES.md opened with "the three gateways use the same
public generation fields", which the live probe contradicts, and listed 2.5
reference media as "held pending verification" while the tool forwarded 30
reference images on 2.5 — the gateway registry has supportsReferenceImages
true and supportsReferenceMedia false there, so the code was right and the
doc was wrong in both directions.

Rewritten around what this repo owns: the rail restriction, the per-model
support matrix, the cost of a reference clip in dollars, and the output
controls. The gateway request shapes, the SDK camelCase field names and the
deployment switches are gone — this repo cannot verify them and they drift
the moment the gateway moves. Moved under docs/ with every other narrative
doc and linked from the README tool table, which was also still describing
the video tool as RealFace-only.

Schema cost re-measured after the description edits: full 13,456 -> 13,981,
media 6,240 -> 6,765, trading unchanged, so the saving is 61% and the badge
reads 14.0K. README rows, overhead doc and both SVG variants follow the
measurement.
The guard move in 1c56e7a stopped half way: the resolution ceiling and the
duration window still sat below the SSRF loop, so a request naming 4K on a
720p model, or 99 seconds on a 15s one, still bought a lookup for every
caller-chosen hostname before the refusal it was always going to get.

Both blocks are pure table reads. They join the rest ahead of the resolver,
and the ordering is pinned by a test that asserts zero hostnames were
resolved for each.
…ody wrote

Round-2 regression pass on the fix itself, both findings mine.

"A single reference video roughly triples a 5s render" was the arithmetic of
the formula I replaced, not the one I wrote. A 15.2s clip is a flat cost, so
the multiple grows as the render shrinks: on 2.0-mini a 5s 720p render goes
~$0.40 to ~$1.61, which is 4x, and 480p is worse because only the render half
takes the discount. The description now gives the dollars instead of a ratio.

The capability doc inherited "unsupported request controls are rejected
before payment rather than silently discarded" from the PR. True of a control
the tool declares for the wrong model — those name the model and the field.
False of a field the tool does not declare at all: zod strips callback_url,
draft and service_tier before the handler runs, so the job proceeds without
them. The doc now says which half is which.

Schema cost re-measured: full 13,981 -> 14,039, media 6,765 -> 6,823.
1339 pass.
…ce job

Round-3 regression pass, all four on my own round-2 fix.

The account-rail refusal ranked below the RealFace and frame-seed guards, so
a wallet-rail request carrying both references and last_frame_url was told to
add image_url, then that frame seeds and references do not mix, and only then
that the rail cannot serve references at all. Three round trips to reach the
one fact that made the other two moot. It is the first guard now, and a test
pins that ordering for four combinations that trip a second guard as well.

The tool description still said one clip "roughly triples" a render, which
was the arithmetic of the formula this branch replaced. It is 4x at 720p and
~24x for three videos plus three audios at 480p, because the resolution
discount reaches the render and not the clip. Dollars now, not ratios.

The capability doc gains the two rows most likely to surprise: a 4K reference
job on seedance-2.0 reserves ~$130.88, and 480p with six clips ~$4.91. Those
are not over-reserves — the gateway's referenceTokens() floors its resolution
factor the same way, so the reserve is the charge.

REFERENCE_IMAGE_LIMIT is read through Object.hasOwn like every other table in
this file, and the two reference-image refusals are split so "wrong model"
and "too many for this model" stop sharing one string. The 480p row in the
reserve table is relabelled: it is what the floor produces, not an observed
charge, and nothing this repo can run will confirm it.

Schema cost: full 14,041, media 6,825. 1340 pass.
Both were redundant with something upstream does anyway, and both were
verified against the gateway's own schema before removal.

`input_type` was a cross-check of an inference this tool makes from fields the
gateway reads too. The gateway declares it `.optional()` and says "omitted
input_type → pure inference", then compares a supplied value against its own
`inferredInputType` and 400s on mismatch. So the parameter could never change
a successful call: its entire reachable effect was to reject a request whose
inputs were fine, and it would reject the CORRECT value first if the two
inferences ever drifted. It is off the schema and off the wire; the gateway
infers.

A clip `role` had one legal value. The gateway's own message reads: 'reference
clip role must be "reference" (the only role the upstream honours); omit it or
set it to "reference"'. A knob with one setting is a decoy for the calling
model, so the shape is now `{ url }` and a stray role is stripped by schema
validation rather than forwarded.

The ten pass-through assignments they sat among collapse to one loop over a
table. Still `!== undefined` and not truthiness — seed 0, camera_fixed false
and watermark false are all meaningful values a truthy test would drop, which
the 1.5-pro test already pins.

Schema cost falls with them: full 14,041 -> 13,971, media 6,825 -> 6,755.
1341 pass.
blockrun_image on the Solana wallet rail sent one POST and never polled.
Past the gateway's ~30s inline window the route answers 202 { id,
poll_url } instead of the image, so the tool returned "No image URL in
response" — and the charge stood, because Solana settles at submit and
cannot settle after a long render (a signed transaction expires with its
blockhash). The user paid and got nothing.

Observed live: google/nano-banana-pro at 4096x4096 booked $0.1575 and
returned no image. The same defect was fixed on the account rail in BlockRunAI#140;
the Solana wallet rail was never moved across. It now uses
solanaPaidAsyncPost, the helper video and music already use, which
handles the inline 200 and the 202 alike and re-signs each poll.

Also fixes a hazard that switch exposed: solanaPaidAsyncPost marked the
payment tracker answered as soon as the submit response arrived,
including a 502/503/504 from the edge. On this rail the submit IS the
paid request, so an origin that never answered may still have settled it
— and the tool said "temporary API issue, try again" on a charge that
already stood. An edge status is no longer the origin's verdict; the
tracker stays armed and the tool books it as a precaution. Video and
music get the same protection.

Two test doubles were letting this through and are repaired rather than
relaxed: quote-guard matched the old helper name, and the async helper's
stand-in never fired onPaidRequest, so the tracker it exists to exercise
was never armed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
VickyXAI and others added 5 commits September 29, 2026 21:10
…firm, say 4x

Review of BlockRunAI#155. Three findings on the reference-media path:

- With references present, the frame-seed conflict now ranks directly under
  the rail refusal. Below the last_frame/RealFace guards, a reference request
  carrying last_frame_url was told to add image_url, then on the resubmit
  that seeds and references never mix.
- The spend confirmation lists reference images/videos/audios. Clips can put
  a 5s render at 4-24x (~$130.88 at 4K with 3+3), and the dialog read
  'video · model · 5s' for all of it.
- reference_videos still said one clip 'roughly triples' a render — the
  replaced formula's arithmetic. It is ~4x at 720p, more at 480p.

Each pinned by a test that fails on the previous source.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Since 0.52.2 the Solana image path can throw BilledJobError (a settled-at-
submit give-up in the async helper). The catch was written for the account
rail: 'the account is charged', 'check user.blockrun.ai/dashboard/activity'
— a dashboard with no record of an on-chain transfer. It now branches on
the rail the way music and video already do.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
full 13,971 -> 14,000, media 6,755 -> 6,784; badge unchanged at 14.0K.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@VickyXAI
VickyXAI merged commit bbd653b into BlockRunAI:main Sep 29, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants