Skip to content

Audio simulation judge reports no agent messages despite five Agent turns in printed transcript #1005

Description

@ak-af

lk agent simulate audio can produce a judge verdict based on a conversation with no agent messages even though the transcript printed by the same command contains multiple agent responses and tool calls.

Environment: lk 2.18.6, livekit-agents 1.6.10, audio simulation with a locally spawned Python agent. Invocation:

lk agent simulate audio --scenarios scenarios.yaml agent/main.py -y --concurrency 1

Observed on 2026-09-30: 11 of 48 runs were ungraded this way. In an earlier 45-run artifact sample, 9 had the contradiction, including a printed pass inferred from the caller's side alone. A printed checkmark/cross cannot reliably identify whether the judge actually received both sides of the conversation.

Concrete example: simulation job SRJ_W4tu7ns5SjmH, scenario new-vehicle-name-solterra.

The command prints:

::group::✗ new-vehicle-name-solterra (SRJ_W4tu7ns5SjmH)
Error: The provided conversation contains only user messages. There are no agent responses present to evaluate whether the agent correctly identified or read back the vehicle model 'Solterra'.

Yet the Transcript: section in that same output contains five Agent turns, interleaved with caller turns, availability/read-back/create tool calls and a successful create result. Two of those Agent turns, with personal/scheduling details omitted, read:

● Agent
  ... Does that work for you?
● You
  Yes. That works for me. I'll drop it off then.
...
● Agent
  Confirming your appointment ... Is that correct?
● You
  Yes. That's correct. ...

The final Agent turn announces the booking. We are omitting raw logs/customer details from this public issue; the full artifact is available for private support review.

Expected: the judge receives the same agent/caller conversation represented by the printed transcript, or the CLI exposes an explicit ungraded result separate from a behavioral failure. A genuine silence failure should still be graded as an agent failure.

Current mitigation: detect the contradiction between a judge claiming no agent responses and the printed transcript containing Agent turns; classify it as a harness failure, exclude it from the automatic graded pass-rate denominator, and require independent review against call/audio and backend effects. We have not established where messages are lost (judge input construction, recording/transcript synchronization, or service-side processing).

Could you inspect the judge input for the job above and confirm whether this is a known issue, and the supported way to identify ungraded jobs programmatically?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions