lk agent simulate audio can produce a judge verdict based on a conversation with no agent messages even though the transcript printed by the same command contains multiple agent responses and tool calls.
Environment: lk 2.18.6, livekit-agents 1.6.10, audio simulation with a locally spawned Python agent. Invocation:
lk agent simulate audio --scenarios scenarios.yaml agent/main.py -y --concurrency 1
Observed on 2026-09-30: 11 of 48 runs were ungraded this way. In an earlier 45-run artifact sample, 9 had the contradiction, including a printed pass inferred from the caller's side alone. A printed checkmark/cross cannot reliably identify whether the judge actually received both sides of the conversation.
Concrete example: simulation job SRJ_W4tu7ns5SjmH, scenario new-vehicle-name-solterra.
The command prints:
::group::✗ new-vehicle-name-solterra (SRJ_W4tu7ns5SjmH)
Error: The provided conversation contains only user messages. There are no agent responses present to evaluate whether the agent correctly identified or read back the vehicle model 'Solterra'.
Yet the Transcript: section in that same output contains five Agent turns, interleaved with caller turns, availability/read-back/create tool calls and a successful create result. Two of those Agent turns, with personal/scheduling details omitted, read:
● Agent
... Does that work for you?
● You
Yes. That works for me. I'll drop it off then.
...
● Agent
Confirming your appointment ... Is that correct?
● You
Yes. That's correct. ...
The final Agent turn announces the booking. We are omitting raw logs/customer details from this public issue; the full artifact is available for private support review.
Expected: the judge receives the same agent/caller conversation represented by the printed transcript, or the CLI exposes an explicit ungraded result separate from a behavioral failure. A genuine silence failure should still be graded as an agent failure.
Current mitigation: detect the contradiction between a judge claiming no agent responses and the printed transcript containing Agent turns; classify it as a harness failure, exclude it from the automatic graded pass-rate denominator, and require independent review against call/audio and backend effects. We have not established where messages are lost (judge input construction, recording/transcript synchronization, or service-side processing).
Could you inspect the judge input for the job above and confirm whether this is a known issue, and the supported way to identify ungraded jobs programmatically?
lk agent simulate audiocan produce a judge verdict based on a conversation with no agent messages even though the transcript printed by the same command contains multiple agent responses and tool calls.Environment: lk 2.18.6, livekit-agents 1.6.10, audio simulation with a locally spawned Python agent. Invocation:
Observed on 2026-09-30: 11 of 48 runs were ungraded this way. In an earlier 45-run artifact sample, 9 had the contradiction, including a printed pass inferred from the caller's side alone. A printed checkmark/cross cannot reliably identify whether the judge actually received both sides of the conversation.
Concrete example: simulation job
SRJ_W4tu7ns5SjmH, scenarionew-vehicle-name-solterra.The command prints:
Yet the
Transcript:section in that same output contains five Agent turns, interleaved with caller turns, availability/read-back/create tool calls and a successful create result. Two of those Agent turns, with personal/scheduling details omitted, read:The final Agent turn announces the booking. We are omitting raw logs/customer details from this public issue; the full artifact is available for private support review.
Expected: the judge receives the same agent/caller conversation represented by the printed transcript, or the CLI exposes an explicit ungraded result separate from a behavioral failure. A genuine silence failure should still be graded as an agent failure.
Current mitigation: detect the contradiction between a judge claiming no agent responses and the printed transcript containing Agent turns; classify it as a harness failure, exclude it from the automatic graded pass-rate denominator, and require independent review against call/audio and backend effects. We have not established where messages are lost (judge input construction, recording/transcript synchronization, or service-side processing).
Could you inspect the judge input for the job above and confirm whether this is a known issue, and the supported way to identify ungraded jobs programmatically?