Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
45 commits
Select commit Hold shift + click to select a range
27a5cd6
feat(harness): generate a suite with one writer per use case, counted…
KarthikAvinashFI Aug 22, 2026
efdb025
feat(harness): write a suite to a plan, review what came back, and ca…
KarthikAvinashFI Aug 22, 2026
40f8fee
Merge remote-tracking branch 'origin/feat/hosted-harness-e2e-runtime'…
KarthikAvinashFI Aug 24, 2026
4e69fe3
perf(harness): make model thinking env-configurable via ALK_HARNESS_T…
KarthikAvinashFI Aug 24, 2026
fa287cd
fix(harness): unblock hosted voice runs (agent Vertex creds mount, pr…
KarthikAvinashFI Aug 24, 2026
170fe05
fix(harness): restore local-env GOOGLE credential fallback and raise …
KarthikAvinashFI Aug 24, 2026
a9db360
feat(harness): drive simulator voice from persona accent, STT languag…
KarthikAvinashFI Aug 24, 2026
4e9111f
feat(harness): vary caller accents and end calls when the agent loops
KarthikAvinashFI Aug 25, 2026
a2d4a1e
feat(simulator): pick Cartesia voice and language by persona accent a…
KarthikAvinashFI Aug 25, 2026
56cfb26
fix(harness): define the review server, stop parallel writers deletin…
KarthikAvinashFI Aug 25, 2026
501114d
Merge remote-tracking branch 'origin/feat/hosted-harness-e2e-runtime'…
KarthikAvinashFI Aug 25, 2026
875e69a
fix(simulator): drop Cartesia voices the provider no longer serves
KarthikAvinashFI Aug 25, 2026
754c179
chore(harness): move credential and runtime limit changes to their ow…
KarthikAvinashFI Aug 25, 2026
76616ca
chore(harness): tighten scenario comments and drop em dashes from gen…
KarthikAvinashFI Aug 25, 2026
8cc41f7
feat(simulator): fill the caller prompt from a template and drive spe…
KarthikAvinashFI Aug 25, 2026
f5e0e3e
feat(simulator): fix the caller prompt template in code instead of ge…
KarthikAvinashFI Aug 25, 2026
ab9076a
feat(scenarios): require the instruction to state an objective and ca…
KarthikAvinashFI Aug 25, 2026
1e2fc2d
chore(simulator): drop the prompt template override nothing sets
KarthikAvinashFI Aug 25, 2026
4331f68
fix(simulator): close background audio on the caller agent and route …
KarthikAvinashFI Aug 25, 2026
82bcf76
feat(simulator): give the caller delivery cues when Cartesia is the v…
KarthikAvinashFI Aug 25, 2026
a84d409
chore(harness): drop the persona guidance reader that nothing calls
KarthikAvinashFI Aug 25, 2026
102050f
fix(harness): let a slice finish before the stage is judged idle
KarthikAvinashFI Aug 25, 2026
24f2b6b
Merge remote-tracking branch 'origin/feat/hosted-harness-e2e-runtime'…
KarthikAvinashFI Aug 25, 2026
813597c
chore(harness): run grading on the same model as everything else
KarthikAvinashFI Aug 25, 2026
2a048e5
fix(harness): size the idle bound for a suite that writes in one tool…
KarthikAvinashFI Aug 25, 2026
67a91d6
fix(harness): pin every model the session can reach to the one the ru…
KarthikAvinashFI Aug 25, 2026
42e2081
fix(scenarios): write a suite one scenario at a time unless fan-out i…
KarthikAvinashFI Aug 25, 2026
8a2181d
feat(scenario): carry the scenario_key and scenario_id a hosted sched…
KarthikAvinashFI Aug 25, 2026
ae982fb
fix(scenarios): stop the instruction telling the caller what the agen…
KarthikAvinashFI Aug 25, 2026
3d612b0
fix(grade): read the speaker off a transcript line instead of leaving…
KarthikAvinashFI Aug 25, 2026
21f0ae6
test(harness): cover scenario identity, deterministic noise and trans…
KarthikAvinashFI Aug 25, 2026
a9463c3
Merge remote-tracking branch 'origin/feat/hosted-harness-e2e-runtime'…
KarthikAvinashFI Aug 25, 2026
64c918b
fix(scenarios): keep the agent's expected moves out of the caller's i…
KarthikAvinashFI Aug 25, 2026
82adf92
fix(simulate): give the caller countable conduct rules instead of one…
KarthikAvinashFI Aug 25, 2026
325db30
fix(scenarios): judge an instruction by whether the caller could say …
KarthikAvinashFI Aug 25, 2026
134a6c4
fix(scenarios): give out-of-band steps a state the caller holds, not …
KarthikAvinashFI Aug 26, 2026
4fe3d32
feat(simulate): add ALK_BACKGROUND_NOISE to silence every call on a run
KarthikAvinashFI Aug 26, 2026
027d46b
fix(scenarios): rename a duplicate folder name instead of dropping th…
KarthikAvinashFI Aug 26, 2026
1441f3a
fix(platform): send the scenario's ascii key rather than its folder name
KarthikAvinashFI Aug 26, 2026
7cf0c18
fix(simulate): close the caller prompt on its objective rather than o…
KarthikAvinashFI Aug 26, 2026
8f1c5c5
fix(simulate): make background noise opt in so a run with no environm…
KarthikAvinashFI Aug 26, 2026
ce4498c
fix(simulate): warn when a missing cartesia key collapses every perso…
KarthikAvinashFI Aug 26, 2026
a536744
fix(simulate): close the caller prompt by naming who the caller is
KarthikAvinashFI Aug 26, 2026
152e02c
fix(simulate): stop handing the caller the grader's pass question as …
KarthikAvinashFI Aug 26, 2026
a18caae
fix(skills): drop the identity claim from the preamble and pin use_ca…
KarthikAvinashFI Aug 26, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
73 changes: 73 additions & 0 deletions src/fi/alk/harness/background_noise.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,73 @@
"""Choose the caller-side ambient noise a scenario should be heard through.

A scenario that sets ``background_noise`` wants the agent to handle a caller phoning from somewhere
real: a car, a street, an office. The clip is chosen here and handed to the voice engine, which
mixes it under the simulated caller's audio.

Two sources, in order. A run may point ``ALK_BACKGROUND_NOISE_CATALOG`` at a JSON file of clips
(each with an ``environment`` tag and a ``url`` or ``path``); the catalog stays a local file so its
asset locations are never committed here. When no catalog matches, a LiveKit builtin clip is used,
which needs no external asset and always works.
"""

from __future__ import annotations

import json
import os
from pathlib import Path

# LiveKit ships these; they are the reliable default when no custom catalog is configured.
_BUILTIN_BY_ENVIRONMENT: dict[str, str] = {
"street": "CITY_AMBIENCE",
"transit": "CITY_AMBIENCE",
"vehicle": "CITY_AMBIENCE",
"outdoors": "FOREST_AMBIENCE",
"retail": "CROWDED_ROOM",
"office": "OFFICE_AMBIENCE",
"home": "OFFICE_AMBIENCE",
}
_DEFAULT_BUILTIN = "OFFICE_AMBIENCE"


def enabled() -> bool:
"""Whether any scenario may be heard through background noise on this run.

Off unless ``ALK_BACKGROUND_NOISE`` opts in, so a run needs no environment at all to be
silent. Continuous ambient audio under the caller competes with endpoint detection, and calls
carrying it end earlier and on fewer turns, so silence is the setting a run should fall into
rather than the one it has to ask for. Opting in still only permits noise: a scenario that
asked for none stays silent either way.
"""
return os.environ.get("ALK_BACKGROUND_NOISE", "0").strip().lower() in (
"1",
"on",
"true",
"yes",
)


def source_for(environment: str = "", seed: str = "") -> str:
"""A background-noise source for a scenario.

Returns a ``url``/``path`` from the configured catalog when one matches the environment, else the
name of a LiveKit builtin clip. The choice is deterministic in ``seed`` so the same scenario
hears the same place across runs.
"""
env = (environment or "").strip().lower()
catalog = os.environ.get("ALK_BACKGROUND_NOISE_CATALOG", "").strip()
if catalog and Path(catalog).is_file():
try:
entries = json.loads(Path(catalog).read_text(encoding="utf-8"))
except (OSError, ValueError):
entries = []
if isinstance(entries, list) and entries:
pool = [
entry
for entry in entries
if str(entry.get("environment", "")).strip().lower() == env
] or entries
chosen = pool[sum(ord(character) for character in (seed or env or "x")) % len(pool)]
located = str(chosen.get("url") or chosen.get("path") or "").strip()
if located:
return located
return _BUILTIN_BY_ENVIRONMENT.get(env, _DEFAULT_BUILTIN)
2 changes: 2 additions & 0 deletions src/fi/alk/harness/build.py
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,7 @@
permission_gate,
provider_env,
provisioning,
thinking_config,
)
from .contract import AgentContract
from .session import Stage
Expand Down Expand Up @@ -202,6 +203,7 @@ def open_stage(
options.disallowed_tools = list(UNWANTED)
options.hooks = gate_hooks(allowed)
options.can_use_tool = permission_gate(ask, allowed)
options.thinking = thinking_config()
return Stage(options, name=SKILL), destination


Expand Down
31 changes: 30 additions & 1 deletion src/fi/alk/harness/config.py
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,24 @@ def chosen_model(model: str | None = None) -> str:
return model or os.environ.get("ALK_HARNESS_MODEL", DEFAULT_MODEL)


def thinking_config() -> dict[str, Any]:
"""How much the model may think, from ALK_HARNESS_THINKING.

The Claude Code CLI defaults to adaptive thinking. In this harness the correctness of what a
stage produces is re-checked by code gates (a scenario is proved against the real world, a
contract is validated), so the model's private reasoning is spent on decisions the gates make
again anyway. Left unset, that reasoning was the majority of generated tokens and the majority
of wall time. Default to disabled for speed; ``adaptive`` restores the old behaviour, and an
integer sets an explicit budget for models that still honour one.
"""
setting = os.environ.get("ALK_HARNESS_THINKING", "disabled").strip().lower()
if setting in {"adaptive", "on", "auto"}:
return {"type": "adaptive", "display": "omitted"}
if setting.isdigit() and int(setting) > 0:
return {"type": "enabled", "budget_tokens": int(setting), "display": "omitted"}
return {"type": "disabled"}


def provisioning(enabled: bool | None = None) -> bool:
"""Compatibility switch for callers selecting the legacy provisioning surface.

Expand All @@ -75,10 +93,20 @@ def provider_env(model: str | None = None) -> dict[str, str]:
Claude Code resolves the GCP project from ``GOOGLE_CLOUD_PROJECT``, the credential file, or
the active gcloud configuration, in that order, so an unset project id is not an error here.
"""
# Every model a session can reach is pinned to the same one. Naming only the main model
# leaves the sub-agent and fast-path settings to the CLI's own preference, and a suite written
# by twenty writers then runs on whatever that preference happens to be rather than on the
# model the run asked for.
chosen = chosen_model(model)
env = {
"CLAUDE_CODE_USE_VERTEX": "1",
"CLOUD_ML_REGION": os.environ.get("CLOUD_ML_REGION", "global"),
"ANTHROPIC_MODEL": chosen_model(model),
"ANTHROPIC_MODEL": chosen,
"ANTHROPIC_DEFAULT_SONNET_MODEL": chosen,
"ANTHROPIC_DEFAULT_OPUS_MODEL": chosen,
"ANTHROPIC_DEFAULT_HAIKU_MODEL": chosen,
"ANTHROPIC_SMALL_FAST_MODEL": chosen,
"CLAUDE_CODE_SUBAGENT_MODEL": chosen,
}
for passthrough in (
"ANTHROPIC_VERTEX_PROJECT_ID",
Expand Down Expand Up @@ -123,6 +151,7 @@ def read_only_session(
options.disallowed_tools = list(UNWANTED)
options.hooks = gate_hooks(allowed)
options.can_use_tool = permission_gate(granted=allowed)
options.thinking = thinking_config()
return options


Expand Down
111 changes: 111 additions & 0 deletions src/fi/alk/harness/data/persona_vocabulary.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,111 @@
{
"GenderChoices": [
"male",
"female"
],
"AgeGroupChoices": [
"18-25",
"25-32",
"32-40",
"40-50",
"50-60",
"60+"
],
"LocationChoices": [
"United States",
"Canada",
"United Kingdom",
"Australia",
"India"
],
"ProfessionChoices": [
"Student",
"Teacher",
"Engineer",
"Doctor",
"Nurse",
"Business Owner",
"Manager",
"Sales Representative",
"Customer Service",
"Technician",
"Consultant",
"Accountant",
"Marketing Professional",
"Retired",
"Homemaker",
"Freelancer",
"Other"
],
"PersonalityChoices": [
"Friendly and cooperative",
"Professional and formal",
"Cautious and skeptical",
"Impatient and direct",
"Detail-oriented",
"Easy-going",
"Anxious",
"Confident",
"Analytical",
"Emotional",
"Reserved",
"Talkative"
],
"CommunicationStyleChoices": [
"Direct and concise",
"Detailed and elaborate",
"Casual and friendly",
"Formal and polite",
"Technical",
"Simple and clear",
"Questioning",
"Assertive",
"Passive",
"Collaborative"
],
"AccentChoices": [
"American",
"Australian",
"Indian",
"Canadian",
"Neutral"
],
"LanguageChoices": [
"Arabic",
"Bulgarian",
"Chinese Simplified",
"Czech",
"Danish",
"Dutch",
"English",
"Finnish",
"French",
"German",
"Greek",
"Hindi",
"Hungarian",
"Indonesian",
"Italian",
"Japanese",
"Korean",
"Malay",
"Norwegian",
"Polish",
"Portuguese",
"Romanian",
"Russian",
"Slovak",
"Spanish",
"Swedish",
"Turkish",
"Ukrainian",
"Vietnamese"
],
"ConversationSpeedChoices": [
"0.5",
"0.75",
"1.0",
"1.25",
"1.5"
]
}
Loading
Loading