Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .agents/skills/debug-agentx-runs/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -154,7 +154,7 @@ Use AgentX phase markers, not total Slurm runtime:

```bash
grep -E \
"Phase warmup progress|WARMUP cache pressure|Phase warmup complete|Phase profiling started|Phase profiling complete|replay_rc=" \
"Phase warmup progress|WARMUP cache pressure|Phase warmup complete|Phase profiling started|Phase profiling complete|process_agentic_result" \
"<LOG_DIR>/benchmark.out"
date -u
```
Expand Down
17 changes: 2 additions & 15 deletions .github/workflows/benchmark-multinode-tmpl.yml
Original file line number Diff line number Diff line change
Expand Up @@ -304,7 +304,6 @@ jobs:

- name: Launch multi-node job script
env:
VALIDATION_BENCHMARK_LIB: ${{ github.workspace }}/.result-tooling/inferencex-e2e/benchmarks/benchmark_lib.sh
PREFILL_ADDITIONAL_SETTINGS: ${{ toJSON(fromJSON(inputs.config).prefill.additional-settings) }}
DECODE_ADDITIONAL_SETTINGS: ${{ toJSON(fromJSON(inputs.config).decode.additional-settings) }}
RUNNER_NAME: ${{ runner.name }}
Expand Down Expand Up @@ -335,18 +334,6 @@ jobs:
while IFS= read -r -d '' setting; do
export "$setting"
done < <(jq -j '.[] + "\u0000"' <<< "$settings_json")
# Each AgentX throughput point owns a fresh server deployment.
# Fixed-sequence and graded eval jobs retain their batching semantics.
if [[ "$IS_AGENTIC" == 1 && "$EVAL_ONLY" != true ]]; then
source "$VALIDATION_BENCHMARK_LIB" --validation-only
check_env_vars CONC CONC_LIST
validate_agentic_concurrency "$CONC"
validate_agentic_concurrency "$CONC_LIST"
if [[ "$CONC" != "$CONC_LIST" ]]; then
echo "ERROR: AgentX CONC must match the single CONC_LIST value" >&2
exit 1
fi
fi
# Resolve the workflow's documented automatic eval-concurrency selection.
if [[ -z "$EVAL_CONC" ]]; then
EVAL_CONC=$(python3 -c 'import os; print(max(map(int, os.environ["CONC_LIST"].split())))')
Expand Down Expand Up @@ -397,7 +384,7 @@ jobs:
PYTHONPATH: ${{ github.workspace }}/.result-tooling/inferencex-e2e
run: |
if [[ -f inferencex-e2e/configs/runners.yaml ]]; then cd inferencex-e2e; fi
source "$PYTHONPATH/benchmarks/benchmark_lib.sh" --validation-only
source "$PYTHONPATH/benchmarks/check_env.sh"
check_env_vars INFERENCEX_RESULTS_PYTHON
# Launcher stamp supplies the producer SHA unless the
# power-producer-sha input is set as a manual override
Expand Down Expand Up @@ -481,7 +468,7 @@ jobs:
PYTHONPATH: ${{ github.workspace }}/.result-tooling/inferencex-e2e
run: |
if [[ -f inferencex-e2e/configs/runners.yaml ]]; then cd inferencex-e2e; fi
source "$PYTHONPATH/benchmarks/benchmark_lib.sh" --validation-only
source "$PYTHONPATH/benchmarks/check_env.sh"
check_env_vars INFERENCEX_RESULTS_PYTHON
expected_concs="${EVAL_CONC}"
if [[ -z "${expected_concs}" ]]; then
Expand Down
6 changes: 3 additions & 3 deletions .github/workflows/benchmark-tmpl.yml
Original file line number Diff line number Diff line change
Expand Up @@ -318,7 +318,7 @@ jobs:
PYTHONPATH: ${{ github.workspace }}/.result-tooling/inferencex-e2e
run: |
if [[ -f inferencex-e2e/configs/runners.yaml ]]; then cd inferencex-e2e; fi
source "$PYTHONPATH/benchmarks/benchmark_lib.sh" --validation-only
source "$PYTHONPATH/benchmarks/check_env.sh"
check_env_vars INFERENCEX_RESULTS_PYTHON AIPERF_FAILED_REQUEST_THRESHOLD
"$INFERENCEX_RESULTS_PYTHON" -P -m infx.results.agentic.validate_agentic_result \
results/aiperf_artifacts \
Expand All @@ -331,7 +331,7 @@ jobs:
PYTHONPATH: ${{ github.workspace }}/.result-tooling/inferencex-e2e
run: |
if [[ -f inferencex-e2e/configs/runners.yaml ]]; then cd inferencex-e2e; fi
source "$PYTHONPATH/benchmarks/benchmark_lib.sh" --validation-only
source "$PYTHONPATH/benchmarks/check_env.sh"
check_env_vars INFERENCEX_RESULTS_PYTHON
if [ ! -f "$RESULT_FILENAME.json" ]; then
echo "no raw result to process: $RESULT_FILENAME.json" >&2
Expand Down Expand Up @@ -440,7 +440,7 @@ jobs:
PYTHONPATH: ${{ github.workspace }}/.result-tooling/inferencex-e2e
run: |
if [[ -f inferencex-e2e/configs/runners.yaml ]]; then cd inferencex-e2e; fi
source "$PYTHONPATH/benchmarks/benchmark_lib.sh" --validation-only
source "$PYTHONPATH/benchmarks/check_env.sh"
check_env_vars INFERENCEX_RESULTS_PYTHON
"$INFERENCEX_RESULTS_PYTHON" -P -m infx.evals.validate_scores

Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/collectivex-sweep.yml
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,7 @@ jobs:
RUN_ATTEMPT: ${{ github.run_attempt }}
run: |
set -eo pipefail
source ../inferencex-e2e/benchmarks/benchmark_lib.sh --validation-only
source ../inferencex-e2e/benchmarks/check_env.sh
check_env_vars INPUT_SUITES INPUT_SWAP_PROFILE INPUT_BACKEND RUN_ID RUN_ATTEMPT GITHUB_OUTPUT
args=(--suites "$INPUT_SUITES" --swap-profile "$INPUT_SWAP_PROFILE" --backend "$INPUT_BACKEND")
[ -n "$INPUT_ONLY_SKU" ] && args+=(--only-sku "$INPUT_ONLY_SKU")
Expand Down
12 changes: 6 additions & 6 deletions .github/workflows/operatorx-sweep.yml
Original file line number Diff line number Diff line change
Expand Up @@ -69,7 +69,7 @@ jobs:
PYTHONPATH: .:inferencex-e2e
run: |
set -eo pipefail
source inferencex-e2e/benchmarks/benchmark_lib.sh --validation-only
source inferencex-e2e/benchmarks/check_env.sh
check_env_vars OPERATORX_POOL OPERATORX_BACKENDS OPERATORX_TESTLISTS OPERATORX_WORLD_SIZES OPERATORX_CHUNK_SIZE GITHUB_RUN_ID GITHUB_RUN_ATTEMPT GITHUB_SHA GITHUB_OUTPUT
python3 -m operatorx.ci plan \
--platform-config operatorx/platforms.json \
Expand Down Expand Up @@ -141,7 +141,7 @@ jobs:
OPERATORX_POOL: ${{ matrix.pool }}
run: |
set -eo pipefail
source inferencex-e2e/benchmarks/benchmark_lib.sh --validation-only
source inferencex-e2e/benchmarks/check_env.sh
check_env_vars RUNNER_TEMP OPERATORX_RECOVERY_RUN OPERATORX_POOL OPERATORX_CLEANUP_SECONDS
python3 -m operatorx.ci recover --artifacts operatorx-recovery \
--run-id "$OPERATORX_RECOVERY_RUN" --pool "$OPERATORX_POOL" \
Expand All @@ -150,7 +150,7 @@ jobs:
- name: Run one Slurm allocation
run: |
set -eo pipefail
source inferencex-e2e/benchmarks/benchmark_lib.sh --validation-only
source inferencex-e2e/benchmarks/check_env.sh
check_env_vars RUNNER_TEMP OPERATORX_SHARD OPERATORX_OUTPUT OPERATORX_CLEANUP_SECONDS GITHUB_RUN_ID GITHUB_RUN_ATTEMPT GITHUB_SHA RUNNER_NAME
python3 -m operatorx.ci execute \
--platform-config operatorx/platforms.json \
Expand All @@ -163,15 +163,15 @@ jobs:
if: always()
run: |
set -eo pipefail
source inferencex-e2e/benchmarks/benchmark_lib.sh --validation-only
source inferencex-e2e/benchmarks/check_env.sh
check_env_vars OPERATORX_OUTPUT OPERATORX_CLEANUP_SECONDS
python3 -m operatorx.ci finalize --output "$OPERATORX_OUTPUT" \
--cleanup-seconds "$OPERATORX_CLEANUP_SECONDS"
- name: Capture scheduler diagnostics on failure
if: failure()
run: |
set -eo pipefail
source inferencex-e2e/benchmarks/benchmark_lib.sh --validation-only
source inferencex-e2e/benchmarks/check_env.sh
check_env_vars OPERATORX_OUTPUT
if [ -d "$OPERATORX_OUTPUT" ]; then
timeout 15 sinfo --Node --noheader --format='%N|%P|%t' \
Expand Down Expand Up @@ -210,7 +210,7 @@ jobs:
PYTHONPATH: .:inferencex-e2e
run: |
set -eo pipefail
source inferencex-e2e/benchmarks/benchmark_lib.sh --validation-only
source inferencex-e2e/benchmarks/check_env.sh
check_env_vars GITHUB_STEP_SUMMARY
python3 -m operatorx.ci summarize \
--manifest operatorx-control/operatorx-manifest.json \
Expand Down
4 changes: 2 additions & 2 deletions .github/workflows/profile.yml
Original file line number Diff line number Diff line change
Expand Up @@ -91,7 +91,7 @@ jobs:
PRIORITY_ROOT="$PRIORITY_ROOT/inferencex-e2e"
fi
env PYTHONPATH="$PRIORITY_ROOT" python3 -P -m infx.workflows.require_launcher "$GITHUB_WORKSPACE"
source "$PRIORITY_ROOT/benchmarks/benchmark_lib.sh" --validation-only
source "$PRIORITY_ROOT/benchmarks/check_env.sh"
check_env_vars INPUTS_CONFIG_FILE INPUTS_CONFIG_KEY INPUTS_CONC
GENERATOR_PATH="${MEASURED_ROOT}"
if [ -f "${MEASURED_ROOT}/infx/matrix/generate.py" ]; then
Expand Down Expand Up @@ -350,7 +350,7 @@ jobs:
PYTHONPATH: ${{ github.workspace }}/.result-tooling/inferencex-e2e
run: |
if [[ -f inferencex-e2e/configs/runners.yaml ]]; then cd inferencex-e2e; fi
source "$PYTHONPATH/benchmarks/benchmark_lib.sh" --validation-only
source "$PYTHONPATH/benchmarks/check_env.sh"
check_env_vars INFERENCEX_RESULTS_PYTHON
"$INFERENCEX_RESULTS_PYTHON" -P -m infx.results.fixed_sequence

Expand Down
6 changes: 3 additions & 3 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,15 +22,15 @@ The end-to-end Python project owns `inferencex-e2e/pyproject.toml`, `inferencex-
- **Klaud Cold reports:** Follow the compact body/comment templates in [`inferencex-e2e/docs/klaud-reporting.md`](inferencex-e2e/docs/klaud-reporting.md), including cleanup and completion reports.
- Commit subjects use conventional English style, while commit bodies include the Chinese translation. Contributor-facing docs use English as the source version and ship with a synchronized `_zh.md` page and language switcher.
- Python under `inferencex-e2e/infx/` uses all stable Ruff rules with reviewed exclusions in `inferencex-e2e/infx/ruff.toml`, line length 100, and the Ruff formatter. The Lint job in `.github/workflows/ci.yml` runs whenever Python files change and fails on any finding. Before pushing Python changes, run the [commands in the testing guide](inferencex-e2e/docs/testing.md#python-lint-and-formatting). Fix findings where practical; justified exceptions use inline `# noqa: CODE` rather than file-wide ignores.
- Follow the nearest existing pattern. Python uses typed signatures and strict Pydantic schemas. YAML uses kebab-case fields. Shared benchmark Bash behavior belongs in `benchmark_lib.sh`, with parameters passed through environment variables.
- Follow the nearest existing pattern. Python uses typed signatures and strict Pydantic schemas. YAML uses kebab-case fields. Shared benchmark behavior that runs inside serving containers belongs in `inferencex-e2e/infx/bench/` as stdlib-only, Python 3.10 compatible commands (`python3 -m infx.bench <command>`), with parameters passed through environment variables or flags. Bash entrypoints stay thin shims.

## Bash conventions (mandatory)

These rules apply to active Bash scripts and shell commands embedded in workflows and recipes. Follow them when adding, changing, or reviewing Bash code. Leave deprecated code alone unless explicitly asked to update it.

- **Configuration flows from the caller.** Workflows, master configs, runtime profiles, and launchers explicitly supply configuration to the scripts they invoke. Receiving scripts consume and validate those inputs; they must not silently choose defaults.
- **No fallback defaults for caller-supplied configuration.** Avoid `${VAR:-default}`, `${VAR:=default}`, their colon-free equivalents, and equivalent "if unset, assign a default" logic. A missing input is a caller error and must fail clearly. Pass values such as `false` and `0` explicitly too.
- **Validate every required environment input with `check_env_vars` before use.** Use the shared helper in `inferencex-e2e/benchmarks/benchmark_lib.sh`. Group required inputs near the beginning, after sourcing the helper; validate inputs used only by a particular execution path when entering that path. The helper rejects both missing and empty values. Do not duplicate it or remove its safe handling of unset variables. Callers needing validation without benchmark initialization can source the library with `--validation-only`.
- **Validate every required environment input with `check_env_vars` before use.** Source the shared helper from `inferencex-e2e/benchmarks/check_env.sh`; sourcing it only defines the function. Group required inputs near the beginning, after sourcing the helper; validate inputs used only by a particular execution path when entering that path. The helper rejects both missing and empty values and lists every missing name. Do not duplicate it or remove its safe handling of unset variables.
- **Do not enable nounset.** No `set -u`, `set -o nounset`, combined flags such as `set -euo pipefail`, or `bash -u` invocation flags. Use explicit validation; preserve other intended shell options, for example `set -eo pipefail`.
- **Preserve configuration precedence and forwarding.** Apply caller-owned settings before recipe-specific overrides, and explicitly forward required inputs across container or job boundaries. Do not replace a supported override with an unconditional assignment in the receiving script.
- Preserve deliberate optional-input handling, runtime-derived values, and unset-safe internal-state probes. These are not permission to invent fallback configuration or replace a documented automatic selection with an arbitrary constant.
Expand Down Expand Up @@ -125,7 +125,7 @@ Deleting a test that fails these questions needs no replacement. Do not preserve
- Every change that can affect benchmark performance and every recipe addition or modification requires a new `inferencex-e2e/perf-changelog.yaml` entry. The file is append-only and byte-sensitive. Preserve all existing bytes and separator whitespace, and append only at the tail.
- New `inferencex-e2e/perf-changelog.yaml` entries must be English-only. Do not add Chinese translations or bilingual descriptions; the bilingual documentation and GitHub-content rules do not apply to these entries. Leave historical entries unchanged.
- Multi-node srt-slurm changes update the recipe YAML and matching master config together. For image bumps, `model.container` must equal `image`.
- Every speculative fixed-sequence benchmark renders prompts with the chat template: single-node srt-slurm recipes that speculate set `benchmark.env.USE_CHAT_TEMPLATE: "true"` (enforced by `inferencex-e2e/infx/srt_slurm/single_node.py::validate_recipe`), which `srt_fixed_sequence.sh` turns into `--use-chat-template` for `run_benchmark_serving`.
- Every speculative fixed-sequence benchmark renders prompts with the chat template: single-node srt-slurm recipes that speculate set `benchmark.env.USE_CHAT_TEMPLATE: "true"` (enforced by `inferencex-e2e/infx/srt_slurm/single_node.py::validate_recipe`), which `python3 -m infx.bench fixed-seq` (run by `srt_fixed_sequence.sh`) turns into `--use-chat-template` for the benchmark client.
- Benchmarks create no new directories under `/workspace`. Root containers must not leave root-owned files in shared AMD runner workspaces.
- Generated configuration is not runtime proof. Run the narrowest local check, then the applicable smoke, sweep, or eval procedure from [`inferencex-e2e/docs/procedures.md`](inferencex-e2e/docs/procedures.md).

Expand Down
Loading
Loading