Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
1411 commits
Select commit Hold shift + click to select a range
b28a170
Support training-only distributed runtimes
FurtherAI Aug 11, 2026
1077c74
Defer trajectory queue shrink during collection
FurtherAI Aug 11, 2026
f91c332
Bound full-model trainer shutdown
FurtherAI Aug 11, 2026
7375005
Preserve live KV across policy updates
FurtherAI Aug 11, 2026
2f29744
Exercise policy transition hash patch
FurtherAI Aug 11, 2026
b845dd7
Allow full-model service teardown to finish
FurtherAI Aug 11, 2026
ed6f555
Rebase policy cache hashes after preemption
FurtherAI Aug 11, 2026
4d3fcaf
Fix trajectory queue lease backpressure
FurtherAI Aug 11, 2026
1ae05bd
Record final cutover acceptance blockers
FurtherAI Aug 11, 2026
1752bc4
Stabilize throughput load evidence
FurtherAI Aug 11, 2026
8465a3f
Make throughput evidence statistically symmetric
FurtherAI Aug 11, 2026
a51ef7b
Stabilize throughput measurement windows
FurtherAI Aug 11, 2026
fb0df2f
Aggregate stable throughput evidence
FurtherAI Aug 11, 2026
82730d3
Update stable-window rejection assertion
FurtherAI Aug 11, 2026
fad26e0
Derive stable settings from execution history
FurtherAI Aug 11, 2026
1dde72e
Handle multi-window throughput evidence
FurtherAI Aug 11, 2026
bb26624
Stabilize matched throughput parity sampling
FurtherAI Aug 11, 2026
2dc328a
Align matched throughput test steps
FurtherAI Aug 11, 2026
b1172de
Bound throughput evidence to measured steps
FurtherAI Aug 11, 2026
7d18a0e
Bound slow workflow acceptance budgets
FurtherAI Aug 11, 2026
1860748
Align DSV4 trainability resource contract
FurtherAI Aug 11, 2026
59d6e62
Install final B300 throughput calibrations
FurtherAI Aug 11, 2026
135e97c
Stabilize real-path parity inputs
FurtherAI Aug 11, 2026
1a5b494
Fix parity rollout inputs deterministically
FurtherAI Aug 11, 2026
0a3ea45
Fix deterministic parity decoding
FurtherAI Aug 11, 2026
917c290
Restore representative train inference retries
FurtherAI Aug 12, 2026
a797373
Trim redundant correctness topologies
FurtherAI Aug 12, 2026
0779671
test: schedule reusable model workflow sessions
FurtherAI Aug 12, 2026
c708fe2
test: stabilize GPT OSS mismatch KL gate
FurtherAI Aug 12, 2026
887cda6
test: shorten steady state throughput validation
FurtherAI Aug 12, 2026
a0f4d44
test: schedule workflows across cluster hosts
FurtherAI Aug 12, 2026
1c1901f
test: pin remote workflow runtime environment
FurtherAI Aug 12, 2026
15e1f51
test: capture matched inputs from packed generations
FurtherAI Aug 12, 2026
d4f7cd9
test: schedule default workflow resources
FurtherAI Aug 12, 2026
22a0e82
test: overlap matched capture under source leases
FurtherAI Aug 12, 2026
5be2fc9
test: share production width functional fixtures
FurtherAI Aug 12, 2026
91ab236
test: account for functional MTP layers
FurtherAI Aug 12, 2026
3c94660
fix: pin vllm runtime build tools
FurtherAI Aug 12, 2026
edba6e7
Optimize sliding context masks
FurtherAI Aug 12, 2026
3ed139a
test: measure throughput after shape warmup
FurtherAI Aug 12, 2026
7bcb7ee
test: cover two post-warmup windows
FurtherAI Aug 12, 2026
cc3c18b
test: update throughput lag percentile fixture
FurtherAI Aug 12, 2026
c04f4c3
test: retain mean activation lag failure case
FurtherAI Aug 12, 2026
4183558
Measure trainer gaps on the CUDA timeline
FurtherAI Aug 12, 2026
28c03ab
Optimize DSV4 prefix compression layouts
FurtherAI Aug 12, 2026
84bc4a1
Reuse functional LoRA vLLM sessions safely
FurtherAI Aug 12, 2026
c8f7f4e
Reuse pinned publication buffers by snapshot slot
FurtherAI Aug 12, 2026
561d9a8
Optimize oracle comparison lifecycle
FurtherAI Aug 12, 2026
2d6f1de
Reuse base Megatron runtime for parity checks
FurtherAI Aug 12, 2026
aeafe31
Avoid repacking fused DSV4 LoRA tensors
FurtherAI Aug 12, 2026
a6d5729
Reuse functional validation runtimes across stages
FurtherAI Aug 12, 2026
564245c
Retry only transient train inference failures
FurtherAI Aug 12, 2026
8a2d630
Handle empty vLLM unload responses
FurtherAI Aug 12, 2026
7f74941
Allocate disjoint shared workflow roles
FurtherAI Aug 12, 2026
bd80bb4
Fix concurrent workflow fixture execution
FurtherAI Aug 12, 2026
567c3e1
Use native DSV4 serving kernels
FurtherAI Aug 12, 2026
cbad7fb
Preserve reused inference runtime contracts
FurtherAI Aug 12, 2026
498cfcd
Fit wide-head FlexAttention backward on Blackwell
FurtherAI Aug 12, 2026
2add349
Handle local DSV4 dummy LoRA experts
FurtherAI Aug 12, 2026
4acfa92
Keep replay-built MoE runtimes stage-local
FurtherAI Aug 12, 2026
5d44ffc
Preserve exact sliding and local expert layouts
FurtherAI Aug 12, 2026
9dfab04
Reduce throughput workflow measurement horizon
FurtherAI Aug 12, 2026
a2cc8f2
Align throughput test with shorter horizon
FurtherAI Aug 12, 2026
8910104
Keep trainability on coherent pretrained models
FurtherAI Aug 12, 2026
3cc563c
Preserve exact workflow session topology
FurtherAI Aug 12, 2026
b0ca2c5
Keep throughput calibrations source-stable
FurtherAI Aug 12, 2026
e7d86b7
Align workflow tests with final protocol
FurtherAI Aug 12, 2026
0bfc300
Keep throughput fixture on measured suffix
FurtherAI Aug 12, 2026
f916bb7
Fence throughput calibrations by runtime contract
FurtherAI Aug 12, 2026
d890981
Refresh B300 throughput contract identities
FurtherAI Aug 12, 2026
9c828e0
Amortize workflow worker imports per host
FurtherAI Aug 12, 2026
c0d2cbf
Stabilize finite-horizon throughput gap sampling
FurtherAI Aug 12, 2026
3e78245
Harden workflow forkserver cleanup
FurtherAI Aug 12, 2026
b103f71
Schedule correctness references independently
FurtherAI Aug 12, 2026
0a3da83
Record exact correctness phase topology
FurtherAI Aug 12, 2026
469c6be
Overlap Megatron trainer preparation
FurtherAI Aug 12, 2026
fff22b9
Overlap workflow preload with preparation
FurtherAI Aug 12, 2026
6aea4ef
Allow variable throughput role counts
FurtherAI Aug 12, 2026
85be028
Use a shorter DSV4 throughput sequence
FurtherAI Aug 12, 2026
b02c59d
Allow trainer startup during vLLM launch
FurtherAI Aug 12, 2026
bc1343e
Remove merged rollout serving
FurtherAI Aug 12, 2026
2e649c8
Restore numerical mismatch retries
FurtherAI Aug 12, 2026
4ab90c3
Align mismatch retry invariant
FurtherAI Aug 12, 2026
560aded
Scale DSV4 throughput work to 32K
FurtherAI Aug 12, 2026
1236714
Reuse resident trainer for functional validation
FurtherAI Aug 12, 2026
85a4e28
Create trainability artifact directory
FurtherAI Aug 12, 2026
b071fee
Prefetch trainer with external inference
FurtherAI Aug 12, 2026
acd2e06
Preserve trainer run identity in actors
FurtherAI Aug 12, 2026
effec27
Handle EP-local dummy MoE adapters generically
FurtherAI Aug 12, 2026
987dce1
Capture MoE routes in resident functional sessions
FurtherAI Aug 12, 2026
77dd8a2
Overlap functional trainer and vLLM startup
FurtherAI Aug 12, 2026
6748204
Fence functional rollouts on serving readiness
FurtherAI Aug 12, 2026
62b5091
Schedule workflow resources without fragmentation
FurtherAI Aug 12, 2026
0cb3e0e
Cache trainer Dynamo graphs per rank
FurtherAI Aug 12, 2026
d5c7a05
Enable Dynamo caching after trainer imports
FurtherAI Aug 12, 2026
69b485c
Support aliased methods in trainer compile cache
FurtherAI Aug 12, 2026
9ccf3d4
Preserve non-strict Dynamo cache bypass
FurtherAI Aug 12, 2026
7261458
Cache late Dynamo identity guards
FurtherAI Aug 12, 2026
e834923
Skip transient CUDA events in trainer cache
FurtherAI Aug 13, 2026
b819e95
Give lazy Qwen bridges stable identities
FurtherAI Aug 13, 2026
29cf697
Key trainer caches by resolved compile plan
FurtherAI Aug 13, 2026
cece4be
Serialize guarded module context variables
FurtherAI Aug 13, 2026
5cf75ab
Reconstruct guarded autograd nodes
FurtherAI Aug 13, 2026
8c61dad
Prune unguarded HybridEP compile state
FurtherAI Aug 13, 2026
56684e5
Register reused Dynamo resume code
FurtherAI Aug 13, 2026
f656f92
Bundle no-op autograd graphs for precompile
FurtherAI Aug 13, 2026
f2bd018
Serialize trainer compile cache sources
FurtherAI Aug 13, 2026
37d4e76
Serialize bundled trainer code objects
FurtherAI Aug 13, 2026
fba84fb
Allow isolated workflow compiler caches
FurtherAI Aug 13, 2026
75c3766
Integrate qualified DSV4 throughput geometry
FurtherAI Aug 13, 2026
c191d05
Scope trainer precompile serialization
FurtherAI Aug 13, 2026
5653399
Match resident vLLM startup deadline
FurtherAI Aug 13, 2026
d71573c
Serialize dynamic Dynamo resume functions
FurtherAI Aug 13, 2026
b9d40f5
Require resident trainability continuation
FurtherAI Aug 13, 2026
0b1be17
Preserve dtype identities in compile guards
FurtherAI Aug 13, 2026
450b52a
Fail closed on workflow operation failures
FurtherAI Aug 13, 2026
ec36a5b
Complete resident compiler cache artifacts
FurtherAI Aug 13, 2026
e72388e
Preserve resident phase evidence after failures
FurtherAI Aug 13, 2026
8994fb1
Namespace regional compile cache entries
FurtherAI Aug 13, 2026
3f1541a
Make HybridEP handles cache serializable
FurtherAI Aug 13, 2026
fd46450
Scope HybridEP handle cache serialization
FurtherAI Aug 13, 2026
2cf319a
Correct DSV4 throughput update geometry
FurtherAI Aug 13, 2026
1de9e47
Serialize HybridEP handles in guard caches
FurtherAI Aug 13, 2026
a9e6b3d
Preserve regional precompile entries
FurtherAI Aug 13, 2026
3d6ff6b
Resolve serialized nested compile frames
FurtherAI Aug 13, 2026
f6d00b9
Classify generated resume frames from origin metadata
FurtherAI Aug 13, 2026
5a83c39
Reuse native compiler cache across trainer runs
FurtherAI Aug 13, 2026
ab57f7a
Cache import-time Megatron compile wrappers
FurtherAI Aug 13, 2026
167ffac
Support compiled callable classes in trainer cache
FurtherAI Aug 13, 2026
8c8356e
Hydrate import-time compiler guards safely
FurtherAI Aug 13, 2026
0f2a916
Canonicalize hydrated nested compile code
FurtherAI Aug 13, 2026
5db8f3d
Canonicalize guarded compile bytecode
FurtherAI Aug 13, 2026
269322b
Handle transformed compile constants safely
FurtherAI Aug 13, 2026
44a436b
Resolve nested compile code across packages
FurtherAI Aug 13, 2026
654e0a0
Canonicalize generated compile frames
FurtherAI Aug 13, 2026
a8b6ae6
Intern hydrated resume code identities
FurtherAI Aug 13, 2026
7fcda63
Keep compile identity interning fast
FurtherAI Aug 13, 2026
f5b56ca
Drop trainer precompile compatibility
FurtherAI Aug 13, 2026
7b146d4
Make throughput groups per step global
FurtherAI Aug 13, 2026
87ee200
Retry inconclusive throughput load once
FurtherAI Aug 13, 2026
9b4b801
Balance DSV4 throughput decode load
FurtherAI Aug 13, 2026
e31b8c9
Bound DSV4 inference concurrency
FurtherAI Aug 13, 2026
d304ff1
Balance DSV4 inference capacity
FurtherAI Aug 13, 2026
151def4
Tune DSV4 inference pressure
FurtherAI Aug 13, 2026
f8326c0
Calibrate DSV4 two-group throughput
FurtherAI Aug 13, 2026
d82490e
Give DSV4 pressure margin
FurtherAI Aug 13, 2026
a2ed9f3
Use representative policy for resident parity
FurtherAI Aug 13, 2026
16c7aba
Calibrate DSV4 balanced throughput load
FurtherAI Aug 13, 2026
eadfe41
Update mismatch policy fence fixture
FurtherAI Aug 13, 2026
6febaea
Calibrate Qwen resident parity variance
FurtherAI Aug 13, 2026
03b0fa1
Align DSV4 HF norm output precision
FurtherAI Aug 13, 2026
3e94cef
Use DP-aware resident score scheduling
FurtherAI Aug 13, 2026
488d0d5
Bound repeated HF route score noise
FurtherAI Aug 13, 2026
146a025
Require utilized throughput workloads
FurtherAI Aug 13, 2026
a6f44a1
Fix throughput utilization test fixture
FurtherAI Aug 13, 2026
88d0ee7
Cover utilization gate combinations
FurtherAI Aug 13, 2026
81b811a
Attribute trainer schedule gap phases
FurtherAI Aug 13, 2026
89b9c12
Reuse repeated prefixes in throughput workloads
FurtherAI Aug 13, 2026
92cf734
Trace pipeline trainer critical path phases
FurtherAI Aug 13, 2026
d9f901e
Decompose trainer schedule gaps
FurtherAI Aug 13, 2026
1ffaf35
Remove nonverbose trainer progress overhead
FurtherAI Aug 13, 2026
8167f90
Allow homogeneous throughput depth expansion
FurtherAI Aug 13, 2026
83b80a1
Stabilize throughput workflow acceptance
FurtherAI Aug 13, 2026
946802f
Retry transient throughput failures
FurtherAI Aug 13, 2026
51d6ce8
Isolate throughput runtime provenance
FurtherAI Aug 13, 2026
3b451e9
Balance DSV4 throughput qualification
FurtherAI Aug 14, 2026
8d76cad
Stabilize DSV4 throughput workload
FurtherAI Aug 14, 2026
2796992
Load DSV4 inference reliably
FurtherAI Aug 14, 2026
e9a5e4d
Handle appended functional fixture layers
FurtherAI Aug 14, 2026
3fd9c08
Restore valid Blackwell wide-head Flex tiles
FurtherAI Aug 14, 2026
42cf811
Fence functional serving baseline before registration
FurtherAI Aug 14, 2026
d958557
Preserve workflow runtime placement invariants
FurtherAI Aug 14, 2026
4c19948
Serialize incompatible base model workflow stages
FurtherAI Aug 14, 2026
ae37ba3
Apply workflow fixture environment before imports
FurtherAI Aug 14, 2026
7a001c2
Size Gemma resident functional context
FurtherAI Aug 14, 2026
5af20f2
Align resident length sequence limits
FurtherAI Aug 14, 2026
6ac0a7b
Align resident scoring sequence length
FurtherAI Aug 14, 2026
a0bda28
Scope wide-head Flex workaround to Blackwell
FurtherAI Aug 14, 2026
2dce744
Select settled throughput execution windows
FurtherAI Aug 14, 2026
054627b
Exercise measured throughput row validation
FurtherAI Aug 14, 2026
17c4501
Stabilize measured workflow gates
FurtherAI Aug 14, 2026
68a62ae
Keep adapter publication off the trainer path
FurtherAI Aug 14, 2026
b652f5f
Calibrate learned-policy parity gates
FurtherAI Aug 14, 2026
cce36cb
Record learned-policy calibration evidence
FurtherAI Aug 14, 2026
b61bab2
Stabilize functional workflow gates
FurtherAI Aug 14, 2026
ac51473
Own packing rendezvous per worker
FurtherAI Aug 14, 2026
7eb93ab
Stabilize GLM length movement
FurtherAI Aug 14, 2026
cc8c6a3
Allow router-free virtual pipeline chunks
FurtherAI Aug 14, 2026
b67893f
Preplan CP batches from packing lookahead
FurtherAI Aug 14, 2026
17d62f6
Overlap controller finalization with training
FurtherAI Aug 14, 2026
58bdfaa
Preserve source times for async pipeline metrics
FurtherAI Aug 14, 2026
2bd401e
Fence controller finalization behind trainer dispatch
FurtherAI Aug 14, 2026
d31ad2a
Stabilize Python GC after trainer compilation
FurtherAI Aug 14, 2026
7e3549c
Allow 230ms workflow trainer gap p50
FurtherAI Aug 14, 2026
fd2a1b5
Preserve resident trainability scenario order
FurtherAI Aug 14, 2026
451143d
Give GPT OSS trainability a reliable horizon
FurtherAI Aug 14, 2026
872e71f
Balance persistent workflow affinities
FurtherAI Aug 14, 2026
98d2bcf
Stabilize workflow host admission
FurtherAI Aug 14, 2026
83ef9ec
Calibrate Gemma MoE learned-policy parity
FurtherAI Aug 15, 2026
265cc8d
Align DSV4 inference wave with training batch
FurtherAI Aug 15, 2026
ea7f51b
Sustain DSV4 throughput wave pressure
FurtherAI Aug 15, 2026
e9fbc31
Calibrate balanced DSV4 throughput workload
FurtherAI Aug 15, 2026
e0dbc26
Disambiguate parity choices by full token path
FurtherAI Aug 15, 2026
9c4e3dd
Isolate direct parity vLLM environments
FurtherAI Aug 15, 2026
349e25b
Match duplicate parity leaves by source scores
FurtherAI Aug 15, 2026
474c232
Calibrate Gemma parity to execution variance
FurtherAI Aug 15, 2026
adcdded
Drain stale packing before throughput evidence
FurtherAI Aug 15, 2026
c1de2e0
Remove stale DSV4 host exclusivity
FurtherAI Aug 15, 2026
3a807db
Sustain DSV4 inference pressure
FurtherAI Aug 15, 2026
34896d3
Fully balance DSV4 throughput load
FurtherAI Aug 15, 2026
4a4bbab
Remove single-rank TCP rendezvous races
FurtherAI Aug 15, 2026
252d7f5
Remove obsolete diagnostics and compatibility helpers
FurtherAI Aug 15, 2026
219f7f6
Promote distributed contracts and drop debug artifacts
FurtherAI Aug 15, 2026
e96b531
Consolidate distributed lifecycle helpers
FurtherAI Aug 15, 2026
cef32e1
Remove residual debug output
FurtherAI Aug 15, 2026
f478ab2
Root-cause and tighten train inference parity gates
FurtherAI Aug 15, 2026
7cf4392
Test architecture-scoped runtime caches
FurtherAI Aug 15, 2026
480124b
Clarify Gemma parity calibration evidence
FurtherAI Aug 15, 2026
67567e3
Merge remote-tracking branch 'origin/main' into austin/monarch_multin…
FurtherAI Aug 17, 2026
590627f
Retry inconclusive throughput evidence once
FurtherAI Aug 17, 2026
756b360
Refresh B300 throughput calibration identities
FurtherAI Aug 17, 2026
fa7aa21
Skip numeric trajectory interning walks
FurtherAI Aug 17, 2026
63861de
Avoid transient packing string interning
FurtherAI Aug 17, 2026
138b06d
Skip interning transient rollout results
FurtherAI Aug 17, 2026
1fdfd14
Calibrate Llama train-inference parity envelope
FurtherAI Aug 17, 2026
efc9e7e
Calibrate Qwen 3.5 packing-order parity variance
FurtherAI Aug 17, 2026
494d119
Record complete post-merge workflow pass
FurtherAI Aug 17, 2026
a155ac6
Make trajectory memory compaction explicit
FurtherAI Aug 17, 2026
1904bf7
Simplify model workflow orchestration
FurtherAI Aug 18, 2026
a02ac70
Remove stale workflow resource helpers
FurtherAI Aug 18, 2026
2d9ad72
Trim workflow scheduler contract tests
FurtherAI Aug 18, 2026
87efc6d
Focus oracle harness invariant coverage
FurtherAI Aug 18, 2026
ca64556
Focus train inference invariant coverage
FurtherAI Aug 18, 2026
e3b547c
Reduce workflow assertion scaffolding
FurtherAI Aug 18, 2026
4837b0a
Remove optional yes-no workflow coverage
FurtherAI Aug 18, 2026
14c534d
Consolidate fast metrics contract tests
FurtherAI Aug 18, 2026
3016691
Share runtime policy request fixtures
FurtherAI Aug 18, 2026
ce487ed
Accept source scripts in Monarch launcher
FurtherAI Aug 18, 2026
fe978c1
Fix workstation multi-node launch portability
FurtherAI Aug 18, 2026
cfa4b10
Pin uv in cluster setup
FurtherAI Aug 18, 2026
2e49ec6
Add typed multi-node launch context
FurtherAI Aug 18, 2026
e9be341
Harden multi-node cluster setup
FurtherAI Aug 18, 2026
bfb2370
Package managed Megatron runtimes for releases
FurtherAI Aug 19, 2026
a1f5158
Build package artifacts only for release PRs
FurtherAI Aug 19, 2026
364b477
Merge remote-tracking branch 'origin/main' into austin/monarch_multin…
FurtherAI Aug 19, 2026
3cfe6a2
Fix managed Megatron CI setup
FurtherAI Aug 19, 2026
74ea2bd
Fingerprint managed runtime cache inputs
FurtherAI Aug 19, 2026
72e353d
Exercise managed Megatron runtime in CI
FurtherAI Aug 19, 2026
5d026ef
Pin TileLang-compatible TVM FFI runtime
FurtherAI Aug 19, 2026
e932bff
Pin the validated FLA kernel contract
FurtherAI Aug 19, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
144 changes: 140 additions & 4 deletions .github/workflows/package-install.yml
Original file line number Diff line number Diff line change
Expand Up @@ -2,15 +2,14 @@ name: Package Install

on:
pull_request:
push:
branches: [main]
workflow_dispatch:

permissions:
contents: read

jobs:
install-smoke-test:
if: startsWith(github.head_ref, 'release/')
runs-on: ubuntu-latest

steps:
Expand All @@ -26,7 +25,7 @@ jobs:
curl -LsSf https://astral.sh/uv/install.sh | sh
echo "$HOME/.cargo/bin" >> "$GITHUB_PATH"

- name: Build wheel
- name: Build managed-runtime wheel
run: python scripts/build_package.py --wheel

- name: Smoke test base wheel import
Expand All @@ -42,7 +41,7 @@ jobs:
cd "$project_dir"
uv init --name art-base-install-smoke --python 3.12 --bare
uv add "openpipe-art @ file://${wheel_path}"
uv run python -c "import importlib.util; assert importlib.util.find_spec('numpy') is None; assert importlib.util.find_spec('torch') is None; import art; from art.pipeline_trainer import PipelineTrainer; print(art.__name__, PipelineTrainer.__name__)"
uv run python -c "import importlib.util; assert importlib.util.find_spec('numpy') is not None; assert importlib.util.find_spec('torch') is None; import art; from art import ServerlessBackend; from art.pipeline_trainer import PipelineTrainer; print(art.__name__, ServerlessBackend.__name__, PipelineTrainer.__name__)"
uv add "weave==0.52.41"
uv run python -c "from weave.trace.settings import override_settings; import art"

Expand All @@ -60,3 +59,140 @@ jobs:
uv init --name art-install-smoke --python 3.12 --bare
uv add "openpipe-art[backend] @ file://${wheel_path}"
uv sync

- name: Resolve published Megatron profiles
run: |
wheel_path="$(python - <<'PY'
from pathlib import Path

print(next(Path("dist").glob("openpipe_art-*.whl")).resolve())
PY
)"

for profile in megatron megatron-cu130; do
case "$profile" in
megatron) index=cu128 ;;
megatron-cu130) index=cu130 ;;
esac
project_dir="$(mktemp -d)"
cd "$project_dir"
uv init --name "art-${profile}-resolve" --python 3.12 --bare
uv add --no-sync \
--index-strategy unsafe-best-match \
--extra-index-url "https://download.pytorch.org/whl/${index}" \
"openpipe-art[${profile}] @ file://${wheel_path}"
done

- name: Resolve published Tinker profile
run: |
wheel_path="$(python - <<'PY'
from pathlib import Path

print(next(Path("dist").glob("openpipe_art-*.whl")).resolve())
PY
)"

project_dir="$(mktemp -d)"
cd "$project_dir"
uv init --name art-tinker-resolve --python 3.12 --bare
uv add --no-sync "openpipe-art[tinker] @ file://${wheel_path}"

- name: Smoke test distributed wheel surface on CPU
env:
ART_VLLM_RUNTIME_CACHE_DIR: ${{ runner.temp }}/art-vllm-runtime-cache
run: |
wheel_path="$(python - <<'PY'
from pathlib import Path

print(next(Path("dist").glob("openpipe_art-*.whl")).resolve())
PY
)"

project_dir="$(mktemp -d)"
cd "$project_dir"
uv init --name art-distributed-install-smoke --python 3.12 --bare
uv venv --python 3.12
uv pip install --python .venv/bin/python \
"openpipe-art @ file://${wheel_path}"
uv pip install --python .venv/bin/python \
--index-url https://download.pytorch.org/whl/cpu \
"torch==2.11.0"
uv pip install --python .venv/bin/python \
"msgspec>=0.21.0" \
"torchmonarch==0.6.0" \
"transformers==5.12.1"

.venv/bin/python - <<'PY'
import sys
from importlib.metadata import metadata
from pathlib import Path
import subprocess
import tempfile
import venv

import art
import art.distributed as distributed

assert Path(art.__file__).resolve().is_relative_to(Path(sys.prefix))
assert "art.distributed.art_runtime" not in sys.modules
from art.distributed import ( # noqa: E402
ArtRuntime,
ClusterSpec,
HostSpec,
InstalledAsyncCallable,
NcclTransportSpec,
PackingRequest,
compile_topology,
)

assert all(
value is not None
for value in (
ArtRuntime,
ClusterSpec,
HostSpec,
InstalledAsyncCallable,
NcclTransportSpec,
PackingRequest,
compile_topology,
)
)
assert "monarch" not in sys.modules
assert "PackingRequest" in distributed.__all__
assert "NcclTransportSpec" in distributed.__all__
assert {"distributed", "megatron"} <= set(
metadata("openpipe-art").get_all("Provides-Extra") or ()
)
bundle = Path(art.__file__).with_name("_megatron_runtime")
assert (bundle / "manifest.json").is_file()
assert (bundle / "uv.lock").is_file()

from art.megatron.runtime.managed import _copy_art

with tempfile.TemporaryDirectory() as temp_dir:
managed = Path(temp_dir) / "managed"
venv.EnvBuilder(with_pip=False).create(managed)
managed_python = managed / "bin" / "python"
_copy_art(managed_python)
managed_site = Path(
subprocess.check_output(
[
str(managed_python),
"-c",
"import sysconfig; print(sysconfig.get_paths()['purelib'])",
],
text=True,
).strip()
)
copied_bundle = managed_site / "art" / "_megatron_runtime"
assert (copied_bundle / "manifest.json").is_file()
assert (copied_bundle / "nixl-de8115ca.tar.gz").is_file()
assert not (managed_site / "art" / "_vllm_runtime").exists()
PY

PYTHONPATH="$GITHUB_WORKSPACE/examples/multinode" timeout 150s \
.venv/bin/art-monarch local \
--program program:main \
--port 0 \
--startup-timeout 90
test ! -e "$ART_VLLM_RUNTIME_CACHE_DIR"
38 changes: 14 additions & 24 deletions .github/workflows/prek.yml
Original file line number Diff line number Diff line change
Expand Up @@ -13,8 +13,7 @@ env:
CI_PYTHON_MM: "3.12"
CI_UV_CACHE_RELEASE_TAG: "prek-uv-cache"
CI_UV_CACHE_ASSET_PREFIX: "prek-uv-cache"
CI_APEX_PARALLEL_BUILD: "8"
CI_APEX_NVCC_THREADS: "1"
CI_BUILD_JOBS: "8"
CI_UV_BUILD_SLOTS: "2"
UV_CACHE_DIR: "/root/.cache/uv"
UV_LINK_MODE: "copy"
Expand All @@ -36,11 +35,11 @@ jobs:
fp="$(python3 scripts/ci/compute_uv_fingerprint.py \
--pyproject pyproject.toml \
--uv-lock uv.lock \
--megatron-pyproject megatron_runtime/pyproject.toml \
--megatron-uv-lock megatron_runtime/uv.lock \
--base-image "${CI_BASE_IMAGE}" \
--python-mm "${CI_PYTHON_MM}" \
--torch-cuda-arch-list "${TORCH_CUDA_ARCH_LIST}" \
--ci-apex-parallel-build "${CI_APEX_PARALLEL_BUILD}" \
--ci-apex-nvcc-threads "${CI_APEX_NVCC_THREADS}")"
--torch-cuda-arch-list "${TORCH_CUDA_ARCH_LIST}")"
echo "fingerprint=${fp}" >> "${GITHUB_OUTPUT}"
echo "Expected uv cache fingerprint: ${fp}"

Expand Down Expand Up @@ -187,14 +186,7 @@ jobs:

- name: Install Megatron dependencies
run: |
original_pyproject="$(mktemp)"
cp pyproject.toml "${original_pyproject}"
cleanup() {
mv "${original_pyproject}" pyproject.toml
}
trap cleanup EXIT

cudnn_path="${GITHUB_WORKSPACE}/.venv/lib/python${CI_PYTHON_MM}/site-packages/nvidia/cudnn"
cudnn_path="${GITHUB_WORKSPACE}/megatron_runtime/.venv/lib/python${CI_PYTHON_MM}/site-packages/nvidia/cudnn"
export CUDNN_PATH="${cudnn_path}"
export CUDNN_HOME="${cudnn_path}"
export CUDNN_INCLUDE_PATH="${cudnn_path}/include"
Expand All @@ -203,16 +195,13 @@ jobs:
export LIBRARY_PATH="${CUDNN_LIBRARY_PATH}${LIBRARY_PATH:+:${LIBRARY_PATH}}"
export LD_LIBRARY_PATH="${CUDNN_LIBRARY_PATH}${LD_LIBRARY_PATH:+:${LD_LIBRARY_PATH}}"
export UV_CONCURRENT_BUILDS="${CI_UV_BUILD_SLOTS}"
export CMAKE_BUILD_PARALLEL_LEVEL="${CI_APEX_PARALLEL_BUILD}"
export MAX_JOBS="${CI_APEX_PARALLEL_BUILD}"
export NINJAFLAGS="-j${CI_APEX_PARALLEL_BUILD}"
python3 scripts/ci/apply_ci_uv_build_overrides.py \
--pyproject pyproject.toml \
--apex-parallel-build "${CI_APEX_PARALLEL_BUILD}" \
--apex-nvcc-threads "${CI_APEX_NVCC_THREADS}"
echo "CI uv build overrides: APEX_PARALLEL_BUILD=${CI_APEX_PARALLEL_BUILD}, NVCC_APPEND_FLAGS=--threads ${CI_APEX_NVCC_THREADS}, UV_CONCURRENT_BUILDS=${CI_UV_BUILD_SLOTS}"
export CMAKE_BUILD_PARALLEL_LEVEL="${CI_BUILD_JOBS}"
export MAX_JOBS="${CI_BUILD_JOBS}"
export NINJAFLAGS="-j${CI_BUILD_JOBS}"
uv --version
uv sync --extra megatron --extra langgraph --extra plotting --group dev --frozen --python "${CI_PYTHON_MM}"
uv sync --extra langgraph --extra plotting --group dev --frozen --python "${CI_PYTHON_MM}"
uv sync --project megatron_runtime --extra cuda12 --group test --frozen --python "${CI_PYTHON_MM}"
uv pip install --python megatron_runtime/.venv/bin/python --no-deps --editable .

- name: Run prek hooks (lint, format, typecheck, uv.lock)
run: |
Expand All @@ -223,9 +212,10 @@ jobs:

- name: Run Megatron lightweight tests
run: |
uv run --no-sync python -c "import megatron.core.packed_seq_params"
uv run --no-sync pytest --nbval --current-env --tb=short \
megatron_runtime/.venv/bin/python -c "import megatron.core.packed_seq_params"
megatron_runtime/.venv/bin/python -m pytest --nbval --current-env --tb=short \
tests/unit/test_megatron_reference_logprobs.py \
tests/unit/test_preprocessing_tokenize.py::test_gemma4_normalizes_json_tool_arguments_for_mapping_template \
tests/unit/test_moe_routing_replay.py \
tests/unit/test_moe_routing_real_path.py \
tests/unit/test_pipeline_trainer_local_backend.py \
Expand Down
10 changes: 6 additions & 4 deletions .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -106,8 +106,10 @@ jobs:
str(runtime_python),
"-c",
"import torch, vllm; "
"from art_vllm_runtime.fast_metrics import FastMetricsSharedWriter; "
"from art_vllm_runtime.policy_spans import "
"_patch_lora_update_coordinator; "
"writer = FastMetricsSharedWriter(); writer.close(); "
"_patch_lora_update_coordinator(); "
"print('runtime compatibility ok')",
],
Expand Down Expand Up @@ -137,22 +139,22 @@ jobs:
git tag v${{ needs.build-package.outputs.version }}
git push origin v${{ needs.build-package.outputs.version }}

- name: Publish draft release
- name: Upload assets to draft release
env:
GH_TOKEN: ${{ github.token }}
run: |
if gh release view v${{ needs.build-package.outputs.version }} --json isDraft | jq -r '.isDraft' | grep -q true; then
gh release edit v${{ needs.build-package.outputs.version }} --draft=false
gh release upload v${{ needs.build-package.outputs.version }} dist/*
else
echo "::error::No draft release found for v${{ needs.build-package.outputs.version }}"
exit 1
fi

- name: Upload assets to release
- name: Publish release
env:
GH_TOKEN: ${{ github.token }}
run: |
gh release upload v${{ needs.build-package.outputs.version }} dist/*
gh release edit v${{ needs.build-package.outputs.version }} --draft=false

- name: Publish to PyPI
uses: pypa/gh-action-pypi-publish@release/v1
Expand Down
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -24,3 +24,4 @@ replays/
!/src/art/wandb/**
/src/art/wandb/__pycache__/
scratch/
/progress_log.md
4 changes: 4 additions & 0 deletions .skyignore
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,9 @@ __pycache__/
.art/
# .env
.venv/
.ruff_cache/
.pytest_cache/
scratch/
grpo_trainer_lora_model/
logs/
shared_cache.db
Expand All @@ -13,5 +16,6 @@ dist/
dev/art-e/data/
replays/
/trajectories/
/progress_log.md
.DS_Store
# .local/
21 changes: 21 additions & 0 deletions THIRD-PARTY-NOTICES
Original file line number Diff line number Diff line change
Expand Up @@ -51,3 +51,24 @@ This project vendors a modified HybridEP runtime derived from DeepEP:
The applicable license text is retained at:

- src/art/megatron/_hybrid_ep/LICENSE

Megatron release wheels also embed pinned source archives used only as headers
when building ART's HybridEP extension:

- NVIDIA NIXL 1.3.2 (commit de8115ca97d3f8fb63a4988e9b4d4a038b2e0f72),
Apache License 2.0, https://github.com/ai-dynamo/nixl
- OpenUCX 1.21.0, BSD 3-Clause License,
https://github.com/openucx/ucx

The complete license and notice files remain in their respective bundled source
archives. ART links HybridEP against the matching NIXL and UCX libraries shipped
by the official `nixl-cu12` or `nixl-cu13` wheel; it does not redistribute those
libraries itself.

On first cross-host HybridEP use, ART may download the checksum-pinned etcd
3.5.33 executable from its official GitHub release. etcd is licensed under the
Apache License 2.0: https://github.com/etcd-io/etcd

The CUDA 12 managed Megatron environment installs NVIDIA Apex from its pinned
25.09 source tag under the BSD 3-Clause License:
https://github.com/NVIDIA/apex
22 changes: 12 additions & 10 deletions dev/sft/sft-from-file.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,8 +9,10 @@


async def main():
backend = MegatronBackend()

art.init_megatron_runtime_config(
topology=art.MegatronTopologyConfig(),
packed_sequence_length=4096,
)
model_name = "run-" + "".join(
random.choices("abcdefghijklmnopqrstuvwxyz0123456789", k=8)
)
Expand All @@ -20,14 +22,14 @@ async def main():
project="sft-from-file",
base_model="Qwen/Qwen3.6-35B-A3B",
)
await model.register(backend)

await train_sft_from_file(
model=model,
file_path="dev/sft/dataset.jsonl",
epochs=1,
peak_lr=2e-4,
)
async with MegatronBackend() as backend:
await model.register(backend)
await train_sft_from_file(
model=model,
file_path="dev/sft/dataset.jsonl",
epochs=1,
peak_lr=2e-4,
)

print("Training complete!")

Expand Down
Loading
Loading