Repository navigation
tests/integration/ is unmarked, so the documented unit-test command provisions billable AWS EKS clusters #109
Description
Activity
- addedP0-criticalBlocks everything; safety or correctness landmineBlocks everything; safety or correctness landminetestingTest suite, CI, coverageTest suite, CI, coverage
on Aug 17, 2026 - added a parent issue
on Aug 17, 2026 Fix opened: PR #129.
The root cause turned out to be worse than this issue described, in a way that changes the fix.
Markers would not have been sufficient. pytest imports a module in order to collect it, and two files here are standalone scripts rather than test modules —
test_eks_permissions.pyandtest_aws_eks_debug.pyfetch credentials, call boto3, and callexit()at module scope. So the AWS calls happen before any marker or skip is consulted, and the module-levelsys.exit(1)crashes the run outright:INTERNALERROR> File ".../tests/integration/test_aws_eks_debug.py", line 63, in <module> INTERNALERROR> sys.exit(1) INTERNALERROR> SystemExit: 1 mainloop: caught unexpected SystemExit!The gate therefore stops collection (
collect_ignore_glob), not execution. It is directory-wide default-deny rather than a per-file allowlist: misclassifying 1 file of 48 costs money, and a file added later must not run for free by default. Opt in withCLUSTRIX_ALLOW_BILLABLE=1.Also worth recording: the credentials to make this real are present on a normal dev machine.
~/.clustrix/.envships uncommentedAWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEYandFlexibleCredentialManagerloads exactly that file (credential_manager.py:304-305), with a 1Password CLI fallback. The only reason the audit run did not reach AWS was that boto3 was absent from that venv.Verification
Full suite run with a socket-blocking
sitecustomizelogging every attempt with a stack trace:Before After Collection crashed ( SystemExit, exit 3)2067 tests collected AWS/EC2/EKS connection attempts live connections confirmed via lsof0 boto3/botocoreframes in any stackpresent 0 Two non-billable network sources remain (not fixed here)
The same instrumentation found 3,078 total connection attempts. Neither is billable:
- 3,072 —
pip installsubprocesses reaching PyPI, becauseClusterConfig.__init__auto-installs packages (config.py:195-208). The test suite mutates the developer's environment. -> Replace silent-failure handling, fix resource leaks and import-time side effects #123 - 6 — live Azure Retail Prices API calls from
tests/test_cost_monitoring.py:259,267viacost_providers/azure.py:148,162-> Replace assertion-free mock tests with tests that actually execute (2,513 mock occurrences) #117
Both have been noted on those issues.
- 3,072 —
- added a commit that references this issue
on Aug 17, 2026 Closed by the merge of PR #129 (merge commit
a9393b7).The suite is now free to run. Verified on merged
masterwith the network hard-blocked and every connection attempt logged with a stack trace:Vector Result pytest tests/ -m "not real_world"(the documented command)0 AWS/EC2/EKS connection attempts explicit file / node-id / directory targeting refused with an actionable message direct import of the two unguarded scripts IMPORTED CLEANLY, no side effects(previously fetched real AWS credentials)gate ON vs OFF 1592 vs 1657 collected — withholds exactly the 65 integration tests, nothing else An adversarial review before merge found 5 defects in the first version of the fix, two serious:
- Explicitly-named paths bypassed
collect_ignore_globentirely and imported the modules —pytest tests/integration/test_aws_eks_debug.pyprintedGot credentials for account: .... Only a missing boto3 stopped it. - The marker hook was tagging the whole suite
expensive(pytest passes that hook every item in the session), so-m "not expensive"would have selected nothing.
Both fixed; full details on PR #129.
Two residual limits, documented rather than fixed:
--noconfteststill collects the directory — no conftest-based gate survives that flag. Mitigated: the modules now import inert.- The opt-in path has no cost ceiling.
CLUSTRIX_ALLOW_BILLABLE=1is a single boolean between a shell and an EKS cluster, andaws_provisioner.py:773swallows cleanup errors into a warning. Tracked on Red-team campaign: trust boundary, script injection, serialization fuzzing, cost safety #126.
- Explicitly-named paths bypassed
- added 5 commits that reference this issue
on Aug 17, 2026
Part of #108 · Phase 0 · Do this before running the test suite again.
Problem
Every file in
tests/integration/has zero pytest markers (verified:grep -rc "pytest.mark" tests/integration/returns0for all 20+ files). The documented "unit test" command therefore selects them:One of the files it selects is
tests/integration/test_aws_eks_real_provision.py, whose own docstring reads:This was confirmed empirically, not inferred. During an audit run of
pytest -m "not real_world",lsofshowed the pytest process holding ESTABLISHED TCP connections toec2-*.compute-1.amazonaws.com:443, and stderr emittedError during cluster cleanup in destructor: Cleanup failed.Anyone who has AWS credentials in their environment and runs the documented command starts billing real EKS infrastructure — and the cleanup path is itself failing, so resources may be orphaned.
This is also what makes the suite hang:
clustrix/kubernetes/aws_provisioner.py:763,771sits in atime.sleep(30)xmax_attemptsretry loop, which is why a full coverage run reproducibly stalls at 98%.Acceptance criteria
tests/integration/carries an explicit marker; anything touching a billable API is markedreal_world(or a newcostlymarker) and is excluded by default.pyproject.tomladdoptsexcludes billable markers by default, so the barepytestcommand is safe.CLUSTRIX_ALLOW_BILLABLE=1) andpytest.skipwithout it — belt and braces with the marker.--strict-markersis already set (pyproject.toml:189); every new marker is registered.tests/integration/lacks a marker.Verification