Kubernetes execution and auto-provisioning are implemented and removed from v0.2.0 as unverified. No job has ever been run against a real cluster.
What exists, and where
| Code |
Lines |
Introduced |
clustrix/executor_kubernetes.py |
594 |
6061247 (2025-08-23) |
clustrix/kubernetes/cluster_provisioner.py |
423 |
d6b4b7b (2025-08-16) |
clustrix/kubernetes/local_provisioner.py |
455 |
d6b4b7b |
clustrix/kubernetes/{aws,azure,gcp,lambda,huggingface}_provisioner.py |
3,611 |
d6b4b7b |
clustrix/kubernetes/__init__.py |
60 |
d6b4b7b |
Plus the widget's Kubernetes section, the k8s_* configuration fields, and docs/source/notebooks/kubernetes_tutorial.ipynb.
Defects found while auditing
Two were real and are worth carrying forward whenever this is restored:
check_k8s_job_status reported failed jobs as successful. Its exception paths returned "completed", so a job that died came back as a normal completion.
- Results were decoded with
ast.literal_eval on the pod log, falling back to returning the raw log text as if it were the function's return value.
Both were fixed during the v0.2.0 sweep, and neither fix has been run against a real cluster.
Note on scope
The five cloud provisioners under clustrix/kubernetes/ are removed with this issue rather than with the per-provider issues, because they are Kubernetes cluster provisioners, not the VM-based cloud execution backends. See the AWS/GCP/Azure/Lambda issues for those.
Why it is being removed rather than fixed
Not because the code is known to be wrong. Because it has never been run against the real thing, and shipping it in the cluster-type dropdown states otherwise. A user who selects it gets a code path no one has ever seen succeed.
v0.2.0 keeps exactly the four backends that have been demonstrated end to end -- local, ssh, slurm, huggingface -- and the documentation now says the rest are planned for a future release rather than currently supported.
Restoring it
Nothing is lost: every line cited above stays reachable in git history at the commits named. Reinstating it means reverting the removal commit and then doing the part that was never done -- running it against real hardware and recording the evidence in this issue.
Definition of done
Kubernetes execution and auto-provisioning are implemented and removed from v0.2.0 as unverified. No job has ever been run against a real cluster.
What exists, and where
clustrix/executor_kubernetes.py6061247(2025-08-23)clustrix/kubernetes/cluster_provisioner.pyd6b4b7b(2025-08-16)clustrix/kubernetes/local_provisioner.pyd6b4b7bclustrix/kubernetes/{aws,azure,gcp,lambda,huggingface}_provisioner.pyd6b4b7bclustrix/kubernetes/__init__.pyd6b4b7bPlus the widget's Kubernetes section, the
k8s_*configuration fields, anddocs/source/notebooks/kubernetes_tutorial.ipynb.Defects found while auditing
Two were real and are worth carrying forward whenever this is restored:
check_k8s_job_statusreported failed jobs as successful. Its exception paths returned"completed", so a job that died came back as a normal completion.ast.literal_evalon the pod log, falling back to returning the raw log text as if it were the function's return value.Both were fixed during the v0.2.0 sweep, and neither fix has been run against a real cluster.
Note on scope
The five cloud provisioners under
clustrix/kubernetes/are removed with this issue rather than with the per-provider issues, because they are Kubernetes cluster provisioners, not the VM-based cloud execution backends. See the AWS/GCP/Azure/Lambda issues for those.Why it is being removed rather than fixed
Not because the code is known to be wrong. Because it has never been run against the real thing, and shipping it in the cluster-type dropdown states otherwise. A user who selects it gets a code path no one has ever seen succeed.
v0.2.0 keeps exactly the four backends that have been demonstrated end to end --
local,ssh,slurm,huggingface-- and the documentation now says the rest are planned for a future release rather than currently supported.Restoring it
Nothing is lost: every line cited above stays reachable in git history at the commits named. Reinstating it means reverting the removal commit and then doing the part that was never done -- running it against real hardware and recording the evidence in this issue.
Definition of done
SUPPORTED_CLUSTER_TYPES, the widget dropdown and the CLI