Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -233,6 +233,8 @@ Fetch only the files relevant to the task. A typical example contains

- **`costguard`** `[cost, cleanup, budgets, automation, iaas, terraform]`
**costguard keeps your STACKIT cloud tidy and installs with one `terraform apply`.** It looks for things that cost money but are not used by anything, reports them in a chat channel and, once you switch deletion on, deletes them. It also watches monthly budgets and posts when one passes a threshold. It runs on a small server that logs in with the service account attached to it, so no key is stored on the server
- **`whisper-fn`** `[functions, serverless, object-storage, s3, ai-ml, scale-to-zero]`
Rebuilds an always-on VM transcription demo as a **STACKIT Function**: a container that only runs while it's handling a request and scales to zero afterwards, instead of a VM that runs (and costs) 24/7

---

Expand Down
3 changes: 3 additions & 0 deletions apps/whisper-fn/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
whisper-fn/.env
__pycache__/
*.pyc
72 changes: 72 additions & 0 deletions apps/whisper-fn/DECISIONS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,72 @@
# Decisions

## 2026-10-02: WHISPER_MODEL small, not large-v3

**Was:** Deploy whisper-fn with `WHISPER_MODEL=small`, not `large-v3` as originally requested.

**Warum:** STACKIT Functions plans cap at `f6` (4096 MB RAM, 1 shared vCPU) --
there is no bigger self-service tier; a larger plan requires contacting
STACKIT directly, with no guaranteed turnaround. Empirically tested all
three model sizes under memory caps matching the real plans:

- `large-v3` (2.88 GB checkpoint): OOM-killed at 4096 MB (`f6`, the max plan),
and even failed to *load* during a local Docker build at 7.75 GB (needed
11.67 GB to succeed there).
- `medium` (1.42 GB checkpoint): also OOM-killed at 4096 MB.
- `small` (~500 MB checkpoint): confirmed working end-to-end (real
transcription of `elephants_dream.mp4`) under caps down to 2048 MB.

**Alternative verworfen:** Hosting `large-v3` on an always-on STACKIT Server/VM
instead of Functions would work memory-wise, but loses Scale-to-Zero and runs
(and bills) continuously.
Not pursued since the user confirmed `small` is acceptable.

## 2026-10-02: Dockerfile custom-build, not Cloud Native Buildpacks

**Was:** Build `whisper-fn`'s image via a hand-written `Dockerfile`
(`sfn function deploy --build=false --push=false`), not the default
`sfn functions build` (buildpacks) path.

**Warum:** The buildpack-built image includes a zero-byte metadata layer
that this project's STACKIT Harbor registry rejects on push
(`error from registry: blob unknown to registry -
sha256:e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855` --
the well-known empty-blob digest). Reproduced consistently across retries;
all other layers pushed fine. A plain Dockerfile build (ordinary
`RUN`/`COPY` layers) does not produce that layer, and pushed successfully.

This was initially built for `large-v3` (to bake the model weights into
the image at build time, since buildpacks have no `apt-get` and openai-whisper
needs `ffmpeg`), then rebuilt for `small` after the model-size decision
above. The `small` Dockerfile dropped the `large-v3`-specific RAM tuning
(Docker Desktop memory bump) since `small` loads fine without it.

**Alternative verworfen:** Pushing to a different registry (ghcr.io, Docker
Hub) instead of fixing the buildpack image -- would have avoided the bug
but moves the image outside STACKIT's own registry; not pursued per user's
choice to stay on the official STACKIT Container Registry path.

## 2026-10-02: Plan f6, concurrency 1

**Was:** Deployed revision uses `plan: f6` (4096 MB) and `concurrency: 1`,
not the template defaults (`f3` / 512 MB, concurrency 50).

**Warum:** `small` was empirically confirmed to need more than the `f3`
(512 MB) and `f4` (1024 MB) tiers -- both OOM-killed. `f5` (2048 MB) passed
a full end-to-end transcription test locally and was initially deployed;
user then asked for `f6` for extra headroom, so that's the live
configuration. Concurrency was dropped from the template's default of 50 to
1, since each `/transcribe` request is a single CPU-bound ~100-150s job on
a 1-shared-vCPU plan -- running several in parallel would contend for the
same CPU/memory budget that's already near its ceiling.

## 2026-10-02: Image tag must be SemVer

**Was:** Pushed image as `whisper-fn:0.1.0`, not `whisper-fn:small`.

**Warum:** `sfn functions deploy` rejected the `:small` tag with
`Error: Error creating function revision: oci image reference invalid` --
STACKIT Functions requires either a SemVer tag (with or without `v` prefix)
or `latest` combined with a digest. Not documented in the Functions docs
mirror at the time of this session; discovered via the deploy error message
itself.
185 changes: 185 additions & 0 deletions apps/whisper-fn/GETTING-STARTED.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,185 @@
# Getting Started

Step-by-step guide to pick this project up and get `whisper-fn` running --
locally first, then deployed to STACKIT. This is the as-built path; see
`DECISIONS.md` for why it diverges from the original plan (short version:
`large-v3` does not fit any STACKIT Functions plan, and the default
buildpack build can't be pushed to this project's Container Registry).

## Current live deployment

- URL: `https://<function-id>.functions.onstackit.cloud`
- Project: `functions` (`<your-stackit-project-id>`)
- Model: `small` (not `large-v3` -- see `DECISIONS.md`)
- Plan: `f6` (4096 MB), concurrency 1
- Image: `registry.onstackit.cloud/<your-registry-project>/whisper-fn:0.1.0`

## Prerequisites

- Docker installed and running (`docker info` should succeed)
- `sfn` CLI >= 1.7.0 (`./install-sfn.sh`; older versions fail auth with
"outdated cli" against this project)
- `curl`
- A STACKIT project with Functions enabled, `Object Storage bucket`, and
`Container Registry` already set up (see below -- all three already exist
for this project as of this writing)

## 1. Install/update the STACKIT Functions CLI (`sfn`)

```bash
cd apps/whisper-fn # this directory
./install-sfn.sh
export PATH="$HOME/.local/bin:$PATH" # add to ~/.zshrc to persist
sfn --version # must be >= 1.7.0
```

## 2. Build the function container (Dockerfile, not buildpacks)

```bash
cd whisper-fn
docker build --platform linux/amd64 -t whisper-fn:local .
```

This is a hand-written `Dockerfile`, deliberately **not**
`sfn functions build` (Cloud Native Buildpacks). The buildpack output
includes a zero-byte metadata layer that this project's STACKIT Harbor
registry rejects on push (`blob unknown to registry`) -- see
`DECISIONS.md`. The Dockerfile bakes the `small` model weights in at build
time so cold starts don't re-download them, creates the required non-root
`uid 1001` user, and serves via `uvicorn asgi:app` (see `asgi.py` for the
hand-rolled ASGI lifespan wrapper that the buildpack path would normally
generate for you).

`--platform linux/amd64` is required even on Apple Silicon -- STACKIT
Functions' Knative runtime only runs `linux/amd64` images.

## 3. Run it locally and smoke-test

```bash
docker run -d --platform linux/amd64 -p 8080:8080 -e PORT=8080 \
--env-file .env --name whisper-fn-local whisper-fn:local
curl http://127.0.0.1:8080/health
# {"status": "ok", "model": "small"}
curl -X POST http://127.0.0.1:8080/transcribe
# downloads elephants_dream.mp4 from the bucket, transcribes, uploads the
# transcript back -- takes 90-150s on CPU
docker rm -f whisper-fn-local
```

`whisper-fn/.env` already has working credentials for the `your-bucket-name`
bucket, which already contains `elephants_dream.mp4`. If you need to
recreate it elsewhere: see `.env.example` for the fields, and `stackit
object-storage bucket create` / `credentials create` to provision a new one.

## 4. (Optional, no Docker needed) Run the pure logic test

```bash
cd whisper-fn
python3 test_handler_local.py
```

Stubs out `boto3`/`whisper`/`imageio_ffmpeg` in memory; checks routing and
the S3 upload flow without real transcription or a real S3 connection.

## 5. Push to the STACKIT Container Registry

The registry itself has no CLI or Terraform support -- it's created once
via the STACKIT Portal (Container Registry -> New Project -> Robot Account
with Pull+Push permissions for CI use). This project already has one:
registry project `<your-registry-project>`.

```bash
docker login registry.onstackit.cloud \
--username 'robot$<your-registry-project>+<robot-name>'
# password: the robot account's secret from the Portal

docker tag whisper-fn:local registry.onstackit.cloud/<your-registry-project>/whisper-fn:<version>
docker push registry.onstackit.cloud/<your-registry-project>/whisper-fn:<version>
```

**The tag must be SemVer** (e.g. `0.1.0`, optionally with a `v` prefix), or
`latest` combined with an explicit digest. Plain tags like `small` or
`local` are rejected by `sfn functions deploy` with `oci image reference
invalid`.

If `docker login` fails with `error storing credentials ... User
interaction is not allowed` (macOS Keychain inaccessible, e.g. from a
non-interactive shell), write the base64 auth directly into an isolated
`DOCKER_CONFIG` directory instead of your real `~/.docker/config.json`:

```bash
mkdir -p /tmp/docker-config
python3 -c "
import base64, json
auth = base64.b64encode(b'robot\$<your-registry-project>+<robot-name>:<secret>').decode()
json.dump({'auths': {'registry.onstackit.cloud': {'auth': auth}}}, open('/tmp/docker-config/config.json', 'w'))
"
docker --config /tmp/docker-config push registry.onstackit.cloud/<your-registry-project>/whisper-fn:<version>
```

## 6. Register the runtime pull-secret (once per project)

```bash
sfn pull-credentials create --address registry.onstackit.cloud \
--pull-credential-name whisper-fn-registry
# prompts for username (robot$...) and password interactively --
# --no-interactive is explicitly disallowed here for security reasons
```

## 7. Update `.stackit-functions/revision.yaml`

Set `spec.image` to the pushed `<registry>/<project>/<image>:<version>`
reference, and `spec.limits.plan` to a plan with enough memory. For
`small`, empirically: `f3` (512 MB) and `f4` (1024 MB) OOM-kill; `f5`
(2048 MB) and `f6` (4096 MB) both work. This project runs `f6`. Test your
own plan choice locally first:

```bash
docker run -d --platform linux/amd64 --memory=<plan-mb>m --memory-swap=<plan-mb>m \
-p 8080:8080 -e PORT=8080 --env-file .env --name whisper-fn-memtest whisper-fn:local
# wait, then:
docker inspect -f 'OOMKilled={{.State.OOMKilled}}' whisper-fn-memtest
docker rm -f whisper-fn-memtest
```

## 8. Log in and deploy

```bash
sfn auth login --project-id <project-id> \
--service-account-key-path ~/.stackit/sa-key.json
# or with a user account via browser login

cd whisper-fn
sfn functions deploy --project-id <project-id> --env-file .env --no-interactive -v
```

No `--build`/`--push` flags needed -- omitting them means "don't rebuild,
don't repush," since the image is already in the registry from step 5.
(Note: these are boolean switches, not `key=value` options --
`--build=false` is a CLI parse error, just omit the flag entirely.)

## 9. Verify

```bash
sfn function describe --function-id <id> # shows the public URL
curl https://<function-url>/health
curl -X POST https://<function-url>/transcribe
```

Then check the bucket for the uploaded transcript.

## Known constraints

- **No model bigger than `small` fits any STACKIT Functions plan.** Plans
cap at `f6` / 4096 MB with 1 shared vCPU; both `large-v3` and `medium`
were OOM-killed even at that ceiling. See `DECISIONS.md`.
- **No documented request-timeout limit** on STACKIT Functions, so the
90-150s transcription time for `small` is not a gateway-timeout risk.
- Concurrency is set to 1: each `/transcribe` call is a single CPU-bound
job that would contend for the same memory/CPU budget if run in parallel.

## If something's unclear

`docs/history/Task.md` has the original design rationale and open questions;
`docs/history/Handout.md` has the session-by-session implementation log;
`DECISIONS.md` has the non-trivial choices made and why.
7 changes: 7 additions & 0 deletions apps/whisper-fn/MAINTAINERS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
# Maintainers

Created by Konstantin Kloster ([@KonstantinMKloster](https://github.com/KonstantinMKloster)).

This app does not currently have an assigned maintainer responsible for
ongoing upkeep. If you rely on it and are interested in maintaining it,
please open an issue or reach out via a Pull Request.
82 changes: 82 additions & 0 deletions apps/whisper-fn/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,82 @@
<!-- tags: functions, serverless, object-storage, s3, ai-ml, scale-to-zero -->

# whisper-fn

Rebuilds an always-on VM transcription demo as a **STACKIT Function**: a
container that only runs while it's handling a request and scales to zero
afterwards, instead of a VM that runs (and costs) 24/7.

## What it does

An HTTP-triggered function (`whisper-fn/`) that, on `POST /transcribe`:

1. downloads a video from an S3 (STACKIT Object Storage) bucket,
2. transcribes it locally with [`openai-whisper`](https://github.com/openai/whisper)
(the open-weights Whisper model running inside the function — **not** a call
to OpenAI's paid API),
3. uploads the resulting transcript back to the same bucket,
4. returns a small JSON summary (bucket, keys, model, duration).

`GET /health` returns a plain status check.

## Why it exists

A Terraform-provisioned-VM version of this demo works, but it runs continuously
via `cloud-init` and only ever leaves its result on local disk — nobody stops
it, so it costs money the whole time it's not transcribing anything. STACKIT
Functions offered a scale-to-zero, HTTP-triggered alternative, so this project
reimplements the same demo (same input video, same Whisper model family) as a
function instead of a VM, with the transcript written back to S3 so the result
is actually retrievable afterwards.

This project provisions its own S3 bucket and Container Registry rather than
assuming you already have one set up — see [GETTING-STARTED.md](GETTING-STARTED.md)
for the one-time manual steps.

## Project layout

| Path | Purpose |
|---|---|
| `whisper-fn/` | The actual function: `handler.py`, dependencies, STACKIT Functions manifests |
| `whisper-fn/.env.example` | Template for the env vars you need (S3 bucket/credentials, model size) — copy to `whisper-fn/.env`, fill in your own values, never commit it |
| `install-sfn.sh` | Installs the `sfn` CLI (STACKIT Functions CLI) locally |
| `setup.sh` | One-command local bootstrap: installs the sfn CLI, prepares `whisper-fn/.env`. Creates no cloud resources. |
| `deploy.sh` | Convenience wrapper: build → auth → confirm → deploy, using `whisper-fn/.env` |
| `GETTING-STARTED.md` | Step-by-step setup and deploy instructions |
| `docs/history/Task.md` | Original design/planning notes (background reading, not required to run this) |
| `docs/history/Handout.md` | Session log — implementation history and the disk-space blocker that paused the first deploy attempt (background reading, not required to run this) |

## Status

Verified end-to-end with a real transcription of `elephants_dream.mp4`
(downloaded from the bucket, transcribed, transcript written back), running
`WHISPER_MODEL=small`, not the originally planned `large-v3` — see
`DECISIONS.md` for why (short version: no STACKIT Functions plan has enough
memory for `large-v3` or even `medium`). Built via a hand-written
`Dockerfile`, not the default buildpack path — see `DECISIONS.md` for that
too (buildpack images get rejected by this registry).

## Quick start

```bash
cd apps/whisper-fn
./setup.sh # installs sfn CLI, prepares whisper-fn/.env (no cloud resources created)
```

Then, after filling in `whisper-fn/.env` with your own S3 bucket/credentials:

```bash
./deploy.sh <project-id> <registry-image-ref> <registry-username>
# e.g. ./deploy.sh <your-stackit-project-id> registry.onstackit.cloud/<reg-project>/whisper-fn:0.1.1 'robot$<reg-project>+<robot-name>'
```

See [`GETTING-STARTED.md`](GETTING-STARTED.md) for the full walkthrough --
creating the Object Storage bucket and Container Registry are one-time,
manual (Portal) steps not covered by `setup.sh` or `deploy.sh`.

## Cost note

This project creates real STACKIT resources once deployed (S3 bucket, a running
function instance while it's warm). Check your STACKIT budget/alerting before
deploying, and tear the function down (`sfn function delete`) when you're done
with it.
Loading