Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
300 changes: 166 additions & 134 deletions .dev-loop/INGEST_REPORT.md

Large diffs are not rendered by default.

8 changes: 4 additions & 4 deletions INDEX.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,10 +12,10 @@ follow the cross-pointers in their index or take the next matching seeded domain
| Domain | Status | Route here when |
|--------|--------|-----------------|
| [databases](wiki/databases/index.md) | **seeded** | Choosing a datastore/database type for a workload (relational vs document vs vector vs graph), designing schemas/tables/keys, choosing or evaluating indexes, writing or optimizing queries, choosing transaction/isolation behavior, multi-row rewrites (reorder, bulk status) on a shared resource, surveying live data to derive a rule, verifying additive migrations |
| [backend](wiki/backend/index.md) | **seeded** | Server-side application code — language-agnostic (`common/`: API contracts, call-site enumeration before a contract change, idempotency, JWT, timeouts/retries, caching, jobs, transactions in app code, shared state/pools, errors, consuming LLM APIs (completion validation, context budgeting), authoring agent-facing artifacts (binding instruction text, agent tool-surface granularity/parity), MAPE-aligned point-prediction calibration, benchmark-relative signal rating, consuming external-API responses, externally-owned defaults, object-storage references, sync-vs-async integration choice, WebSocket/SSE connection lifecycle) plus stack subtrees: `java/` (JPA, Spring proxies, JVM threads/memory), `node/` (event loop, promises, runtime validation, shutdown), `python/` (GIL/asyncio, pydantic, WSGI/ASGI workers, language traps, packaging data files with `importlib.resources`) |
| [frontend](wiki/frontend/index.md) | **seeded** | Web UI code: state placement, rendering performance, in-UI data fetching (races, infinite scroll), auth token handling, forms, XSS-safe output, accessibility, any new or changed user action (this wiki's development standard adds a WebMCP tool surface per action) |
| [infrastructure](wiki/infrastructure/index.md) | **seeded** | CI/CD pipelines, secrets in build/deploy, container image builds, rollout/rollback strategy, observability (logs/metrics/alerting), per-environment/path-valued config, multi-agent orchestration (worker liveness signals, shared run state, tmux pane delivery, completion gates, worktree-isolated workers, autonomous ask-vs-rule decisions, session context/token budgeting, a pre-built code knowledge graph as a freshness-gated orientation layer for planning, the merged-tree gate for parallel branches, the verify command written into a worker brief, a lock owner id inherited from the parent that spawned the session) |
| [testing](wiki/testing/index.md) | **seeded** | Writing or structuring automated tests: level choice, test-before-code ordering, a UI action that is also a registered agent tool, cases/assertions, cross-layer effect scoping, test data, mock decisions, flaky tests, test-infrastructure containers (Testcontainers) failing on the dev host, testing a SwiftPM executable target (release-process quality → qa) |
| [backend](wiki/backend/index.md) | **seeded** | Server-side application code — language-agnostic (`common/`: API contracts, call-site enumeration before a contract change, idempotency, JWT, timeouts/retries, caching, jobs, transactions in app code, shared state/pools, errors, consuming LLM APIs (completion validation, context budgeting, first-call load latency of a self-hosted model server), authoring agent-facing artifacts (binding instruction text, agent tool-surface granularity/parity), MAPE-aligned point-prediction calibration, benchmark-relative signal rating, consuming external-API responses, externally-owned defaults, object-storage references, sync-vs-async integration choice, WebSocket/SSE connection lifecycle) plus stack subtrees: `java/` (JPA, Spring proxies, JVM threads/memory, Kotlin: kotlinx.serialization enum coercion, implicit-receiver shadowing in scope functions), `node/` (event loop, promises, runtime validation, shutdown), `python/` (GIL/asyncio, pydantic, WSGI/ASGI workers, language traps, packaging data files with `importlib.resources`) |
| [frontend](wiki/frontend/index.md) | **seeded** | Web UI code: state placement, rendering performance, in-UI data fetching (races, infinite scroll), auth token handling, forms, XSS-safe output, accessibility, concurrent optimistic updates against one server-owned object (web or mobile view models), scroll-scrubbed animations under a reduced-motion preference, any new or changed user action (this wiki's development standard adds a WebMCP tool surface per action) |
| [infrastructure](wiki/infrastructure/index.md) | **seeded** | CI/CD pipelines, secrets in build/deploy, container image builds, rollout/rollback strategy, observability (logs/metrics/alerting), per-environment/path-valued config, config values a build step writes into generated source, multi-agent orchestration (worker liveness signals, shared run state, tmux pane delivery, completion gates, worktree-isolated workers, autonomous ask-vs-rule decisions, session context/token budgeting, a pre-built code knowledge graph as a freshness-gated orientation layer for planning, the merged-tree gate for parallel branches, the verify command written into a worker brief, a lock owner id inherited from the parent that spawned the session) |
| [testing](wiki/testing/index.md) | **seeded** | Writing or structuring automated tests: level choice, test-before-code ordering, a UI action that is also a registered agent tool, cases/assertions, cross-layer effect scoping, test data, mock decisions (including a fake `IntersectionObserver` under a viewport-animation library), contract tests for a deprecated-alias table, flaky tests, test-infrastructure containers (Testcontainers) failing on the dev host, testing a SwiftPM executable target (release-process quality → qa) |
| [qa](wiki/qa/index.md) | **seeded** | Release-quality process: release gates, regression scoping, bug reports, severity/priority triage, evidence for completion claims, the agent-tool parity gate for a web UI release, acting on code-review feedback, adversarial review of high-risk diffs, exploratory testing (guarded-path coverage, override matrices), scope-purity gates, sourcing deliverable documents from generated artifacts, verifying the quantitative claims in a document before publishing it, a documented claim about a third-party tool's side effects, an obligation row in a tier/policy table that another contract also pins, rationale prose left behind by a config-value change, automated verification of document deliverables (spec/RFC gates), an aging detector for model-coupled agent guidance, capturing an app's own screen content without Screen Recording permission (writing automated test code → testing) |
| [debugging](wiki/debugging/index.md) | **seeded** | Diagnosing a failure — finding what is wrong and why: reproducing, bisection, hypothesis testing, traces/logs, intermittent failures (fixing the diagnosed fault → its owning domain) |
| [security](wiki/security/index.md) | **seeded** | Trust-boundary decisions: input validation, session-vs-token auth choice, per-resource authorization (IDOR), secrets hygiene (including ciphertext orphaned by a regenerated encryption key), dependency trust, PII handling, in-session agent tool exposure (prompt-injection blast radius), the author identity a commit publishes to a public repository, host-compromise triage / incident response (verifying assumed security agents, identifying masquerading processes) (XSS rendering → frontend; CI secrets → infrastructure; JWT implementation → backend/frontend auth) |
Expand Down
18 changes: 18 additions & 0 deletions log.md
Original file line number Diff line number Diff line change
Expand Up @@ -208,3 +208,21 @@ Append-only. Format: `## [YYYY-MM-DD] <ingest|revise|lint|gap|contradiction|drif
## [2026-09-28] ingest | testing-strategy-agent-tool-shared-handler-tests — a UI action that is also a registered tool is tested once at the shared function plus two entry-point tests per tool (Registration incl. AbortSignal teardown, Wiring via spy) against a `document.modelContext` stub; a bug fix that changes the handler's contract changes the `inputSchema` assertion in the same commit; one DevTools Run-tool pass per release.

## [2026-09-28] revise | WebMCP adopted as the development standard (owner decision 2026-09-28): frontend/agent-interfaces/agent-facing-tool-surfaces trigger widened to any new or changed user action in a web UI + bug-fix and exclusion-list edge cases; AGENTS.md routing step 7 gains a web-UI-action row → frontend agent-interfaces then qa parity gate; INDEX.md frontend/qa/testing route lines and the frontend/qa/testing domain indexes updated; related links added both ways (release-gates, cross-layer-effect-tests, in-session-tool-exposure). The standard keeps the human UI primary and the tool layer additive (CG draft; Chrome origin trial + ChatGPT desktop runtimes).

## [2026-09-28] ingest | testing-mocking-fake-intersection-observer-for-viewport-animations — under motion/framer-motion `whileInView`, observers are cached per (root, serialised options) in a module WeakMap so constructor probes read 0 from the second test; the per-element callback is looked up by `entry.target`, so a fake entry without it makes the animation silently never fire while first-paint styles keep the suite green; count `observe(element)`, include `target`, add an end-state positive control and a hardcoded-defaults mutation (framer-motion 13.2.0 source; field evidence cover-letter 2026-09-27)

## [2026-09-28] ingest | testing-quality-alias-table-contract-tests — a deprecated-alias layer is tested row-for-row from the design document's mapping with a size assertion equal to the legacy-name count and one retargeting mutation proven red; a spot-check of 4/18 let un-asserted aliases forward wrong (PIT, Google Testing Blog, JUnit 5 parameterized; field evidence linkly-calendar Android 2026-09-27)

## [2026-09-28] revise | infrastructure-agent-orchestration-checkable-claims-in-an-adopted-plan — +trigger, +Do row, +edge row, +field evidence: an adopt-only plan that freezes test bodies and forbids new failures is probed by one production slice + whole-suite run + failing-set delta against baseline, then reverted and reported (seagrass t174, failures 1→2 inside a frozen test class)

## [2026-09-28] ingest | frontend-state-concurrent-optimistic-updates — concurrent per-field PATCH optimistic updates against one server-owned object: confirmed snapshot + ordered pending patches rendered on top, FIFO-serialised requests, success sets confirmed=response, failure drops only its own patch; invalidate a query cache only when `isMutating() === 1`; overlap test with response gates (TanStack optimistic-updates guide, TkDodo concurrent-updates; field evidence linkly-calendar Android SettingsViewModel 2026-09-27, 4 tests red→green, 131/0)

## [2026-09-28] ingest | infrastructure-config-config-values-emitted-as-source-text — a build step writing user/env input into generated source (AGP `buildConfigField` → `BuildConfig.java`, "must have valid Java content") must escape `\` then `"` in one helper and add a compile probe with hostile characters to the verification step, because javac fails before any runtime validation runs (AGP DSL reference, gradle-tips; field evidence javac 17 `unclosed string literal` 2026-09-27)

## [2026-09-28] ingest | backend-java-kotlin-coerced-enum-defaults-in-kotlinx-serialization — with `coerceInputValues = true`, an unknown enum value on a property with a default is silently replaced by the default; keep must-reject enum properties required (no default), pin with a SerializationException decode test, or make forward-compat explicit with an `UNKNOWN` member (kotlinx.serialization docs/json.md; field evidence linkly-calendar Android TripPlace.category 2026-09-27)

## [2026-09-28] ingest | backend-java-kotlin-implicit-receiver-shadowing-in-scope-functions — inside `apply`/`run`/`with` an unqualified call resolves to the implicit receiver's member before a same-named top-level function (Kotlin spec overload resolution), so a fixture helper `trip(id)` silently calls the fake's `trip(id)`; name helpers distinctly or use `also`/`let` (field evidence linkly-calendar Android TripViewModelTest 2026-09-27, 3 failures → 13 pass)

## [2026-09-28] ingest | backend-common-llm-self-hosted-model-load-latency — the first call to a self-hosted model server after idle pays the disk load (Ollama keeps a model 5 m by default; `OLLAMA_KEEP_ALIVE` / per-request `keep_alive`, request wins); size the client timeout above load+generation, warm up before measuring, set keep_alive longer than the caller's idle interval (Ollama FAQ/API docs; field evidence 8 GB Q8 GGUF judge, 10 s SDK timeout → 31 s recorded failure, warm-up 1588 ms then p50 786 ms, 2026-09-27)

## [2026-09-28] ingest | frontend-design-scrubbed-scroll-animations-under-reduced-motion — a `scrub` ScrollTrigger sets progress from scroll position, so a global `timeScale()` reduced-motion switch never reaches it; read the preference where the effect is created (`gsap.matchMedia` `(prefers-reduced-motion: reduce)`), skip the trigger and set the final state; enumerate `scrub`/progress `onUpdate` sites at task review (GSAP ScrollTrigger, timeScale, matchMedia docs; field evidence cover-letter DurationMeter 2026-09-27)
71 changes: 71 additions & 0 deletions wiki/backend/common/llm/self-hosted-model-load-latency.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,71 @@
---
id: backend-common-llm-self-hosted-model-load-latency
domain: backend
category: llm
applies_to: [general]
confidence: verified
sources:
- https://github.com/ollama/ollama/blob/main/docs/faq.mdx
- https://github.com/ollama/ollama/blob/main/docs/api.md
last_verified: 2026-09-28
related: [backend-common-llm-context-window-budget, backend-common-reliability-timeouts-and-retries, backend-common-reliability-client-side-rate-limiting]
---

# First Request to a Self-Hosted Model Server After Idle

## When this applies

A client calls a locally hosted model server (Ollama or a server with the same
keep-alive model) that loads weights from disk on demand and unloads them after
an idle window, and the caller is intermittent — a hook, a cron job, an
evaluation script, a judge model — or the caller's SDK has a short default
timeout (10 s is common). Also when the first call of a run fails with a
timeout while later calls succeed in under a second.

## Do this

1. **Measure the load once and size the client timeout above load plus
generation.** Time a first call after the model is unloaded (`ollama ps`
shows what is loaded and until when); a multi-gigabyte GGUF takes tens of
seconds from disk. Set the timeout to 60 s or more for such models — a
timeout at 10 s returns an error while the load continues.
2. **Send a warm-up request before anything you measure or batch**, and report
its latency separately from the steady-state numbers.
3. **Set `keep_alive` longer than the caller's idle interval.** Ollama keeps a
model loaded "for 5 minutes before being unloaded" by default; the
`OLLAMA_KEEP_ALIVE` environment variable changes it for all models, and the
`keep_alive` parameter on `/api/generate` and `/api/chat` overrides it per
request (duration string, seconds, a negative number to keep loaded, `0` to
unload after the response).
4. **Treat a timeout on the first call as a load in progress, not a dead
server:** check the server's loaded-model list before retrying — a retry
inside the same short timeout pays the load again and records a second
failure ([backend-common-reliability-timeouts-and-retries]).

| Caller pattern | keep_alive |
|----------------|------------|
| Interactive session, calls seconds apart | Default (5 m) |
| Hook or job firing every N minutes | Longer than N (`OLLAMA_KEEP_ALIVE=30m` or per-request `"30m"`) |
| Dedicated judge/eval box, memory to spare | Negative value (`-1`) — stays loaded |
| One-shot batch, then free the memory | `0` on the last request |

## Edge cases

| Case | Then |
|------|------|
| The SDK retries on timeout | The first result is recorded as a failure after `timeout × attempts` while the server was loading the whole time; disable retry for the warm-up call or raise the timeout first |
| Several models share one server | Each cold model pays its own load; warm up each one you will call |
| The env var and the request parameter disagree | The request parameter wins — the documented precedence |

## Instead of

| If you are about to | Do this instead | Why |
|---------------------|-----------------|-----|
| Read the first-call timeout as "the server is broken" | Check `ollama ps` and re-send after the load | The model was loading; the error is the client's clock, not the server |
| Publish p50 latency from a run that includes the cold call | Warm up first and report the warm-up separately | One 30 s load skews every summary statistic |

## Sources

- https://github.com/ollama/ollama/blob/main/docs/faq.mdx — "By default models are kept in memory for 5 minutes before being unloaded"; `keep_alive` accepts "a duration string (such as "10m" or "24h")", "a number in seconds (such as 3600)", "any negative number which will keep the model loaded in memory (e.g. -1 or "-1m")", "'0' which will unload the model immediately after generating a response"; "The `keep_alive` API parameter with the `/api/generate` and `/api/chat` API endpoints will override the `OLLAMA_KEEP_ALIVE` setting"
- https://github.com/ollama/ollama/blob/main/docs/api.md — `keep_alive`: "controls how long the model will stay loaded into memory following the request (default: `5m`)"
- Field evidence 2026-09-27 (a local evaluation harness against an Ollama-compatible server hosting an 8 GB Q8 GGUF judge model; client SDK default timeout 10 s with retry): the first judgment was recorded as `31334 ms ERROR … Request timed out (timeout=10.0)`; with the timeout raised to 120 s and one warm-up call (`warmup 1588 ms`) the run scored 50/50 with p50 786 ms, and the server's process list showed the model unloading "4 minutes from now"
Loading
Loading