Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -143,6 +143,7 @@ all, and ad-blockers can't tell it apart from your own API traffic.
- [Stack integrations](docs/INTEGRATIONS.md) — React, Node, Next.js, Go, Rust, generic OTel
- [Using Autter Runtime **without npm**](docs/WITHOUT-NPM.md) — any OTel SDK, an OTel Collector, or plain HTTP from any language
- [Architecture & data model](docs/ARCHITECTURE.md)
- [Operation logging and diagnostic context](docs/OPERATION-LOGGING.md)
- [Continuous detection, profiles, and outcomes](docs/CONTINUOUS-DETECTION.md)
- [Roadmap](docs/PLAN.md) · [Releasing](docs/RELEASING.md)

Expand Down
1 change: 1 addition & 0 deletions docs/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,6 +43,7 @@ Two key scopes separate frontend and backend credentials:
| --- | --- | --- | --- |
| `runtime_error_occurrences` | MergeTree | `(org_id, repository_id, fingerprint, occurred_at)` | 14 d |
| `runtime_spans` | MergeTree | `(org_id, repository_id, trace_id, started_at)` | 7 d |
| `runtime_logs` | ReplacingMergeTree | `(org_id, repository_id, occurred_at, event_id)` | 14 d |
| `runtime_metrics_1m` | SummingMergeTree | `(org_id, repository_id, service, environment, release, route, bucket_at)` | 90 d |
| `runtime_llm_calls` | MergeTree | `(org_id, repository_id, started_at)` | 90 d |
| `runtime_profile_samples` | MergeTree | `(org_id, repository_id, service, environment, release, observed_at, profile_id)` | 7 d |
Expand Down
8 changes: 7 additions & 1 deletion docs/GETTING-STARTED.md
Original file line number Diff line number Diff line change
@@ -1,3 +1,9 @@
# Operation logging

For structured messages, completed operation summaries, and explicit outcomes,
see [Operation logging and diagnostic context](OPERATION-LOGGING.md). These APIs
require the updated Node/Next.js SDK and ingester described in that guide.

# Getting started

Already collecting logs in an external provider? See [External sources](EXTERNAL-SOURCES.md) for repository connectors. The SDK/OTel setup below applies when you instrument application code.
Expand Down Expand Up @@ -335,4 +341,4 @@ clickhouse-client --password dev`.
- [ ] Direct browser ingest: your CSP includes
`connect-src https://your-ingester…`, and you accept that ad-blockers
may drop some events (the relay avoids this).
- [ ] The ingester's `/healthz` is wired to your load-balancer health check.
- [ ] The ingester's `/healthz` is wired to your load-balancer health check.
114 changes: 114 additions & 0 deletions docs/OPERATION-LOGGING.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,114 @@
# Operation logging and diagnostic context

Operation logging starts in `@autter/runtime-node` and `@autter/runtime-next`
version **1.4.0**. Update the ingester to **1.4.0** before upgrading SDKs: it creates
`runtime_logs` through migration `0011-runtime-logs` and accepts `/v1/logs`.

## Node and Next.js

Initialize `initAutterServer` once, before the application starts. For Next.js,
use `registerAutter` in `instrumentation.ts` and import the APIs below from
`@autter/runtime-next/server`. Use the Node runtime; these APIs use Node's async
context and do not run in the browser or an edge worker.

```ts
import {
initAutterServer, withRuntimeOperation, runtimeLogger,
} from "@autter/runtime-node";

const runtime = initAutterServer({
apiKey: process.env.AUTTER_RUNTIME_KEY!,
service: "payments-api",
release: process.env.GIT_SHA,
logging: { console: false, minLevel: "info" },
});

await withRuntimeOperation("checkout", async (operation) => {
operation.setContext({ "payment.provider": "stripe", "cart.item_count": 3 });
await operation.step("reserve_inventory", () => reserveInventory());
const payment = await operation.step("confirm_payment", () => confirmPayment());
runtimeLogger.info("Payment confirmation returned", { "payment.attempts": 2 });
if (!payment.confirmed) {
operation.outcome("failed", "Payment was not confirmed; no order created");
return;
}
await operation.step("create_order", () => createOrder(payment));
});

// Flush at the end of a short-lived invocation; shutdown before process exit.
await runtime.shutdown();
```

`withRuntimeOperation(name, fn, attributes?)` runs the callback in an always-recorded
process span and emits one completed operation summary. It returns the callback's
result and rethrows its original error. An unhandled callback error records an
exception in the trace; the log summary is related evidence, not a second issue.

`operation.setContext(attributes)` merges redacted nested context (arrays are
replaced) and adds attributes to the operation and
its active span. `operation.step(name, fn)` records the step result and elapsed
time; it rethrows failures. A caught step error can be recovered by application
code: the final outcome follows the callback's result or an explicit outcome.
Up to 64 steps are recorded. Log names and context should describe the operation,
not contain request bodies, payment data, or personal information.
`RuntimeLogContext` supports nested objects, arrays and scalar values; standard
trace attributes encode structured values as JSON for OTLP compatibility.

`operation.outcome(status, message?)` accepts `succeeded`, `failed`, `degraded`,
`cancelled`, and `pending`. The default is `succeeded` when the callback returns.
Returning normally is not proof of business success: declare an outcome when
the intended result was not achieved. An explicit `failed` outcome also emits
the existing `autter.outcome` trace event, including the reporting call site.
`pending` means the operation has not confirmed its downstream result.

`createRuntimeLogger(attributes?)` creates a logger with `debug`, `info`, `warn`,
and `error` methods. `runtimeLogger` is the default instance. `error` records
diagnostic context; use `captureException` for exception grouping or declare a
failed operation outcome for business-failure grouping. Logs inherit the
active operation ID, name, parent operation ID, and application attributes, plus
the active OTel trace/span IDs. Concurrent operations keep separate contexts.
Async context does not cross a queue or process: propagate an opaque workflow
identifier in the job payload and pass it as a custom attribute in the consumer.

## Collection and privacy

- Logs are OTLP/HTTP log records. Both OTLP JSON and protobuf are accepted.
evlog and other OTLP log exporters can use `/v1/logs` with a server ingest key.
- Source attributes are bounded and scrubbed again at ingestion. Tenant identity
comes from the validated key, never the supplied event. Client/browser keys
cannot send server logs.
- Redaction applies before console output and export. Never intentionally send
secrets or personal data; pattern-based redaction cannot identify every secret.
- `logging.minLevel` filters ordinary messages; completed operation summaries
are retained independently. `logging.console` defaults to true and can be
disabled without disabling export.
- Logs batch in memory (up to 1,000 records or 4 MiB, whichever comes first).
Requests contain up to 50 records or 512 KiB. Context has depth, field, step
and string limits; truncated context is marked. Each flush has a 10-second
budget and makes at most three attempts per batch. Failed batches remain
buffered for a later flush, with automatic retry intervals up to 30 seconds. Buffer overflow and failed shutdown delivery are
reported. `flushRuntimeLogs()` rejects if delivery fails, and
`runtimeLogStats()` reports buffered and dropped counts. This is best-effort
telemetry, not durable delivery or an audit log.
- Records expire after 14 days. Successful operation counts represent captured
summaries, not a guarantee that every application operation was observed.

## Investigation and fixes

Autter's repository Runtime Logs view displays captured messages, operation
summaries, outcomes, steps, and trace IDs. Investigation readers retrieve related
logs and spans using captured IDs and compare bounded successful operations.
They disclose unavailable sources and truncation. Event text is untrusted
diagnostic evidence, never instructions to an agent. Business context improves
an investigation but does not establish a cause by itself.

The observations feed the existing root-cause analysis and draft-fix pipelines,
which retain their repository settings, source checks and validation rules. A reported outcome call site can locate the reporting code; it is not
proof that this location caused the failure. The fix agent must confirm the
cause, reproduce applicable conditions, and validate the intended result.

Log and trace exporters are independent. Operation evidence may arrive after an
initial analysis; refresh the evidence panel or rerun analysis to incorporate
later arrivals. Shutdown attempts log export before closing tracers. Regular Runtime flushing
exports logs and traces independently so a failed log export does not block
exception telemetry.
8 changes: 4 additions & 4 deletions package-lock.json

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion packages/otlp-ingester/package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "@autter/otlp-ingester",
"version": "1.3.4",
"version": "1.4.0",
"description": "Self-hostable OTLP + browser-error ingest service for Autter Runtime: normalises telemetry into a per-repo ClickHouse data model",
"license": "MIT",
"type": "module",
Expand Down
14 changes: 14 additions & 0 deletions packages/otlp-ingester/src/clickhouse.ts
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ import { latencyTableDDL, type LatencyHistogram } from "./latency.js";
import { profileTableDDL, type ProfileSample } from "./profiles.js";
import { memoryTableDDL, platformEventTableDDL, type MemorySample, type PlatformEvent } from "./memory.js";
import { sourceMapTableDDL } from "./source-maps.js";
import { logTableDDL, type RuntimeLogRecord } from "./logs.js";
import {
MIGRATIONS,
migrationsTableDDL,
Expand Down Expand Up @@ -75,6 +76,7 @@ export class ClickHouseStore {
memoryTableDDL(db),
platformEventTableDDL(db),
sourceMapTableDDL(db),
logTableDDL(db),
`CREATE TABLE IF NOT EXISTS ${db}.runtime_error_occurrences (
org_id String,
repository_id String,
Expand Down Expand Up @@ -294,6 +296,18 @@ export class ClickHouseStore {
});
}

async insertLogs(ctx: IngestContext, logs: RuntimeLogRecord[]): Promise<void> {
if (!logs.length) return;
if (!this.configured) throw new Error("CLICKHOUSE_URL is not configured");
await this.ensureSchema();
await this.getClient().insert({ table: this.table("runtime_logs"), format: "JSONEachRow", clickhouse_settings: INSERT_SETTINGS,
values: logs.map((row) => ({ org_id: ctx.orgId, repository_id: ctx.repositoryId, event_id: row.id,
service: row.service, environment: row.environment, release: row.release, trace_id: row.traceId,
span_id: row.spanId, operation_id: row.operationId, operation: row.operation, event_type: row.type,
severity: row.severity, message: row.message, outcome: row.outcome, duration_ms: row.durationMs,
attributes: JSON.stringify(row.attributes), occurred_at: row.occurredAt.toISOString() })) });
}

async insertProfileSamples(ctx: IngestContext, samples: ProfileSample[]): Promise<void> {
if (!samples.length) return;
if (!this.configured) throw new Error("CLICKHOUSE_URL is not configured");
Expand Down
124 changes: 124 additions & 0 deletions packages/otlp-ingester/src/context.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,124 @@
/** Bounded, defensive privacy boundary for custom telemetry from any OTLP SDK. */
export function sanitizeRuntimeContext(
input: unknown,
): Record<string, unknown> {
let budget = 512;
const seen = new WeakSet<object>();
const scrub = (value: string) =>
value
.replace(/[A-Z0-9._%+-]+@[A-Z0-9.-]+\.[A-Z]{2,}/gi, "[redacted]")
.replace(
/\b(?:bearer\s+[A-Za-z0-9._~+\/=~-]{10,}|eyJ[A-Za-z0-9_-]+\.[A-Za-z0-9_-]+\.[A-Za-z0-9_-]+|gh[pousr]_[A-Za-z0-9]{20,}|sk-[A-Za-z0-9_-]{20,}|xox[baprs]-[A-Za-z0-9-]{10,}|autter_(?:rt|pat)_[A-Za-z0-9_-]{10,}|(?:AKIA|ASIA)[A-Z0-9]{16})\b/gi,
"[redacted]",
)
.replace(
/-----BEGIN [A-Z ]*PRIVATE KEY-----[\s\S]*?(?:-----END [A-Z ]*PRIVATE KEY-----|$)/g,
"[redacted]",
)
.replace(/([a-z][a-z0-9+.-]*:\/\/)[^\s/:@]+:[^\s/@]+@/gi, "$1[redacted]@")
.replace(/(https?:\/\/[^\s?#]+)[?#][^\s]*/gi, "$1");
const visit = (value: unknown, key: string, depth: number): unknown => {
if (--budget < 0 || depth > 6) return "[truncated]";
if (/__proto__|constructor|prototype/i.test(key)) return undefined;
const usage =
/(?:^|\.)(?:input|output|total|prompt|completion)_?tokens$/i.test(key) &&
typeof value === "number" &&
Number.isFinite(value) &&
value >= 0;
if (
!usage &&
/password|passwd|secret|token|credential|authorization|cookie|email|phone|ssn|card[._-]?number|connection[._-]?string|api[._-]?key|private[._-]?key|request[._-]?body|response[._-]?body|headers|url[._-]?query/i.test(
key,
)
)
return "[redacted]";
if (typeof value === "string") {
// Serialized custom context crosses the same privacy boundary as nested values.
if (/^\s*[\[{]/.test(value)) {
try {
return visit(JSON.parse(value), key, depth + 1);
} catch {
/* plain text */
}
}
const text = /(?:url|path|route|target)$/i.test(key)
? value.split(/[?#]/)[0]!
: value;
return scrub(text).slice(0, /stack/i.test(key) ? 32000 : 2048);
}
if (typeof value === "number")
return Number.isFinite(value) ? value : undefined;
if (typeof value === "boolean" || value === null) return value;
if (!value || typeof value !== "object") return undefined;
if (seen.has(value)) return "[circular]";
seen.add(value);
if (Array.isArray(value))
return value.slice(0, 64).map((v) => visit(v, key, depth + 1));
const result: Record<string, unknown> = {};
for (const [k, v] of Object.entries(value).slice(0, 128)) {
const safe = visit(v, k, depth + 1);
if (safe !== undefined) result[k.slice(0, 200)] = safe;
}
return result;
};
const result = visit(input, "", 0);
return result && typeof result === "object" && !Array.isArray(result)
? (result as Record<string, unknown>)
: {};
}

export interface OtlpValue {
stringValue?: string;
intValue?: string | number;
doubleValue?: number;
boolValue?: boolean;
arrayValue?: { values?: OtlpValue[] };
kvlistValue?: { values?: OtlpAttribute[] };
}
export interface OtlpAttribute {
key?: string;
value?: OtlpValue;
}

export function decodeOtlpAttributes(
attributes: OtlpAttribute[] | undefined,
): Record<string, unknown> {
const decode = (v: OtlpValue | undefined, depth = 0): unknown => {
if (!v || depth > 6) return undefined;
if (v.stringValue !== undefined) return v.stringValue;
if (v.intValue !== undefined) return Number(v.intValue);
if (v.doubleValue !== undefined) return v.doubleValue;
if (v.boolValue !== undefined) return v.boolValue;
if (v.arrayValue)
return (v.arrayValue.values ?? [])
.slice(0, 64)
.map((value) => decode(value, depth + 1));
if (v.kvlistValue)
return Object.fromEntries(
(v.kvlistValue.values ?? [])
.slice(0, 128)
.filter(
(a) =>
a.key && !/^(?:__proto__|constructor|prototype)$/.test(a.key),
)
.map((a) => [a.key!, decode(a.value, depth + 1)]),
);
return undefined;
};
return sanitizeRuntimeContext(
Object.fromEntries(
(attributes ?? [])
.slice()
.sort(
(a, b) =>
Number(/^autter\.(?:operation|event)\./.test(b.key ?? "")) -
Number(/^autter\.(?:operation|event)\./.test(a.key ?? "")),
)
.slice(0, 128)
.filter(
(a) => a.key && !/^(?:__proto__|constructor|prototype)$/.test(a.key),
)
.map((a) => [a.key!, decode(a.value)]),
),
);
}
2 changes: 1 addition & 1 deletion packages/otlp-ingester/src/continuous-detection.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -50,7 +50,7 @@ test("sampled handled exception markers survive OTLP normalization", () => {
{ key: "autter.sampled", value: { boolValue: true } },
] }],
}] }] }] });
assert.deepEqual(result.occurrences[0]?.attributes, { "autter.handled": true, "autter.sampled": true });
assert.deepEqual(result.occurrences[0]?.attributes, { "exception.type": "ValueError", "autter.handled": true, "autter.sampled": true });
});

test("browser timing is aggregated; failed outcome becomes an issue", () => {
Expand Down
Loading
Loading