Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions packages/ai/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -75,6 +75,13 @@ result = await evals.run(

`LD_API_TOKEN` is required. Configure `LD_SDK_KEY` — or initialize your own client with `init_client(client=...)` — to emit one `$ld:ai:offline-evals:generation` event per generated row, plus one `$ld:ai:offline-evals:criterion` event per `(row, criterion)` when `criteria` are supplied, through the standard SDK event transport. The SDK reports scores; LaunchDarkly rules on them at ingest. A judge served by a different provider than `generation` needs a handler for it in `judge_handlers`. Each row's tool calls are recorded during generation and rendered into the judge's `{{message_history}}`, between the row input and the generated output, so a rubric can grade the tool trajectory as well as the final answer. Use `LD_API_BASE_URI` for staging or local management API traffic; it is separate from the SDK delivery setting `LD_BASE_URI`. Evaluation-run links use the explicit `ui_base_uri` option or `LD_UI_BASE_URI`, defaulting to `https://app.launchdarkly.com`; set it when the project is not in production, or a run created elsewhere still links to the production app. See the [core evaluations guide](../client/README.md#run-an-evaluation-from-code).

## Experimental features

Experimental features are not re-exported from this package. Import them from
`launchdarkly_ai_server.experimental`, which this package installs. For example, Agent Skills
lives in `launchdarkly_ai_server.experimental.skills`; see the
[Agent Skills guide](https://github.com/launchdarkly/python-ai-sdk/blob/main/packages/client/README.md#agent-skills-experimental).

---

All exports, types, and behaviors are identical to `launchdarkly-ai-server`. See the [core client README](../client/README.md) for the full API reference.
63 changes: 46 additions & 17 deletions packages/client/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -200,6 +200,20 @@ asyncio.run(main())
| `shutdown()` | Flush all events and telemetry, then close the client. Await before process exit. |
| `inspect_config(key, context)` | Read an AI Config variation without invoking the model. Never raises. Returns `{"enabled", "config", "meta"}`. |

`init_client` reads these option keys. Each overrides the environment variable of the same
purpose (see [Environment Variables](#environment-variables)); any other key is logged as a
warning and ignored.

| Option | Environment variable | Description |
|---|---|---|
| `sdkKey` | `LD_SDK_KEY` | LaunchDarkly server-side SDK key |
| `baseUri` | `LD_BASE_URI` | Polling base URI. Not used with `client=...` |
| `streamUri` | `LD_STREAM_URI` | Streaming URI. Not used with `client=...` |
| `eventsUri` | `LD_EVENTS_URI` | Events URI. Not used with `client=...` |
| `otlpEndpoint` | `OTEL_EXPORTER_OTLP_ENDPOINT` | OTLP endpoint spans are exported to |
| `serviceName` | `LD_SERVICE_NAME` | OTel `service.name` resource attribute |
| `environment` | `LD_ENVIRONMENT` | `deployment.environment` resource attribute |

### `config(**args)`

The primary entry point for AI config invocations. Accepts either a single handler or a list of handlers and routes to the correct one at invoke-time based on the flag variation's provider and mode.
Expand Down Expand Up @@ -345,7 +359,11 @@ asyncio.run(main())

---

### Agent Skills
### Agent Skills (experimental)

> **Experimental.** Import Agent Skills from `launchdarkly_ai_server.experimental.skills`;
> none of these names is exported from the package root. They may change in a minor release,
> so review the changelog when you upgrade.

Skills are versioned `SKILL.md` documents managed in LaunchDarkly and attached to AI Config
variations by reference. The SDK tells you which skills a config references, retrieves their
Expand All @@ -357,9 +375,9 @@ import asyncio
import hashlib
from pathlib import Path

from launchdarkly_ai_server import (
init_client, inspect_config, skill_refs, get_skill, write_skills,
InMemorySkillStore,
from launchdarkly_ai_server import init_client, inspect_config
from launchdarkly_ai_server.experimental.skills import (
InMemorySkillStore, get_skill, set_skill_store, skill_refs, write_skills,
)

SKILL_MD = "---\nname: PDF Extraction\n---\nExtract text from PDFs.\n"
Expand All @@ -376,10 +394,15 @@ async def main():
# hash does not match is withheld.
"contentHash": hashlib.sha256(SKILL_MD.encode("utf-8")).hexdigest(),
})
await init_client(options={"skillStore": store})
set_skill_store(store)
await init_client()

# 1. Which skills does this config reference? Pure projection — no I/O.
info = await inspect_config("doc-agent", {"kind": "user", "key": "user-123"})
if info["config"] is None:
# The config could not be resolved. Stop here: an empty reference list
# passed to write_skills would prune every skill it manages.
return
refs = skill_refs(info["config"]) # [SkillReference(key='pdf-extraction', version=2)]

# 2. Fetch content. Returns None rather than raising when a skill is unavailable.
Expand All @@ -402,11 +425,12 @@ project library**, which puts every skill's `description` into the agent's conte
skills no AI Config references and skills belonging to other teams. `write_skills(skill_refs(...), root)`,
as above, writes only what the resolved variation asked for.

**`skills` is a validated field.** Config parsing fails closed on a `skills` value that is not
a list of `{key, version}` objects (key matching `^[a-z0-9][a-z0-9-]*$`, version an integer
≥ 1): the whole variation is rejected, `inspect_config` returns `config: None`, and
`extract_variation` raises. If your variations carry a custom `skills` field of a different
shape, rename it before upgrading.
**`skill_refs` validates the `skills` field.** It raises `ValueError` when `skills` is present
but is not a list of `{key, version}` objects (key matching `^[a-z0-9][a-z0-9-]*$`, version an
integer ≥ 1), including `skills: null`. One bad entry rejects the whole field, so
`write_skills` never receives a partial list that would prune skills the config still
references. Config parsing does not check `skills`, so a malformed field never fails
`config().invoke()` or other core calls.

**Integrity is not optional.** Content is returned only when its sha256 (lowercase hex, over
the verbatim UTF-8 bytes) matches the delivered `contentHash`, its key and version revalidate,
Expand Down Expand Up @@ -542,7 +566,7 @@ retrieval, verification, and telemetry as `get_skill`, but reports which of five
happened instead of collapsing them all to `None`.

```python
from launchdarkly_ai_server import get_skill_result
from launchdarkly_ai_server.experimental.skills import get_skill_result

outcome = await get_skill_result("pdf-extraction")

Expand Down Expand Up @@ -590,13 +614,15 @@ authenticated with the environment's server-side SDK key.
```python
import os

from launchdarkly_ai_server import FDv2SkillStore, init_client, watch_skills
from launchdarkly_ai_server.experimental.skills import (
FDv2SkillStore, set_skill_store, watch_skills,
)

store = FDv2SkillStore(os.environ["LD_SDK_KEY"]).start()
if not store.wait_for_skills(timeout=10):
# No payload arrived. Reconciling now would find an empty store; see below.
print(f"skill delivery has not answered yet: {store.failed or 'still waiting'}")
await init_client(options={"skillStore": store})
set_skill_store(store)

# Materialize now, and re-materialize whenever delivery changes.
report, watcher = await watch_skills("*", ".claude/skills")
Expand Down Expand Up @@ -697,7 +723,7 @@ that skips verification.

| Export | Description |
|---|---|
| `skill_refs(config)` | Project a config's `skills` array into `list[SkillReference]`. Pure — no client, store, or network needed. Returns `[]` when the field is absent. A `skills` field that is present but not a list (including `null`) fails the config parse instead, so an unreadable field never reaches a pruning reconcile as "no skills". |
| `skill_refs(config)` | Project a config's `skills` array into `list[SkillReference]`. Pure — no client, store, or network needed. Returns `[]` when the field is absent or the config is not a dict. Raises `ValueError` when the field is present but malformed (including `null`), so an unreadable field never reaches a pruning reconcile as "no skills". |
| `get_skill(key, *, version=None)` | One verified skill, or `None`. `version=None` means newest available; a specific `version` matches exactly. Raises only when no store is configured. |
| `get_skill_result(key, *, version=None)` | The same retrieval, reporting **why**: a frozen `SkillOutcome` with `.skill`, `.reason` (`ok` / `absent` / `integrity_failure` / `store_unavailable` / `wrong_version`), and `.detail`. See *Failing closed on tampering* above. Raises only when no store is configured. |
| `get_skills(refs)` | Batch form. Accepts `SkillReference` values and bare key strings (string = latest). Results follow input order; missing or unverifiable entries are omitted. |
Expand All @@ -709,9 +735,12 @@ that skips verification.
| `watch_skills(skills, root, *, debounce=0.5, on_reconcile=None, …)` | `write_skills` plus a re-reconcile on every delivery change, so revocation takes effect within `debounce` rather than at the next restart. Returns `(initial report, SkillWatcher)`; close the watcher when done. `debounce` is in **seconds**, non-negative and finite. `on_reconcile` receives each *subsequent* report. One watcher per root. |
| `StoreDiagnostics` | What the transport has seen: `payloads_transferred`, `skill_objects_received`, `objects_ignored`, `objects_revoked`, `payloads_ignored`, `hashless_objects`, `connection_failures`, `last_error`. |

Configure the store with `init_client(options={"skillStore": store})`. With none configured,
the accessors raise `RuntimeError` explaining what to do, and `write_skills` reports the
failure (or raises, with `on_unavailable="raise"`). `shutdown()` clears it.
Configure the store with `set_skill_store(store)`. It applies on every call, before or after
`init_client`, and `None` is ignored rather than clearing a configured store. Anything else
without callable `get_object` and `all_objects` raises `TypeError`. Replacing a store does not
close the previous one, and a running watcher keeps the store it started with. With none
configured, the accessors raise `RuntimeError` explaining what to do, and `write_skills`
reports the failure (or raises, with `on_unavailable="raise"`). `shutdown()` clears it.

`ReconcileReport.actions` holds one `ReconcileAction` per outcome (`written`, `updated`,
`skipped_current`, `removed`, or `error`), each with `key`, `version`, the resolved `path`, and
Expand Down
37 changes: 23 additions & 14 deletions packages/client/agents.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,7 +61,6 @@ from launchdarkly_ai_server import (
TrackData, UsageDict, HandlerResult, HandlerStreamEvent,
StreamEvent, StreamChunkEvent, StreamDoneEvent, ExecuteStreamEvent, ExecuteStreamDoneEvent,
VariationMeta, InitClientOptions, JudgeResult, ParseResult, ParseSuccess, ParseFailure,
Skill, SkillReference, ReconcileAction, ReconcileReport,
)

# Utilities
Expand All @@ -79,15 +78,24 @@ from launchdarkly_ai_server import execute_and_track, execute_and_stream, wrap_t
# Entry points
from launchdarkly_ai_server import config, graph, resolve_graph, init_evaluations

# Agent Skills
from launchdarkly_ai_server import (
# Agent Skills (experimental: none of these is exported from the package root)
from launchdarkly_ai_server.experimental.skills import (
set_skill_store, SkillStore, InMemorySkillStore, FDv2SkillStore, StoreDiagnostics,
skill_refs, get_skill, get_skill_result, get_skills, all_skills, write_skills,
SkillStore, InMemorySkillStore, SkillOutcome,
watch_skills, SkillWatcher,
Skill, SkillReference, SkillOutcome, ReconcileAction, ReconcileReport,
SKILL_FILENAME, MANIFEST_FILENAME, MANIFEST_VERSION,
ReconcileActionKind, OnUnavailable, SkillOutcomeReason, # the three closed-set unions
)
```

An experimental feature is exported only from its module under
`launchdarkly_ai_server.experimental`, and its names may change in a minor release. No core
public type, function or `init_client` option may name it. Core reaches it only through an
internal hook (as `shutdown` clears the skill store), and an error there is logged, never
raised into the core call. The experimental part is the SDK's API, not the LaunchDarkly data
model: `AiConfigRep` still documents the `skills` references a config carries.

`MAX_SKILL_CONTENT_BYTES` and `SKILL_OBJECT_KIND` are deliberately **not** exported; both stay
internal to `skills_core`:

Expand All @@ -97,7 +105,7 @@ internal to `skills_core`:
would imply otherwise. An adapter that needs it imports it from
`launchdarkly_ai_server.skills_core`.

When adding a new export, add it to `__init__.py`'s imports and `__all__`. Handler packages must never import from sub-paths (e.g. `launchdarkly_ai_server.client`).
When adding a new core export, add it to `__init__.py`'s imports and `__all__`. An Agent Skills export goes in `experimental/skills.py` instead, and in `EXPERIMENTAL_SKILLS_SURFACE` in `tests/test_skills.py`. Handler packages must never import from sub-paths (e.g. `launchdarkly_ai_server.client`).

---

Expand Down Expand Up @@ -197,11 +205,13 @@ Three layers, in increasing order of blast radius:

1. **Reference discovery** — `skill_refs(config)` projects the config's `skills` array into
typed `SkillReference` values. Pure: no network, no client, no store, no telemetry.
Validation of the array itself lives in `parse_ai_config` and is **fail closed** — one
malformed reference fails the whole config parse.
It also validates the array, and **fails closed**: a present but malformed field
(including `null`, or one bad entry) raises `ValueError` rather than returning a partial
list that would authorize a prune. `parse_ai_config` deliberately does not check `skills`,
so an experimental field cannot fail a core config call (TESTING.md §0.3).
2. **Content accessors** — `get_skill`, `get_skill_result`, `get_skills`, `all_skills` read
through the `SkillStore` seam. Configure a store with
`init_client(options={"skillStore": store})`; with none configured the accessors raise
through the `SkillStore` seam. Configure a store with `set_skill_store(store)`; with
none configured the accessors raise
an actionable `RuntimeError`. A delivery transport can be added behind the seam
without touching the public API.
3. **Materialization** — `write_skills(skills, root)` writes `<root>/<key>/SKILL.md` and
Expand Down Expand Up @@ -519,8 +529,8 @@ Store data is **untrusted input**; the transport is not part of the trust bounda
- **Those two bounds live in `_key_rejection_reason`, not in the key grammar, and must not
move.** `is_valid_skill_key` / `skill_key_rejection_reason` deliberately admit an over-long
or reserved key, because:
- `parse_ai_config` fails closed on a bad `skills` entry, so a grammar rejection would
invalidate the *entire* AI Config for a Linux customer over a Windows-only constraint;
- `skill_refs` fails closed on a bad `skills` entry, so a grammar rejection would
reject *every* skill reference for a Linux customer over a Windows-only constraint;
- it would also shrink `skill_refs`, which authorizes a prune, turning "fails to write on
Windows" into "deleted on Linux".

Expand Down Expand Up @@ -638,9 +648,8 @@ three signals are emitted by the `record_*` functions there; nothing else calls
the allowlist is enforced in one place.

**Injection goes through `skills.py`.** `_set_store`, `_set_emitter_for_testing` and
`_clear_state` delegate to `skills_core`. `init_client`, `shutdown` and tests use those
(`skills._set_store(store)` is the setter `init_client` uses); none should reach into
`skills_core` directly.
`_clear_state` delegate to `skills_core`. `set_skill_store`, `shutdown` and tests use
those; none should reach into `skills_core` directly.

### Descriptor-pinned filesystem access

Expand Down
52 changes: 0 additions & 52 deletions packages/client/src/launchdarkly_ai_server/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -63,24 +63,6 @@
resolve_tools,
)
from .sdk_info import SDK_INFO_CONTEXT, SDK_INFO_EVENT, register_ai_sdk_package
from .skills import (
InMemorySkillStore,
all_skills,
get_skill,
get_skill_result,
get_skills,
skill_refs,
)
from .skills_core import SkillStore
from .skills_fdv2 import FDv2SkillStore, StoreDiagnostics
from .skills_fs import (
MANIFEST_FILENAME,
MANIFEST_VERSION,
SKILL_FILENAME,
OnUnavailable,
write_skills,
)
from .skills_watch import SkillWatcher, watch_skills
from .tracking import execute_and_stream, execute_and_track, wrap_tool_handlers
from .types import (
NATIVE_TOOL_KEY,
Expand Down Expand Up @@ -111,13 +93,6 @@
ProviderGraphResponse,
ProviderHandler,
ProviderResponse,
ReconcileAction,
ReconcileActionKind,
ReconcileReport,
Skill,
SkillOutcome,
SkillOutcomeReason,
SkillReference,
StreamChunkEvent,
StreamDoneEvent,
StreamEvent,
Expand Down Expand Up @@ -181,11 +156,6 @@
"ProviderGraphResponse",
"ProviderHandler",
"ProviderResponse",
"ReconcileAction",
"ReconcileReport",
"Skill",
"SkillOutcome",
"SkillReference",
"StreamChunkEvent",
"StreamDoneEvent",
"StreamEvent",
Expand Down Expand Up @@ -284,28 +254,6 @@
"graph",
"resolve_graph",
"GraphInstance",
# skills
"skill_refs",
"get_skill",
"get_skill_result",
"get_skills",
"all_skills",
"write_skills",
"SkillStore",
"InMemorySkillStore",
# skills — LaunchDarkly delivery store and on-change re-reconcile
"FDv2SkillStore",
"StoreDiagnostics",
"watch_skills",
"SkillWatcher",
# skills — literal types for typed consumers
"ReconcileActionKind",
"OnUnavailable",
"SkillOutcomeReason",
# skills — on-disk filenames and manifest version
"SKILL_FILENAME",
"MANIFEST_FILENAME",
"MANIFEST_VERSION",
]

register_ai_sdk_package("launchdarkly-ai-server", __version__)
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
"""
Experimental features of the LaunchDarkly AI SDK.

Each feature is a submodule, for example ``launchdarkly_ai_server.experimental.skills``.
Names imported from here may change in a minor release, so review the changelog
when you upgrade. Names from the package root change only in a major release.
"""
70 changes: 70 additions & 0 deletions packages/client/src/launchdarkly_ai_server/experimental/skills.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
"""
Agent Skills (experimental).

Configure a store with ``set_skill_store``, then read verified skill content with
``get_skill``, ``get_skills`` or ``all_skills``, or write it to disk with
``write_skills``. These names may change in a minor release, so review the
changelog when you upgrade.
"""

from ..skills import (
InMemorySkillStore,
all_skills,
get_skill,
get_skill_result,
get_skills,
set_skill_store,
skill_refs,
)
from ..skills_core import SkillStore
from ..skills_fdv2 import FDv2SkillStore, StoreDiagnostics
from ..skills_fs import (
MANIFEST_FILENAME,
MANIFEST_VERSION,
SKILL_FILENAME,
OnUnavailable,
write_skills,
)
from ..skills_watch import SkillWatcher, watch_skills
from ..types import (
ReconcileAction,
ReconcileActionKind,
ReconcileReport,
Skill,
SkillOutcome,
SkillOutcomeReason,
SkillReference,
)

__all__ = [ # noqa: RUF022
# configuration
"set_skill_store",
"SkillStore",
"InMemorySkillStore",
# LaunchDarkly delivery store and on-change re-reconcile
"FDv2SkillStore",
"StoreDiagnostics",
"watch_skills",
"SkillWatcher",
# accessors
"skill_refs",
"get_skill",
"get_skill_result",
"get_skills",
"all_skills",
"write_skills",
# value types
"Skill",
"SkillReference",
"SkillOutcome",
"ReconcileAction",
"ReconcileReport",
# literal types for typed consumers
"ReconcileActionKind",
"OnUnavailable",
"SkillOutcomeReason",
# on-disk filenames and manifest version
"SKILL_FILENAME",
"MANIFEST_FILENAME",
"MANIFEST_VERSION",
]
Loading
Loading