Skip to content

feat(mcp): add the MCP server module, and fix the ordering and NIP-11 bugs it exposed - #549

Merged
tcheeric merged 53 commits into
mainfrom
release/2.1.0
Aug 31, 2026
Merged

tcheeric merged 53 commits into
mainfrom
release/2.1.0

Conversation

@tcheeric

Copy link
Copy Markdown
Owner

Summary

Adds nostr-java-mcp, a Model Context Protocol server that exposes the SDK to LLM agents,
and brings the release line from 2.0.8 to 2.3.1. Along the way it fixes two bugs in the existing
transport that were reachable from ordinary use, and puts the documentation under test.

This is large (226 files) because it spans four releases. See Where to start reviewing below.

Type of change

  • feat - New feature (non-breaking)
  • fix - Bug fix (non-breaking)
  • docs - Documentation only
  • test - Adding or updating tests
  • chore - Build, CI, or tooling changes

What changed?

The MCP server (nostr-java-mcp). 22 tools covering reading, publishing, identity
lifecycle, subscriptions and NIP-17 messaging, over stdio or HTTP, plus guided prompts and a
container image. Its safety model is the interesting part, because publishing to Nostr cannot be
undone:

  • Writes are confirmed by default: the agent gets a preview and a token, and nothing is
    published until it calls again with that token. A hallucinated post becomes a no-op.
  • Private keys never cross the tool boundary. nostr_import_identity accepts no key
    material at all
    : the agent names a location the server reads for itself, so a key never
    enters the model's context or the host's logs.
  • Binding a server to one identity decrypts only that key, so another identity's key is absent
    from the process rather than merely out of policy.

Two bugs in existing code, both found by testing against a real relay rather than a fake:

  1. Inbound frames could arrive out of order. Every frame was dispatched on a freshly started
    virtual thread, so an EOSE could overtake the stored events it follows. Since every caller
    ends a query on that signal, any query could silently return a partial answer
    indistinguishable from the relay holding less data. Frames are now drained in order per
    listener.
  2. nostr_relay_info could not read any relay's NIP-11 document. Java's HTTP client offers
    an HTTP/2 upgrade, which a relay reads as a botched websocket handshake and answers 400.
    curl worked, which is why nothing noticed. Now pinned to HTTP/1.1.

-DnoDocker=true never worked. Documented in five places and passed by
scripts/release.sh --no-docker, it set a property nothing reads, so the container-backed tests
ran anyway. Anyone following those instructions without Docker got a failing build and no
explanation. Both now use -Pno-docker, verified by running verify with each.

Documentation. Rewritten and reorganised around Diátaxis, and now tested:
DocumentationAccuracyTest fails when a guide names a type or method that does not exist, a link
resolves to nothing, an install snippet quotes the wrong version, or a page is orphaned from the
index. It immediately caught the reference teaching BaseMessage.read(json), a method that has
never existed. CONTRIBUTING.md is removed: it taught the pre-2.0 architecture, including a
EventNostr base class that no longer exists.

Breaking changes

None. No public member was removed or altered in the already-released modules; the diff against
main for those modules is additions plus the frame-ordering fix.

Testing

mvn verify              → BUILD SUCCESS   147 unit + 67 integration tests
mvn -Pno-docker verify  → BUILD SUCCESS
  • Unit tests pass: mvn test
  • Integration tests pass: mvn verify (requires Docker)

Beyond the suite, three things were checked that unit tests cannot:

  • Every requirement run against the shipped jar. A harness drives the built artefact as an
    MCP host does, over stdio against a real relay, asserting one check per ticket item: 33/33,
    stable across three runs. It found a defect the class-level tests missed, where a bound server
    still advertised an identity argument, defeating the guarantee binding exists to provide.
  • A real model drives the tool surface. OllamaAgentIT offers all 22 tools to a local LLM
    and checks it reaches the right one from a plain-language request, and that it reads a publish
    preview as "not yet published". 25/25. This asks whether a model can use the tools, which a
    correct-but-unusable tool would fail.
  • Every safety guard mutation-tested. Each was broken deliberately to confirm the matching
    test fails: the confirmation token, removal-without-backup, key secrecy, subscription reaping,
    buffer draining, and the DM outcome reporting.

Where to start reviewing

  1. docs/explanation/nostr-java-mcp-spec.md — the design and its reasoning.
  2. nostr-java-mcp/src/main/java/nostr/mcp/identity/IdentityVault.java — the security boundary.
  3. nostr-java-mcp/src/main/java/nostr/mcp/write/WriteGuard.java — the single point every write
    passes through.
  4. nostr-java-client/.../NostrRelayClient.java — the frame-ordering fix, the one change
    affecting existing users.

Review focus

  • The frame-ordering fix touches shared transport code. The regression test was confirmed to
    fail against the old dispatch, but a second opinion on the per-listener queue is worth having.
  • Per-session identity binding over HTTP is deferred, not implemented, and the reason is
    recorded in the ticket: the SDK builds one server across sessions, and a half-enforced
    isolation boundary is worse than an absent one. Process-level binding ships and is stronger.
  • The branch name release/2.1.0 no longer matches its contents, which now end at 2.3.1.
    Happy to rename or re-target if you would prefer a different base.

tcheeric added 30 commits August 30, 2026 10:23
Introduces the NIP-59 rumor: an unsigned event that carries direct message
content before it is sealed and gift wrapped. Rumor deliberately does not
implement ISignable and holds no signature field, so the deniability
property NIP-59 depends on is enforced by the type system rather than by
convention. Ids are derived through EventSerializer, and JSON is handled by
Jackson, so canonical serialization has a single implementation.

Adds GenericEvent.update(long createdAt), which recomputes the id without
consulting the clock. The existing no-arg update() delegates to it, so no
call site changes behaviour. Without this seam a caller cannot set the
randomised past timestamp NIP-59 requires on seals and gift wraps, because
computing the id silently resets it to now.

Adds the NIP-17 and NIP-59 kind constants.

Verified against the worked example published in NIP-59: the derived rumor
id matches the specification exactly.

Refs #540
Adds GiftWrapper, whose two methods hide a rumor and reveal it, and
Nip59GiftWrapper, which seals a rumor to its recipient and wraps the seal
under a key generated for that one event and then discarded.

The receive path authenticates before it returns. A rumor carries no
signature, so the seal's signature is the only evidence of authorship: it
is verified, and the rumor's author is compared against the sealing key.
Without that comparison anyone could attribute a message to anyone by
editing an unsigned field, which NIP-17 calls out explicitly. The rumor id
is recomputed as a third check.

Timestamps are drawn from SecureRandom over the two days preceding the
event and never move into the future, so wraps cannot be correlated by
time and relays will not drop them for being future-dated.

Also fixes GenericEvent.getByteArraySupplier, which called update() and so
reset created_at during signing. That silently discarded any deliberately
chosen timestamp, making correct gift wrapping impossible even for a
caller using update(long).

Verified against the worked example published in NIP-59: wrapping its rumor
with the published ephemeral key reproduces the published wrap author, and
the resulting event opens to the exact rumor id the specification states.

Refs #541
Adds ChatMessage, the kind-14 rumor an application author actually works
with, carrying recipients, an optional conversation subject, and an
optional parent for replies. Recipients plus sender define a conversation,
so a participant added or removed starts a different one.

Adds DirectMessageService and its NIP-17 implementation. Composing a
message yields one gift wrap per participant rather than a single event,
because there is no shared envelope: each copy is separately encrypted,
which is what keeps a conversation's membership private. Making that
plurality visible in the signature stops a caller publishing one event and
silently dropping the rest.

The sender is a participant too. A sender who published only their
recipients' copies could never read the conversation back, since they
cannot decrypt a wrap addressed to someone else. composeByRecipient keeps
each participant paired with their own event, which phase 4 needs to
publish to the relays each recipient nominated.

Sending a message attributed to another identity is refused up front, since
the recipient would reject it as forged.

The service neither publishes nor retains anything, so a message can be
composed and verified without a relay.

Verified with the identities from the worked example published in NIP-17.

Refs #542
Adds DirectMessageRelayList, the kind-10050 event naming where someone
receives private messages, and planDelivery, which pairs each participant's
copy with the relays that should carry it.

NIP-17 permits delivery only to the relays a recipient nominated, and
forbids sending at all to someone who published no list. Publishing
elsewhere is not merely ineffective: it scatters a metadata-bearing event
across relays the recipient never chose while the one they actually read
may never see it. Phases 1 to 3 produced correct messages that were still
not safe to put on the network; this closes that gap.

No wrap is built for an unreachable recipient, so an event that must not be
sent cannot later be published by mistake. Such a recipient still appears
in the plan carrying no event, because silently omitting them is how a
message goes half-delivered without anyone noticing. That also separates
'this person does not accept private messages' from 'the relay was down'.

Relay lists are resolved through DirectMessageRelayLookup rather than by
querying relays directly, keeping message composition a pure function that
can be verified without a network.

Refs #543
Adds a how-to for sending and reading NIP-17 messages, covering the two
things callers most often get wrong: that composing a message yields one
event per participant including a copy addressed to the sender, and that a
kind-1059 subscription delivers wraps you cannot open, so reading an inbox
must skip failures rather than abandon the batch.

Deprecates the NIP-04 types. They hide a message's text but leave the
correspondents, the timing, and the message count public on every relay
that carries the event. Nothing is removed: NIP-04 remains functional for
reading existing conversations and for interoperating with clients that
send nothing else.

Refs #544
Each snippet printed in docs/howto/private-direct-messages.md appears here
in a form close enough to fail if the API changes under it. Documentation
that has drifted from the API is worse than none.

Refs #544
Minor bump: NIP-17 and NIP-59 support is additive, and the NIP-04
deprecation removes nothing.
Add ADRs 0001-0005 covering the introduction of a nostr-java-api
capability layer: module purpose and boundary, multi-relay failure
semantics, v1 service scope, pool concurrency and subscription
lifecycle, and pool membership, EOSE aggregation and resource
ownership.

Add docs/CONTEXT.md as the shared vocabulary for the module chain and
the new terms the design introduces, and link both from the docs index.
Synthesise ADRs 0001-0005 into a spec covering the problem, user
stories, implementation and testing decisions, and explicit
out-of-scope boundaries.

Break the spec into ten tickets in dependency order, starting with the
RelayConnection seam prefactor that every later ticket depends on and
ending with the NostrClient facade and documentation.
NostrRelayClient is the only way to reach a relay today, and it is a
Spring component that opens a real WebSocket. Code coordinating several
relays has nowhere to substitute relay behaviour, so scenarios where
relays disagree, one accepts, one rejects with a reason, one goes
silent, one drops mid-stream, cannot be tested without live relays.

Extract RelayConnection, exposing only what coordinating code needs
rather than mirroring the client's full surface, with a
RelayConnectionFactory resolving a URI to a connection.
NostrRelayClient implements it unchanged; ConnectionState is reused
from the transport package because it was public API before this seam
existed.

Add a scriptable FakeRelay fixture that routes payloads by
subscription id as the real client does, and reproduces the existing
subscription and routing test scenarios with no Mockito.

Refs ADR-0001.
Later multi-relay work is verified almost entirely against FakeRelay,
so a fake that diverges from NostrRelayClient would let those tests
pass while production fails. Add a contract test that drives one
subscription-routing scenario through both implementations via the
RelayConnection interface and asserts they agree; reverting the fake to
broadcasting makes it fail.

Fix the accompanying subscriber-eviction test, which asserted the wrong
ordering: it closed the only subscriber, so registration identifiers
never collided and reintroducing the collision left it green. It now
keeps a subscriber alive across the release, and fails when the bug
returns.
FakeRelay.reconnect() revived a closed connection in place, which
NostrRelayClient cannot do: it has no reopen path, so a closed
connection is terminal and callers must obtain a replacement from the
factory. Tests for auto-resubscribe would have been written against
behaviour production does not have.

Replace it with reconnected(), returning a new connection with the same
URI and scripted behaviour and no subscribers, so re-subscription stays
the caller's responsibility to prove.
Nostr is a multi-relay protocol, but a single connection gives a single
answer, so every caller wanting real delivery had to write fan-out,
result aggregation and timeout handling themselves.

RelayPool publishes concurrently on virtual threads and returns a
PublishResult recording, per relay, acceptance, rejection with the
relay's verbatim reason, a timeout, or unreachability. Partial failure
is ordinary data; a publish that no relay accepted throws
NoRelayAcceptedException carrying the same result, so an event that
reached nobody cannot be mistaken for a published one.

The timeout is a budget for the whole publish rather than for each
relay in turn, since spending it per relay let N stalled relays cost N
times the wait the caller asked for. OK frames are matched on event id,
so another event's acceptance is never credited to this publish, and a
relay that cannot be reached is reported as unreachable rather than as
a timeout.

Refs ADR-0002, ADR-0004.
A connection serves one request at a time, so two threads publishing
through the pool collided with NostrRelayClient's in-flight limit.
Each relay now has its own lock: concurrent publishes to one relay
queue, while publishes to different relays still overlap, so a slow
relay delays only its own queue.

Relay health becomes the pool's responsibility, as docs/CONTEXT.md
assigns it. Each relay's connection state is observable, and downed
relays are retried on a schedule the pool owns rather than one the
caller must remember to drive. Relays that drop after connecting are
returned to the downed set as well, since retrying only startup
failures would let a long-running pool quietly shrink.

Fixes a ConcurrentModificationException found by racing reconnection
against publishing: membership was held in plain LinkedHashMaps that a
publish iterated while a recovering relay rejoined from another thread.
Membership is now held concurrently, with configured order kept
separately so results stay predictable.

Refs ADR-0004.
A subscription spread over five relays receives every event five times,
ends its backlog five times, and loses a relay silently when one drops.
RelayPool.subscribe now presents that as a single stream.

Events are de-duplicated by id through a bounded window and delivered
already parsed, since de-duplication has to parse them anyway.
Bounding matters on exactly the long-lived firehose subscriptions that
most need de-duplication, and an evicted id being seen as new again is
the accepted cost.

Per-relay EOSE frames are aggregated into one end-of-backlog signal,
emitted once every relay has reported or a timeout expires, so a relay
that never replays cannot leave an application showing a loading state
forever. Subscriptions retain their filter so a relay that drops is
re-subscribed when it recovers, and report a malformed payload rather
than letting one relay's nonsense end a stream the others are serving.

Also widens the publish-budget assertion's margin: it measured wall
clock tightly enough to fail under load, and now discriminates on
five stalled relays instead of three, far from the per-relay cost.

Refs ADR-0002, ADR-0004, ADR-0005.
The SDK's modules are each deliberately narrow, so an application had
to assemble connection handling, fan-out, de-duplication and NIP-17
delivery itself, every time. The new module owns that work.

RelayPool gains runtime membership: relays join and leave a running
pool, and one borrowed to reach a recipient is released afterwards by
counting holders, so two overlapping deliveries do not cut each other
off and a long-running process does not accumulate connections.

RelayListLookup gives DirectMessageRelayLookup its first real
implementation, resolving kind-10050 lists through the pool, which is
what nostr-java-identity declared but could not satisfy without
acquiring a transport.

DirectMessagePublisher performs the delivery plan identity can only
produce, sending each gift wrap to its own recipient's relays and
reporting per recipient, since a group message reaching three of four
people has partly succeeded.

NostrClient ties them together behind an identity and a relay set.
Ownership follows construction: a pool it built is closed with it, one
handed in is left for its owner, so a container can share a pool across
clients.

Refs ADR-0001, ADR-0003, ADR-0005.
The transport dispatches each inbound payload on its own thread, so an
EOSE could overtake the stored events it follows. A subscription then
told its caller the backlog was drained while those events were still
arriving, which is exactly the guarantee the signal exists to give.
Subscriptions now handle one payload at a time.

Found by testing against a live relay rather than a fake: a relay list
lookup intermittently reported nothing for a list the relay was plainly
serving, because the lookup stopped waiting the moment EOSE arrived.

Adds the round-trip integration test that caught it, covering both
publishing and a full NIP-17 delivery, plus a wait strategy that holds
the container until it has actually stored an event. Waiting for the
port is not enough: it binds before database migration finishes, and on
some hardware a relay worker panics during startup leaving the relay
accepting connections but answering nothing. Both look like client bugs.

Documents the module in a how-to guide, the README, the architecture
overview and the changelog.
…ounts

Ticket 10 claimed the API reference covered NostrClient, RelayPool and
PublishResult, and that the known limitations were written down. Neither
was true: the reference documented none of the new types, and the
limitations existed only in Javadoc and ADRs where a user would not
find them.

Add reference entries for the client and pool APIs, and an architecture
section for nostr-java-api. Record the three behaviours that will
otherwise look like bugs: throughput to one relay is latency-bound,
a relay borrowed for a direct message is briefly shared, and
de-duplication is windowed so a very late duplicate can reach the
caller twice.

Update the module count and dependency chain in the codebase overview,
architecture and dependency-alignment docs, which still described four
modules. The v2.0.0 highlight keeps its original numbers, since it
describes that release rather than the present.

Every documented snippet was compiled against the real API, and every
documented signature, default and enum constant checked against source.
The test asserts stored events precede the end-of-backlog signal, but
removing the ordering lock leaves it green: FakeRelay delivers frames
synchronously, so the race the lock prevents cannot occur in it.
Claiming otherwise would overstate what the suite protects.

Note where the guarantee is really covered, which is the live-relay
integration test that found the defect to begin with. Making the fake
dispatch concurrently was tried and reverted: it broke the fixture's
own tests without ever forcing the race.
Adds nostr-java-api, the client-facing entry point, together with the
multi-relay pool it is built on: fan-out publishing with per-relay
outcomes, de-duplicated fan-in subscriptions, and NIP-17 direct message
delivery to each recipient's own relays.

Released as 2.2.0 rather than folded into 2.1.0. That version was never
tagged, but its changelog section is written and dated and its artifacts
exist locally, so absorbing a new module into it would change what
2.1.0 means to anyone who already has it.

Also moves the frame-ordering fix into this release's notes; it was
recorded under 2.1.0 but landed in this cycle.
The spec targeted 2.0.x and planned to build relay pooling, subscription
merging and NIP-17 itself. All three now ship in the SDK, so the module
consumes them instead of duplicating them.

Its one declared hard prerequisite is satisfied: NIP-17 landed in 2.1.0
and delivery to a recipient's own relays in 2.2.0, so phase 0 is gone
and the DM phase becomes adapter work over NostrClient.

Resolves two name collisions that would have put two different types of
the same name in one codebase. The spec's RelayPool becomes
RelayDirectory, which maps logical relay names and owns no connections,
and its DirectMessageService becomes McpDirectMessageService.

Records what the module no longer builds, and the three SDK limitations
that shape its design: single-request-in-flight per relay, transient
relay sharing during DM delivery, and windowed de-duplication.

Every SDK symbol the spec names was checked against source.
A specification is not executable, so its assumptions about this SDK
rot silently until somebody implements against them. Prototyping the
spec's own call sequences against a live relay found two claims in it
that were wrong.

subscribe() is asynchronous: it returns before any stored event or the
end-of-backlog signal arrives. The spec said a tool could return after
EOSE, which would have made every subscription stall for the backlog
timeout on an unresponsive relay.

A direct message to one recipient reports two outcomes, and a sender
who has published no relay list sees their own archival copy come back
UNREACHABLE while the recipient is DELIVERED. Reported naively that
reads as 'one of two delivered', telling a user their message failed
when it arrived.

McpSpecAssumptionsIT now asserts both, plus per-relay publish outcomes,
total failure carrying its result, mid-subscription delivery, relay-list
lookup, and signing as a named identity. Verified to fail when the
assumption breaks: making subscribe block turns it red.
…py claim

The isolation argument still reasoned about NostrRelayClient's futures,
from before RelayPool existed. Its conclusion holds but its evidence was
stale: capping an identity at one in-flight operation is now what the
pool does deliberately, per relay, so citing it as a drawback of threads
was misleading.

McpSpecAssumptionsIT asserted that a one-recipient message yields two
outcomes, but not the behaviour §6.2 actually warns about: that a sender
who published no relay list sees their own archival copy come back
UNREACHABLE while the recipient is DELIVERED. That is now asserted, so
the warning rests on an observation rather than on reasoning.

Also re-checked every component name the spec defines against the SDK;
all nine are free.
Records the decisions and the reasoning behind each, since the reasoning
is what a reader needs when a trade-off resurfaces.

os-keychain becomes the default keystore: it needs no passphrase so it
does not block unattended startup, with encrypted-file kept as the
portable fallback and selected explicitly in the compose file, where
there is no keychain to talk to.

Identity lifecycle tools stay. The design already removes the reason to
withhold them: no tool accepts or returns key material, creation is
reversible, and the irreversible operations are two-step guarded.

Bound processes share relay connections through a broker, so N
identities do not open N websockets to one relay. The broker shares
transport only; sharing a RelayPool would share subscriptions and
per-relay state, undoing the isolation the processes exist for.

Identity mutation gains its own identity-policy switch, because
destroying a key and publishing an event have different risk profiles
and collapsing them forces a choice between a useful agent and a
protected keystore. Confirmation tokens do not expire, no identity is
auto-created on first run, and HTTP transport auth is deferred with the
localhost-only constraint stated plainly rather than implied.

Verified the ContactList question: no such type exists, only the
constant Kinds.CONTACT_LIST. A value type over kind-3 belongs in
nostr-java-event following DirectMessageRelayList, making it this
module's one remaining SDK prerequisite.
…honestly

Resolving the open questions asserted a shared relay-connection broker
without defining it, which hid the hard part: bound processes are
separate address spaces, so a shared broker is another process rather
than an object. It gets a name, a place in the key components, and its
real costs stated: its own lifecycle, a single point of failure for
every identity's relay access, and visibility of all their traffic. It
is opt-in rather than default, since a handful of identities should
just open a handful of connections.

The non-goals still claimed no NIP work remained, which the ContactList
answer had just made false. Both that section and the decisions table
now name kind-3 as the module's one SDK prerequisite, blocking the
social phase only.

Audited all nine decisions against the whole document: each is stated
and none is contradicted.
Auditing a specification with greps proves it is self-consistent, not
that its instructions work. Both new claims were testable, so they were
built instead of asserted.

The connection broker: two independent RelayPool instances, fed by one
websocket through a RelayConnection, both publish successfully to a real
relay, and the second keeps working after the first closes its view.
That last part is the claim that matters, since a broker owning the
transport is exactly what distinguishes it from a pool owning it.

The kind-3 contact list: a value type written to the DirectMessageRelayList
pattern round-trips through a real relay, published and recovered with
its p tags intact. This is why §12 can call the prerequisite small: one
value type, no new infrastructure.

Both guards were mutation-verified. The broker test initially passed
even when the shared connection was closed on release, because both
publishes had already completed; it now publishes after a release, and
fails when the connection is torn down.
Deferring HTTP transport authentication is only safe if the constraint
that replaces it is visible. It was stated once, in the decision itself,
while §7.2 tells a deployer to run HTTP inside a container, which is
exactly where it goes wrong: publishing a port reaches every interface
by default.

The constraint now appears where each reader meets it. bind-address
defaults to 127.0.0.1 in the configuration, the safety model lists an
exposed transport as a threat and warns at startup on a non-loopback
bind, and the compose file publishes to the host loopback only.

A deferral that lives in one section of a design document is not a
deferral, it is a trap.
Each ticket cuts a complete path rather than a layer, so every one is
demoable on its own: the skeleton ships a working tool over stdio, the
read path ships three tools with the argument conventions later tools
inherit, and so on.

Two depart from the spec's six phases. Phase 1 is split into skeleton,
keystore and CLI-plus-binding, since it was far too large for one
sitting. Phase 3 is split by risk, keeping publishing separate from
keystore mutation, which is also why the lifecycle tools wait on
WriteGuard: they reuse its two-step confirmation rather than inventing
a second one.

The ContactList prerequisite leads, since it is the only SDK work the
module needs and it unblocks the social phase.

Graph verified acyclic, with no missing blockers and no forward
references.
Kinds.CONTACT_LIST existed but nothing modelled the event, so reading a
follow list meant parsing p tags by hand. This is the only SDK work the
planned MCP module needs.

Each entry keeps all three parts NIP-02 defines rather than the key
alone. The relay hint is how a client finds someone it has never seen,
and the petname is how it shows a readable name without a global
registry; an implementation that reads only the key discards both on
every round trip, which an earlier prototype of this type did. Both are
reported absent rather than blank, since a list routinely carries
[p, key, , ], and a petname is emitted in its third position even
when there is no hint, because NIP-02 reads parameters positionally and
a shifted petname would be read as a relay.

Entries keep their order, since NIP-02 asks that new follows be
appended so the list reads chronologically. A duplicated key keeps its
first entry, and an entry with no key is discarded rather than making a
whole list unreadable.

Guards are mutation-verified: dropping the hint and petname, removing
the kind check, accepting any tag as a contact, and keeping duplicates
each turn a test red.

Closes ticket 01.
tcheeric added 23 commits August 30, 2026 16:42
First slice of nostr-java-mcp: an MCP host launches the server, sees
its tool list, and calls nostr_list_relays. That one tool needs no
keystore and writes nothing, so it proves transport, registration,
configuration and the SDK underneath without risking anything.

The module adapts nostr-java-api rather than reaching past it. Going
straight to NostrRelayClient would rebuild connection management,
result aggregation and de-duplication the SDK already owns, and get
them wrong in the ways it already learned about.

Tools are registered one class each, because policy and single-identity
mode will both work by not registering a tool rather than rejecting a
call, and that only holds if there is one list to review. A duplicate
name is rejected rather than silently resolved, since which of two
tools an agent got would otherwise depend on iteration order.

Pins jackson-annotations to 2.21. Jackson 3 keeps its annotations under
the 2.x coordinate and needs the JsonSerializeAs introduced there;
nearest-wins resolution was handing the SDK's mapper an older jar and
the server died on its first JSON frame. Only the end-to-end test
caught it: every unit test passed while no host could talk to it.

Closes ticket 02.
…keys

An MCP server signs on an agent's behalf, so the private key is the whole
security boundary: the agent must be able to say "sign as alice" without
ever being able to read alice's key. IdentityVault holds that line. It
loads key material once at startup, hands out only IdentitySummary
(alias, public key, source), and wipes the material on close.

Where a key comes from is a KeySource, so operators can choose their own
trust model: the OS keychain by default, an encrypted PKCS#12 file, or
environment variables for containers. KeySources maps configuration to a
backend in one place, so adding a backend is a new case rather than an
edit spread across startup.

Choosing an identity is explicit. A single configured identity is the
default, since naming the only option is ceremony, but several with no
stated default stay ambiguous and the tool asks rather than guessing
which key signs. A configured default naming no identity fails at
startup, where it is a typo, not at first use, where it is a mystery.

ToolSurfaceSecrecyTest sweeps every registered tool's schema, output and
errors for key material, so a future tool that leaks one fails here. It
was checked against a deliberately leaking tool, and the vault's wipe,
ambiguity and startup-validation guards were each checked by breaking
them and watching the corresponding test fail.
… CLI

The multi-identity server guards against posting as the wrong account;
it cannot remove the risk, because once a default exists an agent may
still name any alias it holds. Binding removes it. Setting
nostr.mcp.identity ties a process to one alias, and the restriction is
applied before decryption rather than after: a bound process reads only
its own entry, so another identity's key is absent from the heap instead
of merely out of policy. With one identity there is nothing to name, so
there is no wrong name to give.

Binding presupposes the key exists, so the same jar administers one.
KeyAdminCli offers keygen, import, list and remove, deliberately as a
command rather than a tool, which keeps key administration out of every
agent's reach. Keys are read from stdin, never arguments, since argv is
visible to every user on the host, and a generated key is stored and
wiped without being printed. IdentityStore is the seam the future
lifecycle tools will share, so "remove an identity" has one meaning and
two front doors.

A bound server whose alias is missing refuses to start and says which
command creates it. An unbound one with an empty keystore still starts,
because that is how a person makes their first key.

Two bugs the integration test caught that the unit tests could not:
PKCS#12 rejects the RAW key algorithm, so keygen failed against a real
keystore while passing against an in-memory one, and the macOS keychain
branch passed the secret in argv. The encrypted-file backend now has
tests of its own against a real file. Golden files pin the bound and
unbound tool lists separately; registering an admin tool was shown to
break the unbound file while leaving the bound one green.
An agent can now answer questions about Nostr without being able to
change anything: nostr_query_events runs a bounded query, nostr_get_profile
resolves a public key or a NIP-05 address and decodes the kind-0 JSON, and
nostr_relay_info reads a relay's NIP-11 document so an agent can learn a
relay's rules without breaking one.

This establishes the argument conventions later tools inherit.
NostrIdentifier accepts hex or bech32 and refuses an nsec in a public
argument with advice to rotate, since by the time a tool sees one the key
is already exposed. TimeArgument normalises "24h", ISO-8601 and bare dates
to Unix seconds. ToolArguments coerces the shapes a model actually sends,
because a JSON schema is a suggestion to a language model. Queries are
bounded twice, by event count and by deadline, and reaching either bound is
reported: an agent that cannot tell "nothing matched" from "I stopped
looking" will report the first when the truth was the second.

The integration test against a real relay then failed in a way no fake
could have shown. Chasing it found an ordering bug in the transport, not
in the new code: every inbound frame was dispatched on a freshly started
virtual thread, so EOSE could overtake the stored events it follows. A
probe showed the end-of-backlog signal arriving with one or two of three
events delivered, varying run to run. Every caller ends its query on that
signal, so every such query could silently return a partial answer.

Frames are now queued per listener and drained in order, which restores
ordering without serialising unrelated listeners or blocking the receiving
thread. The regression test was confirmed to fail against the old
dispatch. RelayStoresEventsWaitStrategy moves into the client test-jar so
both modules share one copy rather than a fork.
An event a relay has accepted cannot be reliably withdrawn, since NIP-09
deletion is advisory. Publishing is therefore the one capability where a
model's ordinary failure mode, confidently doing something nobody asked
for, produces a permanent public result. write-policy: confirm is the
default: the first call signs the event, holds it, and returns a preview
with a token; only a second call carrying that token publishes. An agent
that invented the intention does not follow through, so the post never
happens. deny registers no write tool at all, because a tool an agent
cannot see cannot be talked into running, whereas a refusal it can see is
an invitation to rephrase.

The event is signed before it is previewed and the held event is the one
sent, so the content cannot drift between what the user approved and what
goes out. Tokens are random, single-use, and survive until used or until
the process ends: expiring them on a timer would fail an agent that
paused to ask its user, which is the behaviour we want to encourage.

Every write goes through WriteGuard rather than reaching the pool, so
policy, rate limiting, identity resolution and the audit log are
properties of the server rather than habits each tool must remember. A
new write tool inherits them by construction. Results follow the SDK:
any acceptance is success carrying the per-relay list, and only reaching
no relay at all is an error, since reporting a partial success as failure
would push an agent to republish something already on the network.

Verified against a real relay: an unconfirmed note leaves the relay
empty, a confirmed one is stored and readable back, and a published
profile reads as a profile. Both guards were mutation-checked. Removing
the confirmation requirement failed two integration tests, and reusing a
spent token failed its unit test.
An agent that can only use pre-configured keys is half a tool: "make me a
throwaway account for this project" is a natural request. But an identity
is the user's Nostr account, so the tools are shaped by the asymmetry
between operations rather than treated alike. Creating, renaming and
setting a default are reversible and freely allowed. Exporting a backup
and removing a key are not, and are guarded.

Import is the case worth dwelling on. nostr_import_identity accepts no
key material whatsoever, because an nsec in an argument would be written
to the conversation log and very likely sent to a third-party inference
API, which is the worst thing this module could do. Instead source names
a location the server reads for itself, a file it can then shred, its own
environment, or a terminal the agent cannot see. The model arranges an
import it never observes. Pasting a key anyway is refused with advice to
treat it as compromised, since by then it is already exposed.

Removal is the only irreversible tool here: an npub with no nsec is a
dead account and every event ever signed with it is orphaned. It confirms
in two steps and refuses outright when no backup exists, unless the
caller states the key is disposable. An agent told to "clean up the test
accounts" has no basis for that judgement, so the tool makes it say so
where a user can contradict it. Backup state is deliberately not
persisted: after a restart this server cannot know a backup file still
exists, and a stale reassurance is worse than asking again.

identity-policy is separate from write-policy and capped by it, so a
deployment can let an agent post without letting it destroy keys, and a
read-only server can do neither. A backend that cannot be written to
registers no lifecycle tools at all, since a tool that could only fail is
worse than one that is absent.

The secrecy sweep now walks the real tool surface rather than a curated
list, so it covers tools added later; it was confirmed to catch an import
tool that accepts an nsec. Both removal guards were mutation-checked.
"Watch my mentions" cannot be answered by a call that returns, and MCP
has no way to push into a tool result, so a subscription becomes a
stateful server resource: the tool opens it and returns an id, events
land in a buffer, and the agent drains them. Each is also exposed as
nostr://subscription/{id} with update notifications, so a host that
supports resources gets push and one that does not still polls.

The buffer is bounded, and what it drops it counts. That count is the
part that matters: an agent silently handed a gap will summarise a
partial feed as though it were complete, whereas one told it missed
thirty events can say so or narrow its filter. The count is monotonic
across drains for the same reason, since a gap that happened two polls
ago is still a gap in what the agent believes.

Reading drains, which is a command that also answers. Accepted
deliberately: the alternative refills the model's context with the same
events on every poll until nothing else fits.

Subscribing does not wait for the backlog. RelayPool.subscribe returns
before any stored event arrives, so blocking would stall on any relay
that never answers, and the result instead reports backlogDrained so an
agent can tell "nothing matched yet" from "still replaying". Relays that
drop out are recorded rather than silently narrowing coverage.

Idle subscriptions are reaped and the total is capped, because an agent's
session ends whenever its user closes a window without telling this
server. Idleness is measured from the last read, not the last event, so a
busy filter cannot keep a forgotten subscription alive forever.

Verified against a real relay: a note published after subscribing arrives,
a second read returns nothing, an overflow reports its drops, and an
empty read distinguishes a drained backlog from a replaying one. Reaping
and draining were both mutation-checked. The filter schema and decoding
now live in one place shared with the query tool, so an agent that has
learnt to filter for one can use the other unchanged.
An agent can now read a thread, read who someone follows, and send a
message only its recipient can open. Replies carry the NIP-10 tags that
make them replies, since without them a client shows an answer detached
from the conversation it answers.

Direct messages are NIP-17 only. NIP-04 exposes both correspondents and
the conversation to every relay, which is unacceptable for a tool acting
on someone's behalf, so it is not offered at all rather than offered with
a warning. An integration test confirms no gift wrap is signed by the
real sender, which is the property that hides who is talking to whom.

Two delivery facts are reported rather than smoothed over, both of which
were established against a live relay rather than reasoned about. A
recipient who published no kind-10050 relay list genuinely cannot be sent
to, so they are named as unreachable instead of silently skipped. And
every conversation includes a copy addressed to the sender, so a message
to one person yields two outcomes; a sender with no relay list of their
own sees that copy fail while the recipient succeeds. Reported naively as
"1 of 2 delivered" that would tell someone their message failed when it
arrived perfectly well, so the archival copy is reported separately and
its failure is phrased as advice about their other devices.

Reading messages is opt-in per identity. Decryption puts private
correspondence into the model's context and therefore the host's logs, so
it stays a decision for the person whose messages they are.

The vault boundary survives this. NIP-17 sealing derives shared secrets
and so needs an Identity rather than a signature, so rather than release
a key the vault builds the collaborator internally and returns only what
that collaborator exposes, which is composing and reading messages and
never the identity behind them. Both reporting rules were
mutation-checked against the live-relay tests.
…henticated

Adds a streamable HTTP transport for hosted deployments where the MCP
host does not launch the process itself, selected by configuration alone
so moving between transports needs no code change. It runs an embedded
servlet container rather than requiring one, since a server an operator
must deploy into a container is a server most people will not run.

The transport carries no authentication in this version. That deferral is
defensible only while the constraint replacing it stays visible, so the
default binding is loopback, a non-loopback bind warns at startup naming
what is at risk, and an unresolvable address is treated as unsafe rather
than given the generous reading. The same principle as the unprotected
keystore backend: the weaker choice should announce itself, because the
person who made it may not be the person reading the logs.

Per-session identity binding is deferred rather than delivered. The
streamable transport builds one server shared across sessions, so
filtering the surface per session needs a session-scoped tool surface the
SDK does not currently offer, and a half-enforced isolation boundary is
worse than an absent one. Process-level binding from ticket 04 remains
the supported answer and is the stronger guarantee anyway, since it is an
address space rather than a check.

The new how-to guide states the authentication constraint where a deployer
will read it, and its settings table is checked against the code: a test
extracts every key McpConfiguration reads and fails when one is
undocumented. Confirmed by adding a setting and watching it fail.
The module is launched as `java -jar`, both by an MCP host and by its own
key-admin CLI, but nothing built such a jar: the documentation promised a
command that did not work. It now ships one self-contained artefact,
since a host configuration naming a classpath would break whenever a
dependency changed.

The container is where the HTTP transport's missing authentication turns
dangerous, because publishing a port reaches every interface by default
and would silently override the server's own loopback binding. Every
mapping in the compose file therefore names 127.0.0.1 explicitly, with a
comment saying why removing it is unsafe. The image is distroless and
runs as a non-root user: this process holds private keys, so the smallest
surface is worth losing a shell for. A profile demonstrates one bound
container per identity, which is the strongest isolation the module
offers because it is an address space rather than a check.

Packaging exposed a real bug. Environment variable names were derived by
uppercasing a setting and replacing dots, but not hyphens, so
`write-policy` became `NOSTR_MCP_WRITE-POLICY` and every hyphenated
setting was unreachable from the environment. That includes the write
policy, the bind address and all the limits, which is to say the
container could not be configured at all. Hyphens are now translated too,
and a test asserts every setting maps to a name a shell can actually set.

Verified by running it: the shaded jar completes an MCP handshake and
runs the CLI, the image builds, and a container answers a real initialize
request over HTTP as `nonroot`.
A tool surface with no guidance makes a model explore by trial and error,
which on a public and irreversible medium is the wrong way to learn:
the mistakes are permanent and other people see them. Three prompts
encode the sequences that work, and each targets a specific mistake
rather than restating the tool list. compose-note names the confirmation
step models skip and warns against republishing a note that already
reached some relays. catch-up-feed reads the follow list before querying,
because a model left to itself queries broadly and then filters, which
returns strangers and misses the people actually followed. watch-mentions
distinguishes a replaying backlog from an empty one, which is how an
agent comes to report "nobody mentioned you" when it simply read too
early.

Identities and relays are also exposed as resources, so a host can put
the server's own configuration into context rather than spending a tool
call on something that cannot change mid-conversation.

The how-to now leads with single-identity mode, because it removes the
wrong-account risk instead of guarding it and matches how host configs
are written anyway; the unbound server is presented as the advanced case
that answers cross-identity questions. It also documents the limits
inherited from the SDK where someone will meet them: per-relay
throughput, windowed de-duplication, subscriptions not surviving a
restart, and the unauthenticated HTTP transport.

Both documents are kept honest by tests. Every registered prompt must
appear in the guide, and each prompt's warning is asserted by the
behaviour it prevents; both were confirmed by breaking them. Prompts and
resources are also verified over the real protocol through a subprocess
MCP client rather than only in the registry.
Binding a server to one identity is meant to remove the wrong-account
mistake rather than guard it: with one identity there is nothing to name,
so nothing can be named wrongly. But all five signing tools still
advertised an `identity` argument when bound, which left it merely
redundant instead of absent. An argument with exactly one acceptable
value is an invitation to pass a different one, so the guarantee the
ticket describes was not actually being delivered.

Found by running every requirement against the shipped jar through a real
MCP client, rather than against the classes through their own tests. The
existing suite passed because it asserted which tools a bound server
registers, never what those tools ask for. The end-to-end check has been
kept in .scratch alongside a unit test that fails on the old behaviour.
The module's tests exercise classes through their own seams, so they can
all pass while the artefact a user runs behaves differently. This drives
the built jar as an MCP host does, over stdio against a real relay, and
asserts one check per ticket checklist item.

It earned its place immediately by finding a defect the existing suite
could not: a bound server still advertised an `identity` argument, so the
mistake binding exists to make impossible was merely redundant. The
surface tests asserted which tools were registered, never what those
tools asked for.

Two things about the harness itself are worth stating. It starts from an
empty keystore each run, so a repeated run measures the product rather
than the residue of the last one. And it probes relay readiness by
publishing and requiring acceptance, not by connecting, because the
nostr-rs-relay startup panic leaves the port open while the relay
silently answers nothing; a connection check passes against exactly the
relay that will fail every test. Verified stable across three
consecutive runs.
Minor rather than patch or major, derived from what actually changed
rather than from the commit subjects alone: twelve feat commits add the
nostr-java-mcp module and the ContactList/Contact types, one fix corrects
relay frame ordering, and a diff of the already-released modules shows no
public member removed or altered. Nothing breaks for an existing consumer.

The MCP server's reported version was a constant in the source, which is
a copy of the pom that nothing keeps in step: this very bump would have
left it announcing 2.2.0 to every host, and no test would have noticed.
It is now filtered in from the pom at build time, with a test asserting
the two agree, and the running jar was checked to report 2.3.0 in its
initialize response.

Verified at the new version on both profiles, 144 unit and 46 integration
tests. One NostrClientRoundTripIT error during the first run was the
known nostr-rs-relay startup panic, which its own retry then passed and a
clean rerun did not reproduce.
Every other test here asks whether the tools work. This asks the question
the module exists to answer: whether a model can use them. A tool can be
correct and still unusable, because its name misleads, its description
omits what the model needs to decide, or its schema invites an argument
the model cannot supply. None of that is visible to a test that calls the
tool directly, and all of it decides whether the server is any good in an
agent's hands.

It runs a real model through Testcontainers and checks that it picks the
right tool unprompted, distinguishes querying from subscribing, reaches
for publish rather than a neighbour when asked to post, and reads a
confirmation preview as "not yet published". That last one matters most:
the whole write guard depends on a model understanding that a preview is
not a success, and until now nothing verified it.

The host's model cache is mounted rather than pulled, because a test that
downloads several gigabytes is a test people disable, and it skips
cleanly when no cache is present. It is tagged and excluded from the
ordinary build since it takes minutes; the exclusion had to go inside the
failsafe execution rather than beside it, because the parent declares its
configuration there and a plugin-level block is silently ignored.

The class documents what a pass does and does not prove. Names and
descriptions reinforce each other, so gutting one description alone still
passes; renaming both tools to nostr_alpha and nostr_beta shows the
descriptions carry real weight on their own. Read a pass as "the surface
is comprehensible", not "every word is load-bearing".
Asked whether all the tools were tested, the honest answer was no. A
coverage measurement put it precisely: eight of the twenty-two had only
ever had their registration checked. Appearing in the golden tool list
guarantees that a class compiles, not that it works.

Calling them found nostr_relay_info broken against every real relay.
Java's HTTP client offers an HTTP/2 upgrade, and a relay serves its
NIP-11 document from the same host and port as its websocket endpoint, so
it read those headers as a botched websocket handshake and answered 400.
curl worked, which is why nothing noticed. The client is now pinned to
HTTP/1.1. Fixing that surfaced a second one: a relay name that was
neither configured nor a URI reached the HTTP client and failed with an
IllegalArgumentException about an undefined scheme rather than a code the
agent could act on.

A third failure was mine, not the product's. The test assumed two
identities are ambiguous, but the first one created becomes the default,
which is deliberate and better. The test now asserts what the tool is
actually for: moving that choice afterwards, checked by which key signs
the next note.

ToolCoverageTest keeps the gap from reopening by failing when any
registered tool is never called. It is a floor, not a quality claim: it
says each tool has been invoked, not that the interesting cases are
covered.
Asked whether the tools were tested with the LLM and with relays, the
measured answer was 19 of 22 against a relay and 4 of 22 against a model.
Both numbers are now 22, and a coverage test fails when either slips,
because the honest answer to that question should come from a
measurement rather than an impression.

The relay gap closed with tests for the three tools that had only been
called against fakes, and each asserts something a stand-in cannot show:
a follow list keeps its relay hint and petname through a real round trip,
a removed identity can no longer sign while the note it sent before
removal is still on the relay, and the relay list reports a genuinely
connected relay as connected.

The model gap closed with a case per tool, phrased as a user would ask,
with all twenty-two offered at once. Offering the whole surface is the
point: the risk is that two tools read alike to a model, and that only
appears when both are on the table. Two tools are deliberately exempt,
with the reason recorded in the test: the escape hatch overlaps every
other publishing tool by design, and inviting a model to choose the one
irreversible tool is a bad habit to build into a suite.

One case failed and taught me something. Asked to watch for notes
"mentioning me", the model declined to call anything and asked which
identity was meant, which is correct behaviour on an ambiguous request.
The prompt was at fault, not the surface, so it now asks for something
answerable. Re-run afterwards: 25 of 25, no flakes.
scripts/doccheck.py verifies that every relative link resolves, every
backtick-quoted type name matches a real Java or TypeScript type, and every
environment variable bound in application config appears in the configuration
docs. It exits non-zero so it can gate a commit.

Where present, .doccheck-allow declares names that are deliberately not types in
this repository: third-party and platform types, and vocabulary from design
proposals that are not yet built.
Three fixes found while extending the checker across the rest of the stack:

- Index .ts and .tsx declarations and exported React components. Several repos
  ship a TypeScript client whose types are legitimately named in docs; without
  this the checker reported them all as unknown.
- Skip plans/, specs/, superpowers/, and archive/ directories. Those documents
  quote link lines destined for other files, whose relative paths are correct
  only from the target, so checking them reports the quoting rather than a
  broken link.
- Extend the built-in skip list with JDK, Spring, browser, and node types
  commonly named in prose.
The link checker special-cased http, https, and mailto. A nostr: URI in a README
was therefore resolved as a relative path and reported broken. Any scheme before
the first slash now counts as external.

Verified against a fixture: scheme URIs skip, a genuinely broken relative link
is still reported and still exits non-zero.
Rewrote the entry points and reorganised the index around Diataxis, then
did the part that actually matters: checked the documentation against the
code instead of reading it.

That found things reading would not. The API reference taught
BaseMessage.read(json), a method that has never existed anywhere in this
project. Five install snippets used <version><!-- X.Y.Z --></version>,
which Maven reads as an empty version, so nobody could copy them. The
getting-started guide told readers to depend on nostr-java-client while
the README told them nostr-java-api. MIGRATION.md linked to a
nostr-java-examples module that does not exist. Nine documents, including
all three operations guides, were unreachable from the index.

A worse one turned up in passing: -DnoDocker=true, documented in five
places and passed by scripts/release.sh --no-docker, sets a property
nothing reads. Running verify with it proved the container tests ran
anyway, so anyone following those instructions on a machine without
Docker got a failing build and no clue why. Both now use -Pno-docker.

Two fresh agents read the docs with no other context and answered a
newcomer's questions from them alone, which is where the uncopyable
versions and the module contradiction surfaced: I could not see them
because I already knew the answers. Compiling the examples caught one of
mine, a checked NoRelayAcceptedException the quick start silently
ignored.

DocumentationAccuracyTest keeps all of this from returning. It fails when
a guide names a type or method that does not exist, a link resolves to
nothing, an install snippet quotes the wrong version, or a page is
orphaned from the index. Each check was confirmed to fail when the defect
it targets is reintroduced.
A superseded checkout keeps its documentation as written; repairing its links
would imply the tree is maintained. A repository can now declare
'doccheck: skip-links' in .doccheck-allow to record that decision.

Opt-in only, verified against a fixture: without the marker a broken link is
still reported and still exits non-zero.
CONTRIBUTING.md documented an architecture this project abandoned in 2.0.
It told contributors to add a per-NIP facade extending EventNostr, to
update a NIP compliance matrix, and offered NIP01EventBuilder and
UserProfile as naming examples. None of those exist; a search for the
classes returns nothing. A guide that confidently teaches a design the
codebase has removed is worse than no guide, because a contributor
follows it and produces work that cannot be merged.

Its accurate half, the commit and pull request conventions, moved into
the codebase overview rather than being deleted with it. That page also
claimed pull requests target `develop` while CONTRIBUTING.md said `main`;
the remote's default branch is `main`, so the contradiction is resolved
in favour of what the repository actually does.

Patch rather than minor, derived from the diff rather than the commit
subjects: since 2.3.0 the only source changes are two nostr_relay_info
bug fixes and a release-script correction. No API was added or removed.

The bump also demonstrated the documentation test earning its place. It
failed the build because multi-relay-publishing.md still named 2.3.0 in a
snippet a reader would copy. Rather than fix that by hand a second time,
scripts/release.sh now updates copyable install snippets as part of the
bump, leaving placeholders, ranges and a migration guide's "previous
version" alone.
@tcheeric
tcheeric merged commit 7a50ac0 into main Aug 31, 2026
5 of 6 checks passed
@tcheeric
tcheeric deleted the release/2.1.0 branch August 31, 2026 20:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant