Repository navigation
feat(mcp): add the MCP server module, and fix the ordering and NIP-11 bugs it exposed - #549
Merged
Merged
Conversation
added 30 commits
August 30, 2026 10:23
Introduces the NIP-59 rumor: an unsigned event that carries direct message content before it is sealed and gift wrapped. Rumor deliberately does not implement ISignable and holds no signature field, so the deniability property NIP-59 depends on is enforced by the type system rather than by convention. Ids are derived through EventSerializer, and JSON is handled by Jackson, so canonical serialization has a single implementation. Adds GenericEvent.update(long createdAt), which recomputes the id without consulting the clock. The existing no-arg update() delegates to it, so no call site changes behaviour. Without this seam a caller cannot set the randomised past timestamp NIP-59 requires on seals and gift wraps, because computing the id silently resets it to now. Adds the NIP-17 and NIP-59 kind constants. Verified against the worked example published in NIP-59: the derived rumor id matches the specification exactly. Refs #540
Adds GiftWrapper, whose two methods hide a rumor and reveal it, and Nip59GiftWrapper, which seals a rumor to its recipient and wraps the seal under a key generated for that one event and then discarded. The receive path authenticates before it returns. A rumor carries no signature, so the seal's signature is the only evidence of authorship: it is verified, and the rumor's author is compared against the sealing key. Without that comparison anyone could attribute a message to anyone by editing an unsigned field, which NIP-17 calls out explicitly. The rumor id is recomputed as a third check. Timestamps are drawn from SecureRandom over the two days preceding the event and never move into the future, so wraps cannot be correlated by time and relays will not drop them for being future-dated. Also fixes GenericEvent.getByteArraySupplier, which called update() and so reset created_at during signing. That silently discarded any deliberately chosen timestamp, making correct gift wrapping impossible even for a caller using update(long). Verified against the worked example published in NIP-59: wrapping its rumor with the published ephemeral key reproduces the published wrap author, and the resulting event opens to the exact rumor id the specification states. Refs #541
Adds ChatMessage, the kind-14 rumor an application author actually works with, carrying recipients, an optional conversation subject, and an optional parent for replies. Recipients plus sender define a conversation, so a participant added or removed starts a different one. Adds DirectMessageService and its NIP-17 implementation. Composing a message yields one gift wrap per participant rather than a single event, because there is no shared envelope: each copy is separately encrypted, which is what keeps a conversation's membership private. Making that plurality visible in the signature stops a caller publishing one event and silently dropping the rest. The sender is a participant too. A sender who published only their recipients' copies could never read the conversation back, since they cannot decrypt a wrap addressed to someone else. composeByRecipient keeps each participant paired with their own event, which phase 4 needs to publish to the relays each recipient nominated. Sending a message attributed to another identity is refused up front, since the recipient would reject it as forged. The service neither publishes nor retains anything, so a message can be composed and verified without a relay. Verified with the identities from the worked example published in NIP-17. Refs #542
Adds DirectMessageRelayList, the kind-10050 event naming where someone receives private messages, and planDelivery, which pairs each participant's copy with the relays that should carry it. NIP-17 permits delivery only to the relays a recipient nominated, and forbids sending at all to someone who published no list. Publishing elsewhere is not merely ineffective: it scatters a metadata-bearing event across relays the recipient never chose while the one they actually read may never see it. Phases 1 to 3 produced correct messages that were still not safe to put on the network; this closes that gap. No wrap is built for an unreachable recipient, so an event that must not be sent cannot later be published by mistake. Such a recipient still appears in the plan carrying no event, because silently omitting them is how a message goes half-delivered without anyone noticing. That also separates 'this person does not accept private messages' from 'the relay was down'. Relay lists are resolved through DirectMessageRelayLookup rather than by querying relays directly, keeping message composition a pure function that can be verified without a network. Refs #543
Adds a how-to for sending and reading NIP-17 messages, covering the two things callers most often get wrong: that composing a message yields one event per participant including a copy addressed to the sender, and that a kind-1059 subscription delivers wraps you cannot open, so reading an inbox must skip failures rather than abandon the batch. Deprecates the NIP-04 types. They hide a message's text but leave the correspondents, the timing, and the message count public on every relay that carries the event. Nothing is removed: NIP-04 remains functional for reading existing conversations and for interoperating with clients that send nothing else. Refs #544
Each snippet printed in docs/howto/private-direct-messages.md appears here in a form close enough to fail if the API changes under it. Documentation that has drifted from the API is worse than none. Refs #544
Minor bump: NIP-17 and NIP-59 support is additive, and the NIP-04 deprecation removes nothing.
Add ADRs 0001-0005 covering the introduction of a nostr-java-api capability layer: module purpose and boundary, multi-relay failure semantics, v1 service scope, pool concurrency and subscription lifecycle, and pool membership, EOSE aggregation and resource ownership. Add docs/CONTEXT.md as the shared vocabulary for the module chain and the new terms the design introduces, and link both from the docs index.
Synthesise ADRs 0001-0005 into a spec covering the problem, user stories, implementation and testing decisions, and explicit out-of-scope boundaries. Break the spec into ten tickets in dependency order, starting with the RelayConnection seam prefactor that every later ticket depends on and ending with the NostrClient facade and documentation.
NostrRelayClient is the only way to reach a relay today, and it is a Spring component that opens a real WebSocket. Code coordinating several relays has nowhere to substitute relay behaviour, so scenarios where relays disagree, one accepts, one rejects with a reason, one goes silent, one drops mid-stream, cannot be tested without live relays. Extract RelayConnection, exposing only what coordinating code needs rather than mirroring the client's full surface, with a RelayConnectionFactory resolving a URI to a connection. NostrRelayClient implements it unchanged; ConnectionState is reused from the transport package because it was public API before this seam existed. Add a scriptable FakeRelay fixture that routes payloads by subscription id as the real client does, and reproduces the existing subscription and routing test scenarios with no Mockito. Refs ADR-0001.
Later multi-relay work is verified almost entirely against FakeRelay, so a fake that diverges from NostrRelayClient would let those tests pass while production fails. Add a contract test that drives one subscription-routing scenario through both implementations via the RelayConnection interface and asserts they agree; reverting the fake to broadcasting makes it fail. Fix the accompanying subscriber-eviction test, which asserted the wrong ordering: it closed the only subscriber, so registration identifiers never collided and reintroducing the collision left it green. It now keeps a subscriber alive across the release, and fails when the bug returns.
FakeRelay.reconnect() revived a closed connection in place, which NostrRelayClient cannot do: it has no reopen path, so a closed connection is terminal and callers must obtain a replacement from the factory. Tests for auto-resubscribe would have been written against behaviour production does not have. Replace it with reconnected(), returning a new connection with the same URI and scripted behaviour and no subscribers, so re-subscription stays the caller's responsibility to prove.
Nostr is a multi-relay protocol, but a single connection gives a single answer, so every caller wanting real delivery had to write fan-out, result aggregation and timeout handling themselves. RelayPool publishes concurrently on virtual threads and returns a PublishResult recording, per relay, acceptance, rejection with the relay's verbatim reason, a timeout, or unreachability. Partial failure is ordinary data; a publish that no relay accepted throws NoRelayAcceptedException carrying the same result, so an event that reached nobody cannot be mistaken for a published one. The timeout is a budget for the whole publish rather than for each relay in turn, since spending it per relay let N stalled relays cost N times the wait the caller asked for. OK frames are matched on event id, so another event's acceptance is never credited to this publish, and a relay that cannot be reached is reported as unreachable rather than as a timeout. Refs ADR-0002, ADR-0004.
A connection serves one request at a time, so two threads publishing through the pool collided with NostrRelayClient's in-flight limit. Each relay now has its own lock: concurrent publishes to one relay queue, while publishes to different relays still overlap, so a slow relay delays only its own queue. Relay health becomes the pool's responsibility, as docs/CONTEXT.md assigns it. Each relay's connection state is observable, and downed relays are retried on a schedule the pool owns rather than one the caller must remember to drive. Relays that drop after connecting are returned to the downed set as well, since retrying only startup failures would let a long-running pool quietly shrink. Fixes a ConcurrentModificationException found by racing reconnection against publishing: membership was held in plain LinkedHashMaps that a publish iterated while a recovering relay rejoined from another thread. Membership is now held concurrently, with configured order kept separately so results stay predictable. Refs ADR-0004.
A subscription spread over five relays receives every event five times, ends its backlog five times, and loses a relay silently when one drops. RelayPool.subscribe now presents that as a single stream. Events are de-duplicated by id through a bounded window and delivered already parsed, since de-duplication has to parse them anyway. Bounding matters on exactly the long-lived firehose subscriptions that most need de-duplication, and an evicted id being seen as new again is the accepted cost. Per-relay EOSE frames are aggregated into one end-of-backlog signal, emitted once every relay has reported or a timeout expires, so a relay that never replays cannot leave an application showing a loading state forever. Subscriptions retain their filter so a relay that drops is re-subscribed when it recovers, and report a malformed payload rather than letting one relay's nonsense end a stream the others are serving. Also widens the publish-budget assertion's margin: it measured wall clock tightly enough to fail under load, and now discriminates on five stalled relays instead of three, far from the per-relay cost. Refs ADR-0002, ADR-0004, ADR-0005.
The SDK's modules are each deliberately narrow, so an application had to assemble connection handling, fan-out, de-duplication and NIP-17 delivery itself, every time. The new module owns that work. RelayPool gains runtime membership: relays join and leave a running pool, and one borrowed to reach a recipient is released afterwards by counting holders, so two overlapping deliveries do not cut each other off and a long-running process does not accumulate connections. RelayListLookup gives DirectMessageRelayLookup its first real implementation, resolving kind-10050 lists through the pool, which is what nostr-java-identity declared but could not satisfy without acquiring a transport. DirectMessagePublisher performs the delivery plan identity can only produce, sending each gift wrap to its own recipient's relays and reporting per recipient, since a group message reaching three of four people has partly succeeded. NostrClient ties them together behind an identity and a relay set. Ownership follows construction: a pool it built is closed with it, one handed in is left for its owner, so a container can share a pool across clients. Refs ADR-0001, ADR-0003, ADR-0005.
The transport dispatches each inbound payload on its own thread, so an EOSE could overtake the stored events it follows. A subscription then told its caller the backlog was drained while those events were still arriving, which is exactly the guarantee the signal exists to give. Subscriptions now handle one payload at a time. Found by testing against a live relay rather than a fake: a relay list lookup intermittently reported nothing for a list the relay was plainly serving, because the lookup stopped waiting the moment EOSE arrived. Adds the round-trip integration test that caught it, covering both publishing and a full NIP-17 delivery, plus a wait strategy that holds the container until it has actually stored an event. Waiting for the port is not enough: it binds before database migration finishes, and on some hardware a relay worker panics during startup leaving the relay accepting connections but answering nothing. Both look like client bugs. Documents the module in a how-to guide, the README, the architecture overview and the changelog.
…ounts Ticket 10 claimed the API reference covered NostrClient, RelayPool and PublishResult, and that the known limitations were written down. Neither was true: the reference documented none of the new types, and the limitations existed only in Javadoc and ADRs where a user would not find them. Add reference entries for the client and pool APIs, and an architecture section for nostr-java-api. Record the three behaviours that will otherwise look like bugs: throughput to one relay is latency-bound, a relay borrowed for a direct message is briefly shared, and de-duplication is windowed so a very late duplicate can reach the caller twice. Update the module count and dependency chain in the codebase overview, architecture and dependency-alignment docs, which still described four modules. The v2.0.0 highlight keeps its original numbers, since it describes that release rather than the present. Every documented snippet was compiled against the real API, and every documented signature, default and enum constant checked against source.
The test asserts stored events precede the end-of-backlog signal, but removing the ordering lock leaves it green: FakeRelay delivers frames synchronously, so the race the lock prevents cannot occur in it. Claiming otherwise would overstate what the suite protects. Note where the guarantee is really covered, which is the live-relay integration test that found the defect to begin with. Making the fake dispatch concurrently was tried and reverted: it broke the fixture's own tests without ever forcing the race.
Adds nostr-java-api, the client-facing entry point, together with the multi-relay pool it is built on: fan-out publishing with per-relay outcomes, de-duplicated fan-in subscriptions, and NIP-17 direct message delivery to each recipient's own relays. Released as 2.2.0 rather than folded into 2.1.0. That version was never tagged, but its changelog section is written and dated and its artifacts exist locally, so absorbing a new module into it would change what 2.1.0 means to anyone who already has it. Also moves the frame-ordering fix into this release's notes; it was recorded under 2.1.0 but landed in this cycle.
The spec targeted 2.0.x and planned to build relay pooling, subscription merging and NIP-17 itself. All three now ship in the SDK, so the module consumes them instead of duplicating them. Its one declared hard prerequisite is satisfied: NIP-17 landed in 2.1.0 and delivery to a recipient's own relays in 2.2.0, so phase 0 is gone and the DM phase becomes adapter work over NostrClient. Resolves two name collisions that would have put two different types of the same name in one codebase. The spec's RelayPool becomes RelayDirectory, which maps logical relay names and owns no connections, and its DirectMessageService becomes McpDirectMessageService. Records what the module no longer builds, and the three SDK limitations that shape its design: single-request-in-flight per relay, transient relay sharing during DM delivery, and windowed de-duplication. Every SDK symbol the spec names was checked against source.
A specification is not executable, so its assumptions about this SDK rot silently until somebody implements against them. Prototyping the spec's own call sequences against a live relay found two claims in it that were wrong. subscribe() is asynchronous: it returns before any stored event or the end-of-backlog signal arrives. The spec said a tool could return after EOSE, which would have made every subscription stall for the backlog timeout on an unresponsive relay. A direct message to one recipient reports two outcomes, and a sender who has published no relay list sees their own archival copy come back UNREACHABLE while the recipient is DELIVERED. Reported naively that reads as 'one of two delivered', telling a user their message failed when it arrived. McpSpecAssumptionsIT now asserts both, plus per-relay publish outcomes, total failure carrying its result, mid-subscription delivery, relay-list lookup, and signing as a named identity. Verified to fail when the assumption breaks: making subscribe block turns it red.
…py claim The isolation argument still reasoned about NostrRelayClient's futures, from before RelayPool existed. Its conclusion holds but its evidence was stale: capping an identity at one in-flight operation is now what the pool does deliberately, per relay, so citing it as a drawback of threads was misleading. McpSpecAssumptionsIT asserted that a one-recipient message yields two outcomes, but not the behaviour §6.2 actually warns about: that a sender who published no relay list sees their own archival copy come back UNREACHABLE while the recipient is DELIVERED. That is now asserted, so the warning rests on an observation rather than on reasoning. Also re-checked every component name the spec defines against the SDK; all nine are free.
Records the decisions and the reasoning behind each, since the reasoning is what a reader needs when a trade-off resurfaces. os-keychain becomes the default keystore: it needs no passphrase so it does not block unattended startup, with encrypted-file kept as the portable fallback and selected explicitly in the compose file, where there is no keychain to talk to. Identity lifecycle tools stay. The design already removes the reason to withhold them: no tool accepts or returns key material, creation is reversible, and the irreversible operations are two-step guarded. Bound processes share relay connections through a broker, so N identities do not open N websockets to one relay. The broker shares transport only; sharing a RelayPool would share subscriptions and per-relay state, undoing the isolation the processes exist for. Identity mutation gains its own identity-policy switch, because destroying a key and publishing an event have different risk profiles and collapsing them forces a choice between a useful agent and a protected keystore. Confirmation tokens do not expire, no identity is auto-created on first run, and HTTP transport auth is deferred with the localhost-only constraint stated plainly rather than implied. Verified the ContactList question: no such type exists, only the constant Kinds.CONTACT_LIST. A value type over kind-3 belongs in nostr-java-event following DirectMessageRelayList, making it this module's one remaining SDK prerequisite.
…honestly Resolving the open questions asserted a shared relay-connection broker without defining it, which hid the hard part: bound processes are separate address spaces, so a shared broker is another process rather than an object. It gets a name, a place in the key components, and its real costs stated: its own lifecycle, a single point of failure for every identity's relay access, and visibility of all their traffic. It is opt-in rather than default, since a handful of identities should just open a handful of connections. The non-goals still claimed no NIP work remained, which the ContactList answer had just made false. Both that section and the decisions table now name kind-3 as the module's one SDK prerequisite, blocking the social phase only. Audited all nine decisions against the whole document: each is stated and none is contradicted.
Auditing a specification with greps proves it is self-consistent, not that its instructions work. Both new claims were testable, so they were built instead of asserted. The connection broker: two independent RelayPool instances, fed by one websocket through a RelayConnection, both publish successfully to a real relay, and the second keeps working after the first closes its view. That last part is the claim that matters, since a broker owning the transport is exactly what distinguishes it from a pool owning it. The kind-3 contact list: a value type written to the DirectMessageRelayList pattern round-trips through a real relay, published and recovered with its p tags intact. This is why §12 can call the prerequisite small: one value type, no new infrastructure. Both guards were mutation-verified. The broker test initially passed even when the shared connection was closed on release, because both publishes had already completed; it now publishes after a release, and fails when the connection is torn down.
Deferring HTTP transport authentication is only safe if the constraint that replaces it is visible. It was stated once, in the decision itself, while §7.2 tells a deployer to run HTTP inside a container, which is exactly where it goes wrong: publishing a port reaches every interface by default. The constraint now appears where each reader meets it. bind-address defaults to 127.0.0.1 in the configuration, the safety model lists an exposed transport as a threat and warns at startup on a non-loopback bind, and the compose file publishes to the host loopback only. A deferral that lives in one section of a design document is not a deferral, it is a trap.
Each ticket cuts a complete path rather than a layer, so every one is demoable on its own: the skeleton ships a working tool over stdio, the read path ships three tools with the argument conventions later tools inherit, and so on. Two depart from the spec's six phases. Phase 1 is split into skeleton, keystore and CLI-plus-binding, since it was far too large for one sitting. Phase 3 is split by risk, keeping publishing separate from keystore mutation, which is also why the lifecycle tools wait on WriteGuard: they reuse its two-step confirmation rather than inventing a second one. The ContactList prerequisite leads, since it is the only SDK work the module needs and it unblocks the social phase. Graph verified acyclic, with no missing blockers and no forward references.
Kinds.CONTACT_LIST existed but nothing modelled the event, so reading a follow list meant parsing p tags by hand. This is the only SDK work the planned MCP module needs. Each entry keeps all three parts NIP-02 defines rather than the key alone. The relay hint is how a client finds someone it has never seen, and the petname is how it shows a readable name without a global registry; an implementation that reads only the key discards both on every round trip, which an earlier prototype of this type did. Both are reported absent rather than blank, since a list routinely carries [p, key, , ], and a petname is emitted in its third position even when there is no hint, because NIP-02 reads parameters positionally and a shifted petname would be read as a relay. Entries keep their order, since NIP-02 asks that new follows be appended so the list reads chronologically. A duplicated key keeps its first entry, and an entry with no key is discarded rather than making a whole list unreadable. Guards are mutation-verified: dropping the hint and petname, removing the kind check, accepting any tag as a contact, and keeping duplicates each turn a test red. Closes ticket 01.
added 23 commits
August 30, 2026 16:42
First slice of nostr-java-mcp: an MCP host launches the server, sees its tool list, and calls nostr_list_relays. That one tool needs no keystore and writes nothing, so it proves transport, registration, configuration and the SDK underneath without risking anything. The module adapts nostr-java-api rather than reaching past it. Going straight to NostrRelayClient would rebuild connection management, result aggregation and de-duplication the SDK already owns, and get them wrong in the ways it already learned about. Tools are registered one class each, because policy and single-identity mode will both work by not registering a tool rather than rejecting a call, and that only holds if there is one list to review. A duplicate name is rejected rather than silently resolved, since which of two tools an agent got would otherwise depend on iteration order. Pins jackson-annotations to 2.21. Jackson 3 keeps its annotations under the 2.x coordinate and needs the JsonSerializeAs introduced there; nearest-wins resolution was handing the SDK's mapper an older jar and the server died on its first JSON frame. Only the end-to-end test caught it: every unit test passed while no host could talk to it. Closes ticket 02.
…keys An MCP server signs on an agent's behalf, so the private key is the whole security boundary: the agent must be able to say "sign as alice" without ever being able to read alice's key. IdentityVault holds that line. It loads key material once at startup, hands out only IdentitySummary (alias, public key, source), and wipes the material on close. Where a key comes from is a KeySource, so operators can choose their own trust model: the OS keychain by default, an encrypted PKCS#12 file, or environment variables for containers. KeySources maps configuration to a backend in one place, so adding a backend is a new case rather than an edit spread across startup. Choosing an identity is explicit. A single configured identity is the default, since naming the only option is ceremony, but several with no stated default stay ambiguous and the tool asks rather than guessing which key signs. A configured default naming no identity fails at startup, where it is a typo, not at first use, where it is a mystery. ToolSurfaceSecrecyTest sweeps every registered tool's schema, output and errors for key material, so a future tool that leaks one fails here. It was checked against a deliberately leaking tool, and the vault's wipe, ambiguity and startup-validation guards were each checked by breaking them and watching the corresponding test fail.
… CLI The multi-identity server guards against posting as the wrong account; it cannot remove the risk, because once a default exists an agent may still name any alias it holds. Binding removes it. Setting nostr.mcp.identity ties a process to one alias, and the restriction is applied before decryption rather than after: a bound process reads only its own entry, so another identity's key is absent from the heap instead of merely out of policy. With one identity there is nothing to name, so there is no wrong name to give. Binding presupposes the key exists, so the same jar administers one. KeyAdminCli offers keygen, import, list and remove, deliberately as a command rather than a tool, which keeps key administration out of every agent's reach. Keys are read from stdin, never arguments, since argv is visible to every user on the host, and a generated key is stored and wiped without being printed. IdentityStore is the seam the future lifecycle tools will share, so "remove an identity" has one meaning and two front doors. A bound server whose alias is missing refuses to start and says which command creates it. An unbound one with an empty keystore still starts, because that is how a person makes their first key. Two bugs the integration test caught that the unit tests could not: PKCS#12 rejects the RAW key algorithm, so keygen failed against a real keystore while passing against an in-memory one, and the macOS keychain branch passed the secret in argv. The encrypted-file backend now has tests of its own against a real file. Golden files pin the bound and unbound tool lists separately; registering an admin tool was shown to break the unbound file while leaving the bound one green.
An agent can now answer questions about Nostr without being able to change anything: nostr_query_events runs a bounded query, nostr_get_profile resolves a public key or a NIP-05 address and decodes the kind-0 JSON, and nostr_relay_info reads a relay's NIP-11 document so an agent can learn a relay's rules without breaking one. This establishes the argument conventions later tools inherit. NostrIdentifier accepts hex or bech32 and refuses an nsec in a public argument with advice to rotate, since by the time a tool sees one the key is already exposed. TimeArgument normalises "24h", ISO-8601 and bare dates to Unix seconds. ToolArguments coerces the shapes a model actually sends, because a JSON schema is a suggestion to a language model. Queries are bounded twice, by event count and by deadline, and reaching either bound is reported: an agent that cannot tell "nothing matched" from "I stopped looking" will report the first when the truth was the second. The integration test against a real relay then failed in a way no fake could have shown. Chasing it found an ordering bug in the transport, not in the new code: every inbound frame was dispatched on a freshly started virtual thread, so EOSE could overtake the stored events it follows. A probe showed the end-of-backlog signal arriving with one or two of three events delivered, varying run to run. Every caller ends its query on that signal, so every such query could silently return a partial answer. Frames are now queued per listener and drained in order, which restores ordering without serialising unrelated listeners or blocking the receiving thread. The regression test was confirmed to fail against the old dispatch. RelayStoresEventsWaitStrategy moves into the client test-jar so both modules share one copy rather than a fork.
An event a relay has accepted cannot be reliably withdrawn, since NIP-09 deletion is advisory. Publishing is therefore the one capability where a model's ordinary failure mode, confidently doing something nobody asked for, produces a permanent public result. write-policy: confirm is the default: the first call signs the event, holds it, and returns a preview with a token; only a second call carrying that token publishes. An agent that invented the intention does not follow through, so the post never happens. deny registers no write tool at all, because a tool an agent cannot see cannot be talked into running, whereas a refusal it can see is an invitation to rephrase. The event is signed before it is previewed and the held event is the one sent, so the content cannot drift between what the user approved and what goes out. Tokens are random, single-use, and survive until used or until the process ends: expiring them on a timer would fail an agent that paused to ask its user, which is the behaviour we want to encourage. Every write goes through WriteGuard rather than reaching the pool, so policy, rate limiting, identity resolution and the audit log are properties of the server rather than habits each tool must remember. A new write tool inherits them by construction. Results follow the SDK: any acceptance is success carrying the per-relay list, and only reaching no relay at all is an error, since reporting a partial success as failure would push an agent to republish something already on the network. Verified against a real relay: an unconfirmed note leaves the relay empty, a confirmed one is stored and readable back, and a published profile reads as a profile. Both guards were mutation-checked. Removing the confirmation requirement failed two integration tests, and reusing a spent token failed its unit test.
An agent that can only use pre-configured keys is half a tool: "make me a throwaway account for this project" is a natural request. But an identity is the user's Nostr account, so the tools are shaped by the asymmetry between operations rather than treated alike. Creating, renaming and setting a default are reversible and freely allowed. Exporting a backup and removing a key are not, and are guarded. Import is the case worth dwelling on. nostr_import_identity accepts no key material whatsoever, because an nsec in an argument would be written to the conversation log and very likely sent to a third-party inference API, which is the worst thing this module could do. Instead source names a location the server reads for itself, a file it can then shred, its own environment, or a terminal the agent cannot see. The model arranges an import it never observes. Pasting a key anyway is refused with advice to treat it as compromised, since by then it is already exposed. Removal is the only irreversible tool here: an npub with no nsec is a dead account and every event ever signed with it is orphaned. It confirms in two steps and refuses outright when no backup exists, unless the caller states the key is disposable. An agent told to "clean up the test accounts" has no basis for that judgement, so the tool makes it say so where a user can contradict it. Backup state is deliberately not persisted: after a restart this server cannot know a backup file still exists, and a stale reassurance is worse than asking again. identity-policy is separate from write-policy and capped by it, so a deployment can let an agent post without letting it destroy keys, and a read-only server can do neither. A backend that cannot be written to registers no lifecycle tools at all, since a tool that could only fail is worse than one that is absent. The secrecy sweep now walks the real tool surface rather than a curated list, so it covers tools added later; it was confirmed to catch an import tool that accepts an nsec. Both removal guards were mutation-checked.
"Watch my mentions" cannot be answered by a call that returns, and MCP
has no way to push into a tool result, so a subscription becomes a
stateful server resource: the tool opens it and returns an id, events
land in a buffer, and the agent drains them. Each is also exposed as
nostr://subscription/{id} with update notifications, so a host that
supports resources gets push and one that does not still polls.
The buffer is bounded, and what it drops it counts. That count is the
part that matters: an agent silently handed a gap will summarise a
partial feed as though it were complete, whereas one told it missed
thirty events can say so or narrow its filter. The count is monotonic
across drains for the same reason, since a gap that happened two polls
ago is still a gap in what the agent believes.
Reading drains, which is a command that also answers. Accepted
deliberately: the alternative refills the model's context with the same
events on every poll until nothing else fits.
Subscribing does not wait for the backlog. RelayPool.subscribe returns
before any stored event arrives, so blocking would stall on any relay
that never answers, and the result instead reports backlogDrained so an
agent can tell "nothing matched yet" from "still replaying". Relays that
drop out are recorded rather than silently narrowing coverage.
Idle subscriptions are reaped and the total is capped, because an agent's
session ends whenever its user closes a window without telling this
server. Idleness is measured from the last read, not the last event, so a
busy filter cannot keep a forgotten subscription alive forever.
Verified against a real relay: a note published after subscribing arrives,
a second read returns nothing, an overflow reports its drops, and an
empty read distinguishes a drained backlog from a replaying one. Reaping
and draining were both mutation-checked. The filter schema and decoding
now live in one place shared with the query tool, so an agent that has
learnt to filter for one can use the other unchanged.
An agent can now read a thread, read who someone follows, and send a message only its recipient can open. Replies carry the NIP-10 tags that make them replies, since without them a client shows an answer detached from the conversation it answers. Direct messages are NIP-17 only. NIP-04 exposes both correspondents and the conversation to every relay, which is unacceptable for a tool acting on someone's behalf, so it is not offered at all rather than offered with a warning. An integration test confirms no gift wrap is signed by the real sender, which is the property that hides who is talking to whom. Two delivery facts are reported rather than smoothed over, both of which were established against a live relay rather than reasoned about. A recipient who published no kind-10050 relay list genuinely cannot be sent to, so they are named as unreachable instead of silently skipped. And every conversation includes a copy addressed to the sender, so a message to one person yields two outcomes; a sender with no relay list of their own sees that copy fail while the recipient succeeds. Reported naively as "1 of 2 delivered" that would tell someone their message failed when it arrived perfectly well, so the archival copy is reported separately and its failure is phrased as advice about their other devices. Reading messages is opt-in per identity. Decryption puts private correspondence into the model's context and therefore the host's logs, so it stays a decision for the person whose messages they are. The vault boundary survives this. NIP-17 sealing derives shared secrets and so needs an Identity rather than a signature, so rather than release a key the vault builds the collaborator internally and returns only what that collaborator exposes, which is composing and reading messages and never the identity behind them. Both reporting rules were mutation-checked against the live-relay tests.
…henticated Adds a streamable HTTP transport for hosted deployments where the MCP host does not launch the process itself, selected by configuration alone so moving between transports needs no code change. It runs an embedded servlet container rather than requiring one, since a server an operator must deploy into a container is a server most people will not run. The transport carries no authentication in this version. That deferral is defensible only while the constraint replacing it stays visible, so the default binding is loopback, a non-loopback bind warns at startup naming what is at risk, and an unresolvable address is treated as unsafe rather than given the generous reading. The same principle as the unprotected keystore backend: the weaker choice should announce itself, because the person who made it may not be the person reading the logs. Per-session identity binding is deferred rather than delivered. The streamable transport builds one server shared across sessions, so filtering the surface per session needs a session-scoped tool surface the SDK does not currently offer, and a half-enforced isolation boundary is worse than an absent one. Process-level binding from ticket 04 remains the supported answer and is the stronger guarantee anyway, since it is an address space rather than a check. The new how-to guide states the authentication constraint where a deployer will read it, and its settings table is checked against the code: a test extracts every key McpConfiguration reads and fails when one is undocumented. Confirmed by adding a setting and watching it fail.
The module is launched as `java -jar`, both by an MCP host and by its own key-admin CLI, but nothing built such a jar: the documentation promised a command that did not work. It now ships one self-contained artefact, since a host configuration naming a classpath would break whenever a dependency changed. The container is where the HTTP transport's missing authentication turns dangerous, because publishing a port reaches every interface by default and would silently override the server's own loopback binding. Every mapping in the compose file therefore names 127.0.0.1 explicitly, with a comment saying why removing it is unsafe. The image is distroless and runs as a non-root user: this process holds private keys, so the smallest surface is worth losing a shell for. A profile demonstrates one bound container per identity, which is the strongest isolation the module offers because it is an address space rather than a check. Packaging exposed a real bug. Environment variable names were derived by uppercasing a setting and replacing dots, but not hyphens, so `write-policy` became `NOSTR_MCP_WRITE-POLICY` and every hyphenated setting was unreachable from the environment. That includes the write policy, the bind address and all the limits, which is to say the container could not be configured at all. Hyphens are now translated too, and a test asserts every setting maps to a name a shell can actually set. Verified by running it: the shaded jar completes an MCP handshake and runs the CLI, the image builds, and a container answers a real initialize request over HTTP as `nonroot`.
A tool surface with no guidance makes a model explore by trial and error, which on a public and irreversible medium is the wrong way to learn: the mistakes are permanent and other people see them. Three prompts encode the sequences that work, and each targets a specific mistake rather than restating the tool list. compose-note names the confirmation step models skip and warns against republishing a note that already reached some relays. catch-up-feed reads the follow list before querying, because a model left to itself queries broadly and then filters, which returns strangers and misses the people actually followed. watch-mentions distinguishes a replaying backlog from an empty one, which is how an agent comes to report "nobody mentioned you" when it simply read too early. Identities and relays are also exposed as resources, so a host can put the server's own configuration into context rather than spending a tool call on something that cannot change mid-conversation. The how-to now leads with single-identity mode, because it removes the wrong-account risk instead of guarding it and matches how host configs are written anyway; the unbound server is presented as the advanced case that answers cross-identity questions. It also documents the limits inherited from the SDK where someone will meet them: per-relay throughput, windowed de-duplication, subscriptions not surviving a restart, and the unauthenticated HTTP transport. Both documents are kept honest by tests. Every registered prompt must appear in the guide, and each prompt's warning is asserted by the behaviour it prevents; both were confirmed by breaking them. Prompts and resources are also verified over the real protocol through a subprocess MCP client rather than only in the registry.
Binding a server to one identity is meant to remove the wrong-account mistake rather than guard it: with one identity there is nothing to name, so nothing can be named wrongly. But all five signing tools still advertised an `identity` argument when bound, which left it merely redundant instead of absent. An argument with exactly one acceptable value is an invitation to pass a different one, so the guarantee the ticket describes was not actually being delivered. Found by running every requirement against the shipped jar through a real MCP client, rather than against the classes through their own tests. The existing suite passed because it asserted which tools a bound server registers, never what those tools ask for. The end-to-end check has been kept in .scratch alongside a unit test that fails on the old behaviour.
The module's tests exercise classes through their own seams, so they can all pass while the artefact a user runs behaves differently. This drives the built jar as an MCP host does, over stdio against a real relay, and asserts one check per ticket checklist item. It earned its place immediately by finding a defect the existing suite could not: a bound server still advertised an `identity` argument, so the mistake binding exists to make impossible was merely redundant. The surface tests asserted which tools were registered, never what those tools asked for. Two things about the harness itself are worth stating. It starts from an empty keystore each run, so a repeated run measures the product rather than the residue of the last one. And it probes relay readiness by publishing and requiring acceptance, not by connecting, because the nostr-rs-relay startup panic leaves the port open while the relay silently answers nothing; a connection check passes against exactly the relay that will fail every test. Verified stable across three consecutive runs.
Minor rather than patch or major, derived from what actually changed rather than from the commit subjects alone: twelve feat commits add the nostr-java-mcp module and the ContactList/Contact types, one fix corrects relay frame ordering, and a diff of the already-released modules shows no public member removed or altered. Nothing breaks for an existing consumer. The MCP server's reported version was a constant in the source, which is a copy of the pom that nothing keeps in step: this very bump would have left it announcing 2.2.0 to every host, and no test would have noticed. It is now filtered in from the pom at build time, with a test asserting the two agree, and the running jar was checked to report 2.3.0 in its initialize response. Verified at the new version on both profiles, 144 unit and 46 integration tests. One NostrClientRoundTripIT error during the first run was the known nostr-rs-relay startup panic, which its own retry then passed and a clean rerun did not reproduce.
Every other test here asks whether the tools work. This asks the question the module exists to answer: whether a model can use them. A tool can be correct and still unusable, because its name misleads, its description omits what the model needs to decide, or its schema invites an argument the model cannot supply. None of that is visible to a test that calls the tool directly, and all of it decides whether the server is any good in an agent's hands. It runs a real model through Testcontainers and checks that it picks the right tool unprompted, distinguishes querying from subscribing, reaches for publish rather than a neighbour when asked to post, and reads a confirmation preview as "not yet published". That last one matters most: the whole write guard depends on a model understanding that a preview is not a success, and until now nothing verified it. The host's model cache is mounted rather than pulled, because a test that downloads several gigabytes is a test people disable, and it skips cleanly when no cache is present. It is tagged and excluded from the ordinary build since it takes minutes; the exclusion had to go inside the failsafe execution rather than beside it, because the parent declares its configuration there and a plugin-level block is silently ignored. The class documents what a pass does and does not prove. Names and descriptions reinforce each other, so gutting one description alone still passes; renaming both tools to nostr_alpha and nostr_beta shows the descriptions carry real weight on their own. Read a pass as "the surface is comprehensible", not "every word is load-bearing".
Asked whether all the tools were tested, the honest answer was no. A coverage measurement put it precisely: eight of the twenty-two had only ever had their registration checked. Appearing in the golden tool list guarantees that a class compiles, not that it works. Calling them found nostr_relay_info broken against every real relay. Java's HTTP client offers an HTTP/2 upgrade, and a relay serves its NIP-11 document from the same host and port as its websocket endpoint, so it read those headers as a botched websocket handshake and answered 400. curl worked, which is why nothing noticed. The client is now pinned to HTTP/1.1. Fixing that surfaced a second one: a relay name that was neither configured nor a URI reached the HTTP client and failed with an IllegalArgumentException about an undefined scheme rather than a code the agent could act on. A third failure was mine, not the product's. The test assumed two identities are ambiguous, but the first one created becomes the default, which is deliberate and better. The test now asserts what the tool is actually for: moving that choice afterwards, checked by which key signs the next note. ToolCoverageTest keeps the gap from reopening by failing when any registered tool is never called. It is a floor, not a quality claim: it says each tool has been invoked, not that the interesting cases are covered.
Asked whether the tools were tested with the LLM and with relays, the measured answer was 19 of 22 against a relay and 4 of 22 against a model. Both numbers are now 22, and a coverage test fails when either slips, because the honest answer to that question should come from a measurement rather than an impression. The relay gap closed with tests for the three tools that had only been called against fakes, and each asserts something a stand-in cannot show: a follow list keeps its relay hint and petname through a real round trip, a removed identity can no longer sign while the note it sent before removal is still on the relay, and the relay list reports a genuinely connected relay as connected. The model gap closed with a case per tool, phrased as a user would ask, with all twenty-two offered at once. Offering the whole surface is the point: the risk is that two tools read alike to a model, and that only appears when both are on the table. Two tools are deliberately exempt, with the reason recorded in the test: the escape hatch overlaps every other publishing tool by design, and inviting a model to choose the one irreversible tool is a bad habit to build into a suite. One case failed and taught me something. Asked to watch for notes "mentioning me", the model declined to call anything and asked which identity was meant, which is correct behaviour on an ambiguous request. The prompt was at fault, not the surface, so it now asks for something answerable. Re-run afterwards: 25 of 25, no flakes.
scripts/doccheck.py verifies that every relative link resolves, every backtick-quoted type name matches a real Java or TypeScript type, and every environment variable bound in application config appears in the configuration docs. It exits non-zero so it can gate a commit. Where present, .doccheck-allow declares names that are deliberately not types in this repository: third-party and platform types, and vocabulary from design proposals that are not yet built.
Three fixes found while extending the checker across the rest of the stack: - Index .ts and .tsx declarations and exported React components. Several repos ship a TypeScript client whose types are legitimately named in docs; without this the checker reported them all as unknown. - Skip plans/, specs/, superpowers/, and archive/ directories. Those documents quote link lines destined for other files, whose relative paths are correct only from the target, so checking them reports the quoting rather than a broken link. - Extend the built-in skip list with JDK, Spring, browser, and node types commonly named in prose.
The link checker special-cased http, https, and mailto. A nostr: URI in a README was therefore resolved as a relative path and reported broken. Any scheme before the first slash now counts as external. Verified against a fixture: scheme URIs skip, a genuinely broken relative link is still reported and still exits non-zero.
Rewrote the entry points and reorganised the index around Diataxis, then did the part that actually matters: checked the documentation against the code instead of reading it. That found things reading would not. The API reference taught BaseMessage.read(json), a method that has never existed anywhere in this project. Five install snippets used <version><!-- X.Y.Z --></version>, which Maven reads as an empty version, so nobody could copy them. The getting-started guide told readers to depend on nostr-java-client while the README told them nostr-java-api. MIGRATION.md linked to a nostr-java-examples module that does not exist. Nine documents, including all three operations guides, were unreachable from the index. A worse one turned up in passing: -DnoDocker=true, documented in five places and passed by scripts/release.sh --no-docker, sets a property nothing reads. Running verify with it proved the container tests ran anyway, so anyone following those instructions on a machine without Docker got a failing build and no clue why. Both now use -Pno-docker. Two fresh agents read the docs with no other context and answered a newcomer's questions from them alone, which is where the uncopyable versions and the module contradiction surfaced: I could not see them because I already knew the answers. Compiling the examples caught one of mine, a checked NoRelayAcceptedException the quick start silently ignored. DocumentationAccuracyTest keeps all of this from returning. It fails when a guide names a type or method that does not exist, a link resolves to nothing, an install snippet quotes the wrong version, or a page is orphaned from the index. Each check was confirmed to fail when the defect it targets is reintroduced.
A superseded checkout keeps its documentation as written; repairing its links would imply the tree is maintained. A repository can now declare 'doccheck: skip-links' in .doccheck-allow to record that decision. Opt-in only, verified against a fixture: without the marker a broken link is still reported and still exits non-zero.
CONTRIBUTING.md documented an architecture this project abandoned in 2.0. It told contributors to add a per-NIP facade extending EventNostr, to update a NIP compliance matrix, and offered NIP01EventBuilder and UserProfile as naming examples. None of those exist; a search for the classes returns nothing. A guide that confidently teaches a design the codebase has removed is worse than no guide, because a contributor follows it and produces work that cannot be merged. Its accurate half, the commit and pull request conventions, moved into the codebase overview rather than being deleted with it. That page also claimed pull requests target `develop` while CONTRIBUTING.md said `main`; the remote's default branch is `main`, so the contradiction is resolved in favour of what the repository actually does. Patch rather than minor, derived from the diff rather than the commit subjects: since 2.3.0 the only source changes are two nostr_relay_info bug fixes and a release-script correction. No API was added or removed. The bump also demonstrated the documentation test earning its place. It failed the build because multi-relay-publishing.md still named 2.3.0 in a snippet a reader would copy. Rather than fix that by hand a second time, scripts/release.sh now updates copyable install snippets as part of the bump, leaving placeholders, ranges and a migration guide's "previous version" alone.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
nostr-java-mcp, a Model Context Protocol server that exposes the SDK to LLM agents,and brings the release line from 2.0.8 to 2.3.1. Along the way it fixes two bugs in the existing
transport that were reachable from ordinary use, and puts the documentation under test.
This is large (226 files) because it spans four releases. See Where to start reviewing below.
Type of change
feat- New feature (non-breaking)fix- Bug fix (non-breaking)docs- Documentation onlytest- Adding or updating testschore- Build, CI, or tooling changesWhat changed?
The MCP server (
nostr-java-mcp). 22 tools covering reading, publishing, identitylifecycle, subscriptions and NIP-17 messaging, over stdio or HTTP, plus guided prompts and a
container image. Its safety model is the interesting part, because publishing to Nostr cannot be
undone:
published until it calls again with that token. A hallucinated post becomes a no-op.
nostr_import_identityaccepts no keymaterial at all: the agent names a location the server reads for itself, so a key never
enters the model's context or the host's logs.
from the process rather than merely out of policy.
Two bugs in existing code, both found by testing against a real relay rather than a fake:
virtual thread, so an
EOSEcould overtake the stored events it follows. Since every callerends a query on that signal, any query could silently return a partial answer
indistinguishable from the relay holding less data. Frames are now drained in order per
listener.
nostr_relay_infocould not read any relay's NIP-11 document. Java's HTTP client offersan HTTP/2 upgrade, which a relay reads as a botched websocket handshake and answers
400.curlworked, which is why nothing noticed. Now pinned to HTTP/1.1.-DnoDocker=truenever worked. Documented in five places and passed byscripts/release.sh --no-docker, it set a property nothing reads, so the container-backed testsran anyway. Anyone following those instructions without Docker got a failing build and no
explanation. Both now use
-Pno-docker, verified by runningverifywith each.Documentation. Rewritten and reorganised around Diátaxis, and now tested:
DocumentationAccuracyTestfails when a guide names a type or method that does not exist, a linkresolves to nothing, an install snippet quotes the wrong version, or a page is orphaned from the
index. It immediately caught the reference teaching
BaseMessage.read(json), a method that hasnever existed.
CONTRIBUTING.mdis removed: it taught the pre-2.0 architecture, including aEventNostrbase class that no longer exists.Breaking changes
None. No public member was removed or altered in the already-released modules; the diff against
mainfor those modules is additions plus the frame-ordering fix.Testing
mvn testmvn verify(requires Docker)Beyond the suite, three things were checked that unit tests cannot:
MCP host does, over stdio against a real relay, asserting one check per ticket item: 33/33,
stable across three runs. It found a defect the class-level tests missed, where a bound server
still advertised an
identityargument, defeating the guarantee binding exists to provide.OllamaAgentIToffers all 22 tools to a local LLMand checks it reaches the right one from a plain-language request, and that it reads a publish
preview as "not yet published". 25/25. This asks whether a model can use the tools, which a
correct-but-unusable tool would fail.
test fails: the confirmation token, removal-without-backup, key secrecy, subscription reaping,
buffer draining, and the DM outcome reporting.
Where to start reviewing
docs/explanation/nostr-java-mcp-spec.md— the design and its reasoning.nostr-java-mcp/src/main/java/nostr/mcp/identity/IdentityVault.java— the security boundary.nostr-java-mcp/src/main/java/nostr/mcp/write/WriteGuard.java— the single point every writepasses through.
nostr-java-client/.../NostrRelayClient.java— the frame-ordering fix, the one changeaffecting existing users.
Review focus
fail against the old dispatch, but a second opinion on the per-listener queue is worth having.
recorded in the ticket: the SDK builds one server across sessions, and a half-enforced
isolation boundary is worse than an absent one. Process-level binding ships and is stronger.
release/2.1.0no longer matches its contents, which now end at 2.3.1.Happy to rename or re-target if you would prefer a different base.