Skip to content

bmp-in: record why a monitored router's BGP session went down - #13

Draft
amtypaldos wants to merge 4 commits into
FastNetMon:mainfrom
super-trace:bmp-peer-down-reason
Draft

amtypaldos wants to merge 4 commits into
FastNetMon:mainfrom
super-trace:bmp-peer-down-reason

Conversation

@amtypaldos

Copy link
Copy Markdown

Summary

When one of a monitored router's own BGP sessions goes down, the router sends a
BMP Peer Down Notification (RFC 7854 §4.9) with a reason code and, depending on
the reason, the BGP NOTIFICATION it sent or received, or the FSM event that
closed the session. netom read only the per-peer header and discarded the rest.
This PR records it:

  • /api/v1/ingresses: each view of the peer (pre-/post-policy) gets a
    structured last_down with time, reason / reason_code,
    notification_code / notification_subcode, shutdown_communication
    (RFC 8203/9003), fsm_event and a one-line description.
  • /api/v1/bgp/neighbors: BMP-monitored peers get lastError (the
    description, in the same {:?}-of-Details wording that bgp-tcp-in already
    uses for native sessions) and a new lastDownTime.
  • netom-cli show ip bgp neighbors: shows Last error and
    Last down: <time> (<age> ago) for those peers.
  • /metrics: new bmp_state_num_peer_down_notifications{router,reason}
    counter, one per Peer Down message. Every reason is emitted so rate() works
    from the first scrape.

The record is kept after the peer comes back up (update_info only merges set
fields and a PeerUp never sets it), so you can still see why a session last
dropped. A synthesized view does not inherit it from the view it was copied
from.

Live, from a gobgp router whose transit shut the session:

netom> show ip bgp neighbors 10.99.0.13
BGP neighbor is 10.99.0.13, remote AS 65300
  BGP router identifier: 10.99.0.13
  BGP state = Idle
  Learned via: BMP feed
  Monitored router: 10.99.0.10 (ingress 5)
  RIB type: InPre
  Ingress id: 13
  ...
  Last error: remote NOTIFICATION: Cease(AdministrativeShutdown)
  Last down: 2026-10-01T21:37:32Z (00:00:12 ago)

Why

Without this, every BMP-monitored session that drops is just Idle: a
max-prefix teardown, an operator shutdown and a hold-timer expiry look the same,
and finding out which means logging into the router. The router already tells
the collector; netom only had to keep it.

Two notes on routecore (no changes needed there):

  • The reason code is read from the message because PeerDownReason has no
    numeric value and maps RFC 9069's code 6 to Unknown.
  • NOTIFICATION data is bounded by the NOTIFICATION's own length because
    NotificationMessage::data() runs to the end of the BMP message.

Out of scope, by nature of BMP: a session that never reaches Established (e.g.
an OPEN rejected for a bad peer AS) is never reported by the router, so it
doesn't appear. Documented in docs/bmp-tcp-in.md.

Validation

  • python3 pkg/check-build-sources.py: OK
  • python3 -B pkg/test-release-metadata.py: OK
  • cargo build --locked: OK
  • cargo test --locked --lib: 358 passed, 31 ignored, 0 failed (was 345
    passed on main)
  • cargo test --locked --bin netom-cli: 115 passed
  • scripts/e2e-addpath-bmp.sh (extended to send reason 3 with a Cease /
    Administrative Shutdown NOTIFICATION and assert /ingresses and
    /bgp/neighbors): OK
  • python -m sphinx -n -W --keep-going -b html docs docs/_build/html: OK
  • cargo clippy --lib --bin netom-cli: no new warnings
  • rustfmt --check: no new diffs in touched hunks (several touched files were
    already not fmt-clean on main; those hunks are left alone)
  • git diff --check: clean
  • Live: image built with pkg/Dockerfile + Dockerfile (smoke test OK),
    fed by a gobgp 4.9 router over BMP:
    • Its transit shut the session (gobgp neighbor disable on the far side).
      /bgp/neighbors reports lastError: remote NOTIFICATION: Cease(AdministrativeShutdown) and lastDownTime; /ingresses has
      last_down with reason 3 and code/subcode 6/2. netom-cli shows both
      lines, and the counter shows reason="remoteNotification".
    • The router hit a prefix limit. gobgp enforces it by setting the peer
      admin-down, and its BMP encoder reports that as reason 4 rather than
      reason 1 with Cease / Maximum Prefixes, so netom shows remote closed without NOTIFICATION. That's what the router sent. The reason 1 +
      Cease(MaximumPrefixesReached) path is covered by unit tests.
    • After that peer re-established, its row kept the last reason.

Unrelated, seen while testing on macOS: on main,
units::bmp_tcp_in::transport::tests::tls_handshake_has_a_deadline sometimes
hangs indefinitely in a full cargo test --lib run (paused tokio clock with
real TCP). It reproduced on unmodified main; it passes when run alone.

Not in this PR

  • bmp-tcp-out still sends a hard-coded reason 4 when it restreams a Peer
    Down; it could forward the recorded reason and NOTIFICATION.
  • Native bgp-tcp-in sessions' last_error doesn't include the RFC 8203
    shutdown communication; the decoder here could be shared.

amtypaldos and others added 4 commits October 1, 2026 15:56
A router sends a Peer Down Notification (RFC 7854 §4.9) when one of
its own BGP sessions goes down. Besides the per-peer header, it carries
a reason code and, depending on it, the BGP NOTIFICATION the router sent
(reason 1) or received (reason 3), or the FSM event that closed the
session (reason 2). peer_down() read only the per-peer header and threw
the rest away, so a BMP-monitored peer could only ever be shown as Idle:
a max-prefix teardown, an operator shutdown and a hold timer expiry all
looked the same.

Decode the message into a PeerDownInfo (time, reason, NOTIFICATION code
and subcode, RFC 8203/9003 shutdown communication, FSM event, and a
one-line description in the same Debug wording as bgp-tcp-in's
last_error) and store it as `last_down` on every view of the peer while
marking it Disconnected. The record survives the next PeerUp, because
update_info only merges fields that are set; a synthesized view does
not inherit it from the view it was copied from.

/api/v1/ingresses shows the structured record; /api/v1/bgp/neighbors
fills `lastError` and a new `lastDownTime` for BMP-monitored peers.

The reason code is read from the message itself because routecore's
PeerDownReason has no numeric value and folds RFC 9069's code 6 into
Unknown, and NOTIFICATION data is bounded by the NOTIFICATION's length
because routecore's data() runs to the end of the BMP message.

Limits: routers only report sessions that reached Established, so a
session that never comes up (e.g. bad peer AS) is not covered, and the
record goes away with the ingress when the rib GC reaps a peer that
never returns.

Test: decoding of every reason, shutdown communication edge cases, the
record on every view, kept across PeerUp, not inherited by synthesized
views; full library suite 357 passed, 31 ignored.

Signed-off-by: Alex Typaldos <alex@supertrace.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`show ip bgp neighbors` already printed `Last error` for sessions netom
terminates itself. BMP-monitored peers now carry the same field, filled
from the router's Peer Down Notification, plus `lastDownTime`; show the
latter as `Last down: <time> (<age> ago)`, reusing bmp.rs's
uptime_from(), so an operator can tell a session that dropped a minute
ago from one that has been down for a week.

The neighbor and ingress fixtures gain a Disconnected BMP peer that went
down on a Cease Administrative Shutdown with a shutdown communication.

Test: netom-cli suite 115 passed.

Signed-off-by: Alex Typaldos <alex@supertrace.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Add bmp_state_num_peer_down_notifications, a per-router counter of BMP
Peer Down Notifications labelled with the RFC 7854 §4.9 reason
(localNotification, remoteNotification, localFsm, ...). It is bumped
once per message, not once per view of the peer it takes down, and a
Peer Down for a peer that was never up is rejected before it counts.

Every reason is emitted, zero included, so rate() works from the first
scrape and an alert on rising localNotification (a router tearing
sessions down, typically on a prefix limit) needs no special casing.
The label set is the fixed list of reasons, so cardinality stays bounded.

Test: counted per reason, unknown peer not counted; full library suite
358 passed, 31 ignored.

Signed-off-by: Alex Typaldos <alex@supertrace.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The feeder's Peer Down now carries reason 3 with a Cease /
Administrative Shutdown NOTIFICATION and an RFC 8203 shutdown
communication instead of reason 4, and the driver asserts that netom
records it: `last_down` on the session in /api/v1/ingresses (reason,
code/subcode, shutdown communication, a time stamped by netom since the
feeder sends a zero timestamp) and `lastError` / `lastDownTime` in
/api/v1/bgp/neighbors. The existing "Peer Down emitted exactly once"
checks on the bmp-out side are unchanged.

Test: e2e-addpath-bmp: OK.

Signed-off-by: Alex Typaldos <alex@supertrace.ai>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Remote-controlled terminal text is unsanitized, and the metric omits parsed notifications contrary to its documented semantics.

Review effort: Balanced
Findings: 2 Medium severity

Open (2)
What changed in this PR

Records and exposes BMP Peer Down reasons across APIs, CLI output, metrics, and documentation.

Changes:

  • Decodes and persists Peer Down reason details and timestamps.
  • Exposes details through neighbor APIs, CLI output, and metrics.
  • Adds unit, fixture, and end-to-end coverage.
File Description
test-data/​cli/​ingresses.json Adds structured Peer Down fixture data.
test-data/​cli/​bgp-neighbors.json Adds a disconnected BMP neighbor fixture.
src/​units/​bmp_tcp_in/​state_machine/​tests.rs Tests persistence, views, and metrics.
src/​units/​bmp_tcp_in/​state_machine/​status_reporter.rs Increments per-reason counters.
src/​units/​bmp_tcp_in/​state_machine/​peer_down.rs Decodes Peer Down details.
src/​units/​bmp_tcp_in/​state_machine/​mod.rs Registers the decoder module.
src/​units/​bmp_tcp_in/​state_machine/​metrics.rs Defines and exports the counter.
src/​units/​bmp_tcp_in/​state_machine/​machine.rs Records reasons during peer teardown.
src/​units/​bgp_tcp_in/​http_ng.rs Exposes reason and time in neighbor APIs.
src/​tests/​util.rs Adds BMP/BGP notification builders.
src/​ingress/​register.rs Stores persistent Peer Down information.
src/​bin/​netom-cli/​commands/​bmp.rs Shares timestamp-age formatting.
src/​bin/​netom-cli/​commands/​bgp.rs Renders last error and down time.
scripts/​e2e-addpath-bmp.sh Documents expanded E2E scope.
scripts/​e2e-addpath-bmp.py Verifies API Peer Down output.
docs/​rib-query-api.md Documents new API fields.
docs/​cli.md Documents CLI output.
docs/​bmp-tcp-in.md Documents Peer Down handling and metrics.
doc/​netom-cli.1 Updates the CLI manual.
Changelog.md Announces the feature.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +726 to +729
self.status_reporter.peer_down_notification(
self.router_id.clone(),
last_down.reason,
);
Comment on lines +89 to +91
Some(text) => format!(
"{side} NOTIFICATION: {details:?} \"{text}\""
),

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants