Repository navigation
connectd: only charge gossip queries against the CPU budget - #9601
Merged
nGoline merged 3 commits intoOct 6, 2026
Merged
Conversation
jaonoctus
reviewed
Oct 6, 2026
nGoline
force-pushed
the
fix/gossip-throttle-ordinary-gossip
branch
from
October 6, 2026 13:41
92d9b52 to
9dccb6e
Compare
jaonoctus
approved these changes
Oct 6, 2026
nGoline
force-pushed
the
fix/gossip-throttle-ordinary-gossip
branch
from
October 6, 2026 15:27
9dccb6e to
f0c3928
Compare
--dev-throttle-gossip shrinks the traffic limits along with the CPU budget, so a test can't tell which one throttled a peer. This sets only the total CPU budget for answering gossip queries (in usec per second), leaving the traffic limits alone. Changelog-None Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The CPU throttle charges every message connectd handles locally: pings, pongs, onion messages, custom messages and gossip forwarded to gossipd, not only the gossip queries it exists to bound. The per-peer share is the budget divided by the number of connected peers, so on a busy node ordinary traffic exhausts it and we stop reading from the peer. Changelog-None Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The CPU throttle charged every message connectd handles itself against a budget split evenly across all connected peers, and on the way out it charged every check for pending query replies, which runs each time the output queue drains. On a node with many peers each share is tiny, so ordinary traffic (gossip, pings, onion messages) used it up and we stopped reading from the peer for a second or more, delaying its channel messages too, with an UNUSUAL "Throttling ... peer ...: too much CPU". Charge only the gossip queries the throttle exists to bound: reading them, and answering them while one is pending. Onion messages have their own ratelimit. Keep the even split, with no minimum share: a minimum would let enough peers claim more than the whole budget between them, and a small share now only slows that peer's queries. Log throttling at debug: a peer hitting its budget is what it's for. Changelog-Fixed: connectd: ordinary gossip, pings and onion messages no longer count against the gossip query CPU budget, so busy nodes no longer throttle their peers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
nGoline
force-pushed
the
fix/gossip-throttle-ordinary-gossip
branch
from
October 6, 2026 18:42
f0c3928 to
0926d93
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This fixes CI after the gossip CPU throttle that shipped in v26.06.8. Since that change, valgrind runs fail on multi-peer offers tests (
test_offer_with_private_channels_multyhop2,test_renepay.py::test_offersandtest_offer_pathsfailed on #9582) with:The throttle is meant to bound the work of answering gossip queries, but it charged much more than that:
query_channel_range/query_short_channel_ids;Under valgrind those charges add up quickly, and the same thing happens on busy mainnet nodes: connectd stops reading from a peer for a second or more, which also delays its channel messages.
Commits:
lightningd: add --dev-gossip-cpu-budget: sets only the CPU budget.--dev-throttle-gossipalso shrinks the traffic limits, so a test using it can't tell which limit throttled a peer.tests: ordinary messages must not count against the gossip CPU budget(xfail(strict=True)): 200 pings to a node with a 10 usec budget and the normal traffic limits must not get the peer throttled.connectd: only charge gossip queries against the CPU budget: charge only reading the two query messages and answering them while one is pending (onion messages already have their own per-peer ratelimit), keep the even per-peer split with no minimum share (a minimum would let enough peers claim more than the whole budget), and log throttling once at debug. Removes the xfail.Answering queries still counts, so
test_gossip_query_channel_range_cpu_throttlestill sees a flood of queries throttled.Testing (docker):
Throttling incoming/outgoing peer ...: too much CPUon pings alone, and passes after it;test_gossip.py(59 passed),test_askrene.py(37),test_xpay.py(35), all throttle tests, source checks;VALGRIND=1:test_offer_with_private_channels_multyhop2,test_renepay.py::test_offersandtest_offer_pathspass.🤖 Generated with Claude Code