Skip to content

LR2021 continuous mode race condition / packet corruption - #3512

Draft
carlhodder wants to merge 1 commit into
meshcore-dev:mainfrom
carlhodder:lr2021_basic_continuous_mode_fixes
Draft

carlhodder wants to merge 1 commit into
meshcore-dev:mainfrom
carlhodder:lr2021_basic_continuous_mode_fixes

Conversation

@carlhodder

Copy link
Copy Markdown

This draft PR is for discussion and hopefully some help testing first.

I've been running LR2021's for a while and have been seeing it repeat made up packets.

There is also open another PR about this: #3261 - I wasn't able to replicate what was seen there. I did alter my code to enforce and log CMD_DAT for the few calls that read that way but I haven't logged anything for over a week (easily >200K packets, probably more).

There's a chunk of my journey to this point in that ticket - basically I believe this comes down to something odd and a reasonable race condition (given this introduced continuous RX mode).

Solution notes

I've tested a handful of different approaches and I decided that I just dont think it's worth the complexity in order to save what forms ~0.1% of events and instead just:

  • Reject any reads where the buffer length differs from the packet length

I'd still prefer to keep continuous mode as from testing it nets us about 5% more packets over one-shot, but if one-shot RX is desired that's an alternate fall back (still needs the 0-length-fifo check though, so it's not much different outside of the extra changes).

Odd Behaviour + The Suspected Race

Odd Behaviour:
What is not so reasonable is my 2 test units both have valid RX_DONE interrupts, a packet length, but 0 data in the buffer. In these instances reading an empty buffer most frequently just repeats a single byte. Most get rejected, but some make it out (these are what started me down this path):

151515151515151515151515151515151515151515151515151515151515151515151515151515151515151515151515151515

0102010D010101010101010101010101010101010101010101

And this occurs in one-shot mode too, sometimes there just isn't data in the buffer? And I meant we never got it - I've logged events where it's been seconds till the next packet that arrives with no extra data. These form about 2/3rds of events where the fifo level doesn't match the packet.

I would love to know if anyone else sees this? I would've thought it was hardware specific except it occurs on all my boards.

The Race Condition:
The second issue is introduces a race condition - we can have quite large delays between the RX_DONE occurring and the packet being read. In this test code the RX_DONE is printed from within the setFlag IRQ (naughty) and the IRQ gets printed from RadioLibWrapper.cpp's recvRaw:

[1972808]: RX_DONE

[1973725]: DEBUG: IRQ: 262512 (getRxFifoLevel: 153, getPacketLength: 77)

That's almost a second between the IRQ occurring and following call to recvRaw - probably why there's now an extra packet of data in the buffer.

The first thing here is that currently we do not check, and with multiple packet's in the buffer we would read the data for the first packet with the packet length of the current, which may differ.

It can also coincide with one of those pesky 0-buffer events and you see packets like:

0504F917850D638F26DA8B824DDD67314E78D702C629884F470A184DDA91DEB1E16729284727FF0C78FC95D33EEBF0B08818181818181818181818181818181818181818181818181818

050AF917A70F17DE80C4B61216F63C82A4A10CF3D28B65DFBEBEBEBEBEBEBEBEBEBEBEBEBEBE71B2DE8933A14A876C44B85164C191411E0C38B665D1D8E8F1327618894F6BE9EBDD8838BD7DEE4ABBF735AD0E2C460580EE18BC3BB6B43E6A5EEE383A38B60BBC71FE885E72F8F2A0841AE1BDCDA82D19C51515374519D5FB39F0A7499E5EB604A788B43FB0A3C55DF12DB4F87BD1E688AAD4EF56B1B9C7A4A7

We can also end up reading a packet in the middle of it being written to the buffer - and clearing the fifo does not interrupt the existing write. The next read will start with these bytes - this means that if we have too many bytes we can't tell if we need to ignore the extra data at the start as old, or the extra data at the end as new.

'tis why I ultimately settled on simple approach.

Extra changes

There are also some minor alterations to isReceiving to better suit continuous RX mode (from testing).

LR2021::isReceiving: Remove header CRC error reset
This flag isn't used by RadioLib, and as some calls can be quite late if we had good packet -> CRC error or reverse this blows away a good packet we could have decoded (clearing it's valid header IRQ will make RadioLib reject it with -24).

I figure as we use this for what is basically CAD - if the header has a CRC error the channel will still be busy for some time, and these are rare. I've seen this drop ~1% of packets (during busy periods).

LR2021::isReceiving: Do not return if reset state detected
When we detect a reset we may have been a bit slow to get here, so fall through and look for a preamble.

RadioLibWrapper::.recvRaw: Ensure we can't get stuck
I explictly clear the RX_DONE interrupt in RadioLibWrapper.cpp's recvRaw - I have run continuous mode with my own IRQ clearing code for a couple of months and I've seen a couple of instances where my repeaters stop receiving packets and self recover some time later. I managed to see this latch on a test unit - it was in STATE_RX without the flag set but the IRQ (and the DIO pin) was set on the radio. It was the advert going out that kicked it out of this state.

It is possible for readData to return an error code without clearing the IRQs, which would leave us in this state. I haven't confirmed this as I've seen it maybe 3 times in 2 months, but it seem plausable so I've included it here.

A bit cheeky but I also updated the existing comment above the new IRQ clear line. The reason that it throws -706 is because the startReceive path calls setRxPath (datasheet: can only be called from standby mode).

Hardware

The two test units are base LR2021 (Waveshare Core2021-XF) and I have 1 in-field repeater with the same, and 2 with GNiceRF 1W LR2021F33 variant ('cause if you pop the can off you can just hack/fit a 4MHZ SAW on the RX path and basically ignore the nearby LoS LTE b8 band tower).

One of the test units uses a Seeed XIAO NRF52840 board, the other a clone. One in-field repeater uses the same XIAO board, and the other two use ProMicro NRF52840 clones. I also ran standard XIAO w/ SX1262 as a control for various tests when needed.

I have variants with direct-tied SPI and ones with 39R matching resistors - all show clean signals on the scope, and the evidences doesn't really align with SPI corruption so I don't believe that is the cause.

…ks for continuous mode

NOTE:
After trying a bunch of variants, these events are < 0.1% so I don't think it's worth the complexity to try read these, and the better approach is just reject all that risk having issues.
@dkmr27

dkmr27 commented Sep 28, 2026 •

Copy link
Copy Markdown

Nothing really to add to the issue, but as a longer term LR2021 user this is something I see a few times per day, packets with a single repeater byte like this
image

Can by any byte not always the same, if it matches a valid header like 0x15 it will be repeated

Looking at logs for the past 24 hours, around 12 occurrences, one every couple of hours

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants