You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
With SEARCH_EXTRACTOR_TYPE=tika, creating one small file in a space causes the search
service to re-download and re-extract the whole space through Tika, even though the space
finished indexing 80 minutes earlier and nothing else changed in between.
Measured amplification: 1 uploaded file → 3,767 file download events in 7 minutes, still
climbing linearly when I killed the container. The instance is unusable for the duration
(~30 min when left to run), and ends with the OpenCloud process OOM-killed.
The first time I hit this was not through normal use but by running the reindex command that the 8.x.x upgrade documentation
tells you to run after upgrading, which took the host down hard and never completed. Details
in "First occurrence" below.
This is filed separately from #3561 because the observable there is cosmetic (directory
timestamps change during a manual --force-rescan), whereas here the same write path is
reached by ordinary user activity and becomes self-amplifying. If the fix in flight for #3561
also stops this, please close as a duplicate — but the amplification behaviour seems worth
tracking explicitly, see "Why the in-flight fix may not fully close this" below.
Steps to reproduce
Small setup — no large space needed to see the amplification factor, only to make it hurt:
Deploy 8.0.1 via Docker Compose with search/tika.yml, i.e. SEARCH_EXTRACTOR_TYPE=tika.
opencloud 1.513GiB / 3GiB cpu = 140-184%
tika 830MiB cpu = 32-48%
host load average 1.24 -> 7.54
Space involved: personal space, 7.7 GB, 65,992 documents in the bleve index.
This was the third identical storm the same day. The two earlier ones were left to run and each
ended with the container OOM-killed and restarted (RestartCount=2); one logged 9,035 file download events out of 9,080 total log lines for the same space. Their triggers were
equally trivial — one file deletion, and adding a paragraph to an existing text file.
First occurrence: the documented 8.x upgrade reindex
Worth flagging separately, because this is not an exotic configuration — the first and worst
incident came from following the official upgrade procedure.
opencloud search index --all-spaces --force-rescan --insecure
Running exactly that on a 16 GB host exhausted all host RAM, drove the machine into swap
thrash and left it completely unresponsive — no SSH, no web UI, not even a clean shutdown. It
required a hard power reset. (Kernel pstore records show an earlier instance of the same
pattern tripping the hardware watchdog.) This was before I applied any container memory limit.
The run also never completed. Log from a detached retry:
=== started Thu Sep 17 02:12:06 UTC 2026 ===
[1/9] ... 8b9ab109 ... in 7.948795465s
[2/9] ... 7e71937e ... in 8.465562913s
[3/9] ... 50e09751 ... in 1m9.92305518s
[4/9] ... b7c59e27 ... in 2m26.01669506s
[5/9] ... 32f23407 ... in 769.477271ms
[6/9] ... e5800d53 ... in 2m32.327717915s
[7/9] ... 7679356b ... in 2m24.611489319s
rpc error: code = Unavailable desc = error reading from server: EOF
=== exit=1 finished Thu Sep 17 04:28:36 UTC 2026 ===
Seven spaces completed in about nine minutes of reported time. The eighth — a 60 GB personal
space — then ran for roughly two hours before the RPC died with Unavailable ... EOF. A
multi-hour traversal failing that way looks like #3541 (long-running IndexSpace reusing an
expired service token).
Two documentation/UX consequences that seem worth addressing regardless of the indexing bug:
Because the documented command includes --force-rescan, every retry restarts from zero
instead of resuming. Three further attempts all died in the same space. A space that
cannot finish in a single uninterrupted pass can therefore never be indexed at all — it is
stuck at whatever partial state the first pass reached. (Mine still is: 27,868 of its
documents are in the index, and no amount of re-running the documented command improves
that.)
The documented command has no memory bound and no guidance about one, so on a memory-
constrained host the documented upgrade step can take the whole machine down rather than
just failing.
What I ruled out
These rule-outs concern the upload-triggered storms above, not the upgrade reindex:
Not a continuation of the initial index. bleve root.bolt for this space was written at
21:04 UTC; there were zero file download events from 21:06 to 22:24 UTC. The storm starts
within seconds of the upload after 78 minutes of complete silence.
Not --force-rescan. No CLI indexing command was running at the time. The trigger is a
normal web UI upload. (The separate upgrade-reindex incident below did use --force-rescan, as the documentation prescribes.)
Not an unindexed space. This space had just completed a full index (65,992 docs).
That propagation looks like a content change to the event stream.
The search service re-indexes the affected subtree, extracting via Tika again — back to
step 2, now fanned out across siblings.
That would explain the shape of the curve: a slow first minute (7 events) while the cascade
gets going, then a sustained plateau of ~700/min as it spreads across the space, with no
natural termination.
I have not proven step 4→5. The decisive test is whether individual files are downloaded
more than once during a storm; the container was removed before I could extract per-resource
counts, which took its logs with it. It reproduces on demand, so I'm happy to capture whatever
would be most useful — per-resourceid download counts, debug-level search service logs, the
event stream, or a NATS main-queue dump. Just say which.
Why the in-flight fix may not fully close this
PR #806 ("Do not change mtimes when nothing has changed") addresses the spurious writes, and @aduffeck's comment indicates the .flock guard-file path is covered too. That should remove this particular source of self-generated change events.
The amplification itself looks like a separate, latent property though: the indexer consumes
an event stream that its own writes can feed, with no provenance check and no per-space
coalescing. Any future write-back during extraction re-creates the same unbounded loop. Two
things that would make it fail safe rather than fail catastrophically:
Provenance suppression — tag change events originating from the search service's own
metadata writes so the indexer ignores them.
Per-space coalescing / debounce — collapse repeated re-index requests for the same space
within a window, so a burst of propagated changes costs one pass rather than N.
Also worth noting for operators: SEARCH_REINDEX_CONCURRENCY bounds parallelism but not total
work, so it doesn't help here.
Workarounds (for anyone hitting this)
SEARCH_EXTRACTOR_TYPE=basic avoids it entirely, at the cost of content extraction for
anything but plain text.
A container mem_limit confines the damage to OpenCloud instead of the host — but it must
be paired with a matching memswap_limit. With mem_limit alone the container spills past
the cap into swap and thrashes the host, which is worse than no limit at all. Neither
prevents the outage.
Setup
OpenCloud 8.0.1 rolling (opencloudeu/opencloud-rolling:8.0.1, compiled 2026-09-16)
Docker Compose deployment including search/tika.yml
Summary
With
SEARCH_EXTRACTOR_TYPE=tika, creating one small file in a space causes the searchservice to re-download and re-extract the whole space through Tika, even though the space
finished indexing 80 minutes earlier and nothing else changed in between.
Measured amplification: 1 uploaded file → 3,767
file downloadevents in 7 minutes, stillclimbing linearly when I killed the container. The instance is unusable for the duration
(~30 min when left to run), and ends with the OpenCloud process OOM-killed.
The first time I hit this was not through normal use but by running the reindex command that
the 8.x.x upgrade documentation
tells you to run after upgrading, which took the host down hard and never completed. Details
in "First occurrence" below.
This is filed separately from #3561 because the observable there is cosmetic (directory
timestamps change during a manual
--force-rescan), whereas here the same write path isreached by ordinary user activity and becomes self-amplifying. If the fix in flight for #3561
also stops this, please close as a duplicate — but the amplification behaviour seems worth
tracking explicitly, see "Why the in-flight fix may not fully close this" below.
Steps to reproduce
Small setup — no large space needed to see the amplification factor, only to make it hurt:
search/tika.yml, i.e.SEARCH_EXTRACTOR_TYPE=tika.(JPEGs with EXIF, MP3s, PDFs — per @aduffeck's analysis in Re-indexing the search, using Tika, propagates a tree change and updates directories'
mtimeon storage #3561 the trigger is Tikaproducing metadata that gets written back via
SetArbitraryMetadata).file downloadevents are beinglogged.
file downloadevents per minute:Expected: a handful of events for the one new file, then silence. Observed: sustained hundreds
per minute across files that did not change.
Measurements from the live incident
file downloadevents per minute (UTC), from a standing start, all for the space that receivedthe single new file:
Resource use during the storm:
Space involved: personal space, 7.7 GB, 65,992 documents in the bleve index.
This was the third identical storm the same day. The two earlier ones were left to run and each
ended with the container OOM-killed and restarted (
RestartCount=2); one logged 9,035file downloadevents out of 9,080 total log lines for the same space. Their triggers wereequally trivial — one file deletion, and adding a paragraph to an existing text file.
First occurrence: the documented 8.x upgrade reindex
Worth flagging separately, because this is not an exotic configuration — the first and worst
incident came from following the official upgrade procedure.
The 8.x.x upgrade documentation
instructs running:
Running exactly that on a 16 GB host exhausted all host RAM, drove the machine into swap
thrash and left it completely unresponsive — no SSH, no web UI, not even a clean shutdown. It
required a hard power reset. (Kernel pstore records show an earlier instance of the same
pattern tripping the hardware watchdog.) This was before I applied any container memory limit.
The run also never completed. Log from a detached retry:
Seven spaces completed in about nine minutes of reported time. The eighth — a 60 GB personal
space — then ran for roughly two hours before the RPC died with
Unavailable ... EOF. Amulti-hour traversal failing that way looks like #3541 (long-running
IndexSpacereusing anexpired service token).
Two documentation/UX consequences that seem worth addressing regardless of the indexing bug:
--force-rescan, every retry restarts from zeroinstead of resuming. Three further attempts all died in the same space. A space that
cannot finish in a single uninterrupted pass can therefore never be indexed at all — it is
stuck at whatever partial state the first pass reached. (Mine still is: 27,868 of its
documents are in the index, and no amount of re-running the documented command improves
that.)
constrained host the documented upgrade step can take the whole machine down rather than
just failing.
What I ruled out
These rule-outs concern the upload-triggered storms above, not the upgrade reindex:
root.boltfor this space was written at21:04 UTC; there were zero
file downloadevents from 21:06 to 22:24 UTC. The storm startswithin seconds of the upload after 78 minutes of complete silence.
--force-rescan. No CLI indexing command was running at the time. The trigger is anormal web UI upload. (The separate upgrade-reindex incident below did use
--force-rescan, as the documentation prescribes.)io_msshowed roughly 33% utilisation during the storm,so the array is not the bottleneck; this is consistent with Large uploads stuck in "processing" forever on slow-fsync storage (ZFS without SLOG) embedded NATS drops the finalize event #3027's observation that the
limiting factor is fsync count rather than bandwidth.
Hypothesised mechanism
Chaining @aduffeck's analysis in #3561 with what I observed:
SetArbitraryMetadataon the extracted file.mtime, and per @v-scharf's follow-up in Re-indexing the search, using Tika, propagates a tree change and updates directories'mtimeon storage #3561, theetagof all parent folders as well.step 2, now fanned out across siblings.
That would explain the shape of the curve: a slow first minute (7 events) while the cascade
gets going, then a sustained plateau of ~700/min as it spreads across the space, with no
natural termination.
I have not proven step 4→5. The decisive test is whether individual files are downloaded
more than once during a storm; the container was removed before I could extract per-resource
counts, which took its logs with it. It reproduces on demand, so I'm happy to capture whatever
would be most useful — per-
resourceiddownload counts, debug-level search service logs, theevent stream, or a NATS
main-queuedump. Just say which.Why the in-flight fix may not fully close this
PR #806 ("Do not change mtimes when nothing has changed") addresses the spurious writes, and
@aduffeck's comment indicates the
.flockguard-file path is covered too. That should removethis particular source of self-generated change events.
The amplification itself looks like a separate, latent property though: the indexer consumes
an event stream that its own writes can feed, with no provenance check and no per-space
coalescing. Any future write-back during extraction re-creates the same unbounded loop. Two
things that would make it fail safe rather than fail catastrophically:
metadata writes so the indexer ignores them.
within a window, so a burst of propagated changes costs one pass rather than N.
Also worth noting for operators:
SEARCH_REINDEX_CONCURRENCYbounds parallelism but not totalwork, so it doesn't help here.
Workarounds (for anyone hitting this)
SEARCH_EXTRACTOR_TYPE=basicavoids it entirely, at the cost of content extraction foranything but plain text.
mem_limitconfines the damage to OpenCloud instead of the host — but it mustbe paired with a matching
memswap_limit. Withmem_limitalone the container spills pastthe cap into swap and thrashes the host, which is worse than no limit at all. Neither
prevents the outage.
Setup
opencloudeu/opencloud-rolling:8.0.1, compiled 2026-09-16)search/tika.ymlSEARCH_EXTRACTOR_TYPE=tika,FRONTEND_FULL_TEXT_SEARCH_ENABLED=true, defaultSEARCH_REINDEX_CONCURRENCYPossibly related
mtimeon storage #3561 — re-indexing with Tika propagates a tree change and updates directorymtime(likely the same root cause, different observable)
IndexSpacetraversal reuses an expired service token