Skip to content

Single file upload triggers a full re-index of an already-indexed space with Tika, exhausting memory #3589

Description

@helderjgoncalves

Summary

With SEARCH_EXTRACTOR_TYPE=tika, creating one small file in a space causes the search
service to re-download and re-extract the whole space through Tika, even though the space
finished indexing 80 minutes earlier and nothing else changed in between.

Measured amplification: 1 uploaded file → 3,767 file download events in 7 minutes, still
climbing linearly when I killed the container. The instance is unusable for the duration
(~30 min when left to run), and ends with the OpenCloud process OOM-killed.

The first time I hit this was not through normal use but by running the reindex command that
the 8.x.x upgrade documentation
tells you to run after upgrading, which took the host down hard and never completed. Details
in "First occurrence" below.

This is filed separately from #3561 because the observable there is cosmetic (directory
timestamps change during a manual --force-rescan), whereas here the same write path is
reached by ordinary user activity and becomes self-amplifying. If the fix in flight for #3561
also stops this, please close as a duplicate — but the amplification behaviour seems worth
tracking explicitly, see "Why the in-flight fix may not fully close this" below.

Steps to reproduce

Small setup — no large space needed to see the amplification factor, only to make it hurt:

  1. Deploy 8.0.1 via Docker Compose with search/tika.yml, i.e. SEARCH_EXTRACTOR_TYPE=tika.
  2. In a personal space, create a folder with, say, 200 files that Tika extracts metadata from
    (JPEGs with EXIF, MP3s, PDFs — per @aduffeck's analysis in Re-indexing the search, using Tika, propagates a tree change and updates directories' mtime on storage #3561 the trigger is Tika
    producing metadata that gets written back via SetArbitraryMetadata).
  3. Index the space and wait for it to go quiet. Confirm no file download events are being
    logged.
  4. Upload one small text file into that space.
  5. Count file download events per minute:
docker compose logs -f --timestamps opencloud \
  | grep '"file download"' \
  | awk '{print substr($1,12,5)}' | uniq -c

Expected: a handful of events for the one new file, then silence. Observed: sustained hundreds
per minute across files that did not change.

Measurements from the live incident

file download events per minute (UTC), from a standing start, all for the space that received
the single new file:

22:24     7     <- file created here
22:25   417
22:26   698
22:27   706
22:28   702
22:29   644
22:30   600     <- container stopped manually

Resource use during the storm:

opencloud   1.513GiB / 3GiB   cpu = 140-184%
tika        830MiB            cpu = 32-48%
host load average   1.24 -> 7.54

Space involved: personal space, 7.7 GB, 65,992 documents in the bleve index.

This was the third identical storm the same day. The two earlier ones were left to run and each
ended with the container OOM-killed and restarted (RestartCount=2); one logged 9,035
file download events out of 9,080 total log lines for the same space. Their triggers were
equally trivial — one file deletion, and adding a paragraph to an existing text file.

First occurrence: the documented 8.x upgrade reindex

Worth flagging separately, because this is not an exotic configuration — the first and worst
incident came from following the official upgrade procedure.

The 8.x.x upgrade documentation
instructs running:

opencloud search index --all-spaces --force-rescan --insecure

Running exactly that on a 16 GB host exhausted all host RAM, drove the machine into swap
thrash and left it completely unresponsive — no SSH, no web UI, not even a clean shutdown. It
required a hard power reset. (Kernel pstore records show an earlier instance of the same
pattern tripping the hardware watchdog.) This was before I applied any container memory limit.

The run also never completed. Log from a detached retry:

=== started Thu Sep 17 02:12:06 UTC 2026 ===
[1/9] ... 8b9ab109 ... in 7.948795465s
[2/9] ... 7e71937e ... in 8.465562913s
[3/9] ... 50e09751 ... in 1m9.92305518s
[4/9] ... b7c59e27 ... in 2m26.01669506s
[5/9] ... 32f23407 ... in 769.477271ms
[6/9] ... e5800d53 ... in 2m32.327717915s
[7/9] ... 7679356b ... in 2m24.611489319s
rpc error: code = Unavailable desc = error reading from server: EOF
=== exit=1 finished Thu Sep 17 04:28:36 UTC 2026 ===

Seven spaces completed in about nine minutes of reported time. The eighth — a 60 GB personal
space — then ran for roughly two hours before the RPC died with Unavailable ... EOF. A
multi-hour traversal failing that way looks like #3541 (long-running IndexSpace reusing an
expired service token).

Two documentation/UX consequences that seem worth addressing regardless of the indexing bug:

  • Because the documented command includes --force-rescan, every retry restarts from zero
    instead of resuming.
    Three further attempts all died in the same space. A space that
    cannot finish in a single uninterrupted pass can therefore never be indexed at all — it is
    stuck at whatever partial state the first pass reached. (Mine still is: 27,868 of its
    documents are in the index, and no amount of re-running the documented command improves
    that.)
  • The documented command has no memory bound and no guidance about one, so on a memory-
    constrained host the documented upgrade step can take the whole machine down rather than
    just failing.

What I ruled out

These rule-outs concern the upload-triggered storms above, not the upgrade reindex:

  • Not a continuation of the initial index. bleve root.bolt for this space was written at
    21:04 UTC; there were zero file download events from 21:06 to 22:24 UTC. The storm starts
    within seconds of the upload after 78 minutes of complete silence.
  • Not --force-rescan. No CLI indexing command was running at the time. The trigger is a
    normal web UI upload. (The separate upgrade-reindex incident below did use
    --force-rescan, as the documentation prescribes.)
  • Not an unindexed space. This space had just completed a full index (65,992 docs).
  • Not disk saturation. Per-device io_ms showed roughly 33% utilisation during the storm,
    so the array is not the bottleneck; this is consistent with Large uploads stuck in "processing" forever on slow-fsync storage (ZFS without SLOG) embedded NATS drops the finalize event #3027's observation that the
    limiting factor is fsync count rather than bandwidth.

Hypothesised mechanism

Chaining @aduffeck's analysis in #3561 with what I observed:

  1. The upload event makes the search service index the new file.
  2. Tika extraction calls SetArbitraryMetadata on the extracted file.
  3. That bumps the parent directory's mtime, and per @v-scharf's follow-up in Re-indexing the search, using Tika, propagates a tree change and updates directories' mtime on storage #3561, the
    etag of all parent folders as well.
  4. That propagation looks like a content change to the event stream.
  5. The search service re-indexes the affected subtree, extracting via Tika again — back to
    step 2, now fanned out across siblings.

That would explain the shape of the curve: a slow first minute (7 events) while the cascade
gets going, then a sustained plateau of ~700/min as it spreads across the space, with no
natural termination.

I have not proven step 4→5. The decisive test is whether individual files are downloaded
more than once during a storm; the container was removed before I could extract per-resource
counts, which took its logs with it. It reproduces on demand, so I'm happy to capture whatever
would be most useful — per-resourceid download counts, debug-level search service logs, the
event stream, or a NATS main-queue dump. Just say which.

Why the in-flight fix may not fully close this

PR #806 ("Do not change mtimes when nothing has changed") addresses the spurious writes, and
@aduffeck's comment indicates the .flock guard-file path is covered too. That should remove
this particular source of self-generated change events.

The amplification itself looks like a separate, latent property though: the indexer consumes
an event stream that its own writes can feed, with no provenance check and no per-space
coalescing. Any future write-back during extraction re-creates the same unbounded loop. Two
things that would make it fail safe rather than fail catastrophically:

  • Provenance suppression — tag change events originating from the search service's own
    metadata writes so the indexer ignores them.
  • Per-space coalescing / debounce — collapse repeated re-index requests for the same space
    within a window, so a burst of propagated changes costs one pass rather than N.

Also worth noting for operators: SEARCH_REINDEX_CONCURRENCY bounds parallelism but not total
work, so it doesn't help here.

Workarounds (for anyone hitting this)

  • SEARCH_EXTRACTOR_TYPE=basic avoids it entirely, at the cost of content extraction for
    anything but plain text.
  • A container mem_limit confines the damage to OpenCloud instead of the host — but it must
    be paired with a matching memswap_limit
    . With mem_limit alone the container spills past
    the cap into swap and thrashes the host, which is worse than no limit at all. Neither
    prevents the outage.

Setup

  • OpenCloud 8.0.1 rolling (opencloudeu/opencloud-rolling:8.0.1, compiled 2026-09-16)
  • Docker Compose deployment including search/tika.yml
  • SEARCH_EXTRACTOR_TYPE=tika, FRONTEND_FULL_TEXT_SEARCH_ENABLED=true, default
    SEARCH_REINDEX_CONCURRENCY
  • Embedded NATS JetStream, store directory on the same spinning-disk array as user data
  • Host: QNAP NAS, x86_64, 16 GB RAM, kernel 5.10.60-qnap; user data on HDD RAID

Possibly related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions