Skip to content

Hidden <body> adds a fixed 30s to every crawl, even when ignore_body_visibility=True (the default) #2129

Description

@mvletter

Summary

_crawl_web waits for document.body to become visible using a hardcoded 30s timeout (async_crawler_strategy.py#L815-L827 on v0.9.2). The result is only acted on when ignore_body_visibility=False — and that flag defaults to True.

So on any page where the body never becomes visible, every crawl spends 30 seconds producing a value that is then discarded.

This isn't an exotic case: it is exactly what AngularJS ng-cloak (and Vue's v-cloak) leave behind whenever the app fails to bootstrap. The page still renders, the crawl still succeeds — it's just 30s slower, every single time.

Reproduction

Self-contained, no external site needed. The two variants differ only by an ng-cloak attribute on <body>:

import asyncio, threading, time
from http.server import BaseHTTPRequestHandler, HTTPServer
from crawl4ai import AsyncWebCrawler, BrowserConfig, CacheMode, CrawlerRunConfig

PROSE = "<p>Real, indexable prose content that the crawler should return.</p>" * 40
HIDDEN = f"""<!DOCTYPE html><html><head>
<style>[ng-cloak] {{ display: none; }}</style></head>
<body ng-cloak><h1>Hello</h1>{PROSE}</body></html>""".encode()
VISIBLE = HIDDEN.replace(b" ng-cloak", b"", 1)

class Handler(BaseHTTPRequestHandler):
    def do_GET(self):
        body = VISIBLE if self.path == "/visible" else HIDDEN
        self.send_response(200)
        self.send_header("Content-Type", "text/html")
        self.send_header("Content-Length", str(len(body)))
        self.end_headers()
        self.wfile.write(body)
    def log_message(self, *a): pass

async def main():
    threading.Thread(target=HTTPServer(("127.0.0.1", 8897), Handler).serve_forever,
                     daemon=True).start()
    await asyncio.sleep(0.3)
    async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
        for label, path in (("visible body", "/visible"), ("hidden body", "/hidden")):
            t = time.time()
            r = await crawler.arun(url=f"http://127.0.0.1:8897{path}",
                                   config=CrawlerRunConfig(cache_mode=CacheMode.BYPASS))
            print(f"  {label:<14} {time.time()-t:6.1f}s  success={r.success}  html={len(r.html or '')}b")

asyncio.run(main())

Output:

  visible body      0.3s  success=True  html=2834b
  hidden body      30.4s  success=True  html=2846b

Real-world impact

On an Oracle B2C helpdesk we crawl, every URL took 34–35s. For comparison, from the same container: curl returned in 3.9s, and plain Playwright with the same wait_until in 4.7s. Instrumenting the pipeline attributed 30.1s of the 34.7s to csp_compliant_wait. Reference sites on that same container: 0.2s and 1.9s.

Reproduced independently on macOS and on a Linux server, so it is not environment-specific.

With the timeout lowered to 2000ms:

page default (30000) 2000ms extracted words
helpdesk home 35.2s 6.1s 259 → 259
helpdesk article 34.2s 6.5s 962 → 962
intermedia.com (visible body) 1.8s 1.0s unaffected
example.com 0.3s 0.2s unaffected

Same content, ~5x faster.

Suggested fix

Make the timeout configurable rather than changing behaviour — e.g. CrawlerRunConfig.body_visibility_timeout, defaulting to 30000 so existing behaviour is byte-for-byte unchanged. Callers who know their pages cloak the body can lower it.

I deliberately did not propose skipping the wait when ignore_body_visibility=True. I measured that variant too, and it returns less content (160 vs 259 words on the real page), because the wait doubles as render time. A configurable ceiling keeps that trade-off with the caller.

I have a patch (6 lines, all additive: the new field in async_configs.py plus the one call site) and verified against v0.9.2:

  • tests/test_config_defaults.py: 33 passed, identical before and after
  • serialisation round-trips (to_dict, from_kwargs, dump/load, clone) preserve the field
  • default value reproduces current timing exactly (30.2s on the repro above)

Happy to open a PR if you'd like it.

Environment

  • crawl4ai 0.9.2 (verified in source) and 0.8.9 (Docker, where we hit it)
  • Chromium via Playwright, headless
  • macOS 15 and Debian container

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions