<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Arif Iqbal]]></title><description><![CDATA[Arif Iqbal]]></description><link>https://arifiqbal.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>Arif Iqbal</title><link>https://arifiqbal.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Sat, 10 Oct 2026 05:03:26 GMT</lastBuildDate><atom:link href="https://arifiqbal.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Optimizing Wappalyzer + Playwright in Docker: Preventing Runaway CPU and Memory Usage]]></title><description><![CDATA[Browser automation can become surprisingly expensive when it runs continuously in a backend pipeline.
A task that looks simple at the application level:
Website → Detect technologies → Return JSON

ma]]></description><link>https://arifiqbal.hashnode.dev/optimizing-wappalyzer-playwright-in-docker-preventing-runaway-cpu-and-memory-usage</link><guid isPermaLink="true">https://arifiqbal.hashnode.dev/optimizing-wappalyzer-playwright-in-docker-preventing-runaway-cpu-and-memory-usage</guid><category><![CDATA[Docker]]></category><category><![CDATA[Devops]]></category><category><![CDATA[playwright]]></category><category><![CDATA[Python]]></category><category><![CDATA[performance]]></category><dc:creator><![CDATA[Arif Iqbal]]></dc:creator><pubDate>Fri, 02 Oct 2026 06:53:24 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/68d7beaf67224dcede72f02d/e3929353-6f21-4d81-a189-5564d633981a.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Browser automation can become surprisingly expensive when it runs continuously in a backend pipeline.</p>
<p>A task that looks simple at the application level:</p>
<pre><code class="language-text">Website → Detect technologies → Return JSON
</code></pre>
<p>may actually execute something closer to:</p>
<pre><code class="language-text">Application
    ↓
Docker CLI
    ↓
Docker Engine
    ↓
Wappalyzer container
    ↓
Playwright
    ↓
Chromium
    ├── Browser process
    ├── Renderer processes
    ├── GPU process
    └── Utility processes
</code></pre>
<p>I recently investigated a situation where Wappalyzer-based technology detection gradually caused high CPU and memory usage on a server.</p>
<p>The individual scans were supposed to have strict timeouts.</p>
<p>The Docker containers were also started with <code>--rm</code>.</p>
<p>Yet some containers survived far longer than expected, leaving Chromium processes consuming significant resources.</p>
<p>The interesting part was that the problem wasn't simply:</p>
<blockquote>
<p>Chromium uses a lot of memory.</p>
</blockquote>
<p>The real issue was <strong>lifecycle management</strong>.</p>
<p>This article explains the problem, why the obvious timeout implementation wasn't enough, and how to make this kind of workload much safer.</p>
<hr />
<h2>The Architecture</h2>
<p>Each technology-detection request launched Wappalyzer inside its own Docker container.</p>
<p>A simplified command looked like this:</p>
<pre><code class="language-bash">docker run --rm \
  wappalyzer-image \
  -i https://example.com \
  --scan-type full \
  -t 30 \
  -oJ -
</code></pre>
<p>The application launched that command asynchronously:</p>
<pre><code class="language-python">process = await asyncio.create_subprocess_exec(
    *command,
    stdout=asyncio.subprocess.PIPE,
    stderr=asyncio.subprocess.PIPE,
)
</code></pre>
<p>It then waited for the scan to finish:</p>
<pre><code class="language-python">stdout, stderr = await asyncio.wait_for(
    process.communicate(),
    timeout=30,
)
</code></pre>
<p>If the timeout was exceeded:</p>
<pre><code class="language-python">process.kill()
await process.wait()
</code></pre>
<p>At first glance, this looks reasonable.</p>
<p>There is a timeout.</p>
<p>There is process cleanup.</p>
<p>And Docker is using:</p>
<pre><code class="language-bash">--rm
</code></pre>
<p>So why could containers remain alive?</p>
<hr />
<h2>The First Sign of Trouble</h2>
<p>The first indication was unusually high server resource consumption.</p>
<p>For this kind of issue, <code>docker stats</code> is one of the quickest places to start:</p>
<pre><code class="language-bash">docker stats
</code></pre>
<p>It lets you see:</p>
<ul>
<li><p>CPU consumption</p>
</li>
<li><p>memory consumption</p>
</li>
<li><p>memory limits</p>
</li>
<li><p>network activity</p>
</li>
<li><p>block I/O</p>
</li>
<li><p>PID counts</p>
</li>
</ul>
<p>The important discovery was that some temporary Wappalyzer containers were consuming significant resources and had been running much longer than expected.</p>
<p>The next step was checking Chromium processes.</p>
<p>For Playwright installations, something like this can be useful:</p>
<pre><code class="language-bash">pgrep -af '/ms-playwright.*chrome'
</code></pre>
<p>And if Wappalyzer uses identifiable Chromium profile directories:</p>
<pre><code class="language-bash">pgrep -af 'wappalyzer-chromium'
</code></pre>
<p>There were many Chromium processes associated with technology-detection jobs.</p>
<p>Chromium spawning multiple processes is normal.</p>
<p>A single browser instance may create separate processes for renderers, networking, GPU operations, storage, and other browser services.</p>
<p>The abnormal part was the <strong>lifetime</strong> of those processes.</p>
<p>A scan intended to finish in approximately 30 seconds should not leave a container alive for minutes—or potentially much longer.</p>
<hr />
<h1>Looking at the Host, Not Just Docker</h1>
<p>Container metrics only tell part of the story.</p>
<p>I also checked the host using:</p>
<pre><code class="language-bash">vmstat 1
</code></pre>
<p>This is particularly useful when diagnosing whether the machine is experiencing:</p>
<ul>
<li><p>CPU saturation</p>
</li>
<li><p>high runnable process queues</p>
</li>
<li><p>swapping</p>
</li>
<li><p>I/O waiting</p>
</li>
<li><p>virtualization steal time</p>
</li>
</ul>
<p>For example, the important columns include:</p>
<pre><code class="language-text">r   runnable processes
si  swap in
so  swap out
us  user CPU
sy  system CPU
id  idle CPU
wa  I/O wait
st  stolen CPU
</code></pre>
<p>When browser workloads are overwhelming a machine, you may see the runnable queue increase while CPU idle time approaches zero.</p>
<p>That gives a much better picture than looking at application logs alone.</p>
<hr />
<h1>The Important Process Boundary</h1>
<p>The root cause becomes clearer when you look at what the application is actually controlling.</p>
<p>The process tree conceptually looks like this:</p>
<pre><code class="language-text">Python application
        │
        ▼
    docker run
        │
        ▼
   Docker Engine
        │
        ▼
     Container
        │
        ▼
    Wappalyzer
        │
        ▼
    Playwright
        │
        ▼
     Chromium
</code></pre>
<p>This distinction matters.</p>
<p>When Python executes:</p>
<pre><code class="language-python">process = await asyncio.create_subprocess_exec(...)
</code></pre>
<p>the subprocess being managed is:</p>
<pre><code class="language-text">docker run
</code></pre>
<p>It is <strong>not Chromium</strong>.</p>
<p>It is also not the Docker container itself.</p>
<p>The Docker CLI communicates with Docker Engine, which creates and manages the container independently.</p>
<p>Therefore:</p>
<pre><code class="language-python">process.kill()
</code></pre>
<p>means roughly:</p>
<blockquote>
<p>Kill the local Docker CLI process that Python started.</p>
</blockquote>
<p>It does <strong>not provide the stronger guarantee</strong>:</p>
<blockquote>
<p>Destroy the container associated with this job and terminate everything running inside it.</p>
</blockquote>
<p>That difference is easy to miss.</p>
<hr />
<h1>Why <code>docker run --rm</code> Isn't a Watchdog</h1>
<p>Another assumption that can cause problems is treating:</p>
<pre><code class="language-bash">docker run --rm
</code></pre>
<p>as lifecycle protection.</p>
<p><code>--rm</code> is useful, but its job is more specific.</p>
<p>Conceptually:</p>
<pre><code class="language-text">Container exits
      ↓
Docker automatically removes it
</code></pre>
<p>It prevents exited temporary containers from accumulating.</p>
<p>But consider this situation:</p>
<pre><code class="language-text">Python timeout occurs
        ↓
docker CLI process is killed
        ↓
Docker container remains alive
        ↓
Chromium remains alive
</code></pre>
<p>The container hasn't exited.</p>
<p>Therefore there is nothing for <code>--rm</code> to remove yet.</p>
<p>So:</p>
<blockquote>
<p><code>--rm</code> is automatic post-exit cleanup, not a container watchdog.</p>
</blockquote>
<p>That's an important distinction for dynamically created containers.</p>
<hr />
<h1>Give Every Temporary Container an Identity</h1>
<p>The first improvement is simple:</p>
<p><strong>Name every job container.</strong></p>
<p>Generate a unique identifier:</p>
<pre><code class="language-python">import uuid

container_name = f"wappalyzer-scan-{uuid.uuid4().hex}"
</code></pre>
<p>Then launch the container with:</p>
<pre><code class="language-bash">docker run --rm \
  --name wappalyzer-scan-&lt;uuid&gt; \
  ...
</code></pre>
<p>Now there is an explicit relationship between the application job and the Docker resource:</p>
<pre><code class="language-text">Application job
      ↕
Container name
</code></pre>
<p>This becomes extremely useful when something goes wrong.</p>
<p>Instead of trying to discover which anonymous Docker container belongs to the failed scan, the application already knows exactly what to remove.</p>
<hr />
<h1>Use Layered Timeouts</h1>
<p>Another problem was having the internal Wappalyzer timeout and the application watchdog effectively use the same value.</p>
<p>For example:</p>
<pre><code class="language-text">Wappalyzer timeout: 30 seconds
Application timeout: 30 seconds
</code></pre>
<p>That creates a race.</p>
<p>At approximately 30 seconds, Wappalyzer may be trying to:</p>
<pre><code class="language-text">finish the scan
    ↓
collect results
    ↓
serialize JSON
    ↓
close Playwright
    ↓
terminate Chromium
    ↓
exit
</code></pre>
<p>At exactly the same time, the parent process may decide:</p>
<pre><code class="language-text">TIMEOUT
</code></pre>
<p>and kill the Docker CLI.</p>
<p>A cleaner model is to separate the two responsibilities.</p>
<p>For example:</p>
<pre><code class="language-text">Internal scan timeout: 30 seconds
Hard watchdog:         45 seconds
</code></pre>
<p>The internal timeout controls normal Wappalyzer behavior.</p>
<p>The outer timeout protects the application from a container that fails to terminate correctly.</p>
<p>In Python:</p>
<pre><code class="language-python">scan_timeout = 30
hard_timeout_grace = 15

hard_timeout = scan_timeout + hard_timeout_grace
</code></pre>
<p>Then:</p>
<pre><code class="language-python">stdout, stderr = await asyncio.wait_for(
    process.communicate(),
    timeout=hard_timeout,
)
</code></pre>
<p>Now Wappalyzer has time to perform normal cleanup before the application applies the hard limit.</p>
<hr />
<h1>Explicitly Destroy Timed-Out Containers</h1>
<p>This is the most important change.</p>
<p>When the hard timeout occurs, don't only kill the Docker CLI process.</p>
<p>Also tell Docker Engine to remove the corresponding container.</p>
<p>Conceptually:</p>
<pre><code class="language-python">try:
    stdout, stderr = await asyncio.wait_for(
        process.communicate(),
        timeout=hard_timeout,
    )

except TimeoutError:
    await terminate_process(process)
    await remove_container(container_name)
    raise
</code></pre>
<p>The cleanup function can execute:</p>
<pre><code class="language-bash">docker rm -f &lt;container-name&gt;
</code></pre>
<p>For example:</p>
<pre><code class="language-python">async def remove_container(container_name: str) -&gt; None:
    cleanup = await asyncio.create_subprocess_exec(
        "docker",
        "rm",
        "-f",
        container_name,
        stdout=asyncio.subprocess.DEVNULL,
        stderr=asyncio.subprocess.DEVNULL,
    )

    await asyncio.wait_for(
        cleanup.wait(),
        timeout=10,
    )
</code></pre>
<p>This creates a much stronger guarantee.</p>
<p>Instead of saying:</p>
<pre><code class="language-text">Stop waiting for docker run.
</code></pre>
<p>the application now says:</p>
<pre><code class="language-text">This job exceeded its maximum lifetime.

Destroy its container.
</code></pre>
<p>When Docker force-removes the container, the Chromium process tree inside it is terminated as well.</p>
<hr />
<h1>Handle Async Cancellation Too</h1>
<p>Timeout isn't the only abnormal exit path.</p>
<p>An asynchronous task can also be cancelled because of:</p>
<ul>
<li><p>application shutdown</p>
</li>
<li><p>worker shutdown</p>
</li>
<li><p>orchestration changes</p>
</li>
<li><p>task cancellation</p>
</li>
<li><p>deployment</p>
</li>
<li><p>upstream failure</p>
</li>
</ul>
<p>So cleanup should also handle cancellation.</p>
<p>For example:</p>
<pre><code class="language-python">except (TimeoutError, asyncio.CancelledError):
    await terminate_process(process)
    await remove_container(container_name)
    raise
</code></pre>
<p>The goal is to maintain this invariant:</p>
<pre><code class="language-text">Application job no longer exists
             =
Temporary job container no longer exists
</code></pre>
<p>Without cancellation cleanup, you can fix timeout leaks while still leaving another path for orphaned containers.</p>
<hr />
<h1>Add CPU Limits</h1>
<p>Even with perfect cleanup, browser automation can be CPU intensive.</p>
<p>A problematic website can cause Chromium to perform expensive rendering, JavaScript execution, network processing, or browser work.</p>
<p>Without resource limits:</p>
<pre><code class="language-bash">docker run ...
</code></pre>
<p>the container can potentially consume substantial host CPU.</p>
<p>Docker allows you to put a ceiling on this:</p>
<pre><code class="language-bash">--cpus 1
</code></pre>
<p>For example:</p>
<pre><code class="language-bash">docker run \
  --cpus 1 \
  ...
</code></pre>
<p>This doesn't mean one CPU is the correct value for every environment.</p>
<p>The right value depends on:</p>
<ul>
<li><p>server hardware</p>
</li>
<li><p>expected scan duration</p>
</li>
<li><p>concurrency</p>
</li>
<li><p>target websites</p>
</li>
<li><p>acceptable throughput</p>
</li>
</ul>
<p>But the principle is important:</p>
<blockquote>
<p>One technology-detection job shouldn't automatically have unrestricted access to host CPU resources.</p>
</blockquote>
<p>Start with a conservative limit and benchmark it.</p>
<hr />
<h1>Add Memory Limits</h1>
<p>The same principle applies to memory.</p>
<p>Chromium can consume significant RAM, especially on complex websites.</p>
<p>Docker can enforce a memory ceiling:</p>
<pre><code class="language-bash">--memory 1g
</code></pre>
<p>You can also restrict the container from consuming additional swap beyond that allocation:</p>
<pre><code class="language-bash">--memory-swap 1g
</code></pre>
<p>Combined:</p>
<pre><code class="language-bash">docker run --rm \
  --name wappalyzer-scan-&lt;uuid&gt; \
  --memory 1g \
  --memory-swap 1g \
  --cpus 1 \
  wappalyzer-image ...
</code></pre>
<p>The <code>1g</code> value here should be treated as an example starting point.</p>
<p>It should be tested against real workloads.</p>
<p>If legitimate scans repeatedly hit the memory ceiling, increase it based on measurements rather than removing the limit entirely.</p>
<p>The objective is containment.</p>
<p>Without limits:</p>
<pre><code class="language-text">Bad scan
    ↓
Chromium consumes more resources
    ↓
Host becomes unstable
</code></pre>
<p>With limits:</p>
<pre><code class="language-text">Bad scan
    ↓
Container reaches resource boundary
    ↓
Impact remains contained
</code></pre>
<hr />
<h1>Resource Limits Don't Replace Concurrency Control</h1>
<p>Per-container limits solve one part of the problem.</p>
<p>Concurrency solves another.</p>
<p>Suppose each container is allowed:</p>
<pre><code class="language-text">1 CPU
1 GB RAM
</code></pre>
<p>One container has a predictable maximum impact.</p>
<p>But if the application launches:</p>
<pre><code class="language-text">20 containers
</code></pre>
<p>the aggregate workload can still overwhelm the host.</p>
<p>A useful way to think about capacity is:</p>
<pre><code class="language-text">Maximum concurrent jobs
        ×
Per-job resource ceiling
        ≈
Maximum browser workload
</code></pre>
<p>For example:</p>
<pre><code class="language-text">2 concurrent jobs × 1 GB
≈ 2 GB potential browser memory
</code></pre>
<p>versus:</p>
<pre><code class="language-text">10 concurrent jobs × 1 GB
≈ 10 GB potential browser memory
</code></pre>
<p>The limits don't mean every container will consume its maximum allocation, but they give you a useful upper bound.</p>
<p>So production browser workloads generally need both:</p>
<pre><code class="language-text">Per-job resource limits
          +
Concurrency control
</code></pre>
<hr />
<h1>Monitor Temporary Containers</h1>
<p>Short-lived containers can be difficult to observe because they may appear and disappear between normal <code>docker ps</code> commands.</p>
<p>A simple watcher helps:</p>
<pre><code class="language-bash">watch -n 0.5 \
  'docker ps --filter ancestor=wappalyzer-image \
  --format "table {{.Names}}\t{{.Status}}\t{{.RunningFor}}"'
</code></pre>
<p>During normal operation, you should see:</p>
<pre><code class="language-text">wappalyzer-scan-a1b2c3...
</code></pre>
<p>Then a few seconds later:</p>
<pre><code class="language-text">&lt;container disappears&gt;
</code></pre>
<p>That is healthy behavior.</p>
<p>The lifecycle should look like:</p>
<pre><code class="language-text">Container created
      ↓
Wappalyzer starts
      ↓
Chromium starts
      ↓
Detection completes
      ↓
Chromium exits
      ↓
Container exits
      ↓
--rm removes container
</code></pre>
<p>What you don't want is:</p>
<pre><code class="language-text">Container created
      ↓
Expected timeout passes
      ↓
Hard watchdog passes
      ↓
Container still running
</code></pre>
<p>If that happens, the cleanup mechanism needs investigation.</p>
<hr />
<h1>Monitor Application-Level Events Too</h1>
<p>Docker monitoring tells you what's happening at the infrastructure layer.</p>
<p>Application logs tell you whether scans are actually progressing.</p>
<p>For example, logging events such as:</p>
<pre><code class="language-text">technology_detection_started
technology_detection_completed
</code></pre>
<p>makes it much easier to correlate:</p>
<pre><code class="language-text">Application job
      ↕
Docker container
      ↕
Website scan
</code></pre>
<p>A scan might produce:</p>
<pre><code class="language-text">technology_detection_started
...
technology_detection_completed
</code></pre>
<p>within 10–30 seconds.</p>
<p>That is very different from seeing a Docker container alive for several minutes without a corresponding active application task.</p>
<p>Combining application logs with Docker-level monitoring is much more useful than relying on either one alone.</p>
<hr />
<h1>Verify the Docker Limits</h1>
<p>Don't assume that because your application constructed the command correctly, Docker actually received the expected configuration.</p>
<p>Inspect a running scan:</p>
<pre><code class="language-bash">docker inspect wappalyzer-scan-&lt;uuid&gt; \
  --format 'Memory={{.HostConfig.Memory}} MemorySwap={{.HostConfig.MemorySwap}} NanoCpus={{.HostConfig.NanoCpus}}'
</code></pre>
<p>You can also inspect real-time resource consumption:</p>
<pre><code class="language-bash">docker stats --no-stream wappalyzer-scan-&lt;uuid&gt;
</code></pre>
<p>This verifies the configuration at the actual container boundary.</p>
<p>For infrastructure problems, verification at the infrastructure layer is important.</p>
<hr />
<h1>Test the Failure Path</h1>
<p>Lifecycle code is most important when something fails.</p>
<p>So don't only test successful Wappalyzer responses.</p>
<p>Useful test cases include:</p>
<pre><code class="language-text">Successful scan
Default timeout
Custom timeout
Provider failure
Command construction
Hard timeout
Task cancellation
Container cleanup
</code></pre>
<p>For example, the timeout test can deliberately make:</p>
<pre><code class="language-python">process.communicate()
</code></pre>
<p>take longer than the hard watchdog.</p>
<p>Then verify that both operations occur:</p>
<pre><code class="language-text">Docker CLI subprocess terminated
              +
Docker container explicitly removed
</code></pre>
<p>Cancellation should receive similar coverage.</p>
<p>This is especially important because orphan-container bugs often don't appear during ordinary happy-path testing.</p>
<hr />
<h1>A Safer Implementation</h1>
<p>Putting the pieces together, the provider can follow a structure similar to this:</p>
<pre><code class="language-python">import asyncio
import uuid


class TechnologyDetector:
    def __init__(
        self,
        image: str,
        timeout_seconds: int = 30,
        hard_timeout_grace_seconds: int = 15,
        memory_limit: str = "1g",
        cpu_limit: str = "1",
    ):
        self.image = image
        self.timeout_seconds = timeout_seconds
        self.hard_timeout_grace_seconds = hard_timeout_grace_seconds
        self.memory_limit = memory_limit
        self.cpu_limit = cpu_limit


    async def detect(self, website: str) -&gt; str:
        scan_timeout = self.timeout_seconds

        hard_timeout = (
            scan_timeout
            + self.hard_timeout_grace_seconds
        )

        container_name = (
            f"wappalyzer-scan-{uuid.uuid4().hex}"
        )

        command = [
            "docker",
            "run",
            "--rm",
            "--name",
            container_name,
            "--memory",
            self.memory_limit,
            "--memory-swap",
            self.memory_limit,
            "--cpus",
            self.cpu_limit,
            self.image,
            "-i",
            website,
            "--scan-type",
            "full",
            "-t",
            str(scan_timeout),
            "-oJ",
            "-",
        ]

        process = await asyncio.create_subprocess_exec(
            *command,
            stdout=asyncio.subprocess.PIPE,
            stderr=asyncio.subprocess.PIPE,
        )

        try:
            stdout, stderr = await asyncio.wait_for(
                process.communicate(),
                timeout=hard_timeout,
            )

        except (TimeoutError, asyncio.CancelledError):
            await self._terminate_process(process)
            await self._remove_container(container_name)
            raise

        if process.returncode != 0:
            raise RuntimeError(
                stderr.decode(errors="replace")
            )

        return stdout.decode()


    @staticmethod
    async def _terminate_process(process) -&gt; None:
        if process.returncode is not None:
            return

        try:
            process.kill()
        except ProcessLookupError:
            return

        await process.wait()


    async def _remove_container(
        self,
        container_name: str,
    ) -&gt; None:
        try:
            cleanup = await asyncio.create_subprocess_exec(
                "docker",
                "rm",
                "-f",
                container_name,
                stdout=asyncio.subprocess.DEVNULL,
                stderr=asyncio.subprocess.DEVNULL,
            )

            await asyncio.wait_for(
                cleanup.wait(),
                timeout=10,
            )

        except (
            TimeoutError,
            FileNotFoundError,
            OSError,
        ):
            return
</code></pre>
<p>This isn't meant to be a drop-in library implementation.</p>
<p>The important concepts are the lifecycle guarantees:</p>
<pre><code class="language-text">Unique container identity
        +
Internal timeout
        +
Hard watchdog
        +
Explicit container cleanup
        +
CPU limit
        +
Memory limit
        +
Cancellation cleanup
</code></pre>
<hr />
<h1>Before and After</h1>
<p>The difference can be summarized simply.</p>
<h2>Before</h2>
<pre><code class="language-text">Application
    ↓
docker run --rm
    ↓
Wappalyzer
    ↓
Chromium

30-second timeout
    ↓
Kill docker CLI
    ↓
Assume everything disappeared
</code></pre>
<p>The dangerous word there is:</p>
<pre><code class="language-text">Assume
</code></pre>
<hr />
<h2>After</h2>
<pre><code class="language-text">Application
    ↓
Generate unique container name
    ↓
Start resource-limited container
    ↓
Wappalyzer
    ↓
Playwright / Chromium

Internal timeout
    ↓
Grace period
    ↓
Hard watchdog
    ↓
Terminate CLI if necessary
    ↓
docker rm -f &lt;known-container&gt;
</code></pre>
<p>Now cleanup is explicit.</p>
<hr />
<h1>The Broader Lesson</h1>
<p>This problem isn't specific to Wappalyzer.</p>
<p>The same principle applies to workloads involving:</p>
<ul>
<li><p>Playwright</p>
</li>
<li><p>Puppeteer</p>
</li>
<li><p>Selenium</p>
</li>
<li><p>headless Chromium</p>
</li>
<li><p>browser-based crawlers</p>
</li>
<li><p>screenshot services</p>
</li>
<li><p>PDF rendering</p>
</li>
<li><p>automated testing</p>
</li>
<li><p>scraping infrastructure</p>
</li>
</ul>
<p>The important realization is that there are multiple lifecycle boundaries:</p>
<pre><code class="language-text">Application task
      ↓
Subprocess
      ↓
Docker container
      ↓
Browser process tree
</code></pre>
<p>Controlling one does not automatically mean you control the others.</p>
<p>A Python subprocess timeout controls the subprocess.</p>
<p>A Docker container limit controls the container.</p>
<p>A browser timeout controls browser operations.</p>
<p>Each boundary needs to be considered explicitly.</p>
<hr />
<h1>Practical Checklist</h1>
<p>If you're running Wappalyzer or another browser-heavy workload inside temporary Docker containers, check the following:</p>
<ul>
<li><p>Give every temporary container a unique name.</p>
</li>
<li><p>Keep <code>--rm</code>, but don't rely on it as a watchdog.</p>
</li>
<li><p>Set an internal scan timeout.</p>
</li>
<li><p>Give the outer watchdog additional cleanup time.</p>
</li>
<li><p>Explicitly <code>docker rm -f</code> containers that exceed the hard timeout.</p>
</li>
<li><p>Handle async cancellation as well as timeout.</p>
</li>
<li><p>Add CPU limits.</p>
</li>
<li><p>Add memory limits.</p>
</li>
<li><p>Limit overall concurrency.</p>
</li>
<li><p>Monitor Docker resource consumption.</p>
</li>
<li><p>Monitor browser process counts.</p>
</li>
<li><p>Correlate application events with container lifetime.</p>
</li>
<li><p>Test timeout and cancellation paths.</p>
</li>
<li><p>Verify Docker limits with <code>docker inspect</code>.</p>
</li>
<li><p>Benchmark limits against real websites before deploying them broadly.</p>
</li>
</ul>
<hr />
<h1>Final Takeaway</h1>
<p>The biggest lesson from this issue was simple:</p>
<blockquote>
<p><strong>A subprocess timeout, a Docker container lifetime, and a browser lifetime are not the same thing.</strong></p>
</blockquote>
<p>When an application dynamically creates browser containers, it should know exactly which container belongs to each job and be capable of explicitly destroying that container when the job ends abnormally.</p>
<p>Resource limits provide the second layer of protection.</p>
<p>Even if cleanup fails, a single browser job shouldn't be able to consume unrestricted CPU or memory.</p>
<p>The combination of:</p>
<pre><code class="language-text">Unique identity
+ resource limits
+ layered timeouts
+ explicit cleanup
+ concurrency control
+ monitoring
</code></pre>
<p>turns a potentially unpredictable browser workload into something much easier to operate safely.</p>
<p>And one final lesson worth remembering:</p>
<p><code>docker run --rm</code> <strong>is cleanup convenience. It is not lifecycle management.</strong></p>
]]></content:encoded></item></channel></rss>