Chrome Page Rendering Appears to Hang in a Cloud Agent and Is Terminated by Cloudflare 520 After Approximately 120 Seconds

Issue Summary

While running a PPT page generation task in an ANNA App, the page source code was generated successfully, and the backend also produced the page HTML and PNG screenshot. However, the frontend never received a successful rendering result.

Each rendering call failed after running for approximately 133 seconds. The frontend received the following error through ANNA Runtime’s tools.invoke:

Unexpected token '<', "<!DOCTYPE "... is not valid JSON

The task retried three times, and every attempt produced exactly the same error. Because the application currently classifies this as a page rendering error, it repeatedly asked the page Agent to modify TSX code that could already be rendered successfully. Modifying the page code did not resolve the issue.

Based on the rendering artifacts in the diagnostics bundle, the consistent failure duration, and a Cloudflare HTTP 520 HTML response recorded during the same task, I suspect the actual failure is as follows: Chrome or a Puppeteer operation in the cloud Agent hangs at some point after rendering has completed. As a result, the tool invocation cannot return in time and is eventually terminated by an upstream Cloudflare or gateway timeout.

At this point, I cannot confirm whether the Chrome hang is caused by the cloud Agent’s CPU, memory, process limit, or another resource constraint. I would appreciate ANNA’s assistance in reviewing the cloud runtime environment and server-side logs.

Environment and Scenario

  • Application type: ANNA App
  • Backend capability: Executa bundled with the App
  • Executa: ppt-engine
  • Tool method: app_render_workspace_page_preview
  • Execution environment: Cloud Agent provided by ANNA
  • Task: Generate one 1280 × 720 PPT page
  • Workspace ID: ppt-20260730-072715
  • Generation Run ID: 9a5b3e5f-32c7-4336-89a6-525d6a611b53
  • Approximate incident time: July 30, 2026, 07:32–07:42 UTC

Expected Behavior

  1. The frontend calls tools.invoke through ANNA Runtime.
  2. ppt-engine starts managed Chrome instances and generates static HTML and a PNG screenshot.
  3. The Executa returns a JSON Tool Result containing fields such as html_path and screenshot_path.
  4. The frontend receives the result, updates the page status to accepted, and proceeds with final deck rendering.

Actual Behavior

  1. The page Agent successfully generates the TSX source.

  2. ppt-engine starts rendering the page preview.

  3. Valid HTML and PNG files appear in the shadow Workspace, and the PNG can be opened successfully.

  4. The frontend does not receive a successful JSON Tool Result.

  5. After approximately 133 seconds, tools.invoke throws a JSON parsing error:

    Unexpected token '<', "<!DOCTYPE "... is not valid JSON
    
  6. The system records the attempt as a render failure and starts a render-fix Agent.

  7. The page TSX is modified three times, but every attempt still fails with the same error after approximately 133 seconds.

  8. The user eventually abandons the task, and the Generation Run status becomes abandoned.

The duration of all three failures is highly consistent:

Stage Previous page Agent completion Next render-fix start Approximate wait
First render 07:32:33 UTC 07:34:46 UTC ~133 seconds
Second render 07:35:44 UTC 07:37:58 UTC ~134 seconds
Third render 07:38:57 UTC 07:41:10 UTC ~133 seconds

This fixed timing pattern is more consistent with an upstream timeout than with a TSX compilation error or an ordinary page runtime exception.

Application Implementation

Frontend Invocation

The frontend invokes the Executa through ANNA Runtime:

runtime.tools.invoke({
  tool_id: pptEngineToolId,
  method: "app_render_workspace_page_preview",
  args: {
    workspace_dir,
    page_id,
  },
  timeoutMs: 600_000,
});

The application configures a 600-second timeout for this long-running tool invocation. Therefore, the failure after approximately 133 seconds is not triggered by the application’s own 600-second timer.

ANNA Runtime receives an RPC Result or RPC Error from the Host through postMessage. The current failure is returned to the application as an RPC Error. The application itself does not directly call JSON.parse() on the page HTML.

Executa Rendering Flow

The primary steps performed by app_render_workspace_page_preview in ppt-engine are:

  1. Read the Workspace’s manifest.json.
  2. Run a pre-render TypeScript check on the target TSX file.
  3. Build the single-page HTML.
  4. Start a managed Chrome instance to convert the runtime HTML into static HTML.
  5. Close that Chrome page and browser instance.
  6. Start another managed Chrome instance, load the static HTML, and generate a PNG screenshot.
  7. Close the second Chrome page and browser instance.
  8. Return a JSON result containing the HTML and screenshot paths.

The current cleanup logic waits directly for Puppeteer:

await page.close?.().catch(() => undefined);
await browser.close().catch(() => undefined);

Although exceptions thrown during cleanup are ignored, no timeout is applied. If page.close() or browser.close() never resolves, the entire tool invocation will remain pending.

In addition to the cleanup stage, launch(), newPage(), page navigation, screenshot capture, or Chrome DevTools Protocol communication could also be affected by cloud resource constraints or the container runtime environment. The current diagnostics bundle does not record the start and completion times of these individual sub-stages, so I cannot determine with certainty which operation is hanging.

Diagnostic Evidence

1. Page Rendering Artifacts Exist

The shadow Workspace contains:

output/page-preview-html/*.html
output/screenshots/*.png

The PNG can be opened successfully, and both the page content and images are rendered correctly. This indicates that the page TSX, image files, Chrome page loading, and screenshot pipeline successfully reached at least the artifact-writing stage.

2. Modifying the Page Code Does Not Change the Error

The page Agent made the following changes in sequence:

  1. Added an alt attribute to an <img> element.
  2. Replaced a text structure containing <br /> with <p> elements.
  3. Rewrote the page as a simpler React component.

After all three changes, the error remained:

Unexpected token '<', "<!DOCTYPE "... is not valid JSON

Therefore, the issue does not appear to be caused by a particular HTML tag, page layout, or TSX syntax error.

3. A Cloudflare HTTP 520 HTML Response Was Explicitly Recorded During the Same Task

The Storage Transfer Log for the same task records a Host Upload negotiation failure:

upload HTTP 520: {'detail': '<!DOCTYPE html>
<!--[if lt IE 7]> <html class="no-js ie6 oldie" lang="en-US"> ...

This confirms that, during the same cloud execution, an upstream ANNA-related endpoint returned a Cloudflare-style HTML error page instead of the expected JSON response.

The raw HTTP response for the rendering tool invocation was not included in the diagnostics bundle. Therefore, I cannot state conclusively that the rendering failure was caused by the same HTTP 520 response. However, the error shape is consistent, and the rendering failures occur at a highly repeatable interval, suggesting the same category of gateway or upstream timeout failure.

4. Page Progress Remains at rendering

The diagnostics bundle records the following page progress:

{
  "status": "rendering",
  "render_attempts": 3,
  "last_html_path": "",
  "last_screenshot_path": ""
}

This means the application started another rendering attempt but never received the Tool Result, so it had no opportunity to persist the paths of the artifacts that had already been generated.

Current Assessment

The most likely failure sequence is:

ppt-engine starts Chrome
→ HTML/PNG artifacts are generated
→ Chrome, Puppeteer, or process cleanup hangs
→ The Executa Tool Invoke remains pending
→ An upstream gateway terminates the request after approximately 120 seconds and returns a Cloudflare HTML error page
→ The Host attempts to parse the HTML as JSON
→ The frontend receives Unexpected token '<'
→ The application misclassifies an infrastructure failure as a page-code error and repeatedly runs render-fix

I suspect that the cloud Agent may have insufficient resources, or that Chrome cannot exit reliably under the current container constraints. Possible causes include:

  • Insufficient memory, causing a Chrome renderer process to be killed by the OOM killer;
  • A low CPU quota or severe CPU throttling;
  • Exhaustion of the PID or file descriptor limit;
  • Failure to exit or reap Chrome child processes correctly;
  • Slow container disk I/O;
  • Unresponsive IPC between Puppeteer and Chrome during shutdown;
  • Timed-out invocations continuing to run and leaving Chrome processes behind, which then accumulate across retries.

Insufficient resources are still a hypothesis. Confirmation requires the cloud Agent’s resource configuration, container metrics, and server-side logs from ANNA.

Requested Assistance from ANNA

I would appreciate ANNA’s help with the following:

  1. What is the current resource configuration of the cloud Agent?

    • Number of CPU cores or CPU quota;
    • Memory limit;
    • /dev/shm size;
    • PID limit;
    • File descriptor limit;
    • Temporary disk capacity and I/O limits.
  2. Can the cloud Agent configuration for this App/Executa be temporarily or permanently increased?

    • I would like to reproduce the same task with more CPU and memory;
    • If the platform supports multiple resource tiers, please provide the recommended minimum configuration for Chrome/Puppeteer workloads.
  3. Please review the cloud logs and resource metrics for this task.

    • Workspace ID: ppt-20260730-072715;
    • Generation Run ID: 9a5b3e5f-32c7-4336-89a6-525d6a611b53;
    • Primary time range: July 30, 2026, 07:32–07:42 UTC;
    • Check for OOM events, CPU throttling, Chrome crashes, orphaned processes, or Tool backend timeouts;
    • Determine which upstream request Cloudflare or the gateway terminated at approximately 120 seconds.
  4. Please confirm whether the backend task continues running after a Tool Invoke timeout.

    • If the frontend or gateway has already returned an error, do the Executa and Chrome processes continue running in the background?
    • Could this leave Chrome child processes behind and affect subsequent retries?
  5. Please preserve clearer error details at the Host layer.

    • Check the HTTP status and Content-Type before parsing a response;
    • When an HTML error page is received, return a structured HTTP 520/502/504 error;
    • Do not return only Unexpected token '<' to the App;
    • Consider logging a redacted response-body summary, request correlation ID, and Cloudflare Ray ID, if available.

Planned Mitigations on Our Side

We plan to add the following safeguards to ppt-engine, but we still need ANNA to confirm whether the cloud environment has a resource or process-management issue:

  • Record separate durations for Chrome startup, page loading, screenshot capture, page.close(), and browser.close();
  • Add a timeout to browser cleanup and forcibly terminate the Chrome process if cleanup times out;
  • Classify Tool Invoke, gateway, and JSON response-parsing failures as infrastructure errors instead of asking the page Agent to modify TSX;
  • Apply limited retries to retryable Tool/Transport errors;
  • Record each browser process’s PID, exit code, and failure stage in the diagnostics bundle.

I would appreciate details of the cloud Agent’s actual resource specification and assistance determining whether this approximately 120-second interruption was caused by insufficient cloud resources, a Chrome process failing to exit, or a platform-level upstream timeout. If the Executa can temporarily be assigned a higher cloud resource tier, we can rerun the same minimal task to verify the behavior.

Hi @HappyLight! :waving_hand:

What a phenomenal report — your suspected failure chain (Chrome/cleanup hang → invocation held open → ~120s upstream cut → HTML error page parsed as JSON → misclassified as a page-code error) was spot on. :bullseye: Your timing table and the 520 evidence saved us a lot of digging. Thank you! :folded_hands:

TL;DR

We confirmed the issue and shipped fixes across the stack — this was actually several problems stacking on top of each other, and all of them are now addressed :white_check_mark::

  • Platform v1.1.0-beta.107 — error surfacing + timeout alignment
  • Agent v1.1.0-beta.27 — Chrome process cleanup & cancellation
  • @anna-ai/app-runtime 0.15.0 + CLI 0.1.43 — a new async job API built exactly for long-running tools like yours :rocket:

What we fixed

1. No more Unexpected token '<' :broom:

The host bridge now checks status and content type before parsing. If a gateway/edge HTML error page comes back, your app receives a structured http_<status> RPC error (with status, a redacted body snippet, the Cloudflare Ray ID when available, and a retriable flag) — exactly what you asked for in your assistance request #5. Your app can now correctly classify these as infrastructure errors instead of triggering render-fix loops.

2. Timeouts now fire before the edge cuts you off :stopwatch:

The sync tools.invoke ceiling was higher than what the network edge can actually hold open (~100s), so results past that point were physically undeliverable — and your 600s request was being silently clamped. The sync ceiling is now 90s, a structured tool_timeout error fires before the edge does, and clamping is no longer silent (the error reports your requested vs. max timeout).

3. Hung Chrome trees are now reaped :evergreen_tree:

You guessed right in question #4: previously, a timed-out invocation could leave the backend running, and cancellation didn’t reliably reach the agent — orphaned Chrome processes could accumulate across retries. Cancellation is now delivered on a fast path, and the agent terminates the entire process group (SIGTERM → SIGKILL), so hung Puppeteer/Chrome trees are fully cleaned up.

4. Beefier Cloud Agents for browser workloads :flexed_biceps:

We’ve raised the default Cloud Agent resources to 2 vCPU / 2 GB RAM (with /dev/shm ≈ 1 GB) — much more comfortable for Chrome/Puppeteer. This applies to newly created or replaced agent machines.

For renders that legitimately need >90s: async jobs! :sparkles:

With @anna-ai/app-runtime 0.15.0 you can use the new async job API instead of a long-held sync call:

const result = await runtime.tools.invokeAsyncAwait({
  tool_id: pptEngineToolId,
  method: "app_render_workspace_page_preview",
  args: { workspace_dir, page_id },
  timeoutMs: 600_000, // ✅ actually honored on the async path
});

It returns a job immediately, survives edge timeouts and even page reloads (listJobs/getJob), supports progress events and cancellation, and is fully supported in the local dev harness. This is the recommended path for your rendering pipeline. :light_bulb:

Your planned mitigations

Your list (stage timing, cleanup timeouts, error classification, bounded retries) is excellent defense-in-depth — we’d still encourage it! With the fixes above, transport failures now arrive as structured, classifiable errors, which should make that work much easier.

Please rerun your minimal task on the updated stack and tell us how it goes — we’d love to see those decks rendering smoothly! :yellow_heart:

Happy building! :rocket:
— The Anna Team

Thank you for the explanation and recommendation above. Following your guidance, we migrated long-running PPTX/PDF export operations from synchronous tools.invoke calls to the Anna Host Async Job API. The frontend now uses tools.invokeAsyncAwait from the official SDK, with the job timeout set to 10 minutes.

The migrated implementation works correctly in the Anna platform environment. However, while continuing local development and debugging, we found that invokeAsyncAwait is still unavailable in the environment started by anna-app dev, even after upgrading @anna-ai/cli to the latest version currently available, v0.1.45.

Below are the complete environment details, reproduction steps, and findings from our initial investigation. We hope this information helps identify a version compatibility issue in the local development runtime.

Summary

  • The same App code and invokeAsyncAwait usage work correctly in the Anna platform environment.
  • In the local anna-app dev --storage aps environment, the SDK does expose invokeAsyncAwait, but when it internally calls tools.invokeAsync, the local Host dispatcher returns unknown_method.
  • tools.listJobs, which is needed for job recovery, fails with the same type of error.
  • This therefore does not appear to be an issue with how the App calls invokeAsyncAwait. Instead, the Browser SDK, local runtime, and dispatcher capabilities bundled by the latest CLI do not appear to be fully aligned.

Local environment

Locally installed CLI version:

@anna-ai/cli v0.1.45

Startup command:

anna-app dev --storage aps

Relevant startup information printed by the CLI:

Anna App developer CLI
v0.1.45 — dev --storage aps

manifest          /Users/leyouming/company_program/anna/ppt-sdk/ppt-app/manifest.json
bundle            /Users/leyouming/company_program/anna/ppt-sdk/ppt-app/bundle/index.html
runtime           uvx anna-app-runtime-local@0.2.0a20
browser sdk       npm @anna-ai/app-runtime
storage backend   aps (real nexus APS via /api/v1/storage/*)
web backend       local (keyless ddgs + stdlib SSRF fetcher, unbilled)
dashboard         http://localhost:5180/

Inspection of the locally installed CLI package confirms the following related dependency versions:

@anna-ai/cli                 0.1.45
@anna-ai/app-runtime         0.15.0
@anna-ai/app-schema          0.19.0
anna-app-runtime-local       0.2.0a20

CLI v0.1.45 currently pins the local runtime command to:

uvx anna-app-runtime-local@0.2.0a20

How the API is called

The App uses the high-level invokeAsyncAwait helper provided by @anna-ai/app-runtime. The call has the following form:

await anna.tools.invokeAsyncAwait(
  {
    tool_id: toolId,
    method: "app_run_pptx_export",
    args: {
      workspace_dir: workspaceDir,
    },
    timeoutMs: 600_000,
    clientTag,
  },
  {
    timeoutMs: 600_000,
    signal,
    onProgress,
  },
);

PDF export and export-artifact publishing use the same approach, with different Executa methods.

We did not implement invokeAsyncAwait ourselves, nor did we wrap the request in an internal plugin JSON-RPC envelope. This directly uses anna.tools.invokeAsyncAwait(...) as exposed by the official Browser SDK.

Local reproduction steps

  1. Install or upgrade to @anna-ai/cli v0.1.45.

  2. Run the following command in the App project:

    anna-app dev --storage aps
    
  3. Open the local dashboard provided by the CLI.

  4. Start a PPTX or PDF export from the App using anna.tools.invokeAsyncAwait(...).

  5. The SDK fails while attempting to create the Host job.

  6. If the page attempts to recover an active job through anna.tools.listJobs(...), that call also fails.

Actual result

Creating the asynchronous job produces:

tools.invokeAsync is not defined
error: unknown_method

Querying recoverable jobs produces:

tools.listJobs is not defined
error: unknown_method

The apparent call path is:

anna.tools.invokeAsyncAwait(...)
  -> Browser SDK internally calls tools.invokeAsync
  -> local Host dispatcher returns unknown_method

In other words, the client-side invokeAsyncAwait method exists, but the underlying Host RPC methods on which it depends are not registered in the current local dispatcher.

Comparison with the platform environment

Environment Same call Result
Anna platform environment anna.tools.invokeAsyncAwait(...) Successfully creates the job, waits for completion, and returns the result
Local anna-app dev anna.tools.invokeAsyncAwait(...) tools.invokeAsync is not defined / unknown_method
Local anna-app dev anna.tools.listJobs(...) tools.listJobs is not defined / unknown_method

Because the same implementation works in the platform environment, the following parts appear to be valid:

  • The tool declaration in the App manifest;
  • The Executa tool ID and method arguments;
  • The invokeAsyncAwait input structure;
  • The 10-minute timeoutMs value;
  • The long-running Executa operation itself.

The remaining difference appears to be the Host/runtime combination used by the local anna-app dev environment.

Initial investigation of the local packages

The following findings are based on inspection of the locally installed packages. We hope they help narrow down the issue.

The Browser SDK includes the Async Job API

@anna-ai/app-runtime v0.15.0, installed with CLI v0.1.45, provides:

tools.invokeAsync
tools.getJob
tools.listJobs
tools.cancelJob
tools.invokeAsyncAwait

The SDK notes the following coordinated versions:

dispatcher_version 0.19.0
app-schema          0.19.0
anna-app-core       0.18.0

The SDK also states that these capabilities are shipped with Nexus >= 1.1.0-beta.108, with end-to-end asynchronous job execution available from beta.110.

The local runtime pinned by the CLI uses an earlier core version

CLI v0.1.45 pins anna-app-runtime-local v0.2.0a20, whose package metadata requires:

anna-app-core >= 0.16.0, < 0.17

The version resolved locally is:

anna-app-core 0.16.0

In the dispatcher registry for this version, we can currently see only the following tools methods:

tools.invoke
tools.list

The following methods do not appear to be registered:

tools.invokeAsync
tools.getJob
tools.listJobs
tools.cancelJob

This is consistent with the unknown_method response from the local runtime.

runtime-local appears to include a job implementation that is not connected to the dispatcher

The anna-app-runtime-local v0.2.0a20 package already contains jobs.py and job-related capabilities such as:

create_tool_job
get_tool_job
list_tool_jobs
cancel_tool_job

The local runtime therefore appears to contain at least part of the asynchronous job implementation. However, the anna-app-core v0.16.0 dispatcher installed with it does not publicly register the corresponding tools.* methods, so the iframe still receives unknown_method.

These findings are only our initial interpretation of the published packages. We would appreciate confirmation of the intended version combination and release relationship.

Expected result

The environment started by the latest @anna-ai/cli through anna-app dev should expose the same Async Job API as the platform environment, including at least:

tools.invokeAsync
tools.getJob
tools.listJobs
tools.cancelJob

The following call provided by the official SDK should work locally:

await anna.tools.invokeAsyncAwait(...);

This would allow developers to debug the full lifecycle of a long-running job locally, including job creation, progress updates, completion, failure, cancellation, and recovery after a page reload, without having to deploy the App to the platform first.

Questions for the Anna team

Could you please help confirm the following?

  1. Is anna-app-runtime-local v0.2.0a20, as pinned by @anna-ai/cli v0.1.45, expected to support invokeAsyncAwait?
  2. If it is expected to support it, could a runtime-local/core version containing the required dispatcher handlers be released and pinned by the CLI as a compatible set?
  3. If a compatible version already exists but is not public or must be selected explicitly, could you provide a temporary working CLI version, runtime-local version, or startup command?
  4. Could anna-app dev perform a Host API capability check during startup? If the Browser SDK exposes the Async Job API but the local dispatcher does not support it, the CLI could report a clear version incompatibility warning before an export is started, rather than returning unknown_method only at runtime.
  5. Could the official documentation include a compatibility matrix covering the CLI, Browser SDK, app-schema, anna-app-core, runtime-local, and Nexus versions? This would help developers determine whether an API is supported in both the platform and local development environments.

Why falling back to synchronous tools.invoke is not an option

As explained above, synchronous tools.invoke is subject to an approximately 90-second public edge request limit. Increasing the caller-side synchronous timeout does not bypass that limit.

PPTX/PDF export may reasonably take longer than 90 seconds in the cloud because it must launch Chrome, render multiple pages, generate the file, and upload the export artifact. Falling back to synchronous tools.invoke would reintroduce the Cloudflare 520 or terminated-request issue we encountered previously. We therefore want to continue using the officially recommended Host Async Job approach instead of reimplementing a polling-based detached background task inside the App.

If useful, we can provide more complete browser error logs, CLI startup logs, a minimal reproduction App, or the relevant invocation code. Thank you for helping us confirm the version compatibility issue in the local development environment.

Hi @HappyLight! :waving_hand:

First off — thank you for another outstanding report. :trophy: Your diagnosis was exactly right: the local runtime shipped the job implementation, but the version of the dispatcher it resolved at install time didn’t register the async tools.* methods — so the iframe got unknown_method even though everything above and below was in place.

What happened :magnifying_glass_tilted_left:

The local runtime pinned by CLI v0.1.45 declared a dependency range that resolved to an older dispatcher — one released before the Async Job API existed. Our internal test setup ran against the in-repo (newer) dispatcher, which masked the mismatch in the published packages. Your package inspection nailed it. :bullseye:

Fixed & published :white_check_mark:

The fix is live now:

  • anna-app-runtime-local 0.2.0a21 — dependency range corrected so the dispatcher with tools.invokeAsync / getJob / listJobs / cancelJob is always resolved
  • @anna-ai/cli 0.1.46 — pins the fixed runtime

To pick it up, just upgrade the CLI and restart:

​bash npm i -g @anna-ai/cli@latest # → v0.1.46 anna-app dev --storage aps ​

The startup banner should show runtime uvx anna-app-runtime-local@0.2.0a21, and invokeAsyncAwait, progress events, cancellation, and listJobs-based reload recovery all work in the local harness — the full job lifecycle you described, no platform deploy needed. :counterclockwise_arrows_button:

Your suggestions :light_bulb:

  • Release guard: we’ve added an automated check to our release pipeline that verifies the CLI ⇄ runtime ⇄ dispatcher versions form a compatible set before anything ships — this class of drift can’t slip through again. :shield:
  • Startup capability check & compatibility matrix: both great ideas — we’re taking them on board for upcoming releases.

Really glad to hear the platform-side migration to async jobs went smoothly — and thank you for the meticulous local investigation. Reports like yours make the whole ecosystem better. :yellow_heart:

Happy building! :rocket:

— The Anna Team