Issue Summary
While running a PPT page generation task in an ANNA App, the page source code was generated successfully, and the backend also produced the page HTML and PNG screenshot. However, the frontend never received a successful rendering result.
Each rendering call failed after running for approximately 133 seconds. The frontend received the following error through ANNA Runtime’s tools.invoke:
Unexpected token '<', "<!DOCTYPE "... is not valid JSON
The task retried three times, and every attempt produced exactly the same error. Because the application currently classifies this as a page rendering error, it repeatedly asked the page Agent to modify TSX code that could already be rendered successfully. Modifying the page code did not resolve the issue.
Based on the rendering artifacts in the diagnostics bundle, the consistent failure duration, and a Cloudflare HTTP 520 HTML response recorded during the same task, I suspect the actual failure is as follows: Chrome or a Puppeteer operation in the cloud Agent hangs at some point after rendering has completed. As a result, the tool invocation cannot return in time and is eventually terminated by an upstream Cloudflare or gateway timeout.
At this point, I cannot confirm whether the Chrome hang is caused by the cloud Agent’s CPU, memory, process limit, or another resource constraint. I would appreciate ANNA’s assistance in reviewing the cloud runtime environment and server-side logs.
Environment and Scenario
- Application type: ANNA App
- Backend capability: Executa bundled with the App
- Executa:
ppt-engine - Tool method:
app_render_workspace_page_preview - Execution environment: Cloud Agent provided by ANNA
- Task: Generate one 1280 × 720 PPT page
- Workspace ID:
ppt-20260730-072715 - Generation Run ID:
9a5b3e5f-32c7-4336-89a6-525d6a611b53 - Approximate incident time: July 30, 2026, 07:32–07:42 UTC
Expected Behavior
- The frontend calls
tools.invokethrough ANNA Runtime. ppt-enginestarts managed Chrome instances and generates static HTML and a PNG screenshot.- The Executa returns a JSON Tool Result containing fields such as
html_pathandscreenshot_path. - The frontend receives the result, updates the page status to
accepted, and proceeds with final deck rendering.
Actual Behavior
-
The page Agent successfully generates the TSX source.
-
ppt-enginestarts rendering the page preview. -
Valid HTML and PNG files appear in the shadow Workspace, and the PNG can be opened successfully.
-
The frontend does not receive a successful JSON Tool Result.
-
After approximately 133 seconds,
tools.invokethrows a JSON parsing error:Unexpected token '<', "<!DOCTYPE "... is not valid JSON -
The system records the attempt as a
renderfailure and starts arender-fixAgent. -
The page TSX is modified three times, but every attempt still fails with the same error after approximately 133 seconds.
-
The user eventually abandons the task, and the Generation Run status becomes
abandoned.
The duration of all three failures is highly consistent:
| Stage | Previous page Agent completion | Next render-fix start |
Approximate wait |
|---|---|---|---|
| First render | 07:32:33 UTC | 07:34:46 UTC | ~133 seconds |
| Second render | 07:35:44 UTC | 07:37:58 UTC | ~134 seconds |
| Third render | 07:38:57 UTC | 07:41:10 UTC | ~133 seconds |
This fixed timing pattern is more consistent with an upstream timeout than with a TSX compilation error or an ordinary page runtime exception.
Application Implementation
Frontend Invocation
The frontend invokes the Executa through ANNA Runtime:
runtime.tools.invoke({
tool_id: pptEngineToolId,
method: "app_render_workspace_page_preview",
args: {
workspace_dir,
page_id,
},
timeoutMs: 600_000,
});
The application configures a 600-second timeout for this long-running tool invocation. Therefore, the failure after approximately 133 seconds is not triggered by the application’s own 600-second timer.
ANNA Runtime receives an RPC Result or RPC Error from the Host through postMessage. The current failure is returned to the application as an RPC Error. The application itself does not directly call JSON.parse() on the page HTML.
Executa Rendering Flow
The primary steps performed by app_render_workspace_page_preview in ppt-engine are:
- Read the Workspace’s
manifest.json. - Run a pre-render TypeScript check on the target TSX file.
- Build the single-page HTML.
- Start a managed Chrome instance to convert the runtime HTML into static HTML.
- Close that Chrome page and browser instance.
- Start another managed Chrome instance, load the static HTML, and generate a PNG screenshot.
- Close the second Chrome page and browser instance.
- Return a JSON result containing the HTML and screenshot paths.
The current cleanup logic waits directly for Puppeteer:
await page.close?.().catch(() => undefined);
await browser.close().catch(() => undefined);
Although exceptions thrown during cleanup are ignored, no timeout is applied. If page.close() or browser.close() never resolves, the entire tool invocation will remain pending.
In addition to the cleanup stage, launch(), newPage(), page navigation, screenshot capture, or Chrome DevTools Protocol communication could also be affected by cloud resource constraints or the container runtime environment. The current diagnostics bundle does not record the start and completion times of these individual sub-stages, so I cannot determine with certainty which operation is hanging.
Diagnostic Evidence
1. Page Rendering Artifacts Exist
The shadow Workspace contains:
output/page-preview-html/*.html
output/screenshots/*.png
The PNG can be opened successfully, and both the page content and images are rendered correctly. This indicates that the page TSX, image files, Chrome page loading, and screenshot pipeline successfully reached at least the artifact-writing stage.
2. Modifying the Page Code Does Not Change the Error
The page Agent made the following changes in sequence:
- Added an
altattribute to an<img>element. - Replaced a text structure containing
<br />with<p>elements. - Rewrote the page as a simpler React component.
After all three changes, the error remained:
Unexpected token '<', "<!DOCTYPE "... is not valid JSON
Therefore, the issue does not appear to be caused by a particular HTML tag, page layout, or TSX syntax error.
3. A Cloudflare HTTP 520 HTML Response Was Explicitly Recorded During the Same Task
The Storage Transfer Log for the same task records a Host Upload negotiation failure:
upload HTTP 520: {'detail': '<!DOCTYPE html>
<!--[if lt IE 7]> <html class="no-js ie6 oldie" lang="en-US"> ...
This confirms that, during the same cloud execution, an upstream ANNA-related endpoint returned a Cloudflare-style HTML error page instead of the expected JSON response.
The raw HTTP response for the rendering tool invocation was not included in the diagnostics bundle. Therefore, I cannot state conclusively that the rendering failure was caused by the same HTTP 520 response. However, the error shape is consistent, and the rendering failures occur at a highly repeatable interval, suggesting the same category of gateway or upstream timeout failure.
4. Page Progress Remains at rendering
The diagnostics bundle records the following page progress:
{
"status": "rendering",
"render_attempts": 3,
"last_html_path": "",
"last_screenshot_path": ""
}
This means the application started another rendering attempt but never received the Tool Result, so it had no opportunity to persist the paths of the artifacts that had already been generated.
Current Assessment
The most likely failure sequence is:
ppt-engine starts Chrome
→ HTML/PNG artifacts are generated
→ Chrome, Puppeteer, or process cleanup hangs
→ The Executa Tool Invoke remains pending
→ An upstream gateway terminates the request after approximately 120 seconds and returns a Cloudflare HTML error page
→ The Host attempts to parse the HTML as JSON
→ The frontend receives Unexpected token '<'
→ The application misclassifies an infrastructure failure as a page-code error and repeatedly runs render-fix
I suspect that the cloud Agent may have insufficient resources, or that Chrome cannot exit reliably under the current container constraints. Possible causes include:
- Insufficient memory, causing a Chrome renderer process to be killed by the OOM killer;
- A low CPU quota or severe CPU throttling;
- Exhaustion of the PID or file descriptor limit;
- Failure to exit or reap Chrome child processes correctly;
- Slow container disk I/O;
- Unresponsive IPC between Puppeteer and Chrome during shutdown;
- Timed-out invocations continuing to run and leaving Chrome processes behind, which then accumulate across retries.
Insufficient resources are still a hypothesis. Confirmation requires the cloud Agent’s resource configuration, container metrics, and server-side logs from ANNA.
Requested Assistance from ANNA
I would appreciate ANNA’s help with the following:
-
What is the current resource configuration of the cloud Agent?
- Number of CPU cores or CPU quota;
- Memory limit;
/dev/shmsize;- PID limit;
- File descriptor limit;
- Temporary disk capacity and I/O limits.
-
Can the cloud Agent configuration for this App/Executa be temporarily or permanently increased?
- I would like to reproduce the same task with more CPU and memory;
- If the platform supports multiple resource tiers, please provide the recommended minimum configuration for Chrome/Puppeteer workloads.
-
Please review the cloud logs and resource metrics for this task.
- Workspace ID:
ppt-20260730-072715; - Generation Run ID:
9a5b3e5f-32c7-4336-89a6-525d6a611b53; - Primary time range: July 30, 2026, 07:32–07:42 UTC;
- Check for OOM events, CPU throttling, Chrome crashes, orphaned processes, or Tool backend timeouts;
- Determine which upstream request Cloudflare or the gateway terminated at approximately 120 seconds.
- Workspace ID:
-
Please confirm whether the backend task continues running after a Tool Invoke timeout.
- If the frontend or gateway has already returned an error, do the Executa and Chrome processes continue running in the background?
- Could this leave Chrome child processes behind and affect subsequent retries?
-
Please preserve clearer error details at the Host layer.
- Check the HTTP status and
Content-Typebefore parsing a response; - When an HTML error page is received, return a structured HTTP 520/502/504 error;
- Do not return only
Unexpected token '<'to the App; - Consider logging a redacted response-body summary, request correlation ID, and Cloudflare Ray ID, if available.
- Check the HTTP status and
Planned Mitigations on Our Side
We plan to add the following safeguards to ppt-engine, but we still need ANNA to confirm whether the cloud environment has a resource or process-management issue:
- Record separate durations for Chrome startup, page loading, screenshot capture,
page.close(), andbrowser.close(); - Add a timeout to browser cleanup and forcibly terminate the Chrome process if cleanup times out;
- Classify Tool Invoke, gateway, and JSON response-parsing failures as infrastructure errors instead of asking the page Agent to modify TSX;
- Apply limited retries to retryable Tool/Transport errors;
- Record each browser process’s PID, exit code, and failure stage in the diagnostics bundle.
I would appreciate details of the cloud Agent’s actual resource specification and assistance determining whether this approximately 120-second interruption was caused by insufficient cloud resources, a Chrome process failing to exit, or a platform-level upstream timeout. If the Executa can temporarily be assigned a higher cloud resource tier, we can rerun the same minimal task to verify the behavior.