Summary
We use Agent Sessions in an Anna App to generate a PowerPoint presentation one page at a time. A six-page generation task initially worked as expected: multiple Sessions read workspace files, invoked file-writing tools, and modified the target TSX files. After the task had been running for some time, newly created Sessions began consistently returning an abnormal result. The client recorded each call as succeeded, but there was no Assistant output, no tool event, and no error event. The Session History contained only the user message.
The application then compared the target TSX file fingerprints, detected no change, and displayed the message: “The Agent completed without modifying the current page TSX.” Based on the diagnostic logs, this message only describes the final validation failure. The underlying issue appears to be: the Session Run ended without performing any observable Agent work, but the SDK or client treated it as successful.
We would like to clarify the following:
- Why did newly created Sessions end after approximately five seconds with an empty result?
- Did the empty result originate from the Agent service, the Anna Runtime SDK, or the client’s stream-frame parsing?
- Could earlier long-running Agent Run timeouts have left Runs active or continued consuming Session/Worker resources, thereby affecting subsequent Sessions?
- When multiple users use the same App concurrently on the Anna platform, is Session concurrency aggregated by App, account, or tenant? Could this issue have been triggered by high cross-user Session concurrency?
Expected Behavior
After calling session.run({ content: prompt }), we expect at least one of the following outcomes:
- Assistant text is returned, with file tools invoked when needed.
- An explicit error frame is returned, such as a permission error, Session expiration, rate limit, or service error.
- If the stream remains inactive, the client terminates it using the Stream Idle Timeout.
- If the entire Run exceeds the client-side total time limit, the client returns a timeout error.
If the Agent performs no work, we expect the call to fail or return a diagnosable status instead of ending with an empty succeeded result.
Actual Behavior
After the issue began, many Session Interactions were recorded as follows:
{
"status": "succeeded",
"output": "",
"events": [],
"session_history": {
"messages": [
{
"role": "user",
"content": "<full page-authoring prompt>"
}
]
},
"session_retries": 0,
"session_cache_miss_retries": 0,
"stream_idle_retries": 0
}
These calls had the following characteristics:
- The Session was created successfully.
- The Run ended within approximately 4.5 to 13.8 seconds, with an average duration of approximately 5.6 seconds.
- There was no Assistant message.
- There was no
tool_startortool_endevent. - There was no
errorevent. - There was no available Usage information.
- No Session expiration, Session Cache Miss, or stream-idle retry was triggered.
- The SHA-256 hash and file size of the target TSX remained unchanged.
How Sessions Are Invoked
1. Connecting to Anna Runtime
The application connects through the Anna App Runtime SDK:
const runtime = await window.AnnaAppRuntime.connect();
Code location: ppt-app/src/runtime/annaRuntime.ts.
2. Creating an Independent Session for Each Agent Operation
Each page-authoring, render-repair, or visual-inspection operation is an independent logical Agent operation. Under normal conditions, the client creates a new Session for each operation instead of maintaining one long-lived Session across pages:
const session = await runtime.agent.session({ submode: "auto" });
Code location: createSession() in ppt-app/src/agent/agentClient.ts.
After creation, the Session is added to the client’s activeSessions collection. This collection is used to attempt cancellation of active Runs when the user explicitly stops generation.
3. Sending the Full Prompt and Consuming the Asynchronous Stream
The page-authoring prompt is sent as a single content field:
const rawStream = session.run({ content: prompt });
for await (const frame of rawStream) {
// Parse text, tool events, errors, and completion events.
}
Each prompt includes:
- The absolute path of the current page TSX file.
- Paths to the presentation requirements, confirmed outline, and art direction.
- Instructions for reading the Authoring Kit.
- The current page title, page intent, and required content.
- A constraint that only the current page TSX may be modified.
- The expected diagnostic JSON response format.
The summary returned by the Agent does not directly determine whether a page passes. The application reads the target TSX fingerprint before and after the Run and uses a change in its SHA-256 hash or file size as a deterministic gate.
4. Session Lifecycle
A normal call follows this lifecycle:
create session
-> session.run({ content })
-> consume async stream
-> session.history()
-> session.delete()
collectWithSessionRetry() calls session.delete() in a finally block, so it attempts to delete the current Session after normal completion, failure, or a client-side timeout.
Important details:
- Errors from
session.delete()are ignored through.catch(() => undefined). - The logs do not record whether Session deletion succeeded.
- The application calls
cancelActiveRuns()and attemptssession.cancel(runId)only when the user explicitly stops generation. - A normal 600-second Run timeout does not automatically call
session.cancel(runId).
5. Session Retries
The client currently handles the following Session-level conditions:
- Session nearing expiration: delete it and create a new Session.
- Session expired: rebuild the Session at most once.
- Session Cache Miss: delete the Session and create another one using a backoff strategy.
- Stream Idle Timeout: retry at most once.
- No executable tools: return an Agent Infrastructure Error.
- HTTP, authorization, Session-creation, or transport errors: return an infrastructure error.
The empty successful responses in this incident did not trigger any of these branches.
How Session Responses Are Consumed
The client treats the return value of session.run() as an AsyncIterable<AnnaAgentRunFrame>. For each received frame, it attempts to extract the following data.
Text Content
The client checks these fields in order:
frame.textframe.contentframe.deltachoices[].delta.contentchoices[].delta.textchoices[].message.content- Nested
frame.payload
Recognized text is converted into:
{ type: "content", text }
All Content Events are concatenated in order to produce the final output.
Tool Events
The client recognizes the following fields in choices[].delta:
tool_starttool_end
It converts them into internal Activity Events:
{
type: "activity",
tool: "fs_read_file",
path: "/absolute/path/to/file",
size: 1234,
message: "fs_read_file /absolute/path/to/file"
}
Error Events
When frame.event === "error", the client reads:
frame.messageframe.code- Whether the error represents expiration.
- Whether the error represents a Session Cache Miss.
It then throws AgentRunStreamError, which enters either the Session-retry path or the infrastructure-error classification path.
Completion Events
The current implementation recognizes two possible completion signals:
choices[].delta.task_completeis converted into an internalcompleteevent, and itstoken_usageis saved.- If the top-level
frame.event === "complete", the loop ends immediately.
The current handling of a top-level complete frame is:
if (frame.event === "complete") {
break;
}
This check occurs before extractFrameEvents(). Consequently, the top-level complete frame itself does not enter the general parser and is not written to the logs. If that frame also contains text, choices, task_complete, error details, or other Runtime metadata, the current client ignores those fields.
For this reason, we cannot fully rule out a client compatibility issue. The existing logs prove that the client ultimately received an empty output and empty events, but they do not prove that the original top-level complete frame returned by the Runtime contained nothing other than its event field.
Logging Implementation
The logging system has three layers: Interaction, Stream, and Semantic logs.
1. Interaction Log
File: .log/ai-page-agent-interactions.jsonl
At the start of each Agent call, the client records:
operation_idinteraction_id- Operation type, such as
authoringorrender-fix page_idandpage_index- Provider and runtime mode
- Full prompt
- Prompt hash
- Session, Cache Miss, and Stream Idle retry counts
- Start time
At the end of each call, the client records:
status- Start time, end time, and
duration_ms - Aggregated
output output_hash- All parsed internal events
- The return value of
session.history() - Usage
- Error name and error message
- Retry counts
2. Stream Log
File: .log/ai-page-agent-stream.jsonl
The client writes a batch after every 10 internal stream events and forcibly flushes any remaining events at the end of the operation. The log includes:
- Content Events
- Tool start and end events
- Tool names
- File paths and sizes that could be parsed
- Error events
- Client activity messages, such as Cache Miss and Session recovery
- Usage parsed from
task_complete
3. Semantic Log
File: .log/ai-page-agent.jsonl
This log records business-gate results, including:
- Page and operation identifiers
- Parsed Agent summary
- Files the Agent reported reading and modifying
- Path, SHA-256 hash, and file size of the TSX before and after the Run
target_tsx_changed- Business-layer failure reason
4. Large-Payload Sidecars
When a serialized log field exceeds 64 KB, PPT Engine writes it to:
.log/payloads/<channel>/*.json
The primary log retains:
- The sidecar path
- The relative path
- Byte count
- SHA-256 hash
The diagnostic package for this incident contains seven ai-page-agent-interactions sidecar files totaling approximately 700 KB. They preserve some of the larger Session Histories.
Log Completeness
The logs capture information already parsed by the client in considerable detail, but they are not wire-level logs. The following information is not fully recorded:
- Raw frames returned by
session.run() - All fields from the top-level
completeframe - HTTP status, response headers, and underlying response body
- The Runtime SDK’s internal transformation of server events
- Full tool-call arguments
- Full tool return values
appSessionUuid- Run ID
- The Session’s actual
expires_invalue - The complete list of authorized tools and warnings returned during Session creation
- Actual results from
session.delete()andsession.cancel() - Errors raised while reading
session.history()
session.history() is read on a best-effort basis. If it fails, the client returns undefined, and the error is not logged. Log writing is also best-effort; a logging failure does not interrupt generation.
Therefore, the diagnostic package proves that:
The client did not parse any text, tool, or error events, and
session.history()returned only the user message.
However, the diagnostic package cannot independently prove that:
The raw events sent by the Agent service were completely empty.
The raw events may have been empty at the service, or information may have been lost during Runtime SDK transformation or the client’s handling of the top-level complete frame.
Incident Timeline
All timestamps below are in UTC.
1. Initial Page Authoring Worked Normally
From 2026-07-24T06:43:17Z to 06:43:18Z, the application concurrently started authoring Sessions for the first five pages.
These Sessions completed normally between 06:47:37Z and 06:50:54Z. The logs contain:
fs_read_filefs_write_file- Final Assistant JSON
- Non-empty Session History
- Non-empty Usage
- Changed TSX fingerprints
For example, the target TSX for page 1 changed as follows:
before: 134 bytes
after: 7652 bytes
target_tsx_changed: true
This demonstrates that, at the beginning of the task:
- Sessions could be created normally.
- The Agent had access to file tools.
- The absolute paths in the prompts were valid.
- The Shadow Workspace was readable and writable.
- The TSX fingerprint gate worked correctly.
Therefore, the issue does not match a condition in which tool permissions were absent from the beginning or paths were consistently invalid.
2. Two 600-Second Agent Run Timeouts Occurred
The logs contain two explicit total Agent Run timeouts:
Page 1 render-fix
started_at: 2026-07-24T06:48:45.901Z
ended_at: 2026-07-24T06:58:47.822Z
duration: 601921 ms
error: Agent run timed out after 600000ms
Page 5 render-fix
started_at: 2026-07-24T06:54:24.109Z
ended_at: 2026-07-24T07:04:25.764Z
duration: 601655 ms
error: Agent run timed out after 600000ms
Both failed records contain:
{
"session_retries": 0,
"session_cache_miss_retries": 0,
"stream_idle_retries": 0
}
These errors were therefore neither Session Cache Misses nor 180-second Stream Idle Timeouts. They were caused by the client’s 600-second total Run limit.
3. Sessions Began Returning Empty Successful Responses
The first clear empty-success record started at:
started_at: 2026-07-24T07:05:57.285Z
ended_at: 2026-07-24T07:06:02.315Z
duration: 5030 ms
status: succeeded
output: ""
events: []
Equivalent responses subsequently appeared across different pages and operation types. They were not limited to one TSX file or to render-fix. The first authoring call for page 6 also returned only an empty successful response.
Near the end of the task, the application started authoring again for the first five pages. All five calls returned empty successful responses after approximately five seconds. This indicates that the issue had spread to all new calls rather than being caused by the content of one page making the Agent decline to modify it.
4. Aggregate Statistics
ai-page-agent-interactions.jsonl contains:
- 135 completed Interactions
- 2 explicitly marked as
failed - 116 marked as
succeededwhile bothoutputandeventswere empty - Minimum empty-success duration: 4,537 ms
- Maximum empty-success duration: 13,829 ms
- Average empty-success duration: approximately 5,570 ms
All six pages were ultimately recorded as:
{
"agent_failures": 5,
"agent_infrastructure_failures": 0,
"last_error": "Page generation failed: the Agent completed without modifying the current page TSX (...)"
}
Why the Client Records an Empty Result as Successful
1. An Empty Stream Returns Normally
collectRunText() initializes:
let output = "";
const events = [];
If the Runtime quickly sends a top-level complete frame, the loop ends immediately and the function returns:
{
text: "",
events: []
}
The current implementation does not validate:
- Whether at least one meaningful frame was received
- Whether an Assistant message exists
- Whether any tool call occurred
- Whether Usage exists
- Whether Session History contains any Agent/Assistant-side content
2. The Interaction Is Recorded as succeeded
As long as collectRunText() does not throw, runPromptWithSession() calls:
finishInteraction({
status: "succeeded",
output: collected.text,
events: collected.events,
session_history: await session.history()
});
Empty text and empty events are not treated as errors by themselves.
3. Empty Text Is Normalized Into a Renderable Result
The Authoring Result is expected to be JSON. When parsing the empty text fails, the client enters its fallback path and produces:
{
status: "ready_for_render",
changed_files: [],
files_read: [],
authoring_kit_sources_read: [],
summary: "",
needs_render: true,
parsed_json: false
}
An empty Session result therefore does not fail at the Agent Client layer.
4. The File-Fingerprint Gate Eventually Detects No Modification
The page workflow reads the target TSX fingerprint before and after the Run:
const changed =
before.sha256 !== after.sha256 ||
before.size_bytes !== after.size_bytes;
Only when changed === false does the application produce the error stating that the Agent completed without modifying the current page TSX.
This gate successfully prevents the empty result from entering rendering, but the error classification is inaccurate. An empty successful response more closely resembles a Session/Runtime infrastructure failure, yet it is counted as a page-content failure.
5. Empty Successful Responses Trigger Many Retries
The page workflow permits three Local Gate Repair attempts before incrementing agent_failures. The Agent Failure Limit is five.
As a result, when empty successful responses persist, one page may produce approximately 20 ineffective Session calls. With multiple pages running concurrently, the client rapidly creates many new Sessions instead of stopping the entire generation task after detecting the first empty successful response.
This explains why 116 empty-success Interactions were recorded for only six pages.
Client-Side Timeout Mechanisms
Stream Idle Timeout: 180 Seconds
The client applies a 180-second idle limit to each iterator.next() call on the asynchronous stream:
const AGENT_STREAM_IDLE_TIMEOUT_MS = 180_000;
If no new frame arrives for 180 consecutive seconds, the client throws StreamIdleTimeoutError, closes the current Iterator, and retries at most once.
Both long-running failures in this incident have stream_idle_retries: 0, so there is no evidence that either Run had a continuous 180-second period with no frames. A more likely interpretation is that each Run continued producing limited activity but exceeded the 600-second total duration.
Total Agent Run Limit: 600 Seconds
The client sets a fixed 600-second total limit around the entire collectRunText() operation:
const AGENT_RUN_TIMEOUT_MS = 600_000;
await Promise.race([
collectRunText(...),
timeoutPromise
]);
This error is produced by the PPT App client. It is not evidence of an automatic Session timeout:
Agent run timed out after 600000ms
Cancellation Semantics After a Timeout
Promise.race() only rejects the outer wait. It does not automatically cancel the asynchronous Iterator still being awaited inside collectRunText().
The current timeout path proceeds to the outer finally block and calls session.delete(), but it does not call:
session.cancel(runId)
The actual post-timeout behavior therefore depends on the Runtime and service semantics:
- If
session.delete()synchronously terminates all Runs, the underlying work should be cleaned up. - If
session.delete()only removes the client-side Session, or if deletion fails, the original Run may remain active. - If deleting the Session does not terminate the underlying Iterator, the
collectRunText()Promise may continue waiting. - Because deletion errors are ignored and unlogged, the diagnostic package cannot confirm the cleanup result.
Session Lifetime May Also Be 600 Seconds
The client reads session.expires_in or session.expiresIn. If the Runtime provides neither, it uses:
const DEFAULT_AGENT_SESSION_EXPIRES_IN_SECONDS = 600;
Because a Session is created immediately before its Run, a long Run may approach both of the following limits if the actual service-side lifetime is also 600 seconds:
- Service-side Session expiration
- The client’s 600-second total Run limit
The diagnostic logs do not record the actual expires_in value from the incident. We therefore cannot determine whether these limits raced, or whether Session expiration is expected to produce an error, a complete, or simply end the stream.
Confirmed Facts
The following conclusions are directly supported by the diagnostic package:
- Sessions and file tools worked normally at the beginning of the task, and the first five pages had their TSX files modified.
- Before the issue began, two Agent Runs reached the client’s 600-second total limit.
- Stable empty successful responses began approximately 90 seconds after the second long-running timeout ended.
- Empty successful responses occurred across multiple pages and operation types.
- The empty successful responses had no text, tool events, errors, Usage, or retry records.
- The corresponding Session Histories contained only the user prompt, with no Assistant or Tool messages.
- The client currently accepts empty
output/eventsand records the Interaction assucceeded. - The TSX fingerprint gate correctly detected that the files had not changed.
- The user-facing “TSX was not modified” message is a business-gate error, not the original Session-layer error.
- The current logs do not retain raw frames, so they cannot determine whether the empty result originated from the service, the Runtime SDK, or the client’s frame-compatibility logic.
Root-Cause Hypotheses
The following hypotheses are ordered by the current strength of evidence, but each still requires service-side logs or raw frames for confirmation.
Hypothesis 1: The Agent Service or Runtime Returns Only an Empty complete While in an Abnormal State
Supporting evidence:
- Empty calls usually end after approximately five seconds, which does not resemble a network hang.
- Session History contains only the user message.
- There are no tools, Assistant messages, Usage, or errors.
- The issue affects all pages.
- Newly created Sessions continue to exhibit the same behavior.
If this hypothesis is correct, service-side or SDK telemetry for the incident period may show Runs being terminated immediately, scheduling failures, unavailable Workers, throttling, exhausted quotas, or internal errors. However, none of these states were converted into a client-visible error frame.
Hypothesis 2: The Client Exits Too Early on a Top-Level complete Frame and Misses Status Carried by the Same Frame
Supporting evidence:
- The client checks
frame.event === "complete"before parsing the frame. - The raw
completeframe is not logged. - If a newer SDK version changed the frame structure and placed content, errors, or Usage on the top-level
completeframe, the current implementation would ignore those fields.
Counterevidence:
session.history()also contains only the user message. This more strongly suggests that the Agent did not produce an Assistant message at all, rather than only the final frame text being missed by the client.
Hypothesis 3: Runs Were Not Actually Canceled After the 600-Second Timeouts and Caused Abnormal Session/Worker Resource Usage
Supporting evidence:
- Two long-running Runs timed out before empty successful responses became widespread.
- The timeout uses only
Promise.race()and does not callsession.cancel(runId). - The result of
session.delete()is not logged, and deletion failures are ignored. - The first empty successful response began approximately 90 seconds after the second timeout ended.
- Page concurrency and gate retries continued creating more Sessions.
We currently cannot determine:
- Whether
session.delete()always terminates active Runs. - Whether the timed-out Runs continued consuming service-side resources.
- Whether they triggered a Session or Worker quota without an explicit 429 response.
Hypothesis 4: The Session Lifetime Races the Client’s 600-Second Total Run Limit
Supporting evidence:
- The client’s total Run limit is 600 seconds.
- The client’s fallback Session lifetime is also 600 seconds when no lifetime is provided.
- Both failures occurred near the 600-second boundary.
We cannot currently confirm the actual expires_in value for the affected Sessions or the Runtime’s expected stream behavior when a Session expires.
Hypothesis 5: Per-User Concurrency or Aggregate Cross-User Concurrency for the Same App Triggered an Unreported Service-Side Protection Mechanism
Supporting evidence:
- A single generation task initially runs five page Sessions concurrently.
- Normal Sessions last four to nine minutes, perform many file reads, and consume substantial tokens.
- Subsequent render repairs and local-gate retries continue creating new Sessions.
- After the empty-success behavior begins, multiple pages generate additional empty Sessions concurrently.
In addition to concurrency within one client, multiple users may use this App simultaneously. If the platform aggregates concurrency by App, account, tenant, or a shared Worker pool, the actual service-side load may be substantially higher than the five Sessions visible in the local logs. In that case, even moderate per-user concurrency could collectively trigger protections related to Session count, Workers, queues, or model calls.
Counterevidence:
- The logs contain no
HTTP 429, Active Session Limit, or Cache Miss error. - The first five concurrent Sessions all completed page authoring successfully.
The client logs show only the Sessions created by this generation task. They cannot reveal whether other users were using the same App at the same time or the platform-wide active Session count aggregated by App, account, or tenant. This hypothesis must therefore be validated against service-side metrics.
Questions for the Anna Team
- Is
session.run()expected to return only a top-levelcompleteframe with no Message, Tool, Usage, or Error under any normal condition? - What is the complete protocol schema for a top-level
completeframe? Can it containtext,choices, errors, a Stop Reason, or Usage? - Can the Runtime SDK convert any service-side error into an empty
completeframe? - If
session.history()contains only the user message, does that confirm that the Agent never started an Assistant turn? - Does
session.delete()synchronously cancel every active Run belonging to that Session? - If the client calls
session.delete()while a Run is still active, what happens to the underlying Iterator and Worker? - If a client-side Run timeout occurs without a call to
session.cancel(runId), can the Run remain active or continue consuming resources? - Is the default Session lifetime 600 seconds? What event should a long-running Run produce when its Session expires?
- Is a single user running five concurrent
submode: "auto"Sessions subject to concurrency limits or account, App, or Worker quotas? - When multiple users use the same App concurrently, are Session concurrency and Worker resources isolated per user, or aggregated by App, account, tenant, or shared resource pool?
- If aggregate cross-user concurrency for one App reaches a limit, can the platform end Runs with an empty
completeinstead of returning a 429 or another explicit error? - Could you inspect the Agent Session/Worker service logs for the following UTC periods and determine whether other users also had concurrent Sessions for this App?
- Normal operation:
2026-07-24T06:43:17Zto07:05:53Z - First 600-second timeout ended:
2026-07-24T06:58:47Z - Second 600-second timeout ended:
2026-07-24T07:04:25Z - Persistent empty successful responses:
2026-07-24T07:05:57Zto07:14:50Z
- Normal operation:
- Which Session IDs, Run IDs, and raw frame fields should clients record so a future occurrence can be correlated with service-side logs?
Current Assessment
The most certain issue is not that the Agent understood the task but chose not to modify the TSX. Instead:
The Agent Session call ended successfully without producing any observable Agent behavior. The client also lacks infrastructure-error detection for an empty successful response, which amplified the issue through repeated retries and ultimately misreported it at the business layer as an unmodified TSX file.
The timing suggests that the 600-second timeouts may be related to the subsequent empty successful responses, particularly if the timed-out Runs were not fully cleaned up. Another important possibility is that aggregate Session concurrency or Worker load from multiple users of the same App triggered a protection mechanism that did not return an explicit error. The available logs contain neither raw frames, Session UUIDs, Run IDs, nor Session deletion results, and they contain no information about other users or platform-wide aggregate concurrency. They therefore cannot establish either hypothesis as causal.
We plan to add stricter empty-success detection and more complete diagnostic fields to the client. We would still appreciate clarification from the Anna team regarding the source of the empty complete, the lifecycle of a Run after a timeout, cross-user Session concurrency accounting for the same App, and the expected protocol behavior when resources are constrained or a Session enters an abnormal state.