I am now developing Anna Deck. While I was trying to let the model gemini-3.5-flash(x18) generate a long html file, the output was cut down. Here’s a detailed discription.
Scenario
I attempted to generate the candidate page in one llm.complete call.
The request includes:
- a system prompt requiring only a complete HTML document;
- the user’s page-refinement request;
- the current page HTML;
- a screenshot of the current page.
The expected response is a document starting with <!doctype html> and ending with </html>.
For the candidate-generation call, I requested:
maxTokens: 12000
modelPreferences: {
hints: [{ name: "google/gemini-3.5-flash" }]
}
The visible response began in the middle of HTML and ended in the middle of HTML. It did not include either the opening document structure or </html>.
The returned usage was:
{
"inputTokens": 4976,
"outputTokens": 4092,
"totalTokens": 9068
}
stopReason was still reported as endTurn.
Follow-up Tests
To ensure if there really exists a unvisiable limmition, I ask the model to output a long html with different input maxToken. This html contains over 500 paragraphs. The following table shows the final output.
Requested maxTokens |
Actual outputTokens |
Last complete marker | Complete </html> |
|---|---|---|---|
| 4096 | 4092 | OUTPUT-0075 |
No |
| 8192 | 4092 | OUTPUT-0085 |
No |
| 12000 | 4092 | OUTPUT-0078 |
No |
| 16000 | 4092 | OUTPUT-0076 |
No |
Expected Behavior
I suggest that the platform should take at least one of the following measures to prevent users from encountering this situation again.
- Honor
maxTokensvalues above 4,096 when supported by the selected model; - Reject unsupported values with a clear error;
- Return an explicit effective output limit and a truncation-specific finish reason.
Returning endTurn for an incomplete HTML document makes it difficult for developers to distinguish successful completion from output cutting down.
Environment
Anna version: Local Anna CLI package: @anna-ai/cli 0.1.47
App source commit: 05071e20d2dcf32a08fdbbbe50f8f3c5c4aa3af0
API: llm.complete (non-streaming)
Runtime:
- Browser-hosted Anna App runtime obtained through
window.AnnaAppRuntime.connect() - Local app served through
anna-app dev - Model preference:
google/gemini-3.5-flash - Actual returned model:
google/gemini-3.5-flash - Provider metadata returned by the raw response:
openrouter
The token counts and stopReason shown above are taken directly from the raw response, not from a UI-rendered or truncated view.