Llm.complete interface may include unannounced max token limitations

I am now developing Anna Deck. While I was trying to let the model gemini-3.5-flash(x18) generate a long html file, the output was cut down. Here’s a detailed discription.

Scenario

I attempted to generate the candidate page in one llm.complete call.

The request includes:

  • a system prompt requiring only a complete HTML document;
  • the user’s page-refinement request;
  • the current page HTML;
  • a screenshot of the current page.

The expected response is a document starting with <!doctype html> and ending with </html>.

For the candidate-generation call, I requested:

maxTokens: 12000
modelPreferences: {
  hints: [{ name: "google/gemini-3.5-flash" }]
}

The visible response began in the middle of HTML and ended in the middle of HTML. It did not include either the opening document structure or </html>.

The returned usage was:

{
  "inputTokens": 4976,
  "outputTokens": 4092,
  "totalTokens": 9068
}

stopReason was still reported as endTurn.

Follow-up Tests

To ensure if there really exists a unvisiable limmition, I ask the model to output a long html with different input maxToken. This html contains over 500 paragraphs. The following table shows the final output.

Requested maxTokens Actual outputTokens Last complete marker Complete </html>
4096 4092 OUTPUT-0075 No
8192 4092 OUTPUT-0085 No
12000 4092 OUTPUT-0078 No
16000 4092 OUTPUT-0076 No

Expected Behavior

I suggest that the platform should take at least one of the following measures to prevent users from encountering this situation again.

  1. Honor maxTokens values above 4,096 when supported by the selected model;
  2. Reject unsupported values with a clear error;
  3. Return an explicit effective output limit and a truncation-specific finish reason.

Returning endTurn for an incomplete HTML document makes it difficult for developers to distinguish successful completion from output cutting down.

Environment

Anna version: Local Anna CLI package: @anna-ai/cli 0.1.47
App source commit: 05071e20d2dcf32a08fdbbbe50f8f3c5c4aa3af0
API: llm.complete (non-streaming)
Runtime:

  • Browser-hosted Anna App runtime obtained through window.AnnaAppRuntime.connect()
  • Local app served through anna-app dev
  • Model preference: google/gemini-3.5-flash
  • Actual returned model: google/gemini-3.5-flash
  • Provider metadata returned by the raw response: openrouter

The token counts and stopReason shown above are taken directly from the raw response, not from a UI-rendered or truncated view.

Hi @Yinghuo_Mars! :waving_hand:

First off — thank you for this outstanding report. :folded_hands: The maxTokens sweep table (4096/8192/12000/16000 → always ~4092) made the diagnosis instant, and your write-up was spot on: there was an unannounced 4096 output cap, and stopReason: "endTurn" on truncated output was genuinely misleading. We’ve fixed both — and we’re happy to say all three of your suggested measures made it in. :sparkles:

:rocket: Shipping this Friday (Sep 25)

Platform 1.1.0-beta.179

  1. :white_check_mark: maxTokens above 4096 is now honored. The effective cap is resolved per call as min(maxTokens ?? 8192, model output cap, path ceiling) — the path ceiling is 8192 for llm.complete (sync transport budget) and 65536 for llm.stream. The default when maxTokens is omitted doubles to 8192.
  2. :white_check_mark: The effective limit is explicit. Every response now carries _meta.maxTokens: { requested, effective, limitedBy: "request" | "model" | "transport" | "default" } — no more silent clamping, ever.
  3. :white_check_mark: Truncation-specific stop reason. stopReason now maps the provider’s real finish reason: "endTurn" | "maxTokens" | "contentFilter" | "toolUse". Your incomplete-HTML case will report "maxTokens", so you can detect it and continue programmatically. :bullseye:

Bonus: if your balance can’t cover a large request, you now get a clean APP_QUOTA_EXCEEDED (429) before the call runs, with requiredCu / remainingCu / requestedMaxTokens in data — perfect for an automatic lower-and-retry.

SDK @anna-ai/app-runtime 0.17.0 :puzzle_piece:

For your exact Anna Deck scenario (maxTokens: 12000): with the new SDK, llm.complete requests above 8192 are transparently delivered over the streaming channel and resolve with the identical single-result shape — your </html> will arrive intact, no code changes needed. :tada:

CLI @anna-ai/cli 0.1.55 — full anna-app dev parity, including mock mode.

:memo: Heads-up

  • No breaking changes — existing apps keep working unchanged. :handshake:
  • Calls that were previously silently cut at 4096 can now generate more tokens, so quota spend may rise accordingly. Set maxTokens explicitly for a tighter budget.
  • llm.stream remains the recommended first-class path for long outputs.

The /developers docs are being refreshed alongside the release. Reports like this one make the platform better for everyone — please keep them coming, and good luck with Anna Deck! :green_heart:

— The Anna App Platform Team