# Structured tool output isn't reliably reproduced by the agent — a proposal for a "render verbatim" channel

**URL:** <https://forum.anna.partners/t/structured-tool-output-isnt-reliably-reproduced-by-the-agent-a-proposal-for-a-render-verbatim-channel/333>\
**Category:** Developers\
**Created:** [September 21, 2026, 6:39am UTC](https://forum.anna.partners/t/structured-tool-output-isnt-reliably-reproduced-by-the-agent-a-proposal-for-a-render-verbatim-channel/333 "2026-09-21T06:39:55Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![sadi](https://yyz1.discourse-cdn.com/flex033/user_avatar/forum.anna.partners/sadi/32/241_2.png) [@sadi](https://forum.anna.partners/u/sadi)\
**Post date:** [September 21, 2026, 6:39am UTC](https://forum.anna.partners/t/structured-tool-output-isnt-reliably-reproduced-by-the-agent-a-proposal-for-a-render-verbatim-channel/333/1 "2026-09-21T06:39:55Z")

</div>

Follow-up from the storage thread, split out as suggested. This is about presentation, not storage — the storage fix is confirmed working.

## 1. The shape of the response

`diagnose_error` returns a flat JSON object. The presentation-critical fields:

```auto
{
  "fingerprint": "sha256:71bf...",
  "category": "python.module_not_found_error",
  "severity": "medium",
  "confidence": 0.9,
  "root_cause": "...",
  "fix_steps": ["step 1", "step 2", "step 3", "step 4"],
  "verify_command": "python -c '...'",
  "history": {
    "seen_before": true,
    "occurrence_count": 6,
    "first_seen": "2026-08-20T...",
    "known_working_fix": null
  },
  "headline": "This is the 6th time you have hit this — first seen 20 August.",
  "follow_up": "Tell me if this fixes it and it goes in your logbook for next time."
}

```

`headline` and `follow_up` are computed **server-side, specifically so the model doesn’t have to compose them.** `headline` is built once in the tool from `history` — occurrence count, first-seen date, and the last known working fix if any. `follow_up` is a fixed constant string, non-null only when a follow-up is actually warranted (occurrence count ≥ 3, or the diagnosis came from the sampling fallback rather than the curated KB, or no confirmed resolution exists yet).

The intent: these are not content for the model to summarize or draw from. They are lines to print.

**`system_prompt_addendum`** (in `manifest.json`):

> When the user pastes an error… call the error-journal diagnose\_error tool with the raw text verbatim. When it returns, print the `headline` field first and word-for-word — it states whether the user has hit this exact problem before and what fixed it last time, which is the reason this app exists. Then give root\_cause, then fix\_steps as an ordered list without rewording them, then verify\_command, and always include the fingerprint. If `follow_up` is present, print it verbatim as the final sentence — never write your own version, and add nothing when it is null. Never substitute your own diagnosis for the tool’s.

**`SKILL.md`** repeats the same instructions in the “Presenting the result” section, with the ordering spelled out as a numbered list (headline → root\_cause → fix\_steps → verify\_command → fingerprint → follow\_up).

So the instruction exists in both documented channels, worded plainly, in imperative voice, with “verbatim” used explicitly.

## 2. Observed vs. expected

Three consecutive pastes of the identical error in a single conversation, all explicitly `#Error Journal`-mentioned except where noted:

| Paste | Mention | `headline`/history shown? | `fix_steps` verbatim? | `follow_up` shown? | Notes |
| --- | --- | --- | --- | --- | --- |
| 1 | none | No | No | No | Tool not invoked at all — base-agent answer, no fingerprint |
| 2 | `#Error Journal` | **Yes** — “You’ve hit this ModuleNotFoundError before… (fingerprint sha256:71bf…)” | No — reworded, reordered, `uv` commands added that aren’t in my KB | No | Tool invoked, `history` reached the agent, but wording/order not preserved |
| 3 | `#Error Journal`, same conversation | **No** — read as a fresh diagnosis, no history mention at all | No — same rewording pattern | No | Same fingerprint as paste 2, `history.seen_before` was true both times |

A later, separate fresh-chat test (different session) reproduced the same pattern: first `#mention` in a new conversation correctly showed the fingerprint but no `headline`/history/`follow_up`; a second message in that same conversation fell back to the model’s own conversational memory (“same error as before”) rather than re-invoking the tool or using returned fields at all.

So the failure isn’t consistent in either direction — it isn’t “always drops X” or “only fails on repeat.” Fingerprint presence, history presence, and verbatim wording all vary independently, paste to paste, in the same conversation, with the same tool response shape.

What’s consistent: **`follow_up` was never shown once** , across every test I’ve run, including occurrences well past the ≥3 threshold where my tool guarantees it’s non-null.

Note on auto-trigger: the unmentioned paste (row 1) not invoking the tool at all is a related but separate issue I’ve reported before — Skill matching not firing reliably on a literal description match. I’m not asking you to solve that here; flagging it only because it’s the same test sequence and shows the tool-invocation layer and the presentation layer are two independent failure points.

## 3. What I’d want, if I could design it

In order of how much value each would add relative to how invasive it’d be to build:

**A. A `render` block with an explicit contract.** Something like:

```auto
{
  "diagnose_error": { ... normal response ... },
  "_render": {
    "verbatim": ["headline", "follow_up"],
    "order": ["headline", "root_cause", "fix_steps", "verify_command", "fingerprint", "follow_up"],
    "omit_if_null": ["follow_up"]
  }
}

```

A `verbatim` list is the core ask — fields the host **inserts directly into the response** rather than passing to the model to paraphrase, the same way a citation or a code block is handled today. If the model is composing prose around them, that’s fine; the listed fields themselves aren’t touched.

**B. Failing that, a stronger prompt-level guarantee.** If a separate render channel is too big a change, even documenting explicitly what weight `system_prompt_addendum` carries relative to the model’s own judgment on formatting would help — right now it reads as a strong instruction but behaves as an optional style hint that’s honored inconsistently.

**C. Ordering as a softer ask.** Guaranteed verbatim content matters more to me than guaranteed ordering — I’d trade ordering for content fidelity if it has to be one or the other.

**D. Conditional display (`follow_up` shown iff non-null)** is really a consequence of (A) — if verbatim + omit-when-null both work, this falls out for free.

I don’t have a strong opinion on _how_ this is implemented (special JSON key, a distinct MCP-style content block type, a manifest flag that changes how tool results are fed to the model) — happy to prototype against whichever shape is easiest on your end.

## 4. How much this matters

Not cosmetic. `headline` and `follow_up` **are** the differentiated part of my app.

The diagnosis itself (root cause, fix steps) is something any capable model can produce from the raw error with reasonable accuracy — it’s useful, but it’s not distinctive. The reason a per-user incident journal exists at all is the sentence _“this is the 6th time you’ve hit this, here’s what fixed it last time.”_ If that line doesn’t reliably show up, the app is indistinguishable from asking the base agent to explain the traceback, which defeats the purpose of building a tool with persistent state in the first place.

Concretely: an app whose core value proposition depends on a specific string being printed, printed maybe 30-40% of the time based on what I’ve observed so far, is not something I can currently promise works to a user, or point to confidently for the Marketplace review.

* * *

Happy to run more structured tests if a specific comparison would help narrow this down — e.g. same prompt with the app explicitly vs. implicitly invoked, across several models, or with the `_render`-style proposal above stubbed in a mock response to see how differently it’s handled. Just let me know what’s useful.

---

<div class="post-metadata">

**Author:** ![hunter](https://yyz1.discourse-cdn.com/flex033/user_avatar/forum.anna.partners/hunter/32/8_2.png) [@hunter](https://forum.anna.partners/u/hunter)\
**Post date:** [September 24, 2026, 5:39am UTC](https://forum.anna.partners/t/structured-tool-output-isnt-reliably-reproduced-by-the-agent-a-proposal-for-a-render-verbatim-channel/333/2 "2026-09-24T05:39:54Z")

</div>

Hi @sadi 👋

First off — **thank you**. This is genuinely one of the best-written reports we’ve received: a clear problem statement, a reproducible test matrix, and a concrete API proposal. Posts like this make the platform better for every developer, and this one did exactly that. 💛

**You were right on all counts.** We verified your analysis against the platform internals: tool results are re-synthesized by the model, and no prompt wording — `system_prompt_addendum`, SKILL.md, however imperative — can turn a probabilistic instruction into a guarantee. Presentation-critical text needs a channel that doesn’t pass through the model at all.

So we built one. 🎉

## `_display` — verbatim display blocks

Shipping **this Friday** in **1.1.0-beta.178**. Your proposal A, almost verbatim (pun intended 😄):

```json
{
  "fingerprint": "sha256:71bf...",
  "fix_steps": ["step 1", "step 2"],

  "_display": {
    "blocks": [
      { "type": "markdown", "text": "📒 **This is the 6th time you have hit this** — first seen 20 August." },
      { "type": "markdown", "text": "_Tell me if this fixes it and it goes in your logbook for next time._" }
    ]
  }
}

```

Declare blocks under the reserved top-level `_display` key of your tool result, and:

- 🖥 **The chat UI renders them byte-for-byte** — in array order, anchored at the tool call, as attributed cards. **100% presentation rate** , fully decoupled from model behavior. Works identically whether the app was `#`-pinned or the tool fired on a bare paste.
- 🤖 **The model can’t paraphrase them** — it receives a “these blocks were already shown verbatim, don’t restate” note (with short previews so its surrounding prose stays coherent).
- 🔁 **History replay renders them identically** — blocks survive refresh and device switches.
- ✂ **Omit-if-null is just code** : don’t emit a block and it doesn’t render. Your `follow_up` logic needs zero prompt gymnastics.

Limits: 1–8 blocks, ≤4,000 chars each, ≤16,000 total, markdown only (strictly sanitized). Anything over the limits is **rejected whole, never truncated** — you’ll find a machine-readable `_display_rejected` reason in the tool trace.

## Migrating Error Journal

Move `headline` and `follow_up` (and `fix_steps`, if you want those byte-exact too) into `_display.blocks`, keep the structured fields for the model to reason about, and **delete the “print verbatim” instructions** from your addendum — they’re no longer needed.

📚 Docs: [Display Blocks (verbatim output)](https://anna.partners/developers/tools/executa-display-blocks)  
🧪 Runnable example (a mini error journal inspired by yours — occurrence counting, conditional follow-up and all): [`examples/python/display-blocks-demo`](https://github.com/whtcjdtc2007/anna-executa-examples/tree/main/examples/python/display-blocks-demo)

Available also on **this Friday**.

We’d love to hear how it works for your app once it lands — and if you run your test matrix again against beta.178, we’re all ears. 🙌
