Follow-up from the storage thread, split out as suggested. This is about presentation, not storage — the storage fix is confirmed working.
1. The shape of the response
diagnose_error returns a flat JSON object. The presentation-critical fields:
{
"fingerprint": "sha256:71bf...",
"category": "python.module_not_found_error",
"severity": "medium",
"confidence": 0.9,
"root_cause": "...",
"fix_steps": ["step 1", "step 2", "step 3", "step 4"],
"verify_command": "python -c '...'",
"history": {
"seen_before": true,
"occurrence_count": 6,
"first_seen": "2026-08-20T...",
"known_working_fix": null
},
"headline": "This is the 6th time you have hit this — first seen 20 August.",
"follow_up": "Tell me if this fixes it and it goes in your logbook for next time."
}
headline and follow_up are computed server-side, specifically so the model doesn’t have to compose them. headline is built once in the tool from history — occurrence count, first-seen date, and the last known working fix if any. follow_up is a fixed constant string, non-null only when a follow-up is actually warranted (occurrence count ≥ 3, or the diagnosis came from the sampling fallback rather than the curated KB, or no confirmed resolution exists yet).
The intent: these are not content for the model to summarize or draw from. They are lines to print.
system_prompt_addendum (in manifest.json):
When the user pastes an error… call the error-journal diagnose_error tool with the raw text verbatim. When it returns, print the
headlinefield first and word-for-word — it states whether the user has hit this exact problem before and what fixed it last time, which is the reason this app exists. Then give root_cause, then fix_steps as an ordered list without rewording them, then verify_command, and always include the fingerprint. Iffollow_upis present, print it verbatim as the final sentence — never write your own version, and add nothing when it is null. Never substitute your own diagnosis for the tool’s.
SKILL.md repeats the same instructions in the “Presenting the result” section, with the ordering spelled out as a numbered list (headline → root_cause → fix_steps → verify_command → fingerprint → follow_up).
So the instruction exists in both documented channels, worded plainly, in imperative voice, with “verbatim” used explicitly.
2. Observed vs. expected
Three consecutive pastes of the identical error in a single conversation, all explicitly #Error Journal-mentioned except where noted:
| Paste | Mention | headline/history shown? |
fix_steps verbatim? |
follow_up shown? |
Notes |
|---|---|---|---|---|---|
| 1 | none | No | No | No | Tool not invoked at all — base-agent answer, no fingerprint |
| 2 | #Error Journal |
Yes — “You’ve hit this ModuleNotFoundError before… (fingerprint sha256:71bf…)” | No — reworded, reordered, uv commands added that aren’t in my KB |
No | Tool invoked, history reached the agent, but wording/order not preserved |
| 3 | #Error Journal, same conversation |
No — read as a fresh diagnosis, no history mention at all | No — same rewording pattern | No | Same fingerprint as paste 2, history.seen_before was true both times |
A later, separate fresh-chat test (different session) reproduced the same pattern: first #mention in a new conversation correctly showed the fingerprint but no headline/history/follow_up; a second message in that same conversation fell back to the model’s own conversational memory (“same error as before”) rather than re-invoking the tool or using returned fields at all.
So the failure isn’t consistent in either direction — it isn’t “always drops X” or “only fails on repeat.” Fingerprint presence, history presence, and verbatim wording all vary independently, paste to paste, in the same conversation, with the same tool response shape.
What’s consistent: follow_up was never shown once, across every test I’ve run, including occurrences well past the ≥3 threshold where my tool guarantees it’s non-null.
Note on auto-trigger: the unmentioned paste (row 1) not invoking the tool at all is a related but separate issue I’ve reported before — Skill matching not firing reliably on a literal description match. I’m not asking you to solve that here; flagging it only because it’s the same test sequence and shows the tool-invocation layer and the presentation layer are two independent failure points.
3. What I’d want, if I could design it
In order of how much value each would add relative to how invasive it’d be to build:
A. A render block with an explicit contract. Something like:
{
"diagnose_error": { ... normal response ... },
"_render": {
"verbatim": ["headline", "follow_up"],
"order": ["headline", "root_cause", "fix_steps", "verify_command", "fingerprint", "follow_up"],
"omit_if_null": ["follow_up"]
}
}
A verbatim list is the core ask — fields the host inserts directly into the response rather than passing to the model to paraphrase, the same way a citation or a code block is handled today. If the model is composing prose around them, that’s fine; the listed fields themselves aren’t touched.
B. Failing that, a stronger prompt-level guarantee. If a separate render channel is too big a change, even documenting explicitly what weight system_prompt_addendum carries relative to the model’s own judgment on formatting would help — right now it reads as a strong instruction but behaves as an optional style hint that’s honored inconsistently.
C. Ordering as a softer ask. Guaranteed verbatim content matters more to me than guaranteed ordering — I’d trade ordering for content fidelity if it has to be one or the other.
D. Conditional display (follow_up shown iff non-null) is really a consequence of (A) — if verbatim + omit-when-null both work, this falls out for free.
I don’t have a strong opinion on how this is implemented (special JSON key, a distinct MCP-style content block type, a manifest flag that changes how tool results are fed to the model) — happy to prototype against whichever shape is easiest on your end.
4. How much this matters
Not cosmetic. headline and follow_up are the differentiated part of my app.
The diagnosis itself (root cause, fix steps) is something any capable model can produce from the raw error with reasonable accuracy — it’s useful, but it’s not distinctive. The reason a per-user incident journal exists at all is the sentence “this is the 6th time you’ve hit this, here’s what fixed it last time.” If that line doesn’t reliably show up, the app is indistinguishable from asking the base agent to explain the traceback, which defeats the purpose of building a tool with persistent state in the first place.
Concretely: an app whose core value proposition depends on a specific string being printed, printed maybe 30-40% of the time based on what I’ve observed so far, is not something I can currently promise works to a user, or point to confidently for the Marketplace review.
Happy to run more structured tests if a specific comparison would help narrow this down — e.g. same prompt with the app explicitly vs. implicitly invoked, across several models, or with the _render-style proposal above stubbed in a mock response to see how differently it’s handled. Just let me know what’s useful.