Anna LLM Sampling currently appears to lack a strong structured-output constraint such as response_format, JSON mode, or JSON Schema. Plugins can only ask the model via prompt to return JSON, but in real workflows the model still occasionally returns invalid JSON, such as missing commas, extra prose, markdown fences, incomplete structures, or truncated responses. This causes plugin-side parsing failures, retries, fallback behavior, and sometimes downstream workflow failures.
Our current workaround is to repeatedly instruct the model to output JSON only, run local JSON extraction and lightweight repair, call the LLM again to repair the previous invalid output, and finally fall back if parsing still fails. This helps in some cases, but it is only an application-level workaround. It cannot guarantee stability, adds extra tokens, latency, and sampling calls, and introduces risk that the repair step may rewrite content or invent missing facts.
Request: please consider adding structured-output support at the Anna Sampling protocol level, for example responseFormat: { type: “json_object” }, and eventually JSON Schema support. The host could map this to native JSON mode / schema features when the selected provider supports them. If the selected model does not support structured output, the host should return a clear error or an explicit fallback signal. This would let plugins depend on stable machine-readable output in critical workflows, reduce unnecessary retries and fallbacks, and improve the reliability of Anna agents.
Thank you so much for this detailed and well-written request — the pain points you described (markdown fences, truncated JSON, prose mixed into output, repair-loop retries…) were spot on, and they directly shaped our design.
Good news: this has shipped!
sampling/createMessage now accepts an optional responseFormat field, with two levels:
L1 — { "type": "json_object" } JSON mode, broadly compatible with most models.
L2 — { "type": "json_schema", "json_schema": { "name": "...", "strict": true, "schema": {...} } } strict schema-constrained output, for models that support it.
A few highlights, designed around exactly what you asked for:
No silent downgrades — if the selected model doesn’t support json_schema, the host returns a clear error code (SAMPLING_UNSUPPORTED_RESPONSE_FORMAT, -32010) by default.
Explicit fallback is opt-in — pass onUnsupported: "json_object" (or "text") if you’d rather degrade gracefully than fail. The result’s _meta.responseFormat tells you exactly what happened: { requested, applied, structuredValid, downgraded }.
Fully backward compatible — omit the field and everything behaves exactly as before.
SDK support across the board — the Python, Node.js, and Go Executa SDKs all expose it via create_message(..., response_format=..., on_unsupported=...) (and the JS/Go equivalents).
Test it locally without a real model — the latest anna CLI dev harness emulates structured output, including a flag to simulate an unsupported model so you can exercise your -32010 / downgrade handling before publishing.
Quick taste (Python):
result = await sampling.create_message(
messages=[...],
max_tokens=512,
response_format={
"type": "json_schema",
"json_schema": {
"name": "summary",
"strict": True,
"schema": {...},
},
},
on_unsupported="json_object", # downgrade to JSON mode instead of erroring
)
There’s also an updated sampling-summarizer example in the examples repo showing the full pattern end-to-end.
Hopefully you can now delete that extract-repair-retry pipeline for good. Please give it a spin and let us know how it goes — feedback like yours is exactly what makes the platform better. We’d love to hear what you build with it!