Json_schema sampling call fails with empty {success:false,error:""} on Cloud Agent despite onUnsupported=json_object set

Hi team,

Related to the two threads on structured sampling – I’m hitting a third, more opaque failure mode for json_schema-constrained sampling specifically on Cloud Agents, writing it up in case it’s connected.

executa_id: 1091

tool_id: tool-calderbuild-wefinance-investment-recommendations-r2q5jdey

ExecutaVersion: 378 (v0.1.5)

Symptom: every invoke of this tool’s generate_recommendations method fails with:

{“success”: false, “error”: “”, “command_id”: “”}

No jsonrpc/id fields, no error text – doesn’t match our own tool’s JSON-RPC response shape, so it looks like it comes from a layer above our code.

What I’ve ruled out:

  1. Cloud Agent binary caching – reproduced on a brand-new Cloud Agent’s first-ever invocation of this tool.
    1. A broken/corrupted binary – extracted our shipped darwin-arm64 binary and drove it directly over raw JSON-RPC stdio. It responds correctly to describe/initialize/invoke and emits a well-formed sampling/createMessage request.
      1. Sampling broken platform-wide – our sibling tool (Advisor Chat, same llm.sample capability, plain-text response) succeeds immediately on the same Cloud Agent in the same session.
        1. Missing permission grant – the Sampling toggle is enabled for this Executa in Installed Apps > Permissions.
        2. The one real difference between the tool that works and the one that doesn’t: this tool requests a strict json_schema-constrained completion, and correctly sets the documented onUnsupported: “json_object” field alongside responseFormat, so the host should downgrade instead of erroring when the model can’t honor strict schema mode. We’re seeing a hard failure instead of a downgrade.
      2. Example failed call:
    2. Request:
    3. {“tool”: “generate_recommendations”, “arguments”: {“risk_profile”: “moderate”, “investment_goal”: “retirement”, “monthly_income”: 8000, “investment_horizon”: “10-15 years”, “transactions”: [4 sample transactions]}}
  2. Response:
  3. {“success”: false, “error”: “”, “command_id”: “a62c9c02-fe86-4b98-865b-3d1f091b3fec”}

Two more command_ids from earlier failed attempts, same shape: a6d70289-c2fa-478f-b14c-505ce59ec38a, 5253b1c0-287d-4d87-a2df-5f29d38969eb

Given the “Sampling token missing” thread’s root cause involved a stale platform-side capability/manifest record after certain publish flows, and my own publish history for this executa was slightly unusual (had to manually recreate a missing .anna/executa.json identity cache before republishing), I’m wondering if there’s a similar stale-registration angle here specific to json_schema-mode sampling.

Would appreciate any pointers on where to look next, or confirmation of whether onUnsupported-based downgrade is fully wired up for Cloud Agents today.

Thanks,

Calder

Follow-up: two more hypotheses eliminated, without spending any Energy.

I inspected grants rather than making new calls (anna-app apps grants wefinance --json, CLI v0.1.49), so none of this required a live invocation.

1. It is not a grant difference between the bundled tools.

All three executas in this app come back identical – sampling_grant enabled, llm_grant complete: true, max_calls_per_day 1000, max_tokens_per_call 4096, and no missing scopes on any of them:

Bill Scanner    (executa_id 1089)   sampling enabled, complete
Advisor Chat    (executa_id 1090)   sampling enabled, complete   <- works
Investment Rec  (executa_id 1091)   sampling enabled, complete   <- fails

So the failing tool is not under-granted relative to the sibling that works.

2. It is not the per-call token ceiling.

The 4096 cap is not being hit. ask_advisor requests maxTokens 2000; generate_recommendations requests 3000. Both are under the ceiling, and neither call site overrides its default.

What that leaves

Same reverse-RPC path, same Cloud Agent, same session, same grants, same token ceiling. The only remaining structural difference between the sibling call that succeeds and the one that fails is the response format block:

"responseFormat": {"type": "json_schema", "json_schema": {...strict: true...}},
"onUnsupported": "json_object"

ask_advisor sends no responseFormat at all and succeeds (duration_ms 18682). generate_recommendations sends the above and comes back as {"success": false, "error": "", "command_id": "<uuid>"} with no jsonrpc/id fields.

I also verified independently that the model layer itself is not the problem: a plain OpenAI-compatible /chat/completions call to my own provider with response_format: {"type": "json_schema", strict: true} returns valid schema-conforming JSON. So model capability is ruled out too – which is why I think this sits in the sampling layer’s handling of responseFormat rather than downstream of it.

Happy to run any instrumented build or specific payload you want tested. My account is currently at 1,000/1,000 Energy so I am paused on live runs, but I will re-test the moment that clears and post the result here either way.

Follow-up: quota top-up landed, and I re-ran the live comparison with fresh Energy.

Same test payload as before, no BYOK involved, both calls made back-to-back in the same Cloud Agent session:

generate_recommendations -> {"success": false, "error": "", "command_id": "471b68d6-b308-4ccb-8e2b-90af5b5f3815"}

ask_advisor (sibling tool, run immediately after) -> {"success": true, "data": {"success": true, "data": {"advice": "..."}, "duration_ms": 12852}, "command_id": "a346c7aa-b708-4c18-9023-2230ffabf594"}

Identical failure signature to before the top-up: empty error string, no jsonrpc/id fields. This rules out account-level Energy exhaustion as a factor — both calls ran under the exact same live conditions (same session, same Cloud Agent, same grants), and only the one sending responseFormat: {"type": "json_schema", ...} fails. Still points at the sampling layer’s handling of that field rather than anything account- or quota-related.

Happy to test any specific payload or instrumented build if that’s useful for narrowing it down further.

Hi @Calder :waving_hand:

First off — thank you for an absolutely stellar bug report. The systematic
elimination across grants, quotas, siblings and payloads saved us a lot of
time and pointed us straight at the right layer. :folded_hands:

TL;DR

You found a real platform bug. It’s fixed in v1.1.0-beta.141 :tada:
No changes needed on your side — your responseFormat +
onUnsupported: "json_object" request was correct all along.

What was happening :magnifying_glass_tilted_left:

Two things compounded, both on our side of the fence:

  1. A capability flag was wrong. One of the models in our sampling
    routing pool was marked as supporting strict json_schema, but the
    upstream provider actually coerces it to plain JSON mode internally.
    Because the flag said “supported”, the host skipped the
    onUnsupported downgrade you asked for
    and forwarded the request
    as-is.
  2. A provider quirk. That provider hard-requires the literal word
    “json” to appear somewhere in the messages when JSON mode is active —
    otherwise it rejects the call instantly. Your prompt (reasonably!)
    relied on responseFormat instead of prompt wording, so every call
    failed upstream before generation even started.

So: ask_advisor (no responseFormat) sailed through, while
generate_recommendations hit the exact intersection of both issues.
Your instinct that this sat “in the sampling layer’s handling of
responseFormat rather than downstream of it” was spot on. :bullseye:

What we fixed in v1.1.0-beta.141 :hammer_and_wrench:

  • :white_check_mark: Corrected the model capability flag — onUnsupported-based
    downgrade now kicks in exactly as documented for that model.
  • :white_check_mark: The host now automatically satisfies the provider’s “json must be
    mentioned” requirement when a responseFormat is present, so you
    never need to word your prompts around it.
  • :white_check_mark: Better errors: sampling provider failures now include the selected
    model and the applied response format in the error message.

After the fix, your exact scenario returns a valid JSON completion, and
when a downgrade happens you’ll see it reflected in
_meta.responseFormat (applied, downgraded, structuredValid) so
your tool can decide whether to re-validate against your schema. :package:

One small thing on your side :light_bulb:

The error: "" you observed was the platform’s structured error
(code -32003, with a full human-readable message) getting swallowed
somewhere in your tool’s sampling error handling before it reached the
invoke result. Worth a quick look — surfacing err.message there will
make any future hiccup much easier to debug. (We’ve also added a
fallback hint on our side so an empty failure shell can’t happen
silently again.)

Thanks again for the top-tier report — please give it another run once
beta.141 is live and let us know how it goes! :rocket:

Follow-up: the v1.1.0-beta.141 fix confirmed (thank you!), but live retesting surfaced a second, separate issue on our own side of the fence.

The good news first: I re-ran the exact generate_recommendations payload from before with fresh Energy, on three different Cloud Agents (destroyed + recreated twice to rule out binary caching). While the failure signature stayed identical, my investigation found a real bug in our own executa: _request_structured_completion’s q.get(timeout=50) could raise a bare queue.Empty on a sampling timeout, and str(queue.Empty()) is "" – so a genuine timeout was rendering as the exact same silent empty-error shell as the original bug, just from an unrelated cause. Fixed and published as wefinance-recommend v0.1.6 and v0.1.7 (surfaces both a real JSON-RPC error’s message and a diagnosable timeout message now). Verified correct in isolation against the actual shipped binary via the real wire protocol, both with a non-responding fake host (produces the new diagnosable timeout message, not an empty one) and with the exact live payload against a properly-responding fake host (succeeds end-to-end).

The new finding: live dashboard tests against app 172 still return the original empty-shell failure unchanged, even after both fixes were published as the “latest” ExecutaVersion for wefinance-recommend. anna-app apps versions wefinance shows why:

in review : v0.2.0 (pinned at submit-review)

App 172’s in-review version 0.2.0 appears to have pinned wefinance-recommend at whatever ExecutaVersion existed when I ran submit-review on 2026-08-20 (v0.1.5) – publishing v0.1.6/v0.1.7 as standalone ExecutaVersions afterward doesn’t seem to retroactively update that pin, so the Cloud Agent is still running the old, unfixed code regardless of which version is “latest.”

Question: is there a way to refresh an in-review App’s pinned tool versions without a full resubmission (which I’d want to avoid mid-review if possible), or does updating a bundled Executa always require cutting and submitting a new App version to actually take effect? Trying to figure out the right way to get the fix in front of your reviewers without disrupting where app 172 already sits in the queue.