The headline result turned out to be about triggering rather than presentation, and I think it matters to anyone relying on Skills for discovery.
All numbers below are from today, CLI v0.1.49, app installed and enabled, skill installed and enabled.
Test 1 — bare paste, no mention: 0/5
Five separate fresh chats. Each received exactly this and nothing else:
Traceback (most recent call last):
File "/srv/app/main.py", line 12, in <module>
import requests
ModuleNotFoundError: No module named 'requests'
My Skill’s description begins with the literal string Traceback (most recent call last) and also contains ModuleNotFoundError. So the input is a verbatim substring match on two separate terms in the description.
The tool was invoked zero times out of five. No fingerprint in any response. All five were answered by the agent from its own knowledge.
One response contained 严重程度: 中等 置信度: 90% — severity and confidence values that only exist in my knowledge base — which suggests the skill body reached the context but the tool call did not happen.
Test 2 — same input with #Error Journal: invoked
Identical paste, prefixed with the app mention. The tool ran; the response carried sha256:71bfc4e4....
So the tool, the executa, the grants and the storage are all fine. The gap is specifically whether a Skill wins against the agent’s own judgement, and on this evidence it does not.
Why I think this happens
The agent recognises a common Python error, knows the answer, and answers. The Skill has to win that competition on every single message. For a well-known error it apparently never does — which inverts the value: the Skill fails exactly where the agent is confident, and my app’s differentiator is not the diagnosis, it is the journal history the agent cannot see.
Question: is Skill matching intended to be advisory, or is it expected to reliably trigger on a literal description match? If advisory, that is worth stating plainly in the Skills docs, because the current wording reads as though a good description produces reliable triggering. It changes how one plans distribution.
Test 3 — presentation is not followed even when the tool runs
In the #mention case, the tool returned fix_steps whose first step is:
python -c 'import sys; print(sys.executable)' # which interpreter is this?
That ordering is deliberate: installing into the wrong interpreter is the actual failure mode, so identifying it comes first.
The agent demoted that to a bullet under “Quick Tips” at the bottom and led with pip install requests — the advice the entry exists to avoid. It also introduced uv commands, which appear nowhere in my knowledge base.
Both SKILL.md and system_prompt_addendum say to print fix_steps verbatim and in order, and to lead with the headline field. Neither happened.
I have since changed the tool to return a preformatted headline string rather than separate fields, on the theory that a ready-to-print sentence is harder to ignore than an object requiring assembly. That version is pending review, so I cannot yet report whether it helps.
Question: how strongly is system_prompt_addendum weighted relative to the agent’s own presentation preferences? If a returned field is meant to be printed verbatim, is there a supported mechanism for that, or is it always advisory?
Two side observations
Unprompted language switching. Three of the five bare-paste responses came back in Chinese. The input was English, the skill is English, and my account language is English. Not blocking, but surprising, and it would be confusing for an end user.
The agent executed commands unprompted. One response included 上面的安装是在我的工作区演示用的 — “the installation above was demonstrated in my workspace” — meaning it actually ran pip install requests in its own sandbox to check. Impressive, and not something I asked for. Worth knowing that a diagnostic paste can cause real execution.
Why this matters beyond my app
Skills are presented as the discovery mechanism that does not require users to know your app exists. If they only fire when the agent has nothing better to say, then in practice apps are reached by explicit mention, and every builder should plan distribution on that basis rather than on auto-trigger.
I would rather be wrong about this. If there is something in my Skill configuration causing it, I would like to know — happy to share the description, the metadata block, or run any test that would help isolate it.