Replies: 1 comment
|
Good question — I've been doing exactly this kind of validation recently across three skill systems, and the frame that kept paying off is separating four layers, because each fails independently: 1. Format validity. Does the SKILL.md parse (frontmatter, name/description present)? Cheap, fully automatable. One repo I worked in audits this: missing 2. Installability. Do the skill dirs land where the target agent looks ( 3. Activation. Does the skill trigger when it should and stay quiet when it shouldn't? This is harness-sensitive (#561 is exactly this): descriptions that assume one agent's context conventions misfire elsewhere. I test this by reading, not running — check the description for agent-specific assumptions (tool names, path layouts, runtime services). 4. Behavior. Do the skill's tool calls resolve in the target agent? This is where 'installs fine' diverges from 'works'. I checked a wiki-query-style skill that installed cleanly everywhere but invoked five host-runtime tools with no MCP/bin surface — dead on arrival outside its home agent. No install test catches this; you have to map every tool/command the body references against what the target agent provides. So 'works with any model' decomposes: layers 1–2 are automatable per agent, layer 3 needs a per-agent read, layer 4 needs a tool inventory per agent. In my experience layer 4 is where cross-agent claims actually die, and it's the layer least covered by testing. If I were formalizing it: unit tests for 1–2 per target, a checklist review for 3, and a tool-reference matrix for 4. |
Uh oh!
There was an error while loading. Please reload this page.
Hi Matt,
First, thanks for sharing these skills and for your presentation. I’ve learned a lot from both.
I mostly use the skills with Pi Coding Agent and Gemini models. They’ve worked surprisingly well in my setup, even though I understand Claude Code is a major part of your own workflow. The README also says the skills “work with any model,” which made me curious about how you validate that in practice.
A recent video raised a related concern: a skill can work well in its author’s environment but fail when someone uses a different model or agent. That resonated with me precisely because my experience with this repo has been good. I’d love to understand what helps these skills travel well.
I saw #561, which identifies a harness-specific assumption in
code-review. It’s a useful example of the distinction between being able to install a skill in another agent and having it behave as intended there.Do you currently test any of the skills across different agents or models, formally or informally? And do you have plans for cross-agent or cross-model evaluations, perhaps covering whether a skill activates when it should, avoids activating when it shouldn’t, and produces the intended outcome?
I’m asking out of genuine curiosity, not because I’ve run into a problem. My experience so far has been positive, and I’d be interested to learn how you think about portability and what “works with any model” is meant to cover.
All reactions