Skip to content

Skill frontmatter: preserve unknown keys on write, project to known keys on read #657

Description

@rubenvdlinde

Describe the feature you'd like to request

What happens today

The skills API added in assistant#594 treats frontmatter two different ways.

listSkills parses the frontmatter and returns two fields. parseMetadataFields reads only the keys in FRONTMATTER_METADATA_FIELDS and ignores every other key without complaining.

loadSkill returns the raw file. loadSkillFromFolder ends in return $skillFile->getContent();, so every key in the frontmatter goes back to the caller.

storeSkill rebuilds the frontmatter from scratch:

$frontmatter = Yaml::dump([
    'name' => $skillName,
    'description' => $description,
]);

On an overwrite that replaces the whole block, so any other key is lost preventig proper frontmatter use for metadata.

Reproduction

Verified on Nextcloud 35.0.1 RC1, assistant 4.0.0, context_agent 2.8.0.

Write a skill over WebDAV with extra frontmatter keys:

B='http://localhost/remote.php/dav/files/admin/Assistant/Context%20Agent/Skills'
curl -u admin:admin -X MKCOL "$B/rich-skill"
cat > SKILL.md <<'EOF'
---
name: rich-skill
description: Skill carrying extra metadata
maturityLevel: 3
targetLevel: 4
state: active
source: learned
provenance: run-7f3a
levelEvidence:
  - eval-run-221
---

### Steps

1. Governed step.
EOF
curl -u admin:admin -T SKILL.md "$B/rich-skill/SKILL.md"

GET /ocs/v2.php/apps/assistant/api/v1/skills lists it correctly and drops the extra keys:

{"skills":[{"name":"rich-skill","description":"Skill carrying extra metadata"}]}

GET /ocs/v2.php/apps/assistant/api/v1/skills/rich-skill returns the file byte for byte:

{"content":"---\nname: rich-skill\ndescription: Skill carrying extra metadata\nmaturityLevel: 3\ntargetLevel: 4\nstate: active\nsource: learned\nprovenance: run-7f3a\nlevelEvidence:\n  - eval-run-221\n---\n\n## Steps\n\n1. Governed step.\n"}

POST to the same route with skillName=rich-skill reduces the stored frontmatter back to name and description. The other six keys are gone, and nothing reports it (there is also no invallid input error raised).

Why this costs something

The load path feeds a model. In context_agent, load_skill ends in return res['content'] and its docstring says it returns "frontmatter + markdown body". Whatever anyone writes in that frontmatter is paid for in tokens on every load, whether the model can act on it or not.

context_agent#212 already measures the cost on the list path: 20 skills is roughly 5500 tokens of injected metadata, and user skills are combined with the admin provided ones. The load path has the same problem per skill, and no ceiling, because nothing bounds what a frontmatter block may contain.

The write path is worse than a cost. It is data loss. A tool call can destroy metadata that the user or another app put there, and nothing reports it (again no error is raised).

Describe the solution you'd like

What I propose

Keep one rule in both directions.

Preserve unknown keys on write. storeSkill reads the existing frontmatter, updates name and description, and keeps everything else. This is a bug fix and it stands alone.

Project to known keys on read. loadSkill returns the body plus the keys the agent needs, not the raw file.

name and description stay required. Everything else is stored faithfully and left out of what the API hands the model.

What it opens up

Once a key can survive a write, a skill can carry things the API does not need to understand yet.

Lifecycle state, so a skill can be marked draft, active or retired.

Provenance, so a generated skill records what produced it.

Group and user scoping for the admin published library. This is the one I care about most, and the trust boundary is real. getGlobalSkillsFolder() reads from an admin's storage, which ordinary users cannot write. A groups key on a skill in that folder can be enforced, because the server filters before the user sees anything. The same key in a user's own skills folder is a preference and nothing more, since the user owns the file. Worth stating so nobody mistakes one for the other.

Scoping is also a token reduction. A large organisation stops injecting a global library that mostly does not apply to the reader.

It also answers part of context_agent#239, which asks for skill listing and selection that does not hand the model everything at once.

So in short this prepares the skills for more future functionality.

What I am not proposing here

No group filtering in this change. That needs work in the global folder path and deserves its own review.

No new required keys.

No change to listSkills, which already does the right thing.

Next

I have a patch for both halves and will open it against main today, linked to this issue. The storeSkill fix is independent, so it can be taken on its own if the read side needs more discussion.

Describe alternatives you've considered

A sidecar metadata file next to SKILL.md. I actually built it that way first, and it didn't work out (so it was a wrong call).

Skills get copied, moved, shared, zipped and synced to git. Frontmatter survives all of that because it is the same file. A sidecar is one careless copy away from being separated from the skill it describes. The agentskills.io format already treats frontmatter as the metadata channel, so this keeps one channel instead of two.

Fix only the write path. Preserving unknown keys on write, and leaving loadSkill returning the raw file. That stops the data loss but leaves the token cost, and it leaves the API inconsistent with listSkills. Worth taking on its own if the read side needs more discussion, which is why the patch keeps the two commits separate. But I strongly sugest keeping a uniform line.

Do nothing and keep metadata outside Nextcloud. An external store keyed by skill name. It breaks the moment a user renames a folder, and it does not survive an export (we have something like this curently setup and want to move away from it).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions