Skill Audit — Evaluability Inventory¶
Generated: 2026-09-05 | vt-base v6.2.6 | 133 skills | SPEC-156 Wave 1
Every skill in every plugin manifest is classified into exactly one bucket. Wave 1 measures evaluability only — whether a skill could be benchmarked — and makes no keep-or-drop call about any skill.
Headline counts¶
| Bucket | Skills | Share of manifest |
|---|---|---|
| Evaluable | 3 | 2.3% |
| Auto-keep (private data) | 19 | 14.3% |
| Auto-keep (toolkit internal) | 12 | 9.0% |
| Not evaluable | 99 | 74.4% |
The two auto-keep fractions are reported separately on purpose: the toolkit-internal bucket is registry plumbing, not private context the base model lacks, and merging them would overstate how much of the library is protected by private data.
Reading this number¶
A low evaluable share is the expected result, not a failed audit. It is a
statement about eval coverage — how many skills carry a suite under
tests/evals/ — and nothing else.
Evaluable (3)¶
| Skill | Category | Reason |
|---|---|---|
vt-d-spec-from-requirements |
Spec tools | eval suite present: tests/evals/vt-d-spec-from-requirements |
vt-d-2-plan |
Workflow phases | eval suite present: tests/evals/vt-d-2-plan |
vt-d-4-review |
Workflow phases | eval suite present: tests/evals/vt-d-4-review |
Auto-keep — private data (19)¶
| Skill | Category | Reason |
|---|---|---|
vt-b-visitrans-cd |
Uncategorized | private-data marker: BRAND_ASSETS_ROOT |
vt-b-visitrans-design-system |
Uncategorized | private-data marker: assets/logos |
vt-v-audit |
Uncategorized | private-data marker: vault/VMS |
vt-v-cwp |
Uncategorized | private-data marker: vault/ISMSI |
vt-v-draft |
Uncategorized | private-data marker: vault/ISMSI |
vt-v-intake |
Uncategorized | private-data marker: vault/ISMSI |
vt-v-interview |
Uncategorized | private-data marker: Vorlage |
vt-v-plan |
Uncategorized | private-data marker: vault/ISMSI |
vt-v-review |
Uncategorized | private-data marker: vault/IMS |
vt-v-start |
Uncategorized | private-data marker: vault/ISMSI |
vt-v-write |
Uncategorized | private-data marker: plugins/vms/ |
vt-o-document-converter-branded |
Document generation | private-data marker: BRAND_ASSETS_ROOT |
vt-o-meeting-minutes |
Meetings | private-data marker: BRAND_ASSETS_ROOT |
vt-p-research-ingest |
Intake & research | private-data marker: ONEDRIVE_ROOT |
vt-d-pd-0-start |
Product design | private-data marker: Vorlage |
vt-d-pd-2-prd |
Product design | private-data marker: ONEDRIVE_ROOT |
vt-d-pd-3-prototype |
Product design | private-data marker: BRAND_ASSETS_ROOT |
vt-d-kw-prd |
Amendment additions | private-data marker: ONEDRIVE_ROOT |
vt-d-kw-prototype |
Amendment additions | private-data marker: BRAND_ASSETS_ROOT |
Auto-keep — toolkit internal (12)¶
| Skill | Category | Reason |
|---|---|---|
vt-c-project-register |
Repo governance (4) | toolkit-internal marker: intake/projects.yaml |
vt-p-domain-dashboard |
PM aggregation & reporting | toolkit-internal marker: intake/projects.yaml |
vt-p-project-sync |
PM aggregation & reporting | toolkit-internal marker: intake/projects.yaml |
vt-p-pov |
Intake & research | toolkit-internal marker: intake/projects.yaml |
vt-d-scaffold |
Scaffolding | toolkit-internal marker: intake/projects.yaml |
vt-d-complete |
Spec tools | toolkit-internal marker: intake/projects.yaml |
vt-d-bug-report |
Debugging & ops | toolkit-internal marker: intake/projects.yaml |
vt-d-triage-bugs |
Debugging & ops | toolkit-internal marker: intake/projects.yaml |
vt-t-skill-demote |
Skill authoring & lifecycle | toolkit-internal marker: intake/projects.yaml |
vt-t-skill-promote |
Skill authoring & lifecycle | toolkit-internal marker: intake/projects.yaml |
vt-t-toolkit-review |
Toolkit maintenance | toolkit-internal marker: intake/projects.yaml |
vt-t-toolkit-update |
Toolkit maintenance | toolkit-internal marker: intake/projects.yaml |
Not evaluable — needs an eval suite (99)¶
This is a measurement gap, not a judgement. A skill in this list has no eval suite yet, so it cannot be benchmarked — that says nothing about whether it earns its keep. Wave 1 makes no keep-or-drop call about any skill, and nothing in this report removes or downgrades one.
Each skill below needs an eval suite before it can be measured. This inventory flags the gap; it does not author suites.
| Skill | Category | Reason |
|---|---|---|
vt-f-report |
Uncategorized | no eval suite under tests/evals/vt-f-report |
vt-v-compliance-checklist |
Uncategorized | no eval suite under tests/evals/vt-v-compliance-checklist |
vt-v-sprint |
Uncategorized | no eval suite under tests/evals/vt-v-sprint |
vt-c-repo-audit |
Repo governance (4) | no eval suite under tests/evals/vt-c-repo-audit |
vt-c-repo-health |
Repo governance (4) | no eval suite under tests/evals/vt-c-repo-health |
vt-c-repo-init |
Repo governance (4) | no eval suite under tests/evals/vt-c-repo-init |
vt-c-compound-docs |
Memory, learning & compound engineering (8) | no eval suite under tests/evals/vt-c-compound-docs |
vt-c-continuous-learning |
Memory, learning & compound engineering (8) | no eval suite under tests/evals/vt-c-continuous-learning |
vt-c-dream |
Memory, learning & compound engineering (8) | no eval suite under tests/evals/vt-c-dream |
vt-c-knowledge-index |
Memory, learning & compound engineering (8) | no eval suite under tests/evals/vt-c-knowledge-index |
vt-c-learning-metrics |
Memory, learning & compound engineering (8) | no eval suite under tests/evals/vt-c-learning-metrics |
vt-c-openbrain-capture |
Memory, learning & compound engineering (8) | no eval suite under tests/evals/vt-c-openbrain-capture |
vt-c-session-consolidator |
Memory, learning & compound engineering (8) | no eval suite under tests/evals/vt-c-session-consolidator |
vt-c-session-journal |
Memory, learning & compound engineering (8) | no eval suite under tests/evals/vt-c-session-journal |
vt-c-dispatching-parallel-agents |
Parallel work & worktrees (3) | no eval suite under tests/evals/vt-c-dispatching-parallel-agents |
vt-c-git-worktree |
Parallel work & worktrees (3) | no eval suite under tests/evals/vt-c-git-worktree |
vt-c-using-git-worktrees |
Parallel work & worktrees (3) | no eval suite under tests/evals/vt-c-using-git-worktrees |
vt-o-doc-versioning |
Document generation | no eval suite under tests/evals/vt-o-doc-versioning |
vt-o-docs-pipeline-orchestrator |
Document generation | no eval suite under tests/evals/vt-o-docs-pipeline-orchestrator |
vt-o-every-style-editor |
Document generation | no eval suite under tests/evals/vt-o-every-style-editor |
vt-o-c4-diagram |
Diagrams | no eval suite under tests/evals/vt-o-c4-diagram |
vt-o-diagram-validate |
Diagrams | no eval suite under tests/evals/vt-o-diagram-validate |
vt-o-markdown-diagram-processor |
Diagrams | no eval suite under tests/evals/vt-o-markdown-diagram-processor |
vt-o-mermaid-diagrams-branded |
Diagrams | no eval suite under tests/evals/vt-o-mermaid-diagrams-branded |
vt-o-mermaid-to-images |
Diagrams | no eval suite under tests/evals/vt-o-mermaid-to-images |
vt-o-gamma-presentation |
Presentations & media | no eval suite under tests/evals/vt-o-gamma-presentation |
vt-o-gemini-imagegen |
Presentations & media | no eval suite under tests/evals/vt-o-gemini-imagegen |
vt-o-remotion-video |
Presentations & media | no eval suite under tests/evals/vt-o-remotion-video |
vt-p-weekly-planning |
PM aggregation & reporting | no eval suite under tests/evals/vt-p-weekly-planning |
vt-p-autoresearch-agent |
Intake & research | no eval suite under tests/evals/vt-p-autoresearch-agent |
vt-p-content-evaluate |
Intake & research | no eval suite under tests/evals/vt-p-content-evaluate |
vt-p-inbox-qualify |
Intake & research | no eval suite under tests/evals/vt-p-inbox-qualify |
vt-p-repo-evaluate |
Intake & research | no eval suite under tests/evals/vt-p-repo-evaluate |
vt-p-research-implement |
Intake & research | no eval suite under tests/evals/vt-p-research-implement |
vt-p-web-capture |
Intake & research | no eval suite under tests/evals/vt-p-web-capture |
vt-p-project-status-sync |
Absorbed from the project-updates plugin (PD-5) | no eval suite under tests/evals/vt-p-project-status-sync |
vt-p-project-updates |
Absorbed from the project-updates plugin (PD-5) | no eval suite under tests/evals/vt-p-project-updates |
vt-d-0-init |
Entry points | no eval suite under tests/evals/vt-d-0-init |
vt-d-0-start |
Entry points | no eval suite under tests/evals/vt-d-0-start |
vt-d-dev-start |
Entry points | no eval suite under tests/evals/vt-d-dev-start |
vt-d-1-bootstrap |
Scaffolding | no eval suite under tests/evals/vt-d-1-bootstrap |
vt-d-bootstrap |
Scaffolding | no eval suite under tests/evals/vt-d-bootstrap |
vt-d-pd-1-research |
Product design | no eval suite under tests/evals/vt-d-pd-1-research |
vt-d-pd-4-validate |
Product design | no eval suite under tests/evals/vt-d-pd-4-validate |
vt-d-pd-5-specs |
Product design | no eval suite under tests/evals/vt-d-pd-5-specs |
vt-d-pd-6-handoff |
Product design | no eval suite under tests/evals/vt-d-pd-6-handoff |
vt-d-pd-analyze-changes |
Product design | no eval suite under tests/evals/vt-d-pd-analyze-changes |
vt-d-pd-capture-decisions |
Product design | no eval suite under tests/evals/vt-d-pd-capture-decisions |
vt-d-pd-inbox-scan |
Product design | no eval suite under tests/evals/vt-d-pd-inbox-scan |
vt-d-pd-pid |
Product design | no eval suite under tests/evals/vt-d-pd-pid |
vt-d-pd-route-decision |
Product design | no eval suite under tests/evals/vt-d-pd-route-decision |
vt-d-activate |
Spec tools | no eval suite under tests/evals/vt-d-activate |
vt-d-shape |
Spec tools | no eval suite under tests/evals/vt-d-shape |
vt-d-spec-check |
Spec tools | no eval suite under tests/evals/vt-d-spec-check |
vt-d-specs-from-prd |
Spec tools | no eval suite under tests/evals/vt-d-specs-from-prd |
vt-d-3-build |
Workflow phases | no eval suite under tests/evals/vt-d-3-build |
vt-d-5-finalize |
Workflow phases | no eval suite under tests/evals/vt-d-5-finalize |
vt-d-6-operate |
Workflow phases | no eval suite under tests/evals/vt-d-6-operate |
vt-d-phase-checkpoint |
Workflow phases | no eval suite under tests/evals/vt-d-phase-checkpoint |
vt-d-defense-in-depth |
Building | no eval suite under tests/evals/vt-d-defense-in-depth |
vt-d-executing-plans |
Building | no eval suite under tests/evals/vt-d-executing-plans |
vt-d-feature-flags |
Building | no eval suite under tests/evals/vt-d-feature-flags |
vt-d-frontend-design |
Building | no eval suite under tests/evals/vt-d-frontend-design |
vt-d-recursive-criticism |
Building | no eval suite under tests/evals/vt-d-recursive-criticism |
vt-d-safe-migrations |
Building | no eval suite under tests/evals/vt-d-safe-migrations |
vt-d-test-driven-development |
Testing | no eval suite under tests/evals/vt-d-test-driven-development |
vt-d-testing-anti-patterns |
Testing | no eval suite under tests/evals/vt-d-testing-anti-patterns |
vt-d-webapp-testing |
Testing | no eval suite under tests/evals/vt-d-webapp-testing |
vt-d-error-monitoring |
Debugging & ops | no eval suite under tests/evals/vt-d-error-monitoring |
vt-d-root-cause-tracing |
Debugging & ops | no eval suite under tests/evals/vt-d-root-cause-tracing |
vt-d-systematic-debugging |
Debugging & ops | no eval suite under tests/evals/vt-d-systematic-debugging |
vt-d-verification-before-completion |
Debugging & ops | no eval suite under tests/evals/vt-d-verification-before-completion |
vt-d-npm-security |
Dependency security | no eval suite under tests/evals/vt-d-npm-security |
vt-d-python-security |
Dependency security | no eval suite under tests/evals/vt-d-python-security |
vt-d-promote |
Quality & release | no eval suite under tests/evals/vt-d-promote |
vt-d-quality-metrics |
Quality & release | no eval suite under tests/evals/vt-d-quality-metrics |
vt-d-quickfix |
Quality & release | no eval suite under tests/evals/vt-d-quickfix |
vt-d-adr |
Docs in flow | no eval suite under tests/evals/vt-d-adr |
vt-d-user-manual-generate |
Docs in flow | no eval suite under tests/evals/vt-d-user-manual-generate |
vt-d-user-manual-update |
Docs in flow | no eval suite under tests/evals/vt-d-user-manual-update |
vt-d-agent-native-architecture |
Build patterns | no eval suite under tests/evals/vt-d-agent-native-architecture |
vt-d-mcp-builder |
Build patterns | no eval suite under tests/evals/vt-d-mcp-builder |
vt-d-multi-tenancy |
Build patterns | no eval suite under tests/evals/vt-d-multi-tenancy |
vt-d-andrew-kane-gem-writer |
Language packs | no eval suite under tests/evals/vt-d-andrew-kane-gem-writer |
vt-d-dhh-ruby-style |
Language packs | no eval suite under tests/evals/vt-d-dhh-ruby-style |
vt-d-dspy-ruby |
Language packs | no eval suite under tests/evals/vt-d-dspy-ruby |
vt-d-container-logistics-ux-expert |
Amendment additions | no eval suite under tests/evals/vt-d-container-logistics-ux-expert |
vt-d-kw-user-test |
Amendment additions | no eval suite under tests/evals/vt-d-kw-user-test |
vt-d-messegelaende-cleanup |
Amendment additions | no eval suite under tests/evals/vt-d-messegelaende-cleanup |
vt-d-skill-venue-hall-research |
Amendment additions | no eval suite under tests/evals/vt-d-skill-venue-hall-research |
vt-d-statusline-spec-phase |
Amendment additions | no eval suite under tests/evals/vt-d-statusline-spec-phase |
vt-t-create-agent-skills |
Skill authoring & lifecycle | no eval suite under tests/evals/vt-t-create-agent-skills |
vt-t-skill-audit |
Skill authoring & lifecycle | no eval suite under tests/evals/vt-t-skill-audit |
vt-t-skill-creator |
Skill authoring & lifecycle | no eval suite under tests/evals/vt-t-skill-creator |
vt-t-skill-eval |
Skill authoring & lifecycle | no eval suite under tests/evals/vt-t-skill-eval |
vt-t-skill-structural-screen |
Skill authoring & lifecycle | no eval suite under tests/evals/vt-t-skill-structural-screen |
vt-t-claudemd-evolve |
Toolkit maintenance | no eval suite under tests/evals/vt-t-claudemd-evolve |
vt-t-doc-sync |
Toolkit maintenance | no eval suite under tests/evals/vt-t-doc-sync |
vt-t-gc |
Toolkit maintenance | no eval suite under tests/evals/vt-t-gc |
Provenance¶
- Regenerate:
bash scripts/generate-skill-audit.sh - Source of truth:
plugins/vt-base/.claude-plugin/skill-symlinks.manifest - Close a gap: author
tests/evals/<skill-name>/*.yaml(>=5 cases, SPEC-098 US4). A skill becomesevaluableas soon as that directory holds one case. - Wave-2 gate: verdict policy stays closed until >=20 skills carry suites
across >=3 manifest categories. That threshold is why every JSON entry
carries
category. Today: 3 skills, see the JSON for the category spread.