Skills catalog
← index
colibri-skills is Colibri’s read-only runtime consumer for skill
artifacts. Skills are authored and reviewed in this repo’s .agent/skills/
directory (the canonical home since the clawdie-ai retirement, PR #146);
Colibri indexes them, validates checksums, chunks searchable text, and exposes
typed structs to the daemon, CLI, and TUI. This crate does not author skills.
→ crates/colibri-skills/src/lib.rs
Decisions
Razdelek z naslovom „Decisions”Source of truth stays in Colibri
Razdelek z naslovom „Source of truth stays in Colibri”Skill artifacts live in this repository’s .agent/skills/, not in a separate
repo. They are committed reviewed directories containing prose, screenshots,
transcripts, scripts, a manifest, and a checksum file. colibri-skills imports
these artifacts into Colibri’s SQLite store at runtime.
This preserves review discipline: a skill changes through a PR in its home repo, then Colibri re-indexes the checkout. (Skills were migrated here from the now-retiring clawdie-ai repo in PR #146.)
Catalog prune (02.aug.2026): the catalog was pruned from 59 to 47 skills.
Twelve clawdie-ai/NanoClaw TS-era channel and tool-integration skills were
removed — add-discord, add-gmail, add-protonmail, add-slack,
add-telegram, add-telegram-swarm, add-voice-transcription, add-parallel,
customize, x-integration, docs-deployment, docs-localization-pipeline
(channels now live in zot; Crowdin retired). Every remaining SKILL.md and its
reference/intent files carry an okf: frontmatter block; the importer ignores
unknown keys, so the catalog index is unaffected.
Read-only, not authoring
Razdelek z naslovom „Read-only, not authoring”The crate deliberately lacks “create skill” or “edit skill” operations. Those
belong in .agent/skills/ where human review and media pipelines run. Putting
authoring here would duplicate state and split review authority.
The import path is target for Phase 1: scan the .agent/skills/ checkout,
parse manifests, verify checksums, and upsert into SQLite. The type scaffold
exists today; the importer, chunker, and FTS5 index are planned.
Manifest-driven identity
Razdelek z naslovom „Manifest-driven identity”Each skill directory contains a run manifest file. From it the importer derives:
skill_iddisplay_namesource_pathwithin the Colibri repo- pipeline stages and models used
- source media metadata
Any file not listed in the manifest can still be classified and indexed as an artifact, but the manifest is the canonical identity document.
Artifact classification by extension and filename
Razdelek z naslovom „Artifact classification by extension and filename”ArtifactType::from_path classifies files without relying on a sidecar:
- Python or shell files → Script
- paths containing contact_sheet → ContactSheet
- paths containing run_manifest and ending in .json → Manifest
- paths containing sha256 or checksum → Checksum
- paths containing report and ending in .json → Report
- .md → Document
- .jpg / .png / .webp → Image
- .txt transcript files → Transcript
- anything else → Other
This heuristic keeps classification local and fast. Misclassified files can be
fixed by renaming within .agent/skills/.
→ crates/colibri-skills/src/lib.rs (ArtifactType::from_path)
Checksums are validated, then stored
Razdelek z naslovom „Checksums are validated, then stored”The run manifest is accompanied by a checksum file. At import time the runtime
computes SHA-256 of each artifact and compares it to the committed checksum.
Failures are reported in ImportSummary::checksum_failures and prevent
success().
Only the hash is stored in SQLite; image and media blobs stay on disk. The catalog stores relative paths and hashes, not the binary content.
Content is chunked into searchable units
Razdelek z naslovom „Content is chunked into searchable units”The planned chunker turns skill content into SkillChunk rows:
- Markdown sections by heading
- Command blocks
- Code blocks
- Tables
- Transcript segments
Chunks are the unit of search and the unit shown in TUI or CLI results.
SkillChunk carries line_start/line_end so a hit can point back to the
source artifact.
→ crates/colibri-skills/src/lib.rs (SkillChunk, ChunkType)
SQLite + FTS5 as the runtime search backend
Razdelek z naslovom „SQLite + FTS5 as the runtime search backend”The target schema keeps three tables:
system_skills— one row per skillsystem_skill_artifacts— one row per filesystem_skill_chunks— one row per searchable chunk, plus a virtual FTS5 table for ranked text search
This matches the store’s pragmatic relational model. If skill volumes grow beyond tens of thousands of chunks, we can move the FTS index to PostgreSQL pgvector; until then, SQLite keeps the control-plane self-contained.
Status is a lifecycle marker, not a state machine
Razdelek z naslovom „Status is a lifecycle marker, not a state machine”SkillStatus is active, archived, or superseded. There is no pending
review state because review happens in .agent/skills/ before import. Colibri simply
stops returning archived skills in default searches but keeps them in the store
for audit and explicit lookups.
Natural-language verification question
Razdelek z naslovom „Natural-language verification question”Each skill can carry a verification field like “can the user create and run
an Astro project?”. This is not an executable test; it is the acceptance
criterion used during skill review and later during agent self-verification.
Runtime commands are read-only
Razdelek z naslovom „Runtime commands are read-only”The CLI surface is planned as:
colibri list-skillscolibri show-skill <id>colibri search-skills <query>colibri index-skillscolibri verify-skill <id>
index-skills refreshes the catalog from disk. The remaining commands query the
runtime store. None mutate the .agent/skills/ checkout.
Entity shape
Razdelek z naslovom „Entity shape”Skill ├─ skill_id, display_name, source_path, status, verification ├─ SkillManifest │ ├─ run_id, created, notes │ ├─ ManifestSource │ ├─ [PipelineStage] │ └─ [ModelUsage] └─ [SkillArtifact] ├─ artifact_type, relative_path, file_name, mime_type, size_bytes, sha256_hash └─ [SkillChunk] ├─ chunk_type, heading, content, line_start, line_end, tokens_estimateSee also
Razdelek z naslovom „See also”- store-schema — coordination and planned skill catalog tables
- operator-cli — planned skill catalog CLI commands
- task-board — agents will match claimed tasks to skills by capability