Skill++ — the design¶
This is the design document the project started from. Parts of it describe plans that were changed or never built, and some numbers are from before later measurements. For how Skill++ works today, read the README, usage.md and architecture.md; the measurements are in research/. Code comments cite its sections as
docs/design.md §N.
Skill++ is a background-observing knowledge engine for developers and technical teams. It watches how work actually gets done, keeps a searchable ledger of candidate workflows, and — only on explicit human approval — promotes them into modular, enterprise-ready SKILL.md files.
Capture is passive. Promotion is always deliberate.
./examples/demo.sh
That runs the whole loop against a scratch ledger — three captured sessions, a redacted credential, generated questions, a scaffolded skill — and touches nothing real. See §12 for what is built and what is not.
1. Core Vision & Value Proposition¶
- Zero-Friction Capture: Routine developer work (terminal pipelines, MCP tool chains, refactoring sequences) and dictated natural-language workflows land in the ledger automatically, with no interruption and no prompt engineering.
- Bottom-Up Intelligence: Operational knowledge is derived from what the team actually does, rather than from documentation someone was supposed to write.
- Nothing Unreviewed Ships: The ledger holds candidates, not skills. A
SKILL.mdis synthesized only after a human reads what it will do and approves it. - Open Standard Native: Output is the standard
SKILL.mdformat, compatible with Claude Code, Cursor, OpenCode, and Microsoft Agent Framework. (See §7 for the honest limits of that portability.)
2. System Architecture¶
┌──────────────────────────────────────────────────────────────┐
│ DAILY INPUTS │
│ • Passive MCP & terminal execution traces │
│ • Active natural-language text / voice dictation │
└───────────────────────────────┬──────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────┐
│ SEGMENTATION │
│ Session cut into task episodes — one workflow per entry │
│ Boundaries: completion markers · new prompt │
│ No marker and nothing shipped → flagged, not proposed │
└───────────────────────────────┬──────────────────────────────┘
│ summarized + sanitized on write
▼
┌──────────────────────────────────────────────────────────────┐
│ THE LEDGER │
│ Compact markdown candidate entries — never raw traces │
│ Searchable · local-first · unapproved entries expire 7–14d │
└───────────────────────────────┬──────────────────────────────┘
│
┌──────────────────┴──────────────────┐
▼ ▼
recurrence signal (N ≥ 3) user ledger query
batched — never a popup "that rollback last Tuesday"
└──────────────────┬──────────────────┘
▼
┌──────────────────────────────────────────────────────────────┐
│ REVIEW & PROMOTION │
│ Pull-based: review command · session end · weekly digest │
│ Effect summary → source traces → 2-question interview │
└───────────────────────────────┬──────────────────────────────┘
│ approved
▼
┌──────────────────────────────────────────────────────────────┐
│ SYNTHESIS │
│ Parameterize · dedup (embedding match) · compose │
│ Emit: SKILL.md + scripts/ + declared deps in metadata │
└───────────────────────────────┬──────────────────────────────┘
▼
┌──────────────────────────────────────────────────────────────┐
│ DISTRIBUTION & LIFECYCLE │
│ provisional → trusted (earned through successful use) │
│ hot → cold → archived (demoted, never auto-deleted) │
│ Local `.claude/skills/` · team PR · dep check at pull time │
└──────────────────────────────────────────────────────────────┘
3. How It Works¶
Step 1 — Ingestion & Observation¶
- Passive listener: A lightweight background daemon or terminal hook observes execution traces — commands, MCP calls, file edit sequences.
- Active dictation: Workflows can be dictated in plain English ("when we update X, run Y, check Z, then notify the on-call"), for quick-thinking developers and non-technical contributors alike.
Step 2 — Segment into tasks¶
A session is not a workflow. One sitting routinely holds several unrelated tasks — a deploy, an unrelated bug fix, an investigation that goes nowhere — and fingerprinting the whole thing as one unit is why a workflow performed three times can register as three unrelated one-offs that never reach the threshold. Developers reuse tasks, not sessions, so the unit of comparison has to be the unit of reuse.
Sessions are therefore cut into episodes before anything is fingerprinted, on two signals that cost nothing to compute:
- Completion markers — a command whose success means the developer's goal is done, not merely that a step worked. Version-control verbs qualify: nobody commits halfway through a thought. Infrastructure commands (
terraform apply,kubectl apply,npm publish) deliberately do not — see §5b. - A new prompt — the developer stating a fresh goal, recognised by its position in the recorded stream rather than by any clock. No idle-gap timer: a threshold needs tuning per person and misfires the moment somebody reads documentation mid-task.
A boundary that would leave an episode of fewer than two substantive steps is ignored, because one step is not a workflow. An episode that ends only because the session did, with no marker, is flagged rather than proposed — that is an investigation with nothing to show for itself.
Step 3 — Summarize on Write (not on read)¶
Raw traces are never persisted. At capture time each observation is compressed into a compact markdown ledger entry and scrubbed in the same pass:
- Compression keeps the ledger the same order of magnitude as the skill library itself, rather than the tens of megabytes raw MCP payloads and file diffs would consume.
- Sanitization happens once, on write. A regex scan strips API keys, tokens, credentials, internal URLs and email addresses before anything touches disk — so the ledger is never a liability sitting in a buffer waiting to be cleaned later.
- Searchability comes for free, because entries are already text.
Step 4 — Candidate Surfacing (two entry points)¶
The ledger is not just a suggestion queue; it is a searchable record of your own work.
- Recurrence signal: Workflows recurring 3+ times across sessions are flagged as proposal candidates.
- User query: Developers can search the ledger directly — "that migration rollback from last Tuesday" — and promote a one-off themselves. This matters because value and frequency correlate only weakly: the highest-value procedures (incident response, cert rotation, quarterly release) are often rare by nature.
Unapproved candidates expire and self-delete after 7–14 days.
Step 5 — Review & Promotion (pull, never push)¶
No interrupting popups. Candidates persist in the ledger, so review is something the developer pulls when they have attention to spare: an explicit review command, a prompt at session end, a weekly digest, or a nudge at PR time.
Review is designed to take seconds, not minutes:
- Effect summary first. The proposal leads with what the skill will do — commands it runs, paths it writes, anything destructive, any network calls — not what it's for. A purpose summary ("deploys to staging and runs smoke tests") can be perfectly accurate while the underlying steps are wrong; effect summaries are what make a fast review a real one.
- Evidence one keypress away. The candidate is shown against the ledger entries it was derived from, with diverging steps highlighted. Confirming "yes, that's what I did" is far faster than auditing free-floating instructions.
- Targeted clarification. Because synthesis happens at approval, the system can ask what traces can never reveal — but only where the trace is genuinely ambiguous. See §4.
Nothing is written to the skill library without passing this gate.
Step 6 — Synthesis¶
Only after approval is a SKILL.md generated.
- Parameterization: Local paths (
/Users/dev/project/...) and environment-specific values become template variables (${PROJECT_PATH}). - Deduplication: Each saved episode is embedded and compared with every existing entry of the same kind: a run with a conversation by its prompts and the openings of the replies, at
SKILL_PLUS_PLUS_MATCH_FLOOR_TURNS(0.85); a run without one by its steps, one numbered line each, atSKILL_PLUS_PLUS_MATCH_FLOOR(0.93). A match at or above the floor joins that entry rather than spawning a near-duplicate. Each floor is set where wrong merges stop, not where merges are most numerous. - Hierarchical composition: Atomic sub-routines (e.g.
git-commit) are extracted once and invoked as sub-skills by higher-level orchestrators, forming a DAG rather than a flat pile of prompts.
4. Clarification at Approval¶
A trace records what happened, not why. The missing half — the diagnosis behind a retry, the rule behind a parameter, the check that happened in a browser — lives only in the developer's head, and approval is the one moment they are already looking at the workflow. Skill++ uses that moment to close the gap.
It is not a questionnaire. A fixed set of questions gets skipped by the third proposal. Instead the ambiguity in the trace generates the question, which means a clean candidate asks nothing and a messy one asks precisely about the part that is messy.
Where the questions come from¶
| Trace signal | What is missing | Question generated |
|---|---|---|
| Failure then retry — a command exits non-zero, a variant succeeds | The diagnosis, not the fix | "What tells you to reach for --force-lock here?" |
| Divergence across occurrences — run 1 hit staging, runs 2–3 hit prod | The rule behind the variable | "Is this always the current branch, or does it vary?" |
| Workflow ends off-trace — capture stops at deploy | The verification step | "How do you know it worked?" |
| Unparameterizable literal — a bare ID or URL | Whether it is fixed or per-run | "Is this account ID constant across environments?" |
| Step present in some runs only | Whether it is conditional or incidental | "Is the cache clear required, or was that a one-off?" |
The failure-then-retry case is the highest-value one: recovery behavior is the entire difference between a skill and a shell script, and it is the part a trace shows without explaining.
Three disciplines that keep review fast¶
- Cap at three questions. If synthesis has ten, the candidate is not ready — return it to the ledger rather than interrogating the developer. The question count is a quality signal about the candidate, not a budget to spend.
- Pre-fill a guess. "Staging — right?" answered with Enter is a confirmation. An empty text box is composition, and composition is what people skip.
- Skipping never blocks. Unanswered gaps still produce a skill, with an explicit
## Open questionssection, landed as provisional. Better than blocking, and far better than guessing silently — and it gives provisional→trusted promotion something concrete to resolve, since the gap closes the first time someone runs the skill and hits that branch.
When there is no trace¶
A dictated workflow — "create a skill for this: I give you information, you search online about the facts, give me in that format" — has no execution history to mine. Retries and divergence do not exist, so the signal table above has nothing to work on.
What remains mechanically checkable is completeness: whether the description covers trigger, procedure, output shape and failure handling. That requires no understanding of the domain.
| Missing | Detected by | Question generated |
|---|---|---|
| A format named but never given | The text refers to "this format" / "the usual structure" with no example, list or block anywhere in it | "You refer to a format but never give one. What should the output look like — fields, order, an example?" |
| Trigger | No when / if / after / every time condition |
"When should this fire?" |
| Source standard | Research or verification is mentioned with no constraint on what counts | "What counts as a good enough source, and how many need to agree?" |
| Failure handling | No otherwise / if not found / conflict branch |
"What should happen when a step fails or comes back empty?" |
| A procedure that is one step | Fewer than two stages parse out | "What are the actual stages, in order?" |
Dictated candidates skip the recurrence threshold. It exists to filter noise, and an explicit request is not noise.
Ask for an example of the output, never a description of one. A worked example costs the developer less time and specifies more.
The agent answers first¶
Where review runs inside an agent session (see §8), the agent resolves what it can before involving the developer: does deploy.sh still exist, does it accept --env, what does package.json call the test script. Only what the repository cannot answer reaches a human.
On a clean candidate that is often zero questions. When it is three, all three are genuine judgment — which is the only kind worth a developer's attention.
5. Output Formats¶
A SKILL.md is an instruction file, not a tool definition. It cannot declare a tool or provision an MCP server. Skill++ therefore emits along three tracks:
| Captured pattern | Emitted as | Why |
|---|---|---|
| Deterministic command pipeline | scripts/ + thin SKILL.md wrapper |
Captured determinism should come back as code, not as prose an agent re-derives (and re-fumbles) each run. Also far easier to review. |
| Judgment-shaped procedure | Instruction-only SKILL.md |
Decision points, conventions, and escalation paths belong in prose. |
| MCP-dependent workflow | SKILL.md referencing tools by name + declared deps |
Skills reference the host's existing tools; they never install them. |
MCP handling. Skill++ references only MCP servers already connected in the session, and never attempts to bundle or provision one. Three rules keep that safe once a skill travels to a teammate:
- Prefer the portable path. Where the ledger shows the same outcome is reachable through a CLI (
ghinstead of a GitHub MCP,psqlinstead of a Postgres MCP), the shell form is generated — it runs anywhere. MCP references are reserved for capabilities with no CLI equivalent. - Declare dependencies — required servers and CLIs — in the
metadatafrontmatter key, which is already supported in the wild and requires no spec extension. - Check at pull, not at run. On install, declared deps are diffed against the teammate's connected servers and missing ones are reported immediately. A skill whose dependency is absent must state what is missing and stop cleanly — never improvise a workaround.
5b. Why the marker list is short¶
Segmentation (§3, step 2) cuts on completion markers, and the temptation is to treat every "work landed" command as one. Infrastructure verbs are excluded on purpose, and the reason generalises.
Take a deploy that runs terraform apply and then ./scripts/deploy.sh <target>. Treat terraform apply as a marker and the episode closes one step early, leaving deploy.sh as a fragment below the minimum size, which is then discarded. The truncated prefix is identical across every occurrence, so it still merges, still reaches the recurrence threshold, and still presents as a clean candidate — one that builds and provisions but never deploys, with nothing anywhere to flag it as incomplete.
Over-cutting is worse than under-cutting. An under-cut candidate is visibly wrong — a sprawling signature, a title naming the wrong task — and dies at review. An over-cut one looks correct and is silently missing its payload. A marker therefore has to mean the developer's goal is done, not a step succeeded.
Tests going red→green are excluded for the same reason: green tests mean the goal was met only when testing was the goal. Usually they are mid-task verification, and signals.py already mines the failure-then-retry pattern for question generation, which is the right use of it.
6. Lifecycle, Decay & Storage¶
Approved skills are never auto-deleted. Disuse is a poor proxy for value; a timer that removes anything untouched for 60 days preferentially destroys the incident runbook and keeps the command you would have typed from memory. On a shared skill it is worse still, since usage is distributed across a team.
The real cost of an unused skill is index bloat, not disk. So skills are demoted, not deleted:
| Tier | Behavior |
|---|---|
| Hot | Present in the always-loaded name + description index. |
| Cold | Dropped from the index; still discoverable by search and loadable on demand. |
| Archived | Retained, surfaced only by explicit lookup. |
- Staleness ≠ disuse. A skill rots when the script it calls is renamed or the flag it passes is removed. Decay is detected by checking whether referenced paths, commands, and tools still resolve — a cheap, accurate signal that a timer cannot approximate.
- Expiry applies to the ledger, not the library. Unapproved candidates disappear after 7–14 days; anything a human blessed is kept.
- Storage. Skill files average 1.5–3 KB; a full organizational library stays under 5 MB. The ledger stays in the same range because entries are summarized on write rather than stored as raw traces.
- Context cost. Agents load only the lightweight
name+descriptionindex, pulling full instructions into the context window on demand.
7. Portability: The Honest Limits¶
Instruction-shaped skills travel cleanly across Claude Code, Cursor, OpenCode, and Microsoft Agent Framework. Tool-bound skills degrade:
- MCP tool names are host-namespaced.
mcp__github__create_pris not the same identifier in every runtime. - Frontmatter extensions vary.
nameanddescriptionare universal; everything beyond them is host-specific.
Portability is therefore a property of the skill, not of the format. The generation rules in §5 exist to keep as many skills as possible in the portable class.
8. Claude Code Integration¶
Terminal (CLI) is the capture surface. Desktop is the distribution surface.
Claude Code in the terminal is the reference host for passive capture: hooks fire automatically, the ledger builds from every session, and skills land in ~/.claude/skills/ ready to use. Desktop has no passive capture—the desktop app runs its own Claude Code runtime that never reads the host's ~/.claude/settings.json—but it does have an upload path for finished skills.
Installation, verification, and the complete terminal→Desktop workflow are covered in usage.md. This section outlines the architecture.
Capture: hooks, not a daemon¶
Claude Code fires hooks — shell commands receiving JSON on stdin — at PreToolUse, PostToolUse, UserPromptSubmit, SessionStart, SessionEnd, PreCompact, and Stop, configured through settings.json. PostToolUse supplies the tool name, its input, and its result: a structured trace stream, considerably cleaner than parsing shell history. SessionEnd is the natural batching point for ledger writes and the pull-review nudge — it triggers the fold rather than performing it, because a hook that waits on a local model is a hook Claude Code kills. No OS daemon, no separate install, and a far smaller infosec surface than a background listener.
The decisive advantage is UserPromptSubmit. It captures what the developer asked for next to what actually ran. Intent is the half of the picture a raw command log can never recover, and having it in the same session materially improves synthesis — it is what reduces §4's clarification pass from an interview to a confirmation.
It earns its keep twice over, because a prompt is also a task boundary. Prompts are recorded into the ordered step stream, not a separate list, so position alone records which prompt preceded which work — no timestamps, no gap threshold, nothing to tune. A shell-history tool has neither half.
Surface mapping¶
| Skill++ concept | Claude Code primitive |
|---|---|
| Trace capture | PostToolUse / PreToolUse hooks |
| Intent capture | UserPromptSubmit hook |
| Task boundaries | UserPromptSubmit — recorded in the step stream, so a prompt's position marks where one task ends and the next begins |
| Segmentation + ledger write + review nudge | SessionEnd hook — stamps the session and spawns skill-plus-plus fold-session, which does the work detached |
| Banking what an earlier session left behind | SessionStart hook → skill-plus-plus fold-pending |
| Pull-based review UI | .claude/commands/skill-plus-plus-review.md → /skill-plus-plus-review |
| Skill output | .claude/skills/<name>/SKILL.md + scripts/ |
| Dependency check at pull | Diff declared deps against .mcp.json and connected mcp__<server>__<tool> names |
| Progressive disclosure | Native — name + description indexed, body loaded on demand |
Two constraints to design around¶
Capture is terminal-only. Hooks fire in Claude Code CLI sessions; the desktop
and web chat surfaces run their own runtimes that don't read the host's
~/.claude/settings.json, so they produce no ledger entries. This is a hard
boundary (see usage.md, What is captured), not a configuration matter. Measured
directly: a chat session produced no buffer, no error, and no log entry — the
hook was never invoked at all.
There is no built-in cold tier. Claude Code indexes everything under the skills directory, so the hot/cold/archived model in §6 is implemented by physically moving files to a sibling directory (.claude/skill-plus-plus/cold/) with a retrieval skill that searches it. Demotion is a file move, not a flag.
Packaging & Distribution¶
Two formats, two use cases:
| Format | Use | How |
|---|---|---|
| Upload ZIP | Claude Desktop | skill-plus-plus bundle --format upload, then Customize → Skills |
| Plugin | Claude Code terminal, team (Phase 2) | skill-plus-plus bundle --format plugin |
A plugin bundles hooks, the review command, and retrieval skill. It's also the
upgrade path for MCP-dependent skills (§5) — plugins can declare mcpServers,
whereas a bare SKILL.md cannot.
Hook names and payload shapes should be confirmed against current documentation before building. That surface evolves faster than the skill format does.
9. Key Differentiators¶
| Metric | Dust.tt | Superpowers | IDE-native memory (Cursor, Copilot, Claude Code) | Skill++ |
|---|---|---|---|---|
| Primary focus | Team knowledge RAG | Engineering process rules (TDD, planning) | Per-developer context recall | Operational workflow capture |
| Creation effort | High (manual prompting) | Manual (maintainer-authored) | Low, but per-session and personal | Passive capture, deliberate promotion |
| Input source | Web forms & docs | Static GitHub repo | Chat transcripts | Live traces + direct dictation |
| Retrospective search | Documents only | N/A | Limited | Full searchable work ledger |
| Security handling | Space-level ACLs | N/A | Varies by vendor | Sanitized on write, before disk |
| Skill structure | Flat assistant prompts | Flat prompt files | Flat memory entries | Sub-skill composition (DAG) |
| Team distribution | Native | Manual repo sync | Weak / personal by design | One-click PR + dep check at pull |
The competitive pressure worth taking seriously is the fourth column: memory and rule-generation features bundled free with the IDE. Skill++ differentiates on the two things those do not do — a searchable ledger of past work, and team-grade distribution with dependency and lifecycle management.
10. Target Outcomes¶
- Self-building repository: The operational capability library grows as a by-product of daily engineering work.
- Zero context bloat: Tiering plus progressive disclosure keeps agent context windows fast and cheap regardless of library size.
- Nothing unreviewed, nothing lost: Every skill was read and approved by a human; nothing approved is ever silently discarded.
- Portable by default: Generation actively prefers forms that survive the trip to another developer's machine.
11. Open Questions¶
- Does fast review stay real review? The effect-summary and evidence design targets a 20-second review. If approval rates approach 100%, the gate has become a rubber stamp and the design has failed.
- Is N ≥ 3 the right trigger? Recurrence is a weak proxy for value. The searchable ledger hedges this, but the balance between pushed suggestions and pulled searches needs measurement.
- Do developers actually answer the clarifying questions? §4 assumes three pre-filled questions get answered rather than skipped. If the skip rate is high, most skills land permanently provisional with open
## Open questions, and the judgment layer never materializes. - Will infosec approve a background listener? Sanitize-on-write and local-first storage are the mitigations. Hook-based capture (§8) sidesteps this almost entirely by removing the daemon, which is an argument for shipping the Claude Code integration first. This remains the primary enterprise adoption risk for the OS-daemon path.
- Provisional → trusted promotion: Landing skills as hints that earn trust through successful use makes shallow review safe. The promotion threshold is unvalidated.
12. Implementation¶
Python 3.9+, standard library only — no dependencies, because a hook that has to import a third-party package is a hook that breaks somebody's session.
skill_plus_plus/
config.py paths and thresholds, all env-overridable
sanitize.py secret/PII scrubbing, applied on write
segment.py cuts a session into task episodes
normalize.py parameterisation + step shapes
ledger.py candidate entries: markdown body, JSON payload
matching.py same-procedure matching by embedding
similar.py embedded text and folding one entry into another
signals.py gap detection, question generation, effect summaries
capture.py hook handlers (fail-safe: always exit 0)
summary.py review surface, SKILL.md scaffold, dependency check
lifecycle.py hot/cold/archived tiering, staleness, usage tracking
install.py settings.json wiring (dry run by default)
cli.py command dispatch
commands/skill-plus-plus-review.md /skill-plus-plus-review — review captured candidates
commands/skill-plus-plus-new.md /skill-plus-plus-new — build a skill from a description
examples/demo.sh end-to-end walkthrough on a scratch ledger
tests/fixtures/messy_session.py demo.sh's sessions, polluted with unrelated work
tests/test_skill_plus_plus.py 88 tests
Division of labour¶
The CLI does everything deterministic: capture, scrub, deduplicate, detect
gaps, summarise effects, manage tiers. The /skill-plus-plus-review command drives an
agent through everything that needs judgement — resolving what the repository
can answer, asking the developer at most three questions, and writing prose
worth reading. Neither half is useful alone.
Commands¶
| Command | Purpose |
|---|---|
skill-plus-plus install --user\|--project [DIR] |
Wire the hooks, for every project or just this one. Dry run without --apply; --remove takes them back out |
skill-plus-plus doctor |
Whether the hooks are wired, the models answer, and anything is waiting to be banked |
skill-plus-plus hook --event <E> |
Hook entry point; reads JSON on stdin, always exits 0 |
skill-plus-plus fold-session <id> |
Bank one ended session; what the SessionEnd hook spawns |
skill-plus-plus dictate --text "…" |
Create a candidate from a description instead of a trace |
skill-plus-plus review [--all] |
Candidates at or above the recurrence threshold |
skill-plus-plus sift [--apply] |
Ask a local model which candidates are methods rather than one-off jobs; parks the rest. Dry run without --apply. See docs/research/episode-filter.md |
skill-plus-plus reopen <id> |
Undo a sift verdict |
skill-plus-plus merge [--apply] |
Merge candidates that are the same procedure worded differently, by embedding. Dry run without --apply |
skill-plus-plus split <id> --at N |
Split a candidate holding two procedures; the original is kept, not deleted |
skill-plus-plus name <id> --title … --description … |
Give a candidate a task-shaped name; written by the agent during draft |
skill-plus-plus accuracy |
How often the ranker agreed with your own promote/dismiss decisions |
skill-plus-plus keep |
Save the work so far as a candidate, without ending the session |
skill-plus-plus reconcile |
Report promoted skills whose file is gone; reports only, never decides |
skill-plus-plus ignored [--threshold N] |
List parked candidates and how often that work happened anyway |
skill-plus-plus web [--port N] [--no-browser] |
Promote or dismiss candidates recognized ≥ 3 times (dismissed ones can be reinstated), then Draft (runs skill-plus-plus draft --apply, with an optional note on what to look out for) for a promoted one; review finished drafts, revise them through your agent, and download each as a skill folder zip; loopback only, no auth |
skill-plus-plus draft <id> [--note "…"] [--apply] |
Have your own agent write a draft SKILL.md; never installs it. --note tells it what to look out for. Dry run without --apply |
skill-plus-plus revise <id> --instruction "…" [--apply] |
Have your own agent change a draft SKILL.md as instructed, in place; the previous version is kept in .revisions/. Dry run without --apply |
skill-plus-plus show <id> |
Effect summary, evidence, open questions |
skill-plus-plus search <words> |
Search the ledger of your own past work |
skill-plus-plus scaffold <id> --name <n> |
Generate a starting SKILL.md |
skill-plus-plus promote <id> --skill-path <p> |
Mark a candidate promoted |
skill-plus-plus dismiss <id> / expire |
Dismiss one / delete unapproved past TTL |
skill-plus-plus lifecycle / tier <name> <tier> |
Inventory and demotion |
skill-plus-plus check --name <n> |
Dependency check at pull time (exit 2 if missing) |
skill-plus-plus bundle --out <dir> [--format upload\|plugin] |
Package skills: upload = one zip per skill for Customize → Skills; plugin = .claude-plugin/ + skills/ |
skill-plus-plus install [--apply] |
Wire Claude Code hooks; dry run without --apply |
Built¶
Capture with intent, segmentation into task episodes, sanitize-on-write (typed
placeholders that keep step shapes stable), the ledger with embedding dedup and
TTL expiry, all five trace gap signals from §4 with the three-question cap and
duplicate suppression, the dictation path with its completeness check and
threshold bypass, effect-first proposals, scaffolding with declared deps and
## Open questions that close when answered, pull-time dependency checking,
hot/cold/archived demotion, staleness by reference resolution, and usage
tracking driven by observed Skill calls.
Segmentation, measured. tests/fixtures/messy_session.py takes the three
demo.sh deploy sessions, keeps the deploy byte-identical across all three so
it genuinely recurs, and surrounds each occurrence with different unrelated
work. Folded as whole sessions, the deploy scores 0.358–0.475 against the 0.85
threshold and is filed as three unrelated one-offs, each titled after whatever
happened to come first. Segmented, it is recovered as a single candidate at
three occurrences with steps identical to the unpolluted baseline, and
titled after the deploy. Both halves are asserted, because the whole-session
path survives as _fold_steps and is still taken by single-episode sessions —
so the regression is a live test rather than a git archaeology exercise.
Not built¶
Deliberately deferred — see §13 for phasing.
- Team PR sync. Phase 2. Only the receiving half exists today: declared
dependencies and
skill-plus-plus check. - Script extraction. The slash command tells the agent to lift
deterministic pipelines into
scripts/run.sh; the CLI does not do it automatically. - Automatic provisional→trusted promotion. The tier is recorded and readable; usage counts are tracked; nothing promotes on them yet.
- Voice input. Dictation is text-only —
skill-plus-plus dictateand/skill-plus-plus-new. Speech-to-text is somebody else's job; the parser does not care how the words arrive. - The OS-level shell daemon. Capture is Claude Code hooks only, which is the sequencing argued for in §11.
- Episode labelling. Segmentation cuts and titles episodes from their own
prompts, but nothing names the varying parameter —
stagingversusprod— which is what a synthesised skill needs to parameterise. That is judgement, and it belongs with the reviewing agent alongside semantic dedup. - Non-linear segmentation. Cutting is linear, so a task interrupted by a second task and then resumed is mis-attributed. Recorded as a known limitation in the fixture rather than papered over.
Verification status¶
Installed and confirmed firing in a live Claude Code terminal session. All
three hooks verified against real payloads: PostToolUse parses Bash and
Edit calls cleanly, UserPromptSubmit captures prompts verbatim, and
SessionEnd folds a buffer into a ledger entry.
Coverage is narrower than §8 originally claimed: chat-surface sessions are not captured at all. See usage.md.
Two limits worth stating plainly about segmentation:
git commitis the only marker with test coverage. The other eight, and the artifact-delivery path, are implemented but unexercised.- It has never run against a real session. Every result above comes from a fixture written for the purpose. The known risk — a mid-task "continue" or "fix that" cutting an episode in half — is precisely the thing a hand-written fixture cannot demonstrate, since its prompts are one-per-task by construction. Replaying real transcripts is the next thing that could show the design is wrong rather than merely incomplete.
13. Roadmap¶
Phase 1 — single developer (built)¶
Capture, ledger, review, promotion, lifecycle. Everything a developer needs to turn their own work into their own skills, on their own machine. The ledger never leaves the laptop.
Phase 2 — team distribution¶
The receiving half already exists: skills declare their dependencies and
skill-plus-plus check verifies them at pull time. What Phase 2 adds is the sending
half — exporting an approved skill as a pull request against a shared library,
with the lifecycle and dedup machinery extended across a team rather than a
directory.
Two design questions to settle before writing any of it, because both are easier to get right at the start than to retrofit:
What travels, and what stays. A ledger entry holds project paths, session
ids, and the developer's own prompts. None of that belongs in a team
repository. A promoted skill's metadata.provenance currently points at a
local ledger id that a teammate cannot resolve — useful locally, meaningless
after the trip. Phase 2 needs an explicit split between the skill (travels) and
its provenance (stays), rather than letting the current field quietly leak
context into a PR.
Whether provisional skills should travel at all. Tiering assumes trust is
earned through successful use (§6). But use by whom? A skill that one developer
has run twice is not validated for a team, and shipping tier: provisional
into a shared library either means nothing or means "do not rely on this" —
which is not obviously a thing worth distributing. The plausible rule is that
only trusted skills open a PR, which makes automatic promotion a Phase 2
prerequisite rather than a nice-to-have.
Team-level dedup is the real work. Lexical similarity is adequate for one person's ledger. Across a team, the same workflow arrives written five different ways, and catching that is a judgement problem — which points at the reviewing agent, not at a similarity threshold.
Later¶
Voice input, the OS-level shell daemon, automatic script extraction.