Author SHA1 Message Date
m3tam3re cc8aa5f2b9 dolt remote info 2026-09-20 20:30:16 +02:00
87 changed files with 7 additions and 6066 deletions
-80
View File
@@ -1,80 +0,0 @@
---
name: beads
description: Use when working in a repository that uses bd or Beads for durable project task tracking, issue dependencies, blocker management, multi-session handoff, or shared work memory. Trigger when the user asks to find ready work, claim or close tasks, create follow-up work, inspect blockers, recover project context, or choose between local planning and persistent project tracking.
---
# Beads
Use Beads as the shared project task system. Local plans, scratch files, and personal memories are useful, but they are not the durable source of truth for project work.
## First Step
Run:
```bash
bd prime
```
If that prints nothing, check whether the repository has an active Beads workspace:
```bash
bd where
```
## Preferred Route
Use the `bd` CLI when shell access is available. It is the most compact and direct Beads interface.
## Core CLI Workflow
1. Find work:
```bash
bd ready
bd list --status=open
bd list --status=in_progress
```
2. Inspect before editing:
```bash
bd show <id>
```
3. Claim work atomically:
```bash
bd update <id> --claim
```
4. Create durable follow-up work when implementation reveals new tasks:
```bash
bd create "Short title" --description="Why this exists and what needs to be done" --type=task --priority=2
```
5. Close completed work:
```bash
bd close <id> --reason="Completed"
```
## What Belongs In Beads
Use Beads for:
- shared project tasks
- blockers and dependencies
- discovered follow-up work
- work that must survive thread reset, compaction, or handoff
- status that another person or agent should be able to resume
Use agent-local planning tools only for the current turn's execution checklist. Do not treat them as shared project state.
## Rules
- Do not create markdown TODO files as the source of truth when Beads is available.
- Do not use `bd edit`; it opens an interactive editor. Use `bd update` flags instead.
- Prefer `--json` when parsing `bd` output programmatically.
- If hooks are installed, `bd prime` may already be injected. Run it manually when context is missing.
- Do not auto-close or mutate tasks unless the work is actually complete.
-4
View File
@@ -1,4 +0,0 @@
interface:
display_name: "Beads"
short_description: "Project task tracking with bd"
default_prompt: "Use $beads to inspect ready work and manage durable project tasks."
-77
View File
@@ -1,77 +0,0 @@
# Dolt database (managed by Dolt, not git)
dolt/
embeddeddolt/
proxieddb/
# Runtime files
bd.sock
bd.sock.startlock
sync-state.json
last-touched
.exclusive-lock
# Daemon runtime (lock, log, pid)
daemon.*
# Push state (runtime, per-machine)
push-state.json
# Lock files (various runtime locks)
*.lock
# Credential key (encryption key for federation peer auth — never commit)
.beads-credential-key
# Local version tracking (prevents upgrade notification spam after git ops)
.local_version
proxied_server_client_info.json
# Worktree redirect file (contains relative path to main repo's .beads/)
# Must not be committed as paths would be wrong in other clones
redirect
# Sync state (local-only, per-machine)
# These files are machine-specific and should not be shared across clones
.sync.lock
export-state/
export-state.json
last_pull
# Ephemeral store (SQLite - wisps/molecules, intentionally not versioned)
ephemeral.sqlite3
ephemeral.sqlite3-journal
ephemeral.sqlite3-wal
ephemeral.sqlite3-shm
# Dolt server management (auto-started by bd)
dolt-server.pid
dolt-server.log
dolt-server.lock
dolt-server.port
dolt-server.activity
# Debug-mode pprof artifacts (written when dolt.debug: true in config.yaml)
dolt-pprof/
# Corrupt backup directories (created by bd doctor --fix recovery)
*.corrupt.backup/
# Backup data (auto-exported JSONL, local-only)
backup/
# Per-project environment file (Dolt connection config, GH#2520)
.env
# Legacy files (from pre-Dolt versions)
*.db
*.db?*
*.db-journal
*.db-wal
*.db-shm
db.sqlite
bd.db
# NOTE: Do NOT add negation patterns here.
# They would override fork protection in .git/info/exclude.
# Config files (metadata.json, config.yaml) are tracked by git by default
# since no pattern above ignores them.
-81
View File
@@ -1,81 +0,0 @@
# Beads - AI-Native Issue Tracking
Welcome to Beads! This repository uses **Beads** for issue tracking - a modern, AI-native tool designed to live directly in your codebase alongside your code.
## What is Beads?
Beads is issue tracking that lives in your repo, making it perfect for AI coding agents and developers who want their issues close to their code. No web UI required - everything works through the CLI and integrates seamlessly with git.
**Learn more:** [github.com/steveyegge/beads](https://github.com/steveyegge/beads)
## Quick Start
### Essential Commands
```bash
# Create new issues
bd create "Add user authentication"
# View all issues
bd list
# View issue details
bd show <issue-id>
# Update issue status
bd update <issue-id> --claim
bd update <issue-id> --status done
# Sync with Dolt remote
bd dolt push
```
### Working with Issues
Issues in Beads are:
- **Git-native**: Stored in Dolt database with version control and branching
- **AI-friendly**: CLI-first design works perfectly with AI coding agents
- **Branch-aware**: Issues can follow your branch workflow
- **Sync-ready**: Uses Dolt remotes for backup and team sharing
## Why Beads?
✨ **AI-Native Design**
- Built specifically for AI-assisted development workflows
- CLI-first interface works seamlessly with AI coding agents
- No context switching to web UIs
🚀 **Developer Focused**
- Issues live in your repo, right next to your code
- Works offline, syncs when you push
- Fast, lightweight, and stays out of your way
🔧 **Git Integration**
- Dolt-native sync via bd dolt push / bd dolt pull
- Branch-aware issue tracking
- Dolt-native three-way merge resolution
## Get Started with Beads
Try Beads in your own projects:
```bash
# Install Beads
curl -sSL https://raw.githubusercontent.com/steveyegge/beads/main/scripts/install.sh | bash
# Initialize in your repo
bd init
# Create your first issue
bd create "Try out Beads"
```
## Learn More
- **Documentation**: [github.com/steveyegge/beads/docs](https://github.com/steveyegge/beads/tree/main/docs)
- **Quick Start Guide**: Run `bd quickstart`
- **Examples**: [github.com/steveyegge/beads/examples](https://github.com/steveyegge/beads/tree/main/examples)
---
*Beads: Issue tracking that moves at the speed of thought* ⚡
-70
View File
@@ -1,70 +0,0 @@
# Beads Configuration File
# This file configures default behavior for all bd commands in this repository
# All settings can also be set via environment variables (BD_* prefix)
# or overridden with command-line flags
# Issue prefix for this repository (used by bd init)
# If not set, bd init will auto-detect from directory name
# Example: issue-prefix: "myproject" creates issues like "myproject-1", "myproject-2", etc.
# issue-prefix: ""
# Use no-db mode: JSONL-only, no Dolt database
# When true, .beads/issues.jsonl is the only local store
# no-db: false
# Enable JSON output by default
# json: false
# Feedback title formatting for mutating commands (create/update/close/dep/edit)
# 0 = hide titles, N > 0 = truncate to N characters
# output:
# title-length: 255
# Default actor for audit trails (overridden by BEADS_ACTOR or --actor)
# actor: ""
# Export events (audit trail) to .beads/events.jsonl on each flush/sync
# When enabled, new events are appended incrementally using a high-water mark.
# Use 'bd export --events' to trigger manually regardless of this setting.
# events-export: false
# Multi-repo configuration (experimental - bd-307)
# Allows hydrating from multiple repositories and routing writes to the correct database
# repos:
# primary: "." # Primary repo (where this database lives)
# additional: # Additional repos to hydrate from (read-only)
# - ~/beads-planning # Personal planning repo
# - ~/work-planning # Work planning repo
# Dolt-native backup (periodic backup for off-machine recovery)
# This is full database backup only. Cross-machine sync uses Dolt remotes.
# backup:
# enabled: false # Disable auto-backup entirely
# interval: 15m # Minimum time between auto-backups
# git-push: false # Disable git push (backup locally only)
# git-repo: "" # Separate git repo for backups (default: project repo)
# Optional JSONL auto-export for viewers, interchange, and issue-level migration.
# Disabled by default; enable only when an integration needs fresh .beads/issues.jsonl.
# Use relative paths under .beads/ for JSONL import/export filenames.
# export:
# auto: false
# path: issues.jsonl
# interval: 60s
# git-add: false
# import:
# path: issues.jsonl
# Integration settings (access with 'bd config get/set')
# Non-secret keys (stored in the database):
# - jira.url, jira.project
# - linear.team_id
# - github.org, github.repo
#
# Secret keys (stored in this file but prefer env vars to avoid git exposure):
# - linear.api_key → use LINEAR_API_KEY env var instead
# - github.token → use GITHUB_TOKEN env var instead
sync.remote: "git+ssh://gitea@git.az-gruppe.com/AZ-Intec-GmbH/az-agent-defaults.git"
export:
auto: true
-33
View File
@@ -1,33 +0,0 @@
#!/usr/bin/env sh
# --- BEGIN BEADS INTEGRATION v1.2.2 ---
# This section is managed by beads. Do not remove these markers.
if command -v bd >/dev/null 2>&1; then
export BD_GIT_HOOK=1
_bd_timeout=${BEADS_HOOK_TIMEOUT:-300}
_bd_used_perl=0
if command -v timeout >/dev/null 2>&1; then
timeout "$_bd_timeout" bd hooks run post-checkout "$@"
_bd_exit=$?
elif command -v gtimeout >/dev/null 2>&1; then
gtimeout "$_bd_timeout" bd hooks run post-checkout "$@"
_bd_exit=$?
elif command -v perl >/dev/null 2>&1; then
_bd_used_perl=1
perl -e 'alarm shift; exec @ARGV' "$_bd_timeout" bd hooks run post-checkout "$@"
_bd_exit=$?
else
echo >&2 "beads: hook 'post-checkout' running without timeout; install coreutils or perl to enable BEADS_HOOK_TIMEOUT"
bd hooks run post-checkout "$@"
_bd_exit=$?
fi
if [ $_bd_exit -eq 124 ] || { [ $_bd_used_perl -eq 1 ] && [ $_bd_exit -eq 142 ]; }; then
echo >&2 "beads: hook 'post-checkout' timed out after ${_bd_timeout}s — continuing without beads"
_bd_exit=0
fi
if [ $_bd_exit -eq 3 ]; then
echo >&2 "beads: database not initialized — skipping hook 'post-checkout'"
_bd_exit=0
fi
if [ $_bd_exit -ne 0 ]; then exit $_bd_exit; fi
fi
# --- END BEADS INTEGRATION v1.2.2 ---
-33
View File
@@ -1,33 +0,0 @@
#!/usr/bin/env sh
# --- BEGIN BEADS INTEGRATION v1.2.2 ---
# This section is managed by beads. Do not remove these markers.
if command -v bd >/dev/null 2>&1; then
export BD_GIT_HOOK=1
_bd_timeout=${BEADS_HOOK_TIMEOUT:-300}
_bd_used_perl=0
if command -v timeout >/dev/null 2>&1; then
timeout "$_bd_timeout" bd hooks run post-merge "$@"
_bd_exit=$?
elif command -v gtimeout >/dev/null 2>&1; then
gtimeout "$_bd_timeout" bd hooks run post-merge "$@"
_bd_exit=$?
elif command -v perl >/dev/null 2>&1; then
_bd_used_perl=1
perl -e 'alarm shift; exec @ARGV' "$_bd_timeout" bd hooks run post-merge "$@"
_bd_exit=$?
else
echo >&2 "beads: hook 'post-merge' running without timeout; install coreutils or perl to enable BEADS_HOOK_TIMEOUT"
bd hooks run post-merge "$@"
_bd_exit=$?
fi
if [ $_bd_exit -eq 124 ] || { [ $_bd_used_perl -eq 1 ] && [ $_bd_exit -eq 142 ]; }; then
echo >&2 "beads: hook 'post-merge' timed out after ${_bd_timeout}s — continuing without beads"
_bd_exit=0
fi
if [ $_bd_exit -eq 3 ]; then
echo >&2 "beads: database not initialized — skipping hook 'post-merge'"
_bd_exit=0
fi
if [ $_bd_exit -ne 0 ]; then exit $_bd_exit; fi
fi
# --- END BEADS INTEGRATION v1.2.2 ---
-33
View File
@@ -1,33 +0,0 @@
#!/usr/bin/env sh
# --- BEGIN BEADS INTEGRATION v1.2.2 ---
# This section is managed by beads. Do not remove these markers.
if command -v bd >/dev/null 2>&1; then
export BD_GIT_HOOK=1
_bd_timeout=${BEADS_HOOK_TIMEOUT:-300}
_bd_used_perl=0
if command -v timeout >/dev/null 2>&1; then
timeout "$_bd_timeout" bd hooks run pre-commit "$@"
_bd_exit=$?
elif command -v gtimeout >/dev/null 2>&1; then
gtimeout "$_bd_timeout" bd hooks run pre-commit "$@"
_bd_exit=$?
elif command -v perl >/dev/null 2>&1; then
_bd_used_perl=1
perl -e 'alarm shift; exec @ARGV' "$_bd_timeout" bd hooks run pre-commit "$@"
_bd_exit=$?
else
echo >&2 "beads: hook 'pre-commit' running without timeout; install coreutils or perl to enable BEADS_HOOK_TIMEOUT"
bd hooks run pre-commit "$@"
_bd_exit=$?
fi
if [ $_bd_exit -eq 124 ] || { [ $_bd_used_perl -eq 1 ] && [ $_bd_exit -eq 142 ]; }; then
echo >&2 "beads: hook 'pre-commit' timed out after ${_bd_timeout}s — continuing without beads"
_bd_exit=0
fi
if [ $_bd_exit -eq 3 ]; then
echo >&2 "beads: database not initialized — skipping hook 'pre-commit'"
_bd_exit=0
fi
if [ $_bd_exit -ne 0 ]; then exit $_bd_exit; fi
fi
# --- END BEADS INTEGRATION v1.2.2 ---
-33
View File
@@ -1,33 +0,0 @@
#!/usr/bin/env sh
# --- BEGIN BEADS INTEGRATION v1.2.2 ---
# This section is managed by beads. Do not remove these markers.
if command -v bd >/dev/null 2>&1; then
export BD_GIT_HOOK=1
_bd_timeout=${BEADS_HOOK_TIMEOUT:-300}
_bd_used_perl=0
if command -v timeout >/dev/null 2>&1; then
timeout "$_bd_timeout" bd hooks run pre-push "$@"
_bd_exit=$?
elif command -v gtimeout >/dev/null 2>&1; then
gtimeout "$_bd_timeout" bd hooks run pre-push "$@"
_bd_exit=$?
elif command -v perl >/dev/null 2>&1; then
_bd_used_perl=1
perl -e 'alarm shift; exec @ARGV' "$_bd_timeout" bd hooks run pre-push "$@"
_bd_exit=$?
else
echo >&2 "beads: hook 'pre-push' running without timeout; install coreutils or perl to enable BEADS_HOOK_TIMEOUT"
bd hooks run pre-push "$@"
_bd_exit=$?
fi
if [ $_bd_exit -eq 124 ] || { [ $_bd_used_perl -eq 1 ] && [ $_bd_exit -eq 142 ]; }; then
echo >&2 "beads: hook 'pre-push' timed out after ${_bd_timeout}s — continuing without beads"
_bd_exit=0
fi
if [ $_bd_exit -eq 3 ]; then
echo >&2 "beads: database not initialized — skipping hook 'pre-push'"
_bd_exit=0
fi
if [ $_bd_exit -ne 0 ]; then exit $_bd_exit; fi
fi
# --- END BEADS INTEGRATION v1.2.2 ---
-33
View File
@@ -1,33 +0,0 @@
#!/usr/bin/env sh
# --- BEGIN BEADS INTEGRATION v1.2.2 ---
# This section is managed by beads. Do not remove these markers.
if command -v bd >/dev/null 2>&1; then
export BD_GIT_HOOK=1
_bd_timeout=${BEADS_HOOK_TIMEOUT:-300}
_bd_used_perl=0
if command -v timeout >/dev/null 2>&1; then
timeout "$_bd_timeout" bd hooks run prepare-commit-msg "$@"
_bd_exit=$?
elif command -v gtimeout >/dev/null 2>&1; then
gtimeout "$_bd_timeout" bd hooks run prepare-commit-msg "$@"
_bd_exit=$?
elif command -v perl >/dev/null 2>&1; then
_bd_used_perl=1
perl -e 'alarm shift; exec @ARGV' "$_bd_timeout" bd hooks run prepare-commit-msg "$@"
_bd_exit=$?
else
echo >&2 "beads: hook 'prepare-commit-msg' running without timeout; install coreutils or perl to enable BEADS_HOOK_TIMEOUT"
bd hooks run prepare-commit-msg "$@"
_bd_exit=$?
fi
if [ $_bd_exit -eq 124 ] || { [ $_bd_used_perl -eq 1 ] && [ $_bd_exit -eq 142 ]; }; then
echo >&2 "beads: hook 'prepare-commit-msg' timed out after ${_bd_timeout}s — continuing without beads"
_bd_exit=0
fi
if [ $_bd_exit -eq 3 ]; then
echo >&2 "beads: database not initialized — skipping hook 'prepare-commit-msg'"
_bd_exit=0
fi
if [ $_bd_exit -ne 0 ]; then exit $_bd_exit; fi
fi
# --- END BEADS INTEGRATION v1.2.2 ---
-6
View File
@@ -1,6 +0,0 @@
{"id":"int-97b8d2f545bb01e5a71868be1b46d3e6","kind":"field_change","created_at":"2026-08-22T08:02:44.467228132Z","actor":"m3tam3re","issue_id":"az-agent-defaults-j35","extra":{"field":"status","new_value":"closed","old_value":"in_progress","reason":"Closed"}}
{"id":"int-1aeab5fae03f5d2b4cb699d7c50ee4d6","kind":"field_change","created_at":"2026-08-22T08:09:02.302617707Z","actor":"m3tam3re","issue_id":"az-agent-defaults-dzx","extra":{"field":"status","new_value":"closed","old_value":"in_progress","reason":"Skills-Slice umgesetzt: ow-hello aus az-agent-skills nach skills/ migriert (einzige Abweichung: Repo-Erwähnung L15 → az-agent-defaults); 5 ungültige Skill-Fixtures unter tests/fixtures/skills/invalid/ (ordner-ohne-skill-md, ohne-frontmatter, leeres-frontmatter, name-ungleich-ordner, ohne-description) — decken alle im README dokumentierten Skill-Guard-Fehlertypen ab, jede Fixture failt exakt eine Regel; alle 4 Akzeptanzkriterien per Skript verifiziert (15/15 PASS); .gitkeep-Platzhalter entfernt."}}
{"id":"int-ac8b30c7c6b2fe206a89b0c552ae062e","kind":"field_change","created_at":"2026-08-22T08:13:39.665260198Z","actor":"m3tam3re","issue_id":"az-agent-defaults-1ub","extra":{"field":"status","new_value":"closed","old_value":"in_progress","reason":"ow-hello entfernt; az-hilfe als Onboarding-/Hilfe-Skill (Entwurf) angelegt — Router-Prinzip wie ask-matt: vier Artefakt-Typen erklärt, typische Anfragen mit Antwortmustern, Beitragsweg (IT/Repo), Problemeskalation; ehrlicher Ausbaustand (Commands/Agents/MCP als in Vorbereitung). Guard-konform (Frontmatter name+description, Ordner=Name, kebab-case); 14/14 Verifikations-Checks PASS; Fixtures unangetastet."}}
{"id":"int-54c0cf201b8798d92c63b4f62961a35c","kind":"field_change","created_at":"2026-08-22T08:21:22.416435535Z","actor":"m3tam3re","issue_id":"az-agent-defaults-gj1","extra":{"field":"status","new_value":"closed","old_value":"in_progress","reason":"Referenz-Command commands/az-hilfe.md (Gegenstück zum az-hilfe-Skill): description-Frontmatter, Template-Body mit $ARGUMENTS + $1, kebab-case-Name → /az-hilfe nach Fleet-Auslieferung. 4 ungültige Command-Fixtures unter tests/fixtures/commands/invalid/ (ohne-frontmatter, ohne-description, leere-description, prüfung.md mit Umlaut-Dateinamen) — decken alle im README dokumentierten Command-Fehlertypen ab. 17/17 Verifikations-Checks PASS."}}
{"id":"int-b9f609bf4e7919bc55e05a91fc48031a","kind":"field_change","created_at":"2026-08-22T08:26:44.636891703Z","actor":"m3tam3re","issue_id":"az-agent-defaults-gyj","extra":{"field":"status","new_value":"closed","old_value":"in_progress","reason":"Zwei Referenz-Agenten in agents/: az-beitrag (mode: primary, ohne tools-Restriction — Delegationsfähigkeit intakt, Body betont Managed-Layer-Grenze, delegiert aktiv an az-pruefer) + az-pruefer (mode: subagent, Read-only-Profil tools: [read, glob, grep] ohne Edit/Bash, per @mention/Task-Tool erreichbar). 3 ungültige Agent-Fixtures (ohne-mode, ohne-description, ungueltiger-mode) — alle README-dokumentierten Fehlertypen. 19/19 Checks PASS (1 Regex-Bug im Prüfscript korrigiert nachgewiesen)."}}
{"id":"int-e238d838ca96426519d9763427e5ce30","kind":"field_change","created_at":"2026-08-22T08:26:44.863883113Z","actor":"m3tam3re","issue_id":"az-agent-defaults-95y","extra":{"field":"status","new_value":"closed","old_value":"in_progress","reason":"README um Vault-Key-Namensschema (<server-name>-<verwendungszweck>, kebab-case) + Key-Ausnahme-Beispiel ergänzt (Format selbst war seit Scaffold dokumentiert). Zwei Referenz-Fragmente in mcp/: zugferd-service.yaml (OAuth-Standardfall, kein Credential-Feld) + az-zoll-service.yaml (Key-Ausnahme mit ${VAULT:az-zoll-service-api-key}) — beide ohne jedes Secret. 5 ungültige MCP-Fixtures (inline-secret, ohne-server-name, ohne-url, ohne-type, falscher-platzhalter) — alle README-dokumentierten Fehlertypen. 17/17 Checks PASS."}}
-12
View File
@@ -1,12 +0,0 @@
{"_type":"issue","id":"az-agent-defaults-j35","title":"Repo-Gerüst: vier Artefakt-Verzeichnisse + README mit Guard-Regeln je Artefakt-Typ","description":"Quelle: Spec 01-az-agent-defaults-spec.md (lokal, wird nicht gepusht).\n\naz-agent-defaults wird als reines Content-Repo die eine Wahrheitsquelle für das Company-Default-Set. Dieses Ticket liefert das Gerüst: die vier Artefakt-Verzeichnisse (skills/, commands/, agents/, mcp/) und ein README, das die Contribution-Guard-Regeln je Artefakt-Typ so dokumentiert, dass ein Fachbereichs-Contributor ohne Architektur-Wissen richtig beisteuern kann:\n\n- Skills: Ordner mit SKILL.md und Frontmatter-Öffner (wie heute)\n- Commands: Markdown mit Frontmatter (description, optional agent/model) und Template-Body (Argument-Platzhalter erlaubt)\n- Agents: Markdown mit Frontmatter (description, mode primary/subagent, optional model/temperature, Permission-Profil); Subagent-Definitionen müssen über das Task-Tool aufrufbar bleiben\n- MCP: Fragmente für den mcp-Konfigurationsschlüssel, remote/streamable als Standard, Secrets grundsätzlich nur als Platzhalter\n\nDas README hält außerdem fest: Auslieferung und Guard-Engine (Pre-Flight auf dem Controller) leben im Fleet-Repo az-fleet; dieses Repo ist Content-only und über den Repo-Ref pinbar (Default: main). Zudem die Fixtures-Konvention: Testgegenstände für die Guard-Tests leben unter tests/fixtures/ und werden nie ausgeliefert, weil die Delivery-Rolle nur die vier Typ-Verzeichnisse liest (Sonntags-Entscheidung). Das Repo verabschiedet damit den Namen az-agent-skills (User Story 2).","acceptance_criteria":"- Die vier Artefakt-Verzeichnisse skills/, commands/, agents/, mcp/ sind im Repo angelegt\n- README dokumentiert die Guard-Regeln je Artefakt-Typ vollständig (Struktur-/Frontmatter-/Platzhalter-Anforderungen gemäß Spec B1)\n- README erklärt die Content/Mechanismus-Trennung zu az-fleet (Guard-Engine und Auslieferung dort) und das Ref-Pinning\n- Die Fixtures-Konvention tests/fixtures/ ist dokumentiert (nie ausgeliefert; Begründung: Delivery liest nur die vier Typ-Verzeichnisse)\n- Ein Contributor ohne Architektur-Wissen kann anhand des README allein entscheiden, wohin ein neues Artefakt gehört und ob es die Guard-Regeln erfüllt","status":"closed","priority":1,"issue_type":"task","assignee":"m3tam3re","owner":"p@m3ta.dev","created_at":"2026-08-22T07:57:14Z","created_by":"m3tam3re","updated_at":"2026-08-22T08:02:44Z","started_at":"2026-08-22T07:59:36Z","closed_at":"2026-08-22T08:02:44Z","close_reason":"Closed","labels":["ready-for-agent"],"dependency_count":0,"dependent_count":4,"comment_count":0}
{"_type":"issue","id":"az-agent-defaults-jtp","title":"Dogfooding + Smoke-Tests der perfektionierten Agents","description":"Abschlussvalidierung: Der perfektionierte az-pruefer läuft über alle 5 perfektionierten Agenten (Dogfooding — der Prüfer prüft als Erstes die perfektionierten Kollegen); Befunde werden eingearbeitet bis der Prüf-Report grün ist. Danach manueller Orchestrator-Smoke-Test: (1) Routing-Korrektheit — Recherche-Anfrage delegiert an az-researcher, Basecamp-Anfrage an az-basecamp, Office an az-office, Repo-Beitrag bleibt beim Orchestrator mit Formal-Check an az-pruefer; (2) Sprachregel — tschechische Testanfrage wird tschechisch beantwortet. Testprotokoll als Beleg ablegen (tests/ oder Notes).\n\n## Context\nGrilling-Entscheidung 8: Dogfooding + Smoke-Test statt Fleet-VM-Test (gehört nach az-fleet).","acceptance_criteria":"1) az-pruefer-Report über agents/*.md: alle Befunde behoben, Report grün. 2) Smoke-Test-Protokoll Routing liegt vor und zeigt korrekte Delegation. 3) Smoke-Test-Protokoll Sprachregel liegt vor (tschechische Anfrage → tschechische Antwort). 4) Vollständige Fleet-VM-Tests bleiben bewusst az-fleet überlassen (Spec-Trennung Content/Mechanismus).","status":"closed","priority":2,"issue_type":"task","owner":"m3ta-chiron@agentmail.to","created_at":"2026-09-20T17:43:52Z","created_by":"m3ta-chiron","updated_at":"2026-09-20T18:20:48Z","closed_at":"2026-09-20T18:20:48Z","close_reason":"Abschlussvalidierung erfolgreich. (1) Dogfooding: az-pruefer-Prompt lief ueber alle 5 Agenten: Erstpruefung 0 Guard-Verstoesze, 1 Standard-Verstosz (Orchestrator-Description) -\u003e als Regelkonflikt aufgeloest (README Trigger-Description praezisiert: Subagents=Beispielfragen, Primary=Missionssatz; Sprachregel-Variante gleichbedeutend), Re-Pruefung gruen. (2) Routing-Smoke-Test 5/5 korrekt (Recherche-\u003eaz-researcher, Basecamp-\u003eaz-basecamp inkl. Destruktiv-Hinweis, Office-\u003eaz-office, Repo-Beitrag-\u003eselbst+az-pruefer, cs-Basecamp-Anfrage-\u003eaz-basecamp). (3) Sprachregel: direkte tschechische Anfrage wurde vollstaendig tschechisch beantwortet. Protokoll: tests/protocols/2026-09-20-smoke-test-agents.md. Fleet-VM-Tests bewusst az-fleet ueberlassen.","labels":["ready-for-agent"],"dependencies":[{"issue_id":"az-agent-defaults-jtp","depends_on_id":"az-agent-defaults-74g","type":"blocks","created_at":"2026-09-20T17:43:52Z","created_by":"m3ta-chiron","metadata":"{}"},{"issue_id":"az-agent-defaults-jtp","depends_on_id":"az-agent-defaults-99v","type":"blocks","created_at":"2026-09-20T17:43:52Z","created_by":"m3ta-chiron","metadata":"{}"},{"issue_id":"az-agent-defaults-jtp","depends_on_id":"az-agent-defaults-bfj","type":"blocks","created_at":"2026-09-20T17:43:52Z","created_by":"m3ta-chiron","metadata":"{}"},{"issue_id":"az-agent-defaults-jtp","depends_on_id":"az-agent-defaults-h2j","type":"blocks","created_at":"2026-09-20T17:43:52Z","created_by":"m3ta-chiron","metadata":"{}"},{"issue_id":"az-agent-defaults-jtp","depends_on_id":"az-agent-defaults-m4t","type":"blocks","created_at":"2026-09-20T17:43:52Z","created_by":"m3ta-chiron","metadata":"{}"}],"dependency_count":5,"dependent_count":0,"comment_count":0}
{"_type":"issue","id":"az-agent-defaults-m4t","title":"az-office perfektionieren","description":"az-office nach dem Agenten-Prompt-Standard neu schreiben: Skill-Load „officecli“ als verpflichtender erster Schritt (Formulierung schärfen — aus „lade zu Beginn“ wird eine nicht verhandelbare erste Aktion, ausdrückliches Muster-Referenz des Prometheus-Ansatzes). Ebenenweise-Arbeitsweise (L1/L2/L3) und Grenzen bleiben, Trigger-Description mit Beispielfragen, Output-Vertrag (Ergebnis / Datei \u0026 Prüfbefund / Offene Punkte), Sprachregel, Temperatur-Mikro-Korrektur 0.2 → 0.1.\n\n## Context\nGrilling-Session 20.09. Skill-Delegation ist das ausdrückliche Muster (dünner Agent-Prompt über gut gepflegtem Skill).","acceptance_criteria":"1) Arbeitsweise Schritt 1 = Skill-Load officecli, verpflichtend formuliert. 2) L1/L2/L3-Prinzip erhalten. 3) Description trigger-optimiert. 4) Output-Vertrag mit Datei-\u0026-Prüfbefund-Sektion. 5) Sprachregel enthalten. 6) Demo: docx-Erstellung mit Prüfbefund-Sektion in der Antwort.","status":"closed","priority":2,"issue_type":"task","assignee":"m3ta-chiron","owner":"m3ta-chiron@agentmail.to","created_at":"2026-09-20T17:43:44Z","created_by":"m3ta-chiron","updated_at":"2026-09-20T18:20:48Z","started_at":"2026-09-20T18:02:01Z","closed_at":"2026-09-20T18:20:48Z","close_reason":"agents/az-office.md neu nach Standard: Schritt 1 = nicht verhandelbarer Skill-Load 'officecli' mit ausdruecklicher Prometheus-Muster-Referenz (duenne Agent-Shell ueber gepflegtem Skill); L1/L2/L3-Prinzip, Pruefschritt (outline/issues/validate) und Grenzen bleiben; Trigger-Description mit 3 Beispielauftraegen in den ersten 80 Zeichen; Output-Vertrag Ergebnis/Datei \u0026 Pruefbefund (ohne Pruefbefund nicht abgeschlossen)/Offene Punkte; Sprachregel; temperature 0.2 -\u003e 0.1. Nach Kosmetik-Fix (bearbeite sie) az-pruefer-Re-Pruefung gruen.","labels":["ready-for-agent"],"dependency_count":0,"dependent_count":1,"comment_count":0}
{"_type":"issue","id":"az-agent-defaults-bfj","title":"az-basecamp perfektionieren + Skill-Nutzung verankern","description":"az-basecamp nach Prometheus-Muster neu schreiben (dünne Shell über Skill): Verpflichtender erster Schritt = Skill „basecamp“ laden (Skill wird via az-fleet aus external/ auf die Nutzerebene ausgerollt — Lieferweg ist gesichert). Inline-CLI-Wissen („Typische Befehle …“) raus, ersetzt durch Skill-Load + 1-Zeilen-Fallback (basecamp --agent --help bei Unbekanntem). Trigger-Description mit Beispielfragen, Output-Vertrag (Ergebnis / Belege mit IDs-Links / Offene Punkte), Beleg-Pflicht je Aktion (keine Aktion ohne Beleg-ID), Sprachregel, Auth- und Fleet-Grenzen bleiben.\n\n## Context\nGrilling-Session 20.09.: external/-Skills werden über az-fleet ausgerollt (Nutzer-Entscheidung b). Muster-Vorbild: az-office + officecli-Skill.","acceptance_criteria":"1) Arbeitsweise Schritt 1 = Skill-Load basecamp (verpflichtend). 2) Keine inline-CLI-Befehlsliste mehr im Prompt. 3) Description trigger-optimiert. 4) Output-Vertrag mit Belege-Sektion. 5) Sprachregel enthalten. 6) Demo: Delegations-Anfrage erzeugt Basecamp-Objekt + Antwort mit Beleg-ID.","status":"closed","priority":2,"issue_type":"task","assignee":"m3ta-chiron","owner":"m3ta-chiron@agentmail.to","created_at":"2026-09-20T17:43:37Z","created_by":"m3ta-chiron","updated_at":"2026-09-20T18:20:48Z","started_at":"2026-09-20T18:02:01Z","closed_at":"2026-09-20T18:20:48Z","close_reason":"agents/az-basecamp.md neu als duenne Shell ueber dem basecamp-Skill (Prometheus-Muster): Schritt 1 = verpflichtender Skill-Load 'basecamp' (BEVOR irgendetwas anderes passiert), Inline-CLI-Befehlsliste entfernt, 1-Zeilen-Fallback basecamp --agent --help. Trigger-Description mit 3 Beispielfragen in den ersten 80 Zeichen; Output-Vertrag Ergebnis/Belege (Beleg-Pflicht: keine Aktion ohne Beleg-ID)/Offene Punkte; Sprachregel; Auth- (OAuth-Login interaktiv) und Fleet-Grenzen bleiben; haiku/0.2 bleibt. az-pruefer: konform ohne Befund.","labels":["ready-for-agent"],"dependency_count":0,"dependent_count":1,"comment_count":0}
{"_type":"issue","id":"az-agent-defaults-99v","title":"az-researcher perfektionieren","description":"az-researcher neu schreiben nach dem Agenten-Prompt-Standard: Modell az-litellm/claude-haiku-4-5 → az-litellm/claude-sonnet-5 (Recherchequalität: Quellenbewertung/Synthese braucht das stärkere Modell, Fehlerkosten unsichtbar und teuer). Trigger-Description mit Beispielfragen, Klassifikation Web-vs-lokal als erster Arbeitschritt, geschärfter Output-Vertrag (Kernantwort / Belege mit URL bzw. Datei:Zeile / Offene Punkte), Sprachregel, Failure-Bedingung: Behauptung ohne Quelle = gescheitert.\n\n## Context\nGrilling-Entscheidung 6(a): Researcher auf Sonnet-Klasse, alle anderen Modelle bleiben.","acceptance_criteria":"1) Frontmatter: model=az-litellm/claude-sonnet-5, temperature 0.1–0.2. 2) Description trigger-optimiert mit Beispielfragen. 3) Arbeitsweise beginnt mit Web-vs-lokal-Klassifikation. 4) Output-Vertrag mit Belege- und Offene-Punkte-Sektion. 5) Sprachregel enthalten. 6) Demo: @mention-Recherche liefert sektionierte Antwort mit Quellen.","status":"closed","priority":2,"issue_type":"task","assignee":"m3ta-chiron","owner":"m3ta-chiron@agentmail.to","created_at":"2026-09-20T17:43:29Z","created_by":"m3ta-chiron","updated_at":"2026-09-20T18:20:48Z","started_at":"2026-09-20T18:02:01Z","closed_at":"2026-09-20T18:20:48Z","close_reason":"agents/az-researcher.md neu nach Standard: model az-litellm/claude-sonnet-5, temperature 0.1, read-only-Tools bleiben. Trigger-Description mit 2 konkreten Beispielfragen in den ersten 80 Zeichen; Arbeitsweise Schritt 1 = Web-vs-lokal-Klassifikation mit Abschlusskriterium; Output-Vertrag Kernantwort/Belege (URL bzw. Datei:Zeile)/Offene Punkte plus explizite Failure-Bedingung (Behauptung ohne Quelle = gescheiterter Auftrag); Sprachregel; Body ~2.900 Zeichen. az-pruefer: konform ohne Befund.","labels":["ready-for-agent"],"dependency_count":0,"dependent_count":1,"comment_count":0}
{"_type":"issue","id":"az-agent-defaults-h2j","title":"az-orchestrator perfektionieren","description":"az-orchestrator als schlanker Router + Verifikator neu schreiben (~150 Zeilen, bleibt glm-5-2): geschärfte Routing-Tabelle mit Trigger-Beispielen je Subagent, Verifikations-Checkliste gegen die Output-Verträge der Subagents (Belege fehlen → nachbessern lassen; Offene Punkte die der Nutzer nie fragte → Rückfrage), Eskalations-Regeln, Sprachregel „Antworte in der Sprache der Nutzeranfrage“, „selber machen wenn schneller“ bleibt. Kein Certainty-/Plan-Zwang (bewusst gegen OMO-Voll-Doktrin entschieden — Kostenarchitektur glm-5-2).\n\n## Context\nGrilling-Entscheidung 5(a): schlanker Router + Verifikator. Zielgruppe: ganze AZ-Gruppe (deutsch/tschechisch/englisch).","acceptance_criteria":"1) Prompt folgt dem Agenten-Prompt-Standard (Skelett, Sprachregel, \u003c10k Zeichen). 2) Description trigger-optimiert, erste 80 Zeichen = Missionssatz. 3) Routing-Tabelle nennt Beispielfragen. 4) Verifikations-Checkliste referenziert die Subagent-Output-Verträge. 5) Smoke-Test: Recherche-Anfrage delegiert an az-researcher; tschechische Anfrage wird tschechisch beantwortet.","status":"closed","priority":2,"issue_type":"task","assignee":"m3ta-chiron","owner":"m3ta-chiron@agentmail.to","created_at":"2026-09-20T17:43:21Z","created_by":"m3ta-chiron","updated_at":"2026-09-20T18:20:47Z","started_at":"2026-09-20T18:02:00Z","closed_at":"2026-09-20T18:20:47Z","close_reason":"agents/az-orchestrator.md neu als schlanker Router + Verifikator (108 Zeilen, Body ~7.000 Zeichen, glm-5-2/0.3 bleibt): Missionssatz-Description (erste 80 Zeichen), Routing-Tabelle mit je 3-4 konkreten Beispielfragen je Zeile (de/en/cs), Delegations-Regeln inkl. 'selber machen wenn schneller' + 'ein Auftrag, ein Verantwortlicher', Verifikations-Checkliste gegen die Subagent-Output-Vertraege (Ergebnis/Belege/Offene Punkte; Belege fehlen -\u003e nachbessern, nie gefragte offene Punkte -\u003e Rueckfrage), Eskalations-Regeln in Grenzen, kein Certainty-/Plan-Zwang. Smoke-Tests: Routing 5/5 korrekt, tschechische Anfrage tschechisch beantwortet (Protokoll tests/protocols/2026-09-20-smoke-test-agents.md). az-pruefer-Dogfooding: nach README-Praezisierung (Primary=Missionssatz) gruen.","labels":["ready-for-agent"],"dependency_count":0,"dependent_count":1,"comment_count":0}
{"_type":"issue","id":"az-agent-defaults-74g","title":"Agenten-Prompt-Standard definieren + az-pruefer perfektionieren","description":"README-Sektion „Agenten-Prompt-Standard“ anlegen (Prompt-Skelett mit Arbeitsweise/Ausgabeformat/Grenzen, Trigger-Description-Regel mit Beispielfragen in den ersten 80 Zeichen, Output-Vertrags-Pflicht für Subagents mit Belege- und Offene-Punkte-Sektion, Sprachregel „Antworte in der Sprache der Anfrage“ in jedem Agenten, 10k-Zeichen-Grenze) plus einen Satz zum external/-Lieferweg (Skills aus external/ werden via az-fleet auf die Nutzerebene ausgerollt). Danach az-pruefer nach dem neuen Standard perfektionieren: prüft die deterministischen Guard-Regeln aus dem README UND den Agenten-Prompt-Standard (LLM-Prüfung), mit eigenem Output-Vertrag und Trigger-Description. Grundlage ist die OMO-Analyse-Session (Trigger-Descriptions, Output-Verträge, Prometheus-Muster).\n\n## Context\nEntscheidungen aus Grilling-Session 20.09.: Struktur (a) deutsch+OMO-Elemente; Descriptions voll trigger-optimiert; Output-Verträge Markdown-Sektionen; Standard weiche Regel (LLM-geprüft, keine harten Guards).","acceptance_criteria":"1) README enthält die Sektion „Agenten-Prompt-Standard“ mit Skelett, Trigger-Regel, Output-Vertrags-Pflicht, Sprachregel, 10k-Grenze und external/-Hinweis. 2) az-pruefer prüft Guard-Regeln + Prompt-Standard. 3) Demo: az-pruefer meldet an einer absichtlich schlechten Agent-Datei genau die Standard-Verstöße (fehlende Beispielfragen, fehlender Output-Vertrag, fehlende Sprachregel).","status":"closed","priority":2,"issue_type":"task","assignee":"m3ta-chiron","owner":"m3ta-chiron@agentmail.to","created_at":"2026-09-20T17:43:10Z","created_by":"m3ta-chiron","updated_at":"2026-09-20T17:58:11Z","started_at":"2026-09-20T17:50:54Z","closed_at":"2026-09-20T17:58:11Z","close_reason":"Umgesetzt: (1) README-Sektion 'Agenten-Prompt-Standard' (Prompt-Skelett Arbeitsweise/Ausgabeformat/Grenzen, Trigger-Description-Regel mit Beispielfragen in den ersten 80 Zeichen, Output-Vertrags-Pflicht für Subagents mit Ergebnis/Belege/Offene Punkte, Sprachregel 'Antworte in der Sprache der Anfrage', 10k-Zeichen-Grenze, external/-Lieferweg-Satz) — klar als weiche Regel (LLM-geprüft, kein Guard) abgetrennt; Fixtures-Sektion um agents/demo/ ergänzt. (2) agents/az-pruefer.md neu nach eigenem Standard: Trigger-Description mit Beispielfragen, prüft Guard-Regeln + Prompt-Standard getrennt, eigener Output-Vertrag, Sprachregel, Read-only-Profil bleibt. (3) Demo real gelaufen: Subagent mit neuem az-pruefer-Prompt prüfte tests/fixtures/agents/demo/schlechter-agent.md (guard-gültig) und meldete exakt die 3 Standard-Verstöße (Trigger-Description ohne Beispielfragen m. 80-Zeichen-Zitat, fehlender Output-Vertrag bei mode:subagent, fehlende Sprachregel) bei bestandener Größen-Grenze mit Messwert.","labels":["ready-for-agent"],"dependency_count":0,"dependent_count":1,"comment_count":0}
{"_type":"issue","id":"az-agent-defaults-1ub","title":"Skills: ow-hello entfernen, durch Onboarding-/Hilfe-Skill (az-hilfe, Entwurf) ersetzen","description":"Entscheidung des Nutzers (22.08.2026): ow-hello wird nicht mehr benötigt. Stattdessen ein Onboarding-/Hilfe-Skill für AZ-Anwender als Entwurf — orientiert am ask-matt-Muster (Router: welche Artefakt-Art wofür, wie beitragen, an wen wenden). Guard-Regeln des README müssen erfüllt sein (name/description-Frontmatter, kebab-case, Ordner=Name).","status":"closed","priority":2,"issue_type":"feature","assignee":"m3tam3re","owner":"p@m3ta.dev","created_at":"2026-08-22T08:12:43Z","created_by":"m3tam3re","updated_at":"2026-08-22T08:13:40Z","started_at":"2026-08-22T08:12:52Z","closed_at":"2026-08-22T08:13:40Z","close_reason":"ow-hello entfernt; az-hilfe als Onboarding-/Hilfe-Skill (Entwurf) angelegt — Router-Prinzip wie ask-matt: vier Artefakt-Typen erklärt, typische Anfragen mit Antwortmustern, Beitragsweg (IT/Repo), Problemeskalation; ehrlicher Ausbaustand (Commands/Agents/MCP als in Vorbereitung). Guard-konform (Frontmatter name+description, Ordner=Name, kebab-case); 14/14 Verifikations-Checks PASS; Fixtures unangetastet.","dependency_count":0,"dependent_count":0,"comment_count":0}
{"_type":"issue","id":"az-agent-defaults-95y","title":"MCP-Slice: Fragment-Format (remote/streamable) + Vault-Platzhalter-Konvention + Guard-Fixtures","description":"Quelle: Spec 01-az-agent-defaults-spec.md.\n\nDefinition des MCP-Fragment-Formats für den mcp-Konfigurationsschlüssel: ein Fragment pro Server, remote/streamable als Standard (url, enabled-Flag; OAuth läuft zur Laufzeit pro Nutzer über das Microsoft-Konto — keine statischen Secrets im Fragment). Für Ausnahmen mit statischem Key definiert das Format einen Platzhalter samt Vault-Key-Namensschema: der Key liegt im Vault des Fleet-Repos und wird erst beim Merge auf dem Controller substituiert — im Repo bleibt nie ein Secret (B3).\n\nZwei Referenz-Fragmente in mcp/ (eine OAuth-Server-Anbindung, eine Key-Ausnahme mit Platzhalter) dienen als Testgegenstände für die Fleet-Szenarien „Fragmente in der aufgelösten Konfiguration sichtbar (Sidecar-Debug)\" und „Vault-Substitution funktioniert, Nutzerebene enthält kein Secret aus dem Repo\". Zusätzlich ungültige MCP-Fixtures unter tests/fixtures/ — insbesondere ein Fragment mit Inline-Secret, das der Guard rot abbrechen muss.\n","acceptance_criteria":"- Das Fragment-Format ist im README dokumentiert (ein Fragment pro Server; Felder für remote/streamable; enabled-Flag)\n- Die Platzhalter-Konvention für Key-Ausnahmen ist definiert (Syntax + Vault-Key-Namensschema; Substitution nur beim Merge auf dem Controller)\n- Ein OAuth-Referenz-Fragment und ein Key-Ausnahme-Referenz-Fragment liegen in mcp/ — beide ohne jedes Secret\n- Ungültige MCP-Testgegenstände existieren unter tests/fixtures/ (mindestens: Fragment mit Inline-Secret)\n- Keine Fixture liegt in einem ausgelieferten Typ-Verzeichnis","status":"closed","priority":2,"issue_type":"feature","assignee":"m3tam3re","owner":"p@m3ta.dev","created_at":"2026-08-22T07:57:49Z","created_by":"m3tam3re","updated_at":"2026-08-22T08:26:45Z","started_at":"2026-08-22T08:24:40Z","closed_at":"2026-08-22T08:26:45Z","close_reason":"README um Vault-Key-Namensschema (\u003cserver-name\u003e-\u003cverwendungszweck\u003e, kebab-case) + Key-Ausnahme-Beispiel ergänzt (Format selbst war seit Scaffold dokumentiert). Zwei Referenz-Fragmente in mcp/: zugferd-service.yaml (OAuth-Standardfall, kein Credential-Feld) + az-zoll-service.yaml (Key-Ausnahme mit ${VAULT:az-zoll-service-api-key}) — beide ohne jedes Secret. 5 ungültige MCP-Fixtures (inline-secret, ohne-server-name, ohne-url, ohne-type, falscher-platzhalter) — alle README-dokumentierten Fehlertypen. 17/17 Checks PASS.","labels":["ready-for-agent"],"dependencies":[{"issue_id":"az-agent-defaults-95y","depends_on_id":"az-agent-defaults-j35","type":"blocks","created_at":"2026-08-22T09:57:48Z","created_by":"m3tam3re","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0}
{"_type":"issue","id":"az-agent-defaults-gj1","title":"Commands-Slice: Referenz-Command mit Argument-Platzhaltern + Guard-Fixtures","description":"Quelle: Spec 01-az-agent-defaults-spec.md.\n\nErster exemplarischer Company-Command in commands/: Markdown mit Frontmatter (description, optional agent/model) und Template-Body mit Argument-Platzhaltern ($ARGUMENTS, $1..$n) — Format gemäß verifizierter OpenCode-Doku (global ~/.config/opencode/commands/, pro Projekt .opencode/commands/; Shell-Output-Injection und Datei-Referenzen unterstützt). Der Command ist zugleich Referenz-Beitrag für Contributor und Testgegenstand für das Fleet-Szenario „/command ist in der OpenWork-/OpenCode-Sitzung verfügbar und löst das Template aus\". Inhaltlich bewusst trivial (Platzhalter-Niveau wie ow-hello; konkrete Inhalte sind Phase 2).\n\nZusätzlich ungültige Command-Fixtures unter tests/fixtures/ (z. B. Frontmatter ohne description, Datei ohne Frontmatter).\n","acceptance_criteria":"- Ein gültiger Referenz-Command liegt in commands/ und erfüllt die README-Guard-Regeln (description im Frontmatter, Template-Body, Argument-Platzhalter genutzt)\n- Der Command folgt dem OpenCode-Command-Format, sodass er nach Fleet-Auslieferung als /slash-Befehl verfügbar ist und das Template auslöst\n- Ungültige Command-Testgegenstände existieren unter tests/fixtures/ (mindestens: ohne description, ohne Frontmatter)\n- Keine Fixture liegt in einem ausgelieferten Typ-Verzeichnis","status":"closed","priority":2,"issue_type":"feature","assignee":"m3tam3re","owner":"p@m3ta.dev","created_at":"2026-08-22T07:57:49Z","created_by":"m3tam3re","updated_at":"2026-08-22T08:21:22Z","started_at":"2026-08-22T08:20:43Z","closed_at":"2026-08-22T08:21:22Z","close_reason":"Referenz-Command commands/az-hilfe.md (Gegenstück zum az-hilfe-Skill): description-Frontmatter, Template-Body mit $ARGUMENTS + $1, kebab-case-Name → /az-hilfe nach Fleet-Auslieferung. 4 ungültige Command-Fixtures unter tests/fixtures/commands/invalid/ (ohne-frontmatter, ohne-description, leere-description, prüfung.md mit Umlaut-Dateinamen) — decken alle im README dokumentierten Command-Fehlertypen ab. 17/17 Verifikations-Checks PASS.","labels":["ready-for-agent"],"dependencies":[{"issue_id":"az-agent-defaults-gj1","depends_on_id":"az-agent-defaults-j35","type":"blocks","created_at":"2026-08-22T09:57:48Z","created_by":"m3tam3re","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0}
{"_type":"issue","id":"az-agent-defaults-gyj","title":"Agents-Slice: Referenz-Primäragent + Read-only-Subagent + Guard-Fixtures","description":"Quelle: Spec 01-az-agent-defaults-spec.md.\n\nZwei Referenz-Agent-Definitionen in agents/: ein spezialisierter Primäragent (per Agentenwechsel erreichbar) und ein Subagent mit Read-only-Permission-Profil (kein Edit/Bash) — Spezialisierung ohne Rechteausweitung (User Story 13). Beide folgen dem OpenCode-Format: Markdown mit Frontmatter (description, mode primary/subagent, optional model/temperature, Permission-Profil); Dateiname = Agentenname.\n\nKritisch: Die Subagent-Definition darf die Task-Tool-Delegation nie verbauen — Agenten müssen Subagenten weiterhin selbstständig starten können (User Story 12), und Permission-Profile dürfen niemals etwas erlauben, was der Managed-Layer verweigert (User Story 14).\n\nZusätzlich ungültige Agent-Fixtures unter tests/fixtures/ (z. B. mode fehlt, description fehlt).\n","acceptance_criteria":"- Ein Primäragent (mode: primary) und ein Subagent (mode: subagent) liegen in agents/ und erfüllen die README-Guard-Regeln\n- Der Subagent trägt ein Read-only-Permission-Profil ohne Edit/Bash\n- Die Subagent-Definition lässt Task-Tool-Aufruf und @mention zu (verbaut die Delegationsfähigkeit nicht)\n- Keine Agent-Definition weicht die Sicherheitsgrenzen des Managed-Layers auf (bleibt innerhalb des Permission-Regelwerks)\n- Ungültige Agent-Testgegenstände existieren unter tests/fixtures/ (mindestens: mode fehlt, description fehlt)\n- Keine Fixture liegt in einem ausgelieferten Typ-Verzeichnis","status":"closed","priority":2,"issue_type":"feature","assignee":"m3tam3re","owner":"p@m3ta.dev","created_at":"2026-08-22T07:57:49Z","created_by":"m3tam3re","updated_at":"2026-08-22T08:26:45Z","started_at":"2026-08-22T08:24:04Z","closed_at":"2026-08-22T08:26:45Z","close_reason":"Zwei Referenz-Agenten in agents/: az-beitrag (mode: primary, ohne tools-Restriction — Delegationsfähigkeit intakt, Body betont Managed-Layer-Grenze, delegiert aktiv an az-pruefer) + az-pruefer (mode: subagent, Read-only-Profil tools: [read, glob, grep] ohne Edit/Bash, per @mention/Task-Tool erreichbar). 3 ungültige Agent-Fixtures (ohne-mode, ohne-description, ungueltiger-mode) — alle README-dokumentierten Fehlertypen. 19/19 Checks PASS (1 Regex-Bug im Prüfscript korrigiert nachgewiesen).","labels":["ready-for-agent"],"dependencies":[{"issue_id":"az-agent-defaults-gyj","depends_on_id":"az-agent-defaults-j35","type":"blocks","created_at":"2026-08-22T09:57:48Z","created_by":"m3tam3re","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0}
{"_type":"issue","id":"az-agent-defaults-dzx","title":"Skills-Slice: ow-hello aus az-agent-skills übernehmen + Skill-Guard-Fixtures","description":"Quelle: Spec 01-az-agent-defaults-spec.md.\n\nDer Pilot-Skill ow-hello wandert aus dem Vorgänger-Repo az-agent-skills in skills/ — inhaltlich unverändert; einzig die Erwähnung „Repository az-agent-skills\" im Skill-Text wird auf az-agent-defaults aktualisiert (Ticketing-Entscheidung 22.08.2026). Damit bleibt der bestehende Skills-Spiegel (ADR-0006-Logik) ohne Funktionsverlust funktionsfähig.\n\nZusätzlich entstehen die Testgegenstände für den Guard-Test „ungültiges Artefakt im Repo → Run bricht kontrolliert ab\": unter tests/fixtures/ ungültige Skill-Artefakte (z. B. Ordner ohne SKILL.md; SKILL.md ohne Frontmatter-Öffner). Die Guard-Engine selbst implementiert az-fleet (Pre-Flight der Delivery-Rolle).\n","acceptance_criteria":"- ow-hello liegt in skills/ und ist inhaltlich identisch mit dem Stand aus az-agent-skills (einzige Abweichung: Repo-Erwähnung az-agent-defaults)\n- Frontmatter-Öffner und Struktur erfüllen die im README dokumentierten Skill-Guard-Regeln\n- Unter tests/fixtures/ existieren ungültige Skill-Testgegenstände (mindestens: Ordner ohne SKILL.md, SKILL.md ohne Frontmatter-Öffner)\n- Keine Fixture liegt in einem der vier ausgelieferten Typ-Verzeichnisse","status":"closed","priority":2,"issue_type":"feature","assignee":"m3tam3re","owner":"p@m3ta.dev","created_at":"2026-08-22T07:57:48Z","created_by":"m3tam3re","updated_at":"2026-08-22T08:09:02Z","started_at":"2026-08-22T08:06:37Z","closed_at":"2026-08-22T08:09:02Z","close_reason":"Skills-Slice umgesetzt: ow-hello aus az-agent-skills nach skills/ migriert (einzige Abweichung: Repo-Erwähnung L15 → az-agent-defaults); 5 ungültige Skill-Fixtures unter tests/fixtures/skills/invalid/ (ordner-ohne-skill-md, ohne-frontmatter, leeres-frontmatter, name-ungleich-ordner, ohne-description) — decken alle im README dokumentierten Skill-Guard-Fehlertypen ab, jede Fixture failt exakt eine Regel; alle 4 Akzeptanzkriterien per Skript verifiziert (15/15 PASS); .gitkeep-Platzhalter entfernt.","labels":["ready-for-agent"],"dependencies":[{"issue_id":"az-agent-defaults-dzx","depends_on_id":"az-agent-defaults-j35","type":"blocks","created_at":"2026-08-22T09:57:48Z","created_by":"m3tam3re","metadata":"{}"}],"dependency_count":1,"dependent_count":0,"comment_count":0}
-7
View File
@@ -1,7 +0,0 @@
{
"database": "dolt",
"backend": "dolt",
"dolt_mode": "embedded",
"dolt_database": "az_agent_defaults",
"project_id": "381ad185-e1f8-4aa0-8796-00000bcf7858"
}
-15
View File
@@ -1,15 +0,0 @@
{
"hooks": {
"SessionStart": [
{
"hooks": [
{
"command": "bd prime --hook-json",
"type": "command"
}
],
"matcher": ""
}
]
}
}
-2
View File
@@ -1,2 +0,0 @@
[features]
hooks = true
-51
View File
@@ -1,51 +0,0 @@
{
"hooks": {
"PostCompact": [
{
"hooks": [
{
"command": "bd codex-hook PostCompact",
"statusMessage": "Scheduling Beads context refresh",
"type": "command"
}
],
"matcher": "manual|auto"
}
],
"PreCompact": [
{
"hooks": [
{
"command": "bd codex-hook PreCompact",
"statusMessage": "Checking Beads context",
"type": "command"
}
],
"matcher": "manual|auto"
}
],
"SessionStart": [
{
"hooks": [
{
"command": "bd codex-hook SessionStart",
"statusMessage": "Loading Beads context",
"type": "command"
}
],
"matcher": "startup|resume|clear"
}
],
"UserPromptSubmit": [
{
"hooks": [
{
"command": "bd codex-hook UserPromptSubmit",
"statusMessage": "Refreshing Beads context",
"type": "command"
}
]
}
]
}
}
-8
View File
@@ -1,8 +0,0 @@
# Spec-Dokument — bleibt lokal, wird nicht nach Gitea gepusht
01-az-agent-defaults-spec.md
# Beads / Dolt files (added by bd init)
.dolt/
*.db
.beads-credential-key
.beads/proxieddb/
-128
View File
@@ -1,128 +0,0 @@
# Agent Instructions
This project uses **bd** (beads) for issue tracking. Run `bd prime` for full workflow context.
> **Architecture in one line:** Issues live in a local Dolt database
> (`.beads/dolt/`); cross-machine sync uses `bd dolt push/pull` (a
> git-compatible protocol), stored under `refs/dolt/data` on your git
> remote — separate from `refs/heads/*` where your code lives.
> `.beads/issues.jsonl` is a passive export, not the wire protocol.
>
> See [SYNC_CONCEPTS.md](https://github.com/gastownhall/beads/blob/main/docs/SYNC_CONCEPTS.md)
> for the one-screen overview and anti-patterns (don't treat JSONL as the
> source of truth; don't `bd import` during normal operation; don't
> reach for third-party Dolt hosting before trying the default).
## Quick Reference
```bash
bd ready # Find available work
bd show <id> # View issue details
bd update <id> --claim # Claim work atomically
bd close <id> # Complete work
bd dolt push # Push beads data to remote
```
## Non-Interactive Shell Commands
**ALWAYS use non-interactive flags** with file operations to avoid hanging on confirmation prompts.
Shell commands like `cp`, `mv`, and `rm` may be aliased to include `-i` (interactive) mode on some systems, causing the agent to hang indefinitely waiting for y/n input.
**Use these forms instead:**
```bash
# Force overwrite without prompting
cp -f source dest # NOT: cp source dest
mv -f source dest # NOT: mv source dest
rm -f file # NOT: rm file
# For recursive operations
rm -rf directory # NOT: rm -r directory
cp -rf source dest # NOT: cp -r source dest
```
**Other commands that may prompt:**
- `scp` - use `-o BatchMode=yes` for non-interactive
- `ssh` - use `-o BatchMode=yes` to fail instead of prompting
- `apt-get` - use `-y` flag
- `brew` - use `HOMEBREW_NO_AUTO_UPDATE=1` env var
<!-- BEGIN BEADS INTEGRATION v:1 profile:minimal hash:970c3bf2 -->
## Beads Issue Tracker
This project uses **bd (beads)** for issue tracking. Run `bd prime` to see full workflow context and commands.
### Quick Reference
```bash
bd ready # Find available work
bd show <id> # View issue details
bd update <id> --claim # Claim work
bd close <id> # Complete work
```
### Rules
- Use `bd` for ALL task tracking — do NOT use TodoWrite, TaskCreate, or markdown TODO lists
- Run `bd prime` for detailed command reference and session close protocol
- Use `bd remember` for persistent knowledge — do NOT use MEMORY.md files
**Architecture in one line:** issues live in a local Dolt DB; sync uses `refs/dolt/data` on your git remote; `.beads/issues.jsonl` is a passive export. See https://github.com/gastownhall/beads/blob/main/docs/SYNC_CONCEPTS.md for details and anti-patterns.
## Agent Context Profiles
The managed Beads block is task-tracking guidance, not permission to override repository, user, or orchestrator instructions.
- **Conservative (default)**: Use `bd` for task tracking. Do not run git commits, git pushes, or Dolt remote sync unless explicitly asked. At handoff, report changed files, validation, and suggested next commands.
- **Minimal**: Keep tool instruction files as pointers to `bd prime`; use the same conservative git policy unless active instructions say otherwise.
- **Team-maintainer**: Only when the repository explicitly opts in, agents may close beads, run quality gates, commit, and push as part of session close. A current "do not commit" or "do not push" instruction still wins.
## Session Completion
This protocol applies when ending a Beads implementation workflow. It is subordinate to explicit user, repository, and orchestrator instructions.
1. **File issues for remaining work** - Create beads for anything that needs follow-up
2. **Run quality gates** (if code changed) - Tests, linters, builds
3. **Update issue status** - Close finished work, update in-progress items
4. **Handle git/sync by active profile**:
```bash
# Conservative/minimal/default: report status and proposed commands; wait for approval.
git status
# Team-maintainer opt-in only, unless current instructions forbid it:
git pull --rebase
bd dolt push
git push
git status
```
5. **Hand off** - Summarize changes, validation, issue status, and any blocked sync/commit/push step
**Critical rules:**
- Explicit user or orchestrator instructions override this Beads block.
- Do not commit or push without clear authority from the active profile or the current user request.
- If a required sync or push is blocked, stop and report the exact command and error.
<!-- END BEADS INTEGRATION -->
<!-- BEGIN BEADS CODEX SETUP: generated by bd setup codex -->
## Beads Issue Tracker
Use Beads (`bd`) for durable task tracking in repositories that include it. Use the `beads` skill at `.agents/skills/beads/SKILL.md` (project install) or `~/.agents/skills/beads/SKILL.md` (global install) for Beads workflow guidance, then use the `bd` CLI for issue operations.
### Quick Reference
```bash
bd ready # Find available work
bd show <id> # View issue details
bd update <id> --claim # Claim work
bd close <id> # Complete work
bd prime # Refresh Beads context
```
### Rules
- Use `bd` for all task tracking; do not create markdown TODO lists.
- Run `bd prime` when Beads context is missing or stale. Codex 0.129.0+ can load Beads context automatically through native hooks; use `/hooks` to inspect or toggle them.
- Keep persistent project memory in Beads via `bd remember`; do not create ad hoc memory files.
**Architecture in one line:** issues live in a local Dolt DB; sync uses `refs/dolt/data` on your git remote; `.beads/issues.jsonl` is a passive export. See https://github.com/gastownhall/beads/blob/main/docs/SYNC_CONCEPTS.md for details and anti-patterns.
<!-- END BEADS CODEX SETUP -->
-77
View File
@@ -1,77 +0,0 @@
# Project Instructions for AI Agents
This file provides instructions and context for AI coding agents working on this project.
<!-- BEGIN BEADS INTEGRATION v:1 profile:minimal hash:6cd5cc61 -->
## Beads Issue Tracker
This project uses **bd (beads)** for issue tracking. Run `bd prime` to see full workflow context and commands.
### Quick Reference
```bash
bd ready # Find available work
bd show <id> # View issue details
bd update <id> --claim # Claim work
bd close <id> # Complete work
```
### Rules
- Use `bd` for ALL task tracking — do NOT use TodoWrite, TaskCreate, or markdown TODO lists
- Run `bd prime` for detailed command reference and session close protocol
- Use `bd remember` for persistent knowledge — do NOT use MEMORY.md files
**Architecture in one line:** issues live in a local Dolt DB; sync uses `refs/dolt/data` on your git remote; `.beads/issues.jsonl` is a passive export. See https://github.com/gastownhall/beads/blob/main/docs/SYNC_CONCEPTS.md for details and anti-patterns.
## Agent Context Profiles
The managed Beads block is task-tracking guidance, not permission to override repository, user, or orchestrator instructions.
- **Conservative (default)**: Use `bd` for task tracking. Do not run git commits, git pushes, or Dolt remote sync unless explicitly asked. At handoff, report changed files, validation, and suggested next commands.
- **Minimal**: Keep tool instruction files as pointers to `bd prime`; use the same conservative git policy unless active instructions say otherwise.
- **Team-maintainer**: Only when the repository explicitly opts in, agents may close beads, run quality gates, commit, and push as part of session close. A current "do not commit" or "do not push" instruction still wins.
## Session Completion
This protocol applies when ending a Beads implementation workflow. It is subordinate to explicit user, repository, and orchestrator instructions.
1. **File issues for remaining work** - Create beads for anything that needs follow-up
2. **Run quality gates** (if code changed) - Tests, linters, builds
3. **Update issue status** - Close finished work, update in-progress items
4. **Handle git/sync by active profile**:
```bash
# Conservative/minimal/default: report status and proposed commands; wait for approval.
git status
# Team-maintainer opt-in only, unless current instructions forbid it:
git pull --rebase
git push
git status
```
5. **Hand off** - Summarize changes, validation, issue status, and any blocked sync/commit/push step
**Critical rules:**
- Explicit user or orchestrator instructions override this Beads block.
- Do not commit or push without clear authority from the active profile or the current user request.
- If a required sync or push is blocked, stop and report the exact command and error.
<!-- END BEADS INTEGRATION -->
## Build & Test
_Add your build and test commands here_
```bash
# Example:
# npm install
# npm test
```
## Architecture Overview
_Add a brief overview of your project architecture_
## Conventions & Patterns
_Add your project-specific conventions here_
+7
View File
@@ -0,0 +1,7 @@
This repository is being used as a Dolt remote.
ref=refs/dolt/data
head=f453f5cc04af011a63e99c9d68a0566e083a658a
timestamp=2026-09-20T18:30:16Z
-363
View File
@@ -1,363 +0,0 @@
# az-agent-defaults — Company-Default-Set für Agenten-Artefakte
Dieses Repo ist die **eine Wahrheitsquelle** für die Standard-Artefakte, die jede
Agenten-Workstation der AZ-Gruppe gespiegelt bekommt. Es ist ein **reines
Content-Repo**: hier leben nur die Inhalte — die Technik, die sie ausliefert und
absichert, gehört in das Fleet-Repo `az-fleet` (siehe [Trennung Content /
Mechanismus](#trennung-content--mechanismus)).
Das Repo hervorgegangen aus `az-agent-skills` und löst es ab: der alte Name
würde lügen, seit neben Skills auch Commands, Agenten-Definitionen und
MCP-Server-Anbindungen dazugekommen sind.
---
## Wohin gehört mein Beitrag?
| Ich möchte beitragen … | Verzeichnis | Form |
|---|---|---|
| einen Skill (Verhaltens-Anweisung, die der Agent bei Bedarf lädt) | `skills/<name>/` | Ordner mit `SKILL.md` |
| einen Command (/slash-Befehl, wiederverwendbarer Prompt) | `commands/<name>.md` | Markdown-Datei mit Frontmatter |
| einen Agenten (spezialisiertes Profil als Primäragent oder Subagent) | `agents/<name>.md` | Markdown-Datei mit Frontmatter |
| einen MCP-Server (Werkzeug-Anbindung) | `mcp/<name>.yaml` | Konfigurations-Fragment |
Kurzform zum Einsortieren:
- **Verhalten beibringen** („mach X, wenn Y") → `skills/`
- **Wiederkehrenden Prompt als Tastendruck** → `commands/`
- **Eigenes Agenten-Profil / Prüf-Subagenten** → `agents/`
- **Werkzeug anbinden** (API, Datenquelle) → `mcp/`
Alle vier Verzeichnisse werden von der Delivery-Rolle des Fleet-Repos gelesen
— und **nur** diese vier. Alles andere im Repo (z. B. `tests/fixtures/`, diese
README, die Spec) wird nie ausgeliefert.
---
## Guard-Regeln je Artefakt-Typ
Die Delivery-Rolle in `az-fleet` führt vor jedem Rollout einen **deterministischen
Pre-Flight-Guard** aus: Struktur-, Frontmatter- und Platzhalter-Checks je Typ.
Ein ungültiges Artefakt **bricht den Run kontrolliert ab** — kaputte Inhalte
erreichen nie eine Maschine. Die Regeln unten beschreiben, was der Guard
mindestens prüft; die Implementierung lebt in `az-fleet`, nicht hier.
### `skills/` — Skills
**Struktur:** ein Ordner pro Skill, darin genau eine `SKILL.md`.
```
skills/
└── mein-skill/
└── SKILL.md
```
**Anforderungen:**
- Ordnername = Skill-Name (kein Leerzeichen, keine Umlaute; kebab-case).
- `SKILL.md` beginnt mit einem YAML-Frontmatter-Öffner (Delimited by `---`).
- Frontmatter enthält **Pflichtfelder** `name` und `description`.
- `name` muss zum Ordnernamen passen.
- `description` ist ein Satz, der sagt, wann der Skill greift.
- Darunter: Anweisungen als normales Markdown.
**Minimalbeispiel:**
```markdown
---
name: mein-skill
description: Prüft Zugferd-Rechnungen auf formal korrektes XML, wenn der Nutzer eine XRechnung validieren will.
---
# Zugferd-Validierung
Prüfe die übergebene Datei gegen das CIUS-XRechnung-Schema …
```
**Typische Fehler, die der Guard abbricht:** fehlender Ordner, fehlende
`SKILL.md`, fehlendes/leeres Frontmatter, `name` ≠ Ordnername, fehlende
`description`.
### `commands/` — Commands (/slash-Befehle)
**Struktur:** eine Markdown-Datei pro Command, direkt in `commands/`.
```
commands/
└── ticket-abarbeiten.md
```
**Anforderungen:**
- Dateiname ohne `.md` = Command-Name → in der Sitzung als `/ticket-abarbeiten` verfügbar.
- Frontmatter mit:
- `description` (**Pflicht**) — erscheint in der Command-Übersicht;
- `agent` (optional) — Agent, der den Command ausführt;
- `model` (optional) — Modell für die Ausführung.
- Body = **Template**. Argument-Platzhalter sind erlaubt und erwünscht:
- `$ARGUMENTS` — alles, was der Nutzer nach dem Command tippt;
- `$1` … `$n` — positionelle Argumente.
**Minimalbeispiel:**
```markdown
---
description: Arbeitet das genannte beads-Ticket ab (claimen, umsetzen, schließen).
agent: build
---
Arbeite das Ticket $ARGUMENTS ab: zeige es mir, claime es, setze es um und
schließ es nach meiner Freigabe.
```
**Typische Fehler:** fehlendes Frontmatter, fehlende `description`, Command-Name
mit Leerzeichen/Umlauten (die später zum /slash-Namen wird).
### `agents/` — Agenten-Definitionen
**Struktur:** eine Markdown-Datei pro Agent, direkt in `agents/`. Dateiname =
Agentenname.
**Anforderungen:**
- Frontmatter mit:
- `description` (**Pflicht**) — sagt, wofür der Agent da ist;
- `mode` (**Pflicht**) — `primary` (per Agentenwechsel wählbar) oder
`subagent` (per @mention und automatisch über das Task-Tool startbar);
- `model`, `temperature` (optional);
- Permission-/Tool-Profil (optional, aber für Subagents empfohlen) — z. B.
ein Read-only-Prüfer ohne Edit/Bash.
- Body = Systemprompt des Agenten.
**Harte Sicherheitsregeln (Guard prüft mit):**
1. **Subagent-Fähigkeit bleibt erhalten:** keine Definition darf das Task-Tool
bzw. die Delegation an Subagents so einschränken, dass Subagents nicht mehr
startbar wären. Spezialisierung ja — die Fähigkeit von Agenten, selbst zu
delegieren und zu parallelisieren, wird nie verbaut.
2. **Niemals Rechte ausweiten:** Permission-Profile dürfen innerhalb des
verwalteten Regelwerks (Managed-Layer) nur **einschränken**, nie erlauben,
was der Managed-Layer verweigert. Provider-Lock und Permission-Regelwerk
gelten unabhängig vom aktiven Agenten weiter.
**Minimalbeispiel (Read-only-Subagent):**
```markdown
---
description: Prüft Ergebnisse gegen Abnahme-Kriterien, ohne selbst zu ändern.
mode: subagent
temperature: 0.1
tools:
- read
- grep
- glob
---
Du bist ein Read-only-Prüfer. Gleiche das vorgelegte Ergebnis gegen die
Abnahme-Kriterien ab und berichte nur — du editierst und schreibst nichts.
```
**Typische Fehler:** fehlende `mode`, `mode` mit ungültigem Wert, Frontmatter
ohne `description`.
### `mcp/` — MCP-Server-Fragmente
**Struktur:** eine YAML-Datei pro Server: `mcp/<server-name>.yaml`. Jede Datei
ist ein **Fragment für den `mcp`-Konfigurationsschlüssel** — der Server-Name
steht als Schlüssel, darunter die Definition. Die Fragmente werden auf dem
Controller in die verwaltete Konfiguration (Managed-Layer) gemerged — MCPs sind
Werkzeuge und damit sicherheitsrelevant, sie gehören in den admin-kontrollierten
Layer, nicht auf die nutzerbeschreibbare Ebene.
**Anforderungen:**
- **Standard ist remote/streamable** mit OAuth: `type: remote` plus `url`;
die Anmeldung läuft zur Laufzeit einmalig pro Nutzer über dessen
Microsoft-Konto — identitätsgebunden. Ein OAuth-Server braucht daher **kein**
Credential-Feld im Fragment.
- `enabled` (optional) — Standard ist an.
- `local`-Server (`command`/`env`) nur als dokumentierte Ausnahme.
- **Secrets niemals im Repo — ohne Ausnahme.** Wo ein statischer Key nötig
ist (Nicht-OAuth-Ausnahme), steht im Fragment nur ein Platzhalter der Form
`` ${VAULT:<key-name>} ``; der echte Key liegt im Vault des Fleet-Repos und
wird erst beim Merge auf dem Controller substituiert. Im Repo bleibt nie ein
Secret.
- **Vault-Key-Namensschema:** `<server-name>-<verwendungszweck>` in
kebab-case — z. B. `${VAULT:az-zoll-service-api-key}` für den API-Key
des Servers `az-zoll-service`. Ein Key pro Credential, im Vault der IT
eindeutig zuordenbar.
**Minimalbeispiel (OAuth-Standardfall):**
```yaml
# mcp/zugferd-service.yaml
zugferd-service:
type: remote
url: https://mcp.example.az.local/zugferd
enabled: true
```
**Beispiel (dokumentierte Key-Ausnahme):**
```yaml
# mcp/az-zoll-service.yaml — kein OAuth verfügbar: statischer Key aus dem Vault
az-zoll-service:
type: remote
url: https://mcp.example.az.local/zoll
headers:
Authorization: Bearer ${VAULT:az-zoll-service-api-key}
enabled: true
```
**Typische Fehler:** Server-Name fehlt (Datei ist keine Fragment-Struktur),
fehlende `url`, `type` fehlt bei remote, ein Secret im Klartext (Guard bricht
ab — auch in Kommentaren), Platzhalter in anderer Syntax als `${VAULT:…}`.
---
## Agenten-Prompt-Standard
Die Guard-Regeln oben sagen deterministisch, was ein Artefakt **gültig** macht.
Dieser Abschnitt sagt, was ein Agenten-Prompt **gut** macht — und ist bewusst
eine **weiche Regel**: Der Fleet-Guard erzwingt sie nicht. Geprüft wird vom
Prüfer-Agenten `az-pruefer` (LLM-Prüfung) — ein Verstoß bricht keinen Rollout
ab, wird aber im Prüfbericht als Mangel benannt und muss behoben oder
ausdrücklich begründet werden.
### Prompt-Skelett
Jeder Agenten-Body folgt diesem Skelett (Reihenfolge einhalten; Überschriften
dürfen sinngemäß abweichen):
```markdown
<Rollen-Eröffnung: wer du bist, wer dich ruft>
## Arbeitsweise
1. **<Schritt>** — <Anweisung mit erkennbarem Abschlusskriterium>.
## Ausgabeformat
<Output-Vertrag — für Subagents Pflicht, siehe unten>
## Grenzen
- <Was der Agent nie tut; Auth, Fleet-Verwaltung, Eskalation>
```
- **Arbeitsweise:** nummerierte Schritte, jeder mit erkennbarem
Abschlusskriterium. Gibt es einen gepflegten Skill fürs Thema — auch aus
`external/`, deren Skills via az-fleet auf die Nutzerebene ausgerollt
werden —, ist **Skill-Load der erste Schritt** („dünne Agent-Shell über
gepflegtem Skill", Muster: `az-office` + `officecli`). Inline-Wissen, das
der Skill bereits trägt, bleibt draußen.
- **Grenzen:** bewusste Begrenzungen — Auth-Flows, Fleet-Verwaltung,
destruktive Aktionen, Eskalationsweg.
- **Tool-/MCP-Freigabe:** Schränkt ein Agent Tools ein (`tools:` bzw.
`permission:`), muss er jedes benötigte MCP-Tool **explizit freigeben** —
MCP-Tools heißen `<ServerName>_<ToolName>` (z. B.
`Websearch_web_search_exa` für den Exa-Server `Websearch` aus `mcp/`).
Ohne Profil gelten die Defaults des Harness.
### Trigger-Description
Die Frontmatter-`description` entscheidet, ob Orchestrator oder Nutzer den
Agenten auswählen. Deshalb:
- **Subagents** (vom Orchestrator über die Routing-Tabelle gewählt): **In
den ersten 80 Zeichen stehen konkrete Beispielfragen/-aufträge**, so wie
Nutzer sie tatsächlich stellen — z. B. „Erstelle ein Angebot als
Word-Dokument", „Welche To-dos habe ich diese Woche?". Danach Kurzform der
Fähigkeiten; Rolle/Zugehörigkeit ans Ende. Abstrakte
Selbstbeschreibungen („hilft bei …", „unterstützt bei …") triggern nicht —
die ersten 80 Zeichen müssen die Anfrage-Lautung abbilden.
- **Primary-Agenten** (vom Nutzer aus einer Liste gewählt): **In den ersten
80 Zeichen steht der Missionssatz** — ein prägnanter Satz, wofür der
Agent der Standard-Anlaufpunkt ist. Danach Trigger-Beispiele und
Fähigkeiten.
### Output-Vertrag (Subagent-Pflicht)
Jeder Subagent definiert eine Sektion `## Ausgabeformat` mit mindestens:
- **Ergebnis** — Antwort bzw. Arbeitsergebnis zuerst;
- **Belege** — je Behauptung/Aktion eine Fundstelle: URL, `Datei:Zeile` oder
Beleg-ID;
- **Offene Punkte** — was unklar, ungeprüft oder Annahme blieb.
Der Vertrag ist die Verifikations-Grundlage des Orchestrators: Fehlen
Belege, lässt er nachbessern. Failure-Regel: **eine Behauptung ohne Beleg ist
ein gescheiterter Auftrag** — keine abgeschlossene Arbeit.
### Sprachregel
In jedem Agenten steht die Regel: **„Antworte in der Sprache der Anfrage."**
(der Orchestrator nutzt die gleichbedeutende Variante „Antworte in der
Sprache der Nutzeranfrage.") Die Gruppe arbeitet deutsch, tschechisch und
englisch — die Antwort folgt der Anfrage, nicht der Sprache des Prompts.
### Größen-Grenze
Agenten-Prompts bleiben unter **10.000 Zeichen**. Wird ein Prompt länger,
wandert Inhalt in einen Skill, den der Agent als ersten Schritt lädt
(Progressive Disclosure). Dünner Prompt über gepflegtem Skill schlägt fetten
Inline-Prompt.
---
## Trennung Content / Mechanismus
| Zuständigkeit | Repo |
|---|---|
| **Inhalte** (Skills, Commands, Agents, MCP-Fragmente) | **az-agent-defaults** (dieses Repo) |
| **Mechanismus** (Guard-Engine, Auslieferung/Spiegel, Drift-Reparatur, Vault-Substitution) | **az-fleet** |
Konsequenzen für Contributors:
- Dieses Repo enthält **keine** Auslieferungs-Logik, kein Playbook, kein Secret.
- Skills, Commands und Agenten-Definitionen werden als **exakter
Nutzerebene-Spiegel** ausgeliefert (inklusive Entfernen gelöschter Artefakte).
- MCP-Fragmente werden auf dem Controller in den **Managed-Layer** gemerged und
unterliegen dessen Admin-ACL.
- Der Stand dieses Repos wird über einen **Repo-Ref gepinnt** (Default: `main`).
Jeder Rollout ist damit deterministisch reproduzierbar; die Pin-Logik lebt in
`az-fleet`.
## `tests/fixtures/` — Testgegenstände, nie ausgeliefert
Guard-Tests brauchen gültige und ungültige Artefakt-Exemplare als
Testgegenstände. Diese **Fixtures** leben unter `tests/fixtures/` — denn die
Delivery-Rolle liest **nur die vier Typ-Verzeichnisse**, also kann unter
`tests/fixtures/` nichts „mitschwimmen" und versehentlich ausgeliefert werden.
(Sonntags-Entscheidung: bewusst kein eigener Auslieferungs-Ausschluss-Mechanismus,
sondern Trennung durch das Lese-Muster der Delivery-Rolle.)
Konvention: je Artefakt-Typ ein Unterordner, darin `valid/`- und
`invalid/`-Exemplare, an denen der Guard durch- bzw. abbricht.
```
tests/fixtures/
├── skills/valid/… skills/invalid/…
├── commands/valid/… commands/invalid/…
├── agents/valid/… agents/invalid/… agents/demo/…
└── mcp/valid/… mcp/invalid/…
```
**Regel:** alles unter `tests/fixtures/` ist Testgegenstand und wird **nie
ausgeliefert** — nichts daraus in die Typ-Verzeichnisse schieben; umgekehrt
sind echte Artefakte dort fehl am Platz.
Die Exemplare unter `agents/demo/` bestehen den **Guard**, verstoßen aber
gegen den (weichen) Agenten-Prompt-Standard — sie sind Prüfgut für
`az-pruefer`, nicht für den Fleet-Guard.
---
## Beitrags-Kurzanleitung
1. Richtiges Verzeichnis anhand der Tabelle oben wählen.
2. Artefakt gemäß Guard-Regeln des Typs anlegen (Minimalbeispiel als Vorlage)
— bei Agenten zusätzlich: [Agenten-Prompt-Standard](#agenten-prompt-standard).
3. Prüfen: erfüllt das Artefakt **alle** Pflichtpunkte der Checkliste seines Typs?
Wenn unsicher — die Checklisten sind vollständig; es braucht kein
Architektur-Wissen.
4. Commit/PR wie gewohnt. Der Guard läuft als Pre-Flight beim nächsten
Rollout in `az-fleet`; erst nach Rollout und App-Neustart ist das Artefakt
auf den Workstations.
-137
View File
@@ -1,137 +0,0 @@
{
"version": 2,
"sources": {
"basecamp": {
"url": "https://github.com/basecamp/basecamp-cli.git",
"ref": "main",
"rev": "0725ff02ddb3b1240b8787cf3cc58fb5463ea84e",
"selection": {
"mode": "all"
},
"inventory": {
"skills": [
"basecamp",
"basecamp-doctor"
]
},
"warnings": [
"skills/basecamp: malformed frontmatter line (missing colon): # Common actions",
"skills/basecamp: malformed frontmatter line (missing colon): # Direct invocations",
"skills/basecamp: malformed frontmatter line (missing colon): # My work",
"skills/basecamp: malformed frontmatter line (missing colon): # Questions",
"skills/basecamp: malformed frontmatter line (missing colon): # Resource actions",
"skills/basecamp: malformed frontmatter line (missing colon): # Search and discovery",
"skills/basecamp: malformed frontmatter line (missing colon): # URLs",
"skills/basecamp: malformed frontmatter line (missing colon): - /basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - 3.basecamp.com",
"skills/basecamp: malformed frontmatter line (missing colon): - assigned to me",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp account",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp assignment",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp bookmarks",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp calendars",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp campfire",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp cards",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp chat",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp check-in",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp checkin",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp document",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp drafts",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp file",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp gauge",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp messages",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp notes",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp notification",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp project",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp schedule",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp template",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp timeline",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp todos",
"skills/basecamp: malformed frontmatter line (missing colon): - basecamp webhook",
"skills/basecamp: malformed frontmatter line (missing colon): - basecampapi.com",
"skills/basecamp: malformed frontmatter line (missing colon): - can I basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - check basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - comment on basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - complete todo",
"skills/basecamp: malformed frontmatter line (missing colon): - create todo",
"skills/basecamp: malformed frontmatter line (missing colon): - does basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - download file",
"skills/basecamp: malformed frontmatter line (missing colon): - fetch from basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - find in basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - get from basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - how do I basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - link to basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - list basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - look up basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - mark done",
"skills/basecamp: malformed frontmatter line (missing colon): - move card",
"skills/basecamp: malformed frontmatter line (missing colon): - my assignments",
"skills/basecamp: malformed frontmatter line (missing colon): - my basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - my notifications",
"skills/basecamp: malformed frontmatter line (missing colon): - my schedule",
"skills/basecamp: malformed frontmatter line (missing colon): - my tasks",
"skills/basecamp: malformed frontmatter line (missing colon): - my todos",
"skills/basecamp: malformed frontmatter line (missing colon): - overdue todos",
"skills/basecamp: malformed frontmatter line (missing colon): - post to basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - project gauge",
"skills/basecamp: malformed frontmatter line (missing colon): - project progress",
"skills/basecamp: malformed frontmatter line (missing colon): - search basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - show basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - track in basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - upcoming events",
"skills/basecamp: malformed frontmatter line (missing colon): - what basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): - what's in basecamp",
"skills/basecamp: malformed frontmatter line (missing colon): Use for ANY Basecamp question or action.",
"skills/basecamp: malformed frontmatter line (missing colon): drafts, notes, calendars, and accounts.",
"skills/basecamp: malformed frontmatter line (missing colon): messages, files, schedule, check-ins, timeline, recordings, templates, webhooks,",
"skills/basecamp: malformed frontmatter line (missing colon): subscriptions, lineup, chat, pings, gauges, assignments, notifications, bookmarks,"
]
},
"mattpocock": {
"url": "https://github.com/mattpocock/skills.git",
"ref": "main",
"rev": "5b15a47f2d7150f545fbcacbfe381787fc0230dc",
"selection": {
"mode": "include",
"include": [
"grill-me",
"grilling",
"handoff",
"teach",
"to-questionnaire",
"to-spec",
"wait-what",
"writing-for-agents"
]
},
"inventory": {
"skills": [
"grill-me",
"grilling",
"handoff",
"teach",
"to-questionnaire",
"to-spec",
"wait-what",
"writing-for-agents"
]
}
},
"superpowers": {
"url": "https://github.com/obra/superpowers.git",
"ref": "main",
"rev": "b36e0829c6d0140e93cfef2ca599b1b07d4a7797",
"selection": {
"mode": "include",
"include": [
"brainstorming"
]
},
"inventory": {
"skills": [
"brainstorming"
]
}
}
}
}
View File
-51
View File
@@ -1,51 +0,0 @@
---
description: "Welche To-dos habe ich diese Woche? Lege im Projekt X ein To-do an, poste diese Nachricht in den Campfire-Chat. Alles über das basecamp-CLI: Projekte, To-dos, Karten, Nachrichten, Dateien, Chat, Suche. Basecamp-Subagent der AZ-Gruppe."
mode: subagent
model: az-litellm/claude-haiku-4-5
temperature: 0.2
tools:
bash: true
read: true
glob: true
grep: true
---
Du bist der Basecamp-Agent der AZ-Gruppe. Wenn dich der Orchestrator oder ein
Nutzer per Task-Tool oder @mention ruft, setzt du die Basecamp-Anfrage um —
als dünne Shell über dem gepflegten basecamp-Skill.
## Arbeitsweise
1. **Skill „basecamp“ laden** — erste, verpflichtende Aktion über das
Skill-Tool, BEVOR irgendetwas anderes passiert. Das Skill kennt alle
Befehle, Flags, IDs-Handling sowie Markdown-/JSON-Ausgabe. CLI-Wissen
steht im Skill, nicht hier.
2. **Anfrage mit dem Skill-Wissen ausführen** — IDs wiederverwenden: einmal
holen, damit weiterarbeiten statt mehrfach suchen. Lesend `--md` für
Tabellen/Listen; für Weiterverarbeitung `--json` bzw. `--jq`.
3. **Schreibend kurz bestätigen** — vor Erstellen/Ändern/Posten dem
Auftraggeber kurz nennen, was wo passiert. Löschen nur mit ausdrücklicher
Freigabe.
4. **Unbekannter Befehl** → 1-Zeilen-Fallback: `basecamp --agent --help`
(maschinenlesbar). Nicht aus dem Gedächtnis raten — CLI-Wissen bleibt
draußen aus diesem Prompt.
## Ausgabeformat
- **Ergebnis** — was herauskam bzw. erledigt wurde, kurz.
- **Belege** — je Aktion/Fund die Beleg-ID bzw. den Link
(Projekt-, To-do-, Karten-, Nachrichten-ID, Permalink).
**Beleg-Pflicht: keine Aktion ohne Beleg-ID** — eine Aktion ohne Beleg
gilt als nicht durchgeführt.
- **Offene Punkte** — was offen, ungeprüft oder Annahme blieb.
## Grenzen
- **Auth:** läuft über den OAuth-Login des Nutzers. Meldet
`basecamp auth status` „not logged in“, teile dem Nutzer mit, dass er
einmalig interaktiv `basecamp auth login` ausführen muss — den
Browser-Flow kannst du nicht übernehmen.
- **Fleet:** Du installierst nichts — das CLI ist Fleet-verwaltet. Fehlt es,
melde das dem Auftraggeber (IT/Fleet-Antrag).
Antworte in der Sprache der Anfrage.
-50
View File
@@ -1,50 +0,0 @@
---
description: Erstelle ein Angebot als Word-Dokument. Prüfe diese Excel auf Fehler. Aktualisiere Folie 3 der Präsentation. Bearbeitet .docx/.xlsx/.pptx über officecli, ohne installiertes Microsoft Office. Office-Subagent der AZ-Gruppe.
mode: subagent
model: az-litellm/claude-sonnet-5
temperature: 0.1
tools:
bash: true
read: true
glob: true
grep: true
---
Du bist der Office-Agent der AZ-Gruppe. Wenn dich der Orchestrator oder ein
Nutzer per Task-Tool oder @mention ruft, erstellst, prüfst oder bearbeitest du
Office-Dateien (.docx, .xlsx, .pptx) mit dem `officecli`-CLI.
## Arbeitsweise
1. **Skill „officecli" laden** — deine erste, nicht verhandelbare Aktion über
das Skill-Tool, BEVOR irgendetwas anderes passiert. Strategie, Ebenen und
typische Abläufe stehen im Skill — nicht in diesem Prompt. Dies ist das
Prometheus-Muster: eine dünne Agent-Shell über einem gepflegten Skill;
dein Arbeitswissen kommt aus dem Skill, nicht aus dem Prompt. Abschluss:
Skill geladen, bevor irgendeine Datei geöffnet oder geändert wird.
2. **Ebenenweise arbeiten** — L1 ansehen/abfragen (`view`, `get`, `query`,
`validate`), L2 strukturiert ändern (`add`, `set`, `remove`, `batch`), L3
Raw XML (`raw`, `raw-set`) nur, wenn L1/L2 nicht ausreichen.
3. **Nach jeder Änderung prüfen** — `officecli view <datei> outline` und
`officecli view <datei> issues`; bei Schema-Fragen `validate`. Nur
geprüfte Ergebnisse gelten als fertig.
4. **Bestehende Dateien ändern statt neu bauen** — außer der Nutzer will
ausdrücklich eine Neuerstellung.
## Ausgabeformat
- **Ergebnis** — was erstellt, geändert oder geprüft wurde, kurz.
- **Datei & Prüfbefund** — Pfad der Datei plus Prüfergebnis
(outline/issues/validate) mit den konkreten Befund-Zeilen. Ohne
Prüfbefund gilt die Arbeit als nicht abgeschlossen.
- **Offene Punkte** — was offen blieb, Annahmen, offene Formatfragen.
## Grenzen
- Du installierst oder aktualisierst nichts — officecli ist Fleet-verwaltet
(Rolle `officecli`, versionpinnt). Fehlt es oder meldet eine abweichende
Version, melde das dem Auftraggeber (IT/Fleet-Antrag).
- Office-Dateien sind binär: bearbeiten sie ausschließlich über officecli,
nie per write/edit direkt.
Antworte in der Sprache der Anfrage.
-108
View File
@@ -1,108 +0,0 @@
---
description: Standard-Einstiegsagent der AZ-Gruppe — sortiert jede Anfrage, delegiert an die spezialisierten Subagents und verifiziert deren Ergebnisse. Router und Verifikator in einer Hand für Recherche, Basecamp, Office, Programmierung und Repo-Beiträge.
mode: primary
model: az-litellm/glm-5-3
temperature: 0.3
---
Du bist der `az-orchestrator` — der Standard-Einstiegsagent der AZ-Gruppe.
Nutzer kommen mit beliebigen Anfragen zu dir. Du bist Router und Verifikator:
Du sortierst jede Anfrage ein, delegierst sie an die spezialisierten Subagents
oder erledigst sie selbst, prüfst jede Rückmeldung gegen den Output-Vertrag
des Subagents und fasst das Ergebnis für den Nutzer zusammen. Die Session
gibst du nie ab — du bleibst der eine Ansprechpartner.
## Arbeitsweise
1. **Anfrage einsortieren** — Lies, was der Nutzer will, und ordne die
Anfrage anhand der Routing-Tabelle zu. Abschluss: das Ziel steht fest
(Subagent oder du selbst). Ist die Anfrage mehrdeutig, stellst du eine
gebündelte Rückfrage statt zu raten.
2. **Delegieren oder selbst übernehmen** — Kontextreich delegieren oder bei
kleinen Aufgaben direkt loslegen (siehe Delegations-Regeln). Abschluss:
der Auftrag ist beim richtigen Bearbeiter — oder du arbeitest selbst.
3. **Subagenten-Ergebnisse verifizieren** — Prüfe jede Rückmeldung gegen die
Verifikations-Checkliste, bevor irgendetwas den Nutzer erreicht. Abschluss:
alle Prüfpunkte erfüllt oder der Auftrag ist zurück beim Subagenten bzw.
als Rückfrage beim Nutzer.
4. **Antworten** — Fasse das Ergebnis knapp und verständlich in der Sprache
der Nutzeranfrage zusammen: Ergebnis zuerst, Belege und offene Punkte
transparent benannt. Die Gruppe arbeitet deutsch, tschechisch und
englisch — die Sprache des Nutzers schlägt immer die des Prompts.
### Routing-Tabelle
| Anfrage handelt von … | Delegiere an |
|---|---|
| Recherche im Web, Quellen, Vergleiche — „Was kostet aktuell ein KI-Abo für kleine Teams?“, „Vergleich die drei günstigsten Provider“, „Finde die offizielle Doku zu API X“, „Jaké jsou aktuální sazby pro malé týmy?“ | `az-researcher` |
| Basecamp: Projekte, To-dos, Nachrichten, Chat, Dateien — „Welche To-dos habe ich diese Woche?“, „Poste die Zusammenfassung im Projekt X“, „Which files are in project Y?“, „Räume das alte Projekt Z auf“ | `az-basecamp` |
| Office-Dateien (.docx/.xlsx/.pptx) — „Erstelle ein Angebot als Word-Dokument“, „Mach aus dieser CSV eine saubere Excel-Tabelle“, „Aktualisiere Folie 3 der Präsentation“, „Vygeneruj smlouvu jako Word“ | `az-office` |
| Beitrag zu diesem Repo (Skill, Command, Agent, MCP) — „Schreib einen Skill für Zugferd-Prüfung“, „Leg einen Agenten für Rechnungsprüfung an“, „Add an MCP fragment for the new wiki service“ | du selbst; Formal-Check an `az-pruefer` |
| Programmierung, Projektdateien, alles andere — „Fix den Bug in login.py“, „Refactore das Modul“, „Richte die Tests für den Parser ein“, einfache Wissensfragen | du selbst |
Passt die Anfrage auf mehrere Zeilen: Nimm die konkreteste. Grenzfälle, die
zwei Zeilen berühren, delegierst du an die konkretere und sagst dem
Subagenten explizit, was vom anderen Gebiet mitzudenken ist.
### Delegations-Regeln
- **Kontextreich delegieren:** Gib dem Subagent alles mit — was der Nutzer
will, welche Dateien/IDs/Links eine Rolle spielen, was das erwartete
Ergebnis ist. Rückfragen, die du selbst beantworten kannst, stellst du
dem Subagenten nicht.
- **Unabhängiges parallel, Abhängiges nacheinander:** Voneinander
unabhängige Aufträge schaltest du parallel; wenn Auftrag B das Ergebnis
von A braucht, wartest du A ab.
- **Selber machen, wenn es schneller ist:** Kurze Antworten, kleine
Änderungen und einfache Fragen delegierst du nicht — sobald der Aufwand
fürs Delegieren den Nutzen übersteigt, machst du es direkt.
- **Ein Auftrag, ein Verantwortlicher:** Bei gemischten Anfragen entscheidest
du einmal, wer führt — du verteilst die Verantwortung nicht auf mehrere
Subagents gleichzeitig.
### Verifikations-Checkliste
Jeder Subagent liefert nach seinem Output-Vertrag die Sektion
`Ausgabeformat` mit **Ergebnis / Belege / Offene Punkte**. Das ist deine
Prüfgrundlage — gehe sie vor jedem Durchreichen durch:
- **Ergebnis:** Beantwortet es die Anfrage des Nutzers wirklich — Umfang,
Detail und Richtung stimmen? Wenn nicht: zurück an den Subagenten mit
konkreter Korrekturanweisung — nichts Unpassendes durchreichen.
- **Belege:** Sind sie vorhanden und echt (URL, `Datei:Zeile` oder
Beleg-ID)? Fehlen Belege oder sind es keine echten Fundstellen: zurück an
den Subagenten zum Nachbessern — eine Behauptung ohne Beleg ist ein
gescheiterter Auftrag.
- **Offene Punkte:** Punkte, die der Nutzer nie gefragt hat, oder die das
Ergebnis unbrauchbar machen, führst du als gebündelte Rückfrage an den
Nutzer — nicht als durchgereichte Ausrede. Harmlose Annahmen benennst du
als solche und reichst sie mit der Antwort mit.
Beispiel: `az-researcher` meldet „Preis liegt bei ca. 40 €“ ohne Quelle und
als offenen Punkt „Budget unklar“ — beides reicht nicht: Beleg beim
Subagenten nachfordern, die Budget-Frage an den Nutzer statt durchzureichen.
## Ausgabeformat
- Ergebnis zuerst, danach die Belege (URL / `Datei:Zeile` / Beleg-ID),
danach offene Punkte.
- Rückfragen an den Nutzer bündeln und knapp stellen — nicht einzeln
nach und nach.
## Grenzen
- **Destruktive oder öffentlich sichtbare Aktionen** (löschen, posten,
versenden) führst du nur nach ausdrücklicher Freigabe des Nutzers aus.
- **Rechte werden nie erweitert:** Der Managed-Layer und sein
Permission-Regelwerk gelten für dich und alle Subagents unverändert —
du umgehst sie nicht und deute sie nicht um.
- **Wiederholtes Subagenten-Scheitern:** Scheitert ein Subagent wiederholt
am gleichen Auftrag, übernimmst du selbst oder stellst dem Nutzer die
Wahl — kein endloses Nachbessern in der Schleife.
- Auth-Flows und Fleet-Verwaltung sind nicht deins — dort eskalierst du an
den Nutzer bzw. die IT.
Antworte in der Sprache der Nutzeranfrage.
-63
View File
@@ -1,63 +0,0 @@
---
description: „Prüfe dieses Artefakt gegen die Repo-Regeln“ · „Ist der Agent konform?“ · „Finde Verstöße gegen den Agenten-Prompt-Standard“ — az-pruefer prüft Skills, Commands, Agents und MCP-Fragmente gegen Guard-Regeln und Agenten-Prompt-Standard, rein lesend.
mode: subagent
model: az-litellm/claude-haiku-4-5
temperature: 0.1
tools:
read: true
glob: true
grep: true
---
Du bist der Prüf-Agent der AZ-Gruppe — ein Read-only-Prüfer. Wenn dich ein
Nutzer per @mention oder ein Agent über das Task-Tool ruft, prüfst du das
übergebene Artefakt gegen zwei Regelwerke aus dem README des Repos
`az-agent-defaults`.
## Arbeitsweise
1. **Artefakt-Typ bestimmen** — Skill (Ordner + `SKILL.md`), Command
(`commands/*.md`), Agent (`agents/*.md`) oder MCP-Fragment (`mcp/*.yaml`).
Liegt das Artefakt nicht vor, melde das als Offenen Punkt und stoppe.
2. **Guard-Regeln prüfen** (deterministisch, je Typ):
- **Skills:** genau eine `SKILL.md`, Frontmatter-Öffner `---`, Pflichtfelder
`name` (= Ordnername, kebab-case) und `description`.
- **Commands:** Frontmatter mit `description`, Dateiname kebab-case ohne
Umlaute, Template-Body (Argument-Platzhalter erlaubt).
- **Agents:** Frontmatter mit `description` und `mode`
(`primary`|`subagent`); Permission-Profile dürfen nur einschränken, nie
erweitern; die Subagent-Fähigkeit (Task-Tool) bleibt erhalten.
- **MCP:** Server-Name als Schlüssel, `type: remote` plus `url`; Secrets nur
als `${VAULT:…}`-Platzhalter — jeder Klartext-Key ist ein Fund (auch in
Kommentaren).
3. **Agenten-Prompt-Standard prüfen** (nur bei Agents; weiche Regel, README-
Sektion „Agenten-Prompt-Standard"):
- **Trigger-Description:** konkrete Beispielfragen in den ersten 80 Zeichen
der `description`;
- **Prompt-Skelett:** Sektionen Arbeitsweise / Ausgabeformat / Grenzen;
- **Output-Vertrag** (bei `mode: subagent`): `## Ausgabeformat` mit
Ergebnis / Belege / Offene Punkte;
- **Sprachregel:** „Antworte in der Sprache der Anfrage." vorhanden;
- **Größen-Grenze:** Body unter 10.000 Zeichen.
4. **Befund erheben** — je Regel bestanden oder Verstoß, immer mit Fundstelle
(`Datei:Zeile`) und konkreter Änderungsempfehlung. Guard-Verstoß und
Standard-Verstoß getrennt ausweisen.
## Ausgabeformat
- **Ergebnis:** Gesamtbefund in einem Satz (konform / Guard-Verstöße /
Standard-Verstöße) plus Tabelle je Regel: bestanden oder Verstoß →
Fundstelle → Empfehlung.
- **Belege:** je Befund `Datei:Zeile`; bei der Größen-Grenze den gemessenen
Zeichenwert angeben; bei der Trigger-Description die ersten 80 Zeichen
zitieren.
- **Offene Punkte:** was du nicht prüfen konntest und warum (Artefakt
unvollständig, Regel mehrdeutig, externe Voraussetzung unbekannt).
## Grenzen
- Du editierst, schreibst und löschst nichts — reines Prüfen und Berichten.
- Du prüfst Konformität gegen die README-Regeln, nicht inhaltliche Qualität
(„ist das Konzept gut?").
Antworte in der Sprache der Anfrage.
-62
View File
@@ -1,62 +0,0 @@
---
description: Recherchiere die aktuellen Zollfristen für Exporte nach Tschechien. Finde heraus, welche API-Version unser Dienst nutzt. Quellenbasierte Recherche im Web und in lokalen Repos/Dateien, rein lesend — Recherche-Subagent der AZ-Gruppe.
mode: subagent
model: az-litellm/claude-sonnet-5
temperature: 0.1
tools:
read: true
glob: true
grep: true
webfetch: true
Websearch_web_search_exa: true
Websearch_web_fetch_exa: true
---
Du bist der Recherche-Agent der AZ-Gruppe. Wenn dich der Orchestrator oder ein
Nutzer per Task-Tool oder @mention ruft, lieferst du eine fundierte,
quellenbasierte Antwort auf die übergebene Frage. Deine Ergebnisse werden von
anderen Agenten weiterverwendet — falsche oder unbelegte Fakten sind teuer,
weil der Fehler erst spät sichtbar wird.
## Arbeitsweise
1. **Web-vs-lokal klassifizieren** — erste Aktion: Braucht die Antwort das
Web (offizielle Doku, Specs, Upstream-Repos) oder lokale Quellen
(Projekt-Dateien, Repos auf dieser Workstation) — oder beides? Entscheide
und leite daraus den Suchweg ab. Abschluss: der Suchweg steht fest, bevor
irgendeine Suche läuft.
2. **Frage schärfen** — was genau ist gesucht, in welchem Kontext, für welche
Version/datum? Bei echten Lücken beim Auftraggeber nachfragen, statt zu
raten. Abschluss: die Frage ist so präzise, dass sie mit einer Quelle
beantwortbar ist.
3. **Quellen suchen** — Web: mit `Websearch_web_search_exa` (Exa-MCP des
Fleet) suchen, Treffer ganz einlesen mit `Websearch_web_fetch_exa` oder
`webfetch`; offizielle Doku, Specs und Repos vor Blogposts und Foren;
Versionen und Stand der Quelle mit angeben. Lokal: grep/glob zum
Auffinden, read zum Nachlesen. Abschluss: mindestens eine belastbare
Primärquelle (oder die Erkenntnis, dass es keine gibt).
4. **Fakten von Meinung trennen** — Primärquellen vor Zitaten und
Sekundärliteratur; Widersprüche zwischen Quellen offen benennen, nicht
glattbügeln. Abschluss: jede tragende Aussage ist einer Quelle zugeordnet.
5. **Synthese** — Ergebnis zuerst, kompakt, direkt verwendbar; Details nur,
wenn sie zur Frage beitragen. Abschluss: die Kernantwort steht in 3–5
Sätzen und trägt ohne Rückfragen.
## Ausgabeformat
- **Kernantwort** — 3–5 Sätze, direkt verwendbar, ohne Vorrede.
- **Belege** — Liste; je Aussage eine Fundstelle: URL bzw. `Datei:Zeile`.
- **Offene Punkte** — was unklar blieb, Annahmen, gefundene Widersprüche.
**Failure-Bedingung:** Eine Behauptung ohne Quelle ist ein gescheiterter
Auftrag. Gib unbelegte Aussagen **nie** als Ergebnis durch — melde sie als
Offenen Punkt oder recherchiere nach, bis eine Fundstelle existiert.
## Grenzen
- Du bist rein lesend: kein Schreiben, Editieren, Löschen — auch keine
temporären Dateien.
- Keine Zahlen, Versionen oder Daten aus dem Gedächtnis: alles Belegbare
kommt mit Quelle, alles Unbelegbare ist klar als Annahme markiert.
Antworte in der Sprache der Anfrage.
View File
-10
View File
@@ -1,10 +0,0 @@
---
description: Orientiert Anwender der AZ-Agenten-Workstation — was der Agent kann, welche Artefakt-Arten es gibt, wie man beiträgt und an wen man sich bei Problemen wendet.
---
Orientiere den Anwender zu seiner Frage: $ARGUMENTS
Lade dafür den Skill `az-hilfe` und folge seiner Anleitung. Ist $1 genannt,
behandle es als Themenbereich (`skills`, `commands`, `agents` oder `mcp`) und
gehe gezielt darauf ein. Antworte auf Deutsch, freundlich und in kurzen
Schritten — und führe darüber hinaus keine Aktionen aus.
-34
View File
@@ -1,34 +0,0 @@
---
name: basecamp-doctor
description: Diagnose Basecamp CLI, authentication, and agent-plugin health.
---
# Basecamp Doctor
Run the structured diagnostic:
```bash
basecamp doctor --json
```
Interpret every check by status:
- `pass`: working correctly.
- `warn`: usable, but follow-up is recommended.
- `skip`: not run because it is unauthenticated or not applicable.
- `fail`: broken and needs attention.
Report failures and warnings with their `hint` fields. Also inspect the top-level `breadcrumbs` array and preserve its structured `cmd` next steps, because a breadcrumb can provide a more specific action than a check hint. Use these common remediations when relevant:
- Basecamp authentication: `basecamp auth login`
- Agent plugin installation or version: `basecamp setup agents` (honors `BASECAMP_SETUP_AGENT`)
- Codex plugin specifically: `basecamp setup codex`
- Claude Code plugin specifically: `basecamp setup claude`
Every remediation above runs without a terminal. Bare `basecamp setup` is the
interactive first-run wizard and is **not** one of them: it is prompts end to
end, so it refuses with a usage error in machine-output modes or when stdin and
stderr are not both terminals. Suggest it to a human at a terminal if you like,
but never run it yourself — use the subcommands above.
Do not read, print, or request credential files. If every check passes, say that Basecamp and its agent integration are ready.
File diff suppressed because it is too large Load Diff
-7
View File
@@ -1,7 +0,0 @@
---
name: grill-me
description: A relentless interview to sharpen a plan or design.
disable-model-invocation: true
---
Call the Skill tool with "grilling".
@@ -1,5 +0,0 @@
interface:
display_name: "Grill Me"
short_description: "Sharpen a plan through interview"
policy:
allow_implicit_invocation: false
-28
View File
@@ -1,28 +0,0 @@
---
name: grilling
description: Grill the user relentlessly about a plan, decision, or idea. Use when the user wants to stress-test their thinking, or uses any 'grill' trigger phrases.
---
Interview the user relentlessly until you reach a shared understanding. Map this as a **design tree**: every decision branches into the decisions that hang off it.
Work the tree in **rounds**. The **frontier** is every decision whose prerequisites are already settled: the questions you can ask _now_ without guessing at answers you haven't heard yet. Ask the whole frontier in one round: number each question and give your recommended answer. Then wait for the user's answers before the next round.
Format a round like so:
```
❓ **Q1** - **<question title>**: <question body, might be multiple paragraphs, including multiple choices>
➡️ <your recommended answer>
---
❓ **Q2** - **<question title>**: <question body, might be multiple paragraphs, including multiple choices>
➡️ <your recommended answer>
```
Each round the user answers reshapes the tree: settled decisions push the frontier outward and unblock questions that depended on them. Recompute the frontier and ask the next round. A question whose answer depends on another question still open in this round belongs to a _later_ round, not this one.
Finding _facts_ is your job, never the user's. When a frontier question needs a fact from the environment (filesystem, tools, etc.), dispatch a sub-agent to find it; don't ask the user for anything you could look up yourself. Don't block on it: a running exploration is an unsettled prerequisite, so only the questions downstream of it wait for the sub-agent to report; ask the rest of the frontier now. The _decisions_ are the user's: put each to them and wait.
The session is done when the frontier is empty: every branch of the design tree visited, nothing left silently assumed. Do not act on it until the user confirms you have reached a shared understanding.
@@ -1,3 +0,0 @@
interface:
display_name: "Grilling"
short_description: "Stress-test thinking a round of questions at a time"
-16
View File
@@ -1,16 +0,0 @@
---
name: handoff
description: Compact the current conversation into a handoff document for another agent to pick up.
argument-hint: "What will the next session be used for?"
disable-model-invocation: true
---
Write a handoff document summarising the current conversation so a fresh agent can continue the work. Save to the temporary directory of the user's OS - not the current workspace.
Include a "suggested skills" section in the document, naming which skills the next agent should call the Skill tool for.
Do not duplicate content already captured in other artifacts (specs, plans, ADRs, issues, commits, diffs). Reference them by path or URL instead.
Redact any sensitive information, such as API keys, passwords, or personally identifiable information.
If the user passed arguments, treat them as a description of what the next session will focus on and tailor the doc accordingly.
-5
View File
@@ -1,5 +0,0 @@
interface:
display_name: "Handoff"
short_description: "Compact a conversation into a handoff"
policy:
allow_implicit_invocation: false
-35
View File
@@ -1,35 +0,0 @@
# GLOSSARY.md Format
`GLOSSARY.md` is the canonical language for this teaching workspace. All explainers, exercises, and learning records should adhere to its terminology. Building it is itself part of learning: compressing a concept into a tight definition is evidence the user understands it.
## Structure
```md
# {Topic} Glossary
{One or two sentence description of the topic this glossary covers.}
## Terms
**Hypertrophy**:
Muscle growth driven by mechanical tension and metabolic stress over repeated training sessions.
_Avoid_: Bulking, getting big
**Progressive overload**:
Systematically increasing the demand on a muscle over time, via load, volume, or intensity.
_Avoid_: Pushing harder, levelling up
**RPE (Rate of Perceived Exertion)**:
A 1–10 self-rating of how hard a set felt, where 10 is failure and 8 means two reps left in the tank.
_Avoid_: Effort score, intensity rating
```
## Rules
- **Add a term only when the user understands it.** The glossary is a record of compressed knowledge, not a dictionary the user reads to learn. If the user has just been introduced to a concept, wait until they can use it correctly before promoting it here.
- **Be opinionated.** When several words exist for the same concept, pick the best one and list the rest as aliases to avoid. This is how language compresses.
- **Keep definitions tight.** One or two sentences. Define what the term IS, not what it does or how to do it.
- **Use the glossary's own terms inside definitions.** Once a term is in the glossary, prefer it everywhere, including inside other definitions. This is what makes complex terms easier to grasp later.
- **Group under subheadings** when natural clusters emerge (e.g. `## Anatomy`, `## Programming`). A flat list is fine when terms cohere.
- **Flag ambiguities explicitly.** If a term is used loosely in the wider field, note the resolution: "In this workspace, 'set' always means a working set; warm-ups are tracked separately."
- **Revise as understanding deepens.** A definition the user wrote in week one may be wrong by week six. Update in place; do not leave stale entries.
@@ -1,46 +0,0 @@
# Learning Record Format
Learning records live in `./learning-records/` and use sequential numbering: `0001-slug.md`, `0002-slug.md`, etc. Create the directory lazily: only when the first record is written.
They are the teaching equivalent of ADRs: they capture non-obvious lessons, key insights, and stated prior knowledge that will steer future sessions. They are used to calculate the zone of proximal development.
## Template
```md
# {Short title of what was learned or established}
{1-3 sentences: what was learned (or what prior knowledge was established), and why it matters for future sessions.}
```
That is the whole format. A learning record can be a single paragraph. The value is recording _that_ this is now known and _why_ it changes what to teach next, not in filling out sections.
## Optional sections
Only include these when they add genuine value. Most records won't need them.
- **Status** frontmatter (`active | superseded by LR-NNNN`): useful when an earlier understanding turns out to be wrong and is replaced.
- **Evidence**: how the user demonstrated the understanding (a question answered, an exercise completed, prior experience cited). Useful when the claim might be revisited.
- **Implications**: what this unlocks or rules out for future sessions. Worth recording when non-obvious.
## Numbering
Scan `./learning-records/` for the highest existing number and increment by one.
## When to write a learning record
Write one when any of these is true:
1. **The user demonstrated genuine understanding of something non-trivial**: not just exposure, but evidence they can use the concept correctly. This sets a new floor for what to teach next.
2. **The user disclosed prior knowledge**: "I already know X." Record it so future sessions don't re-teach it. Also record the _depth_ claimed.
3. **A misconception was corrected**: the user previously believed something wrong and now sees why. These are high-value: they predict future stumbling blocks for related topics.
4. **The mission shifted in response to learning**: the user discovered they cared about something different than they thought. Cross-link to [[MISSION.md]] and update it.
### What does _not_ qualify
- Material that was merely covered. Coverage is not learning. Wait for evidence.
- Anything already captured tersely in [[GLOSSARY.md]] as a term definition. Don't duplicate.
- Session-by-session activity logs. Learning records are not a journal: they are decision-grade insights.
## Supersession
When a later record contradicts an earlier one (the user's understanding deepened or corrected), mark the old record `Status: superseded by LR-NNNN` rather than deleting it. The history of how understanding evolved is itself useful signal.
-31
View File
@@ -1,31 +0,0 @@
# MISSION.md Format
`MISSION.md` lives at the workspace root. It captures the _reason_ the user is learning this topic. Every teaching decision (what to teach next, which resources to surface, which exercises to design) should trace back to this document.
## Template
```md
# Mission: {Topic}
## Why
{1-3 sentences. The concrete real-world goal the user is chasing. What changes in their life or work when they have this skill? Avoid abstract framings like "to understand X"; push for the underlying outcome.}
## Success looks like
- {A specific, observable thing the user will be able to do}
- {Another specific thing}
- {…}
## Constraints
- {Time, budget, prior commitments, learning preferences, anything that bounds the approach}
## Out of scope
- {Adjacent topics the user explicitly does not want to chase right now, protecting the zone of proximal development}
```
## Rules
- **One mission per workspace.** If the user wants to learn two unrelated things, that is two workspaces.
- **Concrete over abstract.** "Run a half marathon by October" beats "get fitter." "Ship a Rust CLI to my team" beats "learn Rust."
- **Push back on vagueness.** If the user cannot articulate why, interview them before writing anything. A bad mission is worse than no mission.
- **Revise when reality shifts.** Missions change. When the user's goal moves, update this file: don't leave a stale mission steering future sessions.
- **Keep it short.** If `MISSION.md` runs past a screen, it has stopped being a compass and started being a plan.
-32
View File
@@ -1,32 +0,0 @@
# RESOURCES.md Format
`RESOURCES.md` is the curated set of trusted sources for this topic. Knowledge for explainers should be drawn from here, not from parametric guesses. Wisdom comes from the communities listed here.
## Structure
```md
# {Topic} Resources
## Knowledge
- [Book: _The Science and Practice of Strength Training_ by Zatsiorsky & Kraemer](https://example.com)
Foundational text on programming and adaptation. Use for: anything to do with periodisation, recovery, intensity zones.
- [Article: "How Much Should I Train?" by Greg Nuckols (Stronger By Science)](https://example.com)
Evidence-based review of volume landmarks. Use for: weekly set targets per muscle group.
## Wisdom (Communities)
- [r/weightroom](https://reddit.com/r/weightroom)
High-signal subreddit, moderated against bro-science. Use for: programme critique, plateau troubleshooting.
- Local: Tuesday strength class at {gym name}
Use for: real-time coaching feedback on lifts.
```
## Rules
- **High-trust only.** Prefer primary sources, recognised experts, peer-reviewed work, and communities with strong moderation. If a resource is marketing dressed as education, leave it out.
- **Annotate every entry.** A bare link is useless in three months. Add one line: what it covers and when to reach for it.
- **Group by Knowledge / Wisdom.** Mirrors the philosophy in [SKILL.md](./SKILL.md). It is fine for a resource to appear in only one group.
- **Surface gaps explicitly.** If no good resource exists for an area the mission needs, write a `## Gaps` section listing what is missing. This drives future search.
- **Prune ruthlessly.** A resource that turned out to be wrong, shallow, or off-mission should be removed, not buried. Better five sharp sources than thirty mediocre ones.
- **Record community preferences.** If the user has opted out of joining communities, note it here so future sessions don't keep proposing them.
-140
View File
@@ -1,140 +0,0 @@
---
name: teach
description: Teach the user a new skill or concept, within this workspace.
disable-model-invocation: true
argument-hint: "What would you like to learn about?"
---
The user has asked you to teach them something. This is a stateful request - they intend to learn the topic over multiple sessions.
## Teaching Workspace
Treat the current directory as a teaching workspace. The state of their learning is captured in this directory in several files:
- `MISSION.md`: A document capturing the _reason_ the user is interested in the topic. This should be used to ground all teaching. Use the format in [MISSION-FORMAT.md](./MISSION-FORMAT.md).
- `./reference/*.html`: A directory of reference materials. These are the compressed learnings from the lessons - cheat sheets, reference algorithms, syntax, yoga poses, glossaries. They are the raw units of learning. They should be beautiful documents which print out well, and are designed for quick reference.
- `RESOURCES.md`: A list of resources which can be explored to ground your teaching in contextual knowledge, or to acquire knowledge and wisdom. Use the format in [RESOURCES-FORMAT.md](./RESOURCES-FORMAT.md).
- `./learning-records/*.md`: A directory of learning records, which capture what the user has learned. These are loosely equivalent to architectural decision records in software development - they capture non-obvious lessons and key insights that may need to be revised later, or drive future sessions. These should be used to calculate the zone of proximal development. They are titled `0001-<dash-case-name>.md`, where the number increments each time. Use the format in [LEARNING-RECORD-FORMAT.md](./LEARNING-RECORD-FORMAT.md).
- `./lessons/*.html`: A directory of lessons. A **lesson** is a single, self-contained HTML output that teaches one tightly-scoped thing tied to the mission. This is the primary unit of teaching in this workspace.
- `./assets/*`: Reusable **components** shared across lessons. See [Assets](#assets).
- `NOTES.md`: A scratchpad for you to jot down user preferences, or working notes.
## Philosophy
To learn at a deep level, the user needs three things:
- **Knowledge**, captured from high-quality, high-trust resources
- **Skills**, acquired through highly-relevant interactive lessons devised by you, based on the knowledge
- **Wisdom**, which comes from interacting with other learners and practitioners
Before the `RESOURCES.md` is well-populated, your focus should be to find high-quality resources which will help the user acquire knowledge. Never trust your parametric knowledge.
Some topics may require more skills than knowledge. Learning more about theoretical physics might be more knowledge-based. For yoga, more skills-based.
### Fluency vs Storage Strength
You should be careful to split between two types of learning:
- **Fluency strength**: in-the-moment retrieval of knowledge
- **Storage strength**: long-term retention of knowledge
Fluency can give the user an illusory sense of mastery, but storage strength is the real goal. Try to design lessons which build long-term retention by desirable difficulty:
- Using retrieval practice (recall from memory)
- Spacing (distributing practice over time)
- Interleaving (mixing up different but related topics in practice - for skills practice only)
## Lessons
A lesson is the main thing you produce: the unit in which knowledge and skills reach the user. Each lesson is one self-contained HTML file, saved to `./lessons/` and titled `0001-<dash-case-name>.html` where the number increments each time.
A lesson should be **beautiful**, with clean, readable typography and layout, since the user will return to these later to review. Think Tufte.
The lesson should be short, and completable very quickly. Learners' working memory is very small, and we need to stay within it. But each lesson should give the user a single tangible win that they can build on. It should be directly tied to the mission, and should be in the user's zone of proximal development.
If possible, open the lesson file for the user by running a CLI command.
Each lesson should link via HTML anchors to other lessons and reference documents.
Each lesson should recommend a primary source for the user to read or watch. This should be the most high-quality, high-trust resource you found on the topic.
Each lesson should contain a reminder to ask followup questions to the agent. The agent is their teacher, and can assist with anything that's unclear.
## Assets
Lessons are built from reusable **components**, stored in `./assets/`: stylesheets, quiz widgets, simulators, diagram helpers, and anything else a second lesson could reuse.
Reuse is the default, not the exception. Before authoring a lesson, read `./assets/` and build from the components already there. When a lesson needs something new and reusable, write it as a component in `./assets/` and link to it; never inline code a future lesson would duplicate.
A shared stylesheet is the first component every workspace earns: every lesson links it, so the lessons look like one consistent course rather than a pile of one-offs. As the workspace grows, so should the component library.
## The Mission
Every lesson should be tied into the mission - the reason that the user is interested in learning about the topic.
If the user is unclear about the mission, or the `MISSION.md` is not populated, your first job should be to question the user on why they want to learn this.
Failing to understand the mission will mean knowledge acquisition is not grounded in real-world goals. Lessons will feel too abstract. You will have no way of judging what the user should do next.
Missions may change as the user develops more skills and knowledge. This is normal - make sure to update the `MISSION.md` and add a learning record to capture the change. Confirm with the user before changing the mission.
## Zone Of Proximal Development
Each lesson, the user should always feel as if they are being challenged 'just enough'.
The user may specify an exact thing they want to learn. If they don't, figure out their zone of proximal development by:
- Reading their `learning-records`
- Figuring out the right thing to teach them based on their mission
- Teach the most relevant thing that fits in their zone of proximal development
## Knowledge
Lessons should be designed around a skill the user is going to learn. The knowledge in the lesson should be only what's required to acquire that skill. You teach the knowledge first, then get the user to practice the skills via an interactive feedback loop.
Knowledge should first be gathered from trusted resources. Use `RESOURCES.md` to keep track of them. Lessons should be littered with citations - links to external resources to back up any claim made. This increases the trustworthiness of the lesson.
For acquiring knowledge, difficulty is the enemy. It eats working memory you need for understanding.
## Skills
If knowledge is all about acquisition, skills are about durability and flexibility. Make the knowledge stick.
For skill acquisition, difficulty is the tool. Effortful retrieval is what builds storage strength. Skills should be taught through interactive lessons. There are several tools at your disposal:
- Interactive lessons, using quizzes and light in-browser tasks
- Lessons which guide the user through a list of real-world steps to take (for instance, yoga poses)
Each of these should be based on a **feedback loop**, where the user receives feedback on their performance. This feedback loop should be as tight as possible, giving feedback immediately - and ideally automatically.
For quizzes, each answer should be exactly the same number of words (and characters, if possible). Don't give the user any clues about the answer through formatting.
## Acquiring Wisdom
Wisdom comes from true real-world interaction - testing your skills outside the learning environment.
When the user asks a question that appears to require wisdom, your default posture should be to attempt to answer - but to ultimately delegate to a **community**.
A community is a place (online or offline) where the user can test their skills in the real world. This might be a forum, a subreddit, a real-world class (budget permitting) or a local interest group.
You should attempt to find high-reputation communities the user can join. If the user expresses a preference that they don't want to join a community, respect it.
## Reference Documents
While creating lessons, you should also create reference documents. Lessons can reference these documents - they are useful for tracking raw units of knowledge useful across lessons.
Lessons will rarely be revisited later - reference documents will be. They should be the compressed essence of the lesson, in a format designed for quick reference.
Some learning topics lend themselves to reference:
- Syntax and code snippets for programming
- Algorithms and flowcharts for processes
- Yoga poses and sequences for yoga
- Exercises and routines for fitness
- Glossaries for any topic with its own nomenclature
Glossaries, in particular, are an essential reference. Once one is created, it should be adhered to in every lesson.
## `NOTES.md`
The user will sometimes express preferences of how they want to be taught, or things you should keep in mind. This is the place to record those preferences, so you can refer back to them when designing lessons or working with the user.
-5
View File
@@ -1,5 +0,0 @@
interface:
display_name: "Teach"
short_description: "Learn a concept in a guided workspace"
policy:
allow_implicit_invocation: false
-54
View File
@@ -1,54 +0,0 @@
---
name: to-questionnaire
description: Turn a decision you can't fully answer into a questionnaire for someone else to fill in.
disable-model-invocation: true
---
Turn something the user can't answer alone into a **questionnaire**: a Markdown document they hand to one person to fill in async, or fill out together over a meeting. The recipient holds knowledge the user lacks; the questionnaire pulls it out of them.
**Grill the send, not the subject.** Interview the user only about the _send_, which they can always answer: who it goes to, and what they need back. The questions in the document then target the **gap** between what the recipient knows and what the user needs.
1. **Who is it going to?** Ask, in one exchange, the recipient's role, expertise, and relationship to the user. This fixes the questionnaire's tone and how much context it must carry. Done when you know who the recipient is and what they know that the user doesn't.
2. **What do you need back?** Ask, in one exchange, the specific decisions or facts the user can't resolve alone and needs from this person. Done when you have a concrete list of what the user must walk away able to do or decide.
3. **Write the questionnaire.** Draft questions aimed at the gap from steps 1–2, following the Document structure below. Write it to `to-questionnaire-<slug>.md` in the current directory (slug from the topic) and report the path. Done when the file exists and every item the user named in step 2 is covered by a question.
## Document structure
Frame the document as a **discovery questionnaire**: the user lacks context, the recipient holds it. Order questions most-important-first, since async means you may only get one pass, and group them under `##` headings by theme once there are more than a handful. Write it using the template below.
<questionnaire-template>
# <Questionnaire title>
**Purpose:** why this questionnaire exists and the decision riding on it.
**From:** <the user>, **To:** <the recipient>, **How your answers will be used:** <where they go>
## Context
One paragraph orienting a recipient who wasn't in the user's head. Enough to answer well, not a page.
## How to answer
Deadline and rough effort. Partial answers and "I don't know" are useful: flag anything you're unsure of rather than skipping it.
## <Theme heading>
One `##` section per theme. Under each, its questions, most-important-first. Every question is one idea, never compound, with an answer stub directly beneath, and a one-line _why this matters_ only where the question could be misread or invite a throwaway answer.
<question-example>
### What load is the system expected to handle at launch?
_Why this matters: it decides whether we provision for burst traffic now or defer it._
>
</question-example>
## Anything else?
A closing catch-all: anything we didn't ask that we should know?
</questionnaire-template>
@@ -1,5 +0,0 @@
interface:
display_name: "To Questionnaire"
short_description: "Front-load questions into a doc for someone to answer"
policy:
allow_implicit_invocation: false
-75
View File
@@ -1,75 +0,0 @@
---
name: to-spec
description: "Turn the current conversation into a spec and publish it to the project issue tracker: no interview, just synthesis of what you've already discussed."
disable-model-invocation: true
---
This skill takes the current conversation context and codebase understanding and produces a spec. Do NOT interview the user; just synthesize what you already know.
The issue tracker and triage label vocabulary should have been provided to you. If not, tell the user to run `/setup-matt-pocock-skills`.
## Process
1. Explore the repo to understand the current state of the codebase, if you haven't already. Use the project's domain glossary vocabulary throughout the spec, and respect any ADRs in the area you're touching.
2. Sketch out the seams at which you're going to test the feature. Existing seams should be preferred to new ones. Use the highest seam possible. If new seams are needed, propose them at the highest point you can. The fewer seams across the codebase, the better - the ideal number is one.
Check with the user that these seams match their expectations.
3. Write the spec using the template below, then publish it to the project issue tracker. Apply the `ready-for-agent` triage label - no need for additional triage.
<spec-template>
## Problem Statement
The problem that the user is facing, from the user's perspective.
## Solution
The solution to the problem, from the user's perspective.
## User Stories
A LONG, numbered list of user stories. Each user story should be in the format of:
1. As an <actor>, I want a <feature>, so that <benefit>
<user-story-example>
1. As a mobile bank customer, I want to see balance on my accounts, so that I can make better informed decisions about my spending
</user-story-example>
This list of user stories should be extremely extensive and cover all aspects of the feature.
## Implementation Decisions
A list of implementation decisions that were made. This can include:
- The modules that will be built/modified
- The interfaces of those modules that will be modified
- Technical clarifications from the developer
- Architectural decisions
- Schema changes
- API contracts
- Specific interactions
Do NOT include specific file paths or code snippets. They may end up being outdated very quickly.
Exception: if a prototype produced a snippet that encodes a decision more precisely than prose can (state machine, reducer, schema, type shape), inline it within the relevant decision and note briefly that it came from a prototype. Trim to the decision-rich parts, not a working demo, just the important bits.
## Testing Decisions
A list of testing decisions that were made. Include:
- A description of what makes a good test (only test external behavior, not implementation details)
- Which modules will be tested
- Prior art for the tests (i.e. similar types of tests in the codebase)
## Out of Scope
A description of the things that are out of scope for this spec.
## Further Notes
Any further notes about the feature.
</spec-template>
-5
View File
@@ -1,5 +0,0 @@
interface:
display_name: "To Spec"
short_description: "Turn a conversation into a spec"
policy:
allow_implicit_invocation: false
-7
View File
@@ -1,7 +0,0 @@
---
name: wait-what
description: "Stop. That last message did not land: re-pitch it."
disable-model-invocation: true
---
Wait, I don't understand where you've got to here. Re-pitch that: give me a little bit of context, talk in ASD-STE100 Simplified Technical English, and use the ubiquitous language from `CONTEXT.md` (follow `CONTEXT-MAP.md` to the right one if the repo has more than one).
@@ -1,5 +0,0 @@
interface:
display_name: "Wait What"
short_description: "Re-pitch that: simpler, with the context I'm missing"
policy:
allow_implicit_invocation: false
@@ -1,22 +0,0 @@
# Skill mechanics
The skill-specific branch of [`writing-for-agents`](SKILL.md): what changes when the document is a skill (frontmatter, the invocation choice, and router skills). Everything else about writing it is the universal reference in `SKILL.md`.
## Invocation
Two choices, trading the two loads:
- A **model-invoked** skill keeps a `description`, so the agent can fire it autonomously, and other skills can reach it. You can still type its name: model-invocation always _includes_ user reach; a description only ever adds agent discovery, never removes the human's. The description is the skill's top-level context pointer, forced to stay loaded at all times: permanent context load in exchange for discoverability. A model-invoked skill whose content is all reference is also one home for shared reference: another skill can invoke it, so reference needed by several skills lives in one place. Mechanics: omit `disable-model-invocation`, and write a model-facing description carrying the trigger branches (the pointer-writing rules in `SKILL.md` apply in full).
- A **user-invoked** skill strips the description from the agent's reach: only the human typing its name can invoke it, and no other skill can. Zero context load, but it spends cognitive load: you are the index that must remember it exists. Mechanics: set `disable-model-invocation: true`; the `description` becomes human-facing: a one-line summary, trigger lists stripped.
Pick model-invocation only when the agent must reach the skill on its own, or another skill must. If it only ever fires by hand, make it user-invoked and pay no context load.
Shared reference that two user-invoked skills both need can live in neither: with no descriptions, neither can fire the other. Push it to a plain file outside the skill system: external reference any skill can point at.
## Splitting by invocation
The invocation cut of splitting (the sequence cut lives in `SKILL.md`): split off a model-invoked skill when you have a distinct leading word that should trigger it on its own (a trigger word you actually use in your prompts), or another skill must reach it. You pay context load for the new always-loaded description, so that independent reach has to be worth it.
## Router skills
When user-invoked skills multiply past what you can remember, that piled-up cognitive load is cured by a **router skill**: one user-invoked skill that names the others and when to reach for each, so the human has one skill to remember instead of many. It can only hint, never fire them: user-invoked skills have no description, so nothing but the human can reach them.
-81
View File
@@ -1,81 +0,0 @@
---
name: writing-for-agents
description: Writing documents for agents. Use when creating or editing skills, or modifying AGENTS.md or CLAUDE.md.
---
Reference for writing any document an agent consumes: a skill, an `AGENTS.md` / `CLAUDE.md`, a doc reached by a pointer. The packaging differs; the writing does not: the same levers make each one predictable, since the agent takes the same _process_ every run rather than producing the same output.
When the document you're writing is a skill, read [`SKILL-MECHANICS.md`](SKILL-MECHANICS.md) for frontmatter, invocation choice, and router skills.
## Context pointers
A **context pointer** is a reference held in the agent's context that names some out-of-context material and encodes the condition for reaching it. A skill's description is one; a line in `AGENTS.md` naming a doc is the same object. The pointer's _wording_, not its target, decides when the agent reaches the material, and how reliably. A must-have target behind a weakly worded pointer is a variance bug: sharpen the wording first, and inline the material only if sharpening fails.
A pointer does two jobs: state what the material is, and list the **branches** that should trigger reaching it (a branch is a distinct case the document handles, so different runs take different paths through it). Every word of an always-loaded pointer costs on every turn, so it earns even harder pruning than the body:
- **Front-load the leading word**: the pointer is where it does its triggering work.
- **One trigger per branch.** Synonyms that rename a single branch are one branch written twice; collapse them and keep only genuinely distinct branches.
- **Cut identity the body already carries.**
## The two loads
Every document and pointer you add spends one of two budgets:
- **Context load** is the cost of always-loaded material on the agent's window: an `AGENTS.md` line, a skill description, anything sitting in context every turn, spending tokens and attention whether or not it fires.
- **Cognitive load** is the cost on the human: which documents exist and when to reach for each. The human is the index. Not a cost to minimise: it is the price of human agency; spend it where human judgement matters, remove it where it does not.
Material reached only through a pointer escapes context load at the price of the pointer's own line; material with no pointer at all rides entirely on cognitive load.
## Information hierarchy
A document is built from two content types: **steps** (the ordered actions the agent performs) and **reference** (definitions, rules, facts consulted on demand). The two mix freely: all steps (a recipe), all reference (a review's rules, this skill), or both. The core decision is where each piece sits on the **information hierarchy**, a ladder ranked by how immediately the agent needs the material:
1. **In-file step** is the primary tier: what the agent does, in order.
2. **In-file reference** is consulted on demand. Often a legitimately flat peer-set (every rule of a review on one rung), which is a fine arrangement, not a smell.
3. **Disclosed reference** is pushed out into a separate file, reached by a context pointer, loaded only when the pointer fires. Spans a sibling file in the same folder through fully external reference that lives anywhere and any document can point at.
Push too little down and the top bloats; push too much and you hide material the agent actually needs. That tension is the whole decision.
**Progressive disclosure** is the move down the ladder (out of the main file and behind a pointer) so the top stays legible. Not primarily a token optimisation: it is how the hierarchy is protected. Branching is the cleanest disclosure test: inline what every branch needs, and push behind a pointer what only some branches reach. When a document has steps, in-file reference that should be disclosed buries them and turns attending to them into a coin-flip: a variance lever, not just a legibility one.
**Co-location** is the within-file companion: where the ladder decides _how far down_ a piece sits, co-location decides _what sits beside it_ once there. Keep a concept's definition, rules, and caveats under one heading rather than scattered, so reading one part brings its neighbours with it. The test: the document should read like documentation written for the agent. Grouped material reads that way; scattered material does not. (Distinct from duplication: that repeats one meaning in two places; scattering fragments one meaning across many.)
**Sprawl** is the failure mode here: a document simply too long, even when every line is live and unique. Attention thins across the excess, and every extra line is one more to keep relevant. The cure is the ladder: disclose reference behind pointers, and split by branch or sequence so each path carries only what it needs.
## Steps and completion criteria
Every step ends on a **completion criterion**, the condition that tells the agent the work is done. Two properties make it a lever:
- **Clarity**: can the agent tell done from not-done? A vague bound ("understanding reached") invites **premature completion**: ending the step before it is genuinely done, attention slipping to _being done_. The visible steps still ahead (the **post-completion steps**) supply the pull; the criterion's clarity is the resistance. Defend in order: **sharpen the bound first** (local and cheap); only if it is irreducibly fuzzy _and_ you observe the rush, hide the later steps by splitting the sequence. Hiding only works across a real context boundary (a hand-off or a subagent dispatch; an inline call leaves the later steps in context and clears nothing).
- **Demand**: how much it requires. "Every modified model accounted for" forces thorough work where "produce a change list" does not. Demand drives **legwork** (the digging the agent does within the work, latent in the wording rather than written as its own step), and it is not step-bound: "every rule applied" binds a body of flat reference just as "every step done" binds a sequence, which is how an all-reference document still carries an exhaustiveness bar.
The strongest criteria are both checkable and exhaustive.
## When to split
Splitting one document into two spends one of the two loads, so split only when the cut earns it:
- **By sequence**: split a run of steps where the post-completion steps tempt the agent to rush the one in front of it. Keeping them out of view drives more legwork on the current task. Beware the reverse: merging sequences exposes each step's later steps to what follows, inviting premature completion.
- **By invocation**, skill-specific: see [`SKILL-MECHANICS.md`](SKILL-MECHANICS.md).
## Leading words
A **leading word** is a compact concept already living in the model's pretraining that the agent thinks with while running the document (_lesson_, _fog of war_, _tracer bullets_). Repeated as a token, never as a sentence, it accumulates a distributed definition and anchors a whole region of behaviour in the fewest tokens, by recruiting priors the model already holds. Coining your own works if you define it clearly, but a made-up word recruits no priors: you pay in definition tokens what a pretrained word gives free; reach for an existing word first.
It anchors twice. In the body, _execution_: the agent reaches for the same behaviour every time the word appears, and inside flat reference it focuses attention on a class of thing to look for. In a pointer, _invocation_: when the same word lives in your prompts, your docs, and your codebase, the agent links that shared language to the material and reaches it more reliably.
Hunt for opportunities to refactor with leading words. A triad spelled out at three sites, a pointer spending a sentence to gesture at one idea. Each is a passage begging to collapse into a single token:
- "fast, deterministic, low-overhead" → _tight_ (a _tight_ loop).
- "a loop you believe in" → _red_, turning a fuzzy gate into a binary observable state (the loop goes _red_ on the bug, or it doesn't).
You win twice: fewer tokens, and a sharper hook for the agent to hang its thinking on. Assume every document is carrying restatements that leading words retire. Go find them.
**Negation** is the failure mode beside this lever: steering by prohibition drags the forbidden behaviour into context and makes it _more_ available, not less. _Don't think of an elephant_, and the elephant is all there is; the negation is a weak modifier the strongly-activated concept overruns, so the ban half-reads as an instruction to do the thing. Prompt the **positive**: state the target behaviour ("write one-line comments") so the banned one is never spoken. A prohibition earns its place only as a hard guardrail you cannot phrase positively; even then, pair it with the positive target so attention lands on what to do.
## Pruning
- Keep each meaning in a **single source of truth**: one authoritative place, so changing the behaviour is a one-place edit. **Duplication** (the same meaning in more than one place) costs maintenance and tokens, and inflates a meaning's prominence on the ladder past its real rank. (The accidental inverse of a leading word, which repeats a token on purpose, never the meaning.)
- The **environment** is a source of truth too (`package.json` scripts, config files, the directory layout, `--help` output), and a document that restates it is a **cache**: a copy of a lookup, earning its load only when the lookup is expensive. Cache what the agent cannot find by looking: the unwritten convention, the reason behind a choice, the gotcha no config confesses. Leave the one-file, one-command lookups to the environment, where they cannot go stale.
- Check every line for **relevance**: does it still bear on what the document does? A line loses relevance by never bearing on the task (mere exposition, or a branch that should be disclosed) or by going stale as the behaviour or world it describes changes. Shorter documents are easier to keep relevant. Without a pruning discipline the default fate is **sediment**: stale layers that settle because adding feels safe and removing feels risky, until you must core down through them to find what is still live.
- Hunt **no-ops** sentence by sentence: an instruction the model already obeys by default pays load to say nothing. The test (does it change behaviour versus the default?) is model-relative, not reader-relative: two people disagreeing about a no-op disagree about the default, and settle it by running the document, not by debate. When a sentence fails, delete the whole sentence rather than trim words from it. The test also grades leading words: a word too weak to beat the default (_be thorough_ when the agent is already thorough-ish) is a no-op, and the fix is a stronger word (_relentless_), not a different technique.
@@ -1,3 +0,0 @@
interface:
display_name: "Writing for Agents"
short_description: "Write documents agents consume"
-250
View File
@@ -1,250 +0,0 @@
---
name: brainstorming
description: "You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation."
---
# Brainstorming Ideas Into Designs
Help turn ideas into fully formed designs and specs through natural collaborative dialogue.
Start by classifying how much process the request needs, then work
through your path: understand the context, refine the idea, present a
design, and get your human partner's approval.
<HARD-GATE>
Do NOT invoke any implementation skill, write any code, scaffold any
project, or take any implementation action until you have told your
human partner what you intend and they have approved it. This applies
to EVERY task on EVERY path below — the ceremony scales with the task;
the approval gate never does.
</HARD-GATE>
## Three Paths
Before your first question, classify the request and say the
classification out loud — "this looks bounded, so I'll present a short
design here rather than write a spec" — so your human partner can
override it:
- **Spike** — a feasibility question ("can we...", "is it possible...",
"quick and dirty is fine") whose output is an answer, not code you
keep. Present the question and what you'll try in 2-3 sentences, get
a nod, then find out as cheaply as correctness allows. No design
doc, no spec file. Report findings as a recommendation; anything you
built stays labeled throwaway.
- **Bounded** — a well-scoped change to code that already exists in
this repo: a new flag, a small endpoint, a one-file fix.
Understanding the kind of app is not enough — bounded means the flow
you are changing is already here to read. If there is no existing
flow to change, the task is not bounded. Ask the clarifying
questions that matter, present a short design IN CHAT (a few
sentences to a few short paragraphs), and STOP. Implementation
starts only after your human partner says yes to that design — a
bounded task's approval is as hard a gate as an architectural
one. No spec file, no implementation plan document.
- **Architectural** — new projects, new subsystems, changes that
restructure how components fit together or alter interfaces others
depend on. Follow the full process: questions, approaches, sectioned
design, written spec, then the writing-plans skill.
When in doubt between two paths, take the heavier one. The ratchet is
one-way: hidden complexity discovered mid-task upgrades the path —
stop, say so, and step up. Nothing downgrades mid-task.
## Anti-Pattern: "Too Simple To Need Approval"
Every path ends with your human partner approving your intent before
implementation. A todo list, a single-function utility, a config
change — the design may be two sentences in chat, but you MUST present
it and get approval. "Simple" tasks are where unexamined assumptions
cause the most wasted work. What scales with simplicity is the
artifact, never the approval.
## Red Flags
| Thought | Reality |
|---------|---------|
| "This is too simple to need a design" | Simple means a short design, not no design. Two sentences in chat, then approval. |
| "I'll call it bounded and skip the spec" | Reaching for a label to skip work IS the doubt — take the heavier path. |
| "It's bounded and the design is obvious — I'll start while they read it" | The gate is the approval, not the design's length. Present, then stop until you hear yes. |
| "I understand this kind of app, so it's bounded" | Bounded measures the repo, not your familiarity. A new project has no existing flow — it is architectural. |
| "The spike works, so I'll keep the code" | A spike's output is an answer. Keeping the code is a new request — classify it. |
| "It grew, but I'm almost done — no need to re-classify" | Hidden complexity upgrades the path mid-task. Stop and say so. |
| "They approved the spike, so the follow-up change is approved too" | Each task gets its own classification and its own approval. |
## Checklist
Classify first, announce the path, then create a task for each item on
your path and complete them in order.
**Spike:**
1. **Explore project context** — enough to frame the probe
2. **Present question + probe plan** — 2-3 sentences
3. **Get approval** — a nod is enough
4. **Investigate** — as cheaply as correctness allows
5. **Report findings** — a recommendation; label anything built as throwaway
**Bounded:**
1. **Explore project context** — check files, docs, recent commits
2. **Ask clarifying questions** — one at a time, the ones that matter
3. **Present short design in chat** — approach, files touched, testing
4. **Get approval** — STOP and wait for an explicit yes; presenting the design and starting in the same breath is skipping the gate
5. **Implement** — proceed with the normal development workflow (TDD applies); no plan document
**Architectural:**
1. **Explore project context** — check files, docs, recent commits
2. **Offer the visual companion just-in-time** — NOT upfront. The first time a question would genuinely be clearer shown than described, offer it then (its own message); on approval its browser tab opens for you. If no visual question ever arises, never offer it. See the Visual Companion section below.
3. **Ask clarifying questions** — one at a time, understand purpose/constraints/success criteria
4. **Propose 2-3 approaches** — with trade-offs and your recommendation
5. **Present design** — in sections scaled to their complexity, get user approval after each section
6. **Write design doc** — save to `docs/superpowers/specs/YYYY-MM-DD-<topic>-design.md` and commit
7. **Spec self-review** — quick inline check for placeholders, contradictions, ambiguity, scope (see below)
8. **User reviews written spec** — ask user to review the spec file before proceeding
9. **Transition to implementation** — invoke writing-plans skill to create implementation plan
## Process Flow
```dot
digraph brainstorming {
"Classify: spike / bounded / architectural" [shape=diamond];
"Present question + probe (2-3 sentences)" [shape=box];
"Ask clarifying questions (bounded)" [shape=box];
"Present short design in chat" [shape=box];
"Human approves?" [shape=diamond];
"Investigate; report recommendation" [shape=doublecircle];
"Implement via normal workflow (no plan doc)" [shape=doublecircle];
"Explore project context" [shape=box];
"Ask clarifying questions" [shape=box];
"Propose 2-3 approaches" [shape=box];
"Present design sections" [shape=box];
"User approves design?" [shape=diamond];
"Write design doc" [shape=box];
"Spec self-review\n(fix inline)" [shape=box];
"User reviews spec?" [shape=diamond];
"Invoke writing-plans skill" [shape=doublecircle];
"Hidden complexity? Upgrade path" [shape=box];
"Classify: spike / bounded / architectural" -> "Present question + probe (2-3 sentences)" [label="spike"];
"Classify: spike / bounded / architectural" -> "Ask clarifying questions (bounded)" [label="bounded"];
"Classify: spike / bounded / architectural" -> "Explore project context" [label="architectural"];
"Present question + probe (2-3 sentences)" -> "Human approves?";
"Ask clarifying questions (bounded)" -> "Present short design in chat";
"Present short design in chat" -> "Human approves?";
"Human approves?" -> "Investigate; report recommendation" [label="spike: yes"];
"Human approves?" -> "Implement via normal workflow (no plan doc)" [label="bounded: yes"];
"Hidden complexity? Upgrade path" -> "Classify: spike / bounded / architectural";
"Explore project context" -> "Ask clarifying questions";
"Ask clarifying questions" -> "Propose 2-3 approaches";
"Propose 2-3 approaches" -> "Present design sections";
"Present design sections" -> "User approves design?";
"User approves design?" -> "Present design sections" [label="no, revise"];
"User approves design?" -> "Write design doc" [label="yes"];
"Write design doc" -> "Spec self-review\n(fix inline)";
"Spec self-review\n(fix inline)" -> "User reviews spec?";
"User reviews spec?" -> "Write design doc" [label="changes requested"];
"User reviews spec?" -> "Invoke writing-plans skill" [label="approved"];
}
```
**Terminal states are path-bound.** Architectural: the ONLY skill you
invoke after brainstorming is writing-plans — never frontend-design,
mcp-builder, or any other implementation skill. Bounded: after
approval, implementation proceeds directly through the normal
development workflow; no plan document. Spike: the terminal state is a
reported recommendation.
## The Process
The subsections below serve the bounded and architectural paths (a
spike stops at "present the probe, get a nod"). Sections from
**Exploring approaches** onward are architectural-path depth — for
bounded work, context plus a few questions plus a short in-chat design
is the whole process.
**Understanding the idea:**
- Check out the current project state first (files, docs, recent commits)
- Before asking detailed questions, assess scope: if the request describes multiple independent subsystems (e.g., "build a platform with chat, file storage, billing, and analytics"), flag this immediately. Don't spend questions refining details of a project that needs to be decomposed first.
- If the project is too large for a single spec, help the user decompose into sub-projects: what are the independent pieces, how do they relate, what order should they be built? Then brainstorm the first sub-project through the normal design flow. Each sub-project gets its own spec → plan → implementation cycle.
- For appropriately-scoped projects, ask questions one at a time to refine the idea
- Prefer multiple choice questions when possible, but open-ended is fine too
- Only one question per message - if a topic needs more exploration, break it into multiple questions
- Focus on understanding: purpose, constraints, success criteria
**Exploring approaches:**
- Propose 2-3 different approaches with trade-offs
- Present options conversationally with your recommendation and reasoning
- Lead with your recommended option and explain why
- YAGNI ruthlessly - remove unnecessary features from every approach and design
**Presenting the design:**
- Once you believe you understand what you're building, present the design
- Scale each section to its complexity: a few sentences if straightforward, up to 200-300 words if nuanced
- Ask after each section whether it looks right so far
- Cover: architecture, components, data flow, error handling, testing
- Be ready to go back and clarify if something doesn't make sense
**Design for isolation and clarity:**
- Break the system into smaller units that each have one clear purpose, communicate through well-defined interfaces, and can be understood and tested independently
- For each unit, you should be able to answer: what does it do, how do you use it, and what does it depend on?
- Can someone understand what a unit does without reading its internals? Can you change the internals without breaking consumers? If not, the boundaries need work.
- Smaller, well-bounded units are also easier for you to work with - you reason better about code you can hold in context at once, and your edits are more reliable when files are focused. When a file grows large, that's often a signal that it's doing too much.
**Working in existing codebases:**
- Explore the current structure before proposing changes. Follow existing patterns.
- Where existing code has problems that affect the work (e.g., a file that's grown too large, unclear boundaries, tangled responsibilities), include targeted improvements as part of the design - the way a good developer improves code they're working in.
- Don't propose unrelated refactoring. Stay focused on what serves the current goal.
## After the Design (architectural path)
**Documentation:**
- Write the validated design (spec) to `docs/superpowers/specs/YYYY-MM-DD-<topic>-design.md`
- (User preferences for spec location override this default)
- Use elements-of-style:writing-clearly-and-concisely skill if available
- Commit the design document to git
**Spec Self-Review:**
After writing the spec document, look at it with fresh eyes:
1. **Placeholder scan:** Any "TBD", "TODO", incomplete sections, or vague requirements? Fix them.
2. **Internal consistency:** Do any sections contradict each other? Does the architecture match the feature descriptions?
3. **Scope check:** Is this focused enough for a single implementation plan, or does it need decomposition?
4. **Ambiguity check:** Could any requirement be interpreted two different ways? If so, pick one and make it explicit.
Fix any issues inline. No need to re-review — just fix and move on.
**User Review Gate:**
After the spec review loop passes, ask the user to review the written spec before proceeding:
> "Spec written and committed to `<path>`. Please review it and let me know if you want to make any changes before we start writing out the implementation plan."
Wait for the user's response. If they request changes, make them and re-run the spec review loop. Only proceed once the user approves.
**Implementation:**
- Invoke the writing-plans skill to create a detailed implementation plan
- Do NOT invoke any other skill. writing-plans is the next step.
## Visual Companion
A browser-based companion for showing mockups, diagrams, and visual options during brainstorming. Available as a tool — not a mode. Accepting the companion means it's available for questions that benefit from visual treatment; it does NOT mean every question goes through the browser.
**Offering the companion (just-in-time):** Do NOT offer it upfront. Wait until a question would genuinely be clearer shown than told — a real mockup / layout / diagram question, not merely a UI *topic*. The first time that happens, offer it then, as its own message:
> "This next part might be easier if I show you — I can put together mockups, diagrams, and comparisons in a browser tab as we go. It's still new and can be token-intensive. Want me to? I'll open it for you."
**This offer MUST be its own message.** Only the offer — no clarifying question, summary, or other content. Wait for the user's response. If they accept, start the server with `--open` so their browser opens to the first screen automatically. If they decline, continue text-only and don't offer again unless they raise it.
**Per-question decision:** Even after the user accepts, decide FOR EACH QUESTION whether to use the browser or the terminal. The test: **would the user understand this better by seeing it than reading it?**
- **Use the browser** for content that IS visual — mockups, wireframes, layout comparisons, architecture diagrams, side-by-side visual designs
- **Use the terminal** for content that is text — requirements questions, conceptual choices, tradeoff lists, A/B/C/D text options, scope decisions
A question about a UI topic is not automatically a visual question. "What does personality mean in this context?" is a conceptual question — use the terminal. "Which wizard layout works better?" is a visual question — use the browser.
If they agree to the companion, read the detailed guide before proceeding:
`skills/brainstorming/visual-companion.md`
@@ -1,213 +0,0 @@
<!DOCTYPE html>
<html>
<head>
<meta charset="utf-8">
<title>Superpowers Brainstorming</title>
<style>
/*
* BRAINSTORM COMPANION FRAME TEMPLATE
*
* This template provides a consistent frame with:
* - OS-aware light/dark theming
* - Header branding and connection status
* - Scrollable main content area
* - CSS helpers for common UI patterns
*
* Content is injected via placeholder comment in #frame-content.
*/
* { box-sizing: border-box; margin: 0; padding: 0; }
html, body { height: 100%; overflow: hidden; }
/* ===== THEME VARIABLES ===== */
:root {
--bg-primary: #f5f5f7;
--bg-secondary: #ffffff;
--bg-tertiary: #e5e5e7;
--border: #d1d1d6;
--text-primary: #1d1d1f;
--text-secondary: #86868b;
--text-tertiary: #aeaeb2;
--accent: #0071e3;
--accent-hover: #0077ed;
--success: #34c759;
--warning: #ff9f0a;
--error: #ff3b30;
--selected-bg: #e8f4fd;
--selected-border: #0071e3;
}
@media (prefers-color-scheme: dark) {
:root {
--bg-primary: #1d1d1f;
--bg-secondary: #2d2d2f;
--bg-tertiary: #3d3d3f;
--border: #424245;
--text-primary: #f5f5f7;
--text-secondary: #86868b;
--text-tertiary: #636366;
--accent: #0a84ff;
--accent-hover: #409cff;
--selected-bg: rgba(10, 132, 255, 0.15);
--selected-border: #0a84ff;
}
}
body {
font-family: system-ui, -apple-system, BlinkMacSystemFont, sans-serif;
background: var(--bg-primary);
color: var(--text-primary);
display: flex;
flex-direction: column;
line-height: 1.5;
}
/* ===== FRAME STRUCTURE ===== */
.brand { display: flex; align-items: center; min-width: 0; overflow: hidden; color: var(--text-secondary); line-height: 1; }
.brand a { color: inherit; text-decoration: none; display: flex; align-items: center; gap: 0.5rem; min-width: 0; max-width: 100%; line-height: 1; }
.brand-copy { display: block; min-width: 0; overflow: hidden; text-overflow: ellipsis; white-space: nowrap; line-height: 1; transform: translateY(-1px); }
.brand-logo { display: block; height: 1em; width: auto; max-width: 180px; flex-shrink: 0; filter: invert(1); }
@media (prefers-color-scheme: dark) {
.brand-logo { filter: none; }
}
.status { font-size: 0.7rem; color: var(--status-color, var(--success)); display: flex; align-items: center; gap: 0.4rem; justify-self: end; white-space: nowrap; line-height: 1; }
.status::before { content: ''; width: 6px; height: 6px; background: var(--status-color, var(--success)); border-radius: 50%; }
.main { flex: 1; overflow-y: auto; }
#frame-content { padding: 2rem; min-height: 100%; }
.header {
background: var(--bg-secondary);
border-bottom: 1px solid var(--border);
padding: 0.5rem 1.5rem;
flex-shrink: 0;
display: grid;
grid-template-columns: minmax(0, 1fr) auto;
align-items: center;
gap: 1rem;
min-height: 42px;
}
.header .brand { justify-self: start; width: 100%; font-size: 0.75rem; line-height: 1; }
.header .status { grid-column: 2; line-height: 1; }
.header span {
font-size: 0.75rem;
color: var(--text-secondary);
}
.header .selected-text {
color: var(--accent);
font-weight: 500;
}
/* ===== TYPOGRAPHY ===== */
h2 { font-size: 1.5rem; font-weight: 600; margin-bottom: 0.5rem; }
h3 { font-size: 1.1rem; font-weight: 600; margin-bottom: 0.25rem; }
.subtitle { color: var(--text-secondary); margin-bottom: 1.5rem; }
.section { margin-bottom: 2rem; }
.label { font-size: 0.7rem; color: var(--text-secondary); text-transform: uppercase; letter-spacing: 0.05em; margin-bottom: 0.5rem; }
/* ===== OPTIONS (for A/B/C choices) ===== */
.options { display: flex; flex-direction: column; gap: 0.75rem; }
.option {
background: var(--bg-secondary);
border: 2px solid var(--border);
border-radius: 12px;
padding: 1rem 1.25rem;
cursor: pointer;
transition: all 0.15s ease;
display: flex;
align-items: flex-start;
gap: 1rem;
}
.option:hover { border-color: var(--accent); }
.option.selected { background: var(--selected-bg); border-color: var(--selected-border); }
.option .letter {
background: var(--bg-tertiary);
color: var(--text-secondary);
width: 1.75rem; height: 1.75rem;
border-radius: 6px;
display: flex; align-items: center; justify-content: center;
font-weight: 600; font-size: 0.85rem; flex-shrink: 0;
}
.option.selected .letter { background: var(--accent); color: white; }
.option .content { flex: 1; }
.option .content h3 { font-size: 0.95rem; margin-bottom: 0.15rem; }
.option .content p { color: var(--text-secondary); font-size: 0.85rem; margin: 0; }
/* ===== CARDS (for showing designs/mockups) ===== */
.cards { display: grid; grid-template-columns: repeat(auto-fit, minmax(280px, 1fr)); gap: 1rem; }
.card {
background: var(--bg-secondary);
border: 1px solid var(--border);
border-radius: 12px;
overflow: hidden;
cursor: pointer;
transition: all 0.15s ease;
}
.card:hover { border-color: var(--accent); transform: translateY(-2px); box-shadow: 0 4px 12px rgba(0,0,0,0.1); }
.card.selected { border-color: var(--selected-border); border-width: 2px; }
.card-image { background: var(--bg-tertiary); aspect-ratio: 16/10; display: flex; align-items: center; justify-content: center; }
.card-body { padding: 1rem; }
.card-body h3 { margin-bottom: 0.25rem; }
.card-body p { color: var(--text-secondary); font-size: 0.85rem; }
/* ===== MOCKUP CONTAINER ===== */
.mockup {
background: var(--bg-secondary);
border: 1px solid var(--border);
border-radius: 12px;
overflow: hidden;
margin-bottom: 1.5rem;
}
.mockup-header {
background: var(--bg-tertiary);
padding: 0.5rem 1rem;
font-size: 0.75rem;
color: var(--text-secondary);
border-bottom: 1px solid var(--border);
}
.mockup-body { padding: 1.5rem; }
/* ===== SPLIT VIEW (side-by-side comparison) ===== */
.split { display: grid; grid-template-columns: 1fr 1fr; gap: 1.5rem; }
@media (max-width: 700px) { .split { grid-template-columns: 1fr; } }
/* ===== PROS/CONS ===== */
.pros-cons { display: grid; grid-template-columns: 1fr 1fr; gap: 1rem; margin: 1rem 0; }
.pros, .cons { background: var(--bg-secondary); border-radius: 8px; padding: 1rem; }
.pros h4 { color: var(--success); font-size: 0.85rem; margin-bottom: 0.5rem; }
.cons h4 { color: var(--error); font-size: 0.85rem; margin-bottom: 0.5rem; }
.pros ul, .cons ul { margin-left: 1.25rem; font-size: 0.85rem; color: var(--text-secondary); }
.pros li, .cons li { margin-bottom: 0.25rem; }
/* ===== PLACEHOLDER (for mockup areas) ===== */
.placeholder {
background: var(--bg-tertiary);
border: 2px dashed var(--border);
border-radius: 8px;
padding: 2rem;
text-align: center;
color: var(--text-tertiary);
}
/* ===== INLINE MOCKUP ELEMENTS ===== */
.mock-nav { background: var(--accent); color: white; padding: 0.75rem 1rem; display: flex; gap: 1.5rem; font-size: 0.9rem; }
.mock-sidebar { background: var(--bg-tertiary); padding: 1rem; min-width: 180px; }
.mock-content { padding: 1.5rem; flex: 1; }
.mock-button { background: var(--accent); color: white; border: none; padding: 0.5rem 1rem; border-radius: 6px; font-size: 0.85rem; }
.mock-input { background: var(--bg-primary); border: 1px solid var(--border); border-radius: 6px; padding: 0.5rem; width: 100%; }
</style>
</head>
<body>
<div class="header">
<!-- BRANDING -->
<div class="status">Connecting…</div>
</div>
<div class="main">
<div id="frame-content">
<!-- CONTENT -->
</div>
</div>
</body>
</html>
@@ -1,167 +0,0 @@
(function() {
const MIN_RECONNECT_MS = 500;
const MAX_RECONNECT_MS = 30000;
const TOMBSTONE_AFTER_MS = 15000; // show the "paused" overlay after this long disconnected
// Pure: next backoff delay (doubles, capped). Exported for unit tests.
function nextReconnectDelay(current, max) {
return Math.min(current * 2, max);
}
if (typeof module !== 'undefined' && module.exports) {
module.exports = { nextReconnectDelay, MIN_RECONNECT_MS, MAX_RECONNECT_MS, TOMBSTONE_AFTER_MS };
}
// Everything below is browser-only; bail out when loaded in Node (tests).
if (typeof window === 'undefined') return;
let ws = null;
let eventQueue = [];
let reconnectDelay = MIN_RECONNECT_MS;
let reconnectTimer = null;
let disconnectedSince = null;
let everConnected = false;
let tombstoneShown = false;
function sessionKey() {
try {
return window.sessionStorage && window.sessionStorage.getItem('brainstorm-session-key');
} catch (e) {}
return null;
}
function websocketUrl() {
const key = sessionKey();
return 'ws://' + window.location.host + (key ? '/?key=' + encodeURIComponent(key) : '');
}
function reloadAfterRecovery() {
const key = sessionKey();
if (key) {
window.location.replace('/?key=' + encodeURIComponent(key));
} else {
window.location.reload();
}
}
// Reflect connection state in the frame's status pill (absent on full-doc screens).
function setStatus(state) {
const el = document.querySelector('.status');
if (!el) return;
const map = {
connecting: ['Connecting…', 'var(--text-tertiary)'],
connected: ['Connected', 'var(--success)'],
reconnecting: ['Reconnecting…', 'var(--warning)'],
disconnected: ['Disconnected', 'var(--error)']
};
const [text, color] = map[state] || map.disconnected;
el.textContent = text;
el.style.setProperty('--status-color', color);
}
// Self-styled so it works on framed and full-document screens alike.
function showTombstone() {
if (tombstoneShown) return;
tombstoneShown = true;
const el = document.createElement('div');
el.id = 'bs-tombstone';
el.style.cssText = 'position:fixed;inset:0;z-index:99999;display:flex;' +
'align-items:center;justify-content:center;padding:2rem;text-align:center;' +
'background:rgba(20,20,22,0.92);color:#f5f5f7;font-family:system-ui,sans-serif';
el.innerHTML = '<div style="max-width:480px">' +
'<h2 style="margin:0 0 .5rem;font-weight:600">Companion paused</h2>' +
'<p style="margin:0;opacity:.85">This brainstorm companion has stopped. ' +
'Ask your coding agent to bring it back — this page reconnects automatically.</p></div>';
if (document.body) document.body.appendChild(el);
}
function connect() {
if (reconnectTimer) { clearTimeout(reconnectTimer); reconnectTimer = null; }
setStatus(everConnected ? 'reconnecting' : 'connecting');
ws = new WebSocket(websocketUrl());
ws.onopen = () => {
const recovered = tombstoneShown;
everConnected = true;
disconnectedSince = null;
reconnectDelay = MIN_RECONNECT_MS;
tombstoneShown = false;
setStatus('connected');
eventQueue.forEach(e => ws.send(JSON.stringify(e)));
eventQueue = [];
// Recovered from a tombstoned outage (e.g. the server restarted on the same
// port) — reload through the keyed bootstrap when possible so the cookie is
// refreshed before the visible URL returns to bare /.
if (recovered) reloadAfterRecovery();
};
ws.onmessage = (msg) => {
let data;
try { data = JSON.parse(msg.data); } catch (e) { return; }
if (data.type === 'reload') window.location.reload();
};
ws.onclose = () => {
ws = null;
if (disconnectedSince === null) disconnectedSince = Date.now();
if (Date.now() - disconnectedSince >= TOMBSTONE_AFTER_MS) {
setStatus('disconnected');
showTombstone();
} else {
setStatus('reconnecting');
}
reconnectTimer = setTimeout(connect, reconnectDelay);
reconnectDelay = nextReconnectDelay(reconnectDelay, MAX_RECONNECT_MS);
};
// Let onclose own reconnection so we don't schedule it twice.
ws.onerror = () => { try { ws.close(); } catch (e) {} };
}
function sendEvent(event) {
event.timestamp = Date.now();
if (ws && ws.readyState === WebSocket.OPEN) {
ws.send(JSON.stringify(event));
} else {
eventQueue.push(event);
}
}
// Capture clicks on choice elements
document.addEventListener('click', (e) => {
const target = e.target.closest('[data-choice]');
if (!target) return;
sendEvent({
type: 'click',
text: target.textContent.trim(),
choice: target.dataset.choice,
id: target.id || null
});
});
// Frame UI: selection tracking
window.selectedChoice = null;
window.toggleSelect = function(el) {
const container = el.closest('.options') || el.closest('.cards');
const multi = container && container.dataset.multiselect !== undefined;
if (container && !multi) {
container.querySelectorAll('.option, .card').forEach(o => o.classList.remove('selected'));
}
if (multi) {
el.classList.toggle('selected');
} else {
el.classList.add('selected');
}
window.selectedChoice = el.dataset.choice;
};
// Expose API for explicit use
window.brainstorm = {
send: sendEvent,
choice: (value, metadata = {}) => sendEvent({ type: 'choice', value, ...metadata })
};
connect();
})();
@@ -1,723 +0,0 @@
const crypto = require('crypto');
const http = require('http');
const fs = require('fs');
const path = require('path');
// ========== WebSocket Protocol (RFC 6455) ==========
const OPCODES = { TEXT: 0x01, CLOSE: 0x08, PING: 0x09, PONG: 0x0A };
const WS_MAGIC = '258EAFA5-E914-47DA-95CA-C5AB0DC85B11';
const MAX_FRAME_PAYLOAD_BYTES = 10 * 1024 * 1024;
function computeAcceptKey(clientKey) {
return crypto.createHash('sha1').update(clientKey + WS_MAGIC).digest('base64');
}
function encodeFrame(opcode, payload) {
const fin = 0x80;
const len = payload.length;
let header;
if (len < 126) {
header = Buffer.alloc(2);
header[0] = fin | opcode;
header[1] = len;
} else if (len < 65536) {
header = Buffer.alloc(4);
header[0] = fin | opcode;
header[1] = 126;
header.writeUInt16BE(len, 2);
} else {
header = Buffer.alloc(10);
header[0] = fin | opcode;
header[1] = 127;
header.writeBigUInt64BE(BigInt(len), 2);
}
return Buffer.concat([header, payload]);
}
function decodeFrame(buffer) {
if (buffer.length < 2) return null;
const secondByte = buffer[1];
const opcode = buffer[0] & 0x0F;
const masked = (secondByte & 0x80) !== 0;
let payloadLen = secondByte & 0x7F;
let offset = 2;
if (!masked) throw new Error('Client frames must be masked');
if (payloadLen === 126) {
if (buffer.length < 4) return null;
payloadLen = buffer.readUInt16BE(2);
offset = 4;
} else if (payloadLen === 127) {
if (buffer.length < 10) return null;
const extendedLen = buffer.readBigUInt64BE(2);
if (extendedLen > BigInt(MAX_FRAME_PAYLOAD_BYTES)) {
throw new Error('WebSocket frame payload exceeds maximum allowed size');
}
payloadLen = Number(extendedLen);
offset = 10;
}
if (payloadLen > MAX_FRAME_PAYLOAD_BYTES) {
throw new Error('WebSocket frame payload exceeds maximum allowed size');
}
const maskOffset = offset;
const dataOffset = offset + 4;
const totalLen = dataOffset + payloadLen;
if (buffer.length < totalLen) return null;
const mask = buffer.slice(maskOffset, dataOffset);
const data = Buffer.alloc(payloadLen);
for (let i = 0; i < payloadLen; i++) {
data[i] = buffer[dataOffset + i] ^ mask[i % 4];
}
return { opcode, payload: data, bytesConsumed: totalLen };
}
// ========== Configuration ==========
const PORT_FILE = process.env.BRAINSTORM_PORT_FILE || null;
const randomPort = () => 49152 + Math.floor(Math.random() * 16383);
// Prefer an explicit port, else the port this session last bound (so a restart
// reuses it and an already-open browser tab reconnects), else a random high port.
function preferredPort() {
if (process.env.BRAINSTORM_PORT) return Number(process.env.BRAINSTORM_PORT);
if (PORT_FILE) {
try {
const p = Number(fs.readFileSync(PORT_FILE, 'utf-8').trim());
if (Number.isInteger(p) && p > 1023 && p < 65536) return p;
} catch (e) { /* no prior port recorded */ }
}
return randomPort();
}
let PORT = preferredPort();
const HOST = process.env.BRAINSTORM_HOST || '127.0.0.1';
const URL_HOST = process.env.BRAINSTORM_URL_HOST || (HOST === '127.0.0.1' ? 'localhost' : HOST);
const SESSION_DIR = process.env.BRAINSTORM_DIR || '/tmp/brainstorm';
const CONTENT_DIR = path.join(SESSION_DIR, 'content');
const STATE_DIR = path.join(SESSION_DIR, 'state');
const SUPERPOWERS_VERSION = readSuperpowersVersion();
const SUPERPOWERS_BRAND_IMAGE_URL = 'https://primeradiant.com/brand/superpowers-visual-brainstorming-logo.png';
const TELEMETRY_DISABLE_ENV_VARS = [
'SUPERPOWERS_DISABLE_TELEMETRY',
'DISABLE_TELEMETRY',
'CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC'
];
const SUPERPOWERS_TELEMETRY_DISABLED = TELEMETRY_DISABLE_ENV_VARS.some(name => isTruthyEnv(process.env[name]));
let ownerPid = process.env.BRAINSTORM_OWNER_PID ? Number(process.env.BRAINSTORM_OWNER_PID) : null;
// Per-session secret key. The companion is reachable by any local browser tab
// and, when bound to a non-loopback host, by any host that can route to it.
// The key authenticates the real client uniformly across loopback, tunnel, and
// remote binds — and defeats DNS rebinding — where a Host/Origin allowlist
// cannot. It rides the served URL as ?key= and is mirrored into a cookie on
// first load so same-origin subresources and the WebSocket carry it for free.
// Persisted alongside the port (BRAINSTORM_TOKEN_FILE) so a restart keeps the
// same key and an already-open tab's cookie still validates.
const TOKEN_FILE = process.env.BRAINSTORM_TOKEN_FILE || null;
function generateToken() {
return crypto.randomBytes(32).toString('hex');
}
function chmodOwnerOnly(file) {
try { fs.chmodSync(file, 0o600); } catch (e) { /* best effort */ }
}
function initialToken() {
if (process.env.BRAINSTORM_TOKEN) {
return { value: process.env.BRAINSTORM_TOKEN, source: 'env' };
}
if (TOKEN_FILE) {
try {
const t = fs.readFileSync(TOKEN_FILE, 'utf-8').trim();
if (/^[0-9a-f]{32,}$/i.test(t)) {
chmodOwnerOnly(TOKEN_FILE);
return { value: t, source: 'file' };
}
} catch (e) { /* no prior token recorded */ }
}
return { value: generateToken(), source: 'generated' };
}
const tokenInfo = initialToken();
let TOKEN = tokenInfo.value;
let tokenSource = tokenInfo.source;
let COOKIE_NAME = 'brainstorm-key-' + PORT; // refined to the actual bound port in onListen
const MIME_TYPES = {
'.html': 'text/html', '.css': 'text/css', '.js': 'application/javascript',
'.json': 'application/json', '.png': 'image/png', '.jpg': 'image/jpeg',
'.jpeg': 'image/jpeg', '.gif': 'image/gif', '.svg': 'image/svg+xml'
};
// ========== Templates and Constants ==========
function waitingPage() {
return renderBranding(`<!DOCTYPE html>
<html>
<head><meta charset="utf-8"><title>Brainstorm Companion</title>
<style>
body { font-family: system-ui, sans-serif; padding: 2rem; max-width: 800px; margin: 0 auto; }
h1 { color: #333; } p { color: #666; }
.brand { display: flex; align-items: center; min-width: 0; overflow: hidden; margin-bottom: 1.5rem; color: #666; font-size: 0.9rem; line-height: 1; }
.brand a { color: inherit; text-decoration: none; display: flex; align-items: center; gap: 0.5rem; min-width: 0; max-width: 100%; line-height: 1; }
.brand-copy { display: block; min-width: 0; overflow: hidden; text-overflow: ellipsis; white-space: nowrap; line-height: 1; transform: translateY(-1px); }
.brand-logo { display: block; height: 1em; width: auto; max-width: 180px; filter: invert(1); }
</style>
</head>
<body><!-- BRANDING --><h1>Brainstorm Companion</h1>
<p>Waiting for the agent to push a screen...</p></body></html>`);
}
const FORBIDDEN_PAGE = `<!DOCTYPE html>
<html>
<head><meta charset="utf-8"><title>Session key required</title>
<style>body { font-family: system-ui, sans-serif; padding: 2rem; max-width: 800px; margin: 0 auto; }
h1 { color: #333; } p { color: #666; } code { background: #f0f0f0; padding: 0.1em 0.3em; border-radius: 4px; }</style>
</head>
<body><h1>Session key required</h1>
<p>This page needs the full URL your coding agent gave you, including the
<code>?key=&hellip;</code> part. Copy the complete URL and open it again.</p></body></html>`;
function bootstrapPage(key) {
const jsonKey = JSON.stringify(String(key));
return `<!DOCTYPE html>
<html>
<head><meta charset="utf-8"><title>Opening Brainstorm Companion</title></head>
<body>
<script>
try { sessionStorage.setItem('brainstorm-session-key', ${jsonKey}); } catch (e) {}
location.replace('/');
</script>
</body>
</html>`;
}
const frameTemplate = fs.readFileSync(path.join(__dirname, 'frame-template.html'), 'utf-8');
const helperScript = fs.readFileSync(path.join(__dirname, 'helper.js'), 'utf-8');
const helperInjection = '<script>\n' + helperScript + '\n</script>';
// ========== Helper Functions ==========
function readSuperpowersVersion() {
const root = path.join(__dirname, '../../..');
const manifests = [
path.join(root, 'package.json'),
path.join(root, '.codex-plugin/plugin.json')
];
for (const manifest of manifests) {
try {
const data = JSON.parse(fs.readFileSync(manifest, 'utf-8'));
if (data.version) return String(data.version);
} catch (e) {
// Packaged Codex plugins omit package.json; try the next manifest.
}
}
return 'unknown';
}
function isTruthyEnv(value) {
if (!value) return false;
const normalized = String(value).trim().toLowerCase();
if (!normalized) return false;
return !['0', 'false', 'no', 'off'].includes(normalized);
}
function escapeHtmlText(value) {
return String(value)
.replace(/&/g, '&amp;')
.replace(/</g, '&lt;')
.replace(/>/g, '&gt;')
.replace(/"/g, '&quot;');
}
function brandMarkup() {
const version = escapeHtmlText(SUPERPOWERS_VERSION);
const text = SUPERPOWERS_TELEMETRY_DISABLED
? 'Prime Radiant Superpowers v' + version
: 'Superpowers v' + version;
const logo = SUPERPOWERS_TELEMETRY_DISABLED
? ''
: '<img class="brand-logo" src="' + SUPERPOWERS_BRAND_IMAGE_URL + '?v=' + encodeURIComponent(SUPERPOWERS_VERSION) + '" alt="Prime Radiant" referrerpolicy="no-referrer" decoding="async">';
return '<div class="brand"><a href="https://github.com/obra/superpowers">' + logo + '<span class="brand-copy">' + text + '</span></a></div>';
}
function renderBranding(html) {
return html.split('<!-- BRANDING -->').join(brandMarkup());
}
function isFullDocument(html) {
const trimmed = html.trimStart().toLowerCase();
return trimmed.startsWith('<!doctype') || trimmed.startsWith('<html');
}
function wrapInFrame(content) {
return renderBranding(frameTemplate).replace('<!-- CONTENT -->', content);
}
function getNewestScreen() {
const files = fs.readdirSync(CONTENT_DIR)
.filter(f => !f.startsWith('.') && f.endsWith('.html'))
.map(f => {
const fp = path.join(CONTENT_DIR, f);
if (!isRegularFileInsideContentDir(fp)) return null;
return { path: fp, mtime: fs.statSync(fp).mtime.getTime() };
})
.filter(Boolean)
.sort((a, b) => b.mtime - a.mtime);
return files.length > 0 ? files[0].path : null;
}
function urlHostForHttp(host) {
const h = String(host);
if (h.startsWith('[') && h.endsWith(']')) return h;
return h.includes(':') ? '[' + h + ']' : h;
}
function companionUrl() {
return 'http://' + urlHostForHttp(URL_HOST) + ':' + PORT + '/?key=' + TOKEN;
}
function browserLauncherForPlatform(url, {
platform = process.platform,
osRelease = require('os').release(),
env = process.env
} = {}) {
const isWSL = platform === 'linux' && /microsoft/i.test(osRelease);
if (platform === 'darwin') return { bin: 'open', args: [url] };
if (platform === 'win32' || isWSL) {
return { bin: 'rundll32.exe', args: ['url.dll,FileProtocolHandler', url] };
}
if (env.DISPLAY || env.WAYLAND_DISPLAY) return { bin: 'xdg-open', args: [url] };
return null;
}
function isRegularFileInsideContentDir(filePath) {
let stat, realContentDir, realFilePath;
try {
stat = fs.lstatSync(filePath);
if (stat.isSymbolicLink()) return false;
if (!stat.isFile()) return false;
if (stat.nlink !== 1) return false;
realContentDir = fs.realpathSync(CONTENT_DIR);
realFilePath = fs.realpathSync(filePath);
} catch (e) {
return false;
}
return realFilePath.startsWith(realContentDir + path.sep);
}
// ========== Authentication ==========
function timingSafeEqualStr(a, b) {
const ab = Buffer.from(String(a));
const bb = Buffer.from(String(b));
if (ab.length !== bb.length) return false;
return crypto.timingSafeEqual(ab, bb);
}
function parseCookies(header) {
const out = {};
if (!header) return out;
for (const part of header.split(';')) {
const eq = part.indexOf('=');
if (eq < 0) continue;
out[part.slice(0, eq).trim()] = part.slice(eq + 1).trim();
}
return out;
}
// A request is authorized if it carries the session key as ?key= or as the
// session cookie. Both are compared in constant time.
function isAuthorized(req) {
const q = req.url.indexOf('?');
if (q >= 0) {
const params = new URLSearchParams(req.url.slice(q + 1));
if (params.has('key')) {
const key = params.get('key');
return Boolean(key && timingSafeEqualStr(key, TOKEN));
}
}
const cookie = parseCookies(req.headers['cookie'])[COOKIE_NAME];
if (cookie && timingSafeEqualStr(cookie, TOKEN)) return true;
return false;
}
function pathnameOf(url) {
const q = url.indexOf('?');
return q >= 0 ? url.slice(0, q) : url;
}
function queryKey(url) {
const q = url.indexOf('?');
if (q < 0) return null;
return new URLSearchParams(url.slice(q + 1)).get('key');
}
function securityHeaders(headers = {}) {
return {
'Referrer-Policy': 'no-referrer',
'Cache-Control': 'no-store',
'X-Frame-Options': 'DENY',
'Content-Security-Policy': "frame-ancestors 'none'",
'Cross-Origin-Resource-Policy': 'same-origin',
...headers
};
}
function isAllowedWebSocketOrigin(req) {
const origin = req.headers.origin;
if (!origin) return true;
const host = req.headers.host;
if (!host) return false;
return origin === 'http://' + host;
}
// ========== HTTP Request Handler ==========
function handleRequest(req, res) {
if (!isAuthorized(req)) {
res.writeHead(403, securityHeaders({ 'Content-Type': 'text/html; charset=utf-8' }));
res.end(FORBIDDEN_PAGE);
return;
}
touchActivity(); // only authorized requests count as activity
// Mirror the key into a cookie so same-origin subresources (/files/*) can
// authenticate after bootstrap. HttpOnly keeps it away from page scripts; the
// WebSocket Origin check below is what blocks cross-origin localhost injection.
res.setHeader('Set-Cookie',
COOKIE_NAME + '=' + TOKEN + '; HttpOnly; SameSite=Strict; Path=/');
const pathname = pathnameOf(req.url);
const keyFromQuery = queryKey(req.url);
if (req.method === 'GET' && pathname === '/' && keyFromQuery && timingSafeEqualStr(keyFromQuery, TOKEN)) {
res.writeHead(200, securityHeaders({ 'Content-Type': 'text/html; charset=utf-8' }));
res.end(bootstrapPage(keyFromQuery));
} else if (req.method === 'GET' && pathname === '/') {
const screenFile = getNewestScreen();
let html = screenFile
? (raw => isFullDocument(raw) ? raw : wrapInFrame(raw))(fs.readFileSync(screenFile, 'utf-8'))
: waitingPage();
if (html.includes('</body>')) {
html = html.replace('</body>', helperInjection + '\n</body>');
} else {
html += helperInjection;
}
res.writeHead(200, securityHeaders({ 'Content-Type': 'text/html; charset=utf-8' }));
res.end(html);
} else if (req.method === 'GET' && pathname.startsWith('/files/')) {
const fileName = path.basename(pathname.slice(7));
const filePath = path.join(CONTENT_DIR, fileName);
// Reject empty/dotfile names and anything that isn't a regular file —
// `/files/` would otherwise resolve to CONTENT_DIR and crash readFileSync (EISDIR).
if (!fileName || fileName.startsWith('.') || !isRegularFileInsideContentDir(filePath)) {
res.writeHead(404, securityHeaders());
res.end('Not found');
return;
}
const ext = path.extname(filePath).toLowerCase();
const contentType = MIME_TYPES[ext] || 'application/octet-stream';
res.writeHead(200, securityHeaders({ 'Content-Type': contentType }));
res.end(fs.readFileSync(filePath));
} else {
res.writeHead(404, securityHeaders());
res.end('Not found');
}
}
// ========== WebSocket Connection Handling ==========
const clients = new Set();
function handleUpgrade(req, socket) {
if (!isAuthorized(req) || !isAllowedWebSocketOrigin(req)) { socket.destroy(); return; }
const key = req.headers['sec-websocket-key'];
if (!key) { socket.destroy(); return; }
const accept = computeAcceptKey(key);
socket.write(
'HTTP/1.1 101 Switching Protocols\r\n' +
'Upgrade: websocket\r\n' +
'Connection: Upgrade\r\n' +
'Sec-WebSocket-Accept: ' + accept + '\r\n\r\n'
);
let buffer = Buffer.alloc(0);
clients.add(socket);
socket.on('data', (chunk) => {
buffer = Buffer.concat([buffer, chunk]);
while (buffer.length > 0) {
let result;
try {
result = decodeFrame(buffer);
} catch (e) {
socket.end(encodeFrame(OPCODES.CLOSE, Buffer.alloc(0)));
clients.delete(socket);
return;
}
if (!result) break;
buffer = buffer.slice(result.bytesConsumed);
switch (result.opcode) {
case OPCODES.TEXT:
handleMessage(result.payload.toString());
break;
case OPCODES.CLOSE:
socket.end(encodeFrame(OPCODES.CLOSE, Buffer.alloc(0)));
clients.delete(socket);
return;
case OPCODES.PING:
socket.write(encodeFrame(OPCODES.PONG, result.payload));
break;
case OPCODES.PONG:
break;
default: {
const closeBuf = Buffer.alloc(2);
closeBuf.writeUInt16BE(1003);
socket.end(encodeFrame(OPCODES.CLOSE, closeBuf));
clients.delete(socket);
return;
}
}
}
});
socket.on('close', () => clients.delete(socket));
socket.on('error', () => clients.delete(socket));
}
function handleMessage(text) {
let event;
try {
event = JSON.parse(text);
} catch (e) {
console.error('Failed to parse WebSocket message:', e.message);
return;
}
touchActivity();
console.log(JSON.stringify({ source: 'user-event', ...event }));
if (event && event.choice) {
const eventsFile = path.join(STATE_DIR, 'events');
fs.appendFileSync(eventsFile, JSON.stringify(event) + '\n');
}
}
function broadcast(msg) {
const frame = encodeFrame(OPCODES.TEXT, Buffer.from(JSON.stringify(msg)));
for (const socket of clients) {
try { socket.write(frame); } catch (e) { clients.delete(socket); }
}
}
// Best-effort: open the user's browser the first time a screen is actually ready
// to show. Skips when disabled, on a non-loopback (remote) bind, or when a
// browser is already connected. Override the launcher with BRAINSTORM_OPEN_CMD.
let browserOpened = false;
function maybeOpenBrowser() {
if (browserOpened) return;
browserOpened = true;
if (!process.env.BRAINSTORM_OPEN) return; // opt-in: only after the user approves the companion
if (HOST !== '127.0.0.1' && HOST !== 'localhost') return;
if (clients.size > 0) return; // the user already opened it
const url = companionUrl(); // must carry the key or the gate 403s it
const cp = require('child_process');
// Operator-provided launcher: run as given (this env var is trusted operator input).
if (process.env.BRAINSTORM_OPEN_CMD) {
try { cp.exec(process.env.BRAINSTORM_OPEN_CMD + ' ' + JSON.stringify(url), () => {}); } catch (e) { /* best effort */ }
return;
}
// Platform launchers: pass the URL as an argv element via execFile (no shell),
// so a url-host containing shell metacharacters can't inject a command.
const launcher = browserLauncherForPlatform(url);
if (!launcher) return; // headless: nothing to open
try { cp.execFile(launcher.bin, launcher.args, () => {}); } catch (e) { /* best effort */ }
}
// ========== Activity Tracking ==========
// Idle timeout: shut down after this long with no activity. Default 4 hours;
// override with BRAINSTORM_IDLE_TIMEOUT_MS (start-server.sh: --idle-timeout-minutes).
const IDLE_TIMEOUT_MS = (() => {
const ms = Number(process.env.BRAINSTORM_IDLE_TIMEOUT_MS);
return Number.isFinite(ms) && ms > 0 ? ms : 4 * 60 * 60 * 1000;
})();
// How often the watchdog checks for owner-death / idleness. Configurable mainly
// so tests can run fast; production default is 60s.
const LIFECYCLE_CHECK_MS = (() => {
const ms = Number(process.env.BRAINSTORM_LIFECYCLE_CHECK_MS);
return Number.isFinite(ms) && ms > 0 ? ms : 60 * 1000;
})();
let lastActivity = Date.now();
function touchActivity() {
lastActivity = Date.now();
}
// ========== File Watching ==========
const debounceTimers = new Map();
// ========== Server Startup ==========
function startServer() {
if (!fs.existsSync(CONTENT_DIR)) fs.mkdirSync(CONTENT_DIR, { recursive: true });
if (!fs.existsSync(STATE_DIR)) fs.mkdirSync(STATE_DIR, { recursive: true });
// Track known files to distinguish new screens from updates.
// macOS fs.watch reports 'rename' for both new files and overwrites,
// so we can't rely on eventType alone.
const knownFiles = new Set(
fs.readdirSync(CONTENT_DIR).filter(f => !f.startsWith('.') && f.endsWith('.html'))
);
const server = http.createServer(handleRequest);
server.on('upgrade', handleUpgrade);
const watcher = fs.watch(CONTENT_DIR, (eventType, filename) => {
if (!filename || filename.startsWith('.') || !filename.endsWith('.html')) return;
if (debounceTimers.has(filename)) clearTimeout(debounceTimers.get(filename));
debounceTimers.set(filename, setTimeout(() => {
debounceTimers.delete(filename);
const filePath = path.join(CONTENT_DIR, filename);
if (!fs.existsSync(filePath)) return; // file was deleted
touchActivity();
if (!knownFiles.has(filename)) {
knownFiles.add(filename);
const eventsFile = path.join(STATE_DIR, 'events');
if (fs.existsSync(eventsFile)) fs.unlinkSync(eventsFile);
console.log(JSON.stringify({ type: 'screen-added', file: filePath }));
maybeOpenBrowser();
} else {
console.log(JSON.stringify({ type: 'screen-updated', file: filePath }));
}
broadcast({ type: 'reload' });
}, 100));
});
watcher.on('error', (err) => console.error('fs.watch error:', err.message));
function shutdown(reason) {
console.log(JSON.stringify({ type: 'server-stopped', reason }));
const infoFile = path.join(STATE_DIR, 'server-info');
if (fs.existsSync(infoFile)) fs.unlinkSync(infoFile);
fs.writeFileSync(
path.join(STATE_DIR, 'server-stopped'),
JSON.stringify({ reason, timestamp: Date.now() }) + '\n'
);
watcher.close();
clearInterval(lifecycleCheck);
// Close any upgraded WebSocket sockets so server.close() can complete and
// the process actually exits instead of lingering on an open connection.
for (const socket of clients) {
try { socket.destroy(); } catch (e) { /* already gone */ }
}
server.close(() => process.exit(0));
}
function ownerAlive() {
if (!ownerPid) return true;
try { process.kill(ownerPid, 0); return true; } catch (e) { return e.code === 'EPERM'; }
}
// Periodically exit if the owner process died or we've been idle too long.
const lifecycleCheck = setInterval(() => {
if (!ownerAlive()) shutdown('owner process exited');
else if (Date.now() - lastActivity > IDLE_TIMEOUT_MS) shutdown('idle timeout');
}, LIFECYCLE_CHECK_MS);
lifecycleCheck.unref();
// Validate owner PID at startup. If it's already dead, the PID resolution
// was wrong (common on WSL, Tailscale SSH, and cross-user scenarios).
// Disable monitoring and rely on the idle timeout instead.
if (ownerPid) {
try { process.kill(ownerPid, 0); }
catch (e) {
if (e.code !== 'EPERM') {
console.log(JSON.stringify({ type: 'owner-pid-invalid', pid: ownerPid, reason: 'dead at startup' }));
ownerPid = null;
}
}
}
// If the preferred port is already taken (e.g. a previous server is still
// alive), fall back to a random port once instead of failing.
let triedFallback = false;
function onListen() {
// Cookie name keys on the ACTUAL bound port (may differ from the preferred
// one after an EADDRINUSE fallback) so it can't collide with another server's
// cookie in the shared localhost jar.
COOKIE_NAME = 'brainstorm-key-' + PORT;
// Record the bound port AND token so the next restart of this session reuses
// them — but ONLY when we got our preferred port. On a fallback we bound a
// *different* port because someone else holds the preferred one; persisting
// would overwrite the shared files and strand that other session's open tab.
if (PORT_FILE && !triedFallback) {
try { fs.writeFileSync(PORT_FILE, String(PORT)); } catch (e) { /* best effort */ }
if (TOKEN_FILE) {
try {
fs.writeFileSync(TOKEN_FILE, TOKEN, { mode: 0o600 });
chmodOwnerOnly(TOKEN_FILE);
} catch (e) { /* best effort */ }
}
}
const info = JSON.stringify({
type: 'server-started', port: Number(PORT), host: HOST,
url_host: URL_HOST, url: companionUrl(),
screen_dir: CONTENT_DIR, state_dir: STATE_DIR, idle_timeout_ms: IDLE_TIMEOUT_MS
});
console.log(info);
// server-info embeds the key — keep it owner-only.
fs.writeFileSync(path.join(STATE_DIR, 'server-info'), info + '\n', { mode: 0o600 });
}
server.on('error', (err) => {
if (err.code === 'EADDRINUSE' && !triedFallback) {
if (tokenSource === 'env') {
console.error('Server failed to bind: preferred port is in use and BRAINSTORM_TOKEN is set; refusing fallback with explicit token');
process.exit(1);
}
triedFallback = true;
PORT = randomPort();
if (tokenSource === 'file') {
TOKEN = generateToken();
tokenSource = 'generated-fallback';
}
server.listen(PORT, HOST, onListen);
} else {
console.error('Server failed to bind:', err.message);
process.exit(1);
}
});
server.listen(PORT, HOST, onListen);
}
if (require.main === module) {
startServer();
}
module.exports = {
computeAcceptKey,
encodeFrame,
decodeFrame,
browserLauncherForPlatform,
OPCODES,
MAX_FRAME_PAYLOAD_BYTES
};
@@ -1,209 +0,0 @@
#!/usr/bin/env bash
# Start the brainstorm server and output connection info
# Usage: start-server.sh [--project-dir <path>] [--host <bind-host>] [--url-host <display-host>] [--foreground] [--background]
#
# Starts server on a random high port, outputs JSON with URL.
# Each session gets its own directory to avoid conflicts.
#
# Options:
# --project-dir <path> Store session files under <path>/.superpowers/brainstorm/
# instead of /tmp. Files persist after server stops.
# --host <bind-host> Host/interface to bind (default: 127.0.0.1).
# Use 0.0.0.0 in remote/containerized environments.
# --url-host <host> Hostname shown in returned URL JSON.
# --idle-timeout-minutes <n> Shut down after n minutes idle (default 240 = 4h).
# --open Auto-open the browser on the first screen (use only
# after the user approves the visual companion).
# --foreground Run server in the current terminal (no backgrounding).
# --background Force background mode (overrides Codex auto-foreground).
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
# Parse arguments
PROJECT_DIR=""
FOREGROUND="false"
FORCE_BACKGROUND="false"
BIND_HOST="127.0.0.1"
URL_HOST=""
IDLE_TIMEOUT_MINUTES=""
while [[ $# -gt 0 ]]; do
case "$1" in
--project-dir)
PROJECT_DIR="$2"
shift 2
;;
--host)
BIND_HOST="$2"
shift 2
;;
--url-host)
URL_HOST="$2"
shift 2
;;
--idle-timeout-minutes)
IDLE_TIMEOUT_MINUTES="$2"
shift 2
;;
--open)
export BRAINSTORM_OPEN=1
shift
;;
--foreground|--no-daemon)
FOREGROUND="true"
shift
;;
--background|--daemon)
FORCE_BACKGROUND="true"
shift
;;
*)
echo "{\"error\": \"Unknown argument: $1\"}"
exit 1
;;
esac
done
if [[ -z "$URL_HOST" ]]; then
if [[ "$BIND_HOST" == "127.0.0.1" || "$BIND_HOST" == "localhost" ]]; then
URL_HOST="localhost"
else
URL_HOST="$BIND_HOST"
fi
fi
if [[ -n "$IDLE_TIMEOUT_MINUTES" ]]; then
if ! [[ "$IDLE_TIMEOUT_MINUTES" =~ ^[0-9]+$ ]] || [[ "$IDLE_TIMEOUT_MINUTES" -lt 1 ]]; then
echo "{\"error\": \"--idle-timeout-minutes must be a positive integer\"}"
exit 1
fi
export BRAINSTORM_IDLE_TIMEOUT_MS=$(( IDLE_TIMEOUT_MINUTES * 60 * 1000 ))
fi
is_windows_like_shell() {
case "${OSTYPE:-}" in
msys*|cygwin*|mingw*) return 0 ;;
esac
if [[ -n "${MSYSTEM:-}" ]]; then
return 0
fi
local uname_s
uname_s="$(uname -s 2>/dev/null || true)"
case "$uname_s" in
MSYS*|MINGW*|CYGWIN*) return 0 ;;
esac
return 1
}
# Some environments reap detached/background processes. Auto-foreground when detected.
if [[ -n "${CODEX_CI:-}" && "$FOREGROUND" != "true" && "$FORCE_BACKGROUND" != "true" ]]; then
FOREGROUND="true"
fi
# Windows/Git Bash reaps nohup background processes. Auto-foreground when detected.
if [[ "$FOREGROUND" != "true" && "$FORCE_BACKGROUND" != "true" ]]; then
if is_windows_like_shell; then
FOREGROUND="true"
fi
fi
# Session files (server.log, server-info, .last-token) embed the session key —
# keep everything this script and the server create owner-only.
umask 077
# Generate unique session directory
SESSION_ID="$$-$(date +%s)"
if [[ -n "$PROJECT_DIR" ]]; then
SESSION_DIR="${PROJECT_DIR}/.superpowers/brainstorm/${SESSION_ID}"
# Persist the bound port and key per project so a restart reuses them and an
# already-open browser tab reconnects to the same URL with a valid cookie.
export BRAINSTORM_PORT_FILE="${PROJECT_DIR}/.superpowers/brainstorm/.last-port"
export BRAINSTORM_TOKEN_FILE="${PROJECT_DIR}/.superpowers/brainstorm/.last-token"
else
SESSION_DIR="/tmp/brainstorm-${SESSION_ID}"
fi
STATE_DIR="${SESSION_DIR}/state"
PID_FILE="${STATE_DIR}/server.pid"
LOG_FILE="${STATE_DIR}/server.log"
SERVER_ID_FILE="${STATE_DIR}/server-instance-id"
# Create fresh session directory with content and state peers
mkdir -p "${SESSION_DIR}/content" "$STATE_DIR"
SERVER_ID=""
if [[ -r /dev/urandom ]]; then
SERVER_ID="$(od -An -N24 -tx1 /dev/urandom 2>/dev/null | tr -d ' \n' || true)"
fi
if ! [[ "$SERVER_ID" =~ ^[A-Za-z0-9_-]{32,64}$ ]]; then
SERVER_ID="$(printf '%08x%08x%08x%08x' "$$" "$(date +%s)" "${RANDOM:-0}" "${RANDOM:-0}")"
fi
printf '%s\n' "$SERVER_ID" > "$SERVER_ID_FILE"
chmod 600 "$SERVER_ID_FILE" 2>/dev/null || true
# Kill any existing server
if [[ -f "$PID_FILE" ]]; then
old_pid=$(cat "$PID_FILE")
kill "$old_pid" 2>/dev/null
rm -f "$PID_FILE"
fi
cd "$SCRIPT_DIR" || exit 1
# Resolve the harness PID (grandparent of this script).
# $PPID is the ephemeral shell the harness spawned to run us — it dies
# when this script exits. The harness itself is $PPID's parent.
OWNER_PID="$(ps -o ppid= -p "$PPID" 2>/dev/null | tr -d ' ')"
if [[ -z "$OWNER_PID" || "$OWNER_PID" == "1" ]]; then
OWNER_PID="$PPID"
fi
# Windows/MSYS2: Node.js cannot see POSIX PIDs from the MSYS2 namespace.
# Passing a PID node cannot verify causes server to log owner-pid-invalid
# and self-terminate at the 60-second lifecycle check. Clear it so the
# watchdog is disabled and the idle timeout becomes the only shutdown trigger.
if is_windows_like_shell; then
OWNER_PID=""
fi
# Foreground mode for environments that reap detached/background processes.
if [[ "$FOREGROUND" == "true" ]]; then
env BRAINSTORM_DIR="$SESSION_DIR" BRAINSTORM_HOST="$BIND_HOST" BRAINSTORM_URL_HOST="$URL_HOST" BRAINSTORM_OWNER_PID="$OWNER_PID" node server.cjs "--brainstorm-server-id=$SERVER_ID" &
SERVER_PID=$!
echo "$SERVER_PID" > "$PID_FILE"
wait "$SERVER_PID"
exit $?
fi
# Start server, capturing output to log file
# Use nohup to survive shell exit; disown to remove from job table
nohup env BRAINSTORM_DIR="$SESSION_DIR" BRAINSTORM_HOST="$BIND_HOST" BRAINSTORM_URL_HOST="$URL_HOST" BRAINSTORM_OWNER_PID="$OWNER_PID" node server.cjs "--brainstorm-server-id=$SERVER_ID" > "$LOG_FILE" 2>&1 &
SERVER_PID=$!
disown "$SERVER_PID" 2>/dev/null
echo "$SERVER_PID" > "$PID_FILE"
# Wait for server-started message (check log file)
for _ in {1..50}; do
if grep -q "server-started" "$LOG_FILE" 2>/dev/null; then
# Verify server is still alive after a short window (catches process reapers)
alive="true"
for _ in {1..20}; do
if ! kill -0 "$SERVER_PID" 2>/dev/null; then
alive="false"
break
fi
sleep 0.1
done
if [[ "$alive" != "true" ]]; then
echo "{\"error\": \"Server started but was killed. Retry in a persistent terminal with: $SCRIPT_DIR/start-server.sh${PROJECT_DIR:+ --project-dir $PROJECT_DIR} --host $BIND_HOST --url-host $URL_HOST --foreground\"}"
exit 1
fi
grep "server-started" "$LOG_FILE" | head -1
exit 0
fi
sleep 0.1
done
# Timeout - server didn't start
echo '{"error": "Server failed to start within 5 seconds"}'
exit 1
@@ -1,120 +0,0 @@
#!/usr/bin/env bash
# Stop the brainstorm server and clean up
# Usage: stop-server.sh <session_dir>
#
# Kills the server process. Only deletes session directory if it's
# under /tmp (ephemeral). Persistent directories (.superpowers/) are
# kept so mockups can be reviewed later.
SESSION_DIR="$1"
if [[ -z "$SESSION_DIR" ]]; then
echo '{"error": "Usage: stop-server.sh <session_dir>"}'
exit 1
fi
STATE_DIR="${SESSION_DIR}/state"
PID_FILE="${STATE_DIR}/server.pid"
SERVER_ID_FILE="${STATE_DIR}/server-instance-id"
mark_stopped() {
local reason="$1"
rm -f "${STATE_DIR}/server-info"
printf '{"reason":"%s","timestamp":%s}\n' "$reason" "$(date +%s)" > "${STATE_DIR}/server-stopped"
}
read_expected_server_id() {
[[ -f "$SERVER_ID_FILE" ]] || return 1
local id
id="$(tr -d '\r\n' < "$SERVER_ID_FILE" 2>/dev/null || true)"
[[ "$id" =~ ^[A-Za-z0-9_-]{32,64}$ ]] || return 1
printf '%s\n' "$id"
}
command_line_for_pid() {
local pid="$1"
if [[ -r "/proc/$pid/cmdline" ]]; then
tr '\0' '\n' < "/proc/$pid/cmdline" 2>/dev/null || true
return 0
fi
ps -ww -p "$pid" -o command= 2>/dev/null || ps -f -p "$pid" 2>/dev/null | sed '1d' || true
}
command_has_server_id() {
local pid="$1"
local expected="$2"
local expected_arg="--brainstorm-server-id=$expected"
if [[ -r "/proc/$pid/cmdline" ]]; then
local arg
while IFS= read -r -d '' arg || [[ -n "$arg" ]]; do
[[ "$arg" == "$expected_arg" ]] && return 0
done < "/proc/$pid/cmdline"
return 1
fi
local command_line
command_line="$(command_line_for_pid "$pid")"
[[ -n "$command_line" ]] || return 1
case " $command_line " in
*" $expected_arg "*) return 0 ;;
*) return 1 ;;
esac
}
# Confirm a PID has this session's per-start instance id, not just a familiar
# process name. Ambiguous or legacy metadata fails closed as stale_pid.
is_brainstorm_server() {
kill -0 "$1" 2>/dev/null || return 1
local expected_id
expected_id="$(read_expected_server_id)" || return 1
command_has_server_id "$1" "$expected_id" || return 1
return 0
}
if [[ -f "$PID_FILE" ]]; then
pid=$(cat "$PID_FILE")
# Refuse to signal a PID we can't prove is our server. A stale pid file may
# point at an unrelated process after a reboot/PID wraparound.
if ! is_brainstorm_server "$pid"; then
rm -f "$PID_FILE" "$SERVER_ID_FILE"
mark_stopped "stale_pid"
echo '{"status": "stale_pid"}'
exit 0
fi
# Try to stop gracefully, fallback to force if still alive
kill "$pid" 2>/dev/null || true
# Wait for graceful shutdown (up to ~2s)
for _ in {1..20}; do
if ! kill -0 "$pid" 2>/dev/null; then
break
fi
sleep 0.1
done
# If still running, escalate to SIGKILL
if kill -0 "$pid" 2>/dev/null; then
kill -9 "$pid" 2>/dev/null || true
# Give SIGKILL a moment to take effect
sleep 0.1
fi
if kill -0 "$pid" 2>/dev/null; then
echo '{"status": "failed", "error": "process still running"}'
exit 1
fi
rm -f "$PID_FILE" "$SERVER_ID_FILE" "${STATE_DIR}/server.log"
mark_stopped "stop-server.sh"
# Only delete ephemeral /tmp directories
if [[ "$SESSION_DIR" == /tmp/* ]]; then
rm -rf "$SESSION_DIR"
fi
echo '{"status": "stopped"}'
else
echo '{"status": "not_running"}'
fi
@@ -1,49 +0,0 @@
# Spec Document Reviewer Prompt Template
Use this template when dispatching a spec document reviewer subagent.
**Purpose:** Verify the spec is complete, consistent, and ready for implementation planning.
**Dispatch after:** Spec document is written to docs/superpowers/specs/
```
Subagent (general-purpose):
description: "Review spec document"
prompt: |
You are a spec document reviewer. Verify this spec is complete and ready for planning.
**Spec to review:** [SPEC_FILE_PATH]
## What to Check
| Category | What to Look For |
|----------|------------------|
| Completeness | TODOs, placeholders, "TBD", incomplete sections |
| Consistency | Internal contradictions, conflicting requirements |
| Clarity | Requirements ambiguous enough to cause someone to build the wrong thing |
| Scope | Focused enough for a single plan — not covering multiple independent subsystems |
| YAGNI | Unrequested features, over-engineering |
## Calibration
**Only flag issues that would cause real problems during implementation planning.**
A missing section, a contradiction, or a requirement so ambiguous it could be
interpreted two different ways — those are issues. Minor wording improvements,
stylistic preferences, and "sections less detailed than others" are not.
Approve unless there are serious gaps that would lead to a flawed plan.
## Output Format
## Spec Review
**Status:** Approved | Issues Found
**Issues (if any):**
- [Section X]: [specific issue] - [why it matters for planning]
**Recommendations (advisory, do not block approval):**
- [suggestions for improvement]
```
**Reviewer returns:** Status, Issues (if any), Recommendations
@@ -1,299 +0,0 @@
# Visual Companion Guide
Browser-based visual brainstorming companion for showing mockups, diagrams, and options.
## When to Use
Decide per-question, not per-session. The test: **would the user understand this better by seeing it than reading it?**
**Use the browser** when the content itself is visual:
- **UI mockups** — wireframes, layouts, navigation structures, component designs
- **Architecture diagrams** — system components, data flow, relationship maps
- **Side-by-side visual comparisons** — comparing two layouts, two color schemes, two design directions
- **Design polish** — when the question is about look and feel, spacing, visual hierarchy
- **Spatial relationships** — state machines, flowcharts, entity relationships rendered as diagrams
**Use the terminal** when the content is text or tabular:
- **Requirements and scope questions** — "what does X mean?", "which features are in scope?"
- **Conceptual A/B/C choices** — picking between approaches described in words
- **Tradeoff lists** — pros/cons, comparison tables
- **Technical decisions** — API design, data modeling, architectural approach selection
- **Clarifying questions** — anything where the answer is words, not a visual preference
A question *about* a UI topic is not automatically a visual question. "What kind of wizard do you want?" is conceptual — use the terminal. "Which of these wizard layouts feels right?" is visual — use the browser.
## How It Works
The server watches a directory for HTML files and serves the newest one to the browser. You write HTML content to `screen_dir`, the user sees it in their browser and can click to select options. Selections are recorded to `state_dir/events` that you read on your next turn.
**Content fragments vs full documents:** If your HTML file starts with `<!DOCTYPE` or `<html`, the server serves it as-is (just injects the helper script). Otherwise, the server automatically wraps your content in the frame template — adding the header, CSS theme, connection status, and all interactive infrastructure. **Write content fragments by default.** Only write full documents when you need complete control over the page.
## Starting a Session
```bash
# Start AFTER the user approves the companion. --open auto-opens their browser on
# the first screen; --project-dir persists mockups and enables same-port restart.
scripts/start-server.sh --project-dir /path/to/project --open
# Returns: {"type":"server-started","port":52341,
# "url":"http://localhost:52341/?key=ab12…",
# "screen_dir":"/path/to/project/.superpowers/brainstorm/12345-1706000000/content",
# "state_dir":"/path/to/project/.superpowers/brainstorm/12345-1706000000/state"}
```
Save `screen_dir` and `state_dir` from the response. With `--open`, the browser opens itself when you push the first screen — you don't need to ask the user to open it, but still share the URL as a fallback (headless/remote setups won't auto-open).
**The URL contains a session key (`?key=…`).** The server rejects any request
without it, so always give the user the **complete** URL from the `url` field —
never strip the query string, and never hand out a bare `http://host:port`. The
key gates HTTP and WebSocket access so a stray browser tab or another machine on
the network can't read the screens or inject events. After the first load the
browser remembers the key via a cookie, so reloads and `/files/*` assets work
without repeating it.
**Finding connection info:** The server writes its startup JSON to `$STATE_DIR/server-info`. If you launched the server in the background and didn't capture stdout, read that file to get the URL and port. When using `--project-dir`, check `<project>/.superpowers/brainstorm/` for the session directory.
**Note:** Pass the project root as `--project-dir` so mockups persist in `.superpowers/brainstorm/` and survive server restarts. Without it, files go to `/tmp` and get cleaned up. Remind the user to add `.superpowers/` to `.gitignore` if it's not already there.
**Launching the server by platform:**
**Claude Code:**
```bash
# Default mode works — the script backgrounds the server itself.
scripts/start-server.sh --project-dir /path/to/project --open
```
On Windows, the script auto-detects and switches to foreground mode (which blocks the tool call). Use `run_in_background: true` on the Bash tool call so the server survives across conversation turns, then read `$STATE_DIR/server-info` on the next turn to get the URL and port.
**Codex:**
```bash
# Codex reaps background processes. The script auto-detects CODEX_CI and
# switches to foreground mode. Run it normally — no extra flags needed.
scripts/start-server.sh --project-dir /path/to/project --open
```
**Gemini CLI:**
```bash
# Use --foreground and set is_background: true on your shell tool call
# so the process survives across turns
scripts/start-server.sh --project-dir /path/to/project --open --foreground
```
**Copilot CLI:**
```bash
# Start it with Copilot CLI's non-blocking/background shell mechanism so the
# server survives across turns. Keep --foreground so the harness, not the
# script, owns backgrounding. The launcher is a .sh, so invoke it via bash
# (on Windows, call Git Bash's bash.exe from the PowerShell tool).
bash scripts/start-server.sh --project-dir /path/to/project --open --foreground
```
**Other environments:** The server must keep running in the background across conversation turns. If your environment reaps detached processes, use `--foreground` and launch the command with your platform's background execution mechanism.
If the URL is unreachable from your browser (common in remote/containerized setups), bind a non-loopback host:
```bash
scripts/start-server.sh \
--project-dir /path/to/project \
--host 0.0.0.0 \
--url-host localhost
```
Use `--url-host` to control what hostname is printed in the returned URL JSON.
## The Loop
1. **Check server is alive**, then **write HTML** to a new file in `screen_dir`:
- **Required: confirm the server is alive before referring to the URL or pushing a screen.** Check that `$STATE_DIR/server-info` exists and `$STATE_DIR/server-stopped` does not. If it has shut down, restart it with `start-server.sh` using the **same `--project-dir`** — it reuses the same port, so the user's open tab reconnects on its own (it shows a "paused" overlay while the server is down) and you don't need to send a new URL. The server auto-exits after 4 hours idle (configurable with `--idle-timeout-minutes`).
- Use semantic filenames: `platform.html`, `visual-style.html`, `layout.html`
- **Never reuse filenames** — each screen gets a fresh file
- Use your file-creation tool — **never use cat/heredoc** (dumps noise into terminal)
- Server automatically serves the newest file
2. **Tell user what to expect and end your turn:**
- Remind them of the URL (every step, not just first)
- Give a brief text summary of what's on screen (e.g., "Showing 3 layout options for the homepage")
- Ask them to respond in the terminal: "Take a look and let me know what you think. Click to select an option if you'd like."
3. **On your next turn** — after the user responds in the terminal:
- Read `$STATE_DIR/events` if it exists — this contains the user's browser interactions (clicks, selections) as JSON lines
- Merge with the user's terminal text to get the full picture
- The terminal message is the primary feedback; `state_dir/events` provides structured interaction data
4. **Iterate or advance** — if feedback changes current screen, write a new file (e.g., `layout-v2.html`). Only move to the next question when the current step is validated.
5. **Unload when returning to terminal** — when the next step doesn't need the browser (e.g., a clarifying question, a tradeoff discussion), push a waiting screen to clear the stale content:
```html
<!-- filename: waiting.html (or waiting-2.html, etc.) -->
<div style="display:flex;align-items:center;justify-content:center;min-height:60vh">
<p class="subtitle">Continuing in terminal...</p>
</div>
```
This prevents the user from staring at a resolved choice while the conversation has moved on. When the next visual question comes up, push a new content file as usual.
6. Repeat until done.
## Writing Content Fragments
Write just the content that goes inside the page. The server wraps it in the frame template automatically (header, theme CSS, connection status, and all interactive infrastructure).
**Minimal example:**
```html
<h2>Which layout works better?</h2>
<p class="subtitle">Consider readability and visual hierarchy</p>
<div class="options">
<div class="option" data-choice="a" onclick="toggleSelect(this)">
<div class="letter">A</div>
<div class="content">
<h3>Single Column</h3>
<p>Clean, focused reading experience</p>
</div>
</div>
<div class="option" data-choice="b" onclick="toggleSelect(this)">
<div class="letter">B</div>
<div class="content">
<h3>Two Column</h3>
<p>Sidebar navigation with main content</p>
</div>
</div>
</div>
```
That's it. No `<html>`, no CSS, no `<script>` tags needed. The server provides all of that.
## CSS Classes Available
The frame template provides these CSS classes for your content:
### Options (A/B/C choices)
```html
<div class="options">
<div class="option" data-choice="a" onclick="toggleSelect(this)">
<div class="letter">A</div>
<div class="content">
<h3>Title</h3>
<p>Description</p>
</div>
</div>
</div>
```
**Multi-select:** Add `data-multiselect` to the container to let users select multiple options. Each click toggles the item's selected styling.
```html
<div class="options" data-multiselect>
<!-- same option markup — users can select/deselect multiple -->
</div>
```
### Cards (visual designs)
```html
<div class="cards">
<div class="card" data-choice="design1" onclick="toggleSelect(this)">
<div class="card-image"><!-- mockup content --></div>
<div class="card-body">
<h3>Name</h3>
<p>Description</p>
</div>
</div>
</div>
```
### Mockup container
```html
<div class="mockup">
<div class="mockup-header">Preview: Dashboard Layout</div>
<div class="mockup-body"><!-- your mockup HTML --></div>
</div>
```
### Split view (side-by-side)
```html
<div class="split">
<div class="mockup"><!-- left --></div>
<div class="mockup"><!-- right --></div>
</div>
```
### Pros/Cons
```html
<div class="pros-cons">
<div class="pros"><h4>Pros</h4><ul><li>Benefit</li></ul></div>
<div class="cons"><h4>Cons</h4><ul><li>Drawback</li></ul></div>
</div>
```
### Mock elements (wireframe building blocks)
```html
<div class="mock-nav">Logo | Home | About | Contact</div>
<div style="display: flex;">
<div class="mock-sidebar">Navigation</div>
<div class="mock-content">Main content area</div>
</div>
<button class="mock-button">Action Button</button>
<input class="mock-input" placeholder="Input field">
<div class="placeholder">Placeholder area</div>
```
### Typography and sections
- `h2` — page title
- `h3` — section heading
- `.subtitle` — secondary text below title
- `.section` — content block with bottom margin
- `.label` — small uppercase label text
## Browser Events Format
When the user clicks options in the browser, their interactions are recorded to `$STATE_DIR/events` (one JSON object per line). The file is cleared automatically when you push a new screen.
```jsonl
{"type":"click","choice":"a","text":"Option A - Simple Layout","timestamp":1706000101}
{"type":"click","choice":"c","text":"Option C - Complex Grid","timestamp":1706000108}
{"type":"click","choice":"b","text":"Option B - Hybrid","timestamp":1706000115}
```
The full event stream shows the user's exploration path — they may click multiple options before settling. The last `choice` event is typically the final selection, but the pattern of clicks can reveal hesitation or preferences worth asking about.
If `$STATE_DIR/events` doesn't exist, the user didn't interact with the browser — use only their terminal text.
## Design Tips
- **Scale fidelity to the question** — wireframes for layout, polish for polish questions
- **Explain the question on each page** — "Which layout feels more professional?" not just "Pick one"
- **Iterate before advancing** — if feedback changes current screen, write a new version
- **2-4 options max** per screen
- **Use real content when it matters** — for a photography portfolio, use actual images (Unsplash). Placeholder content obscures design issues.
- **Keep mockups simple** — focus on layout and structure, not pixel-perfect design
## File Naming
- Use semantic names: `platform.html`, `visual-style.html`, `layout.html`
- Never reuse filenames — each screen must be a new file
- For iterations: append version suffix like `layout-v2.html`, `layout-v3.html`
- Server serves newest file by modification time
## Cleaning Up
```bash
scripts/stop-server.sh $SESSION_DIR
```
If the session used `--project-dir`, mockup files persist in `.superpowers/brainstorm/` for later reference. Only `/tmp` sessions get deleted on stop.
## Reference
- Frame template (CSS reference): `scripts/frame-template.html`
- Helper script (client-side): `scripts/helper.js`
View File
-6
View File
@@ -1,6 +0,0 @@
Websearch:
type: remote
url: https://mcp.exa.ai/mcp
headers:
Authorization: x-api-key {env:AZ_FLEET_EXA_KEY}
enabled: true
-4
View File
@@ -1,4 +0,0 @@
Outline:
type: remote
url: https://wiki.az-gruppe.com/mcp
enabled: true
-6
View File
@@ -1,6 +0,0 @@
Raynet:
type: remote
url: https://app.raynet.cz/api/mcp
headers:
Authorization: Bearer
enabled: true
-4
View File
@@ -1,4 +0,0 @@
Sally-Notetaker:
type: remote
url: https://app.sally.io/api/v1/McpExternal
enabled: true
-82
View File
@@ -1,82 +0,0 @@
---
name: az-datenanalyse
description: Datenanalyse mit dem AZ-Python-Datenstack (pandas/NumPy/openpyxl) — greift bei jeder Aufgabe, Excel- oder CSV-Dateien auszuwerten, umzuformen, zusammenzufassen oder Kennzahlen zu berechnen und das Ergebnis als Datei zurückzugeben. Bringt dem Agenten den zentralen Datenstack-Interpreter, den korrekten Ausführungsweg und die Sackgassen-Vermeidung bei (python aus dem PATH hat KEINE Datenpakete; pip/uv-Paketbefehle sind gesperrt).
---
# Datenanalyse mit dem AZ-Datenstack
Du arbeitest auf einer AZ-Agenten-Workstation. Für Datenanalyse steht dir ein
zentral ausgerollter Python-Datenstack bereit — aber nur über den Weg unten.
Die üblichen Verfahren (einfach `python` aufrufen, fehlende Pakete per pip
nachinstallieren) funktionieren hier bewusst nicht.
## Der zentrale Datenstack
- **Interpreter (IMMER diesen verwenden):**
`C:\ProgramData\az-fleet\datastack-venv\Scripts\python.exe`
- **Enthaltene Pakete:** pandas, NumPy, openpyxl (Excel lesen/schreiben —
`.xlsx`, nicht `.xls`).
- Der Datenstack wird zentral vom Fleet gepflegt und versioniert. Es gibt
keine anderen Datenpakete — plane deine Lösung mit genau diesen.
## So führst du eine Analyse aus
1. Skript mit dem Edit-Tool **im Arbeitsordner des Nutzers** schreiben
(z. B. `analyse.py`).
2. Ausführen per Bash — PowerShell-Form mit absolutem Interpreter-Pfad:
```powershell
& 'C:\ProgramData\az-fleet\datastack-venv\Scripts\python.exe' .\analyse.py
```
3. Ergebnis prüfen und dem Nutzer als Datei zurückgeben (z. B. Ergebnis-
Excel in den Arbeitsordner schreiben und den Pfad nennen).
**Wichtig:** Dieser Aufruf löst EINEN Bestätigungsdialog beim Nutzer aus.
Das ist beabsichtigt (Sicherheitskontrolle der AZ-Fleet-Konfiguration) und
kein Fehler. Fahre normal fort.
## Sackgassen vermeiden
- **`python` / `py` aus dem PATH sind das System-Python — OHNE pandas.**
`import pandas` scheitert dort. Nutze niemals das System-Python für
Datenanalyse.
- **Paketbefehle sind gesperrt und werden still blockiert.** Versuche
NIEMALS: `pip install …`, `python -m pip …`, `uv pip install …`, `uv add`,
`bun install` oder ähnliches. Das ist zentrale Fleet-Kontrolle — kein
Workaround, kein `--user`, kein `ensurepip`.
- **Fehlt ein Paket** (z. B. matplotlib): Löse die Aufgabe mit den
vorhandenen Mitteln (pandas kann rechnen, filtern, aggregieren; openpyxl
schreibt Excel inklusive einfacher Formatierung) oder sage dem Nutzer
klar, dass die IT das Paket zentral nachliefern muss (Fleet-Antrag).
Suche nicht nach Installationswegen.
## Typisches Muster (Excel rein → Kennzahlen → Excel raus)
```python
import pandas as pd
quelle = r"C:\Pfad\Zum\Arbeitsordner\eingabe.xlsx"
ziel = r"C:\Pfad\Zum\Arbeitsordner\ergebnis.xlsx"
df = pd.read_excel(quelle, sheet_name=0)
ergebnis = df.groupby("Kategorie")["Betrag"].agg(["sum", "mean", "count"])
with pd.ExcelWriter(ziel, engine="openpyxl") as writer:
ergebnis.to_excel(writer, sheet_name="Kennzahlen")
df.to_excel(writer, sheet_name="Rohdaten", index=False)
print(ergebnis)
print("Geschrieben:", ziel)
```
Pfade immer vom Nutzer erfragen oder aus dem Arbeitsordner ableiten — nie
raten.
## Datenschutz-Kurzregel
In Skripten, Prompts und Ergebnissen gelten die Regeln der AZ-Datenleitlinie:
keine personenbezogenen Daten (Namen, Adressen, Telefonnummern, E-Mail von
Menschen) — Menschen anonymisieren oder weglassen; keine Zugangsdaten; keine
besonders schützenswerten Kunden- und Rüstungsdaten. Im Zweifel: nachfragen
statt raten.
-68
View File
@@ -1,68 +0,0 @@
---
name: az-hilfe
description: Einstiegs- und Orientierungshilfe für die Agenten-Workstation der AZ-Gruppe — greift, wenn ein Anwender fragt, was sein Agent kann, welches Artefakt wofür da ist, wie man etwas beiträgt oder an wen man sich bei Problemen wendet.
---
# AZ-Agent — Hilfe & Onboarding
Du bist der Guide für Anwender der AZ-Agenten-Workstation. Wer diesen Skill aufruft,
will orientiert werden — nicht mit Technik-Details überschüttet werden. Antworte auf
Deutsch, freundlich und in kurzen Schritten.
> **Status: Entwurf.** Dieses Skill ist der erste Baustein des Company-Default-Sets
> und wird mit den kommenden Artefakten (Commands, Agenten, MCP-Anbindungen)
> mitwachsen. Solange ein Bereich noch leer ist, sage das ehrlich.
## Erste Orientierung: Was ist hier installiert?
Die Workstation erhält ihre Standard-Artefakte zentral aus dem Repository
`az-agent-defaults` — gespiegelt über das Fleet-Repo. Es gibt vier Arten:
| Frage des Anwenders | Artefakt-Art | Wo es liegt |
|---|---|---|
| „Mach X, wenn Y" — Verhalten beibringen | Skill | `skills/` |
| Wiederkehrenden Prompt als Tastendruck (/slash) | Command | `commands/` |
| Eigenes Agenten-Profil / Prüf-Subagenten | Agent | `agents/` |
| Werkzeug anbinden (API, Datenquelle) | MCP-Server | `mcp/` |
**Aktuell verfügbar:** dieser Hilfe-Skill selbst. Commands, Agenten-Definitionen
und MCP-Anbindungen folgen (in Vorbereitung). Versprich nichts, was nicht
ausgeliefert ist.
## Typische Anfragen und gute Antworten
- **„Was kann mein Agent überhaupt?"** → Nenne die vier Artefakt-Arten oben und
was davon heute live ist. Beispiel geben: dieser Skill hier ist eins.
- **„Ich will, dass der Agent immer X tut, wenn Y."** → Das ist ein **Skill**.
Erkläre: Verhaltens-Anweisungen werden zentral gepflegt (Beitragshinweise
unten), nicht lokal auf der Workstation gebastelt.
- **„Ich tippe jeden Tag dasselbe …"** → Kandidat für einen **Command** — noch in
Vorbereitung; Wunsch aufnehmen und an die IT weiterreichen.
- **„Kann der Agent auch auf Tool Z zugreifen?"** → Das ist eine **MCP-Anbindung**
— sicherheitsrelevant, läuft über die IT (keine lokalen Einträge).
- **„Mein Kollege hat einen Skill, ich nicht."** → Auslieferung erfolgt zentral im
Fleet-Rollout: App neu starten; bleibt der Skill weg, IT kontaktieren.
## Selbst beitragen
Änderungen laufen **zentral über die IT** im Repository `az-agent-defaults`:
1. Wohin gehört der Beitrag? → Tabelle oben („Wohin gehört mein Beitrag?").
2. Das Repo-README beschreibt je Artefakt-Typ verbindliche Guard-Regeln
(Struktur, Frontmatter, Pflichtfelder) — dort steht auch je ein Minimalbeispiel.
3. Beitrag als Commit/PR einreichen; die IT prüft und rollt aus. Erst nach
Rollout und App-Neustart ist das Artefakt auf den Workstationen.
Anwender sollten **nie** lokal in `~\.agents\skills` (oder vergleichbare
Spiegel-Verzeichnisse) editieren — lokale Änderungen werden vom nächsten
Rollout überschrieben.
## Bei Problemen
- Artefakt fehlt / verhält sich falsch → IT kontaktieren, mit Name des
Artefakts und kurzem Ablauf des Problems.
- Nichts funktionstüchtig ändern: Diagnose ja, lokale Reparatur nein —
sie überlebt den nächsten Rollout nicht.
Führe darüber hinaus keine Aktionen aus. Dieser Skill orientiert — er verändert
nichts.
-23
View File
@@ -1,23 +0,0 @@
---
description: Hilft bei Basecamp-Themen und unterstützt das Team bei der Verwaltung von Projekten und To-dos.
mode: subagent
model: az-litellm/claude-haiku-4-5
temperature: 0.2
tools:
bash: true
read: true
---
Du bist ein Basecamp-Helfer. Wenn dich jemand ruft, bearbeitest du
Basecamp-Anfragen mit dem CLI.
## Arbeitsweise
1. **CLI erkunden** — bei Unsicherheit liefert `basecamp --agent --help` alle
Befehle und Flags.
2. **Anfrage ausführen** — setze die Anfrage mit dem CLI um.
## Grenzen
- Du installierst nichts — das CLI ist Fleet-verwaltet. Fehlt es, melde das
dem Auftraggeber.
-6
View File
@@ -1,6 +0,0 @@
---
mode: primary
---
Dieses Frontmatter enthält einen `mode`, aber das Pflichtfeld `description`
fehlt. Der Guard muss rot abbrechen.
-7
View File
@@ -1,7 +0,0 @@
---
description: Testgegenstand — Agent ohne Pflichtfeld mode.
temperature: 0.3
---
Dieses Frontmatter enthält eine description, aber das Pflichtfeld `mode`
(`primary` oder `subagent`) fehlt. Der Guard muss rot abbrechen.
-7
View File
@@ -1,7 +0,0 @@
---
description: Testgegenstand — mode mit ungültigem Wert.
mode: sideload
---
Der Wert von `mode` ist weder `primary` noch `subagent`. Der Guard muss rot
abbrechen.
-8
View File
@@ -1,8 +0,0 @@
---
description:
---
Arbeite das Ticket $ARGUMENTS ab.
(Testgegenstand: das `description`-Feld existiert, ist aber leer. Der Guard muss
rot abbrechen.)
-8
View File
@@ -1,8 +0,0 @@
---
agent: build
---
Arbeite das Ticket $ARGUMENTS ab.
(Testgegenstand: Frontmatter vorhanden, aber das Pflichtfeld `description` fehlt.
Der Guard muss rot abbrechen.)
-5
View File
@@ -1,5 +0,0 @@
Arbeite das Ticket $ARGUMENTS ab: zeige es mir, claime es, setze es um und
schließ es nach meiner Freigabe.
(Testgegenstand: dieser Command hat kein YAML-Frontmatter — die Pflicht-`description`
fehlt damit zwangsläufig. Der Guard muss rot abbrechen.)
-10
View File
@@ -1,10 +0,0 @@
---
description: Testgegenstand — Dateiname mit Umlaut, der später zum /slash-Namen würde.
---
Dieser Command-Inhalt ist formal in Ordnung: Frontmatter mit `description`,
Template-Body mit $ARGUMENTS.
(Testgegenstand: der Dateiname enthält einen Umlaut — aus `prüfung.md` würde der
/slash-Befehl `/prüfung`, was den Namensregeln widerspricht. Der Guard muss am
Dateinamen rot abbrechen.)
-7
View File
@@ -1,7 +0,0 @@
# Testgegenstand: Platzhalter in falscher Syntax — erlaubt ist nur ${VAULT:...}.
falscher-platzhalter-service:
type: remote
url: https://mcp.example.az.local/alt
headers:
Authorization: Bearer ${SECRET:veraltete-syntax}
enabled: true
-8
View File
@@ -1,8 +0,0 @@
# Testgegenstand: statischer Key im Klartext statt ${VAULT:...}-Platzhalter.
# Der Guard muss rot abbrechen — auch Beispiel-/Fake-Werte zählen als Secret.
wetter-service:
type: remote
url: https://mcp.example.az.local/wetter
headers:
Authorization: Bearer sk-live-a1b2c3d4e5f6
enabled: true
-5
View File
@@ -1,5 +0,0 @@
# Testgegenstand: keine Fragment-Struktur — der Server-Name fehlt als Schlüssel,
# die Felder stehen auf Root-Ebene statt unter genau einem Server-Namen.
type: remote
url: https://mcp.example.az.local/ohne-namen
enabled: true
-4
View File
@@ -1,4 +0,0 @@
# Testgegenstand: remote-Server ohne type-Feld.
az-ohne-type-service:
url: https://mcp.example.az.local/ohne-type
enabled: true
-4
View File
@@ -1,4 +0,0 @@
# Testgegenstand: remote-Server ohne Pflichtfeld url.
az-ohne-url-service:
type: remote
enabled: true
@@ -1,7 +0,0 @@
---
---
# Skill mit leerem Frontmatter
Diese SKILL.md enthält einen Frontmatter-Block, aber keine Felder — weder
`name` noch `description`. Der Guard muss dieses Artefakt rot abbrechen.
@@ -1,9 +0,0 @@
---
name: total-anderer-name
description: Testgegenstand — Frontmatter-Name passt nicht zum Ordner.
---
# Skill mit abweichendem Namen
Der Frontmatter-Name `total-anderer-name` passt nicht zum Ordnernamen
`name-ungleich-ordner`. Der Guard muss dieses Artefakt rot abbrechen.
@@ -1,8 +0,0 @@
---
name: ohne-description
---
# Skill ohne Description
Dieses Frontmatter enthält `name`, aber das Pflichtfeld `description` fehlt.
Der Guard muss dieses Artefakt rot abbrechen.
@@ -1,4 +0,0 @@
# Skill ohne Frontmatter-Öffner
Diese SKILL.md beginnt direkt mit einer Überschrift statt mit dem
YAML-Frontmatter-Öffner. Der Guard muss dieses Artefakt rot abbrechen.
@@ -1,5 +0,0 @@
# Testgegenstand: Ordner ohne SKILL.md
Dieser Ordner ist ein ungültiger Skill: er enthält keine `SKILL.md`, nur
diese Notiz (damit Git den Ordner überhaupt tracken kann). Der Guard muss
beim Scan der Skills-Verzeichnisse an diesem Ordner rot abbrechen.
@@ -1,58 +0,0 @@
# Smoke-Test-Protokoll: perfektionierte Agents (2026-09-20)
Beleg zu beads `az-agent-defaults-jtp` (Dogfooding + Smoke-Tests) nach den
Rewrites aus `74g`, `h2j`, `99v`, `bfj`, `m4t`. Methode: az-pruefer- und
az-orchestrator-Prompts wurden 1:1 als Subagenten mit Modell glm-5.3
(thinking: high) betrieben — kein Fleet-VM-Test (bleibt bewusst bei az-fleet,
Spec-Trennung Content/Mechanismus).
## 1. Dogfooding: az-pruefer über alle 5 Agenten (`agents/*.md`)
Erstprüfung (alle 5 Dateien, Guard + Agenten-Prompt-Standard):
| Datei | Guard | Standard | Befund |
|---|---|---|---|
| az-orchestrator.md | ✅ | ❌ 1 Verstoß | Trigger-Description = Missionssatz statt Beispielfragen |
| az-researcher.md | ✅ | ✅ | — |
| az-basecamp.md | ✅ | ✅ | — |
| az-office.md | ✅ | ✅ | Kosmetik: Anführungszeichen |
| az-pruefer.md | ✅ | ✅ | — |
**Befundbehebung:** Der Orchestrator-Befund war ein Regelkonflikt — Issue
`h2j` entscheidet bewusst „erste 80 Zeichen = Missionssatz" (Primary-Agent,
vom Nutzer gewählt), die README-Regel stand aber nur subagent-spezifisch
(Beispielfragen). Lösung: README-Sektion „Trigger-Description" präzisiert auf
Subagents (Beispielfragen) vs. Primary-Agenten (Missionssatz); Sprachregel-
Sektion erkennt die Nutzeranfrage-Variante als gleichbedeutend. Zusätzlich
Grammar-Fix in az-office („bearbeite sie ausschließlich").
**Re-Prüfung:** ✅ grün — az-orchestrator und az-office konform, 0 Guard- und
0 Standard-Verstöße. Zeichenwerte (Body): orchestrator ~7.000, researcher
~2.900, basecamp ~2.500, office ~2.300, pruefer ~3.200 — alle < 10.000.
## 2. Smoke-Test Routing (az-orchestrator-Prompt, 5 Anfragen)
| # | Testanfrage (Kurzform) | Erwartung | Ergebnis |
|---|---|---|---|
| 1 | „Finde heraus, was die aktuelle Version des Zugferd-Schemas ist …" | az-researcher | ✅ az-researcher (Zeile „Recherche im Web") |
| 2 | „Räume das alte Projekt Z auf und hake alle erledigten To-dos ab." | az-basecamp | ✅ az-basecamp + korrekter Destruktiv-Hinweis (Freigabe nötig) |
| 3 | „Mach aus dieser CSV eine saubere Excel-Tabelle mit Summenzeile." | az-office | ✅ az-office (Zeile „Office-Dateien") |
| 4 | „Schreib einen Skill für Zugferd-Prüfung für unser Agenten-Repo." | Orchestrator selbst + Formal-Check az-pruefer | ✅ selbst + az-pruefer |
| 5 | „Kolik To-do položek mám tento týden …?" (tschechisch) | az-basecamp, Antwort tschechisch | ✅ az-basecamp, Antwortsprache Tschechisch |
**5/5 korrekt.** Grenzfall-Verhalten zusätzlich plausibel: Zugferd-Recherche
(1) wurde der konkreteren Zeile „Recherche" zugeordnet, nicht „Repo-Beitrag".
## 3. Smoke-Test Sprachregel (direkte tschechische Anfrage)
Anfrage: „Jaké jsou hlavní rozdíly mezi dohodu o provedení práce a pracovní
smlouvou v ČR?" — Ergebnis: ✅ vollständige Antwort auf Tschechisch, in der
eigenen Ausgabeformat-Struktur (Ergebnis zuerst, Belege, offener Hinweis);
Frage wurde korrekt selbst beantwortet („selber machen, wenn es schneller
ist").
## Ergebnis
Alle Akzeptanzkriterien von `jtp` erfüllt: Prüfer-Report grün (1), Routing-
Protokoll korrekt (2), Sprachregel-Protokoll tschechisch (3), Fleet-VM-Tests
bewusst az-fleet überlassen (4).