The problem
Model accuracy drops once more than about 15–20 tools are visible at once. Real agent platforms have hundreds.
The solution
Register tools once, then ask /advise for a shortlist per request:
Registry (SQLite, versioned)
└─ Retrieval index (cosine over local embeddings)
└─ Selection pipeline (permissions → shortcut → router → candidates)
└─ Execution & feedback (arg validation, confirmation gate, audit, usage stats)
- Versioned tools: they are never mutated in place, and the full history is kept.
- Permission checks first, before any ranking.
- Gated, audited execution, with outcome capture that feeds back into ranking.
- Runs fully locally: local embedding model, no API keys.