Sovereign Desk: a private knowledge agent that runs on one workstation
The problem
Law firms, clinics, and industrial companies sit on documents they cannot send to cloud AI under client or regulatory policies. Current local stacks (Ollama plus small open models) keep data in-house but answer like a weak model, so teams quietly keep pasting into ChatGPT anyway. Demand is real but the tooling is DIY: TeleAgent just shipped a private-deployment edition for sovereign environments, and GitHub hosts multiple sovereign-AI reference projects trying to solve exactly this.
The idea
Why now
Two unlocks landed within days of each other. Colibri demonstrated 744B-2.8T-class MoE models running on consumer hardware by treating VRAM, RAM, and storage as one inference hierarchy. WeKnora shipped a production-grade self-hosted knowledge stack (29 model vendors, swappable local backends, Langfuse tracing) that already supports local Ollama. Together they turn a frontier-quality private knowledge agent from a GPU-cluster project into a workstation purchase.
What it combines
Colibri's AI memory multitiering (giant models on consumer hardware) plus WeKnora's self-hosted knowledge platform (RAG plus agent plus self-maintaining wiki). The mix matters because WeKnora solves the knowledge-management half that Colibri lacks, and Colibri solves the model-quality half that WeKnora's local-Ollama path lacks; neither alone delivers a frontier-quality private agent on cheap hardware.
MVP
Docker Compose bundle: WeKnora plus Colibri serving one supported MoE (e.g. GLM-5.2), a setup wizard that ingests a folder of PDFs, and a published benchmark of answer quality and 5-concurrent-user latency on a mid-range workstation. Deliberately skip multi-tenancy and fine-tuning.
Distribution
B2B2C through MSPs and IT services firms that already sell into law firms, clinics, and industrial companies. They charge for the workstation, the setup, and a maintenance contract; the software bundle is the wedge that wins the engagement.
Why it wins
AnythingLLM and Open WebUI plus Ollama already serve small firms, but they cap out around 70B-class models on a workstation. GrapeUp Aiboostr and SecureLLMs sell private AI as managed per-seat services. Sovereign Desk is neither: a one-time hardware-plus-software package, sold through MSPs, that brings 744B-class brains to the same box the firm already owns.
Risks
Interactive latency: Colibri's honestly reported numbers are slow on weak hardware (fractions of a token per second in streaming CPU mode). The MVP de-risks this by publishing measured latency on a defined hardware spec before selling anything; if it cannot beat 5 tok/s on a $2,000 box, the product thesis fails.
Build it with
- Colibri: pure-C engine that runs 744B-2.8T MoE models on consumer hardwareMemory-multitiering inference engine that runs 744B-2.8T MoE on consumer hardware
- Tencent WeKnora: open-source LLM knowledge platform (RAG + agent + self-maintaining Wiki)Self-hosted knowledge platform: RAG, agent, self-maintaining wiki, swappable local models
Repo to start from
sovereign-desk: Docker Compose bundle plus setup wizard and published benchmarks for private knowledge agents on consumer hardware.
Evidence
- Colibri
- Tencent WeKnora
- Sovereign AI project charter
- TeleAgent hits 1.5M users, private-deployment edition
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.