docs: sync status, hermes setup, and cutover docs with box reality
This commit is contained in:
+42
-62
@@ -6,72 +6,52 @@
|
||||
## Wo das Projekt lebt
|
||||
- **Code:** `F:\Coding Stuff\mission-control-2` (Windows-Dev-PC) + Gitea-Remote
|
||||
`https://git.tobisniceshomelab.ddnsfree.com/Hitonabi/mission-control-v2` (Branch `main`).
|
||||
- **v1** (`F:\Coding Stuff\mission-control`) bleibt unangetastet bis zum Cutover.
|
||||
- **v1** (`F:\Coding Stuff\mission-control`) wird abgelöst (Cutover läuft, s.u.).
|
||||
- **Box** (Bosgame, `192.168.178.151`): **MC2 LIVE auf :9001** als sudo-freier systemd-USER-Dienst
|
||||
(`~/mission-control-v2`, Update via `deploy/deploy.sh`). Verifiziert gegen echtes llama-swap:
|
||||
6 Modelle inkl. Caps, **GPU/GTT (133 GB) + Temps**, Services-Health. Parallel zu v1 (unangetastet).
|
||||
(`~/mission-control-v2`, Update via `deploy/deploy.sh`).
|
||||
|
||||
## Gateway = builtin (NICHT LiteLLM)
|
||||
LiteLLM verlangt Python <3.14; die Box hat nur 3.14 (uvloop+orjson scheitern). Der **eingebaute
|
||||
Gateway** (`:9001/v1`, OpenAI-kompatibel) liefert `model:auto` (kurz→`fast`, komplex→`heavy`) +
|
||||
Streaming und ist E2E verifiziert. Vertrag (OpenAI-API) bleibt austauschbar.
|
||||
|
||||
## Phasen-Fortschritt
|
||||
- [x] **Phase 0 — Gerüst:** FastAPI (`/api/health`,`/api/models`) + React/shadcn-Shell (Cmd+K, Dark, PWA). Verifiziert.
|
||||
- [x] **Phase 1 — Engine + Routing:** Compute-Module (fit/caps/sources), Discover (live HF), Engine-Write
|
||||
(register + groups/Ko-Residenz), LiteLLM-Gateway-Config + Service, Frontend Modelle&Routing
|
||||
(Caps-Chips, Fit, Discover-Tab, Routing-View). Lokal verifiziert (Backend-Smoke + Frontend-Build + Browser).
|
||||
- [x] **Phase 2 — System/OS + Connect:** System-Status (psutil CPU/RAM/Disk, sysfs GPU/Temp guarded),
|
||||
Wartung (restart/self-update als systemd-USER-Dienst, sudo-frei), Connect-Snippets (Cline/OpenCode/
|
||||
Zed/Continue/Claude Code/Memory-MCP → Gateway `model:auto` + LAN-IP-Override). Lokal verifiziert.
|
||||
- [x] **Phase 3 — Memory + MCP:** Memory-Service (SQLite/WAL, 5 Kategorien, Dedupe-Kurator), Router
|
||||
(CRUD/export/dedupe), `mcp/mcp_memory.py` (Guard-Beschreibungen gegen Loop) + `mcp/mcp_mc.py` NEU
|
||||
(Stack-Management-Tools für Hermes: list/discover/register/route/restart/status). Frontend
|
||||
MemoryView (Add/Filter/Suche/Delete/Aufräumen). Lokal verifiziert (CRUD+Dedupe; MCP syntax-OK).
|
||||
- [~] **Phase 4 — Hermes-Schicht:** MC-Seite ✅ (services/agent.py + /api/agent/status, AgentView mit
|
||||
Status-Tiles + „Hermes öffnen"), `deploy/hermes-webui.service`, **Runbook `docs/HERMES_SETUP.md`**.
|
||||
**Box-Ausführung offen** (hermes-webui installieren, Brain=auto, Tools/MCP verdrahten — Runbook).
|
||||
- [x] **Phase 5 — Betrieb/Observability/Politur:** Backup (Memory-SQLite + Configs, retain 7;
|
||||
`/api/system/backup` + `deploy/backup.sh`), Services-Health (`/api/system/services`), Observability-
|
||||
Links (llama-swap /ui, Gateway), **Theme-Toggle Hell/Dunkel**. Lokal verifiziert.
|
||||
- [ ] **Phase 6 — Cutover** (/opt → v2, v1 aus). **Braucht Box.**
|
||||
- Offen aus Phase 0/1: **Box-Deploy** (Gitea-Repo da; Box-Setup als systemd-USER-Dienst noch offen),
|
||||
LiteLLM **Complexity-Auto-Router** gegen installierte Version verifizieren.
|
||||
- [x] **Phase 0 — Gerüst:** FastAPI + React/shadcn-Shell (Cmd+K, Dark, PWA). Live auf Box.
|
||||
- [x] **Phase 1 — Engine + Routing:** Compute (fit/caps/sources), Discover (live HF), Engine-Write
|
||||
(register + groups), **builtin Gateway** `model:auto`, Modelle&Routing-UI. Box-verifiziert.
|
||||
- [x] **Phase 2 — System/OS + Connect:** Metriken, Dienste, Self-Update, Connect-Snippets → Gateway.
|
||||
- [x] **Phase 3 — Memory + MCP:** Memory-UI, `mcp_memory.py` + `mcp_mc.py` (Stack-Management).
|
||||
- [x] **Phase 4 — Hermes-Schicht:** Agent-Status + AgentView; **Box: Hermes verdrahtet** (eigenes Hirn,
|
||||
MCP memory+stack). hermes-webui (nesquena) = Block C offen.
|
||||
- [x] **Phase 5 — Betrieb/Politur:** Backup, Services-Health, Observability-Links, Theme-Toggle.
|
||||
- [x] **W1–W8** (Audit-Arbeitspaket): Thrash-Fix, HF-Install, Rollen-UX, BEDIENUNG.md, Wartung. ✅
|
||||
- [~] **Phase 6 — Cutover:** läuft (s.u.). v1-Dashboard+Zombies weg; v1 :9000 Stop offen (User-sudo).
|
||||
|
||||
## Hermes-Hirn (Entscheidung 2026-06-25)
|
||||
Hermes hat ein **eigenes festes Hirn = Hermes-4-14B** (`model.model: hermes`, ttl 99999 = immer warm)
|
||||
+ **interne Delegation an `heavy`** (Qwen3.5-122B) für harte Teilaufgaben. **`model:auto` ist NUR
|
||||
Gateway/Vibe-Coding**, nicht Hermes. Verifiziert: „bist du da?" → **3s, sauber, kein Thrash**.
|
||||
|
||||
## Box-Stand (Session 2026-06-25)
|
||||
- **8 Modelle**, Rollen sauber: `fast` (Qwen3.6-35B-A3B), `heavy` (Qwen3.5-122B-A10B), `coder`
|
||||
(Qwen3-Coder-30B), `vision` (Qwen3-VL-8B), `scout` (Qwen3-8B), `hermes` (Hermes-4-14B). Legacy
|
||||
`manager`/`reviewer`-Aliase entfernt (Modelle bleiben per Realname ladbar).
|
||||
- **Thrash behoben** (war NICHT die Session): kaputte Tools global via `agent.disabled_toolsets`
|
||||
abgeschaltet (browser/vision/computer_use/image_gen/tts/video*/memory). Details: Memory
|
||||
`hermes-thrash-rootcause-fix`. **`hermes tools disable` greift NICHT am api_server** — nur die config.
|
||||
- **Cutover teilweise:** `hermes-dashboard` (:9119) disabled (killte 4 v1-MCP-Zombies);
|
||||
**v1 `mission-control.service` (:9000) läuft noch** → `sudo systemctl disable --now mission-control`.
|
||||
- MCP nutzt `/opt/mission-control/.venv/bin/python` (hat `mcp`-Modul) für v2-Skripte → /opt-Dir bleibt.
|
||||
|
||||
## Nächste Schritte (siehe Plan „Nächste Schritte" + docs/CUTOVER.md)
|
||||
- **Block C:** nesquena hermes-webui (:8787) installieren → „Hermes öffnen" zeigt darauf.
|
||||
- **Block E:** Ko-Residenz `hermes`+`fast` (swap:false) testen, GTT beobachten.
|
||||
- **User-sudo:** v1 :9000 stoppen; sudoers für apt-get/reboot erweitern (OS-Update/Reboot).
|
||||
- Offen/user-seitig: SSH→Windows (OpenSSH am PC), v1-Passwort-Hygiene (Git-Historie).
|
||||
|
||||
## Lokal entwickeln/verifizieren
|
||||
```bash
|
||||
# Backend
|
||||
cd backend && .venv/Scripts/python -m uvicorn app:app --port 9000
|
||||
# Frontend (Dev)
|
||||
cd frontend && npm run dev # http://localhost:5173 (proxyt /api → :9000)
|
||||
# Frontend (Build → wird vom Backend ausgeliefert, committet)
|
||||
cd frontend && npm run build
|
||||
cd backend && .venv/Scripts/python -m uvicorn app:app --port 9000 # Backend
|
||||
cd frontend && npm run dev # http://localhost:5173
|
||||
cd frontend && npm run build # Prod-Build (committet)
|
||||
```
|
||||
|
||||
## API (Stand Phase 1)
|
||||
- `GET /api/health` — Status + engine/gateway reachable
|
||||
- `GET /api/models` — installierte Modelle inkl. `capabilities`
|
||||
- `GET /api/discover[?force=1]` — beste Modelle je Kategorie (live HF, Fit, Caps, ⭐recommended)
|
||||
- `GET /api/fit?params_b=&quant=&ctx=&name=` — Hardware-Fit-Schätzung
|
||||
- `POST /api/models/register` — GGUF als llama-swap-Modell eintragen (cmd + Rolle-Alias, optional jinja)
|
||||
- `GET/PUT /api/groups` — llama-swap-`groups` (Ko-Residenz `swap:false`)
|
||||
- `GET /api/routing`, `PUT /api/routing/route` — LiteLLM-Gateway-Mapping
|
||||
- `GET /api/system/status` — CPU/RAM/GPU/Disk/Temp
|
||||
- `POST /api/system/restart` (Whitelist), `POST /api/system/self-update` — Wartung (Box, sudo-frei)
|
||||
- `GET /api/connect?host=` — IDE-/Agent-Snippets (Gateway + Memory-MCP)
|
||||
- `GET/POST/PUT/DELETE /api/memory[...]` + `/api/memory/export` + `/api/memory/dedupe` — geteiltes Gedächtnis
|
||||
- `mcp/mcp_memory.py` (geteiltes Memory) + `mcp/mcp_mc.py` (Stack-Management für Hermes) — stdio-MCP
|
||||
- `GET /api/agent/status` — Hermes Gateway/WebUI-Erreichbarkeit + Verdrahtungs-Hinweise
|
||||
|
||||
## Box-Stack LIVE (Phase 6 ausgeführt) — Stand & Cutover: siehe `docs/CUTOVER.md`
|
||||
- Modelle `fast` (Qwen3.6-35B-A3B) + `heavy` (Qwen3.5-122B-A10B) geladen & lauffähig.
|
||||
- Eingebauter Gateway `:9001/v1` `model: auto` (fast↔heavy) — E2E verifiziert.
|
||||
- Hermes: Brain via Gateway, MCP (memory+stack) verdrahtet, Memory vereinheitlicht (v1-DB, 16 Einträge).
|
||||
- Offene Tuning-Punkte (kein Blocker): Hermes-Agent-Brain für die Loop (Qwen3.6 thrasht → `coder`
|
||||
empfohlen), SSH→Windows (Windows-seitig), nesquena-WebUI optional. Details: `docs/CUTOVER.md`.
|
||||
|
||||
## Nächster sinnvoller Schritt (user-gated — braucht dich / Downloads / Entscheidungen)
|
||||
Phasen 0–5 sind fertig **und MC2 läuft live auf der Box** (`:9001`). Es bleiben **bewusst von dir
|
||||
auszulösende** Schritte (schwere Downloads, Passwörter, Eingriff in die laufende v1/Hermes-Produktion):
|
||||
1. **Hirne laden (~90 GB):** Qwen3.6-35B-A3B (`fast`, `--jinja`) + Qwen3.5-122B-A10B (`heavy`) eintragen
|
||||
+ Gruppe `brains` (`swap:false`). UI: Modelle & Routing → `http://192.168.178.151:9001`.
|
||||
2. **LiteLLM-Gateway** installieren/starten (`docs/HERMES_SETUP.md` §1), `model:auto` verifizieren.
|
||||
3. **Hermes-Schicht** verdrahten (`docs/HERMES_SETUP.md`): hermes-webui + Passwort, Brain=auto, Tools/MCP
|
||||
(`mcp_mc`+`mcp_memory`), SSH→Windows.
|
||||
4. **Cutover (Phase 6):** wenn v2 dir reicht — v1 stilllegen, v2 ggf. auf den Hauptport. (Aktuell läuft
|
||||
v2 risikofrei parallel auf :9001, v1 unangetastet.)
|
||||
|
||||
Reference in New Issue
Block a user