Compare commits
262 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 4e5a7eed35 | |||
| 612ce127fa | |||
| 357285e798 | |||
| b21794a750 | |||
| 97344f1ad2 | |||
| a682e3c4e7 | |||
| 8b0e14db7c | |||
| c8e408cdd8 | |||
| fa8f776910 | |||
| 2b124d5b0a | |||
| 019bd6d18b | |||
| a23f4c0652 | |||
| a024ea1b2f | |||
| d2ab662146 | |||
| 50867b2781 | |||
| a2091b4f5f | |||
| dd3d66156b | |||
| 18f36f52ce | |||
| 02fac7f55f | |||
| 9e5392fa4f | |||
| ff6c1daf6c | |||
| 14b5c6d6a0 | |||
| 2c8d4a081e | |||
| ab8d651efc | |||
| dd947298d0 | |||
| 2c95f1c5bc | |||
| cd3c5b5eed | |||
| 2db0ce080f | |||
| df8ec46018 | |||
| a8d6f3f598 | |||
| 2263e298fd | |||
| 1adbb19572 | |||
| 41aeb111d7 | |||
| d887a6c248 | |||
| 4eaa11c32a | |||
| 8c7dd5b30f | |||
| 49bf11886e | |||
| 996fea224a | |||
| 9822af7640 | |||
| 9e7cc6612f | |||
| 4a680b863e | |||
| d7de1f47d4 | |||
| 646031e5c0 | |||
| 195b07deef | |||
| 2a4eb51bd7 | |||
| 7704ffddab | |||
| 588d032335 | |||
| e84386b03e | |||
| a16287e19c | |||
| f88c2a95c1 | |||
| ed55612cfa | |||
| 47ade040e6 | |||
| 852e723cc8 | |||
| f90e5f4519 | |||
| c3d1cc87e2 | |||
| 88c2659a9c | |||
| 6f7498a956 | |||
| a1cea1a1ea | |||
| 60765469aa | |||
| 68ad291a8f | |||
| e08d812598 | |||
| 45fb61000a | |||
| 403961da51 | |||
| 64efe18500 | |||
| 98360ad31f | |||
| 09a1c98514 | |||
| aff0105700 | |||
| 2db839aef9 | |||
| 9967155332 | |||
| 7e857a546b | |||
| a341d254e1 | |||
| 6db1224d95 | |||
| 3c0a24feb7 | |||
| 38f8c96b66 | |||
| c3c7c2ca91 | |||
| feba73a11a | |||
| 16f4683859 | |||
| 41f46569bf | |||
| ac475ba509 | |||
| ef1e7eb8e6 | |||
| cb4d7102b5 | |||
| 984224959a | |||
| f655f09ce1 | |||
| dd99401f3e | |||
| 087eb2259d | |||
| 8108d3fd28 | |||
| a894a4e952 | |||
| c1fac1e688 | |||
| 9e432dbd2d | |||
| 3611defe5d | |||
| c77be502f9 | |||
| ec9d1bff5b | |||
| bb922b2a21 | |||
| 8a802241f7 | |||
| a9a26cf856 | |||
| b31736cde7 | |||
| 50e04e0aee | |||
| 68cad9c0f1 | |||
| 9c3dd9be82 | |||
| d1ea505b7a | |||
| 63a83c69c3 | |||
| 181db38b49 | |||
| b1fa8bea43 | |||
| d73075bea0 | |||
| 71d1d99c44 | |||
| 4ee016d3a7 | |||
| 9923743219 | |||
| 637cbd9135 | |||
| 88661545b4 | |||
| de3c452f20 | |||
| d832f90ad6 | |||
| 7bc20a6302 | |||
| e2b3bb7088 | |||
| 209afebefa | |||
| 9332971384 | |||
| 799a0c72e9 | |||
| d07fe1de66 | |||
| 49c9e27502 | |||
| b4a8f92cfa | |||
| 15d5598eec | |||
| dd4c4c8c0f | |||
| 6d7a17e765 | |||
| 53d0796bd5 | |||
| eeb397f066 | |||
| 75727c6b36 | |||
| f08595910d | |||
| 11f0066b47 | |||
| aa62c98247 | |||
| cf587616dd | |||
| 8e7ce1b1d3 | |||
| 2360ad173a | |||
| f82729f88e | |||
| d3157d2535 | |||
| 58dda66f84 | |||
| b38e3360c5 | |||
| 46f108b6f3 | |||
| 8554e7b29c | |||
| 9fb321e45e | |||
| b383711f6d | |||
| 689ea48d72 | |||
| 00209dedff | |||
| 56243e1835 | |||
| 5c3f50dfa5 | |||
| 965d7b2002 | |||
| 6d552b035a | |||
| 05bef8642d | |||
| 019f08093d | |||
| 45d635afae | |||
| 03933ade35 | |||
| dcfede7e69 | |||
| cf65589b93 | |||
| 707b812f8b | |||
| 694aff9801 | |||
| 89cdc60b6e | |||
| ec9488d94e | |||
| 35189fde0e | |||
| ea256fca0d | |||
| 3dc878f012 | |||
| 578d09ac7c | |||
| c4708d7d5d | |||
| 763e634dfd | |||
| 1e011714dd | |||
| 6d529e3f79 | |||
| bfa6844124 | |||
| 14325e7690 | |||
| c074d977ce | |||
| 123303240d | |||
| 31d5e5d727 | |||
| 6758bbfbd9 | |||
| 45e74a97bd | |||
| 46777cb99a | |||
| 75a1be4a71 | |||
| 43880b1965 | |||
| 5b699aaa79 | |||
| 510de69250 | |||
| 8bb4e11f31 | |||
| 38f0394166 | |||
| 65d8ab5fe3 | |||
| c48e583790 | |||
| 530d77ff1b | |||
| 463aae19a9 | |||
| 45bede6579 | |||
| 0afddcdfdc | |||
| 6a464bf653 | |||
| 483eb0b2fb | |||
| a64b30c429 | |||
| 8881d65c8d | |||
| 037c4b11df | |||
| b51fc899ca | |||
| e19831f7d5 | |||
| 9b68db9e82 | |||
| 879afbb1d4 | |||
| abd9392c90 | |||
| 689bf3eed3 | |||
| b51f30b49a | |||
| 1e8e114550 | |||
| 93111c6eaf | |||
| e2547ec301 | |||
| 825fd60972 | |||
| eafeaf333d | |||
| f1b0d61ada | |||
| 7a11fd2846 | |||
| 0266dc9e92 | |||
| 74f64731ab | |||
| 341ea870bb | |||
| a7c3f8f516 | |||
| 35dcc69ba5 | |||
| 2c60caf790 | |||
| 2e4cddc840 | |||
| 311f4d7b68 | |||
| 77daa38cf0 | |||
| d286dbf203 | |||
| 066feee3ea | |||
| 0c7b0b19af | |||
| 563e7837b9 | |||
| f32e4baf5b | |||
| da929845a9 | |||
| 4acea97f03 | |||
| 18f361d485 | |||
| dbe6e4b4f3 | |||
| 42d2e58570 | |||
| 1e879ca3c4 | |||
| 3fa23d16c6 | |||
| bc3abb127a | |||
| b90cef3968 | |||
| e88dfb8f98 | |||
| b6af1c1b4a | |||
| e9fb72fdc7 | |||
| 8d5885d683 | |||
| f1e0503c73 | |||
| 97f8bc34bc | |||
| be5341fb03 | |||
| 85d8371261 | |||
| ef93b7f919 | |||
| b08e867c1a | |||
| bd6aacf3ed | |||
| cb02ede5fb | |||
| d15744a812 | |||
| 6a8e55cc43 | |||
| e1da5c797d | |||
| 807c2c6194 | |||
| 2536d91430 | |||
| 1f391644ca | |||
| 2ea3d01b58 | |||
| c863f01a78 | |||
| 501ba36b89 | |||
| 77b6dee02f | |||
| ba435fb1d7 | |||
| ceca2ae8e3 | |||
| db1f62227b | |||
| fc0153d0de | |||
| af46a7b041 | |||
| f1cbfa8e67 | |||
| 2e6655c398 | |||
| 1b332f86e6 | |||
| 1f4c987652 | |||
| c773dd7eae | |||
| 81468df9c0 | |||
| 6c8b6d81fe | |||
| cff3f0b1a8 | |||
| 1b421e30f9 | |||
| b805c294eb |
@@ -0,0 +1,28 @@
|
||||
{
|
||||
"version": "0.0.1",
|
||||
"configurations": [
|
||||
{
|
||||
"name": "mc2",
|
||||
"runtimeExecutable": "F:\\Coding Stuff\\mission-control-2\\backend\\.venv\\Scripts\\python.exe",
|
||||
"runtimeArgs": [
|
||||
"-m",
|
||||
"uvicorn",
|
||||
"app:app",
|
||||
"--app-dir",
|
||||
"F:\\Coding Stuff\\mission-control-2\\backend",
|
||||
"--port",
|
||||
"9000"
|
||||
],
|
||||
"port": 9000
|
||||
},
|
||||
{
|
||||
"name": "frontend",
|
||||
"runtimeExecutable": "npm",
|
||||
"runtimeArgs": ["run", "dev", "--", "--port", "5180", "--strictPort"],
|
||||
"cwd": "F:\\Coding Stuff\\mission-control-2\\frontend",
|
||||
"env": { "MC_API_TARGET": "http://192.168.178.151:9001" },
|
||||
"autoPort": false,
|
||||
"port": 5180
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,12 @@
|
||||
# Shell-/Deploy-Skripte MÜSSEN LF behalten — sonst bricht der Deploy auf der Box
|
||||
# (CRLF macht `set -euo pipefail` zu `pipefail\r` → "invalid option name").
|
||||
*.sh text eol=lf
|
||||
*.bash text eol=lf
|
||||
|
||||
# systemd-Units und Service-Configs ebenfalls LF.
|
||||
*.service text eol=lf
|
||||
*.timer text eol=lf
|
||||
|
||||
# Windows-Batch-Wrapper bleiben CRLF.
|
||||
*.cmd text eol=crlf
|
||||
*.bat text eol=crlf
|
||||
+22
@@ -0,0 +1,22 @@
|
||||
# Python
|
||||
backend/.venv/
|
||||
__pycache__/
|
||||
*.pyc
|
||||
|
||||
# Node / Vite
|
||||
frontend/node_modules/
|
||||
# frontend/dist wird committet (kein Node-Build auf der Box) — siehe deploy/
|
||||
|
||||
# Avatar-VRM (groß + lizenz-/redistributionssensibel) — liegt lokal + auf der Box, nicht in git.
|
||||
# Wird per Direkt-Deploy auf die Box gespielt (dist/avatar.vrm), nicht über git.
|
||||
frontend/public/avatar.vrm
|
||||
frontend/dist/avatar.vrm
|
||||
|
||||
# Env / local
|
||||
*.env
|
||||
.DS_Store
|
||||
|
||||
# Box-Recon-/Scratch-Skripte (lokale Diagnose, nicht fürs Repo)
|
||||
box_recon*
|
||||
gemma_swap*
|
||||
|
||||
@@ -0,0 +1,68 @@
|
||||
# Mission Control 2.0
|
||||
|
||||
Komponierbarer Local-AI-Stack für den Bosgame M5. Greenfield-Neuaufbau —
|
||||
siehe Architektur-Plan (`docs/` bzw. der genehmigte Plan).
|
||||
|
||||
**Schichten:** Engine (llama-swap, **Vulkan/RADV** auf Strix Halo) · **Builtin-Routing-Gateway**
|
||||
in MC2 (`model: auto`, kein externer LiteLLM-Dienst — scheitert auf Python 3.14) ·
|
||||
Mission Control 2.0 (FastAPI + React/shadcn) · Hermes Agent + hermes-webui ·
|
||||
Shared Memory (SQLite via MCP). Jede Schicht hinter stabilem Vertrag austauschbar.
|
||||
|
||||
## Status: Phasen 0–5 ✅ · MC2 **live auf der Box** (:9001) · Modelle/Hermes-Wiring + Cutover offen
|
||||
|
||||
Fortschritt & Resume-Guide: siehe [`docs/STATUS.md`](docs/STATUS.md).
|
||||
|
||||
- **Phase 0** — FastAPI-Skeleton + React/shadcn-Shell (Cmd+K, Dark, PWA).
|
||||
- **Phase 1** — Compute-Module (fit/caps/sources, portiert), **Discover** (live HF + Fit + Caps),
|
||||
**Engine-Write** (register + `groups`/Ko-Residenz + vocab-geprüfte Spec-Drafts),
|
||||
**Builtin-Gateway** (`model: auto` + Fallbacks), Frontend **Modelle & Routing** (Caps-Chips, Fit, Discover, Routing-View).
|
||||
|
||||
- **Phase 2** — System-Status (CPU/RAM/GPU/Disk), Wartung (restart/self-update, sudo-frei),
|
||||
**Connect** (saubere IDE-Snippets → Gateway `model:auto`, LAN-IP-Override).
|
||||
- **Phase 3** — Geteiltes **Gedächtnis** (SQLite/WAL, 5 Kategorien, Dedupe-Kurator) + **MCP-Server**
|
||||
(`mcp/mcp_memory.py` shared, `mcp/mcp_mc.py` Stack-Management für Hermes), MemoryView.
|
||||
|
||||
- **Phase 4** — Hermes-**Agent-Status** (`/api/agent/status`, AgentView mit Tiles + „Hermes öffnen"),
|
||||
`deploy/hermes-webui.service`, **Box-Runbook** [`docs/HERMES_SETUP.md`](docs/HERMES_SETUP.md)
|
||||
(hermes-webui, Brain=`auto`, Tools/MCP-Verdrahtung). Box-Ausführung steht noch aus.
|
||||
|
||||
- **Phase 5** — **Backup** (Memory + Configs), **Services-Health** + Observability-Links,
|
||||
**Theme-Toggle** (Hell/Dunkel). Box-Deploy/-Wiring + Cutover (Phase 6) brauchen die Box.
|
||||
|
||||
API: `health · models · discover · fit · models/register · groups · routing · system/* · connect ·
|
||||
memory/* · agent/status` (Details in `docs/STATUS.md`). MCP: `mcp/` (siehe `mcp/requirements.txt`).
|
||||
Box-Runbooks: `docs/HERMES_SETUP.md` + `deploy/` (Units, deploy.sh, backup.sh).
|
||||
|
||||
## Entwickeln
|
||||
|
||||
**Backend:**
|
||||
```bash
|
||||
cd backend
|
||||
python -m venv .venv && .venv/Scripts/python -m pip install -r requirements.txt # Windows
|
||||
.venv/Scripts/python -m uvicorn app:app --port 9000
|
||||
```
|
||||
|
||||
**Frontend (Dev, proxyt /api → :9000):**
|
||||
```bash
|
||||
cd frontend
|
||||
npm install
|
||||
npm run dev # http://localhost:5173
|
||||
```
|
||||
|
||||
**Frontend (Build → wird vom Backend ausgeliefert):**
|
||||
```bash
|
||||
cd frontend && npm run build # → frontend/dist
|
||||
```
|
||||
|
||||
## Env-Vars (Auswahl)
|
||||
|
||||
| Variable | Default | Zweck |
|
||||
|---|---|---|
|
||||
| `MC_LLAMA_SWAP_URL` | `http://127.0.0.1:8080` | Engine |
|
||||
| `MC_CONFIG_PATH` | `/etc/llama-swap/config.yaml` | llama-swap Config |
|
||||
| `MC_GATEWAY_URL` | `http://127.0.0.1:$MC_PORT` | Builtin-Gateway (Teil von MC2, kein externer Dienst) |
|
||||
| `MC_PORT` | `9000` | MC-Backend-Port |
|
||||
| `MC_ENGINE_PATH` | `/opt/llamacpp-vulkan` | Aktive Engine-Binary (Vulkan-Build) |
|
||||
| `MC_ENGINE_REPO` | `ggml-org/llama.cpp` | Quelle für Engine-Update-Check |
|
||||
| `MC_DRAFTS_DIR` | `$MODELS/drafts` | Spec-Draft-Modelle (vocab-geprüft) |
|
||||
| `MC_SPEC_TYPE` | `draft-simple` | Speculative-Decoding-Typ (llama.cpp) |
|
||||
@@ -0,0 +1,98 @@
|
||||
"""
|
||||
Mission Control 2.0 — dünner FastAPI-Einstieg.
|
||||
|
||||
Hängt die Router ein, liefert (in Prod) das gebaute React-Frontend aus und
|
||||
setzt eine no-cache-Middleware. Im Dev läuft das Frontend über den Vite-Dev-
|
||||
Server (proxyt /api hierher), daher CORS für localhost offen.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
import os
|
||||
from contextlib import asynccontextmanager
|
||||
|
||||
from fastapi import FastAPI
|
||||
from fastapi.middleware.cors import CORSMiddleware
|
||||
from fastapi.responses import FileResponse
|
||||
from fastapi.staticfiles import StaticFiles
|
||||
from starlette.requests import Request
|
||||
|
||||
from config import FRONTEND_DIST, VERSION
|
||||
from routers import agent, connect, gateway_proxy, health, maintenance, memory, models, reminders as reminders_router, routing, system, voice
|
||||
from services import reminders, sentry, warmer
|
||||
|
||||
# Zentrales Logging — Level via MC_LOG_LEVEL (INFO default). Eine Konfiguration
|
||||
# für alle Module (logging.getLogger(__name__)).
|
||||
logging.basicConfig(
|
||||
level=os.environ.get("MC_LOG_LEVEL", "INFO").upper(),
|
||||
format="%(asctime)s %(levelname)-7s %(name)s: %(message)s",
|
||||
)
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
@asynccontextmanager
|
||||
async def lifespan(app: FastAPI):
|
||||
"""Hintergrund-Tasks an den App-Lebenszyklus binden: Re-Warm-Wächter fürs Agent-Hirn
|
||||
+ Health-Wächter (meldet Ausfälle/Erholung in den Lucy-Briefkasten und auf Telegram)."""
|
||||
tasks = []
|
||||
if warmer.ENABLED:
|
||||
tasks.append(asyncio.create_task(warmer.rewarm_loop()))
|
||||
log.info("Hirn-Re-Warm-Wächter aktiv (Intervall %ss, Hirn dynamisch aus Hermes-Config)", warmer.INTERVAL)
|
||||
if sentry.ENABLED:
|
||||
tasks.append(asyncio.create_task(sentry.sentry_loop()))
|
||||
tasks.append(asyncio.create_task(reminders.reminders_loop()))
|
||||
try:
|
||||
yield
|
||||
finally:
|
||||
for task in tasks:
|
||||
task.cancel()
|
||||
|
||||
|
||||
app = FastAPI(title="Mission Control 2.0", version=VERSION, lifespan=lifespan)
|
||||
|
||||
# Dev: Vite-Dev-Server (5173) ruft das Backend per /api auf.
|
||||
app.add_middleware(
|
||||
CORSMiddleware,
|
||||
allow_origins=["http://localhost:5173", "http://127.0.0.1:5173"],
|
||||
allow_methods=["*"],
|
||||
allow_headers=["*"],
|
||||
)
|
||||
|
||||
|
||||
@app.middleware("http")
|
||||
async def no_cache(request: Request, call_next):
|
||||
resp = await call_next(request)
|
||||
if request.url.path.startswith("/api"):
|
||||
resp.headers["Cache-Control"] = "no-cache"
|
||||
return resp
|
||||
|
||||
|
||||
app.include_router(health.router)
|
||||
app.include_router(models.router)
|
||||
app.include_router(routing.router)
|
||||
app.include_router(system.router)
|
||||
app.include_router(connect.router)
|
||||
app.include_router(memory.router)
|
||||
app.include_router(agent.router)
|
||||
app.include_router(voice.router) # Sprache: STT/TTS-Proxy + Hermes-Agent-Chat (Voice-Tab)
|
||||
app.include_router(reminders_router.router) # Erinnerungen/Routinen (A3) — feuern in den Briefkasten
|
||||
app.include_router(gateway_proxy.router) # OpenAI-kompatibler /v1-Gateway (model:auto)
|
||||
app.include_router(maintenance.router)
|
||||
|
||||
|
||||
# Prod: gebautes Frontend ausliefern (falls vorhanden). SPA-Fallback auf index.html.
|
||||
if FRONTEND_DIST.exists():
|
||||
app.mount("/assets", StaticFiles(directory=FRONTEND_DIST / "assets"), name="assets")
|
||||
|
||||
@app.get("/{full_path:path}")
|
||||
def spa(full_path: str):
|
||||
# Falls die Datei direkt in FRONTEND_DIST liegt (z.B. manifest.webmanifest, favicon.ico), liefere sie aus
|
||||
target = FRONTEND_DIST / full_path
|
||||
if target.is_file():
|
||||
return FileResponse(target)
|
||||
|
||||
index = FRONTEND_DIST / "index.html"
|
||||
if index.exists():
|
||||
# index.html nie cachen → Browser zieht nach jedem Deploy das aktuelle (gehashte) Bundle.
|
||||
return FileResponse(index, headers={"Cache-Control": "no-cache, must-revalidate"})
|
||||
return {"detail": "frontend not built"}
|
||||
|
||||
@@ -0,0 +1,118 @@
|
||||
"""
|
||||
Zentrale Konfiguration für Mission Control 2.0.
|
||||
|
||||
Eine Quelle der Wahrheit für Pfade, URLs und Defaults — alles über Env-Vars
|
||||
überschreibbar. Bewusst schlank: MC 2.0 ist ein Glue-Cockpit, das vorhandene
|
||||
Dienste (llama-swap, LiteLLM-Gateway, Hermes) steuert, statt sie nachzubauen.
|
||||
"""
|
||||
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
from ruamel.yaml import YAML
|
||||
|
||||
# --- Engine (llama-swap) -----------------------------------------------------
|
||||
LLAMA_SWAP_URL = os.environ.get("MC_LLAMA_SWAP_URL", "http://127.0.0.1:8080").rstrip("/")
|
||||
CONFIG_PATH = Path(os.environ.get("MC_CONFIG_PATH", "/etc/llama-swap/config.yaml"))
|
||||
MODELS_DIR = Path(os.environ.get("MC_MODELS_DIR", "/srv/models"))
|
||||
# Cache der Modell-Entdeckung ("aktuell beste Modelle", live von HuggingFace).
|
||||
# Persistent neben den Modellen (übersteht Deploys). TTL = Frische-Fenster.
|
||||
DISCOVER_CACHE_PATH = Path(os.environ.get("MC_DISCOVER_CACHE", str(MODELS_DIR / "mc2-discover.json")))
|
||||
DISCOVER_TTL = int(os.environ.get("MC_DISCOVER_TTL", "43200")) # 12 h
|
||||
# Geteiltes Gedächtnis (SQLite, WAL). Persistent neben den Modellen.
|
||||
# Hinweis: nur noch für die einmalige Mem0-Migration relevant — das aktive Gedächtnis
|
||||
# liegt jetzt in Mem0/Chroma hinter dem Sidecar (siehe MEM0_SERVICE_URL).
|
||||
MEMORY_DB = Path(os.environ.get("MC_MEMORY_DB", str(MODELS_DIR / "mc2-memory.db")))
|
||||
# Mem0-Sidecar (auto-lernendes, semantisches Gedächtnis). Läuft im ~/.mem0/venv (Python 3.12),
|
||||
# weil mem0+chromadb unter dem 3.14-Backend nicht laufen. MC2 spricht ihn lokal per HTTP an.
|
||||
MEM0_SERVICE_URL = os.environ.get("MC_MEM0_SERVICE_URL", "http://127.0.0.1:8765").rstrip("/")
|
||||
# Befehl-Vorlage für llama-swap: {model}=GGUF-Pfad, {ctx}=Kontext, ${PORT} bleibt stehen.
|
||||
# Hinweis: --prompt-cache/--prompt-cache-all sind llama-CLI-Flags, NICHT llama-server —
|
||||
# llama-server lehnt sie ab ("invalid argument") und startet dann nicht. Prompt-Caching
|
||||
# macht llama-server ohnehin automatisch pro Slot (KV-Reuse).
|
||||
_DEFAULT_CMD_TEMPLATE = (
|
||||
"llama-server -m {model} --host 127.0.0.1 --port ${PORT} "
|
||||
"-c {ctx} -ngl 999 -fa on --no-mmap"
|
||||
)
|
||||
CMD_TEMPLATE = os.environ.get("MC_CMD_TEMPLATE", _DEFAULT_CMD_TEMPLATE)
|
||||
if "{model}" not in CMD_TEMPLATE:
|
||||
CMD_TEMPLATE = _DEFAULT_CMD_TEMPLATE
|
||||
DEFAULT_TTL = int(os.environ.get("MC_DEFAULT_TTL", "300"))
|
||||
# Verzeichnis mit Draft-Modellen für Speculative Decoding. Beim Hinzufügen eines
|
||||
# fast/coder-Modells wird hieraus automatisch ein **vocab-kompatibler** Draft gewählt
|
||||
# (Vocab-Check via services.gguf_meta; ein inkompatibler Draft lässt llama.cpp scheitern).
|
||||
DRAFTS_DIR = Path(os.environ.get("MC_DRAFTS_DIR", str(MODELS_DIR / "drafts")))
|
||||
# Optionaler expliziter Default-Draft (leer = Auto-Erkennung aus DRAFTS_DIR). Wird nur
|
||||
# verwendet, wenn er zum Ziel-Modell vocab-kompatibel ist. (Früher fix qwen2.5 → entfernt,
|
||||
# weil das mit neueren Vocabs wie Qwen3.6 inkompatibel ist und Spec stillschweigend brach.)
|
||||
SPEC_DRAFT_MODEL_PATH = os.environ.get("MC_SPEC_DRAFT_MODEL", "")
|
||||
# Speculative-Decoding-Typ (llama.cpp dieser Generation braucht --spec-type zusätzlich
|
||||
# zu --spec-draft-model, sonst ist Spec inaktiv).
|
||||
SPEC_TYPE = os.environ.get("MC_SPEC_TYPE", "draft-simple")
|
||||
# MTP-Speculative-Decoding (Multi-Token-Prediction): manche Modelle bringen einen eigenen
|
||||
# MTP-Kopf mit (z.B. gemma-4 → arch 'gemma4-assistant', Datei 'mtp-*.gguf'). Der wird mit
|
||||
# `--model-draft <mtp.gguf> --spec-type draft-mtp --spec-draft-n-max N` geladen (NICHT
|
||||
# --spec-draft-model/draft-simple). 1,5–2× Durchsatz bei null Qualitätsverlust.
|
||||
SPEC_DRAFT_N_MAX = int(os.environ.get("MC_SPEC_DRAFT_N_MAX", "4"))
|
||||
# Env für HuggingFace-Downloads: XET deaktivieren (Hänger bei ~6 MB, siehe v1-Gotcha).
|
||||
HF_DOWNLOAD_ENV = {"HF_HUB_DISABLE_XET": "1"}
|
||||
|
||||
# --- Routing-Gateway (builtin in MC2, model: auto) ---------------------------
|
||||
# MC2 IST der Gateway (services/gateway.py + routers/gateway_proxy.py). KEIN externer
|
||||
# LiteLLM-Dienst (scheitert auf Python 3.14). Daher keine Gateway-Config-Datei mehr.
|
||||
GATEWAY_URL = os.environ.get("MC_GATEWAY_URL", f"http://127.0.0.1:{os.environ.get('MC_PORT', '9000')}").rstrip("/")
|
||||
|
||||
# --- Hermes Agent (eigener Dienst auf der Box) -------------------------------
|
||||
# Gateway (OpenAI-API des Agenten) + interaktives Web-Terminal (ttyd → `hermes chat`).
|
||||
HERMES_API_URL = os.environ.get("HERMES_API_URL", "http://127.0.0.1:8642").rstrip("/")
|
||||
# API-Key der Hermes-`api_server`-Plattform (~/.hermes/.env: API_SERVER_KEY). Nötig für
|
||||
# /v1/chat/completions (Voice-Pipeline) — Bearer-Auth, sonst 401. Derselbe volle Agent
|
||||
# (Tools + geteiltes Mem0) wie CLI/Telegram, nur über HTTP.
|
||||
def _read_hermes_env(key: str) -> str:
|
||||
"""Liest einen Schlüssel aus ~/.hermes/.env (Fallback, falls nicht in der Prozess-Env).
|
||||
Der MC2-Dienst erbt die Hermes-Secrets sonst nicht."""
|
||||
try:
|
||||
env_path = Path(os.path.expanduser(os.environ.get("HERMES_HOME", "~/.hermes"))) / ".env"
|
||||
for line in env_path.read_text(encoding="utf-8").splitlines():
|
||||
line = line.strip()
|
||||
if line.startswith(f"{key}="):
|
||||
return line.split("=", 1)[1].strip().strip('"').strip("'")
|
||||
except OSError:
|
||||
pass
|
||||
return ""
|
||||
|
||||
|
||||
HERMES_API_KEY = (
|
||||
os.environ.get("HERMES_API_KEY")
|
||||
or os.environ.get("API_SERVER_KEY")
|
||||
or _read_hermes_env("API_SERVER_KEY")
|
||||
)
|
||||
# Modellfeld im OpenAI-Request; die api_server-Plattform nutzt ihr konfiguriertes Hirn,
|
||||
# das Feld ist i.d.R. kosmetisch. Override via Env, falls die Plattform strikt prüft.
|
||||
HERMES_API_MODEL = os.environ.get("HERMES_API_MODEL", "hermes")
|
||||
|
||||
# --- Voice-Sidecar (STT faster-whisper + TTS Piper/Chatterbox) ---------------
|
||||
# Eigenes Python-3.12-venv (~/.voice/venv), analog Mem0-Sidecar. MC2 proxyt nach außen.
|
||||
VOICE_SERVICE_URL = os.environ.get("MC_VOICE_SERVICE_URL", "http://127.0.0.1:8650").rstrip("/")
|
||||
# Hermes-Terminal: ttyd-Web-Terminal der interaktiven Agent-CLI (Ersatz für AnythingLLM-Chat).
|
||||
# Wird in MC2 per iframe eingebettet (Terminal-Seite). Siehe deploy/hermes-terminal.service.
|
||||
HERMES_TERMINAL_URL = os.environ.get("MC_HERMES_TERMINAL_URL", "http://192.168.178.151:7681").rstrip("/")
|
||||
# GitHub-Repo für Update-Checks.
|
||||
HERMES_AGENT_REPO = os.environ.get("MC_HERMES_AGENT_REPO", "NousResearch/hermes-agent")
|
||||
HERMES_HOME = Path(os.path.expanduser(os.environ.get("HERMES_HOME", "~/.hermes")))
|
||||
# PC Executor — läuft auf dem Windows-PC, erreichbar über LAN.
|
||||
PC_EXECUTOR_URL = os.environ.get("MC_PC_EXECUTOR_URL", "http://192.168.178.98:7777").rstrip("/")
|
||||
|
||||
# --- Server ------------------------------------------------------------------
|
||||
HOST = os.environ.get("MC_HOST", "0.0.0.0")
|
||||
PORT = int(os.environ.get("MC_PORT", "9000"))
|
||||
# Gebautes React-Frontend (frontend/dist). In Prod liefert FastAPI es statisch aus;
|
||||
# im Dev läuft der Vite-Dev-Server separat und proxyt /api hierher.
|
||||
FRONTEND_DIST = Path(os.environ.get("MC_FRONTEND_DIST", str(Path(__file__).resolve().parent.parent / "frontend" / "dist")))
|
||||
|
||||
# Version (Phase 0 — Greenfield-Skeleton).
|
||||
VERSION = "2.0.0-w8"
|
||||
|
||||
# Gemeinsame YAML-Instanz (preserve_quotes hält Kommentare/Quotes in config.yaml).
|
||||
yaml = YAML()
|
||||
yaml.preserve_quotes = True
|
||||
@@ -0,0 +1,56 @@
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
# Add backend directory to sys.path so we can import services
|
||||
sys.path.append(str(Path(__file__).resolve().parent))
|
||||
|
||||
from services.llamaswap import read_config, write_config, spec_draft_flags, _PATH_RE
|
||||
from config import CONFIG_PATH
|
||||
|
||||
def migrate():
|
||||
print(f"Reading config from {CONFIG_PATH}...")
|
||||
if not CONFIG_PATH.exists():
|
||||
print(f"Config path {CONFIG_PATH} does not exist. Skipping.")
|
||||
return
|
||||
|
||||
cfg = read_config()
|
||||
models = cfg.get("models", {})
|
||||
|
||||
for name, spec in models.items():
|
||||
if not isinstance(spec, dict):
|
||||
continue
|
||||
cmd = spec.get("cmd", "")
|
||||
if not cmd:
|
||||
continue
|
||||
|
||||
print(f"Migrating model: {name}")
|
||||
|
||||
# 1. Defektes --prompt-cache/--prompt-cache-all entfernen (llama-CLI-Flags,
|
||||
# die llama-server ablehnt → Start scheitert). Caching macht llama-server
|
||||
# automatisch pro Slot.
|
||||
cmd = cmd.replace(" --prompt-cache-all", "").replace(" --prompt-cache", "")
|
||||
|
||||
# 2. Extract aliases/role
|
||||
aliases = spec.get("aliases", [])
|
||||
role = aliases[0] if aliases else None
|
||||
|
||||
# 3. Add parallel + (nur vocab-kompatibles) Speculative Decoding für fast/coder.
|
||||
# spec_draft_flags() prüft die Vocab-Kompatibilität und hängt --spec-type an;
|
||||
# ein inkompatibler Draft (z.B. qwen2.5 ↔ Qwen3.6) wird NICHT gesetzt.
|
||||
if role in ("fast", "coder"):
|
||||
if "--parallel" not in cmd:
|
||||
cmd = cmd.strip() + " --parallel 2"
|
||||
if "--spec-draft-model" not in cmd:
|
||||
target = mt.group(1) if (mt := _PATH_RE.search(cmd)) else ""
|
||||
cmd = cmd.strip() + spec_draft_flags(target)
|
||||
|
||||
# Update cmd
|
||||
from ruamel.yaml.scalarstring import LiteralScalarString
|
||||
spec["cmd"] = LiteralScalarString(cmd.strip() + "\n")
|
||||
|
||||
print(f"Writing updated config back to {CONFIG_PATH}...")
|
||||
write_config(cfg)
|
||||
print("Migration completed successfully!")
|
||||
|
||||
if __name__ == "__main__":
|
||||
migrate()
|
||||
@@ -0,0 +1,60 @@
|
||||
{
|
||||
"_comment": "Kuratierter Modell-Katalog (Cookbook) für Strix Halo / Ryzen AI MAX+ 395 — 128GB unified, bandbreiten-limitiert (256 GB/s). MoE-first. EINE Quelle der Wahrheit für KORREKTE Metadaten (total/active params, moe, generation) → präzise Empfehlungen ohne Namens-Raterei. Inspiriert vom Odysseus-Cookbook (statischer, validierter Katalog statt Live-Scraping). Erweiterbar: neue Modelle hier eintragen. Felder: name (Match-Identifier), repo (HF org/name für Install), family (+Subtyp), generation (numerisch, für Upgrade-Vergleich), total_params_b, active_params_b (=total bei dense), moe, quant, ctx (empfohlen), tools, vision.",
|
||||
"version": "2026-06-27",
|
||||
"models": [
|
||||
{
|
||||
"role": "fast", "name": "Qwen3.6-35B-A3B", "repo": "Qwen/Qwen3.6-35B-A3B-GGUF",
|
||||
"family": "qwen", "generation": 3.6, "total_params_b": 35, "active_params_b": 3,
|
||||
"moe": true, "quant": "Q4_K_M", "ctx": 32768, "tools": true, "vision": true
|
||||
},
|
||||
{
|
||||
"role": "fast", "name": "Qwen3-30B-A3B-Instruct", "repo": "unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF",
|
||||
"family": "qwen", "generation": 3.0, "total_params_b": 30, "active_params_b": 3,
|
||||
"moe": true, "quant": "Q4_K_M", "ctx": 32768, "tools": true, "vision": false
|
||||
},
|
||||
|
||||
{
|
||||
"role": "heavy", "name": "Qwen3.5-122B-A10B", "repo": "Qwen/Qwen3.5-122B-A10B-GGUF",
|
||||
"family": "qwen", "generation": 3.5, "total_params_b": 122, "active_params_b": 10,
|
||||
"moe": true, "quant": "Q4_K_M", "ctx": 32768, "tools": true, "vision": false
|
||||
},
|
||||
{
|
||||
"role": "heavy", "name": "gpt-oss-120b", "repo": "ggml-org/gpt-oss-120b-GGUF",
|
||||
"family": "gpt-oss", "generation": 1.0, "total_params_b": 120, "active_params_b": 5,
|
||||
"moe": true, "quant": "MXFP4", "ctx": 32768, "tools": true, "vision": false
|
||||
},
|
||||
|
||||
{
|
||||
"role": "coder", "name": "Qwen3-Coder-Next", "repo": "Qwen/Qwen3-Coder-Next-GGUF",
|
||||
"family": "qwen-coder", "generation": 3.0, "total_params_b": 84, "active_params_b": 3,
|
||||
"moe": true, "quant": "Q4_K_M", "ctx": 65536, "tools": true, "vision": false
|
||||
},
|
||||
{
|
||||
"role": "coder", "name": "Qwen3-Coder-30B-A3B-Instruct", "repo": "unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF",
|
||||
"family": "qwen-coder", "generation": 3.0, "total_params_b": 30, "active_params_b": 3,
|
||||
"moe": true, "quant": "Q4_K_M", "ctx": 65536, "tools": true, "vision": false
|
||||
},
|
||||
|
||||
{
|
||||
"role": "vision", "name": "Qwen3-VL-8B-Instruct", "repo": "Qwen/Qwen3-VL-8B-Instruct-GGUF",
|
||||
"family": "qwen-vl", "generation": 3.0, "total_params_b": 8, "active_params_b": 8,
|
||||
"moe": false, "quant": "Q4_K_M", "ctx": 32768, "tools": false, "vision": true
|
||||
},
|
||||
{
|
||||
"role": "vision", "name": "Qwen3-VL-2B-Instruct", "repo": "Qwen/Qwen3-VL-2B-Instruct-GGUF",
|
||||
"family": "qwen-vl", "generation": 3.0, "total_params_b": 2, "active_params_b": 2,
|
||||
"moe": false, "quant": "Q4_K_M", "ctx": 32768, "tools": false, "vision": true
|
||||
},
|
||||
|
||||
{
|
||||
"role": "scout", "name": "gemma-4-26B-A4B-it", "repo": "google/gemma-4-26B-A4B-it-GGUF",
|
||||
"family": "gemma", "generation": 4.0, "total_params_b": 26, "active_params_b": 4,
|
||||
"moe": true, "quant": "Q4_K_M", "ctx": 32768, "tools": false, "vision": true
|
||||
},
|
||||
{
|
||||
"role": "scout", "name": "gemma-4-31B-it", "repo": "google/gemma-4-31B-it-GGUF",
|
||||
"family": "gemma", "generation": 4.0, "total_params_b": 31, "active_params_b": 31,
|
||||
"moe": false, "quant": "Q4_K_M", "ctx": 32768, "tools": false, "vision": true
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,8 @@
|
||||
fastapi>=0.115
|
||||
uvicorn[standard]>=0.30
|
||||
httpx>=0.27
|
||||
ruamel.yaml>=0.18
|
||||
psutil>=5.9
|
||||
huggingface_hub>=0.27
|
||||
mcp>=1.2.0
|
||||
tzdata>=2024.1
|
||||
@@ -0,0 +1,44 @@
|
||||
"""Agent-Endpoint: Hermes-Status + WebUI-Link (MC verlinkt nur, betreibt nicht)."""
|
||||
|
||||
from fastapi import APIRouter
|
||||
from pydantic import BaseModel
|
||||
|
||||
from fastapi import HTTPException
|
||||
|
||||
from services.agent import agent_status, hermes_brain_info, set_agent_brain, update_brain_model
|
||||
|
||||
router = APIRouter(prefix="/api")
|
||||
|
||||
|
||||
class BrainReq(BaseModel):
|
||||
model: str
|
||||
|
||||
|
||||
class SetBrainReq(BaseModel):
|
||||
model_id: str
|
||||
|
||||
|
||||
@router.get("/agent/status")
|
||||
def status() -> dict:
|
||||
return agent_status()
|
||||
|
||||
|
||||
@router.get("/agent/brain")
|
||||
def brain_info() -> dict:
|
||||
"""Aktuelles Agent-Hirn (hermes) + bestes NousResearch-Hermes-Update."""
|
||||
return hermes_brain_info()
|
||||
|
||||
|
||||
@router.post("/agent/brain/set")
|
||||
def set_brain(body: SetBrainReq) -> dict:
|
||||
"""Setzt ein installiertes Modell als Agent-Hirn (Alias + warm-Gruppe + Config)."""
|
||||
res = set_agent_brain(body.model_id)
|
||||
if not res.get("ok"):
|
||||
raise HTTPException(400, res.get("reason", "Fehler beim Setzen des Agent-Hirns"))
|
||||
return res
|
||||
|
||||
|
||||
@router.post("/agent/brain")
|
||||
def set_brain_model(body: BrainReq) -> dict:
|
||||
ok = update_brain_model(body.model)
|
||||
return {"ok": ok}
|
||||
@@ -0,0 +1,21 @@
|
||||
"""Connect-Endpoint: erzeugt IDE-/Agent-Snippets (auf den Gateway + Memory-MCP)."""
|
||||
|
||||
from fastapi import APIRouter
|
||||
|
||||
from services.connect import DEFAULT_HOST, build_snippets, check_health
|
||||
|
||||
router = APIRouter(prefix="/api")
|
||||
|
||||
|
||||
@router.get("/connect")
|
||||
def connect(host: str = DEFAULT_HOST, mcp_path: str | None = None) -> dict:
|
||||
kwargs = {}
|
||||
if mcp_path:
|
||||
kwargs["mcp_script_path"] = mcp_path
|
||||
return build_snippets(host=host, **kwargs)
|
||||
|
||||
|
||||
@router.get("/connect/health")
|
||||
def connect_health() -> dict:
|
||||
"""Live-Status der zwei Leitungen (Gateway + Gedächtnis) für den Verbinden-Tab."""
|
||||
return check_health()
|
||||
@@ -0,0 +1,123 @@
|
||||
import os
|
||||
|
||||
import httpx
|
||||
from fastapi import APIRouter, Request
|
||||
from fastapi.responses import JSONResponse, StreamingResponse
|
||||
|
||||
from config import LLAMA_SWAP_URL
|
||||
from services.gateway_stream import record_stream_chunk, record_usage
|
||||
from services.router_logic import LANES, choose_for_lane
|
||||
from services.routing_policy import load_policy
|
||||
|
||||
router = APIRouter(prefix="/v1")
|
||||
|
||||
# Antwortsprache für IDE-/Lane-Traffic: die Coding-Modelle antworten sonst englisch
|
||||
# (User-Anforderung 03.07.2026). Leerer String (MC_GATEWAY_LANG_DIRECTIVE="") schaltet ab.
|
||||
_LANG_DIRECTIVE = os.environ.get(
|
||||
"MC_GATEWAY_LANG_DIRECTIVE",
|
||||
"Antworte dem Nutzer grundsätzlich auf Deutsch (Erklärungen, Pläne, Rückfragen, "
|
||||
"Zusammenfassungen) — auch wenn die Frage oder Tool-Anweisungen englisch sind. "
|
||||
"Quellcode, Bezeichner und Shell-Befehle bleiben unverändert.")
|
||||
|
||||
|
||||
def _inject_language(body: dict, alias: str) -> None:
|
||||
"""Deutsch-Direktive anhängen. An die ERSTE System-Message (viele Chat-Templates
|
||||
erwarten nur eine), sonst als neue System-Message. `hermes` ausgenommen — Lucys
|
||||
Persona (SOUL.md) regelt die Sprache selbst."""
|
||||
if not _LANG_DIRECTIVE or alias == "hermes":
|
||||
return
|
||||
msgs = body.get("messages")
|
||||
if not isinstance(msgs, list):
|
||||
return
|
||||
first_sys = next((m for m in msgs if isinstance(m, dict) and m.get("role") == "system"), None)
|
||||
if first_sys is None:
|
||||
msgs.insert(0, {"role": "system", "content": _LANG_DIRECTIVE})
|
||||
elif isinstance(first_sys.get("content"), str):
|
||||
first_sys["content"] = first_sys["content"].rstrip() + "\n\n" + _LANG_DIRECTIVE
|
||||
elif isinstance(first_sys.get("content"), list):
|
||||
first_sys["content"].append({"type": "text", "text": _LANG_DIRECTIVE})
|
||||
|
||||
# Virtuelle Lanes, die der Gateway zusätzlich zu den echten Modellen als „Modell" anbietet.
|
||||
_LANE_LABELS = {"coding": "Coding (Router → coder/heavy/fast)", "chat": "Chat (Router → fast/heavy)"}
|
||||
|
||||
|
||||
@router.get("/models")
|
||||
async def models():
|
||||
async with httpx.AsyncClient(timeout=10) as c:
|
||||
r = await c.get(f"{LLAMA_SWAP_URL}/v1/models")
|
||||
data = r.json()
|
||||
# Lanes ganz oben einblenden, damit IDEs einfach „coding"/„chat" wählen können.
|
||||
lanes = [{"id": lane, "object": "model", "owned_by": "mc2-router",
|
||||
"description": _LANE_LABELS.get(lane, lane)} for lane in LANES]
|
||||
# Kontextlänge je Modell mitliefern (aus der llama-swap-Config geparst). Ohne sie
|
||||
# budgetieren Clients blind — Hermes-Subagents nahmen 256k an, schickten passende
|
||||
# max_tokens und rissen damit den echten Server-Kontext (Radar-Lauf 02.07.).
|
||||
# Rollen-Aliase (heavy/coder/hermes …) tauchen bei llama-swap NICHT als Einträge auf,
|
||||
# Clients fragen aber genau damit an → als eigene Einträge einblenden.
|
||||
ctx_map: dict[str, int] = {}
|
||||
alias_entries: list[dict] = []
|
||||
try:
|
||||
from services import llamaswap
|
||||
for m in llamaswap.list_models():
|
||||
ctx = m.get("ctx")
|
||||
if ctx:
|
||||
for api_id in m.get("api_ids", []):
|
||||
ctx_map[api_id] = ctx
|
||||
for alias in m.get("aliases", []):
|
||||
entry = {"id": alias, "object": "model", "owned_by": "mc2-alias",
|
||||
"description": f"Alias für {m['name']}"}
|
||||
if ctx:
|
||||
entry["context_length"] = ctx
|
||||
alias_entries.append(entry)
|
||||
except Exception:
|
||||
pass
|
||||
if isinstance(data, dict) and isinstance(data.get("data"), list):
|
||||
for entry in data["data"]:
|
||||
if (ctx := ctx_map.get(entry.get("id"))):
|
||||
entry.setdefault("context_length", ctx)
|
||||
data["data"] = lanes + alias_entries + data["data"]
|
||||
return JSONResponse(data, status_code=r.status_code)
|
||||
|
||||
|
||||
async def _proxy(path: str, request: Request):
|
||||
body = await request.json()
|
||||
requested = str(body.get("model") or "auto")
|
||||
if requested.lower() in ("auto", "chat", "coding"):
|
||||
lane = requested.lower()
|
||||
alias, reason = choose_for_lane(lane, body)
|
||||
body["model"] = alias
|
||||
routed = {"x-mc-routed-to": alias, "x-mc-route-reason": reason, "x-mc-lane": lane}
|
||||
else:
|
||||
alias = requested
|
||||
routed = {"x-mc-routed-to": requested}
|
||||
# fast-Spur: Thinking aus für flotte Antworten (sofern Client es nicht selbst setzt).
|
||||
pol = load_policy()
|
||||
if pol["fast_no_think"] and alias == pol["fast"] and "chat_template_kwargs" not in body:
|
||||
body["chat_template_kwargs"] = {"enable_thinking": False}
|
||||
_inject_language(body, alias)
|
||||
url = f"{LLAMA_SWAP_URL}{path}"
|
||||
|
||||
if body.get("stream"):
|
||||
async def gen():
|
||||
async with httpx.AsyncClient(timeout=None) as c:
|
||||
async with c.stream("POST", url, json=body) as r:
|
||||
async for chunk in r.aiter_raw():
|
||||
record_stream_chunk(chunk, alias)
|
||||
yield chunk
|
||||
return StreamingResponse(gen(), media_type="text/event-stream", headers=routed)
|
||||
|
||||
async with httpx.AsyncClient(timeout=600) as c:
|
||||
r = await c.post(url, json=body)
|
||||
resp_json = r.json()
|
||||
record_usage(resp_json.get("usage") if isinstance(resp_json, dict) else None, alias)
|
||||
return JSONResponse(resp_json, status_code=r.status_code, headers=routed)
|
||||
|
||||
|
||||
@router.post("/chat/completions")
|
||||
async def chat_completions(request: Request):
|
||||
return await _proxy("/v1/chat/completions", request)
|
||||
|
||||
|
||||
@router.post("/completions")
|
||||
async def completions(request: Request):
|
||||
return await _proxy("/v1/completions", request)
|
||||
@@ -0,0 +1,21 @@
|
||||
"""Health-/Status-Endpoint — schlanker Lebenszeichen-Check für MC 2.0."""
|
||||
|
||||
from fastapi import APIRouter
|
||||
|
||||
from config import VERSION
|
||||
from services import gateway, llamaswap
|
||||
|
||||
router = APIRouter(prefix="/api")
|
||||
|
||||
|
||||
@router.get("/health")
|
||||
def health() -> dict:
|
||||
return {
|
||||
"status": "ok",
|
||||
"version": VERSION,
|
||||
"engine_reachable": llamaswap.engine_reachable(),
|
||||
"gateway_reachable": gateway.gateway_reachable(),
|
||||
# Echte Hirn-Bereitschaft: Engine kann erreichbar sein, das Agent-Hirn ('fast') aber tot
|
||||
# (Crash/OOM nach Engine-Update). Das wäre sonst ein silent fail (App-Fehler statt Status).
|
||||
"brain": llamaswap.brain_status(),
|
||||
}
|
||||
@@ -0,0 +1,80 @@
|
||||
"""Wartungs-Endpoints: Update-Badge, OS-/Engine-Update, Reboot, Restart, Logs."""
|
||||
|
||||
from fastapi import APIRouter, HTTPException, Header
|
||||
from pydantic import BaseModel
|
||||
|
||||
from services import maintenance
|
||||
|
||||
router = APIRouter(prefix="/api")
|
||||
|
||||
|
||||
class SudoReq(BaseModel):
|
||||
sudo_password: str | None = None
|
||||
|
||||
|
||||
class RestartReq(BaseModel):
|
||||
service: str
|
||||
sudo_password: str | None = None
|
||||
|
||||
|
||||
@router.get("/maintenance/updates")
|
||||
def updates() -> dict:
|
||||
return maintenance.updates()
|
||||
|
||||
|
||||
@router.get("/maintenance/update-details")
|
||||
def update_details(kind: str) -> dict:
|
||||
if kind not in ("os", "engine", "swap", "hermes"):
|
||||
raise HTTPException(400, "Unbekannte Update-Art.")
|
||||
return maintenance.update_details(kind)
|
||||
|
||||
@router.post("/maintenance/check-updates")
|
||||
def check_updates(body: SudoReq) -> dict:
|
||||
res = maintenance.check_updates_job(body.sudo_password)
|
||||
if isinstance(res, dict) and not res.get("ok", True):
|
||||
return res
|
||||
return res
|
||||
|
||||
|
||||
@router.post("/maintenance/os-update")
|
||||
def os_update(body: SudoReq) -> dict:
|
||||
res = maintenance.os_update_job(body.sudo_password)
|
||||
if isinstance(res, dict) and not res.get("ok", True):
|
||||
return res
|
||||
return res
|
||||
|
||||
|
||||
@router.post("/maintenance/engine-update")
|
||||
def engine_update(body: SudoReq) -> dict:
|
||||
res = maintenance.engine_update_job(body.sudo_password)
|
||||
if not res:
|
||||
raise HTTPException(400, "Kein Engine-Update-Befehl gesetzt (MC_ENGINE_UPDATE_CMD).")
|
||||
return res
|
||||
|
||||
|
||||
@router.post("/maintenance/swap-update")
|
||||
def swap_update(body: SudoReq) -> dict:
|
||||
res = maintenance.swap_update_job(body.sudo_password)
|
||||
if not res:
|
||||
raise HTTPException(400, "Kein Router-Update-Befehl gesetzt (MC_SWAP_UPDATE_CMD).")
|
||||
return res
|
||||
|
||||
|
||||
@router.post("/maintenance/hermes-update")
|
||||
def hermes_update() -> dict:
|
||||
return maintenance.hermes_update_job()
|
||||
|
||||
|
||||
@router.post("/maintenance/reboot")
|
||||
def reboot(body: SudoReq) -> dict:
|
||||
return maintenance.reboot(body.sudo_password)
|
||||
|
||||
|
||||
@router.post("/maintenance/restart")
|
||||
def restart(body: RestartReq) -> dict:
|
||||
return maintenance.restart_service(body.service, body.sudo_password)
|
||||
|
||||
|
||||
@router.get("/maintenance/logs")
|
||||
def logs(service: str, lines: int = 200, x_sudo_password: str | None = Header(None)) -> dict:
|
||||
return maintenance.logs(service, lines, x_sudo_password)
|
||||
@@ -0,0 +1,81 @@
|
||||
"""Memory-Endpoints (geteiltes Gedächtnis). LAN-only, kein Token in 2.0-Phase 3."""
|
||||
|
||||
from fastapi import APIRouter, HTTPException
|
||||
from pydantic import BaseModel
|
||||
|
||||
from services import memory
|
||||
|
||||
router = APIRouter(prefix="/api")
|
||||
|
||||
|
||||
class MemIn(BaseModel):
|
||||
content: str
|
||||
category: str = "knowledge"
|
||||
source: str = "manual"
|
||||
|
||||
|
||||
class MemUp(BaseModel):
|
||||
content: str | None = None
|
||||
category: str | None = None
|
||||
|
||||
|
||||
class DedupeIn(BaseModel):
|
||||
apply: bool = False
|
||||
threshold: float = 0.85
|
||||
|
||||
|
||||
class LearnIn(BaseModel):
|
||||
text: str | None = None
|
||||
messages: list[dict] | None = None
|
||||
source: str = "auto"
|
||||
category: str = "knowledge"
|
||||
|
||||
|
||||
@router.get("/memory/export")
|
||||
def export() -> dict:
|
||||
return memory.export_text()
|
||||
|
||||
|
||||
@router.get("/memory/graph")
|
||||
def graph(min_score: float = 0.45, top_k: int = 3) -> dict:
|
||||
"""Fakten als Ähnlichkeits-Graph (Knoten + semantische Kanten) für die Visualisierung."""
|
||||
return memory.graph(min_score=min_score, top_k=top_k)
|
||||
|
||||
|
||||
@router.post("/memory/learn", status_code=201)
|
||||
def learn(body: LearnIn) -> dict:
|
||||
"""Auto-Lernen: Gesprächs-Turns/Text durchreichen → Mem0 extrahiert Fakten selbst."""
|
||||
return memory.learn(text=body.text, messages=body.messages,
|
||||
source=body.source, category=body.category)
|
||||
|
||||
|
||||
@router.post("/memory/dedupe")
|
||||
def dedupe(body: DedupeIn) -> dict:
|
||||
return memory.dedupe(apply=body.apply, threshold=body.threshold)
|
||||
|
||||
|
||||
@router.get("/memory")
|
||||
def list_mem(q: str = "", category: str = "") -> list[dict]:
|
||||
return memory.list_memories(q=q, category=category)
|
||||
|
||||
|
||||
@router.post("/memory", status_code=201)
|
||||
def add(body: MemIn) -> dict:
|
||||
if body.category not in memory.CATEGORIES:
|
||||
raise HTTPException(400, f"Kategorie '{body.category}' unbekannt.")
|
||||
return memory.add_memory(body.content, body.category, body.source)
|
||||
|
||||
|
||||
@router.put("/memory/{mid}")
|
||||
def update(mid: str, body: MemUp) -> dict:
|
||||
res = memory.update_memory(mid, content=body.content, category=body.category)
|
||||
if not res:
|
||||
raise HTTPException(404, "Eintrag nicht gefunden")
|
||||
return res
|
||||
|
||||
|
||||
@router.delete("/memory/{mid}")
|
||||
def delete(mid: str) -> dict:
|
||||
if not memory.delete_memory(mid):
|
||||
raise HTTPException(404, "Eintrag nicht gefunden")
|
||||
return {"ok": True}
|
||||
@@ -0,0 +1,293 @@
|
||||
"""Modelle-Endpoints: Liste (mit Caps), Discover, Fit, Register, Groups."""
|
||||
|
||||
import psutil
|
||||
from fastapi import APIRouter, HTTPException
|
||||
from pydantic import BaseModel
|
||||
|
||||
from config import HF_DOWNLOAD_ENV, MODELS_DIR
|
||||
from services import budget, discover, hf, jobengine, llamaswap
|
||||
from services.fit import evaluate_fit, max_ctx_for
|
||||
|
||||
router = APIRouter(prefix="/api")
|
||||
|
||||
|
||||
def _ram_gb() -> float:
|
||||
return psutil.virtual_memory().total / (1024 ** 3)
|
||||
|
||||
|
||||
@router.get("/models")
|
||||
def models() -> dict:
|
||||
items = llamaswap.list_models()
|
||||
return {"models": items, "count": len(items), "running": llamaswap.get_running_models()}
|
||||
|
||||
|
||||
@router.get("/discover")
|
||||
def discover_models(force: bool = False) -> dict:
|
||||
ram = _ram_gb()
|
||||
data = discover.refresh_discover(ram) if force else discover.safe_discover(ram)
|
||||
if not data:
|
||||
raise HTTPException(502, "Modell-Quellen gerade nicht erreichbar — später erneut.")
|
||||
return {**data, "sys_ram_gb": round(ram, 1)}
|
||||
|
||||
|
||||
@router.get("/fit")
|
||||
def fit(params_b: float = 0, quant: str = "Q4_K_M", ctx: int = 8192,
|
||||
name: str = "", role: str = "") -> dict:
|
||||
"""Hardware-Fit-Vorschau. params_b<=0 → aus KATALOG (echte Metadaten, MoE-bewusst)
|
||||
oder sonst aus dem Namen geschätzt. assigned_ctx = der ctx, der TATSÄCHLICH vergeben
|
||||
würde: SETUP-BEWUSST (neben Hirn/warmem Set), nicht nur gegen den Gesamt-RAM.
|
||||
So sieht die 'Erweiterte Ansicht' vor dem Download Ampel + echten ctx."""
|
||||
ram = _ram_gb()
|
||||
pb = params_b if params_b > 0 else budget.params_b_for(name)
|
||||
saw = budget.setup_aware_ctx(pb, quant, role=role or None)
|
||||
return {
|
||||
"params_b": round(pb, 1),
|
||||
"fit": evaluate_fit(pb, quant, ctx, ram, name=name),
|
||||
"optimal_ctx": max_ctx_for(pb, quant, ram), # Roh-Obergrenze (Modell allein)
|
||||
"assigned_ctx": saw["ctx"], # setup-bewusst vergeben
|
||||
"budget": {"gtt_gb": saw["gtt_gb"], "reserved_gb": saw["reserved_gb"],
|
||||
"budget_gb": saw["budget_gb"], "mode": saw["mode"]},
|
||||
"sys_ram_gb": round(ram, 1),
|
||||
}
|
||||
|
||||
|
||||
class RegisterReq(BaseModel):
|
||||
model_path: str
|
||||
role: str | None = None
|
||||
ctx: int = 8192
|
||||
ttl: int | None = None
|
||||
mmproj_path: str | None = None
|
||||
jinja: bool = False
|
||||
|
||||
|
||||
@router.post("/models/register")
|
||||
def register(req: RegisterReq) -> dict:
|
||||
try:
|
||||
model_id = llamaswap.register_model(
|
||||
req.model_path, role=req.role, ctx=req.ctx, ttl=req.ttl,
|
||||
mmproj_path=req.mmproj_path, jinja=req.jinja,
|
||||
)
|
||||
except PermissionError as exc:
|
||||
raise HTTPException(500, str(exc))
|
||||
return {"ok": True, "model_id": model_id}
|
||||
|
||||
|
||||
class InstallReq(BaseModel):
|
||||
repo: str
|
||||
role: str | None = None
|
||||
quant: str = "Q4_K_M"
|
||||
ctx: int | None = None
|
||||
jinja: bool = False
|
||||
hf_token: str | None = None
|
||||
|
||||
|
||||
@router.get("/hf/search")
|
||||
def hf_search(q: str = "") -> dict:
|
||||
return {"results": hf.search(q)}
|
||||
|
||||
|
||||
@router.get("/hf/quants")
|
||||
def hf_quants(repo: str) -> dict:
|
||||
repo = hf.normalize_repo(repo)
|
||||
return {"repo": repo, "quants": hf.list_quants(repo)}
|
||||
|
||||
|
||||
@router.post("/models/install")
|
||||
def install(req: InstallReq) -> dict:
|
||||
"""Lädt ein Modell von HuggingFace (Hintergrund-Job) UND trägt es sofort in
|
||||
llama-swap ein (cmd + Rolle-Alias). llama-swap (-watch-config) lädt es, sobald
|
||||
die Datei da ist. Split-GGUFs werden komplett geladen, registriert wird der
|
||||
erste Teil (-00001-of-…). Akzeptiert volle HF-URL ODER org/repo."""
|
||||
repo = hf.normalize_repo(req.repo)
|
||||
info = hf.resolve_gguf(repo, req.quant)
|
||||
if not info["first"]:
|
||||
raise HTTPException(404, f"Keine GGUF-Datei für Quant '{req.quant}' in {repo} gefunden.")
|
||||
|
||||
subdir = repo.split("/")[-1]
|
||||
target = MODELS_DIR / subdir
|
||||
target.mkdir(parents=True, exist_ok=True)
|
||||
model_path = str(target / info["first"])
|
||||
mmproj_path = str(target / info["mmproj"]) if info["mmproj"] else None
|
||||
|
||||
ctx = req.ctx
|
||||
if ctx is None:
|
||||
# SETUP-BEWUSST: größter ctx, der neben Hirn/warmem Set passt (nicht nur Modell allein).
|
||||
ctx = budget.setup_aware_ctx(budget.params_b_for(repo), req.quant, role=req.role)["ctx"]
|
||||
|
||||
# Sofort registrieren (Datei kommt gleich) — robust gegen -watch-config.
|
||||
try:
|
||||
model_id = llamaswap.register_model(
|
||||
model_path, role=req.role, ctx=ctx, mmproj_path=mmproj_path, jinja=req.jinja)
|
||||
except PermissionError as exc:
|
||||
raise HTTPException(500, str(exc))
|
||||
|
||||
# Download-Job: alle GGUF-Teile (+ mmproj) per --include holen.
|
||||
args = [hf.hf_bin(), "download", repo]
|
||||
for f in info["files"]:
|
||||
args.append(f)
|
||||
if info["mmproj"]:
|
||||
args.append(info["mmproj"])
|
||||
args += ["--local-dir", str(target)]
|
||||
env = dict(HF_DOWNLOAD_ENV)
|
||||
if req.hf_token:
|
||||
env["HF_TOKEN"] = req.hf_token
|
||||
job_id = jobengine.start_job(args, f"download {req.repo}", env=env)
|
||||
jobengine.attach_download_progress(job_id, str(target), info["total_bytes"])
|
||||
return {"ok": True, "job_id": job_id, "model_id": model_id, "model_path": model_path,
|
||||
"total_bytes": info["total_bytes"], "files": len(info["files"])}
|
||||
|
||||
|
||||
@router.get("/jobs")
|
||||
def jobs() -> dict:
|
||||
return {"jobs": jobengine.public_jobs()}
|
||||
|
||||
|
||||
@router.post("/jobs/{job_id}/cancel")
|
||||
def cancel(job_id: str) -> dict:
|
||||
return {"ok": jobengine.cancel_job(job_id)}
|
||||
|
||||
|
||||
class RoleReq(BaseModel):
|
||||
role: str | None = None
|
||||
|
||||
|
||||
@router.get("/roles/{role}/recommend")
|
||||
def recommend_role(role: str) -> dict:
|
||||
"""Welches installierte Modell passt am besten auf diese Rolle? (Capability + setup-
|
||||
bewusster Fit). Basis für 'Empfohlen'-Hinweis + Auto-Pick im Rollen-Zuweisungs-Modal."""
|
||||
from services import roles
|
||||
return roles.recommend_for_role(role)
|
||||
|
||||
|
||||
@router.post("/models/{model_id}/role")
|
||||
def set_model_role(model_id: str, body: RoleReq) -> dict:
|
||||
# Das Agent-Hirn (Rolle 'hermes') braucht den warm-bewussten Flow (Alias + brains-Gruppe +
|
||||
# ttl 0 + Hermes config.default + Gateway-Restart) — Single Source of Truth UI ↔ Hermes.
|
||||
if (body.role or "").strip().lower() == "hermes":
|
||||
from services.agent import set_agent_brain
|
||||
res = set_agent_brain(model_id)
|
||||
if not res.get("ok"):
|
||||
raise HTTPException(400, res.get("reason", "Fehler beim Setzen des Agent-Hirns"))
|
||||
return res
|
||||
if not llamaswap.set_role(model_id, body.role):
|
||||
raise HTTPException(404, "Modell nicht gefunden")
|
||||
return {"ok": True}
|
||||
|
||||
|
||||
class CtxReq(BaseModel):
|
||||
ctx: int
|
||||
|
||||
|
||||
@router.get("/models/{model_id}/ctx/auto")
|
||||
def auto_ctx(model_id: str) -> dict:
|
||||
"""Setup-bewusster Optimal-ctx für ein bestehendes Modell (Rolle/Params/Quant +
|
||||
aktuelles Setup). Basis für den 'Auto'-Button an der Modellkarte."""
|
||||
m = next((x for x in llamaswap.list_models() if x["name"] == model_id), None)
|
||||
if not m:
|
||||
raise HTTPException(404, "Modell nicht gefunden")
|
||||
saw = budget.setup_aware_ctx_for_model(m)
|
||||
return {"model_id": model_id, "current_ctx": m.get("ctx"),
|
||||
"params_b": round(budget.params_of_model(m), 1), "quant": m.get("quant"),
|
||||
"role": m.get("role"), **saw}
|
||||
|
||||
|
||||
@router.post("/models/{model_id}/ctx")
|
||||
def set_model_ctx(model_id: str, body: CtxReq) -> dict:
|
||||
if not llamaswap.set_ctx(model_id, body.ctx):
|
||||
raise HTTPException(404, "Modell nicht gefunden")
|
||||
return {"ok": True}
|
||||
|
||||
|
||||
@router.get("/models/drafts")
|
||||
def list_drafts(target: str = "") -> dict:
|
||||
"""Verfügbare Draft-Modelle + ihre Vocab-Kompatibilität zum Ziel-Modell
|
||||
(target = GGUF-Pfad). Basis für die idiotensichere Spec-Draft-Auswahl im UI."""
|
||||
return llamaswap.drafts_for(target)
|
||||
|
||||
|
||||
class DraftReq(BaseModel):
|
||||
draft_path: str | None = None
|
||||
|
||||
|
||||
@router.post("/models/{model_id}/draft")
|
||||
def set_model_draft(model_id: str, body: DraftReq) -> dict:
|
||||
"""Setzt/entfernt den Speculative-Decoding-Draft eines Modells. Inkompatible
|
||||
(oder nicht prüfbare) Drafts werden serverseitig abgelehnt."""
|
||||
try:
|
||||
res = llamaswap.set_spec_draft(model_id, body.draft_path)
|
||||
except PermissionError as exc:
|
||||
raise HTTPException(500, str(exc))
|
||||
if not res["ok"]:
|
||||
raise HTTPException(400 if "kompatib" in res["reason"].lower() else 404, res["reason"])
|
||||
return res
|
||||
|
||||
|
||||
@router.post("/models/unload")
|
||||
def unload_all_models() -> dict:
|
||||
import httpx
|
||||
from config import LLAMA_SWAP_URL
|
||||
try:
|
||||
with httpx.Client(timeout=10.0) as c:
|
||||
r = c.post(f"{LLAMA_SWAP_URL}/api/models/unload")
|
||||
return {"ok": r.status_code == 200}
|
||||
except Exception as exc:
|
||||
raise HTTPException(500, str(exc))
|
||||
|
||||
|
||||
@router.post("/models/{model_id}/unload")
|
||||
def unload_model(model_id: str) -> dict:
|
||||
import httpx
|
||||
from config import LLAMA_SWAP_URL
|
||||
try:
|
||||
with httpx.Client(timeout=10.0) as c:
|
||||
r = c.post(f"{LLAMA_SWAP_URL}/api/models/unload/{model_id}")
|
||||
return {"ok": r.status_code == 200}
|
||||
except Exception as exc:
|
||||
raise HTTPException(500, str(exc))
|
||||
|
||||
|
||||
@router.post("/models/{model_id}/load")
|
||||
def load_model(model_id: str) -> dict:
|
||||
import httpx
|
||||
from config import LLAMA_SWAP_URL
|
||||
try:
|
||||
# Trigger load by sending a lightweight completion request.
|
||||
body = {
|
||||
"model": model_id,
|
||||
"messages": [{"role": "user", "content": "ping"}],
|
||||
"max_tokens": 1
|
||||
}
|
||||
# High timeout because model loading might take time
|
||||
with httpx.Client(timeout=60.0) as c:
|
||||
c.post(f"{LLAMA_SWAP_URL}/v1/chat/completions", json=body)
|
||||
return {"ok": True}
|
||||
except Exception as exc:
|
||||
raise HTTPException(500, str(exc))
|
||||
|
||||
|
||||
@router.delete("/models/{model_id}")
|
||||
def delete(model_id: str) -> dict:
|
||||
if not llamaswap.delete_model(model_id):
|
||||
raise HTTPException(404, "Modell nicht gefunden")
|
||||
return {"ok": True}
|
||||
|
||||
|
||||
@router.get("/groups")
|
||||
def groups() -> dict:
|
||||
return {"groups": llamaswap.list_groups()}
|
||||
|
||||
|
||||
class GroupReq(BaseModel):
|
||||
group: str
|
||||
members: list[str]
|
||||
swap: bool = False
|
||||
persist: bool = False
|
||||
|
||||
|
||||
@router.put("/groups")
|
||||
def set_group(req: GroupReq) -> dict:
|
||||
try:
|
||||
llamaswap.set_group(req.group, req.members, swap=req.swap, persist=req.persist)
|
||||
except PermissionError as exc:
|
||||
raise HTTPException(500, str(exc))
|
||||
return {"ok": True}
|
||||
@@ -0,0 +1,37 @@
|
||||
"""Erinnerungen & Routinen (A3) — dünner REST-Layer über services/reminders.py.
|
||||
Konsument ist v.a. der Hermes-Agent via mcp_mc.py (reminder_create/list/delete);
|
||||
LAN-only wie alle MC2-Endpoints."""
|
||||
|
||||
from fastapi import APIRouter, HTTPException
|
||||
from pydantic import BaseModel
|
||||
|
||||
from services import reminders
|
||||
|
||||
router = APIRouter(prefix="/api")
|
||||
|
||||
|
||||
class ReminderIn(BaseModel):
|
||||
text: str # was angesagt werden soll
|
||||
when: str # ISO-8601; ohne Offset = Commander-Zeitzone (Europe/Berlin)
|
||||
repeat: str = "" # '' einmalig | daily | weekdays | weekly
|
||||
|
||||
|
||||
@router.get("/reminders")
|
||||
def list_reminders() -> dict:
|
||||
return {"items": reminders.list_all()}
|
||||
|
||||
|
||||
@router.post("/reminders")
|
||||
def create_reminder(body: ReminderIn) -> dict:
|
||||
try:
|
||||
return {"ok": True, "item": reminders.create(body.text, body.when, body.repeat)}
|
||||
except ValueError as exc:
|
||||
raise HTTPException(400, str(exc))
|
||||
|
||||
|
||||
@router.delete("/reminders/{reminder_id}")
|
||||
def delete_reminder(reminder_id: int) -> dict:
|
||||
try:
|
||||
return {"ok": True, "deleted": reminders.delete(reminder_id)}
|
||||
except KeyError as exc:
|
||||
raise HTTPException(404, str(exc.args[0]))
|
||||
@@ -0,0 +1,43 @@
|
||||
"""Routing-Endpoints: Lane-Summary (chat/coding) + UI-editierbare Policy (hot-reload)."""
|
||||
|
||||
from fastapi import APIRouter, HTTPException
|
||||
from pydantic import BaseModel
|
||||
|
||||
from services import gateway
|
||||
from services.routing_policy import policy_meta, save_policy
|
||||
|
||||
router = APIRouter(prefix="/api")
|
||||
|
||||
|
||||
@router.get("/routing")
|
||||
def routing() -> dict:
|
||||
return {**gateway.routing_summary(), "gateway_reachable": gateway.gateway_reachable()}
|
||||
|
||||
|
||||
@router.get("/routing/policy")
|
||||
def get_policy() -> dict:
|
||||
"""Aktuelle Policy + Defaults (für „Zurücksetzen“) + Feld-Spezifikation für den Editor."""
|
||||
return policy_meta()
|
||||
|
||||
|
||||
class PolicyPatch(BaseModel):
|
||||
fast: str | None = None
|
||||
heavy: str | None = None
|
||||
coder: str | None = None
|
||||
coder_lite: str | None = None
|
||||
heavy_chars: int | None = None
|
||||
coding_escalate_chars: int | None = None
|
||||
fast_no_think: bool | None = None
|
||||
|
||||
|
||||
@router.put("/routing/policy")
|
||||
def put_policy(patch: PolicyPatch) -> dict:
|
||||
"""Teil-Update der Routing-Policy. Validiert, persistiert atomar, sofort wirksam (hot-reload)."""
|
||||
fields = {k: v for k, v in patch.model_dump().items() if v is not None}
|
||||
if not fields:
|
||||
raise HTTPException(status_code=400, detail="Keine Felder zum Aktualisieren.")
|
||||
try:
|
||||
new_policy = save_policy(fields)
|
||||
except (ValueError, TypeError) as e:
|
||||
raise HTTPException(status_code=400, detail=f"Ungültige Policy: {e}")
|
||||
return {"policy": new_policy}
|
||||
@@ -0,0 +1,128 @@
|
||||
"""System-Endpoints: Live-Status + Wartung (Restart/Self-Update — auf der Box).
|
||||
|
||||
Wartung läuft als systemd-USER-Dienst → KEIN sudo/Passwort (Nordstern).
|
||||
Lokal (Windows) schlagen die Shell-Befehle harmlos fehl und werden als Fehler
|
||||
zurückgegeben statt zu crashen.
|
||||
"""
|
||||
|
||||
import logging
|
||||
import os
|
||||
import subprocess
|
||||
|
||||
from fastapi import APIRouter, HTTPException
|
||||
from pydantic import BaseModel
|
||||
|
||||
import httpx
|
||||
|
||||
from config import GATEWAY_URL, HERMES_API_URL, LLAMA_SWAP_URL, MEM0_SERVICE_URL, VOICE_SERVICE_URL
|
||||
from services import backup as backup_svc
|
||||
from services.agent import agent_status
|
||||
from services.gateway import gateway_reachable
|
||||
from services.llamaswap import engine_reachable, list_models
|
||||
from services.pricing import compute_savings
|
||||
from services.system import system_status
|
||||
from services.token_stats import get_stats
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(prefix="/api")
|
||||
|
||||
# Nur diese User-Dienste dürfen neugestartet werden.
|
||||
ALLOWED_SERVICES = {"mission-control-2", "hermes-gateway", "hermes-webui", "mem0-service", "voice-service"}
|
||||
# Quelle für Self-Update (auf der Box ~/mission-control-v2).
|
||||
SOURCE_DIR = os.path.expanduser(os.environ.get("MC2_SOURCE_DIR", "~/mission-control-v2"))
|
||||
|
||||
|
||||
@router.get("/system/status")
|
||||
def status() -> dict:
|
||||
return system_status()
|
||||
|
||||
|
||||
def _mem0_reachable() -> bool:
|
||||
try:
|
||||
return httpx.get(f"{MEM0_SERVICE_URL}/health", timeout=2).status_code == 200
|
||||
except Exception:
|
||||
return False
|
||||
|
||||
|
||||
def _voice_reachable() -> bool:
|
||||
try:
|
||||
return httpx.get(f"{VOICE_SERVICE_URL}/health", timeout=2).status_code == 200
|
||||
except Exception:
|
||||
return False
|
||||
|
||||
|
||||
@router.get("/system/services")
|
||||
def services() -> dict:
|
||||
"""Aggregierte Erreichbarkeit aller Stack-Dienste (für die Health-Anzeige)."""
|
||||
a = agent_status()
|
||||
gw_url = f"{GATEWAY_URL}/v1"
|
||||
return {
|
||||
"services": [
|
||||
{"name": "Engine (llama-swap)", "unit": "llama-swap", "url": LLAMA_SWAP_URL, "ok": engine_reachable()},
|
||||
{"name": "Gateway (integriert)", "unit": "mission-control-2", "url": gw_url, "ok": gateway_reachable()},
|
||||
{"name": "Hermes-Gateway", "unit": "hermes-gateway", "url": HERMES_API_URL, "ok": a["gateway_reachable"]},
|
||||
{"name": "Hermes-Terminal", "unit": "hermes-terminal", "url": a["terminal_url"], "ok": a["terminal_reachable"]},
|
||||
{"name": "Mem0 (Gedächtnis)", "unit": "mem0-service", "url": MEM0_SERVICE_URL, "ok": _mem0_reachable()},
|
||||
{"name": "Voice (STT/TTS)", "unit": "voice-service", "url": VOICE_SERVICE_URL, "ok": _voice_reachable()},
|
||||
],
|
||||
"links": {
|
||||
"engine_ui": f"{LLAMA_SWAP_URL}/ui",
|
||||
"gateway": gw_url,
|
||||
"hermes_terminal": a["terminal_url"],
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
@router.post("/system/backup")
|
||||
def backup() -> dict:
|
||||
return backup_svc.backup_now()
|
||||
|
||||
|
||||
@router.get("/system/backups")
|
||||
def backups() -> dict:
|
||||
return {"backups": backup_svc.list_backups()}
|
||||
|
||||
|
||||
def _run(cmd: list[str], cwd: str | None = None) -> dict:
|
||||
try:
|
||||
p = subprocess.run(cmd, cwd=cwd, capture_output=True, text=True, timeout=180)
|
||||
return {"ok": p.returncode == 0, "code": p.returncode,
|
||||
"out": (p.stdout or "")[-2000:], "err": (p.stderr or "")[-2000:]}
|
||||
except Exception as exc: # noqa: BLE001
|
||||
return {"ok": False, "code": -1, "out": "", "err": str(exc)}
|
||||
|
||||
|
||||
class RestartReq(BaseModel):
|
||||
service: str
|
||||
|
||||
|
||||
@router.post("/system/restart")
|
||||
def restart(req: RestartReq) -> dict:
|
||||
if req.service not in ALLOWED_SERVICES:
|
||||
raise HTTPException(400, f"Dienst '{req.service}' nicht erlaubt.")
|
||||
return _run(["systemctl", "--user", "restart", req.service])
|
||||
|
||||
|
||||
@router.post("/system/self-update")
|
||||
def self_update() -> dict:
|
||||
"""git pull (Source) → venv-Deps → Dienst-Restart. Auf der Box; lokal Fehler."""
|
||||
pull = _run(["git", "fetch", "--all"], cwd=SOURCE_DIR)
|
||||
reset = _run(["git", "reset", "--hard", "origin/main"], cwd=SOURCE_DIR)
|
||||
restart_res = _run(["systemctl", "--user", "restart", "mission-control-2"])
|
||||
return {"pull": pull, "reset": reset, "restart": restart_res}
|
||||
|
||||
|
||||
@router.get("/system/token-stats")
|
||||
def token_stats() -> dict:
|
||||
"""Token-Verbrauch + Cloud-Ersparnis. Logik im pricing-Service (SSoT)."""
|
||||
# Rolle je Modell/Alias (lowercase) für die Tarif-Auflösung auflösen.
|
||||
role_map: dict[str, str | None] = {}
|
||||
try:
|
||||
for m in list_models():
|
||||
role_map[m["name"].lower()] = m.get("role")
|
||||
for alias in m.get("aliases", []):
|
||||
role_map[alias.lower()] = m.get("role")
|
||||
except Exception:
|
||||
log.warning("token_stats: list_models fehlgeschlagen, Tarife per Name", exc_info=True)
|
||||
return compute_savings(get_stats(), role_map)
|
||||
@@ -0,0 +1,322 @@
|
||||
"""
|
||||
Voice-Endpoints für „Mit Hermes reden" (Browser-Voice + 3D-Avatar).
|
||||
|
||||
Dünner Layer: STT/TTS werden zum Voice-Sidecar (:8650) geproxyt; der Chat geht an den
|
||||
Hermes-`api_server` (:8642, OpenAI-kompatibel) — denselben vollen Agenten mit Tools +
|
||||
geteiltem Mem0 wie CLI/Telegram. Mit stabilem `X-Hermes-Session-Id` hält die Plattform den
|
||||
Transcript server-seitig, daher schickt der Client je Turn nur die neue User-Nachricht.
|
||||
|
||||
LAN-only (kein Token in der 2.0-Phase), wie die übrigen MC2-Endpoints.
|
||||
"""
|
||||
|
||||
import logging
|
||||
import os
|
||||
import time
|
||||
|
||||
import httpx
|
||||
from fastapi import APIRouter, File, Form, HTTPException, UploadFile
|
||||
from fastapi.responses import Response, StreamingResponse
|
||||
from pydantic import BaseModel
|
||||
|
||||
from config import HERMES_API_KEY, HERMES_API_MODEL, HERMES_API_URL, LLAMA_SWAP_URL, VOICE_SERVICE_URL
|
||||
from services import announce
|
||||
from services.voice_metrics import Timer, get_metrics, record_stage # Per-Stage-Latenz (C2)
|
||||
|
||||
# Injection-Schutz (Stufe 0): guard.py liegt im mcp/-Verzeichnis. Per Pfad laden (eigene MC2-Venv).
|
||||
import sys as _sys
|
||||
_GUARD_DIR = os.path.join(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))), "mcp")
|
||||
if _GUARD_DIR not in _sys.path:
|
||||
_sys.path.insert(0, _GUARD_DIR)
|
||||
try:
|
||||
from guard import wrap_untrusted
|
||||
except Exception: # den Voice-Pfad nie wegen des Filters lahmlegen
|
||||
def wrap_untrusted(text: str, label: str = "") -> str:
|
||||
return text
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
router = APIRouter(prefix="/api")
|
||||
|
||||
# Bildschirm-Sicht: das DEDIZIERTE Vision-Modell (Qwen3-VL-8B) beschreibt das Bild; die Beschreibung
|
||||
# geht als TEXT an Hermes -> Lucy behält ihr volles Hirn/Gedächtnis UND nutzt das bessere VL-Modell
|
||||
# (statt der schwächeren Vision der fast-MoE). Per Env abschaltbar/umstellbar.
|
||||
VISION_MODEL = os.environ.get("MC_VISION_MODEL", "vision")
|
||||
# Knappe Beschreibung = schnellere VL-Generierung UND weniger Hermes-Kontext-Bloat (B2).
|
||||
VISION_MAX_TOKENS = int(os.environ.get("MC_VISION_MAX_TOKENS", "280"))
|
||||
|
||||
|
||||
async def _describe_images(image_urls: list[str], hint: str) -> str:
|
||||
"""Lässt das Vision-Modell die Screenshots (1 je Monitor) knapp beschreiben (Deutsch).
|
||||
Mehrere Bilder gehen in EINER Nachricht ans VL-Modell. Leerer String bei Fehler."""
|
||||
multi = len(image_urls) > 1
|
||||
intro = (f"Hier sind {len(image_urls)} Screenshots (je ein Monitor). Beschreibe auf Deutsch in höchstens "
|
||||
"5 kurzen Sätzen das Wesentliche (pro Monitor: App/Fenster, wichtige Inhalte, sichtbarer Text/Code). "
|
||||
"Keine Einleitung, keine Wiederholung der Frage. "
|
||||
if multi else
|
||||
"Beschreibe auf Deutsch in höchstens 5 kurzen Sätzen das Wesentliche auf diesem Screenshot "
|
||||
"(App/Fenster, wichtige Inhalte, sichtbarer Text/Code). Keine Einleitung. ")
|
||||
content: list = [{"type": "text", "text": intro + "Frage des Nutzers dazu: " + hint}]
|
||||
for u in image_urls:
|
||||
content.append({"type": "image_url", "image_url": {"url": u}})
|
||||
try:
|
||||
# 45 s statt 120 s: Qwen3-VL braucht warm ~5 s; wenn es 45 s nicht schafft, ist etwas
|
||||
# kaputt und Lucy soll lieber ohne Bildschirm-Kontext antworten als ewig hängen.
|
||||
async with httpx.AsyncClient(timeout=httpx.Timeout(float(os.environ.get("MC_VISION_TIMEOUT", "45")), connect=5.0)) as client:
|
||||
r = await client.post(f"{LLAMA_SWAP_URL}/v1/chat/completions", json={
|
||||
"model": VISION_MODEL, "max_tokens": VISION_MAX_TOKENS, "stream": False,
|
||||
"messages": [{"role": "user", "content": content}],
|
||||
})
|
||||
r.raise_for_status()
|
||||
return (r.json().get("choices") or [{}])[0].get("message", {}).get("content", "").strip()
|
||||
except Exception as exc:
|
||||
log.warning("Vision-Beschreibung fehlgeschlagen: %s", exc)
|
||||
return ""
|
||||
|
||||
_TIMEOUT = httpx.Timeout(120.0, connect=5.0) # Chatterbox-TTS auf CPU darf dauern
|
||||
|
||||
|
||||
class TTSIn(BaseModel):
|
||||
text: str
|
||||
engine: str = "piper"
|
||||
voice: str = ""
|
||||
language: str = ""
|
||||
ref_path: str = ""
|
||||
|
||||
|
||||
class ChatIn(BaseModel):
|
||||
text: str # die neue User-Äußerung (STT-Ergebnis)
|
||||
session_id: str # stabiler Voice-Faden → server-seitiger Transcript
|
||||
session_key: str = "" # optional: Langzeit-Memory-Scope
|
||||
system: str = "" # optionaler ephemerer System-Prompt (z.B. „antworte knapp/gesprochen")
|
||||
model: str = ""
|
||||
images: list[str] = [] # optionale Bildschirm-Sicht: ein data:-URL je Monitor (Lucys „Augen")
|
||||
|
||||
|
||||
class AnnounceIn(BaseModel):
|
||||
text: str # die Meldung (wird von Lucy gesprochen)
|
||||
subject: str = "" # kurze Betreffzeile (z.B. "[Update]")
|
||||
source: str = "" # Absender (sentry/notify/cron …) — nur fürs Log/Panel
|
||||
priority: str = "normal" # 'silent' = nur im Verlauf zeigen, nicht sprechen
|
||||
|
||||
|
||||
class AlarmIn(BaseModel):
|
||||
text: str # die Alarm-Meldung
|
||||
subject: str = "[Alarm]" # Betreff (Telegram-Präfix)
|
||||
source: str = "alarm" # Absender fürs Log/Panel (z.B. "lucy-watchdog")
|
||||
|
||||
|
||||
@router.post("/alarm")
|
||||
def alarm(body: AlarmIn) -> dict:
|
||||
"""Lucy-UNABHÄNGIGER Alarm-Weg: schickt direkt auf Telegram (und legt die Meldung in den
|
||||
Briefkasten). Für Absender, die NICHT auf die sprechende Lucy zählen können — allen voran
|
||||
der PC-seitige Lucy-Watchdog, wenn die Desktop-App selbst hängt (dann nützt der Briefkasten
|
||||
nichts, weil niemand ihn vorliest → Telegram ist der einzige verlässliche Kanal). LAN-only
|
||||
wie alle MC2-Endpoints."""
|
||||
text = (body.text or "").strip()
|
||||
if not text:
|
||||
raise HTTPException(400, "Leere Meldung.")
|
||||
subject = (body.subject or "[Alarm]").strip()
|
||||
try:
|
||||
item = announce.add(text, subject, body.source or "alarm", "normal")
|
||||
except ValueError as exc:
|
||||
raise HTTPException(400, str(exc))
|
||||
announce.notify_telegram(subject, text) # best-effort Telegram (posix/bash; Windows = No-op)
|
||||
return {"ok": True, "item": item}
|
||||
|
||||
|
||||
@router.post("/voice/announce")
|
||||
def voice_announce(body: AnnounceIn) -> dict:
|
||||
"""Meldung in den Briefkasten legen (Lucy-Proaktivität). Absender: Health-Wächter,
|
||||
notify.sh (Updates/Radar/Telegram-Spiegel), Hermes-cron. LAN-only wie alle MC2-Endpoints."""
|
||||
try:
|
||||
return {"ok": True, "item": announce.add(body.text, body.subject, body.source, body.priority)}
|
||||
except ValueError as exc:
|
||||
raise HTTPException(400, str(exc))
|
||||
|
||||
|
||||
@router.get("/voice/announcements")
|
||||
def voice_announcements(after: int | None = None, limit: int = 20) -> dict:
|
||||
"""Neue Meldungen nach Cursor `after` abholen (Lucy pollt). Ohne `after` nur den
|
||||
aktuellen Cursor-Stand (latest) — Erststart plappert so keine alten Meldungen nach."""
|
||||
return announce.list_after(after, limit)
|
||||
|
||||
|
||||
@router.get("/voice/metrics")
|
||||
def voice_metrics() -> dict:
|
||||
"""Per-Stage-Latenz (STT/Vision/Chat-TTFB/TTS) — rollende Statistik, macht die Voice-Pipeline
|
||||
messbar (C2). Anzeige im Frontend-Overhaul (E)."""
|
||||
return get_metrics()
|
||||
|
||||
|
||||
@router.get("/voice/health")
|
||||
def voice_health() -> dict:
|
||||
"""Erreichbarkeit des Voice-Sidecars + ob der Hermes-API-Key gesetzt ist."""
|
||||
out: dict = {"sidecar": False, "hermes_key": bool(HERMES_API_KEY)}
|
||||
try:
|
||||
r = httpx.get(f"{VOICE_SERVICE_URL}/health", timeout=httpx.Timeout(5.0))
|
||||
out["sidecar"] = r.status_code == 200
|
||||
out["detail"] = r.json() if r.status_code == 200 else None
|
||||
except Exception as exc: # noqa: BLE001
|
||||
out["error"] = str(exc)
|
||||
return out
|
||||
|
||||
|
||||
@router.get("/voice/voices")
|
||||
def voice_voices() -> dict:
|
||||
try:
|
||||
r = httpx.get(f"{VOICE_SERVICE_URL}/voices", timeout=httpx.Timeout(10.0))
|
||||
r.raise_for_status()
|
||||
return r.json()
|
||||
except Exception as exc: # noqa: BLE001
|
||||
raise HTTPException(502, f"Voice-Sidecar nicht erreichbar: {exc}")
|
||||
|
||||
|
||||
@router.post("/voice/stt")
|
||||
async def voice_stt(audio: UploadFile = File(...), language: str = Form(default="")) -> dict:
|
||||
"""Mikro-Audio → Text (Proxy auf Sidecar /stt)."""
|
||||
data = await audio.read()
|
||||
if not data:
|
||||
raise HTTPException(400, "Leeres Audio.")
|
||||
files = {"audio": (audio.filename or "rec.webm", data, audio.content_type or "audio/webm")}
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=_TIMEOUT) as client:
|
||||
with Timer("stt"):
|
||||
r = await client.post(f"{VOICE_SERVICE_URL}/stt", files=files, data={"language": language})
|
||||
r.raise_for_status()
|
||||
return r.json()
|
||||
except httpx.HTTPError as exc:
|
||||
raise HTTPException(502, f"STT fehlgeschlagen: {exc}")
|
||||
|
||||
|
||||
@router.post("/voice/turn")
|
||||
async def voice_turn(audio: UploadFile = File(...)) -> dict:
|
||||
"""Semantische Turn-Detection (Smart Turn v3): war die Äußerung fertig? Proxy → Sidecar."""
|
||||
data = await audio.read()
|
||||
if not data:
|
||||
raise HTTPException(400, "Leeres Audio.")
|
||||
files = {"audio": (audio.filename or "rec.wav", data, audio.content_type or "audio/wav")}
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=httpx.Timeout(10.0, connect=3.0)) as client:
|
||||
with Timer("turn"):
|
||||
r = await client.post(f"{VOICE_SERVICE_URL}/turn", files=files)
|
||||
r.raise_for_status()
|
||||
return r.json()
|
||||
except httpx.HTTPError as exc:
|
||||
# Turn-Check ist eine Optimierung — bei Ausfall lieber sofort antworten als hängen.
|
||||
log.warning("Turn-Check fehlgeschlagen: %s", exc)
|
||||
return {"complete": True, "probability": 1.0, "engine": "fallback"}
|
||||
|
||||
|
||||
@router.post("/voice/reference")
|
||||
async def voice_set_reference(audio: UploadFile = File(...)) -> dict:
|
||||
"""Klon-Referenz (z.B. ElevenLabs-Erzeugnis) hochladen → Chatterbox nutzt sie. Proxy → Sidecar."""
|
||||
data = await audio.read()
|
||||
if not data:
|
||||
raise HTTPException(400, "Leeres Audio.")
|
||||
files = {"audio": (audio.filename or "ref.wav", data, audio.content_type or "audio/mpeg")}
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=_TIMEOUT) as client:
|
||||
r = await client.post(f"{VOICE_SERVICE_URL}/reference", files=files)
|
||||
r.raise_for_status()
|
||||
return r.json()
|
||||
except httpx.HTTPError as exc:
|
||||
raise HTTPException(502, f"Referenz-Upload fehlgeschlagen: {exc}")
|
||||
|
||||
|
||||
@router.get("/voice/reference")
|
||||
def voice_get_reference() -> dict:
|
||||
try:
|
||||
r = httpx.get(f"{VOICE_SERVICE_URL}/reference", timeout=httpx.Timeout(8.0))
|
||||
r.raise_for_status()
|
||||
return r.json()
|
||||
except Exception as exc: # noqa: BLE001
|
||||
return {"active": False, "error": str(exc)}
|
||||
|
||||
|
||||
@router.delete("/voice/reference")
|
||||
def voice_clear_reference() -> dict:
|
||||
try:
|
||||
r = httpx.delete(f"{VOICE_SERVICE_URL}/reference", timeout=httpx.Timeout(8.0))
|
||||
r.raise_for_status()
|
||||
return r.json()
|
||||
except httpx.HTTPError as exc:
|
||||
raise HTTPException(502, f"Löschen fehlgeschlagen: {exc}")
|
||||
|
||||
|
||||
@router.post("/voice/tts")
|
||||
async def voice_tts(body: TTSIn) -> Response:
|
||||
"""Text → Sprache (Proxy auf Sidecar /tts), liefert WAV-Bytes."""
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=_TIMEOUT) as client:
|
||||
with Timer("tts"):
|
||||
r = await client.post(f"{VOICE_SERVICE_URL}/tts", json=body.model_dump())
|
||||
r.raise_for_status()
|
||||
return Response(content=r.content, media_type=r.headers.get("content-type", "audio/wav"))
|
||||
except httpx.HTTPError as exc:
|
||||
raise HTTPException(502, f"TTS fehlgeschlagen: {exc}")
|
||||
|
||||
|
||||
@router.post("/voice/chat")
|
||||
async def voice_chat(body: ChatIn) -> StreamingResponse:
|
||||
"""Neue User-Äußerung → Hermes-Agent (api_server, streamend). SSE wird 1:1 durchgereicht.
|
||||
|
||||
Mit `X-Hermes-Session-Id` hält die Plattform den Verlauf — wir senden nur die neue Nachricht.
|
||||
Auth per Bearer (API_SERVER_KEY); ohne Key liefert :8642 ein 401."""
|
||||
if not HERMES_API_KEY:
|
||||
raise HTTPException(503, "HERMES_API_KEY/API_SERVER_KEY nicht gesetzt — Agent-Auth fehlt.")
|
||||
|
||||
headers = {
|
||||
"Authorization": f"Bearer {HERMES_API_KEY}",
|
||||
"X-Hermes-Session-Id": body.session_id,
|
||||
}
|
||||
if body.session_key:
|
||||
headers["X-Hermes-Session-Key"] = body.session_key
|
||||
|
||||
async def gen():
|
||||
t0 = time.perf_counter()
|
||||
first = True
|
||||
first_content = True
|
||||
# Bildschirm-Sicht INNERHALB des Streams (C2-Fix): so startet die SSE-Antwort sofort und
|
||||
# der Client bekommt ein Progress-Event (-> Lucy kann eine Warte-Ansage sprechen), statt
|
||||
# dass der Request bis zu 120 s "tot" hängt, während das Vision-Modell beschreibt.
|
||||
user_text = body.text
|
||||
imgs = [u for u in (body.images or []) if u]
|
||||
if imgs:
|
||||
yield b'event: hermes.vision.progress\ndata: {"note": "Bildschirm wird angeschaut"}\n\n'
|
||||
with Timer("vision"):
|
||||
desc = await _describe_images(imgs, body.text)
|
||||
if desc:
|
||||
safe_desc = wrap_untrusted(desc, "BILDSCHIRM")
|
||||
user_text = f"[Bildschirm-Sicht — das ist gerade auf dem/den Schirm(en) zu sehen:\n{safe_desc}\n]\n\n{body.text}"
|
||||
messages = []
|
||||
if body.system:
|
||||
messages.append({"role": "system", "content": body.system})
|
||||
messages.append({"role": "user", "content": user_text})
|
||||
payload = {"model": body.model or HERMES_API_MODEL, "messages": messages, "stream": True}
|
||||
# Lucys Hirn (Qwen3.6) ist ein Thinking-Modell -> für die gesprochene Assistentin Thinking AUS,
|
||||
# sonst generiert es tausende Reasoning-Token VOR der kurzen Antwort (gemessen: 11k Token, ~30s TTFB).
|
||||
# Gleiches Muster wie die fast-Spur im Gateway (gateway_proxy.py) und die Mem0-Extraktion.
|
||||
if os.environ.get("MC_VOICE_NO_THINK", "1") not in ("0", "false", "False"):
|
||||
payload["chat_template_kwargs"] = {"enable_thinking": False}
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=httpx.Timeout(None, connect=5.0)) as client:
|
||||
async with client.stream(
|
||||
"POST", f"{HERMES_API_URL}/v1/chat/completions", json=payload, headers=headers,
|
||||
) as r:
|
||||
if r.status_code != 200:
|
||||
detail = (await r.aread()).decode("utf-8", "replace")[:500]
|
||||
yield f"data: {{\"error\": \"Hermes {r.status_code}: {detail}\"}}\n\n".encode()
|
||||
return
|
||||
async for chunk in r.aiter_raw():
|
||||
if first: # Time-To-First-Byte des Hermes-Streams (Verbindungs-Overhead)
|
||||
record_stage("chat_ttfb", (time.perf_counter() - t0) * 1000.0)
|
||||
first = False
|
||||
# Erster CONTENT-Delta = echte Hirn-Latenz (Agent-Overhead + LLM-TTFT) —
|
||||
# chat_ttfb misst nur den SSE-Start (~5 ms) und ist dafür blind.
|
||||
if first_content and b'"content"' in chunk:
|
||||
record_stage("chat_first_content", (time.perf_counter() - t0) * 1000.0)
|
||||
first_content = False
|
||||
yield chunk
|
||||
except httpx.HTTPError as exc:
|
||||
yield f"data: {{\"error\": \"Verbindung zu Hermes fehlgeschlagen: {exc}\"}}\n\n".encode()
|
||||
|
||||
return StreamingResponse(gen(), media_type="text/event-stream")
|
||||
@@ -0,0 +1,254 @@
|
||||
"""
|
||||
Hermes-Agent-Status (Control-Plane-Read). MC betreibt Hermes NICHT — es zeigt nur
|
||||
Status + verlinkt das standalone hermes-webui. Voller Zugriff + Tools/MCP werden in
|
||||
Hermes' eigener Config verdrahtet (siehe docs/HERMES_SETUP.md).
|
||||
"""
|
||||
|
||||
import logging
|
||||
import os
|
||||
import re
|
||||
|
||||
import httpx
|
||||
import psutil
|
||||
|
||||
from config import HERMES_TERMINAL_URL, HERMES_API_URL, HERMES_HOME, PC_EXECUTOR_URL
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
|
||||
def _hermes_version(name: str) -> float | None:
|
||||
"""Versionszahl aus 'Hermes-4.3', 'Hermes-4', 'Nous-Hermes-2' → 4.3/4.0/2.0."""
|
||||
low = (name or "").lower()
|
||||
if "hermes" not in low:
|
||||
return None
|
||||
m = re.search(r"hermes[-_ ]?(\d+(?:\.\d+)?)", low)
|
||||
return float(m.group(1)) if m else None
|
||||
|
||||
|
||||
def _active_brain_name() -> str:
|
||||
"""Aktives Agent-Hirn aus Hermes' Config: model.default (sonst model.model)."""
|
||||
try:
|
||||
from ruamel.yaml import YAML
|
||||
p = HERMES_HOME / "config.yaml"
|
||||
if p.exists():
|
||||
with p.open(encoding="utf-8") as f:
|
||||
cfg = YAML().load(f) or {}
|
||||
m = (cfg.get("model") or {}) if isinstance(cfg, dict) else {}
|
||||
return str(m.get("default") or m.get("model") or "auto")
|
||||
except Exception:
|
||||
log.debug("_active_brain_name: Lesefehler", exc_info=True)
|
||||
return "auto"
|
||||
|
||||
|
||||
def hermes_brain_info() -> dict:
|
||||
"""Aktuelles Agent-Hirn = Modell/Alias, das Hermes laut Config nutzt (model.default),
|
||||
plus Budget-Check. Zeigt das REAL genutzte Hirn — unabhängig von einer 'hermes'-Rolle."""
|
||||
from services import llamaswap
|
||||
|
||||
models = llamaswap.list_models()
|
||||
brain = _active_brain_name() # z.B. "fast" (Alias) oder ein Modellname
|
||||
bl = brain.lower()
|
||||
cur = next((m for m in models if (m.get("role") or "").lower() == bl), None) \
|
||||
or next((m for m in models if bl in (m["name"] or "").lower()), None)
|
||||
cur_params = (cur.get("capabilities") or {}).get("params_b") if cur else None
|
||||
current = None
|
||||
if cur:
|
||||
current = {"name": cur["name"], "alias": brain, "filename": cur.get("filename"),
|
||||
"params_b": cur_params, "quant": cur.get("quant"),
|
||||
"size_bytes": cur.get("size_bytes"),
|
||||
"gguf_path": cur.get("gguf_path"), "incomplete": cur.get("incomplete")}
|
||||
|
||||
# Fit-Check: passt das (immer warme) Hirn + das größte on-demand-Modell zusammen ins Budget?
|
||||
budget = None
|
||||
try:
|
||||
from services.budget import footprint_gb, gtt_budget_gb
|
||||
groups = llamaswap.list_groups()
|
||||
persist = set()
|
||||
for g in groups.values():
|
||||
if isinstance(g, dict) and (g.get("persist") or g.get("persistent")):
|
||||
persist.update(g.get("members") or [])
|
||||
|
||||
cur_name = cur["name"] if cur else None
|
||||
brain_gb = footprint_gb(cur) if cur else 0.0
|
||||
# voller Always-Warm-Footprint (alle persist, Brain=Empfehlung) — nur Info
|
||||
warm = brain_gb + sum(footprint_gb(m) for m in models
|
||||
if m["name"] in persist and m["name"] != cur_name)
|
||||
largest_od = max((footprint_gb(m) for m in models if m["name"] not in persist), default=0.0)
|
||||
gtt = gtt_budget_gb()
|
||||
# Seit dem `persistent`-Fix (03.07.) bleibt das GANZE Warmset (Hirn+embed+vision)
|
||||
# resident, wenn ein on-demand-Modell DANEBEN lädt → der reale Peak ist Warmset +
|
||||
# größtes on-demand, nicht nur Hirn + größtes. Genau daran wird `fits` gemessen.
|
||||
budget = {
|
||||
"gtt_gb": gtt,
|
||||
"brain_gb": round(brain_gb, 1),
|
||||
"warm_projected_gb": round(warm, 1),
|
||||
"largest_ondemand_gb": round(largest_od, 1),
|
||||
"fits": (warm + largest_od) <= gtt,
|
||||
"free_after_gb": round(gtt - warm - largest_od, 1),
|
||||
}
|
||||
except Exception:
|
||||
log.debug("hermes_brain_info: Budget-Berechnung fehlgeschlagen", exc_info=True)
|
||||
|
||||
return {"current": current, "recommended": None, "update_available": False, "budget": budget}
|
||||
|
||||
|
||||
def _reach(url: str, path: str = "") -> bool:
|
||||
try:
|
||||
with httpx.Client(timeout=3.0) as c:
|
||||
return c.get(f"{url}{path}").status_code < 500
|
||||
except httpx.HTTPError:
|
||||
return False
|
||||
|
||||
|
||||
def _count_enabled_mcp_servers() -> int:
|
||||
config_path = HERMES_HOME / "config.yaml"
|
||||
if not config_path.exists():
|
||||
return 0
|
||||
try:
|
||||
from ruamel.yaml import YAML
|
||||
r_yaml = YAML()
|
||||
with config_path.open("r", encoding="utf-8") as f:
|
||||
cfg = r_yaml.load(f) or {}
|
||||
mcp_servers = cfg.get("mcp_servers", {}) if isinstance(cfg, dict) else {}
|
||||
if not isinstance(mcp_servers, dict):
|
||||
return 0
|
||||
return sum(1 for v in mcp_servers.values() if isinstance(v, dict) and v.get("enabled", True))
|
||||
except Exception:
|
||||
log.debug("_count_enabled_mcp_servers: Fehler", exc_info=True)
|
||||
return 0
|
||||
|
||||
|
||||
def agent_status() -> dict:
|
||||
"""Erreichbarkeit von Gateway (:8642) + WebUI (:8787) + lokale Hinweise."""
|
||||
home = HERMES_HOME
|
||||
brain_model = "auto"
|
||||
config_path = home / "config.yaml"
|
||||
if config_path.exists():
|
||||
try:
|
||||
from ruamel.yaml import YAML
|
||||
r_yaml = YAML()
|
||||
with config_path.open("r", encoding="utf-8") as f:
|
||||
cfg = r_yaml.load(f) or {}
|
||||
if isinstance(cfg, dict):
|
||||
# Hermes nutzt model.default als aktives Modell (model.model = Provider-Param).
|
||||
m = cfg.get("model", {}) or {}
|
||||
brain_model = m.get("default") or m.get("model") or "auto"
|
||||
except Exception:
|
||||
log.debug("agent_status: Hermes-config.yaml nicht lesbar", exc_info=True)
|
||||
|
||||
|
||||
return {
|
||||
"gateway_url": HERMES_API_URL,
|
||||
# Interaktives Web-Terminal (ttyd → `hermes chat`), eingebettet in MC2.
|
||||
"terminal_url": HERMES_TERMINAL_URL,
|
||||
"gateway_reachable": _reach(HERMES_API_URL, "/health"),
|
||||
"terminal_reachable": _reach(HERMES_TERMINAL_URL, "/"),
|
||||
"home_exists": home.exists(),
|
||||
"brain_model": brain_model,
|
||||
# Best-effort: welche Verdrahtung lokal sichtbar ist (auf der Box aussagekräftig).
|
||||
"has_config": (home / "config.yaml").exists() or (home / "config.json").exists(),
|
||||
"has_skills": (home / "skills").exists(),
|
||||
"has_memories": (home / "memories").exists(),
|
||||
# Neue Felder: Telegram, MCP-Server-Anzahl, PC-Executor-Erreichbarkeit.
|
||||
"telegram_enabled": bool(os.environ.get("TELEGRAM_BOT_TOKEN", "")),
|
||||
"mcp_server_count": _count_enabled_mcp_servers(),
|
||||
"pc_executor_reachable": _reach(PC_EXECUTOR_URL, "/health"),
|
||||
}
|
||||
|
||||
|
||||
def set_agent_brain(model_id: str) -> dict:
|
||||
"""Setzt ein (bereits installiertes) Modell als Agent-Hirn — WARM-bewusst:
|
||||
1) vergibt den 'hermes'-Alias (das Agent-Hirn-Slot),
|
||||
2) tauscht es in die residente brains-Gruppe (altes Hirn raus, fast/vision bleiben),
|
||||
3) zeigt die Hermes-Config auf den 'hermes'-Alias + Gateway-Restart.
|
||||
So bleibt das neue Hirn warm und der Agent nutzt es sofort."""
|
||||
from services import llamaswap
|
||||
models = {m["name"]: m for m in llamaswap.list_models()}
|
||||
if model_id not in models:
|
||||
return {"ok": False, "reason": "Modell nicht installiert — erst über Modelle-finden laden."}
|
||||
old = next((m["name"] for m in models.values() if m.get("role") == "hermes"), None)
|
||||
if model_id == old:
|
||||
# Idempotent härten: auch wenn schon Hirn, warm (brains) + ttl 0 sicherstellen.
|
||||
try:
|
||||
from services.llamaswap import set_ttl
|
||||
brains = (llamaswap.list_groups().get("brains") or {}).get("members") or []
|
||||
if model_id not in brains:
|
||||
llamaswap.set_group("brains", brains + [model_id], swap=False, persist=True)
|
||||
set_ttl(model_id, 0)
|
||||
except PermissionError as exc:
|
||||
return {"ok": False, "reason": str(exc)}
|
||||
return {"ok": True, "old": old, "new": model_id, "note": "ist bereits das Agent-Hirn"}
|
||||
try:
|
||||
llamaswap.set_role(model_id, "hermes") # 1) Alias
|
||||
brains = (llamaswap.list_groups().get("brains") or {}).get("members") or []
|
||||
new_members = [x for x in brains if x not in (old, model_id)] + [model_id]
|
||||
llamaswap.set_group("brains", new_members, swap=False, persist=True) # 2) warm
|
||||
# 2b) TTL härten: neues Hirn nie auto-entladen; altes Hirn auf Default entspannen.
|
||||
from services.llamaswap import set_ttl, DEFAULT_TTL
|
||||
set_ttl(model_id, 0)
|
||||
if old:
|
||||
set_ttl(old, DEFAULT_TTL)
|
||||
except PermissionError as exc:
|
||||
return {"ok": False, "reason": str(exc)}
|
||||
update_brain_model("hermes") # 3) Config + Restart
|
||||
# Weiche Budget-Warnung (kein Hard-Block): passt Hirn + größtes on-demand zusammen ins GTT?
|
||||
warning = None
|
||||
try:
|
||||
b = hermes_brain_info().get("budget") or {}
|
||||
if b and not b.get("fits", True):
|
||||
warning = (f"Speicher-Warnung: Hirn (~{b.get('brain_gb')} GB) + größtes on-demand-"
|
||||
f"Modell (~{b.get('largest_ondemand_gb')} GB) übersteigen das GTT-Budget "
|
||||
f"(~{b.get('gtt_gb')} GB) — das Hirn ist persistent, heavy/coder laden "
|
||||
f"DANEBEN: Überlauf droht (Lade-Crash/Swapping statt Verdrängung).")
|
||||
except Exception:
|
||||
log.debug("set_agent_brain: Budget-Check fehlgeschlagen", exc_info=True)
|
||||
return {"ok": True, "old": old, "new": model_id, "warning": warning}
|
||||
|
||||
|
||||
def update_brain_model(new_model: str) -> bool:
|
||||
from config import HERMES_HOME
|
||||
home = HERMES_HOME
|
||||
config_path = home / "config.yaml"
|
||||
|
||||
# Ensure home directory exists
|
||||
home.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
cfg = {}
|
||||
if config_path.exists():
|
||||
try:
|
||||
from ruamel.yaml import YAML
|
||||
r_yaml = YAML()
|
||||
with config_path.open("r", encoding="utf-8") as f:
|
||||
cfg = r_yaml.load(f) or {}
|
||||
except Exception:
|
||||
log.debug("update_brain_model: bestehende config.yaml nicht lesbar", exc_info=True)
|
||||
cfg = {}
|
||||
|
||||
if not isinstance(cfg, dict):
|
||||
cfg = {}
|
||||
|
||||
if "model" not in cfg or not isinstance(cfg["model"], dict):
|
||||
cfg["model"] = {}
|
||||
|
||||
# Hermes liest model.default als aktives Modell; model.model ist der Provider-Param.
|
||||
# Beide setzen, sonst greift die Umschaltung nicht (latenter Bug: nur model.model gesetzt).
|
||||
cfg["model"]["default"] = new_model
|
||||
cfg["model"]["model"] = new_model
|
||||
|
||||
try:
|
||||
from ruamel.yaml import YAML
|
||||
r_yaml = YAML()
|
||||
with config_path.open("w", encoding="utf-8") as f:
|
||||
r_yaml.dump(cfg, f)
|
||||
|
||||
# Restart the user-space service to apply changes
|
||||
try:
|
||||
import services.maintenance as maintenance
|
||||
maintenance.restart_service("hermes-gateway")
|
||||
except Exception:
|
||||
log.warning("update_brain_model: hermes-gateway-Restart fehlgeschlagen", exc_info=True)
|
||||
|
||||
return True
|
||||
except Exception:
|
||||
log.warning("update_brain_model: Schreiben der config.yaml fehlgeschlagen", exc_info=True)
|
||||
return False
|
||||
@@ -0,0 +1,99 @@
|
||||
"""
|
||||
Melde-Briefkasten der Box (Lucy-Proaktivität, Faden A3).
|
||||
|
||||
Alles, was die Box dem Commander aktiv sagen will (Health-Wächter, Auto-Updates,
|
||||
Radar, Hermes-cron via notify.sh), landet als Eintrag hier. Die Lucy-Desktop-App
|
||||
pollt `/api/voice/announcements` und SPRICHT neue Einträge von sich aus.
|
||||
|
||||
Persistenz als JSON neben den Modellen (übersteht Deploys/Neustarts, wie der
|
||||
Discover-Cache). Bewusst klein: fortlaufende IDs als Cursor, Ring der letzten
|
||||
MAX_ITEMS Einträge, ein Lock für die FastAPI-Threadpool-Worker.
|
||||
"""
|
||||
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
import shutil
|
||||
import subprocess
|
||||
import threading
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
from config import MODELS_DIR
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
NOTIFY_SH = str(Path(__file__).resolve().parent.parent.parent / "deploy" / "notify.sh")
|
||||
|
||||
STORE_PATH = Path(os.environ.get("MC_ANNOUNCE_STORE", str(MODELS_DIR / "mc2-announce.json")))
|
||||
MAX_ITEMS = int(os.environ.get("MC_ANNOUNCE_MAX", "200"))
|
||||
|
||||
_lock = threading.Lock()
|
||||
_state: dict | None = None # {"next_id": int, "items": [...]}
|
||||
|
||||
|
||||
def _load() -> dict:
|
||||
global _state
|
||||
if _state is None:
|
||||
try:
|
||||
_state = json.loads(STORE_PATH.read_text(encoding="utf-8"))
|
||||
assert isinstance(_state.get("next_id"), int) and isinstance(_state.get("items"), list)
|
||||
except Exception:
|
||||
_state = {"next_id": 1, "items": []}
|
||||
return _state
|
||||
|
||||
|
||||
def _save(state: dict) -> None:
|
||||
try:
|
||||
tmp = STORE_PATH.with_suffix(".tmp")
|
||||
tmp.write_text(json.dumps(state, ensure_ascii=False), encoding="utf-8")
|
||||
tmp.replace(STORE_PATH)
|
||||
except OSError:
|
||||
# Briefkasten darf den Absender nie blockieren — dann eben nur in-memory.
|
||||
log.warning("announce: Store %s nicht schreibbar", STORE_PATH, exc_info=True)
|
||||
|
||||
|
||||
def add(text: str, subject: str = "", source: str = "", priority: str = "normal") -> dict:
|
||||
"""Eintrag anhängen. priority: 'normal' (sprechen) | 'silent' (nur Verlauf/Panel)."""
|
||||
text = (text or "").strip()
|
||||
if not text:
|
||||
raise ValueError("Leere Meldung.")
|
||||
with _lock:
|
||||
state = _load()
|
||||
item = {
|
||||
"id": state["next_id"],
|
||||
"ts": time.time(),
|
||||
"subject": (subject or "").strip()[:120],
|
||||
"text": text[:4000],
|
||||
"source": (source or "").strip()[:60],
|
||||
"priority": priority if priority in ("normal", "silent") else "normal",
|
||||
}
|
||||
state["next_id"] += 1
|
||||
state["items"].append(item)
|
||||
del state["items"][:-MAX_ITEMS]
|
||||
_save(state)
|
||||
log.info("announce #%s [%s] %s: %.80s", item["id"], item["source"] or "-", item["subject"] or "-", text)
|
||||
return item
|
||||
|
||||
|
||||
def notify_telegram(subject: str, text: str) -> None:
|
||||
"""Best-effort auch auf Telegram (User ist evtl. nicht am PC). MC_NOTIFY_NO_ANNOUNCE=1
|
||||
verhindert, dass notify.sh die Meldung ZURÜCK in den Briefkasten spiegelt — der Absender
|
||||
(Wächter/Erinnerung) hat sie dort schon selbst abgelegt. Auf Windows (Dev) ein No-op."""
|
||||
if not (os.name == "posix" and shutil.which("bash")):
|
||||
return
|
||||
try:
|
||||
subprocess.run(["bash", NOTIFY_SH, "-s", subject, text],
|
||||
timeout=30, capture_output=True, env={**os.environ, "MC_NOTIFY_NO_ANNOUNCE": "1"})
|
||||
except Exception:
|
||||
log.warning("notify_telegram: notify.sh fehlgeschlagen", exc_info=True)
|
||||
|
||||
|
||||
def list_after(after: int | None, limit: int = 20) -> dict:
|
||||
"""Einträge NACH Cursor `after` (aufsteigend). Ohne Cursor nur den aktuellen
|
||||
Stand liefern (latest) — so initialisiert Lucy ihren Cursor, ohne Altes nachzuplappern."""
|
||||
with _lock:
|
||||
state = _load()
|
||||
latest = state["next_id"] - 1
|
||||
items = [] if after is None else [i for i in state["items"] if i["id"] > after][: max(1, min(limit, 100))]
|
||||
return {"latest": latest, "items": items}
|
||||
@@ -0,0 +1,73 @@
|
||||
"""
|
||||
Voll-Zustands-Backup (mem0 + Hermes-Configs/Secrets + llama-swap config).
|
||||
Delegiert an deploy/backup.sh (eine Quelle der Wahrheit, identisch zum systemd-Timer);
|
||||
Restore läuft bewusst nur per CLI (deploy/restore.sh) — siehe docs/BACKUP.md.
|
||||
"""
|
||||
|
||||
import subprocess
|
||||
import tarfile
|
||||
from pathlib import Path
|
||||
|
||||
from config import MODELS_DIR
|
||||
|
||||
BACKUP_DIR = Path(MODELS_DIR) / "mc2-backups"
|
||||
SRC_ROOT = Path(__file__).resolve().parents[2]
|
||||
BACKUP_SH = SRC_ROOT / "deploy" / "backup.sh"
|
||||
|
||||
|
||||
def _ts(p: Path) -> str:
|
||||
"""Zeitstempel aus 'mc2-state-<ts>.tar.gz' (Path.stem ließe '.tar' stehen)."""
|
||||
return p.name[len("mc2-state-"):-len(".tar.gz")]
|
||||
|
||||
|
||||
def _latest() -> Path | None:
|
||||
if not BACKUP_DIR.exists():
|
||||
return None
|
||||
snaps = sorted(BACKUP_DIR.glob("mc2-state-*.tar.gz"), reverse=True)
|
||||
return snaps[0] if snaps else None
|
||||
|
||||
|
||||
def _components(tarball: Path) -> list[str]:
|
||||
"""Top-Level-Einträge im Tarball (zur Anzeige im UI)."""
|
||||
try:
|
||||
with tarfile.open(tarball, "r:gz") as t:
|
||||
top = {m.name.split("/")[1] for m in t.getmembers()
|
||||
if m.name.startswith("./") and "/" in m.name[2:]}
|
||||
top |= {m.name[2:] for m in t.getmembers()
|
||||
if m.name.startswith("./") and "/" not in m.name[2:] and m.isfile()}
|
||||
return sorted(x for x in top if x)
|
||||
except Exception:
|
||||
return []
|
||||
|
||||
|
||||
def backup_now() -> dict:
|
||||
"""Erstellt einen Voll-Zustands-Snapshot via deploy/backup.sh."""
|
||||
try:
|
||||
r = subprocess.run(["/bin/bash", str(BACKUP_SH)], capture_output=True, text=True, timeout=180)
|
||||
if r.returncode != 0:
|
||||
return {"ok": False, "snapshot": "", "files": [], "error": (r.stderr or r.stdout).strip()[-300:]}
|
||||
except Exception as exc: # noqa: BLE001
|
||||
return {"ok": False, "snapshot": "", "files": [], "error": str(exc)}
|
||||
|
||||
latest = _latest()
|
||||
if not latest:
|
||||
return {"ok": False, "snapshot": "", "files": [], "error": "Kein Backup erzeugt"}
|
||||
return {
|
||||
"ok": True,
|
||||
"snapshot": _ts(latest),
|
||||
"files": _components(latest),
|
||||
"size_mb": round(latest.stat().st_size / 1_000_000, 2),
|
||||
}
|
||||
|
||||
|
||||
def list_backups() -> list[dict]:
|
||||
if not BACKUP_DIR.exists():
|
||||
return []
|
||||
out = []
|
||||
for p in sorted(BACKUP_DIR.glob("mc2-state-*.tar.gz"), reverse=True):
|
||||
out.append({
|
||||
"snapshot": _ts(p),
|
||||
"file": p.name,
|
||||
"size_mb": round(p.stat().st_size / 1_000_000, 2),
|
||||
})
|
||||
return out
|
||||
@@ -0,0 +1,205 @@
|
||||
"""
|
||||
Speicher-Budget & SETUP-BEWUSSTE ctx-Vergabe — EINE Quelle der Wahrheit.
|
||||
|
||||
Modelliert die auf der Box VERIFIZIERTE Residenz-Realität (llama-swap, GTT ~124 GB):
|
||||
• Die `persistent`-Gruppe (brains = Hirn+embed+vision) bleibt IMMER resident —
|
||||
seit dem Key-Fix 03.07.2026 greift der Schutz wirklich (vorher stand `persist`
|
||||
in der Config, das llama-swap stillschweigend ignorierte; on-demand-Last
|
||||
verdrängte damals die ganze Gruppe).
|
||||
• Ein on-demand-Modell (heavy/coder/…) lädt NEBEN die brains-Gruppe und muss
|
||||
deren Footprint mit einplanen.
|
||||
Daraus folgt, wie viel Speicher NEBEN einem Zielmodell reserviert bleiben muss —
|
||||
und damit der größte Kontext, der wirklich passt (nicht nur für das Modell allein).
|
||||
|
||||
Vorher rechnete nur der Hirn-Wechsel (agent.py) setup-bewusst; die allgemeine
|
||||
ctx-Vergabe nahm den Gesamt-RAM in Isolation. Dieses Modul vereint beides.
|
||||
"""
|
||||
|
||||
import re
|
||||
|
||||
import psutil
|
||||
|
||||
from services.fit import (
|
||||
QUANT_BYTES_PER_PARAM,
|
||||
estimate_memory_gb,
|
||||
extract_params_b,
|
||||
max_ctx_in_budget,
|
||||
)
|
||||
|
||||
HEADROOM_GB = 4.0 # OS/Treiber/Fragmentierung
|
||||
|
||||
|
||||
def gtt_budget_gb() -> float:
|
||||
"""GPU-adressierbarer Speicher (GTT) in GB — die harte Obergrenze. Liest
|
||||
amdgpu.gttsize aus /proc/cmdline, sonst RAM minus OS-Reserve."""
|
||||
try:
|
||||
with open("/proc/cmdline") as f:
|
||||
m = re.search(r"amdgpu\.gttsize=(\d+)", f.read())
|
||||
if m:
|
||||
return round(int(m.group(1)) / 1024.0, 1)
|
||||
except Exception:
|
||||
pass
|
||||
return round(psutil.virtual_memory().total / (1024 ** 3) - 6.0, 1)
|
||||
|
||||
|
||||
def params_of_model(model: dict) -> float:
|
||||
"""Robuste Params (Mrd.) eines INSTALLIERTEN Modells: MAXIMUM aus Caps-Schätzung und
|
||||
Dateigröße. Deckt 'Coder-Next' ohne Größe im Namen (→ aus Datei) und Split-GGUFs
|
||||
(size_bytes = nur erster Teil → ignoriert) ab."""
|
||||
caps = model.get("capabilities") or {}
|
||||
quant = model.get("quant") or "Q4_K_M"
|
||||
bpp = QUANT_BYTES_PER_PARAM.get(quant.upper(), 0.55)
|
||||
size_gb = (model.get("size_bytes") or 0) / (1024 ** 3)
|
||||
pb_size = (size_gb / bpp) if size_gb > 1.0 else 0.0
|
||||
return max(float(caps.get("params_b") or 0), pb_size, 7.0)
|
||||
|
||||
|
||||
_CTK_RE = re.compile(r"(?:--cache-type-k|(?<![\w-])-ctk)\s+(\S+)")
|
||||
_CTV_RE = re.compile(r"(?:--cache-type-v|(?<![\w-])-ctv)\s+(\S+)")
|
||||
|
||||
|
||||
def _cache_types(cmd: str) -> tuple[str | None, str | None]:
|
||||
"""K/V-Cache-Quantisierung aus dem llama-server-Cmd (Default f16 → None)."""
|
||||
ck = m.group(1) if (m := _CTK_RE.search(cmd or "")) else None
|
||||
cv = m.group(1) if (m := _CTV_RE.search(cmd or "")) else None
|
||||
return ck, cv
|
||||
|
||||
|
||||
def _real_kv_gb(model: dict, ctx: int) -> float | None:
|
||||
"""ECHTE KV-Cache-Größe (GiB) aus den GGUF-Architektur-Metadaten (Layer × KV-Heads ×
|
||||
Head-Dim) + der cache-type-Quantisierung des Cmds. None, wenn das GGUF nicht lesbar ist
|
||||
→ Aufrufer fällt auf die params-basierte Heuristik zurück."""
|
||||
from services import gguf_meta
|
||||
path = model.get("gguf_path")
|
||||
if not path:
|
||||
return None
|
||||
meta = gguf_meta.arch_meta(path)
|
||||
if not meta:
|
||||
return None
|
||||
ck, cv = _cache_types(model.get("cmd") or "")
|
||||
return gguf_meta.kv_cache_gb(meta, ctx, ck, cv)
|
||||
|
||||
|
||||
def footprint_gb(model: dict) -> float:
|
||||
"""Loaded-Footprint eines Modells = Gewichte + KV-Cache (bei seinem aktuellen ctx).
|
||||
KV kommt aus den ECHTEN Architektur-Metadaten des GGUF (nicht mehr params-geschätzt) —
|
||||
entscheidend bei MoE (A3B): die alte Schätzung hing an den Gesamt-Params und überschätzte
|
||||
grob (z.B. „68 GB reserviert" statt real ~25 GB). Heuristik bleibt Fallback."""
|
||||
quant = model.get("quant") or "Q4_K_M"
|
||||
ctx = int(model.get("ctx") or 32768)
|
||||
bpp = QUANT_BYTES_PER_PARAM.get(quant.upper(), 0.55)
|
||||
size_gb = (model.get("size_bytes") or 0) / (1024 ** 3)
|
||||
pb = params_of_model(model)
|
||||
weights = max(pb * bpp, size_gb)
|
||||
kv = _real_kv_gb(model, ctx)
|
||||
if kv is None:
|
||||
kv = estimate_memory_gb(pb, quant, ctx) - pb * bpp
|
||||
return weights + max(kv, 0.0)
|
||||
|
||||
|
||||
def params_b_for(name: str) -> float:
|
||||
"""Parameter (Mrd.) für einen Modell-/Repo-Namen: KATALOG (echte Metadaten) zuerst,
|
||||
sonst Namens-Schätzung. Gemeinsam für Fit-Vorschau und ctx-Vergabe."""
|
||||
from services import catalog
|
||||
meta = catalog.meta_for_name(name) if name else None
|
||||
if meta and meta.get("total_params_b"):
|
||||
return float(meta["total_params_b"])
|
||||
return extract_params_b(name)
|
||||
|
||||
|
||||
def _coresident_members(groups: dict) -> set:
|
||||
"""Modelle, die GLEICHZEITIG warm sind: Mitglieder aller `swap:false`-Gruppen
|
||||
(Ko-Residenz, z.B. brains = Hirn+embed+vision). Seit dem `persistent`-Fix (03.07.2026)
|
||||
überlebt die Gruppe auch on-demand-Last: Coder/heavy laden DANEBEN, nicht an ihre
|
||||
Stelle (live verifiziert: Coder + Qwen3.6 gleichzeitig `ready`). Die frühere
|
||||
Beobachtung „heavy verdrängt die Gruppe" war der ignorierte `persist`-Key."""
|
||||
out: set = set()
|
||||
for g in (groups or {}).values():
|
||||
if isinstance(g, dict) and g.get("swap") is False:
|
||||
out.update(g.get("members") or [])
|
||||
return out
|
||||
|
||||
|
||||
def reserved_gb(role: str | None) -> dict:
|
||||
"""Speicher, der NEBEN einem Zielmodell der gegebenen Rolle resident bleibt — gemäß der
|
||||
seit dem `persistent`-Fix (03.07.2026) geltenden Semantik: die ko-residente
|
||||
`swap:false`-Gruppe (brains) bleibt IMMER geladen, on-demand-Modelle laden daneben.
|
||||
|
||||
- Modell IN der Ko-Residenz-Gruppe (Hirn/embed/vision): koexistiert mit den ÜBRIGEN
|
||||
Gruppen-Mitgliedern → reserviert deren Summe.
|
||||
- Modell AUSSERHALB (heavy/coder/coder-lite/scout): lädt NEBEN die Gruppe →
|
||||
reserviert deren GESAMTE Summe (früher 0.0, weil der kaputte `persist`-Key die
|
||||
Gruppe verdrängen ließ — diese Rechnung erlaubte zu große Kontexte).
|
||||
"""
|
||||
from services import llamaswap
|
||||
models = llamaswap.list_models()
|
||||
groups = llamaswap.list_groups()
|
||||
cores = _coresident_members(groups)
|
||||
brain = next((m for m in models if (m.get("role") == "hermes")), None)
|
||||
brain_gb = footprint_gb(brain) if brain else 0.0
|
||||
role = (role or "").strip().lower()
|
||||
|
||||
holder = next((m for m in models if (m.get("role") == role)), None) if role else None
|
||||
holder_name = holder["name"] if holder else None
|
||||
# Hirn (hermes) ist per Definition Teil der Ko-Residenz-Gruppe; sonst Gruppen-Mitgliedschaft prüfen.
|
||||
in_group = role == "hermes" or bool(holder_name and holder_name in cores)
|
||||
|
||||
if in_group:
|
||||
others = sum(footprint_gb(m) for m in models
|
||||
if m["name"] in cores and m["name"] != holder_name)
|
||||
return {"reserved_gb": others, "mode": "co-resident", "brain_gb": brain_gb}
|
||||
# on-demand: lädt neben die (persistente) Ko-Residenz-Gruppe → deren Summe reservieren.
|
||||
warm = sum(footprint_gb(m) for m in models if m["name"] in cores)
|
||||
return {"reserved_gb": warm, "mode": "ondemand-beside-warmset", "brain_gb": brain_gb}
|
||||
|
||||
|
||||
def setup_aware_ctx(params_b: float, quant: str, role: str | None = None) -> dict:
|
||||
"""Größter Kontext, der für ein Modell (params_b/quant) der gegebenen Rolle NEBEN dem
|
||||
bestehenden Setup passt. Gibt ctx + die Budget-Herleitung zurück (für UI/Transparenz)."""
|
||||
gtt = gtt_budget_gb()
|
||||
r = reserved_gb(role)
|
||||
budget = max(gtt - r["reserved_gb"] - HEADROOM_GB, 0.0)
|
||||
ctx = max_ctx_in_budget(params_b, quant, budget)
|
||||
return {
|
||||
"ctx": ctx,
|
||||
"gtt_gb": gtt,
|
||||
"reserved_gb": round(r["reserved_gb"], 1),
|
||||
"budget_gb": round(budget, 1),
|
||||
"mode": r["mode"],
|
||||
}
|
||||
|
||||
|
||||
def _snap_ctx(raw_ctx: float, cap: int | None = None) -> int:
|
||||
"""Größter 'schöner' Kontext ≤ raw_ctx (und ≤ Trainings-Kontext des Modells, falls bekannt)."""
|
||||
from services.fit import _NICE_CTX
|
||||
if cap:
|
||||
raw_ctx = min(raw_ctx, cap)
|
||||
best = _NICE_CTX[0]
|
||||
for c in _NICE_CTX:
|
||||
if c <= raw_ctx:
|
||||
best = c
|
||||
return best
|
||||
|
||||
|
||||
def setup_aware_ctx_for_model(model: dict) -> dict:
|
||||
"""Setup-bewusster Optimal-ctx für ein INSTALLIERTES Modell. Für den 'Auto'-Button an der
|
||||
Modellkarte. Nutzt die ECHTE KV-Größe des GGUF (gleiche Zahlensprache wie footprint_gb) —
|
||||
Fallback auf die params-Heuristik nur, wenn das GGUF nicht lesbar ist."""
|
||||
from services import gguf_meta
|
||||
quant = model.get("quant") or "Q4_K_M"
|
||||
path = model.get("gguf_path")
|
||||
meta = gguf_meta.arch_meta(path) if path else None
|
||||
if not meta:
|
||||
return setup_aware_ctx(params_of_model(model), quant, role=model.get("role"))
|
||||
|
||||
gtt = gtt_budget_gb()
|
||||
r = reserved_gb(model.get("role"))
|
||||
budget = max(gtt - r["reserved_gb"] - HEADROOM_GB, 0.0)
|
||||
bpp = QUANT_BYTES_PER_PARAM.get(quant.upper(), 0.55)
|
||||
size_gb = (model.get("size_bytes") or 0) / (1024 ** 3)
|
||||
weights = max(params_of_model(model) * bpp, size_gb)
|
||||
ck, cv = _cache_types(model.get("cmd") or "")
|
||||
per_tok = gguf_meta.kv_gb_per_token(meta, ck, cv)
|
||||
ctx = _snap_ctx((budget - weights) / per_tok, cap=meta.get("n_ctx_train")) if per_tok > 0 else 2048
|
||||
return {"ctx": ctx, "gtt_gb": gtt, "reserved_gb": round(r["reserved_gb"], 1),
|
||||
"budget_gb": round(budget, 1), "mode": r["mode"]}
|
||||
@@ -0,0 +1,135 @@
|
||||
"""
|
||||
Modell-Capabilities — EINE Quelle der Wahrheit für Modell-Eigenschaften
|
||||
(MoE / Tools / Vision / Coder / Reasoning / Embedding / Kontext).
|
||||
Portiert aus Mission Control v1 (model_caps.py).
|
||||
|
||||
Quellen, geschichtet: GGUF-Header (offline, authoritativ) → cmd-Flags
|
||||
(--jinja/--mmproj) → HF-Block (tags + chat_template) → Familien-Fallback.
|
||||
Tool-Fähigkeit dreistufig: yes (bestätigt) | likely (Familie) | no.
|
||||
"""
|
||||
|
||||
import re
|
||||
import struct
|
||||
|
||||
from services.fit import extract_active_params_b, extract_params_b
|
||||
|
||||
_GGUF_FIXED = {0: 1, 1: 1, 2: 2, 3: 2, 4: 4, 5: 4, 6: 4, 7: 1, 10: 8, 11: 8, 12: 8}
|
||||
|
||||
|
||||
def _read_gguf_meta(path: str) -> dict:
|
||||
"""Liest nur den GGUF-Metadaten-Header (architecture/context_length/expert_count/
|
||||
parameter_count). Bricht vor dem Tokenizer-Array ab → schnell, lädt NICHT das Modell."""
|
||||
out: dict = {}
|
||||
try:
|
||||
with open(path, "rb") as f:
|
||||
if f.read(4) != b"GGUF":
|
||||
return {}
|
||||
struct.unpack("<I", f.read(4))[0]
|
||||
f.read(8)
|
||||
kv = struct.unpack("<Q", f.read(8))[0]
|
||||
|
||||
def ru32() -> int: return struct.unpack("<I", f.read(4))[0]
|
||||
def ru64() -> int: return struct.unpack("<Q", f.read(8))[0]
|
||||
def rstr() -> str: return f.read(ru64()).decode("utf-8", "replace")
|
||||
|
||||
def rval(t: int):
|
||||
if t == 8: return rstr()
|
||||
if t == 0: return struct.unpack("<B", f.read(1))[0]
|
||||
if t == 1: return struct.unpack("<b", f.read(1))[0]
|
||||
if t == 2: return struct.unpack("<H", f.read(2))[0]
|
||||
if t == 3: return struct.unpack("<h", f.read(2))[0]
|
||||
if t == 4: return struct.unpack("<I", f.read(4))[0]
|
||||
if t == 5: return struct.unpack("<i", f.read(4))[0]
|
||||
if t == 6: return struct.unpack("<f", f.read(4))[0]
|
||||
if t == 7: return f.read(1) != b"\x00"
|
||||
if t == 10: return struct.unpack("<Q", f.read(8))[0]
|
||||
if t == 11: return struct.unpack("<q", f.read(8))[0]
|
||||
if t == 12: return struct.unpack("<d", f.read(8))[0]
|
||||
if t == 9:
|
||||
et = ru32(); cnt = ru64()
|
||||
if et == 8:
|
||||
for _ in range(cnt):
|
||||
f.seek(ru64(), 1)
|
||||
elif et == 9:
|
||||
for _ in range(cnt):
|
||||
rval(9)
|
||||
else:
|
||||
f.seek(cnt * _GGUF_FIXED.get(et, 0), 1)
|
||||
return None
|
||||
raise ValueError(f"unbekannter GGUF-Typ {t}")
|
||||
|
||||
want = {"architecture", "context_length", "expert_count", "parameter_count"}
|
||||
for _ in range(kv):
|
||||
key = rstr()
|
||||
t = ru32()
|
||||
if key == "tokenizer.ggml.tokens":
|
||||
break
|
||||
v = rval(t)
|
||||
short = key.split(".")[-1]
|
||||
if short in want and short not in out:
|
||||
out[short] = v
|
||||
except Exception:
|
||||
return out
|
||||
return out
|
||||
|
||||
|
||||
_TOOL_FAMILIES = (
|
||||
"qwen2.5", "qwen3", "qwen2", "hermes", "mistral", "mixtral", "devstral",
|
||||
"command-r", "command_r", "llama-3.1", "llama3.1", "llama-3.3", "llama-4", "llama4",
|
||||
"functionary", "watt", "firefunction", "granite", "glm-4", "glm-5", "ministral",
|
||||
)
|
||||
_REASON_KW = (
|
||||
"-r1", "deepseek-r1", "qwq", "magistral", "-think", "thinking", "-o1",
|
||||
"gpt-oss", "reasoning", "exaone-deep", "phi-4-reasoning", "phi-4-mini-reasoning",
|
||||
)
|
||||
_CODE_KW = ("coder", "-code", "code-", "codestral", "starcoder", "deepseek-coder")
|
||||
_VISION_KW = ("-vl", "vision", "llava", "pixtral", "multimodal", "-mm-", "qwen3vl", "qwen2-vl")
|
||||
_EMBED_KW = ("bge", "e5-", "gte-", "nomic-embed", "embed")
|
||||
_MOE_ARCH = ("moe", "mixtral", "deepseek2", "deepseek3", "llama4", "qwen3moe", "grok")
|
||||
|
||||
|
||||
def capabilities(name: str = "", cmd: str = "", gguf_path: str = "", hf: dict | None = None) -> dict:
|
||||
"""Capability-Tag-Set für ein Modell. Alle Quellen optional — nutzt, was da ist."""
|
||||
low = (name or "").lower()
|
||||
cmdl = (cmd or "").lower()
|
||||
hf = hf or {}
|
||||
|
||||
meta = _read_gguf_meta(gguf_path) if gguf_path else {}
|
||||
arch = str(meta.get("architecture") or hf.get("architecture") or "").lower()
|
||||
tags = [str(t).lower() for t in (hf.get("tags") or [])]
|
||||
chat_tpl = str(hf.get("chat_template") or "")
|
||||
|
||||
expert_count = int(meta.get("expert_count") or 0)
|
||||
moe = (
|
||||
expert_count > 1
|
||||
or any(a in arch for a in _MOE_ARCH)
|
||||
or bool(re.search(r"\d+x\d+\.?\d*b", low))
|
||||
or bool(re.search(r"a\d+\.?\d*b", low))
|
||||
)
|
||||
active_b = extract_active_params_b(name)
|
||||
|
||||
pcount = int(meta.get("parameter_count") or 0)
|
||||
params_b = round(pcount / 1e9, 1) if pcount else extract_params_b(name)
|
||||
ctx = meta.get("context_length")
|
||||
if not ctx:
|
||||
m = re.search(r"-(?:c|-ctx-size)\s+(\d+)", cmdl)
|
||||
ctx = int(m.group(1)) if m else None
|
||||
|
||||
# Cap context length at 131072 for Qwen / Hermes models to prevent reporting scaled RoPE context of 256k+ which might OOM or be unstable.
|
||||
if ctx and ctx > 131072 and ("qwen" in low or "hermes" in low):
|
||||
ctx = 131072
|
||||
|
||||
tool_confirmed = "--jinja" in cmdl or "tool_call" in chat_tpl or "<tools>" in chat_tpl
|
||||
tool_family = any(fam in low for fam in _TOOL_FAMILIES) or "function-calling" in tags
|
||||
tools = "yes" if tool_confirmed else ("likely" if tool_family else "no")
|
||||
|
||||
vision = "--mmproj" in cmdl or "vl" in arch or "clip" in arch or any(k in low for k in _VISION_KW)
|
||||
coder = any(k in low for k in _CODE_KW)
|
||||
reasoning = any(k in low for k in _REASON_KW) or "reasoning" in tags
|
||||
embedding = "bert" in arch or any(k in low for k in _EMBED_KW)
|
||||
|
||||
return {
|
||||
"moe": moe, "active_b": active_b, "tools": tools, "vision": vision,
|
||||
"coder": coder, "reasoning": reasoning, "embedding": embedding,
|
||||
"ctx": ctx, "params_b": params_b or None, "arch": arch or None,
|
||||
}
|
||||
@@ -0,0 +1,115 @@
|
||||
"""
|
||||
Kuratierter Modell-Katalog ("Cookbook", inspiriert von Odysseus): EINE Quelle der
|
||||
Wahrheit für KORREKTE Metadaten (total/active params, moe, generation) statt
|
||||
Namens-Raterei. Macht Empfehlung + Upgrade-Erkennung präzise und MoE-bewusst
|
||||
für die bandbreiten-limitierte Strix-Halo-Box.
|
||||
|
||||
Daten: backend/models_catalog.json. Fällt sanft aus (leerer Katalog), wenn die
|
||||
Datei fehlt → discover nutzt dann nur die HF-Dynamik.
|
||||
"""
|
||||
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import re
|
||||
|
||||
from services.fit import estimate_memory_gb, estimate_speed
|
||||
|
||||
_CATALOG_PATH = os.path.join(os.path.dirname(__file__), "..", "models_catalog.json")
|
||||
_cache: dict = {"mtime": 0.0, "models": []}
|
||||
|
||||
|
||||
def _load() -> list[dict]:
|
||||
try:
|
||||
mt = os.path.getmtime(_CATALOG_PATH)
|
||||
if mt != _cache["mtime"]:
|
||||
with open(_CATALOG_PATH, encoding="utf-8") as f:
|
||||
data = json.load(f) or {}
|
||||
_cache.update(mtime=mt, models=[m for m in data.get("models", []) if m.get("name")])
|
||||
except (OSError, ValueError):
|
||||
_cache.update(mtime=0.0, models=[])
|
||||
return _cache["models"]
|
||||
|
||||
|
||||
def _norm(name: str) -> str:
|
||||
"""Vergleichs-Stamm: kleingeschrieben, Org-Prefix/Quant/GGUF/Split entfernt."""
|
||||
s = (name or "").lower().split("/")[-1]
|
||||
s = re.sub(r"\.gguf$", "", s)
|
||||
s = re.sub(r"-\d+-of-\d+$", "", s)
|
||||
s = re.sub(r"[-_](ud-)?(i?q\d[\w]*|f16|bf16|fp16|f32|mxfp4)$", "", s)
|
||||
return s.strip("-_ ")
|
||||
|
||||
|
||||
def entries() -> list[dict]:
|
||||
return list(_load())
|
||||
|
||||
|
||||
def entries_for_role(role: str) -> list[dict]:
|
||||
return [e for e in _load() if e.get("role") == role]
|
||||
|
||||
|
||||
def meta_for_name(name: str) -> dict | None:
|
||||
"""Katalog-Metadaten zu einem Modell(namen) — matcht lokalen Namen ODER HF-Repo."""
|
||||
n = _norm(name)
|
||||
if not n:
|
||||
return None
|
||||
for e in _load():
|
||||
cand = {_norm(e.get("name", "")), _norm(e.get("repo", ""))}
|
||||
if n in cand or any(c and (c in n or n in c) for c in cand):
|
||||
return e
|
||||
return None
|
||||
|
||||
|
||||
def fit_of(e: dict, ram_gb: float) -> dict:
|
||||
"""Hardware-Fit eines Katalog-Eintrags (MoE-bewusst über active_params_b)."""
|
||||
total = float(e.get("total_params_b") or 7)
|
||||
active = float(e.get("active_params_b") or total)
|
||||
quant = e.get("quant") or "Q4_K_M"
|
||||
ctx = int(e.get("ctx") or 32768)
|
||||
req_gb = estimate_memory_gb(total, quant, ctx)
|
||||
tps = estimate_speed(req_gb, ram_gb, (active / total) if total else 1.0)
|
||||
usable = max(ram_gb - 4.0, 0)
|
||||
if req_gb > usable:
|
||||
level, text = "too_tight", "Zu groß (OOM)"
|
||||
elif req_gb > usable * 0.8:
|
||||
level, text = "marginal", "Könnte knapp werden"
|
||||
else:
|
||||
level, text = "perfect", "Passt perfekt"
|
||||
return {"level": level, "text": text, "req_gb": round(req_gb, 1), "tps": round(tps, 0)}
|
||||
|
||||
|
||||
def stack_score(e: dict, ram_gb: float) -> float:
|
||||
"""Score für DIESE Hardware: muss passen, dann Wissen (total params) + Tempo
|
||||
(tps — belohnt MoE durch niedrige aktive Params automatisch). Bandbreiten-Box
|
||||
→ MoE gewinnt bei vergleichbarem Wissen gegen dense."""
|
||||
fit = fit_of(e, ram_gb)
|
||||
if fit["level"] == "too_tight":
|
||||
return -100.0
|
||||
total = float(e.get("total_params_b") or 7)
|
||||
fit_bonus = 3.0 if fit["level"] == "perfect" else 1.0
|
||||
knowledge = math.log2(total + 1) / 8.0 # ~0..1 (bis ~256B)
|
||||
speed = min((fit["tps"] or 0) / 80.0, 1.0) # normalisiert; MoE = hohe tps
|
||||
return fit_bonus + 1.2 * knowledge + 1.0 * speed
|
||||
|
||||
|
||||
def to_model_dict(e: dict, ram_gb: float) -> dict:
|
||||
"""Katalog-Eintrag → discover-kompatibles Modell-Dict (echte Metadaten)."""
|
||||
total = float(e.get("total_params_b") or 7)
|
||||
active = e.get("active_params_b")
|
||||
repo = e.get("repo") or e.get("name")
|
||||
role = e.get("role")
|
||||
caps = {
|
||||
"moe": bool(e.get("moe")), "active_b": active,
|
||||
"tools": "yes" if e.get("tools") else "no",
|
||||
"vision": bool(e.get("vision")), "coder": role == "coder",
|
||||
"reasoning": role == "heavy", "embedding": False,
|
||||
"ctx": e.get("ctx"), "params_b": total, "arch": e.get("family"),
|
||||
}
|
||||
return {
|
||||
"name": e.get("name"), "author": repo.split("/")[0] if "/" in repo else "catalog",
|
||||
"repo": repo, "role": role, "params_b": total, "active_b": active,
|
||||
"moe": bool(e.get("moe")), "generation": e.get("generation"),
|
||||
"family": e.get("family"), "quant": e.get("quant") or "Q4_K_M",
|
||||
"tags": ["catalog"], "downloads": 0, "fit": fit_of(e, ram_gb),
|
||||
"optimal_ctx": int(e.get("ctx") or 32768), "caps": caps, "curated": True,
|
||||
}
|
||||
@@ -0,0 +1,137 @@
|
||||
"""
|
||||
Connect: erzeugt saubere, getestete Konfig-Snippets für IDEs/Agenten auf dem
|
||||
LOKALEN PC (separate Maschine im LAN). Alle zeigen auf den **Gateway** der Box
|
||||
(Lanes `coding`/`chat`, Cockpit-Port :9001/v1) + den **Shared-Memory-MCP** (MC :9001).
|
||||
|
||||
Wichtig: Host ist die LAN-IP der Box (NICHT eine Proxy-Domain) — das war in v1
|
||||
die häufigste Fehlerquelle. Der Aufrufer übergibt den Host explizit.
|
||||
"""
|
||||
|
||||
import json
|
||||
|
||||
import httpx
|
||||
|
||||
from config import LLAMA_SWAP_URL, MEM0_SERVICE_URL, PORT
|
||||
|
||||
DEFAULT_HOST = "192.168.178.151"
|
||||
|
||||
# IDEs bekommen NUR die 'coding'-Lane zu sehen — die Lane routet intern selbst auf das
|
||||
# passende Modell (Coder/Heavy/…). Ein einziger Eintrag, kein manuelles Modell-Wählen mehr.
|
||||
IDE_MODEL = "coding"
|
||||
|
||||
|
||||
def _gw(host: str) -> str:
|
||||
# Eingebauter Gateway: MC2 serviert /v1 selbst (gleicher Port wie das Cockpit).
|
||||
return f"http://{host}:{PORT}/v1"
|
||||
|
||||
|
||||
def build_snippets(host: str = DEFAULT_HOST,
|
||||
mcp_script_path: str = r"F:\\Coding Stuff\\mission-control-2\\mcp\\mcp_memory.py",
|
||||
mcp_python: str = "python") -> dict:
|
||||
gw = _gw(host)
|
||||
mc_url = f"http://{host}:{PORT}"
|
||||
|
||||
# Kilo Code = aktiver Nachfolger der Roo/Cline-Linie (Roo im April 2026 eingestellt).
|
||||
# Gleiche Provider-Settings-Struktur; Multi-Agent kommt aus dem Tool (Orchestrator-Modus).
|
||||
kilo = json.dumps({
|
||||
"apiProvider": "openai",
|
||||
"openAiBaseUrl": gw,
|
||||
"openAiApiKey": "local",
|
||||
"openAiModelId": "coding",
|
||||
}, indent=2)
|
||||
|
||||
# Zed = Rust-nativ, ~16x leichter als VS Code, Agent Panel mit parallelen Agenten,
|
||||
# Windows stabil seit Okt 2025 — das Haupt-Tool für den User (Anti-Bloat + klickbar).
|
||||
zed = json.dumps({
|
||||
"language_models": {
|
||||
"openai_compatible": {
|
||||
"bosgame": {
|
||||
"api_url": gw,
|
||||
"available_models": [
|
||||
{"name": IDE_MODEL, "display_name": "Box / coding", "max_tokens": 131072,
|
||||
"capabilities": {"tools": True}}
|
||||
],
|
||||
}
|
||||
}
|
||||
},
|
||||
"agent": {
|
||||
"default_model": {"provider": "openai_compatible", "model": IDE_MODEL}
|
||||
}
|
||||
}, indent=2)
|
||||
|
||||
# Claude Code spricht das Anthropic-Format; der Gateway ist OpenAI-kompatibel und
|
||||
# bietet KEIN /v1/messages (verifiziert). Daher braucht es einen kleinen Übersetzer
|
||||
# (Anthropic ⇄ OpenAI) als Aufsatz. Die env-Vars sind Claude Codes echte Schnittstelle.
|
||||
claude_code = (
|
||||
f"# Claude Code spricht das Anthropic-Format — der Gateway ist OpenAI-kompatibel ({gw})\n"
|
||||
f"# und hat kein /v1/messages. Dazwischen muss ein Übersetzer (Anthropic ⇄ OpenAI) laufen:\n"
|
||||
f"# • claude-code-router (leichtgewichtig, npm)\n"
|
||||
f"# • oder LiteLLM mit /v1/messages-Bridge\n"
|
||||
f"# Den Übersetzer auf den Gateway zeigen lassen: baseURL={gw}, model=coding, apiKey=local.\n"
|
||||
f"# Dann Claude Code auf den lokalen Übersetzer richten (Beispiel-Port 3456):\n"
|
||||
f"\n"
|
||||
f'export ANTHROPIC_BASE_URL="http://localhost:3456"\n'
|
||||
f'export ANTHROPIC_AUTH_TOKEN="local"\n'
|
||||
f'export ANTHROPIC_MODEL="coding"'
|
||||
)
|
||||
|
||||
memory_mcp = json.dumps({
|
||||
"mcpServers": {
|
||||
"mission-control-memory": {
|
||||
"command": mcp_python,
|
||||
"args": [mcp_script_path],
|
||||
"env": {"MC_URL": mc_url},
|
||||
}
|
||||
}
|
||||
}, indent=2)
|
||||
|
||||
return {
|
||||
"host": host,
|
||||
"gateway_url": gw,
|
||||
"mc_url": mc_url,
|
||||
# Leitung 1 — das MODELL. Bewusst NUR 3 Tools (Verdikt 03.07.2026): Zed (Haupt-Tool:
|
||||
# Rust-nativ/leichtgewichtig, Agent Panel, parallele Agenten, Windows stabil seit 10/2025),
|
||||
# Kilo Code (VS-Code-Fallback mit Orchestrator-Modus), Claude Code (starke Sessions).
|
||||
# Cursor/Continue/Roo/OpenCode entfernt — Roo eingestellt, Rest cloud-first, Terminal
|
||||
# oder von Kilo abgedeckt. Das Evolution-Radar (AUTONOMIE_PLAN E4) überwacht die Kategorie.
|
||||
"tools": {
|
||||
"zed": {"label": "Zed (empfohlen)", "lang": "json", "snippet": zed,
|
||||
"note": "Leichtgewicht-Editor (Rust, ~16x weniger RAM als VS Code). settings.json → "
|
||||
"language_models.openai_compatible; Agent Panel nutzt dann die Box. Lane 'coding' routet intern."},
|
||||
"kilo": {"label": "Kilo Code", "lang": "json", "snippet": kilo,
|
||||
"note": "VS Code/JetBrains (schwergewichtiger). OpenAI-Provider → Gateway. "
|
||||
"API-Profile je Modus: Code→coding · Architect/Orchestrator→heavy · Ask/Debug→chat."},
|
||||
"claude_code": {"label": "Claude Code", "lang": "bash", "snippet": claude_code,
|
||||
"note": "Für die starken Sessions. Lokal-Betrieb bräuchte einen Anthropic⇄OpenAI-Übersetzer vor dem Gateway; Gedächtnis-Anbindung via Memory-MCP unten."},
|
||||
},
|
||||
# Leitung 2 — das GEDÄCHTNIS. Separater MCP-Server, gilt zusätzlich zu jedem Tool oben.
|
||||
"memory": {"label": "Shared Memory (MCP)", "lang": "json", "snippet": memory_mcp,
|
||||
"note": "Eigene Leitung: MCP-Block für jedes MCP-fähige Tool. mcp_memory.py muss lokal liegen."},
|
||||
}
|
||||
|
||||
|
||||
def check_health() -> dict:
|
||||
"""Live-Erreichbarkeit der beiden Leitungen, aus Sicht der Box:
|
||||
Leitung 1 = Gateway/Engine (llama-swap), Leitung 2 = Gedächtnis-Sidecar (Mem0)."""
|
||||
gateway = {"ok": False, "detail": "nicht erreichbar"}
|
||||
try:
|
||||
with httpx.Client(timeout=3.0) as c:
|
||||
r = c.get(f"{LLAMA_SWAP_URL}/v1/models")
|
||||
if r.status_code == 200:
|
||||
n = len(r.json().get("data", []))
|
||||
gateway = {"ok": True, "detail": f"{n} Modelle verfügbar" if n else "bereit"}
|
||||
else:
|
||||
gateway = {"ok": False, "detail": f"HTTP {r.status_code}"}
|
||||
except Exception: # noqa: BLE001
|
||||
pass
|
||||
|
||||
memory = {"ok": False, "detail": "nicht erreichbar"}
|
||||
try:
|
||||
with httpx.Client(timeout=3.0) as c:
|
||||
r = c.get(f"{MEM0_SERVICE_URL}/health")
|
||||
memory = ({"ok": True, "detail": "bereit"} if r.status_code == 200
|
||||
else {"ok": False, "detail": f"HTTP {r.status_code}"})
|
||||
except Exception: # noqa: BLE001
|
||||
pass
|
||||
|
||||
return {"gateway": gateway, "memory": memory}
|
||||
@@ -0,0 +1,179 @@
|
||||
"""
|
||||
Automatische Modell-Entdeckung ("aktuell beste Modelle"): fragt vertrauenswürdige
|
||||
HF-Orgs live ab, kategorisiert per Stichwort, rankt nach Hardware-Fit + Beliebtheit
|
||||
und cached. Portiert aus Mission Control v1 (cookbook.py-Discover).
|
||||
|
||||
Wichtig (Greenfield-Fix gegen v1): EIN gemeinsamer Ranking-Helfer `rank_runnable`
|
||||
ist die Quelle der Wahrheit — sowohl die „beste Empfehlung" je Kategorie als auch
|
||||
spätere Auto-Setups nutzen ihn, damit sie nie auseinanderlaufen.
|
||||
"""
|
||||
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import time
|
||||
from datetime import datetime
|
||||
|
||||
import httpx
|
||||
|
||||
import logging
|
||||
|
||||
from config import DISCOVER_CACHE_PATH, DISCOVER_TTL
|
||||
from services import catalog
|
||||
from services.caps import capabilities
|
||||
from services.fit import evaluate_fit, extract_params_b, max_ctx_for
|
||||
from services.sources import CATEGORIES, SKIP_TOKENS, TRUSTED_AUTHORS
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
_FIT_ORDER = {"perfect": 0, "marginal": 1, "too_tight": 2}
|
||||
|
||||
|
||||
def _categorize(repo_id: str) -> str:
|
||||
low = repo_id.lower()
|
||||
for cat in CATEGORIES:
|
||||
if any(k in low for k in cat["kw"]):
|
||||
return cat["role"]
|
||||
return "scout"
|
||||
|
||||
|
||||
def _fetch_author_models(author: str) -> list:
|
||||
url = (f"https://huggingface.co/api/models?author={author}"
|
||||
f"&filter=gguf&sort=downloads&direction=-1&limit=40")
|
||||
try:
|
||||
with httpx.Client(timeout=12.0) as c:
|
||||
data = c.get(url).json()
|
||||
return data if isinstance(data, list) else []
|
||||
except Exception:
|
||||
log.debug("discover: Abfrage für Autor %s fehlgeschlagen", author, exc_info=True)
|
||||
return []
|
||||
|
||||
|
||||
def _age_days(last_modified, now_ts: float) -> float:
|
||||
"""Alter eines HF-Modells in Tagen (lastModified ISO). Unbekannt → ~1.5 Jahre."""
|
||||
if not last_modified:
|
||||
return 540.0
|
||||
try:
|
||||
dt = datetime.fromisoformat(str(last_modified).replace("Z", "+00:00"))
|
||||
return max((now_ts - dt.timestamp()) / 86400.0, 0.0)
|
||||
except Exception:
|
||||
return 540.0
|
||||
|
||||
|
||||
def _score(m: dict, now_ts: float) -> float:
|
||||
"""Zukunftssicherer Rang-Score für DIESE Hardware. Kombiniert:
|
||||
- Fit: perfect dominiert (Bonus 3.0 > Summe der übrigen Terme → passt-komfortabel zuerst),
|
||||
- Recency: neuere Generationen bevorzugt (Halbwertszeit ~9 Monate über lastModified),
|
||||
- Capability: mehr Parameter (log-skaliert),
|
||||
- Popularity: Downloads (log-skaliert).
|
||||
So gewinnt bei vergleichbarer Größe die NEUERE Generation (z.B. Qwen3-Coder vor
|
||||
Qwen2.5-Coder), ohne dass kleine Populär-Modelle große verdrängen."""
|
||||
fit_bonus = 3.0 if m["fit"]["level"] == "perfect" else 0.0
|
||||
recency = 0.5 ** (_age_days(m.get("lastModified"), now_ts) / 270.0)
|
||||
cap = math.log2(max(float(m.get("params_b") or 1.0), 1.0) + 1.0) / 8.0
|
||||
pop = math.log10(float(m.get("downloads") or 0) + 1.0) / 7.0
|
||||
return fit_bonus + 1.2 * recency + 1.2 * cap + 0.5 * pop
|
||||
|
||||
|
||||
def rank_runnable(models: list[dict]) -> list[dict]:
|
||||
"""EINE Quelle der Wahrheit fürs Ranking lauffähiger Modelle für DIESE Hardware.
|
||||
Nur was passt (too_tight fliegt raus), dann nach `_score` (Fit + Recency + Capability
|
||||
+ Popularity). Bevorzugt neuere, fähige Modelle → zukunftssicher; „Modelle finden"
|
||||
schlägt nie ein Downgrade vor (Downgrade-Sperre zusätzlich in maintenance)."""
|
||||
now_ts = time.time()
|
||||
return sorted(
|
||||
[m for m in models if m["fit"]["level"] != "too_tight"],
|
||||
key=lambda m: -_score(m, now_ts),
|
||||
)
|
||||
|
||||
|
||||
def refresh_discover(ram_gb: float) -> dict:
|
||||
"""Quellen live abfragen, kategorisieren, ranken, cachen. Wirft nur, wenn KEINE
|
||||
Quelle erreichbar war."""
|
||||
raw, seen, ok = [], set(), 0
|
||||
for author in TRUSTED_AUTHORS:
|
||||
models = _fetch_author_models(author)
|
||||
if models:
|
||||
ok += 1
|
||||
for m in models:
|
||||
rid = m.get("id")
|
||||
if not rid or rid in seen:
|
||||
continue
|
||||
seen.add(rid)
|
||||
raw.append(m)
|
||||
if ok == 0 and not catalog.entries():
|
||||
raise RuntimeError("Keine Quelle erreichbar.")
|
||||
|
||||
by_cat: dict[str, list] = {c["role"]: [] for c in CATEGORIES}
|
||||
for m in raw:
|
||||
rid = m["id"]
|
||||
low = rid.lower()
|
||||
if any(tok in low for tok in SKIP_TOKENS):
|
||||
continue
|
||||
role = _categorize(rid)
|
||||
params_b = extract_params_b(rid)
|
||||
quant = "Q4_K_M" # Referenz-Quant für die Fit-Einschätzung
|
||||
fit = evaluate_fit(params_b, quant, 8192, ram_gb, name=rid)
|
||||
tags = [str(t) for t in (m.get("tags") or [])]
|
||||
by_cat[role].append({
|
||||
"name": rid.split("/")[-1], "author": rid.split("/")[0], "repo": rid,
|
||||
"role": role, "params_b": params_b, "quant": quant, "tags": tags,
|
||||
"downloads": int(m.get("downloads") or 0), "likes": int(m.get("likes") or 0),
|
||||
"lastModified": m.get("lastModified"),
|
||||
"fit": fit, "optimal_ctx": max_ctx_for(params_b, quant, ram_gb),
|
||||
"caps": capabilities(name=rid, hf={"tags": tags}),
|
||||
})
|
||||
|
||||
cats = []
|
||||
for c in CATEGORIES:
|
||||
role = c["role"]
|
||||
# 1) KATALOG zuerst (kuratierte, korrekte Metadaten, MoE-bewusst gerankt) —
|
||||
# macht die Empfehlung präzise statt Namens-Raterei.
|
||||
cat_entries = sorted(catalog.entries_for_role(role),
|
||||
key=lambda e: -catalog.stack_score(e, ram_gb))
|
||||
cat_models = [m for m in (catalog.to_model_dict(e, ram_gb) for e in cat_entries)
|
||||
if m["fit"]["level"] != "too_tight"]
|
||||
# 2) HF-Dynamik als Ergänzung (nicht-kuratierte Funde), dedupliziert.
|
||||
hf_ranked = rank_runnable(by_cat[role])
|
||||
seen = {catalog._norm(m["repo"]) for m in cat_models}
|
||||
extra = [h for h in hf_ranked if catalog._norm(h["repo"]) not in seen]
|
||||
combined = cat_models + extra
|
||||
if combined:
|
||||
cats.append({
|
||||
"role": role, "title": c["title"], "icon": c["icon"],
|
||||
"models": combined[:6],
|
||||
# Empfehlung = bester KURATIERTER Eintrag, sonst beste HF-Fundstelle.
|
||||
"recommended": (cat_models[0]["repo"] if cat_models
|
||||
else (hf_ranked[0]["repo"] if hf_ranked else None)),
|
||||
})
|
||||
|
||||
data = {"updated": time.time(), "categories": cats}
|
||||
try:
|
||||
DISCOVER_CACHE_PATH.parent.mkdir(parents=True, exist_ok=True)
|
||||
tmp = DISCOVER_CACHE_PATH.with_name(DISCOVER_CACHE_PATH.name + ".tmp")
|
||||
tmp.write_text(json.dumps(data, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
os.replace(tmp, DISCOVER_CACHE_PATH)
|
||||
except Exception:
|
||||
log.debug("discover: Cache-Schreiben fehlgeschlagen (nur Beschleunigung)", exc_info=True)
|
||||
return data
|
||||
|
||||
|
||||
def load_discover() -> dict | None:
|
||||
try:
|
||||
if DISCOVER_CACHE_PATH.exists():
|
||||
return json.loads(DISCOVER_CACHE_PATH.read_text(encoding="utf-8"))
|
||||
except Exception:
|
||||
log.debug("discover: Cache-Lesen fehlgeschlagen", exc_info=True)
|
||||
return None
|
||||
|
||||
|
||||
def safe_discover(ram_gb: float) -> dict | None:
|
||||
"""Aus Cache (wenn frisch) oder live; wirft nie — None wenn nichts da."""
|
||||
cached = load_discover()
|
||||
if cached and (time.time() - cached.get("updated", 0) < DISCOVER_TTL):
|
||||
return cached
|
||||
try:
|
||||
return refresh_discover(ram_gb)
|
||||
except Exception:
|
||||
log.warning("discover: Live-Refresh fehlgeschlagen, nutze Cache", exc_info=True)
|
||||
return cached
|
||||
@@ -0,0 +1,106 @@
|
||||
"""
|
||||
Hardware-Fit-Mathe (VRAM/RAM, tps-Schätzung) für APUs mit Unified Memory
|
||||
(Bosgame M5 / Strix Halo). Portiert aus Mission Control v1 (hw_math.py).
|
||||
"""
|
||||
|
||||
import re
|
||||
|
||||
# Bytes pro Parameter je GGUF-Quant (Annahme).
|
||||
QUANT_BYTES_PER_PARAM = {
|
||||
"Q2_K": 0.35, "Q3_K_S": 0.38, "Q3_K_M": 0.42, "Q3_K_L": 0.45,
|
||||
"Q4_0": 0.50, "Q4_1": 0.55, "Q4_K_S": 0.50, "Q4_K_M": 0.55,
|
||||
"Q5_0": 0.62, "Q5_1": 0.68, "Q5_K_S": 0.62, "Q5_K_M": 0.65,
|
||||
"Q6_K": 0.75, "Q8_0": 1.00, "F16": 2.00, "BF16": 2.00,
|
||||
"MXFP4": 0.55, "FP8": 1.05, "AWQ": 0.55,
|
||||
}
|
||||
|
||||
|
||||
def estimate_memory_gb(params_b: float, quant: str, ctx: int) -> float:
|
||||
"""Geschätzter Speicherbedarf in GB (Gewichte + Kontext-KV).
|
||||
KV-Cache skaliert NICHT linear mit den Gesamt-Parametern (er hängt an
|
||||
Layern × KV-Heads, gedämpft durch GQA) → sqrt-Skalierung, kalibriert am
|
||||
gemessenen Punkt Hermes-4-14B @ 128K ≈ 19 GB KV."""
|
||||
bpp = QUANT_BYTES_PER_PARAM.get(quant.upper(), 0.65)
|
||||
weights = params_b * bpp
|
||||
context_vram = (ctx / 8192) * (max(params_b, 7) / 7) ** 0.5 * 0.84
|
||||
return weights + context_vram
|
||||
|
||||
|
||||
def extract_active_params_b(name: str) -> float | None:
|
||||
"""Aktive Parameter bei MoE ('30B-A3B' → 3.0). None bei Dense."""
|
||||
m = re.search(r"(?<![a-zA-Z])a(\d+(?:\.\d+)?)b\b", name.lower())
|
||||
return float(m.group(1)) if m else None
|
||||
|
||||
|
||||
def estimate_speed(req_gb: float, sys_ram_gb: float, moe_active_ratio: float = 1.0) -> float:
|
||||
"""Geschätzte t/s anhand der ~273 GB/s Bandbreite der APU.
|
||||
moe_active_ratio = aktive/gesamt Params; < 1 bei MoE."""
|
||||
bw = 273 if sys_ram_gb > 8 else 70
|
||||
if req_gb <= 0:
|
||||
return 0.0
|
||||
raw_tps = (bw / req_gb) * 0.55
|
||||
if moe_active_ratio < 0.8:
|
||||
raw_tps *= (1.0 / moe_active_ratio) ** 0.5
|
||||
return raw_tps
|
||||
|
||||
|
||||
def evaluate_fit(params_b: float, quant: str, ctx: int, sys_ram_gb: float, name: str = "") -> dict:
|
||||
"""Fit für ein Shared-Memory-System (APU). name → MoE-Erkennung (optional)."""
|
||||
req_gb = estimate_memory_gb(params_b, quant, ctx)
|
||||
active_b = extract_active_params_b(name) if name else None
|
||||
moe_ratio = (active_b / params_b) if (active_b and params_b > 0) else 1.0
|
||||
tps = estimate_speed(req_gb, sys_ram_gb, moe_ratio)
|
||||
usable_ram = max(sys_ram_gb - 4.0, 0)
|
||||
if req_gb > usable_ram:
|
||||
fit_level, text = "too_tight", "Zu groß (OOM)"
|
||||
elif req_gb > usable_ram * 0.8:
|
||||
fit_level, text = "marginal", "Könnte knapp werden"
|
||||
else:
|
||||
fit_level, text = "perfect", "Passt perfekt"
|
||||
return {"level": fit_level, "text": text, "req_gb": round(req_gb, 1), "tps": round(tps, 0)}
|
||||
|
||||
|
||||
def extract_params_b(name: str) -> float:
|
||||
"""Parametergröße (Mrd.) aus Repo-/Dateiname. 8x7B (MoE) → 56."""
|
||||
moe = re.search(r"(\d+)x(\d+(?:\.\d+)?)[bB]", name)
|
||||
if moe:
|
||||
return float(moe.group(1)) * float(moe.group(2))
|
||||
m = re.search(r"(\d+(?:\.\d+)?)[bB](?![a-zA-Z])", name)
|
||||
return float(m.group(1)) if m else 7.0
|
||||
|
||||
|
||||
_NICE_CTX = [2048, 4096, 8192, 16384, 32768, 49152, 65536, 98304, 131072]
|
||||
|
||||
|
||||
def max_ctx_in_budget(params_b: float, quant: str, budget_gb: float) -> int:
|
||||
"""Größter 'schöner' Kontext, dessen Gewichte + KV in budget_gb passen.
|
||||
Budget-basierter Kern → wird von der setup-bewussten ctx-Vergabe
|
||||
(services.budget) mit dem ECHTEN freien Budget gefüttert."""
|
||||
bpp = QUANT_BYTES_PER_PARAM.get(quant.upper(), 0.65)
|
||||
weights = params_b * bpp
|
||||
ctx_budget = budget_gb - weights
|
||||
if ctx_budget <= 0:
|
||||
return 2048
|
||||
# KV pro 8k — EXAKTE Inverse von estimate_memory_gb (sqrt, kalibriert an
|
||||
# Hermes-14B@128K≈19GB). Vorher linear → für große Modelle viel zu konservativ.
|
||||
per_8k = (max(params_b, 7) / 7) ** 0.5 * 0.84
|
||||
raw_ctx = (ctx_budget / per_8k) * 8192
|
||||
best = _NICE_CTX[0]
|
||||
for c in _NICE_CTX:
|
||||
if c <= raw_ctx:
|
||||
best = c
|
||||
return best
|
||||
|
||||
|
||||
def max_ctx_for(params_b: float, quant: str, sys_ram_gb: float) -> int:
|
||||
"""Roh-Obergrenze: größter Kontext für dieses Modell ALLEIN gegen den
|
||||
Gesamt-RAM (80 % nutzbar). Ignoriert bewusst das übrige Setup —
|
||||
setup-bewusst rechnet services.budget.setup_aware_ctx."""
|
||||
return max_ctx_in_budget(params_b, quant, max(sys_ram_gb - 4.0, 0) * 0.8)
|
||||
|
||||
|
||||
def recommend_ctx(params_b: float, quant: str, sys_ram_gb: float) -> dict:
|
||||
ctx = max_ctx_for(params_b, quant, sys_ram_gb)
|
||||
k = ctx // 1024
|
||||
return {"ctx": ctx, "k": k,
|
||||
"note": f"Bis ~{k}k Kontext passt komfortabel auf deine Hardware ({round(sys_ram_gb)} GB)."}
|
||||
@@ -0,0 +1,46 @@
|
||||
"""
|
||||
Routing-Gateway-Status (eingebauter Modus). MC2 IST der Gateway: serviert
|
||||
`/v1/*` mit `model: auto`-Komplexitäts-Routing vor llama-swap. Kein externer
|
||||
LiteLLM-Dienst nötig (baut auf Python 3.14 nicht); bleibt später austauschbar.
|
||||
"""
|
||||
|
||||
from config import PORT
|
||||
from services.llamaswap import engine_reachable
|
||||
from services.routing_policy import load_policy
|
||||
|
||||
|
||||
def routing_summary() -> dict:
|
||||
p = load_policy()
|
||||
coding_default = p["coder_lite"] or p["coder"]
|
||||
return {
|
||||
"mode": "builtin",
|
||||
"endpoint": f":{PORT}/v1 (OpenAI-kompatibel)",
|
||||
# Virtuelle Lanes, die Clients/IDEs als „Modell" wählen (Router pickt das echte Alias).
|
||||
"lanes": [
|
||||
{
|
||||
"name": "chat",
|
||||
"aka": "auto",
|
||||
"target": f"{p['fast']} ↔ {p['heavy']} (nach Komplexität)",
|
||||
"threshold_chars": p["heavy_chars"],
|
||||
},
|
||||
{
|
||||
"name": "coding",
|
||||
"target": f"{coding_default} ↔ {p['coder']} (Eskalation)",
|
||||
"escalate_chars": p["coding_escalate_chars"],
|
||||
},
|
||||
],
|
||||
# Rückwärtskompatible Flach-Liste (alte UI/Clients).
|
||||
"routes": [
|
||||
{"name": "chat", "target": f"{p['fast']} ↔ {p['heavy']} (nach Komplexität)"},
|
||||
{"name": "coding", "target": f"{coding_default} ↔ {p['coder']} (Eskalation)"},
|
||||
{"name": "<alias>", "target": "llama-swap-Passthrough (lädt bei Bedarf)"},
|
||||
],
|
||||
"heavy_threshold_chars": p["heavy_chars"],
|
||||
"fallbacks": [],
|
||||
"context_window_fallbacks": [],
|
||||
}
|
||||
|
||||
|
||||
def gateway_reachable() -> bool:
|
||||
# Der eingebaute Gateway lebt in MC und proxyt llama-swap → erreichbar, wenn Engine läuft.
|
||||
return engine_reachable()
|
||||
@@ -0,0 +1,41 @@
|
||||
"""Token-Erfassung für den Builtin-Gateway.
|
||||
|
||||
Parst die `usage`-Felder aus llama-swap-Antworten (Stream + Non-Stream) und meldet
|
||||
sie an token_stats. Hält den gateway_proxy-Router dünn und ersetzt die zuvor inline
|
||||
verstreute, still scheiternde String-Suche durch einen testbaren SSE-Zeilenparser.
|
||||
"""
|
||||
|
||||
import json
|
||||
import logging
|
||||
|
||||
from services.token_stats import increment_tokens
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
|
||||
def record_usage(usage: dict | None, model: str) -> None:
|
||||
"""Ein usage-Objekt verbuchen (no-op bei None/leer)."""
|
||||
if not usage:
|
||||
return
|
||||
prompt = usage.get("prompt_tokens", 0)
|
||||
completion = usage.get("completion_tokens", 0)
|
||||
if prompt or completion:
|
||||
increment_tokens(prompt, completion, model=model)
|
||||
|
||||
|
||||
def record_stream_chunk(chunk: bytes, model: str) -> None:
|
||||
"""Rohen SSE-Chunk auf `usage` prüfen und Tokens verbuchen. Fehler werden
|
||||
geloggt (debug) statt verschluckt — ein defekter Chunk bricht den Stream nicht."""
|
||||
if b'"usage"' not in chunk:
|
||||
return
|
||||
text = chunk.decode("utf-8", errors="ignore")
|
||||
for line in text.splitlines():
|
||||
if not line.startswith("data:"):
|
||||
continue
|
||||
data_str = line[5:].strip()
|
||||
if not data_str or data_str == "[DONE]":
|
||||
continue
|
||||
try:
|
||||
record_usage(json.loads(data_str).get("usage"), model)
|
||||
except json.JSONDecodeError:
|
||||
log.debug("gateway stream: usage-Parsing fehlgeschlagen: %s", data_str[:120])
|
||||
@@ -0,0 +1,266 @@
|
||||
"""
|
||||
GGUF-Tokenizer-Fingerprint — liest die Tokenizer-Identität direkt aus dem
|
||||
GGUF-Header (ohne das Modell zu laden), um zu entscheiden, ob ein Draft-Modell
|
||||
**vocab-kompatibel** mit einem Ziel-Modell ist (Voraussetzung für Speculative
|
||||
Decoding in llama.cpp — sonst: "draft model vocab type must match target").
|
||||
|
||||
Wir lesen nur die Metadaten-KV-Sektion am Dateianfang und brechen ab, sobald
|
||||
`tokenizer.ggml.tokens` erreicht ist (dessen Länge = n_vocab). model+pre+n_vocab
|
||||
identifizieren den Tokenizer eindeutig genug, um die in der Praxis relevanten
|
||||
Fälle zu unterscheiden (Qwen2.5 vs Qwen3 vs Qwen3.6 etc.). Die llama.cpp-Prüfung
|
||||
beim Laden bleibt der letzte Schiedsrichter.
|
||||
"""
|
||||
|
||||
import hashlib
|
||||
import struct
|
||||
from functools import lru_cache
|
||||
|
||||
# GGUF value types (https://github.com/ggml-org/ggml/blob/master/docs/gguf.md)
|
||||
_T_UINT8, _T_INT8, _T_UINT16, _T_INT16, _T_UINT32, _T_INT32, _T_FLOAT32, \
|
||||
_T_BOOL, _T_STRING, _T_ARRAY, _T_UINT64, _T_INT64, _T_FLOAT64 = range(13)
|
||||
|
||||
_SCALAR_FMT = {
|
||||
_T_UINT8: "<B", _T_INT8: "<b", _T_UINT16: "<H", _T_INT16: "<h",
|
||||
_T_UINT32: "<I", _T_INT32: "<i", _T_FLOAT32: "<f", _T_BOOL: "<?",
|
||||
_T_UINT64: "<Q", _T_INT64: "<q", _T_FLOAT64: "<d",
|
||||
}
|
||||
_SCALAR_SIZE = {t: struct.calcsize(f) for t, f in _SCALAR_FMT.items()}
|
||||
|
||||
_WANT_STRINGS = {"tokenizer.ggml.model", "tokenizer.ggml.pre", "general.architecture"}
|
||||
|
||||
|
||||
class _Reader:
|
||||
def __init__(self, f):
|
||||
self.f = f
|
||||
|
||||
def read(self, n: int) -> bytes:
|
||||
b = self.f.read(n)
|
||||
if len(b) != n:
|
||||
raise EOFError("unerwartetes Dateiende beim GGUF-Parsen")
|
||||
return b
|
||||
|
||||
def u32(self) -> int:
|
||||
return struct.unpack("<I", self.read(4))[0]
|
||||
|
||||
def u64(self) -> int:
|
||||
return struct.unpack("<Q", self.read(8))[0]
|
||||
|
||||
def gstr(self) -> str:
|
||||
n = self.u64()
|
||||
return self.read(n).decode("utf-8", "replace")
|
||||
|
||||
def scalar(self, vtype: int):
|
||||
"""Liest einen Skalar-Wert (für die Architektur-Metadaten). None bei Nicht-Skalar."""
|
||||
fmt = _SCALAR_FMT.get(vtype)
|
||||
if not fmt:
|
||||
self.skip_value(vtype)
|
||||
return None
|
||||
return struct.unpack(fmt, self.read(_SCALAR_SIZE[vtype]))[0]
|
||||
|
||||
def skip_value(self, vtype: int) -> None:
|
||||
"""Liest einen Wert und verwirft ihn (um den Datei-Pointer korrekt
|
||||
weiterzuschieben). Arrays werden elementweise konsumiert."""
|
||||
if vtype == _T_STRING:
|
||||
self.f.seek(self.u64(), 1)
|
||||
elif vtype in _SCALAR_SIZE:
|
||||
self.f.seek(_SCALAR_SIZE[vtype], 1)
|
||||
elif vtype == _T_ARRAY:
|
||||
etype = self.u32()
|
||||
count = self.u64()
|
||||
if etype == _T_STRING:
|
||||
for _ in range(count):
|
||||
self.f.seek(self.u64(), 1)
|
||||
elif etype in _SCALAR_SIZE:
|
||||
self.f.seek(_SCALAR_SIZE[etype] * count, 1)
|
||||
else:
|
||||
raise ValueError(f"unbekannter Array-Elementtyp {etype}")
|
||||
else:
|
||||
raise ValueError(f"unbekannter GGUF-Wertetyp {vtype}")
|
||||
|
||||
|
||||
def _read_fingerprint(path: str) -> dict | None:
|
||||
"""Liest model/pre/n_vocab aus dem GGUF-Header. None bei Fehler/kein GGUF."""
|
||||
try:
|
||||
with open(path, "rb") as fh:
|
||||
r = _Reader(fh)
|
||||
if r.read(4) != b"GGUF":
|
||||
return None
|
||||
r.u32() # version
|
||||
r.u64() # tensor_count
|
||||
kv_count = r.u64()
|
||||
fp: dict = {"model": None, "pre": None, "arch": None, "n_vocab": None,
|
||||
"tokens_sha": None}
|
||||
for _ in range(kv_count):
|
||||
key = r.gstr()
|
||||
vtype = r.u32()
|
||||
if key == "tokenizer.ggml.tokens" and vtype == _T_ARRAY:
|
||||
etype = r.u32()
|
||||
count = r.u64()
|
||||
fp["n_vocab"] = count
|
||||
if etype != _T_STRING:
|
||||
return None
|
||||
# ECHTE Vocab-Identität: sha256 über die tatsächliche Token-Liste
|
||||
# (familienunabhängig — funktioniert für Qwen, Llama, Mistral, …).
|
||||
h = hashlib.sha256()
|
||||
h.update(count.to_bytes(8, "little"))
|
||||
for _ in range(count):
|
||||
n = r.u64()
|
||||
h.update(r.read(n))
|
||||
fp["tokens_sha"] = h.hexdigest()
|
||||
# model/pre kommen vor tokens → wir haben alles. Abbrechen.
|
||||
break
|
||||
if key in _WANT_STRINGS and vtype == _T_STRING:
|
||||
val = r.gstr()
|
||||
if key == "tokenizer.ggml.model":
|
||||
fp["model"] = val
|
||||
elif key == "tokenizer.ggml.pre":
|
||||
fp["pre"] = val
|
||||
else:
|
||||
fp["arch"] = val
|
||||
else:
|
||||
r.skip_value(vtype)
|
||||
if fp["model"] is None and fp["n_vocab"] is None:
|
||||
return None
|
||||
return fp
|
||||
except (OSError, EOFError, ValueError, struct.error):
|
||||
return None
|
||||
|
||||
|
||||
@lru_cache(maxsize=256)
|
||||
def _cached(path: str, mtime: float, size: int) -> tuple | None:
|
||||
fp = _read_fingerprint(path)
|
||||
if fp is None:
|
||||
return None
|
||||
return (fp.get("model"), fp.get("pre"), fp.get("n_vocab"), fp.get("arch"), fp.get("tokens_sha"))
|
||||
|
||||
|
||||
def fingerprint(path: str) -> dict | None:
|
||||
"""Tokenizer-Fingerprint eines GGUF (gecacht nach Pfad+mtime+size).
|
||||
Returns dict(model, pre, n_vocab, arch, tokens_sha) oder None wenn nicht lesbar."""
|
||||
import os
|
||||
try:
|
||||
st = os.stat(path)
|
||||
except OSError:
|
||||
return None
|
||||
t = _cached(path, st.st_mtime, st.st_size)
|
||||
if t is None:
|
||||
return None
|
||||
return {"model": t[0], "pre": t[1], "n_vocab": t[2], "arch": t[3], "tokens_sha": t[4]}
|
||||
|
||||
|
||||
def vocab_key(path: str) -> tuple | None:
|
||||
"""ECHTER Vergleichsschlüssel für Vocab-Kompatibilität: (model, pre, n_vocab, sha256
|
||||
der vollständigen Token-Liste). Vergleicht den TATSÄCHLICHEN Vokabular-Inhalt, nicht
|
||||
nur Metadaten — familienunabhängig (Qwen, Llama, Mistral, …). Genau diese Identität
|
||||
verlangt llama.cpp für Speculative Decoding."""
|
||||
fp = fingerprint(path)
|
||||
if not fp or fp["n_vocab"] is None or not fp.get("tokens_sha"):
|
||||
return None
|
||||
return (fp["model"], fp["pre"], fp["n_vocab"], fp["tokens_sha"])
|
||||
|
||||
|
||||
def compatible(target_path: str, draft_path: str) -> bool | None:
|
||||
"""True/False ob draft vocab-kompatibel zum target ist. None = unbestimmbar
|
||||
(eine Datei nicht lesbar) → UI behandelt das als 'nicht bestätigt'."""
|
||||
a = vocab_key(target_path)
|
||||
b = vocab_key(draft_path)
|
||||
if a is None or b is None:
|
||||
return None
|
||||
return a == b
|
||||
|
||||
|
||||
# ── Architektur-Metadaten für EHRLICHE KV-Cache-Größen ──────────────────────────────
|
||||
# Der KV-Cache hängt an (Layer × KV-Heads × Head-Dim), NICHT an den Gesamt-Parametern.
|
||||
# Bei MoE (z.B. Qwen3.6-35B-A3B) ist das entscheidend: die alte params-basierte Schätzung
|
||||
# überschätzte grob (aktive vs. gesamte Params + GQA), reale KV liest man direkt hier.
|
||||
# Schlüssel sind arch-präfixiert ('qwen3moe.block_count', 'llama.attention.head_count_kv' …),
|
||||
# gegen echte GGUFs verifiziert. Wir sammeln die gewünschten Skalar-Schlüssel per Suffix.
|
||||
_ARCH_WANT = (
|
||||
".block_count", ".attention.head_count_kv", ".attention.head_count",
|
||||
".attention.key_length", ".attention.value_length", ".embedding_length",
|
||||
".context_length",
|
||||
)
|
||||
|
||||
|
||||
def _read_arch_meta(path: str) -> dict | None:
|
||||
try:
|
||||
with open(path, "rb") as fh:
|
||||
r = _Reader(fh)
|
||||
if r.read(4) != b"GGUF":
|
||||
return None
|
||||
r.u32() # version
|
||||
r.u64() # tensor_count
|
||||
kv_count = r.u64()
|
||||
raw: dict = {}
|
||||
arch = None
|
||||
for _ in range(kv_count):
|
||||
key = r.gstr()
|
||||
vtype = r.u32()
|
||||
if key == "general.architecture" and vtype == _T_STRING:
|
||||
arch = r.gstr()
|
||||
continue
|
||||
if key == "tokenizer.ggml.tokens":
|
||||
break # Arch-Metadaten stehen davor → fertig, Rest überspringen
|
||||
suf = next((s for s in _ARCH_WANT if key.endswith(s)), None)
|
||||
if suf is not None and vtype in _SCALAR_SIZE:
|
||||
raw[suf] = r.scalar(vtype)
|
||||
else:
|
||||
r.skip_value(vtype)
|
||||
n_layers = raw.get(".block_count")
|
||||
n_head = raw.get(".attention.head_count")
|
||||
n_head_kv = raw.get(".attention.head_count_kv") or n_head # GQA fehlt → MHA
|
||||
n_embd = raw.get(".embedding_length")
|
||||
hd_k = raw.get(".attention.key_length") \
|
||||
or (int(n_embd / n_head) if (n_embd and n_head) else None)
|
||||
hd_v = raw.get(".attention.value_length") or hd_k
|
||||
if not (n_layers and n_head_kv and hd_k and hd_v):
|
||||
return None # unvollständig → Aufrufer nutzt Heuristik-Fallback
|
||||
return {"arch": arch, "n_layers": int(n_layers), "n_head_kv": int(n_head_kv),
|
||||
"head_dim_k": int(hd_k), "head_dim_v": int(hd_v),
|
||||
"n_ctx_train": int(raw[".context_length"]) if raw.get(".context_length") else None}
|
||||
except (OSError, EOFError, ValueError, struct.error):
|
||||
return None
|
||||
|
||||
|
||||
@lru_cache(maxsize=128)
|
||||
def _arch_cached(path: str, mtime: float, size: int) -> dict | None:
|
||||
return _read_arch_meta(path)
|
||||
|
||||
|
||||
def arch_meta(path: str) -> dict | None:
|
||||
"""Architektur-Metadaten eines GGUF (gecacht nach Pfad+mtime+size):
|
||||
{arch, n_layers, n_head_kv, head_dim_k, head_dim_v, n_ctx_train}. None wenn nicht lesbar
|
||||
oder unvollständig."""
|
||||
import os
|
||||
try:
|
||||
st = os.stat(path)
|
||||
except OSError:
|
||||
return None
|
||||
return _arch_cached(path, st.st_mtime, st.st_size)
|
||||
|
||||
|
||||
# Bytes pro KV-Cache-Element je cache-type (inkl. Block-Overhead der k-Quants).
|
||||
_KV_BPE = {
|
||||
"f32": 4.0, "f16": 2.0, "bf16": 2.0,
|
||||
"q8_0": 1.0625, "q5_1": 0.75, "q5_0": 0.6875,
|
||||
"q4_1": 0.625, "q4_0": 0.5625, "iq4_nl": 0.5625,
|
||||
}
|
||||
_GIB = 1024 ** 3
|
||||
|
||||
|
||||
def _bpe(cache_type: str | None) -> float:
|
||||
return _KV_BPE.get((cache_type or "f16").lower(), 2.0)
|
||||
|
||||
|
||||
def kv_cache_gb(meta: dict, ctx: int, ck: str | None = None, cv: str | None = None) -> float:
|
||||
"""Echte KV-Cache-Größe (GiB) für ctx Tokens, K/V ggf. quantisiert. Formel wie llama.cpp:
|
||||
je Layer & Token hält der Cache n_head_kv × head_dim Elemente für K und für V."""
|
||||
per_tok = meta["n_layers"] * meta["n_head_kv"] * ctx
|
||||
k = per_tok * meta["head_dim_k"] * _bpe(ck)
|
||||
v = per_tok * meta["head_dim_v"] * _bpe(cv)
|
||||
return (k + v) / _GIB
|
||||
|
||||
|
||||
def kv_gb_per_token(meta: dict, ck: str | None = None, cv: str | None = None) -> float:
|
||||
"""KV-GiB pro Kontext-Token — für den analytischen ctx-Solver (linear in ctx)."""
|
||||
return kv_cache_gb(meta, 1, ck, cv)
|
||||
@@ -0,0 +1,103 @@
|
||||
"""HuggingFace-Helfer: GGUF-Dateien eines Repos auflösen (inkl. Split-Teile) + Größen,
|
||||
freie Suche, Repo-URL→ID, verfügbare Quants."""
|
||||
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
|
||||
import httpx
|
||||
|
||||
|
||||
def normalize_repo(s: str) -> str:
|
||||
"""Akzeptiert volle HF-URL oder `org/repo` → liefert immer `org/repo`."""
|
||||
s = (s or "").strip()
|
||||
m = re.search(r"huggingface\.co/([^/\s]+/[^/\s?#]+)", s)
|
||||
if m:
|
||||
return m.group(1)
|
||||
return s.strip("/")
|
||||
|
||||
|
||||
def list_quants(repo: str) -> list[str]:
|
||||
"""Verfügbare Quant-Stufen eines Repos (aus den GGUF-Dateinamen, ohne mmproj)."""
|
||||
quants: set[str] = set()
|
||||
for e in _tree(repo):
|
||||
p = str(e.get("path", ""))
|
||||
if p.lower().endswith(".gguf") and "mmproj" not in p.lower():
|
||||
m = re.search(r"(I?Q\d[\w]*|F16|BF16|FP16|F32)", p, re.IGNORECASE)
|
||||
if m:
|
||||
quants.add(m.group(1).upper())
|
||||
# gängige Reihenfolge zuerst
|
||||
order = {"Q4_K_M": 0, "Q4_K_S": 1, "Q5_K_M": 2, "Q6_K": 3, "Q8_0": 4, "Q3_K_M": 5, "Q2_K": 6}
|
||||
return sorted(quants, key=lambda q: (order.get(q, 99), q))
|
||||
|
||||
|
||||
def search(q: str = "", limit: int = 24) -> list[dict]:
|
||||
"""Freie HF-Suche nach GGUF-Repos. Ohne q → Top-GGUF nach Downloads (Stöbern)."""
|
||||
url = (f"https://huggingface.co/api/models?filter=gguf"
|
||||
f"&sort=downloads&direction=-1&limit={limit}")
|
||||
if q and q.strip():
|
||||
url += f"&search={q.strip()}"
|
||||
try:
|
||||
with httpx.Client(timeout=12.0) as c:
|
||||
data = c.get(url).json()
|
||||
except Exception:
|
||||
return []
|
||||
out = []
|
||||
for m in (data if isinstance(data, list) else []):
|
||||
rid = m.get("id")
|
||||
if rid:
|
||||
out.append({"repo": rid, "downloads": int(m.get("downloads") or 0),
|
||||
"likes": int(m.get("likes") or 0)})
|
||||
return out
|
||||
|
||||
|
||||
def hf_bin() -> str:
|
||||
"""Pfad zur `hf`-CLI (bevorzugt neben dem laufenden Python im venv)."""
|
||||
cand = os.path.join(os.path.dirname(sys.executable), "hf")
|
||||
return cand if os.path.exists(cand) else "hf"
|
||||
|
||||
|
||||
def _tree(repo: str) -> list[dict]:
|
||||
url = f"https://huggingface.co/api/models/{repo}/tree/main?recursive=true"
|
||||
with httpx.Client(timeout=20.0) as c:
|
||||
data = c.get(url).json()
|
||||
return data if isinstance(data, list) else []
|
||||
|
||||
|
||||
def _size(entry: dict) -> int:
|
||||
return int(entry.get("size") or (entry.get("lfs") or {}).get("size") or 0)
|
||||
|
||||
|
||||
def resolve_gguf(repo: str, quant: str = "Q4_K_M") -> dict:
|
||||
"""Beste GGUF-Auswahl eines Repos für einen Quant. Behandelt Split-GGUFs
|
||||
(-00001-of-000NN) als Gruppe. Liefert die Datei-/Pattern-Infos für den Download.
|
||||
|
||||
Rückgabe: {files:[paths], first:path, total_bytes:int, mmproj:path|None, split:bool}
|
||||
"""
|
||||
tree = _tree(repo)
|
||||
ggufs = [e for e in tree if str(e.get("path", "")).lower().endswith(".gguf")]
|
||||
q = quant.lower()
|
||||
# mmproj separat (Vision-Projektor)
|
||||
mmproj = next((e["path"] for e in ggufs if "mmproj" in e["path"].lower()), None)
|
||||
model = [e for e in ggufs if "mmproj" not in e["path"].lower()]
|
||||
# bevorzugt den gewünschten Quant
|
||||
pref = [e for e in model if q in e["path"].lower()]
|
||||
chosen = pref or model
|
||||
if not chosen:
|
||||
return {"files": [], "first": None, "total_bytes": 0, "mmproj": mmproj, "split": False}
|
||||
# Split? Wenn die gewählten Dateien -of- enthalten → alle Teile dieser Gruppe.
|
||||
split = any("-of-" in e["path"].lower() for e in chosen)
|
||||
if split:
|
||||
parts = sorted([e for e in chosen if "-of-" in e["path"].lower()], key=lambda e: e["path"])
|
||||
files = [e["path"] for e in parts]
|
||||
first = files[0]
|
||||
total = sum(_size(e) for e in parts)
|
||||
else:
|
||||
# ein einzelnes File: nimm das kleinste passende (typisch genau eins)
|
||||
chosen.sort(key=lambda e: _size(e))
|
||||
first = chosen[0]["path"]
|
||||
files = [first]
|
||||
total = _size(chosen[0])
|
||||
if mmproj:
|
||||
total += next((_size(e) for e in ggufs if e["path"] == mmproj), 0)
|
||||
return {"files": files, "first": first, "total_bytes": total, "mmproj": mmproj, "split": split}
|
||||
@@ -0,0 +1,195 @@
|
||||
"""
|
||||
Mini-Job-System: Hintergrund-Prozesse mit Live-Log + Download-Fortschritt.
|
||||
Portiert aus Mission Control v1 (jobengine.py). In-Memory, ein Daemon-Thread je Job.
|
||||
"""
|
||||
|
||||
import glob
|
||||
import os
|
||||
import shlex
|
||||
import subprocess
|
||||
import threading
|
||||
import time
|
||||
import uuid
|
||||
|
||||
JOBS: dict[str, dict] = {}
|
||||
_PROCS: dict[str, subprocess.Popen] = {}
|
||||
_LOG_CAP = 400
|
||||
|
||||
|
||||
def _append_log(job: dict, line: str) -> None:
|
||||
job["log"].append(line)
|
||||
if len(job["log"]) > _LOG_CAP:
|
||||
del job["log"][0]
|
||||
|
||||
|
||||
def _pump_output(job: dict, stream) -> None:
|
||||
"""Liest byteweise; `\\r` (tqdm/hf-Fortschritt) überschreibt die letzte Zeile."""
|
||||
buf = b""
|
||||
overwrite = False
|
||||
pending_cr = False
|
||||
|
||||
def commit():
|
||||
line = buf.decode("utf-8", "replace")
|
||||
if overwrite and job["log"]:
|
||||
job["log"][-1] = line
|
||||
else:
|
||||
_append_log(job, line)
|
||||
|
||||
while True:
|
||||
ch = stream.read(1)
|
||||
if not ch:
|
||||
break
|
||||
if pending_cr:
|
||||
pending_cr = False
|
||||
if ch == b"\n":
|
||||
commit(); overwrite = False; buf = b""
|
||||
continue
|
||||
commit(); overwrite = True; buf = b""
|
||||
if ch == b"\r":
|
||||
pending_cr = True
|
||||
elif ch == b"\n":
|
||||
commit(); overwrite = False; buf = b""
|
||||
else:
|
||||
buf += ch
|
||||
if pending_cr:
|
||||
commit(); overwrite = True; buf = b""
|
||||
if buf:
|
||||
commit()
|
||||
|
||||
|
||||
def _run_job(job_id: str, args: list[str], env: dict | None = None, sudo_password: str | None = None):
|
||||
job = JOBS[job_id]
|
||||
job["state"] = "running"
|
||||
try:
|
||||
actual_args = list(args)
|
||||
if sudo_password is not None:
|
||||
for i, arg in enumerate(actual_args):
|
||||
if isinstance(arg, str):
|
||||
actual_args[i] = arg.replace("sudo -n", "sudo -S").replace("sudo ", "sudo -S ")
|
||||
|
||||
proc = subprocess.Popen(
|
||||
actual_args, stdout=subprocess.PIPE, stderr=subprocess.STDOUT,
|
||||
stdin=subprocess.PIPE if sudo_password is not None else None,
|
||||
bufsize=0,
|
||||
env={**os.environ, **(env or {})},
|
||||
)
|
||||
_PROCS[job_id] = proc
|
||||
|
||||
if sudo_password is not None and proc.stdin:
|
||||
proc.stdin.write((sudo_password + "\n").encode("utf-8"))
|
||||
proc.stdin.flush()
|
||||
proc.stdin.close()
|
||||
|
||||
_pump_output(job, proc.stdout)
|
||||
proc.wait()
|
||||
job["returncode"] = proc.returncode
|
||||
job["state"] = "canceled" if job.get("canceled") else ("done" if proc.returncode == 0 else "failed")
|
||||
|
||||
# Check if failed due to sudo authorization failure
|
||||
if proc.returncode != 0 and job["log"]:
|
||||
log_str = "\n".join(job["log"])
|
||||
if "a password is required" in log_str or "password" in log_str.lower() or "sudo:" in log_str:
|
||||
job["sudo_failed"] = True
|
||||
except Exception as exc: # noqa: BLE001
|
||||
_append_log(job, f"[mc] Fehler: {exc}")
|
||||
job["state"] = "failed"
|
||||
job["returncode"] = -1
|
||||
finally:
|
||||
_PROCS.pop(job_id, None)
|
||||
job["finished_at"] = time.time()
|
||||
cb = job.pop("_on_done", None)
|
||||
if cb and job["state"] == "done":
|
||||
try:
|
||||
cb()
|
||||
except Exception as exc: # noqa: BLE001
|
||||
_append_log(job, f"[mc] Nachbearbeitung-Fehler: {exc}")
|
||||
|
||||
|
||||
def attach_download_progress(job_id: str, local_dir: str, total_bytes: int) -> None:
|
||||
"""Fortschritt in % aus wachsenden *.incomplete-Dateien (hf schreibt sie)."""
|
||||
if not total_bytes or total_bytes <= 0:
|
||||
return
|
||||
job = JOBS.get(job_id)
|
||||
if job is not None:
|
||||
job["progress"] = 0
|
||||
job["total_bytes"] = total_bytes
|
||||
|
||||
def _watch():
|
||||
pat = os.path.join(local_dir, ".cache", "huggingface", "download", "**", "*.incomplete")
|
||||
prev_t = prev_b = None
|
||||
rate = 0.0
|
||||
while True:
|
||||
j = JOBS.get(job_id)
|
||||
if not j or j["state"] in ("done", "failed", "canceled"):
|
||||
break
|
||||
try:
|
||||
inc = glob.glob(pat, recursive=True)
|
||||
cur = sum(os.path.getsize(f) for f in inc) if inc else 0
|
||||
if cur:
|
||||
j["progress"] = min(99, int(cur * 100 / total_bytes))
|
||||
j["done_bytes"] = cur
|
||||
now = time.time()
|
||||
if prev_t is not None and now > prev_t and cur >= prev_b:
|
||||
inst = (cur - prev_b) / (now - prev_t)
|
||||
rate = inst if rate == 0 else 0.3 * inst + 0.7 * rate
|
||||
if rate > 0:
|
||||
j["rate_bps"] = rate
|
||||
j["eta_s"] = int((total_bytes - cur) / rate)
|
||||
prev_t, prev_b = now, cur
|
||||
except Exception: # noqa: BLE001
|
||||
pass
|
||||
time.sleep(1.0)
|
||||
j = JOBS.get(job_id)
|
||||
if j and j["state"] == "done":
|
||||
j["progress"] = 100
|
||||
j.pop("eta_s", None)
|
||||
|
||||
threading.Thread(target=_watch, daemon=True).start()
|
||||
|
||||
|
||||
def start_job(args: list[str], label: str, env: dict | None = None, on_done=None,
|
||||
sudo_password: str | None = None, group: str | None = None) -> str:
|
||||
job_id = uuid.uuid4().hex[:12]
|
||||
# Mask password in log if present in args
|
||||
log_args = list(args)
|
||||
JOBS[job_id] = {
|
||||
"id": job_id, "label": label, "state": "queued", "group": group,
|
||||
"log": ["$ " + " ".join(shlex.quote(a) for a in log_args)],
|
||||
"returncode": None, "started_at": time.time(), "finished_at": None,
|
||||
}
|
||||
if on_done:
|
||||
JOBS[job_id]["_on_done"] = on_done
|
||||
threading.Thread(target=_run_job, args=(job_id, args, env, sudo_password), daemon=True).start()
|
||||
return job_id
|
||||
|
||||
|
||||
def active_in_group(group: str) -> dict | None:
|
||||
"""Erster laufender/wartender Job einer Gruppe (z.B. 'maintenance'), sonst None.
|
||||
Basis für den Wartungs-Riegel: nur EIN System-Update gleichzeitig."""
|
||||
for j in JOBS.values():
|
||||
if j.get("group") == group and j.get("state") in ("running", "queued"):
|
||||
return j
|
||||
return None
|
||||
|
||||
|
||||
def cancel_job(job_id: str) -> bool:
|
||||
job = JOBS.get(job_id)
|
||||
if not job or job["state"] in ("done", "failed", "canceled"):
|
||||
return False
|
||||
job["canceled"] = True
|
||||
_append_log(job, "[mc] Abbruch angefordert…")
|
||||
proc = _PROCS.get(job_id)
|
||||
if proc is not None:
|
||||
try:
|
||||
proc.terminate()
|
||||
except Exception: # noqa: BLE001
|
||||
pass
|
||||
else:
|
||||
job["state"] = "canceled"
|
||||
job["finished_at"] = time.time()
|
||||
return True
|
||||
|
||||
|
||||
def public_jobs() -> list[dict]:
|
||||
"""Jobs ohne interne Felder (_on_done) für die API."""
|
||||
return [{k: v for k, v in j.items() if not k.startswith("_")} for j in JOBS.values()]
|
||||
@@ -0,0 +1,565 @@
|
||||
"""
|
||||
Engine-Service: liest/schreibt die llama-swap config.yaml und spricht die
|
||||
llama-swap-API. Portiert & erweitert aus Mission Control v1.
|
||||
|
||||
NEU in 2.0: `groups` für Ko-Residenz (schnell + schwer gleichzeitig geladen,
|
||||
`swap:false`) → Multi-Model-Delegation ohne Nachlade-Latenz.
|
||||
"""
|
||||
|
||||
import logging
|
||||
import os
|
||||
import re
|
||||
|
||||
import httpx
|
||||
from ruamel.yaml.scalarstring import LiteralScalarString
|
||||
|
||||
from config import (
|
||||
CMD_TEMPLATE, CONFIG_PATH, DEFAULT_TTL, DRAFTS_DIR, LLAMA_SWAP_URL,
|
||||
SPEC_DRAFT_MODEL_PATH, SPEC_DRAFT_N_MAX, SPEC_TYPE,
|
||||
)
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# Kanonische Serving-Rollen — EINE Quelle der Wahrheit (identisch zu sources.ROLE_IDS,
|
||||
# maintenance, frontend ModelBadges.ROLES). `hermes` = Lucys Agent-Hirn (warm + ko-resident
|
||||
# in der `brains`-Gruppe); UI-Label „Hirn".
|
||||
ROLE_IDS = {"fast", "heavy", "coder", "vision", "scout", "hermes"}
|
||||
|
||||
_CTX_RE = re.compile(r"-(?:c|-ctx-size)\s+(\d+)")
|
||||
_PATH_RE = re.compile(r"-(?:m|-model)\s+([^\s]+)")
|
||||
_QUANT_RE = re.compile(r"(Q\d_[A-Z0-9_]+|IQ\d_[A-Z0-9_]+|fp16|bf16)\.gguf", re.IGNORECASE)
|
||||
_SPLIT_RE = re.compile(r"-(\d+)-of-(\d+)\.gguf$", re.IGNORECASE)
|
||||
|
||||
|
||||
def _gguf_total_size(path: str) -> int | None:
|
||||
"""Gesamtgröße eines GGUF inkl. ALLER Split-Teile (…-00001-of-00003.gguf).
|
||||
Die Größe nur des ersten Teils ist bei Splits irreführend (oft nur ein Header)."""
|
||||
try:
|
||||
base = os.path.basename(path)
|
||||
m = _SPLIT_RE.search(base)
|
||||
if not m:
|
||||
return os.path.getsize(path)
|
||||
prefix, dirn = base[:m.start()], os.path.dirname(path)
|
||||
total = sum(os.path.getsize(os.path.join(dirn, f))
|
||||
for f in os.listdir(dirn)
|
||||
if f.startswith(prefix) and _SPLIT_RE.search(f))
|
||||
return total or os.path.getsize(path)
|
||||
except OSError:
|
||||
return None
|
||||
|
||||
|
||||
# --- Lesen -------------------------------------------------------------------
|
||||
def read_config() -> dict:
|
||||
if not CONFIG_PATH.exists():
|
||||
return {"models": {}}
|
||||
from ruamel.yaml import YAML
|
||||
r_yaml = YAML()
|
||||
r_yaml.preserve_quotes = True
|
||||
with CONFIG_PATH.open("r", encoding="utf-8") as f:
|
||||
data = r_yaml.load(f) or {}
|
||||
if not data.get("models"):
|
||||
data["models"] = {}
|
||||
return data
|
||||
|
||||
|
||||
def _parse_model(name: str, spec: dict) -> dict:
|
||||
spec = spec or {}
|
||||
cmd = str(spec.get("cmd", "")).strip()
|
||||
ctx = int(m.group(1)) if (m := _CTX_RE.search(cmd)) else None
|
||||
|
||||
path = filename = quant = ""
|
||||
size_bytes = None
|
||||
if (m := _PATH_RE.search(cmd)):
|
||||
path = m.group(1).replace("'", "").replace('"', "")
|
||||
filename = os.path.basename(path)
|
||||
if os.path.exists(path):
|
||||
size_bytes = _gguf_total_size(path)
|
||||
if (q := _QUANT_RE.search(path)):
|
||||
quant = q.group(1).upper()
|
||||
|
||||
aliases = spec.get("aliases") or []
|
||||
if isinstance(aliases, str):
|
||||
aliases = [aliases]
|
||||
aliases = [str(a) for a in aliases]
|
||||
role = aliases[0].lower() if aliases else (name.lower() if name.lower() in ROLE_IDS else None)
|
||||
|
||||
prompt_cache = "--prompt-cache " in cmd or cmd.endswith("--prompt-cache") or "--prompt-cache-all" in cmd
|
||||
# Draft-Modell erkennen — klassisch (--spec-draft-model) ODER MTP (--model-draft / -md).
|
||||
spec_draft = None
|
||||
if (m_draft := re.search(r"--(?:spec-draft-model|model-draft)\s+([^\s]+)", cmd)) \
|
||||
or (m_draft := re.search(r"(?<![\w-])-md\s+([^\s]+)", cmd)):
|
||||
spec_draft = os.path.basename(m_draft.group(1).replace("'", "").replace('"', ""))
|
||||
spec_type = None
|
||||
if (m_st := re.search(r"--spec-type\s+([^\s]+)", cmd)):
|
||||
spec_type = m_st.group(1)
|
||||
# Spec ist nur AKTIV, wenn BEIDES gesetzt ist (Draft-Modell UND --spec-type).
|
||||
spec_active = bool(spec_draft and spec_type)
|
||||
parallel_match = re.search(r"--parallel\s+(\d+)", cmd)
|
||||
parallel_slots = int(parallel_match.group(1)) if parallel_match else 1
|
||||
|
||||
from services.caps import capabilities
|
||||
return {
|
||||
"name": name,
|
||||
"role": role,
|
||||
"aliases": aliases,
|
||||
"api_ids": [name] + aliases,
|
||||
"ctx": ctx,
|
||||
"ttl": spec.get("ttl"),
|
||||
"cmd": cmd,
|
||||
"gguf_path": path,
|
||||
"filename": filename,
|
||||
"quant": quant,
|
||||
"size_bytes": size_bytes,
|
||||
"incomplete": not path,
|
||||
"prompt_cache": prompt_cache,
|
||||
"spec_draft_model": spec_draft,
|
||||
"spec_type": spec_type,
|
||||
"spec_active": spec_active,
|
||||
"parallel_slots": parallel_slots,
|
||||
"capabilities": capabilities(
|
||||
name=filename or name, cmd=cmd,
|
||||
gguf_path=(path if (path and os.path.exists(path)) else ""),
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
def list_models() -> list[dict]:
|
||||
cfg = read_config()
|
||||
return [_parse_model(name, spec) for name, spec in (cfg.get("models") or {}).items()]
|
||||
|
||||
|
||||
def engine_reachable() -> bool:
|
||||
try:
|
||||
with httpx.Client(timeout=3.0) as c:
|
||||
return c.get(f"{LLAMA_SWAP_URL}/v1/models").status_code == 200
|
||||
except Exception:
|
||||
return False
|
||||
|
||||
|
||||
# --- Schreiben ---------------------------------------------------------------
|
||||
def model_id_from_path(model_path: str) -> str:
|
||||
"""Sprechende Modell-ID (= API-Name) aus dem GGUF-Pfad: Repo-Ordnername ohne
|
||||
'-GGUF'. Fallback: Dateiname ohne Quant-Suffix.
|
||||
Split-GGUFs liegen oft in einem Quant-Unterordner (…/Q4_K_M/file-00001-of-…) →
|
||||
dann eine Ebene höher (Repo-Ordner) nehmen, sonst hieße das Modell 'Q4_K_M'."""
|
||||
d = os.path.basename(os.path.dirname(model_path))
|
||||
if re.fullmatch(r"(I?Q\d[\w]*|UD-Q\d[\w]*|F16|BF16|FP16|F32)", d, flags=re.I):
|
||||
d = os.path.basename(os.path.dirname(os.path.dirname(model_path)))
|
||||
name = re.sub(r"[-_]?GGUF$", "", d, flags=re.I).strip("-_")
|
||||
if not name:
|
||||
fn = re.sub(r"\.gguf$", "", os.path.basename(model_path), flags=re.I)
|
||||
fn = re.sub(r"-\d+-of-\d+$", "", fn)
|
||||
name = re.sub(r"[-_](Q\d[\w]*|IQ\d[\w]*|F16|BF16|FP16|F32)$", "", fn, flags=re.I)
|
||||
return name or "modell"
|
||||
|
||||
|
||||
def set_role_alias(cfg: dict, model_id: str, role: str | None) -> None:
|
||||
"""Rolle als eindeutigen llama-swap-`aliases`-Eintrag setzen (vorher bei allen
|
||||
anderen Modellen entfernen). role=None/leer entfernt den Alias."""
|
||||
models = cfg.get("models") or {}
|
||||
role = (role or "").strip().lower()
|
||||
if role:
|
||||
for mid, spec in models.items():
|
||||
if mid == model_id or not isinstance(spec, dict):
|
||||
continue
|
||||
al = [a for a in (spec.get("aliases") or []) if str(a).lower() != role]
|
||||
if al:
|
||||
spec["aliases"] = al
|
||||
else:
|
||||
spec.pop("aliases", None)
|
||||
spec = models.get(model_id)
|
||||
if isinstance(spec, dict):
|
||||
if role and role != model_id.lower():
|
||||
spec["aliases"] = [role]
|
||||
else:
|
||||
spec.pop("aliases", None)
|
||||
|
||||
|
||||
def _augment_vision(cmd: str, model_path: str, mmproj_path: str | None) -> str:
|
||||
"""Vision-Modelle brauchen --mmproj <projektor> und --jinja."""
|
||||
if mmproj_path:
|
||||
if "--mmproj" not in cmd:
|
||||
cmd += f" --mmproj {mmproj_path}"
|
||||
if "--jinja" not in cmd:
|
||||
cmd += " --jinja"
|
||||
return cmd
|
||||
|
||||
|
||||
def write_config(cfg: dict) -> None:
|
||||
"""Atomar schreiben (tmp + os.replace), damit llama-swap mit -watch-config nie
|
||||
eine halbe Datei sieht. Fehlende Schreibrechte → klare Meldung."""
|
||||
try:
|
||||
CONFIG_PATH.parent.mkdir(parents=True, exist_ok=True)
|
||||
tmp = CONFIG_PATH.with_name(CONFIG_PATH.name + ".tmp")
|
||||
from ruamel.yaml import YAML
|
||||
r_yaml = YAML()
|
||||
r_yaml.preserve_quotes = True
|
||||
with tmp.open("w", encoding="utf-8") as f:
|
||||
r_yaml.dump(cfg, f)
|
||||
os.replace(tmp, CONFIG_PATH)
|
||||
# llama-swap (-watch-config) lädt jetzt neu und verwirft dabei alle Modelle inkl. Hirn.
|
||||
# Den Re-Warm-Wächter anstoßen, damit das Hirn nicht bis zum nächsten Tick kalt liegt.
|
||||
try:
|
||||
from services import warmer
|
||||
warmer.nudge()
|
||||
except Exception: # noqa: BLE001 — Vorwärmen ist Komfort, nie ein Schreib-Blocker
|
||||
pass
|
||||
except PermissionError as exc:
|
||||
raise PermissionError(
|
||||
f"Mission Control darf '{CONFIG_PATH}' nicht schreiben. "
|
||||
f"Einmalig: sudo chown -R hitonabi:hitonabi {CONFIG_PATH.parent}"
|
||||
) from exc
|
||||
|
||||
|
||||
|
||||
def register_model(model_path: str, role: str | None = None, ctx: int = 8192,
|
||||
ttl: int | None = None, mmproj_path: str | None = None,
|
||||
jinja: bool = False) -> str:
|
||||
"""Ein GGUF als llama-swap-Modell eintragen (cmd + Rolle-Alias). Gibt die
|
||||
Modell-ID zurück. jinja=True erzwingt --jinja (Tool-Calling, z.B. fürs Agent-Hirn)."""
|
||||
cfg = read_config()
|
||||
model_id = model_id_from_path(model_path)
|
||||
cmd = CMD_TEMPLATE.replace("{model}", model_path).replace("{ctx}", str(ctx))
|
||||
cmd = _augment_vision(cmd, model_path, mmproj_path)
|
||||
if jinja and "--jinja" not in cmd:
|
||||
cmd += " --jinja"
|
||||
|
||||
role_lower = (role or "").strip().lower()
|
||||
# KV-Cache-Reuse über Turns (Prompt-Cache wiederverwenden) — hilft allen Chat-Modellen
|
||||
# (Agent-Hirn, Coding, Multi-Turn). Spiegelt die auf der Box bewährten Flags wider, damit
|
||||
# neu installierte Modelle nicht hinter dem hand-getunten Stand zurückbleiben (Drift-Fix).
|
||||
# NICHT bei Vision-Modellen: --cache-reuse + --mmproj bricht llama-server (live verifiziert,
|
||||
# deshalb fahren vision/scout auf der Box ohne cache-reuse).
|
||||
if "--cache-reuse" not in cmd and "--mmproj" not in cmd:
|
||||
cmd += " --cache-reuse 256 -cram 16384"
|
||||
# IDE-Coding profitiert von Nebenläufigkeit; sonst Default 1 Slot = voller Kontext/Anfrage
|
||||
# (--parallel teilt den Kontext HART auf die Slots auf, s. docs/OPTIMIZATION_PLAN.md §9.4 V6).
|
||||
if role_lower == "coder" and "--parallel" not in cmd:
|
||||
cmd += " --parallel 2"
|
||||
# Vocab-kompatiblen Draft automatisch anhängen — klassisch (DRAFTS_DIR) ODER MTP-Kopf neben
|
||||
# dem Modell. Self-guarding: ohne kompatiblen/vorhandenen Draft passiert nichts (später im UI
|
||||
# setzbar). Bei frischem Install existiert die Modell-GGUF noch nicht → ebenfalls kein Draft.
|
||||
if "--spec-draft-model" not in cmd and "--model-draft" not in cmd:
|
||||
cmd += spec_draft_flags(model_path)
|
||||
|
||||
cfg.setdefault("models", {})[model_id] = {
|
||||
"cmd": LiteralScalarString(cmd + "\n"),
|
||||
"ttl": ttl if ttl is not None else DEFAULT_TTL,
|
||||
}
|
||||
set_role_alias(cfg, model_id, role)
|
||||
write_config(cfg)
|
||||
return model_id
|
||||
|
||||
|
||||
# --- Speculative-Draft / Vocab-Kompatibilität --------------------------------
|
||||
def list_drafts() -> list[dict]:
|
||||
"""Alle Draft-GGUFs in DRAFTS_DIR mit Tokenizer-Fingerprint."""
|
||||
from services import gguf_meta
|
||||
out = []
|
||||
if DRAFTS_DIR.is_dir():
|
||||
for p in sorted(DRAFTS_DIR.glob("*.gguf")):
|
||||
out.append({
|
||||
"path": str(p), "filename": p.name,
|
||||
"size_bytes": p.stat().st_size if p.exists() else None,
|
||||
"vocab": gguf_meta.fingerprint(str(p)),
|
||||
})
|
||||
return out
|
||||
|
||||
|
||||
def _is_mtp_draft(draft_path: str) -> bool:
|
||||
"""Ist dieser Draft ein MTP-Kopf (Multi-Token-Prediction) statt eines klassischen
|
||||
Draft-Modells? MTP-Köpfe (z.B. gemma-4) laden mit `--model-draft … --spec-type
|
||||
draft-mtp` statt `--spec-draft-model … --spec-type draft-simple`. Erkennung am Arch
|
||||
('…-assistant' / 'mtp') oder Dateinamen ('mtp-*', '*-MTP', '*-assistant')."""
|
||||
base = os.path.basename(draft_path).lower()
|
||||
if base.startswith("mtp-") or "-mtp" in base or "assistant" in base:
|
||||
return True
|
||||
from services import gguf_meta
|
||||
arch = ((gguf_meta.fingerprint(draft_path) or {}).get("arch") or "").lower()
|
||||
return arch.endswith("-assistant") or "mtp" in arch
|
||||
|
||||
|
||||
def _spec_flags_for_draft(draft_path: str) -> str:
|
||||
"""Korrekte llama-server-Spec-Flags für einen (vocab-kompatiblen) Draft. MTP-Kopf →
|
||||
`--model-draft … --spec-type draft-mtp --spec-draft-n-max N`; klassischer Draft →
|
||||
`--spec-draft-model … --spec-type draft-simple`. (Beides nötig, sonst Spec inaktiv.)"""
|
||||
if _is_mtp_draft(draft_path):
|
||||
return (f" --model-draft {draft_path} --spec-type draft-mtp"
|
||||
f" --spec-draft-n-max {SPEC_DRAFT_N_MAX}")
|
||||
return f" --spec-draft-model {draft_path} --spec-type {SPEC_TYPE}"
|
||||
|
||||
|
||||
def _sibling_mtp_drafters(target_path: str) -> list[str]:
|
||||
"""MTP-Kopf-GGUFs NEBEN dem Zielmodell (gleicher Ordner): 'mtp-*.gguf', '*-MTP.gguf',
|
||||
'*-assistant*.gguf'. Per Konstruktion vocab-identisch zum Modell → idealer Draft."""
|
||||
out: list[str] = []
|
||||
d = os.path.dirname(target_path)
|
||||
if os.path.isdir(d):
|
||||
for f in sorted(os.listdir(d)):
|
||||
fl = f.lower()
|
||||
if fl.endswith(".gguf") and (fl.startswith("mtp-") or "-mtp" in fl or "assistant" in fl):
|
||||
p = os.path.join(d, f)
|
||||
if p != target_path:
|
||||
out.append(p)
|
||||
return out
|
||||
|
||||
|
||||
def find_compatible_draft(target_path: str) -> str | None:
|
||||
"""Pfad eines vocab-kompatiblen Drafts für target_path, oder None.
|
||||
Bevorzugt einen MTP-Kopf NEBEN dem Modell (höchste Qualität, by-construction),
|
||||
dann MC_SPEC_DRAFT_MODEL (falls gesetzt+kompatibel), sonst der erste kompatible
|
||||
Draft in DRAFTS_DIR. None auch, wenn target (noch) fehlt (nicht verifizierbar →
|
||||
bewusst KEIN Draft anhängen)."""
|
||||
if not target_path or not os.path.exists(target_path):
|
||||
return None
|
||||
from services import gguf_meta
|
||||
candidates: list[str] = list(_sibling_mtp_drafters(target_path))
|
||||
if SPEC_DRAFT_MODEL_PATH and os.path.exists(SPEC_DRAFT_MODEL_PATH):
|
||||
candidates.append(SPEC_DRAFT_MODEL_PATH)
|
||||
for d in list_drafts():
|
||||
if d["path"] not in candidates:
|
||||
candidates.append(d["path"])
|
||||
for c in candidates:
|
||||
if gguf_meta.compatible(target_path, c) is True:
|
||||
return c
|
||||
return None
|
||||
|
||||
|
||||
def spec_draft_flags(target_path: str) -> str:
|
||||
"""llama-server-Flags für Speculative Decoding (Draft + --spec-type), oder ''
|
||||
wenn kein kompatibler Draft existiert. MTP-bewusst (s. _spec_flags_for_draft)."""
|
||||
d = find_compatible_draft(target_path)
|
||||
return _spec_flags_for_draft(d) if d else ""
|
||||
|
||||
|
||||
def drafts_for(target_path: str) -> dict:
|
||||
"""Für die UI: alle Drafts + ihre Kompatibilität zum Ziel-Modell. Schließt MTP-Köpfe
|
||||
NEBEN dem Zielmodell ein (DRAFTS_DIR kennt sie nicht). `mtp:true` markiert MTP-Drafts.
|
||||
compatible=None heißt 'nicht prüfbar' (Ziel- oder Draft-GGUF fehlt)."""
|
||||
from services import gguf_meta
|
||||
exists = bool(target_path and os.path.exists(target_path))
|
||||
drafts = list_drafts()
|
||||
seen = {d["path"] for d in drafts}
|
||||
for p in _sibling_mtp_drafters(target_path):
|
||||
if p not in seen:
|
||||
drafts.append({"path": p, "filename": os.path.basename(p),
|
||||
"size_bytes": os.path.getsize(p) if os.path.exists(p) else None,
|
||||
"vocab": gguf_meta.fingerprint(p)})
|
||||
for d in drafts:
|
||||
d["compatible"] = gguf_meta.compatible(target_path, d["path"]) if exists else None
|
||||
d["mtp"] = _is_mtp_draft(d["path"])
|
||||
return {
|
||||
"target_path": target_path,
|
||||
"target_exists": exists,
|
||||
"target_vocab": gguf_meta.fingerprint(target_path) if exists else None,
|
||||
"drafts": drafts,
|
||||
}
|
||||
|
||||
|
||||
def set_spec_draft(model_id: str, draft_path: str | None) -> dict:
|
||||
"""Setzt (oder entfernt mit draft_path=None) den Spec-Draft eines Modells.
|
||||
Validiert die Vocab-Kompatibilität — ein inkompatibler/unprüfbarer Draft wird
|
||||
abgelehnt (idiotensicher). Returns {ok, reason}."""
|
||||
cfg = read_config()
|
||||
spec = (cfg.get("models") or {}).get(model_id)
|
||||
if not isinstance(spec, dict):
|
||||
return {"ok": False, "reason": "Modell nicht gefunden"}
|
||||
cmd = str(spec.get("cmd", ""))
|
||||
# vorhandene Spec-Flags entfernen (idempotent) — klassisch UND MTP.
|
||||
cmd = re.sub(r"\s+--(?:spec-draft-model|model-draft)\s+\S+", "", cmd)
|
||||
cmd = re.sub(r"\s+-md\s+\S+", "", cmd)
|
||||
cmd = re.sub(r"\s+--spec-type\s+\S+", "", cmd)
|
||||
cmd = re.sub(r"\s+--spec-draft-n-(?:max|min)\s+\S+", "", cmd)
|
||||
|
||||
if draft_path:
|
||||
# relative Angabe (nur Dateiname) gegen DRAFTS_DIR auflösen
|
||||
if not os.path.isabs(draft_path) and "/" not in draft_path:
|
||||
draft_path = str(DRAFTS_DIR / draft_path)
|
||||
if not os.path.exists(draft_path):
|
||||
return {"ok": False, "reason": "Draft-Datei nicht gefunden"}
|
||||
from services import gguf_meta
|
||||
target = ""
|
||||
if (mt := _PATH_RE.search(cmd)):
|
||||
target = mt.group(1).replace("'", "").replace('"', "")
|
||||
comp = gguf_meta.compatible(target, draft_path) if os.path.exists(target) else None
|
||||
if comp is not True:
|
||||
reason = ("Draft ist NICHT vocab-kompatibel zum Modell — Speculative Decoding "
|
||||
"würde beim Laden scheitern."
|
||||
if comp is False else
|
||||
"Kompatibilität nicht prüfbar (Modell-GGUF fehlt) — Draft nicht gesetzt.")
|
||||
return {"ok": False, "reason": reason}
|
||||
cmd = cmd.rstrip() + _spec_flags_for_draft(draft_path)
|
||||
|
||||
spec["cmd"] = LiteralScalarString(cmd.rstrip() + "\n")
|
||||
write_config(cfg)
|
||||
return {"ok": True, "reason": ""}
|
||||
|
||||
|
||||
# --- Groups (Ko-Residenz) ----------------------------------------------------
|
||||
def set_group(group: str, members: list[str], swap: bool = False, persist: bool = False) -> None:
|
||||
"""llama-swap-`groups`-Eintrag setzen. swap=False → alle Mitglieder dürfen
|
||||
GLEICHZEITIG laufen (Ko-Residenz, keine Nachlade-Latenz). persist=True →
|
||||
Mitglieder werden nie von anderen Gruppen verdrängt.
|
||||
|
||||
llama-swap kennt dafür AUSSCHLIESSLICH den Key `persistent` — `persist` wird von ihm
|
||||
stillschweigend ignoriert (so verlor das Hirn seinen Verdrängungsschutz; live gefunden
|
||||
03.07.2026: Coder-Last warf die brains-Gruppe raus). `persist` wird zusätzlich weiter
|
||||
geschrieben, weil MC2-API/UI (routers/models.py, agent.py, Frontend) diesen Key lesen."""
|
||||
cfg = read_config()
|
||||
groups = cfg.setdefault("groups", {})
|
||||
groups[group] = {"swap": swap, "persist": persist, "persistent": persist,
|
||||
"members": list(members)}
|
||||
write_config(cfg)
|
||||
|
||||
|
||||
def list_groups() -> dict:
|
||||
return read_config().get("groups") or {}
|
||||
|
||||
|
||||
def set_role(model_id: str, role: str | None) -> bool:
|
||||
"""Rolle (llama-swap-Alias) eines bestehenden Modells setzen/ändern. So tauscht man
|
||||
z.B. das `fast`-Hirn: Rolle `fast` auf ein anderes Modell legen (Alias wandert)."""
|
||||
cfg = read_config()
|
||||
if model_id not in (cfg.get("models") or {}):
|
||||
return False
|
||||
set_role_alias(cfg, model_id, role)
|
||||
write_config(cfg)
|
||||
return True
|
||||
|
||||
|
||||
def set_ctx(model_id: str, ctx: int) -> bool:
|
||||
"""Kontextlänge (-c) eines bestehenden Modells ändern."""
|
||||
cfg = read_config()
|
||||
spec = (cfg.get("models") or {}).get(model_id)
|
||||
if not spec:
|
||||
return False
|
||||
cmd = str(spec.get("cmd", ""))
|
||||
if _CTX_RE.search(cmd):
|
||||
cmd = re.sub(r"-(?:c|-ctx-size)\s+\d+", f"-c {ctx}", cmd)
|
||||
else:
|
||||
cmd = cmd.rstrip() + f" -c {ctx}"
|
||||
spec["cmd"] = LiteralScalarString(cmd if cmd.endswith("\n") else cmd + "\n")
|
||||
write_config(cfg)
|
||||
return True
|
||||
|
||||
|
||||
def set_ttl(model_id: str, ttl: int) -> bool:
|
||||
"""Idle-TTL (Sekunden) eines bestehenden Modells setzen. ttl=0 → nie automatisch
|
||||
entladen (für das Agent-Hirn, das dauerhaft warm bleiben muss)."""
|
||||
cfg = read_config()
|
||||
spec = (cfg.get("models") or {}).get(model_id)
|
||||
if not isinstance(spec, dict):
|
||||
return False
|
||||
spec["ttl"] = int(ttl)
|
||||
write_config(cfg)
|
||||
return True
|
||||
|
||||
|
||||
def delete_model(model_id: str) -> bool:
|
||||
"""Entfernt einen Modell-Eintrag aus der config.yaml, löscht die zugehörigen
|
||||
GGUF-Dateien (auch Splits) vom Datenträger und bereinigt leere Ordner.
|
||||
"""
|
||||
cfg = read_config()
|
||||
models = cfg.get("models") or {}
|
||||
if model_id not in models:
|
||||
return False
|
||||
|
||||
model_spec = models[model_id] or {}
|
||||
cmd = str(model_spec.get("cmd", "")).strip()
|
||||
if (m := _PATH_RE.search(cmd)):
|
||||
path = m.group(1).replace("'", "").replace('"', "")
|
||||
if path:
|
||||
# 1. Haupt-GGUF-Datei löschen
|
||||
if os.path.exists(path):
|
||||
try:
|
||||
os.remove(path)
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
# 2. Split-GGUF-Teile löschen (z.B. dateiname-00001-of-00005.gguf etc.)
|
||||
dirname = os.path.dirname(path)
|
||||
basename = os.path.basename(path)
|
||||
if os.path.isdir(dirname):
|
||||
split_idx = basename.find("-00001-of-")
|
||||
if split_idx != -1:
|
||||
prefix = basename[:split_idx]
|
||||
for f in os.listdir(dirname):
|
||||
if f.startswith(prefix) and f.endswith(".gguf"):
|
||||
try:
|
||||
os.remove(os.path.join(dirname, f))
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
# mmproj-Datei (Vision adapter) aus dem Befehl parsen & löschen
|
||||
if "mmproj" in cmd:
|
||||
mmproj_match = re.search(r'--mmproj\s+[\'"]?([^\s\'"]+)[\'"]?', cmd)
|
||||
if mmproj_match:
|
||||
m_path = mmproj_match.group(1)
|
||||
if os.path.exists(m_path):
|
||||
try:
|
||||
os.remove(m_path)
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
# 3. Eltern-Ordner löschen, falls er leer ist und nicht der Modelle-Wurzelordner selbst ist
|
||||
try:
|
||||
if not os.listdir(dirname) and os.path.basename(dirname) != "models":
|
||||
os.rmdir(dirname)
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
del models[model_id]
|
||||
for g in (cfg.get("groups") or {}).values():
|
||||
if isinstance(g, dict) and model_id in (g.get("members") or []):
|
||||
g["members"] = [m for m in g["members"] if m != model_id]
|
||||
write_config(cfg)
|
||||
return True
|
||||
|
||||
|
||||
def brain_model_name() -> str | None:
|
||||
"""Modellname von Lucys Agent-Hirn. Bevorzugt das Modell mit dem 'hermes'-Alias/-Rolle;
|
||||
fällt auf Hermes' aktives `model.default` zurück (deckt den Fall ab, dass die Config direkt
|
||||
auf einen Modellnamen statt den Alias zeigt)."""
|
||||
models = list_models()
|
||||
for m in models:
|
||||
names = {str(a).lower() for a in (m.get("aliases") or [])}
|
||||
if m.get("role"):
|
||||
names.add(str(m["role"]).lower())
|
||||
if "hermes" in names:
|
||||
return m["name"]
|
||||
# Fallback: das real von Hermes genutzte Hirn (model.default), per Alias/Name auflösen.
|
||||
try:
|
||||
from services.agent import _active_brain_name
|
||||
brain = (_active_brain_name() or "").lower()
|
||||
if brain and brain != "auto":
|
||||
cur = next((m for m in models if (m.get("role") or "").lower() == brain), None) \
|
||||
or next((m for m in models if brain in (m["name"] or "").lower()), None)
|
||||
if cur:
|
||||
return cur["name"]
|
||||
except Exception:
|
||||
log.debug("brain_model_name: Hermes-Fallback fehlgeschlagen", exc_info=True)
|
||||
return None
|
||||
|
||||
|
||||
def brain_status() -> dict:
|
||||
"""Ist Lucys Agent-Hirn (Rolle 'hermes') WIRKLICH geladen & bereit? Prüft /running — ein
|
||||
abgestürztes Modell (z.B. OOM/Crash nach Engine-Update) erscheint dort NICHT als running.
|
||||
Fängt damit den Fall 'Engine erreichbar, aber Hirn tot', den engine_reachable() nicht sieht."""
|
||||
name = brain_model_name()
|
||||
running = get_running_models()
|
||||
return {"role": "hermes", "model": name, "ready": bool(name and name in running)}
|
||||
|
||||
|
||||
def get_running_models() -> list[str]:
|
||||
"""Fragt den /running Endpunkt von llama-swap ab. Gibt die Namen der geladenen
|
||||
Modelle zurück. Neuere llama-swap-Versionen liefern Objekte ({model, state, ...})
|
||||
statt Strings — beide Formen werden auf Namens-Strings normalisiert."""
|
||||
try:
|
||||
with httpx.Client(timeout=2.0) as c:
|
||||
r = c.get(f"{LLAMA_SWAP_URL}/running")
|
||||
if r.status_code == 200:
|
||||
data = r.json().get("running") or []
|
||||
return [x.get("model", "") if isinstance(x, dict) else x for x in data]
|
||||
except Exception:
|
||||
log.warning("get_running_models fehlgeschlagen", exc_info=True)
|
||||
return []
|
||||
@@ -0,0 +1,707 @@
|
||||
"""
|
||||
Wartung: Updates (OS/Engine/Modelle), Dienst-Neustart (system- vs user-aware),
|
||||
Reboot, Logs. Portiert/modernisiert aus Mission Control v1 (routers/maintenance.py).
|
||||
|
||||
Passwortfrei über NOPASSWD-Whitelist (sudo -n). OS-Update/Reboot brauchen einmalig
|
||||
erweiterte sudoers (siehe docs/BEDIENUNG.md). Lange Ops laufen als jobengine-Job.
|
||||
"""
|
||||
|
||||
import os
|
||||
import re
|
||||
import subprocess
|
||||
import time
|
||||
from datetime import datetime
|
||||
|
||||
import httpx
|
||||
import psutil
|
||||
|
||||
from services import catalog, discover, jobengine, llamaswap, system
|
||||
|
||||
# System-Dienste (root, via sudo -n NOPASSWD) vs. User-Dienste (systemctl --user).
|
||||
SYSTEM_SERVICES = {"llama-swap"}
|
||||
USER_SERVICES = {"mission-control-2", "hermes-gateway", "hermes-terminal", "mem0-service", "voice-service"}
|
||||
|
||||
# Engine-Update: lädt den neuesten Vulkan-Build (deploy/update-engine.sh, läuft als root).
|
||||
_REPO_ROOT = os.path.abspath(os.path.join(os.path.dirname(__file__), "..", ".."))
|
||||
ENGINE_UPDATE_CMD = os.environ.get(
|
||||
"MC_ENGINE_UPDATE_CMD", f"sudo bash {_REPO_ROOT}/deploy/update-engine.sh")
|
||||
# Engine = offizieller Vulkan-Build von ggml-org/llama.cpp (RADV auf Strix Halo).
|
||||
ENGINE_PATH = os.environ.get("MC_ENGINE_PATH", "/opt/llamacpp-vulkan")
|
||||
ENGINE_REPO = os.environ.get("MC_ENGINE_REPO", "ggml-org/llama.cpp")
|
||||
_engine_cache = {"ts": 0.0, "avail": False}
|
||||
|
||||
# Router = llama-swap (mostlygeek): proxyt Anfragen und wechselt die Modelle heiß. Eigenes
|
||||
# Upstream-Projekt mit eigenem Release-Zyklus → getrennt von der Engine geführt.
|
||||
SWAP_UPDATE_CMD = os.environ.get(
|
||||
"MC_SWAP_UPDATE_CMD", f"sudo bash {_REPO_ROOT}/deploy/update-swap.sh")
|
||||
SWAP_BIN = os.environ.get("MC_SWAP_BIN", "/usr/local/bin/llama-swap")
|
||||
SWAP_REPO = os.environ.get("MC_SWAP_REPO", "mostlygeek/llama-swap")
|
||||
_swap_cache = {"ts": 0.0, "avail": False}
|
||||
|
||||
# Stack-Funktionsprüfung NACH jedem Update (OS/Engine/Router): verifiziert per echter
|
||||
# Inferenz, dass der Stack noch läuft → Job wird rot, wenn ein Update etwas zerschossen hat.
|
||||
STACK_POSTCHECK = os.path.join(_REPO_ROOT, "deploy", "stack-postcheck.sh")
|
||||
|
||||
|
||||
def _installed_engine_build() -> int | None:
|
||||
"""Build-Nummer der installierten llama-server-Binary (z.B. 9821), oder None.
|
||||
Vulkan-Build braucht LD_LIBRARY_PATH=ENGINE_PATH zum Start von --version."""
|
||||
bin_path = os.path.join(ENGINE_PATH, "llama-server")
|
||||
if not os.path.exists(bin_path):
|
||||
return None
|
||||
try:
|
||||
env = dict(os.environ, LD_LIBRARY_PATH=ENGINE_PATH)
|
||||
out = subprocess.run([bin_path, "--version"], capture_output=True, text=True,
|
||||
timeout=20, env=env)
|
||||
txt = (out.stderr or "") + (out.stdout or "")
|
||||
# Formate je nach Build: "version: 9821 (hash)" (aktuell), "build: <hash> (9821)", "b9821".
|
||||
if (m := re.search(r"version:\s*(\d{3,})", txt)) \
|
||||
or (m := re.search(r"build:\s*\S+\s*\((\d+)\)", txt)) \
|
||||
or (m := re.search(r"\bb(\d{3,})\b", txt)):
|
||||
return int(m.group(1))
|
||||
except Exception:
|
||||
return None
|
||||
return None
|
||||
|
||||
|
||||
def _ram_gb() -> float:
|
||||
return psutil.virtual_memory().total / (1024 ** 3)
|
||||
|
||||
|
||||
def _os_upgradable() -> int:
|
||||
try:
|
||||
# LC_ALL=C erzwingt englische apt-Ausgabe ("[upgradable from: ...]") — sonst zählt
|
||||
# grep auf einer deutschen Box ("[aktualisierbar von:]") nichts und meldet faelschlich 0.
|
||||
out = subprocess.run(
|
||||
["bash", "-c", "LC_ALL=C apt list --upgradable 2>/dev/null | grep -c upgradable || true"],
|
||||
capture_output=True, text=True, timeout=10)
|
||||
return int((out.stdout or "0").strip() or 0)
|
||||
except Exception:
|
||||
return 0
|
||||
|
||||
|
||||
def _engine_update_available() -> bool:
|
||||
now = time.time()
|
||||
if now - _engine_cache["ts"] < 3600:
|
||||
return _engine_cache["avail"]
|
||||
avail = False
|
||||
try:
|
||||
rel = httpx.get(f"https://api.github.com/repos/{ENGINE_REPO}/releases/latest",
|
||||
timeout=6, headers={"User-Agent": "MissionControl2"}).json()
|
||||
tag = str(rel.get("tag_name", ""))
|
||||
latest = int(m.group(1)) if (m := re.search(r"(\d{3,})", tag)) else None
|
||||
installed = _installed_engine_build()
|
||||
if latest is not None and installed is not None:
|
||||
avail = latest > installed # präziser Build-Nummer-Vergleich
|
||||
else: # Fallback: Release-Datum vs. Engine-mtime
|
||||
pub = datetime.fromisoformat(rel["published_at"].replace("Z", "+00:00")).timestamp()
|
||||
avail = pub > os.path.getmtime(ENGINE_PATH) + 86400
|
||||
except Exception:
|
||||
avail = False
|
||||
_engine_cache.update(ts=now, avail=avail)
|
||||
return avail
|
||||
|
||||
|
||||
def _installed_swap_version() -> int | None:
|
||||
"""Versions-Nummer der installierten llama-swap-Binary (z.B. 228), oder None."""
|
||||
if not os.path.exists(SWAP_BIN):
|
||||
return None
|
||||
try:
|
||||
out = subprocess.run([SWAP_BIN, "--version"], capture_output=True, text=True, timeout=15)
|
||||
txt = (out.stdout or "") + (out.stderr or "")
|
||||
if (m := re.search(r"version:\s*(\d+)", txt)) or (m := re.search(r"\bv?(\d{2,})\b", txt)):
|
||||
return int(m.group(1))
|
||||
except Exception:
|
||||
return None
|
||||
return None
|
||||
|
||||
|
||||
def _swap_update_available() -> bool:
|
||||
now = time.time()
|
||||
if now - _swap_cache["ts"] < 3600:
|
||||
return _swap_cache["avail"]
|
||||
avail = False
|
||||
try:
|
||||
rel = httpx.get(f"https://api.github.com/repos/{SWAP_REPO}/releases/latest",
|
||||
timeout=6, headers={"User-Agent": "MissionControl2"}).json()
|
||||
tag = str(rel.get("tag_name", ""))
|
||||
latest = int(m.group(1)) if (m := re.search(r"(\d{2,})", tag)) else None
|
||||
installed = _installed_swap_version()
|
||||
if latest is not None and installed is not None:
|
||||
avail = latest > installed
|
||||
except Exception:
|
||||
avail = False
|
||||
_swap_cache.update(ts=now, avail=avail)
|
||||
return avail
|
||||
|
||||
|
||||
_comp_cache = {"ts": 0.0, "data": []}
|
||||
|
||||
|
||||
def _hermes_agent_update() -> dict:
|
||||
"""Hermes-Agent wird aus **git** aktualisiert (CLI `hermes update` = git pull origin <branch>).
|
||||
Darum HEAD vs. origin/<branch> prüfen (fetch + behind-count) — NICHT GitHub-Releases: die
|
||||
werden selten getaggt, main läuft ihnen voraus → sonst zeigt das UI nie ein Update an."""
|
||||
info = {"key": "hermes_agent", "name": "Hermes Agent", "current": None,
|
||||
"latest": None, "update": False, "reachable": None}
|
||||
git = system.find_hermes_agent_git()
|
||||
if not git or not git.get("path"):
|
||||
return info
|
||||
path = git["path"]
|
||||
info["current"] = git.get("hash")
|
||||
try:
|
||||
branch = (subprocess.run(["git", "-C", path, "rev-parse", "--abbrev-ref", "HEAD"],
|
||||
capture_output=True, text=True, timeout=8).stdout.strip() or "main")
|
||||
fetch = subprocess.run(["git", "-C", path, "fetch", "-q", "origin", branch],
|
||||
capture_output=True, text=True, timeout=25)
|
||||
info["reachable"] = (fetch.returncode == 0)
|
||||
if fetch.returncode == 0:
|
||||
cnt = subprocess.run(["git", "-C", path, "rev-list", "--count", f"HEAD..origin/{branch}"],
|
||||
capture_output=True, text=True, timeout=8)
|
||||
behind = int(cnt.stdout.strip() or "0") if cnt.returncode == 0 else 0
|
||||
info["behind"] = behind
|
||||
info["update"] = behind > 0
|
||||
oh = subprocess.run(["git", "-C", path, "rev-parse", "--short", f"origin/{branch}"],
|
||||
capture_output=True, text=True, timeout=8).stdout.strip()
|
||||
info["latest"] = (f"{oh} ({behind} neu)" if behind else (oh or info["current"]))
|
||||
except Exception:
|
||||
info["reachable"] = False
|
||||
return info
|
||||
|
||||
|
||||
def _components_cached() -> list[dict]:
|
||||
"""Update-Status von Hermes-Agent (1h-Cache → GitHub schonen)."""
|
||||
now = time.time()
|
||||
if now - _comp_cache["ts"] < 3600 and _comp_cache["data"]:
|
||||
return _comp_cache["data"]
|
||||
data = [_hermes_agent_update()]
|
||||
_comp_cache.update(ts=now, data=data)
|
||||
return data
|
||||
|
||||
|
||||
def _params_of(m: dict) -> float:
|
||||
"""Größen-bewusste Parameterzahl eines installierten Modells: max aus Namens-Schätzung
|
||||
und Dateigröße (fängt namenlose wie 'Qwen3-Coder-Next' UND Split-GGUFs ab)."""
|
||||
from services.fit import QUANT_BYTES_PER_PARAM
|
||||
caps = m.get("capabilities") or {}
|
||||
bpp = QUANT_BYTES_PER_PARAM.get((m.get("quant") or "Q4_K_M").upper(), 0.55)
|
||||
size_gb = (m.get("size_bytes") or 0) / (1024 ** 3)
|
||||
pb_size = (size_gb / bpp) if size_gb > 1.0 else 0.0
|
||||
return max(float(caps.get("params_b") or 0), pb_size, 0.0)
|
||||
|
||||
|
||||
# Familien-Subtyp + Generations-Version aus dem Modellnamen (für „echtes Upgrade?").
|
||||
_FAM_PATS = (("qwen", r"qwen(\d+(?:\.\d+)?)"), ("gemma", r"gemma[-_ ]?(\d+(?:\.\d+)?)"),
|
||||
("llama", r"llama[-_ ]?(\d+(?:\.\d+)?)"), ("phi", r"phi[-_ ]?(\d+(?:\.\d+)?)"),
|
||||
("mistral", r"mistral"), ("hermes", r"hermes[-_ ]?(\d+(?:\.\d+)?)"))
|
||||
|
||||
|
||||
def _gen_key(name: str):
|
||||
"""(Familie+Subtyp, Generations-Version) oder None. Z.B. 'Qwen3-VL-2B' → ('qwen-vl', 3.0),
|
||||
'Qwen2.5-VL-7B' → ('qwen-vl', 2.5). Nur gleiche Familie ist sinnvoll vergleichbar."""
|
||||
low = (name or "").lower()
|
||||
sub = "-vl" if any(k in low for k in ("-vl", "vl-", "vision", "llava", "pixtral")) else \
|
||||
"-coder" if ("coder" in low or "-code" in low) else ""
|
||||
for fam, pat in _FAM_PATS:
|
||||
m = re.search(pat, low)
|
||||
if m:
|
||||
ver = float(m.group(1)) if (m.groups() and m.group(1)) else 0.0
|
||||
return (fam + sub, ver)
|
||||
return None
|
||||
|
||||
|
||||
def _meta(name: str, model_dict: dict | None = None, im: dict | None = None) -> dict:
|
||||
"""Metadaten (family, gen, total, active, moe) — bevorzugt den kuratierten Katalog,
|
||||
sonst die Felder eines Discover-/Modell-Dicts, sonst Namens-/Größen-Heuristik."""
|
||||
cm = catalog.meta_for_name(name)
|
||||
if cm:
|
||||
return {"family": cm.get("family"), "gen": cm.get("generation"),
|
||||
"total": float(cm.get("total_params_b") or 0),
|
||||
"active": cm.get("active_params_b"), "moe": bool(cm.get("moe"))}
|
||||
d = model_dict or {}
|
||||
g = _gen_key(name)
|
||||
total = float(d.get("params_b") or 0) or (_params_of(im) if im else 0.0)
|
||||
return {"family": (d.get("family") or (g[0] if g else None)),
|
||||
"gen": (d.get("generation") if d.get("generation") is not None else (g[1] if g else None)),
|
||||
"total": total, "active": d.get("active_b"), "moe": bool(d.get("moe"))}
|
||||
|
||||
|
||||
def model_upgrades() -> list[dict]:
|
||||
"""Je Rolle ein ECHTES Upgrade — nur wenn die Empfehlung wirklich besser ist:
|
||||
gleiche Familie UND (neuere Generation ODER deutlich größer) UND kein Tempo-Downgrade
|
||||
(MoE-first für die bandbreiten-limitierte Box: dense ersetzt MoE nur bei großem Wissens-
|
||||
Sprung). Metadaten kommen aus dem kuratierten Katalog → keine Namens-Raterei."""
|
||||
disc = discover.safe_discover(_ram_gb())
|
||||
if not disc:
|
||||
return []
|
||||
|
||||
installed = llamaswap.list_models()
|
||||
inst_by_role = {m["role"]: m for m in installed if m.get("role")}
|
||||
cmds = " ".join(str(s.get("cmd", "")).lower()
|
||||
for s in (llamaswap.read_config().get("models") or {}).values())
|
||||
out = []
|
||||
|
||||
for c in disc.get("categories", []):
|
||||
role = c["role"]
|
||||
im = inst_by_role.get(role)
|
||||
if im is None:
|
||||
continue
|
||||
rec = c.get("recommended")
|
||||
if not rec:
|
||||
continue
|
||||
rec_model = next((x for x in c.get("models", []) if x.get("repo") == rec), None)
|
||||
|
||||
i = _meta(im["name"], im=im)
|
||||
r = _meta(rec, model_dict=rec_model)
|
||||
|
||||
if not i["family"] or not r["family"] or i["family"] != r["family"]:
|
||||
continue # andere/unbekannte Familie → kein Upgrade
|
||||
if r["gen"] is not None and i["gen"] is not None and r["gen"] < i["gen"] - 1e-6:
|
||||
continue # ältere Generation → niemals
|
||||
same_gen = (r["gen"] is None or i["gen"] is None or abs(r["gen"] - i["gen"]) < 1e-6)
|
||||
if same_gen:
|
||||
if r["total"] and i["total"] and r["total"] < i["total"] * 1.05:
|
||||
continue # gleiche Gen, nicht größer → kein Upgrade
|
||||
# MoE-first: ein MoE durch dense ersetzen nur bei deutlichem Wissens-Sprung
|
||||
if i["moe"] and not r["moe"] and r["total"] < i["total"] * 1.5:
|
||||
continue
|
||||
# Tempo nicht verschlechtern (aktive Params), außer großer Wissens-Gewinn
|
||||
ia, ra = (i["active"] or i["total"]), (r["active"] or r["total"])
|
||||
if ia and ra > ia * 1.3 and r["total"] < i["total"] * 1.3:
|
||||
continue
|
||||
|
||||
base = rec.split("/")[-1].lower()
|
||||
stem = base[:-5] if base.endswith("-gguf") else base
|
||||
if base in cmds or (stem and stem in cmds):
|
||||
continue # schon installiert
|
||||
out.append({"role": role, "title": c["title"], "repo": rec})
|
||||
return out
|
||||
|
||||
|
||||
def _last_apt_update() -> float | None:
|
||||
for path in ["/var/lib/apt/periodic/update-success-stamp", "/var/cache/apt/pkgcache.bin"]:
|
||||
if os.path.exists(path):
|
||||
try:
|
||||
return os.path.getmtime(path)
|
||||
except Exception:
|
||||
pass
|
||||
return None
|
||||
|
||||
|
||||
def updates() -> dict:
|
||||
ups = model_upgrades()
|
||||
return {"os": _os_upgradable(), "engine": 1 if _engine_update_available() else 0,
|
||||
"swap": 1 if _swap_update_available() else 0,
|
||||
"models": len(ups), "model_list": ups, "last_check": _last_apt_update(),
|
||||
"components": _components_cached()}
|
||||
|
||||
|
||||
# ── Update-Details (was genau wird aktualisiert) — on-demand beim Öffnen des Fensters ──
|
||||
|
||||
def _os_held_back() -> list[dict]:
|
||||
"""Pakete, die apt aktuell NICHT einspielt, obwohl es sie gäbe — mit ehrlichem Grund.
|
||||
Zwei Fälle, aus `apt-get -s upgrade` (Simulation) gelesen:
|
||||
• Phasen-Rollout (Ubuntu staffelt Updates prozentual aus) → 'deferred due to phasing'.
|
||||
• zurückgehalten wegen neuer Abhängigkeiten → 'kept back'.
|
||||
Beides ist normal und löst sich von selbst — verhindert nur das 'hängt fest'-Gefühl,
|
||||
wenn nach 'Fertig' noch aktualisierbare Pakete übrig scheinen."""
|
||||
held: list[dict] = []
|
||||
sections = {
|
||||
"The following upgrades have been deferred due to phasing:": "phasing",
|
||||
"The following packages have been kept back:": "kept_back",
|
||||
}
|
||||
try:
|
||||
out = subprocess.run(
|
||||
["bash", "-c", "LC_ALL=C apt-get -s upgrade 2>/dev/null"],
|
||||
capture_output=True, text=True, timeout=25)
|
||||
reason: str | None = None
|
||||
for line in (out.stdout or "").splitlines():
|
||||
hdr = sections.get(line.strip())
|
||||
if hdr: # Abschnitts-Kopf → folgende Zeilen sammeln
|
||||
reason = hdr
|
||||
continue
|
||||
if reason and line.startswith((" ", "\t")): # eingerückt = Paketnamen des Abschnitts
|
||||
for name in line.split():
|
||||
held.append({"name": name, "reason": reason})
|
||||
elif reason: # nicht eingerückt → Abschnitt zu Ende
|
||||
reason = None
|
||||
except Exception: # noqa: BLE001 — nur Zusatzinfo, nie ein Blocker
|
||||
pass
|
||||
held.sort(key=lambda p: p["name"])
|
||||
return held
|
||||
|
||||
|
||||
def os_update_details() -> dict:
|
||||
"""Liste der aktualisierbaren apt-Pakete (Name, installiert → Kandidat) + ehrliche
|
||||
Anzeige der vom System zurückgestellten Pakete (Phasen-Rollout / kept back)."""
|
||||
out_pkgs: list[dict] = []
|
||||
try:
|
||||
# LC_ALL=C → englische Ausgabe, damit der Regex "[upgradable from: ...]" greift
|
||||
# (deutsche Box meldet sonst "[aktualisierbar von:]" und die Liste bliebe leer).
|
||||
out = subprocess.run(["bash", "-c", "LC_ALL=C apt list --upgradable 2>/dev/null"],
|
||||
capture_output=True, text=True, timeout=20)
|
||||
for line in (out.stdout or "").splitlines():
|
||||
# Format: name/repo neue_version arch [upgradable from: alte_version]
|
||||
m = re.match(r"^([^/\s]+)/\S+\s+(\S+)\s+\S+\s+\[upgradable from:\s*([^\]]+)\]",
|
||||
line.strip())
|
||||
if m:
|
||||
out_pkgs.append({"name": m.group(1), "candidate": m.group(2),
|
||||
"current": m.group(3).strip()})
|
||||
out_pkgs.sort(key=lambda p: p["name"])
|
||||
return {"kind": "os", "count": len(out_pkgs), "packages": out_pkgs,
|
||||
"held_back": _os_held_back()}
|
||||
except Exception as exc: # noqa: BLE001
|
||||
return {"kind": "os", "count": len(out_pkgs), "packages": out_pkgs, "error": str(exc)}
|
||||
|
||||
|
||||
def engine_update_details() -> dict:
|
||||
"""Installierte vs. neueste Engine-Build-Nummer + Release-Name/-Notizen/-Link."""
|
||||
info: dict = {"kind": "engine", "installed_build": _installed_engine_build(),
|
||||
"latest_build": None, "latest_tag": None, "name": None,
|
||||
"url": None, "body": None}
|
||||
try:
|
||||
rel = httpx.get(f"https://api.github.com/repos/{ENGINE_REPO}/releases/latest",
|
||||
timeout=8, headers={"User-Agent": "MissionControl2"}).json()
|
||||
tag = str(rel.get("tag_name", ""))
|
||||
info["latest_tag"] = tag
|
||||
info["latest_build"] = int(m.group(1)) if (m := re.search(r"(\d{3,})", tag)) else None
|
||||
info["name"] = rel.get("name") or tag
|
||||
info["url"] = rel.get("html_url")
|
||||
body = (rel.get("body") or "").strip()
|
||||
info["body"] = body[:2000] if body else None
|
||||
# Nur zusammenfassen, wenn wirklich ein neuerer Build ansteht (spart einen LLM-Call,
|
||||
# wenn das Modal bei aktuellem Stand geöffnet wird).
|
||||
if body and info["latest_build"] and info["installed_build"] \
|
||||
and info["latest_build"] > info["installed_build"]:
|
||||
ctx = ("Es geht um ein Update der Inferenz-Engine llama.cpp (Vulkan-Build, treibt alle "
|
||||
"Sprachmodelle der Box auf der AMD-Strix-Halo-GPU).")
|
||||
info.update(_summarize_release("engine", tag, ctx, body[:6000]))
|
||||
except Exception as exc: # noqa: BLE001
|
||||
info["error"] = str(exc)
|
||||
return info
|
||||
|
||||
|
||||
def swap_update_details() -> dict:
|
||||
"""Installierte vs. neueste llama-swap-Version + Release-Name/-Notizen/-Link."""
|
||||
info: dict = {"kind": "swap", "installed_build": _installed_swap_version(),
|
||||
"latest_build": None, "latest_tag": None, "name": None,
|
||||
"url": None, "body": None}
|
||||
try:
|
||||
rel = httpx.get(f"https://api.github.com/repos/{SWAP_REPO}/releases/latest",
|
||||
timeout=8, headers={"User-Agent": "MissionControl2"}).json()
|
||||
tag = str(rel.get("tag_name", ""))
|
||||
info["latest_tag"] = tag
|
||||
info["latest_build"] = int(m.group(1)) if (m := re.search(r"(\d{2,})", tag)) else None
|
||||
info["name"] = rel.get("name") or tag
|
||||
info["url"] = rel.get("html_url")
|
||||
body = (rel.get("body") or "").strip()
|
||||
info["body"] = body[:2000] if body else None
|
||||
if body and info["latest_build"] and info["installed_build"] \
|
||||
and info["latest_build"] > info["installed_build"]:
|
||||
ctx = ("Es geht um ein Update von llama-swap (der Router, der Anfragen an die Box "
|
||||
"verteilt und Sprachmodelle heiß nachlädt).")
|
||||
info.update(_summarize_release("swap", tag, ctx, body[:6000]))
|
||||
except Exception as exc: # noqa: BLE001
|
||||
info["error"] = str(exc)
|
||||
return info
|
||||
|
||||
|
||||
# LLM-Zusammenfassung anstehender Updates in Lucys Stimme (die Box liest ihre Release-Notes
|
||||
# selbst). Jede Zusammenfassung beginnt mit einem klaren Aktions-Verdikt für den Besitzer
|
||||
# (kein Entwickler): "Musst du etwas tun? NEIN — das Fangnetz regelt das / JA: …".
|
||||
# Gecacht je Komponente auf den neuesten Commit-Hash/Release-Tag — das Modal darf beliebig
|
||||
# oft geöffnet werden, ohne das Modell jedes Mal neu zu befragen.
|
||||
_relnotes_caches: dict = {"hermes": {}, "engine": {}, "swap": {}}
|
||||
|
||||
# Erste Zeile jeder Modell-Antwort: "AKTION: NEIN" oder "AKTION: JA — <grund>".
|
||||
# Toleriert Markdown-Deko (**fett**, Bullet, Überschrift), die 'fast' gern einstreut.
|
||||
_ACTION_RX = re.compile(
|
||||
r"^[\s*_>#\-]*AKTION:?\s*\**\s*(JA|NEIN)\b[\s—:,.\-*_]*(.*?)[\s*_]*$", re.IGNORECASE)
|
||||
|
||||
|
||||
def _summarize_release(kind: str, key: str, context: str, changes: str) -> dict:
|
||||
"""Fasst Release-Notes/Commits in Lucys Stimme zusammen und trennt das Aktions-Verdikt
|
||||
ab. Rückgabe: {summary, action_needed, action_text}. Cache je Komponente auf `key`."""
|
||||
empty = {"summary": "", "action_needed": None, "action_text": ""}
|
||||
if not key:
|
||||
return empty
|
||||
cache = _relnotes_caches.setdefault(kind, {})
|
||||
if cache.get("key") == key and cache.get("data"):
|
||||
return cache["data"]
|
||||
from config import LLAMA_SWAP_URL
|
||||
prompt = (
|
||||
"Du bist Lucy, die Stimme einer lokalen AI-Box, und erklärst dem Besitzer (KEIN Entwickler, "
|
||||
"fasst nie eine Konsole an) ein anstehendes Update ruhig und verständlich.\n"
|
||||
f"{context}\n\n"
|
||||
"Anstehende Änderungen:\n" + changes + "\n\n"
|
||||
"Antworte auf DEUTSCH, knapp und klar. HALTE DICH GENAU an dieses Format:\n"
|
||||
"Zeile 1 ist das Aktions-Verdikt und beginnt mit 'AKTION: ':\n"
|
||||
" • 'AKTION: NEIN' — wenn der Besitzer nichts tun muss (die Box spielt es selbst ein, das "
|
||||
"Fangnetz prüft danach automatisch und rollt bei Problemen von allein zurück). Das ist der "
|
||||
"Normalfall.\n"
|
||||
" • 'AKTION: JA — <was er konkret tun/entscheiden muss>' — NUR wenn er wirklich selbst "
|
||||
"handeln muss.\n"
|
||||
"Danach maximal 5 kurze Stichpunkte (je mit '- '), Wichtigstes zuerst: Breaking Changes oder "
|
||||
"geänderte/entfernte Config-Schlüssel ZUERST und mit '⚠️' markiert, dann was unser Setup "
|
||||
"betrifft, dann lohnende neue Features. Keine Einleitung, keine Überschrift."
|
||||
)
|
||||
try:
|
||||
r = httpx.post(f"{LLAMA_SWAP_URL}/v1/chat/completions", timeout=90.0, json={
|
||||
"model": "fast", "max_tokens": 550, "temperature": 0.2,
|
||||
"chat_template_kwargs": {"enable_thinking": False},
|
||||
"messages": [{"role": "user", "content": prompt}],
|
||||
})
|
||||
r.raise_for_status()
|
||||
resp = r.json()
|
||||
choice = (resp.get("choices") or [{}])[0]
|
||||
text = ((choice.get("message") or {}).get("content") or "").strip()
|
||||
finish_reason = choice.get("finish_reason") or ""
|
||||
|
||||
# finish_reason == "length" → Antwort wurde wegen Token-Limits abgeschnitten →
|
||||
# letzten (unvollständigen) Stichpunkt entfernen. Bei "stop"/None → vollständig.
|
||||
if finish_reason == "length" and text:
|
||||
head, _, _tail = text.rpartition("\n")
|
||||
text = head if head else ""
|
||||
|
||||
# Aktions-Verdikt herauslösen (strukturiertes Feld fürs UI). Normalerweise Zeile 1,
|
||||
# aber tolerant: erste passende Zeile suchen (Modell startet manchmal mit Leerzeile).
|
||||
action_needed: bool | None = None
|
||||
action_text = ""
|
||||
lines = text.splitlines()
|
||||
for idx, ln in enumerate(lines):
|
||||
if m := _ACTION_RX.match(ln):
|
||||
action_needed = m.group(1).upper() == "JA"
|
||||
action_text = (m.group(2) or "").strip()
|
||||
text = "\n".join(lines[:idx] + lines[idx + 1:]).strip()
|
||||
break
|
||||
|
||||
data = {"summary": text, "action_needed": action_needed, "action_text": action_text}
|
||||
if text:
|
||||
cache.update(key=key, data=data)
|
||||
return data
|
||||
except Exception as exc: # noqa: BLE001 — Zusammenfassung ist Komfort, nie Blocker
|
||||
return {"summary": f"(Zusammenfassung nicht verfügbar: {exc})",
|
||||
"action_needed": None, "action_text": ""}
|
||||
|
||||
|
||||
def _summarize_hermes_commits(commits: list[dict]) -> dict:
|
||||
subjects = "\n".join(f"- {c['subject']}" for c in commits[:100])
|
||||
context = ("Es geht um ein Update des Hermes-Agenten (das Gehirn/Werkzeug-System der Box auf "
|
||||
"Strix Halo; genutzt werden: api_server/Gateway, memory-provider-Plugin 'mc2-memory', "
|
||||
"terminal-/web-Tools, approvals, cron).")
|
||||
return _summarize_release("hermes", commits[0]["hash"] if commits else "", context, subjects)
|
||||
|
||||
|
||||
def hermes_update_details() -> dict:
|
||||
"""Commits, die ein Hermes-Update einspielen würde (HEAD..origin/<branch>) + LLM-Zusammenfassung."""
|
||||
info: dict = {"kind": "hermes", "branch": None, "behind": 0, "commits": []}
|
||||
git = system.find_hermes_agent_git()
|
||||
if not git or not git.get("path"):
|
||||
info["error"] = "Hermes-Agent-Repo nicht gefunden."
|
||||
return info
|
||||
path = git["path"]
|
||||
try:
|
||||
branch = (subprocess.run(["git", "-C", path, "rev-parse", "--abbrev-ref", "HEAD"],
|
||||
capture_output=True, text=True, timeout=8).stdout.strip() or "main")
|
||||
info["branch"] = branch
|
||||
subprocess.run(["git", "-C", path, "fetch", "-q", "origin", branch],
|
||||
capture_output=True, text=True, timeout=25)
|
||||
log = subprocess.run(["git", "-C", path, "log", "--pretty=format:%h\x1f%s\x1f%cr",
|
||||
f"HEAD..origin/{branch}"], capture_output=True, text=True, timeout=10)
|
||||
commits = []
|
||||
for line in (log.stdout or "").splitlines():
|
||||
parts = line.split("\x1f")
|
||||
if len(parts) == 3:
|
||||
commits.append({"hash": parts[0], "subject": parts[1], "when": parts[2]})
|
||||
info["commits"] = commits
|
||||
info["behind"] = len(commits)
|
||||
if commits:
|
||||
info.update(_summarize_hermes_commits(commits))
|
||||
except Exception as exc: # noqa: BLE001
|
||||
info["error"] = str(exc)
|
||||
return info
|
||||
|
||||
|
||||
def update_details(kind: str) -> dict:
|
||||
return {"os": os_update_details, "engine": engine_update_details,
|
||||
"swap": swap_update_details,
|
||||
"hermes": hermes_update_details}.get(kind, lambda: {"error": "unbekannt"})()
|
||||
|
||||
|
||||
def _run(cmd: list[str], sudo_password: str | None = None) -> dict:
|
||||
actual_cmd = list(cmd)
|
||||
has_sudo = False
|
||||
|
||||
if cmd and cmd[0] == "sudo":
|
||||
has_sudo = True
|
||||
# If we have a password, use -S instead of -n
|
||||
if sudo_password is not None:
|
||||
if "-n" in actual_cmd:
|
||||
actual_cmd = [x for x in actual_cmd if x != "-n"]
|
||||
if "-S" not in actual_cmd:
|
||||
actual_cmd.insert(1, "-S")
|
||||
else:
|
||||
# Force -n to fail cleanly if password is required
|
||||
if "-S" in actual_cmd:
|
||||
actual_cmd = [x for x in actual_cmd if x != "-S"]
|
||||
if "-n" not in actual_cmd:
|
||||
actual_cmd.insert(1, "-n")
|
||||
|
||||
try:
|
||||
input_data = (sudo_password + "\n") if (has_sudo and sudo_password is not None) else None
|
||||
p = subprocess.run(actual_cmd, input=input_data, capture_output=True, text=True, timeout=120)
|
||||
|
||||
err_msg = p.stderr or ""
|
||||
if p.returncode != 0 and ("a password is required" in err_msg or "password" in err_msg.lower() or "sudo:" in err_msg):
|
||||
if sudo_password is not None:
|
||||
return {"ok": False, "status": "incorrect_password", "out": p.stdout or "", "err": "Falsches Sudo-Passwort."}
|
||||
return {"ok": False, "status": "password_required", "out": p.stdout or "", "err": "Sudo-Passwort erforderlich."}
|
||||
|
||||
return {"ok": p.returncode == 0, "out": (p.stdout or "")[-4000:], "err": (p.stderr or "")[-2000:]}
|
||||
except Exception as exc: # noqa: BLE001
|
||||
return {"ok": False, "out": "", "err": str(exc)}
|
||||
|
||||
|
||||
def check_sudo_needs_password(sudo_password: str | None = None) -> dict | None:
|
||||
"""Checks if sudo needs a password. Returns error dict if password required/incorrect, else None."""
|
||||
res = _run(["sudo", "true"], sudo_password=sudo_password)
|
||||
if not res["ok"]:
|
||||
return res
|
||||
return None
|
||||
|
||||
|
||||
def restart_service(name: str, sudo_password: str | None = None) -> dict:
|
||||
if name in SYSTEM_SERVICES:
|
||||
if err := check_sudo_needs_password(sudo_password):
|
||||
return err
|
||||
return _run(["sudo", "systemctl", "restart", name], sudo_password=sudo_password)
|
||||
if name in USER_SERVICES:
|
||||
return _run(["systemctl", "--user", "restart", name])
|
||||
return {"ok": False, "err": f"Dienst '{name}' nicht erlaubt."}
|
||||
|
||||
|
||||
def logs(service: str, lines: int = 200, sudo_password: str | None = None) -> dict:
|
||||
lines = max(1, min(lines, 1000))
|
||||
if service in USER_SERVICES:
|
||||
r = _run(["journalctl", "--user", "-u", service, "-n", str(lines), "--no-pager"])
|
||||
return {"ok": r["ok"], "text": r["out"] or r["err"]}
|
||||
if service in SYSTEM_SERVICES:
|
||||
# Journal-Lesen braucht i.d.R. KEIN sudo (User ist in Gruppe adm/systemd-journal).
|
||||
# Erst ohne sudo versuchen; nur bei fehlenden Rechten auf sudo zurückfallen.
|
||||
r = _run(["journalctl", "-u", service, "-n", str(lines), "--no-pager"])
|
||||
if r["ok"]:
|
||||
return {"ok": True, "text": r["out"] or "(keine Log-Einträge)"}
|
||||
if err := check_sudo_needs_password(sudo_password):
|
||||
return err
|
||||
r = _run(["sudo", "journalctl", "-u", service, "-n", str(lines), "--no-pager"],
|
||||
sudo_password=sudo_password)
|
||||
return {"ok": r["ok"], "text": r["out"] or r["err"]}
|
||||
return {"ok": False, "text": "", "err": "Dienst nicht erlaubt."}
|
||||
|
||||
def check_updates_job(sudo_password: str | None = None) -> dict:
|
||||
if err := check_sudo_needs_password(sudo_password):
|
||||
return err
|
||||
|
||||
def on_done():
|
||||
_engine_cache.update(ts=0.0, avail=False)
|
||||
_comp_cache.update(ts=0.0, data=[]) # Hermes-Status ebenfalls neu berechnen lassen
|
||||
|
||||
cmd = "sudo apt-get update"
|
||||
job_id = jobengine.start_job(["bash", "-c", cmd], "Nach Updates suchen", on_done=on_done, sudo_password=sudo_password)
|
||||
return {"ok": True, "job_id": job_id}
|
||||
|
||||
|
||||
def _maintenance_busy() -> dict | None:
|
||||
"""Wartungs-Riegel: nur EIN binär-/dienst-veränderndes Update gleichzeitig. Verhindert
|
||||
Doppelklick UND parallele Updates aus zwei Tabs/Sessions (racende .bak-Sicherung/Restarts)."""
|
||||
if j := jobengine.active_in_group("maintenance"):
|
||||
return {"ok": False, "status": "busy", "running": j.get("label")}
|
||||
return None
|
||||
|
||||
|
||||
def os_update_job(sudo_password: str | None = None) -> dict:
|
||||
if busy := _maintenance_busy():
|
||||
return busy
|
||||
if err := check_sudo_needs_password(sudo_password):
|
||||
return err
|
||||
# Nach dem apt-Upgrade den Stack funktional prüfen (Job wird rot, wenn etwas kaputt ging).
|
||||
cmd = ("sudo apt-get update && sudo DEBIAN_FRONTEND=noninteractive apt-get upgrade -y "
|
||||
f"&& bash {STACK_POSTCHECK}")
|
||||
job_id = jobengine.start_job(["bash", "-c", cmd], "OS-Update (apt)",
|
||||
group="maintenance", sudo_password=sudo_password)
|
||||
return {"ok": True, "job_id": job_id}
|
||||
|
||||
|
||||
def engine_update_job(sudo_password: str | None = None) -> dict | None:
|
||||
if not ENGINE_UPDATE_CMD:
|
||||
return None
|
||||
if busy := _maintenance_busy():
|
||||
return busy
|
||||
if err := check_sudo_needs_password(sudo_password):
|
||||
return err
|
||||
|
||||
def on_done():
|
||||
_engine_cache.update(ts=0.0, avail=False) # Cache leeren → frischer Build-Vergleich
|
||||
|
||||
# update-engine.sh sichert den alten Build, aktualisiert, startet llama-swap neu, prüft den
|
||||
# Stack (stack-postcheck.sh) und rollt bei Fehler selbst zurück. Exit 0 nur bei verifiziertem
|
||||
# neuen Build → on_done (Cache leeren) läuft nur dann; bei Rollback (Exit 1/2) bleibt das Badge.
|
||||
job_id = jobengine.start_job(["bash", "-c", ENGINE_UPDATE_CMD],
|
||||
"Engine-Update (llama.cpp Vulkan)",
|
||||
group="maintenance", on_done=on_done, sudo_password=sudo_password)
|
||||
return {"ok": True, "job_id": job_id}
|
||||
|
||||
|
||||
def swap_update_job(sudo_password: str | None = None) -> dict | None:
|
||||
if not SWAP_UPDATE_CMD:
|
||||
return None
|
||||
if busy := _maintenance_busy():
|
||||
return busy
|
||||
if err := check_sudo_needs_password(sudo_password):
|
||||
return err
|
||||
|
||||
def on_done():
|
||||
_swap_cache.update(ts=0.0, avail=False) # Cache leeren → frischer Versions-Vergleich
|
||||
|
||||
# update-swap.sh sichert die alte Binary, aktualisiert, startet llama-swap neu, prüft den Stack
|
||||
# (stack-postcheck.sh) und rollt bei Fehler selbst zurück. Exit 0 nur bei verifizierter neuer
|
||||
# Version → on_done (Cache leeren) läuft nur dann; bei Rollback (Exit 1/2) bleibt das Badge.
|
||||
job_id = jobengine.start_job(["bash", "-c", SWAP_UPDATE_CMD],
|
||||
"Router-Update (llama-swap)",
|
||||
group="maintenance", on_done=on_done, sudo_password=sudo_password)
|
||||
return {"ok": True, "job_id": job_id}
|
||||
|
||||
|
||||
def hermes_update_job() -> dict:
|
||||
"""Hermes-Agent aktualisieren wie die CLI (`hermes update` = git pull + Deps), danach
|
||||
den Gateway neu starten. Davor ein Sicherheits-Backup (unser deploy/backup.sh). Kein sudo
|
||||
(alles im User-Space). Läuft als Hintergrund-Job (kann ~1 Min dauern)."""
|
||||
if busy := _maintenance_busy():
|
||||
return busy
|
||||
git = system.find_hermes_agent_git()
|
||||
path = (git or {}).get("path") or os.path.expanduser("~/.hermes/hermes-agent")
|
||||
py = os.path.join(path, "venv", "bin", "python")
|
||||
backup = os.path.join(_REPO_ROOT, "deploy", "backup.sh")
|
||||
postcheck = os.path.join(_REPO_ROOT, "deploy", "hermes-postcheck.sh")
|
||||
# Backup → update → Gateway-Neustart → Gehirn-Check (Job wird rot, wenn Mem0/Plugin kaputt).
|
||||
# Backup → update → doctor (Config/Deps-Diagnose der NEUEN Version) → Gateway-Neustart →
|
||||
# Gehirn-Check (erweitert um Config-Drift-Scan, Tool- und Voice-Smoke — Lehre aus v0.18:
|
||||
# 'approvals.mode auto' wurde still ungültig und legte alle Tools in pending_approval).
|
||||
cmd = (f"bash {backup} || true; "
|
||||
f"cd {path} && {py} -m hermes_cli.main update --yes "
|
||||
f"&& {py} -m hermes_cli.main doctor "
|
||||
f"&& systemctl --user restart hermes-gateway "
|
||||
f"&& sleep 6 && bash {postcheck}")
|
||||
|
||||
def on_done():
|
||||
_comp_cache.update(ts=0.0, data=[]) # Update-Status neu berechnen lassen
|
||||
|
||||
job_id = jobengine.start_job(["bash", "-c", cmd], "Hermes-Agent-Update",
|
||||
group="maintenance", on_done=on_done)
|
||||
return {"ok": True, "job_id": job_id}
|
||||
|
||||
|
||||
def reboot(sudo_password: str | None = None) -> dict:
|
||||
if err := check_sudo_needs_password(sudo_password):
|
||||
return err
|
||||
return _run(["sudo", "reboot"], sudo_password=sudo_password)
|
||||
@@ -0,0 +1,147 @@
|
||||
"""
|
||||
Geteiltes Gedächtnis (die „Verfassung") — jetzt auto-lernend & semantisch über Mem0.
|
||||
|
||||
Dieser Service ist nur noch ein dünner HTTP-Client auf den Mem0-Sidecar (mem0_service/app.py,
|
||||
läuft im ~/.mem0/venv unter Python 3.12). Die `/api/memory`-API-Form bleibt unverändert, damit
|
||||
UI und MCP-Server kompatibel bleiben. Neu gegenüber der alten flachen SQLite:
|
||||
|
||||
- search (q gesetzt) ist SEMANTISCH (Vektor/Embeddings) statt LIKE-Textsuche, mit Relevanz-Score.
|
||||
- learn() reicht Gesprächs-Turns durch → Mem0 EXTRAHIERT Fakten selbst (Auto-Lernen).
|
||||
- Dedup macht Mem0 beim Auto-Lernen selbst; der manuelle Kurator unten bleibt als Komfort.
|
||||
|
||||
5 Kategorien (user · instruction · stable · versioned · ephemeral) bleiben als Metadaten erhalten.
|
||||
"""
|
||||
|
||||
import re
|
||||
from difflib import SequenceMatcher
|
||||
|
||||
import httpx
|
||||
|
||||
from config import MEM0_SERVICE_URL
|
||||
|
||||
CATEGORIES = ("identity", "knowledge", "rules", "events")
|
||||
|
||||
_TIMEOUT = httpx.Timeout(60.0, connect=5.0) # LLM-Extraktion kann ein paar Sekunden dauern
|
||||
|
||||
|
||||
def _get(path: str, **params) -> list | dict:
|
||||
r = httpx.get(f"{MEM0_SERVICE_URL}{path}", params=params, timeout=_TIMEOUT)
|
||||
r.raise_for_status()
|
||||
return r.json()
|
||||
|
||||
|
||||
def _post(path: str, data: dict) -> dict:
|
||||
r = httpx.post(f"{MEM0_SERVICE_URL}{path}", json=data, timeout=_TIMEOUT)
|
||||
r.raise_for_status()
|
||||
return r.json()
|
||||
|
||||
|
||||
def _put(path: str, data: dict) -> dict:
|
||||
r = httpx.put(f"{MEM0_SERVICE_URL}{path}", json=data, timeout=_TIMEOUT)
|
||||
r.raise_for_status()
|
||||
return r.json()
|
||||
|
||||
|
||||
def _delete(path: str) -> dict:
|
||||
r = httpx.delete(f"{MEM0_SERVICE_URL}{path}", timeout=_TIMEOUT)
|
||||
r.raise_for_status()
|
||||
return r.json()
|
||||
|
||||
|
||||
def list_memories(q: str = "", category: str = "") -> list[dict]:
|
||||
"""Alle Fakten oder — wenn q gesetzt — die semantisch ähnlichsten (mit `score`)."""
|
||||
return _get("/memory", **{k: v for k, v in (("q", q), ("category", category)) if v})
|
||||
|
||||
|
||||
def add_memory(content: str, category: str = "stable", source: str = "manual") -> dict:
|
||||
"""Einen Fakt VERBATIM speichern (keine LLM-Umformung). Auto-Lernen → learn()."""
|
||||
return _post("/memory", {"content": content.strip(), "category": category, "source": source})
|
||||
|
||||
|
||||
def update_memory(mid: str, content: str | None = None, category: str | None = None) -> dict | None:
|
||||
try:
|
||||
return _put(f"/memory/{mid}", {"content": content, "category": category})
|
||||
except httpx.HTTPStatusError as exc:
|
||||
if exc.response.status_code == 404:
|
||||
return None
|
||||
raise
|
||||
|
||||
|
||||
def delete_memory(mid: str) -> bool:
|
||||
try:
|
||||
_delete(f"/memory/{mid}")
|
||||
return True
|
||||
except httpx.HTTPStatusError as exc:
|
||||
if exc.response.status_code == 404:
|
||||
return False
|
||||
raise
|
||||
|
||||
|
||||
def learn(text: str | None = None, messages: list[dict] | None = None,
|
||||
source: str = "auto", category: str = "stable") -> dict:
|
||||
"""Auto-Lernen: Text/Gesprächs-Turns durchreichen → Mem0 extrahiert die Fakten selbst."""
|
||||
return _post("/learn", {"text": text, "messages": messages,
|
||||
"source": source, "category": category})
|
||||
|
||||
|
||||
def graph(min_score: float = 0.45, top_k: int = 3) -> dict:
|
||||
"""Fakten als Ähnlichkeits-Graph (Knoten + semantische Kanten) für die UI-Visualisierung."""
|
||||
return _get("/graph", min_score=min_score, top_k=top_k)
|
||||
|
||||
|
||||
def export_text() -> dict:
|
||||
rows = sorted(list_memories(), key=lambda r: (r.get("category", ""), r.get("updated_at", "")))
|
||||
lines = ["# Mission Control — Gedächtnis\n"]
|
||||
current = ""
|
||||
for r in rows:
|
||||
if r.get("category") != current:
|
||||
current = r.get("category", "")
|
||||
lines.append(f"\n## {current}\n")
|
||||
src = r.get("source", "")
|
||||
when = (r.get("updated_at") or "")[:10]
|
||||
lines.append(f"- {r.get('content', '')} _(Quelle: {src}, {when})_")
|
||||
return {"text": "\n".join(lines), "count": len(rows)}
|
||||
|
||||
|
||||
# --- Manueller Kurator (deterministisch, kein LLM) ---------------------------
|
||||
def _norm(s: str) -> str:
|
||||
s = re.sub(r"[^\w\s]", " ", s.lower(), flags=re.UNICODE)
|
||||
return re.sub(r"\s+", " ", s).strip()
|
||||
|
||||
|
||||
def dedupe(apply: bool = False, threshold: float = 0.85) -> dict:
|
||||
"""Findet Dubletten (exakt/enthalten/ähnlich) je Kategorie, behält den längsten
|
||||
Eintrag. Mem0 dedupliziert beim Auto-Lernen schon semantisch — das hier ist der
|
||||
manuelle Komfort-Knopf fürs UI (z.B. nach vielen Verbatim-Importen)."""
|
||||
rows = sorted(list_memories(), key=lambda r: (-len(r.get("content", "")), r.get("created_at", "")))
|
||||
used: set[str] = set()
|
||||
groups: list[dict] = []
|
||||
for i, a in enumerate(rows):
|
||||
if a["id"] in used:
|
||||
continue
|
||||
na = _norm(a.get("content", ""))
|
||||
if not na:
|
||||
continue
|
||||
dups = []
|
||||
for b in rows[i + 1:]:
|
||||
if b["id"] in used or b.get("category") != a.get("category"):
|
||||
continue
|
||||
nb = _norm(b.get("content", ""))
|
||||
if not nb:
|
||||
continue
|
||||
if nb in na or na in nb or SequenceMatcher(None, na, nb).ratio() >= threshold:
|
||||
dups.append(b); used.add(b["id"])
|
||||
if dups:
|
||||
used.add(a["id"])
|
||||
groups.append({
|
||||
"keep": {"id": a["id"], "content": a.get("content"), "category": a.get("category")},
|
||||
"remove": [{"id": d["id"], "content": d.get("content")} for d in dups],
|
||||
})
|
||||
dup_count = sum(len(g["remove"]) for g in groups)
|
||||
removed = 0
|
||||
if apply:
|
||||
for g in groups:
|
||||
for d in g["remove"]:
|
||||
if delete_memory(d["id"]):
|
||||
removed += 1
|
||||
return {"groups": groups, "duplicate_count": dup_count, "removed": removed, "applied": apply}
|
||||
@@ -0,0 +1,58 @@
|
||||
"""Kosten-/Ersparnis-Berechnung für die Token-Statistik (eine Quelle der Wahrheit).
|
||||
|
||||
Vergleicht die lokal verbrauchten Tokens gegen die Cloud-Listenpreise vergleichbarer
|
||||
Modellklassen (Stand Juni 2026, USD pro 1M Tokens, in/out) und liefert die so
|
||||
eingesparte Summe. Wird vom System-Router dünn aufgerufen.
|
||||
"""
|
||||
|
||||
import os
|
||||
|
||||
# Cloud-Listenpreise je Rolle/Modellklasse: (input_usd_per_1M, output_usd_per_1M).
|
||||
PRICING: dict[str, tuple[float, float]] = {
|
||||
"heavy": (15.0, 75.0),
|
||||
"coder": (3.0, 15.0),
|
||||
"hermes": (1.0, 5.0),
|
||||
"fast": (0.15, 0.60),
|
||||
"scout": (0.15, 0.60),
|
||||
"vision": (0.15, 0.60),
|
||||
"reasoning": (0.15, 0.60),
|
||||
}
|
||||
# Tarif für nicht zuordenbare Tokens (Default-/Fallback-Klasse).
|
||||
DEFAULT_RATE: tuple[float, float] = (0.15, 0.60)
|
||||
# Baseline/Legacy-Tokens (vor modellspezifischem Logging) am Premium-Tarif bewerten,
|
||||
# damit historische Ersparnis erhalten bleibt.
|
||||
BASELINE_RATE: tuple[float, float] = PRICING["heavy"]
|
||||
USD_TO_EUR = float(os.environ.get("MC_USD_TO_EUR", "0.92"))
|
||||
|
||||
|
||||
def compute_savings(stats: dict, role_map: dict[str, str | None]) -> dict:
|
||||
"""Aggregiert Tokens und berechnet die Cloud-Ersparnis.
|
||||
|
||||
role_map: Modell-/Alias-Name (lowercase) -> Rolle, zur Tarif-Auflösung.
|
||||
"""
|
||||
prompt = stats.get("prompt_tokens", 0)
|
||||
completion = stats.get("completion_tokens", 0)
|
||||
|
||||
modeled_p = modeled_c = 0
|
||||
saved_usd = 0.0
|
||||
for m_name, m_tokens in (stats.get("models") or {}).items():
|
||||
mp = m_tokens.get("prompt", 0)
|
||||
mc = m_tokens.get("completion", 0)
|
||||
modeled_p += mp
|
||||
modeled_c += mc
|
||||
role = role_map.get(m_name, m_name)
|
||||
rate_in, rate_out = PRICING.get(role, DEFAULT_RATE)
|
||||
saved_usd += (mp * rate_in + mc * rate_out) / 1_000_000.0
|
||||
|
||||
baseline_p = max(0, prompt - modeled_p)
|
||||
baseline_c = max(0, completion - modeled_c)
|
||||
saved_usd += (baseline_p * BASELINE_RATE[0] + baseline_c * BASELINE_RATE[1]) / 1_000_000.0
|
||||
|
||||
return {
|
||||
"prompt_tokens": prompt,
|
||||
"completion_tokens": completion,
|
||||
"total_tokens": prompt + completion,
|
||||
"saved_usd": round(saved_usd, 2),
|
||||
"saved_eur": round(saved_usd * USD_TO_EUR, 2),
|
||||
"pricing": {role: {"in": r[0], "out": r[1]} for role, r in PRICING.items()},
|
||||
}
|
||||
@@ -0,0 +1,170 @@
|
||||
"""
|
||||
Erinnerungen & Routinen (Lucy-Proaktivität, Faden A3).
|
||||
|
||||
„Lucy, erinner mich morgen um 9 an …" — der Hermes-Agent legt per Tool (mcp_mc.py:
|
||||
reminder_create) einen Eintrag an; der Wecker-Loop hier feuert ihn zur Zeit in den
|
||||
Melde-Briefkasten (announce.py → Lucy spricht) und parallel auf Telegram.
|
||||
|
||||
Bewusst OHNE Hermes-cron: der Wecker läuft deterministisch im MC2-Backend (systemd,
|
||||
überlebt Gateway-Ausfälle) und bleibt in UNSEREN Schichten (Hermes-Quellcode/Config
|
||||
bleibt Auto-Update-Kanal, Verdikt Faden 11c).
|
||||
|
||||
Zeiten: ISO-8601. Ohne Offset gilt die Zeitzone des Commanders (MC_LOCAL_TZ,
|
||||
Default Europe/Berlin) — die Box selbst läuft auf UTC, naive Zeiten dürfen daher
|
||||
NIE als Box-Lokalzeit interpretiert werden. Wiederholung: daily | weekdays | weekly.
|
||||
"""
|
||||
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
import threading
|
||||
import time
|
||||
from datetime import datetime, timedelta
|
||||
from pathlib import Path
|
||||
from zoneinfo import ZoneInfo
|
||||
|
||||
from config import MODELS_DIR
|
||||
from services import announce
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
STORE_PATH = Path(os.environ.get("MC_REMINDERS_STORE", str(MODELS_DIR / "mc2-reminders.json")))
|
||||
LOCAL_TZ = ZoneInfo(os.environ.get("MC_LOCAL_TZ", "Europe/Berlin"))
|
||||
INTERVAL = int(os.environ.get("MC_REMINDERS_INTERVAL", "20")) # Wecker-Tick (Sekunden)
|
||||
MAX_ITEMS = int(os.environ.get("MC_REMINDERS_MAX", "100"))
|
||||
REPEATS = ("", "daily", "weekdays", "weekly")
|
||||
LATE_NOTE_S = 600 # feuert >10 min zu spät (Box war aus) → ehrlich dazusagen
|
||||
|
||||
_lock = threading.Lock()
|
||||
_state: dict | None = None # {"next_id": int, "items": [...]}
|
||||
|
||||
|
||||
def _load() -> dict:
|
||||
global _state
|
||||
if _state is None:
|
||||
try:
|
||||
_state = json.loads(STORE_PATH.read_text(encoding="utf-8"))
|
||||
assert isinstance(_state.get("next_id"), int) and isinstance(_state.get("items"), list)
|
||||
except Exception:
|
||||
_state = {"next_id": 1, "items": []}
|
||||
return _state
|
||||
|
||||
|
||||
def _save(state: dict) -> None:
|
||||
try:
|
||||
tmp = STORE_PATH.with_suffix(".tmp")
|
||||
tmp.write_text(json.dumps(state, ensure_ascii=False), encoding="utf-8")
|
||||
tmp.replace(STORE_PATH)
|
||||
except OSError:
|
||||
log.warning("reminders: Store %s nicht schreibbar", STORE_PATH, exc_info=True)
|
||||
|
||||
|
||||
def _parse_when(when: str) -> datetime:
|
||||
"""ISO-8601 → bewusste Zeit. Naive Angaben = Commander-Zeitzone (NICHT Box-UTC)."""
|
||||
try:
|
||||
dt = datetime.fromisoformat(when.strip())
|
||||
except ValueError:
|
||||
raise ValueError("Zeit nicht lesbar — bitte ISO-8601 mit Datum UND Uhrzeit, z.B. 2026-07-04T09:00.")
|
||||
if dt.tzinfo is None:
|
||||
dt = dt.replace(tzinfo=LOCAL_TZ)
|
||||
return dt
|
||||
|
||||
|
||||
def _fmt(ts: float) -> str:
|
||||
"""Menschlich lesbare Commander-Lokalzeit (fürs Tool-Echo/Listing)."""
|
||||
return datetime.fromtimestamp(ts, LOCAL_TZ).strftime("%a %d.%m.%Y %H:%M")
|
||||
|
||||
|
||||
def _advance(ts: float, repeat: str) -> float:
|
||||
"""Nächste Wiederholung NACH jetzt (holt verpasste Termine ohne Mehrfach-Feuern auf)."""
|
||||
dt = datetime.fromtimestamp(ts, LOCAL_TZ)
|
||||
now = time.time()
|
||||
while dt.timestamp() <= now:
|
||||
dt += timedelta(days=7 if repeat == "weekly" else 1)
|
||||
if repeat == "weekdays":
|
||||
while dt.weekday() > 4: # Sa/So überspringen
|
||||
dt += timedelta(days=1)
|
||||
return dt.timestamp()
|
||||
|
||||
|
||||
def create(text: str, when: str, repeat: str = "") -> dict:
|
||||
text = (text or "").strip()
|
||||
if not text:
|
||||
raise ValueError("Leerer Erinnerungs-Text.")
|
||||
repeat = (repeat or "").strip().lower()
|
||||
if repeat == "once":
|
||||
repeat = ""
|
||||
if repeat not in REPEATS:
|
||||
raise ValueError("repeat muss leer, daily, weekdays oder weekly sein.")
|
||||
ts = _parse_when(when).timestamp()
|
||||
if repeat:
|
||||
if ts <= time.time(): # Startpunkt vorbei → auf die nächste Wiederholung rücken (Uhrzeit bleibt exakt)
|
||||
ts = _advance(ts, repeat)
|
||||
elif ts <= time.time() - 60:
|
||||
raise ValueError(f"Die Zeit liegt in der Vergangenheit ({_fmt(ts)}) — bitte neu umrechnen "
|
||||
f"(jetzt ist {_fmt(time.time())}).")
|
||||
with _lock:
|
||||
state = _load()
|
||||
if len(state["items"]) >= MAX_ITEMS:
|
||||
raise ValueError(f"Zu viele offene Erinnerungen (max {MAX_ITEMS}).")
|
||||
item = {"id": state["next_id"], "text": text[:500], "next_ts": ts,
|
||||
"repeat": repeat, "created_ts": time.time()}
|
||||
state["next_id"] += 1
|
||||
state["items"].append(item)
|
||||
_save(state)
|
||||
log.info("reminder #%s angelegt: %s → %s%s", item["id"], _fmt(ts), text[:60],
|
||||
f" (Routine {repeat})" if repeat else "")
|
||||
return {**item, "when_local": _fmt(ts)}
|
||||
|
||||
|
||||
def list_all() -> list[dict]:
|
||||
with _lock:
|
||||
items = sorted(_load()["items"], key=lambda i: i["next_ts"])
|
||||
return [{**i, "when_local": _fmt(i["next_ts"])} for i in items]
|
||||
|
||||
|
||||
def delete(reminder_id: int) -> dict:
|
||||
with _lock:
|
||||
state = _load()
|
||||
for i, item in enumerate(state["items"]):
|
||||
if item["id"] == reminder_id:
|
||||
state["items"].pop(i)
|
||||
_save(state)
|
||||
return {**item, "when_local": _fmt(item["next_ts"])}
|
||||
raise KeyError(f"Erinnerung {reminder_id} nicht gefunden.")
|
||||
|
||||
|
||||
def _fire_due() -> None:
|
||||
now = time.time()
|
||||
with _lock:
|
||||
state = _load()
|
||||
due = [i for i in state["items"] if i["next_ts"] <= now]
|
||||
if not due:
|
||||
return
|
||||
for item in due:
|
||||
if item["repeat"]:
|
||||
item["next_ts"] = _advance(item["next_ts"], item["repeat"])
|
||||
else:
|
||||
state["items"].remove(item)
|
||||
_save(state)
|
||||
for item in due:
|
||||
late = now - item["next_ts"] > LATE_NOTE_S if not item["repeat"] else False
|
||||
text = f"Erinnerung, Commander: {item['text']}"
|
||||
if late:
|
||||
text += f" (Eigentlich fällig um {_fmt(item['next_ts'])} — die Box war wohl aus.)"
|
||||
announce.add(text, subject="[Erinnerung]", source="reminder")
|
||||
announce.notify_telegram("[Erinnerung]", item["text"])
|
||||
log.info("reminder #%s gefeuert%s", item["id"], " (Routine)" if item["repeat"] else "")
|
||||
|
||||
|
||||
async def reminders_loop() -> None:
|
||||
"""Wecker (Hintergrund-Task im MC2-Lifespan). Verpasste einmalige Erinnerungen
|
||||
(Box war aus) feuern beim nächsten Tick nach — ehrlich als verspätet markiert."""
|
||||
import asyncio
|
||||
log.info("reminders: Wecker aktiv (Tick %ss, Zeitzone %s)", INTERVAL, LOCAL_TZ)
|
||||
while True:
|
||||
try:
|
||||
await asyncio.to_thread(_fire_due)
|
||||
except Exception:
|
||||
log.debug("reminders: Tick fehlgeschlagen", exc_info=True)
|
||||
await asyncio.sleep(INTERVAL)
|
||||
@@ -0,0 +1,143 @@
|
||||
"""
|
||||
Rollen-Empfehlung: welches INSTALLIERTE Modell passt am besten auf eine Serving-Rolle?
|
||||
Capability-getrieben (Vision/Coder/Tools/MoE aus services.caps) + setup-bewusster Fit
|
||||
(services.budget). Speist den 'Empfohlen'-Hinweis + Auto-Pick im Rollen-Zuweisungs-Modal.
|
||||
|
||||
EINE Quelle der Wahrheit mit der ctx-/Fit-Logik: nutzt budget.setup_aware_ctx_for_model
|
||||
und fit.evaluate_fit — dieselbe Mathematik wie Install-Automatik und Auto-ctx-Button.
|
||||
"""
|
||||
|
||||
import psutil
|
||||
|
||||
from services import budget, catalog, llamaswap
|
||||
from services.fit import evaluate_fit
|
||||
|
||||
|
||||
def _ram_gb() -> float:
|
||||
return psutil.virtual_memory().total / (1024 ** 3)
|
||||
|
||||
|
||||
def _capability_suit(role: str, caps: dict, name: str) -> float:
|
||||
"""0..1 — Capability-Eignung (HARTE Gates). 0 = grundsätzlich falsch für die Rolle.
|
||||
Größe/Tempo bewertet getrennt _pref(), damit z.B. 'fast' nicht das größte Modell zieht."""
|
||||
role = (role or "").lower()
|
||||
low = (name or "").lower()
|
||||
vision = bool(caps.get("vision"))
|
||||
coder = bool(caps.get("coder"))
|
||||
tools = caps.get("tools") != "no"
|
||||
|
||||
if role == "vision":
|
||||
if not vision:
|
||||
return 0.0 # harte Anforderung
|
||||
return 1.0 if ("vl" in low or "llava" in low or "pixtral" in low) else 0.7 # dediziert > omni
|
||||
if role == "coder":
|
||||
return 1.0 if coder else 0.4 # Coder-Modell Pflicht für Empfehlung
|
||||
if role == "hermes":
|
||||
# Agent-Hirn: Hermes-Familie am robustesten; sonst natives Tool-Calling Pflicht.
|
||||
if "hermes" in low:
|
||||
return 1.0
|
||||
return 0.6 if tools else 0.1
|
||||
if role == "fast":
|
||||
# Alltags-Hirn braucht zuverlässige Tools; Größe/Tempo macht _pref.
|
||||
return 1.0 if tools else 0.6
|
||||
if role == "heavy":
|
||||
return 1.0
|
||||
if role == "scout":
|
||||
return 0.9 if vision else 0.7
|
||||
return 0.5
|
||||
|
||||
|
||||
def _pref(role: str, params: float, tps: float) -> float:
|
||||
"""0..1 — rollengerechte GRÖSSEN-/TEMPO-Präferenz. 'fast' belohnt Tempo & Kleinheit,
|
||||
'heavy' Größe (Wissen), 'hermes' moderate Größe (muss warm + ko-resident bleiben)."""
|
||||
role = (role or "").lower()
|
||||
if role == "fast":
|
||||
speed = min(tps / 25.0, 1.0)
|
||||
size_ok = 1.0 if params <= 50 else 50.0 / params
|
||||
return speed * size_ok
|
||||
if role == "heavy":
|
||||
return min(params / 120.0, 1.0)
|
||||
if role == "hermes":
|
||||
return 1.0 if params <= 24 else max(0.15, 24.0 / params) # 7–24B ideal als Hirn
|
||||
if role == "vision":
|
||||
return 1.0 if params <= 12 else 0.7 # klein/günstig bevorzugt
|
||||
if role == "coder":
|
||||
return 0.5 + 0.5 * min(params / 80.0, 1.0)
|
||||
if role == "scout":
|
||||
return 1.0 if params <= 40 else 0.5
|
||||
return 0.5
|
||||
|
||||
|
||||
def _catalog_role_match(role: str, name: str) -> bool:
|
||||
"""Ist dieses Modell im kuratierten Katalog (Cookbook) genau für DIESE Rolle gelistet?
|
||||
Dann ist es der prinzipien-konforme Pick → starker Bonus."""
|
||||
meta = catalog.meta_for_name(name)
|
||||
return bool(meta and (meta.get("role") or "").lower() == (role or "").lower())
|
||||
|
||||
|
||||
def _reason(role: str, caps: dict, name: str, fit: dict, fits: bool,
|
||||
incomplete: bool, cat_match: bool) -> str:
|
||||
if incomplete:
|
||||
return "Download unvollständig"
|
||||
if role == "vision" and not caps.get("vision"):
|
||||
return "keine Vision-Fähigkeit"
|
||||
if role == "coder" and not caps.get("coder"):
|
||||
return "kein Coder-Modell"
|
||||
if role == "hermes" and "hermes" not in (name or "").lower() and caps.get("tools") == "no":
|
||||
return "kein natives Tool-Calling"
|
||||
if not fits:
|
||||
return "passt nicht ins Budget (OOM)"
|
||||
bits = []
|
||||
if cat_match:
|
||||
bits.append("Katalog-Pick ✓")
|
||||
if role == "vision":
|
||||
bits.append("Vision ✓")
|
||||
if role == "coder" and caps.get("coder"):
|
||||
bits.append("Coder ✓")
|
||||
if role == "hermes":
|
||||
bits.append("Hermes" if "hermes" in (name or "").lower()
|
||||
else ("Tools ✓" if caps.get("tools") != "no" else "ohne Tools"))
|
||||
if caps.get("moe"):
|
||||
bits.append("MoE")
|
||||
bits.append(f"{fit['text']}, ~{fit['tps']:.0f} t/s")
|
||||
return " · ".join(bits)
|
||||
|
||||
|
||||
def recommend_for_role(role: str) -> dict:
|
||||
"""Rankt alle installierten Modelle für eine Rolle. Empfohlen = bester geeigneter,
|
||||
passender Eintrag. Liefert pro Modell Fit/Eignung/Begründung fürs UI."""
|
||||
role = (role or "").strip().lower()
|
||||
ram = _ram_gb()
|
||||
out = []
|
||||
for m in llamaswap.list_models():
|
||||
caps = m.get("capabilities") or {}
|
||||
params = budget.params_of_model(m)
|
||||
quant = m.get("quant") or "Q4_K_M"
|
||||
ctx = budget.setup_aware_ctx_for_model(m)["ctx"]
|
||||
fit = evaluate_fit(params, quant, ctx, ram, name=m["name"])
|
||||
incomplete = bool(m.get("incomplete"))
|
||||
fits = (fit["level"] != "too_tight") and not incomplete
|
||||
tps = fit["tps"] or 0
|
||||
suit = _capability_suit(role, caps, m["name"])
|
||||
cat_match = _catalog_role_match(role, m["name"])
|
||||
suitable = suit >= 0.5 and fits
|
||||
|
||||
fit_term = {"perfect": 1.0, "marginal": 0.3}.get(fit["level"], -2.0)
|
||||
# Eignung dominiert (×2), rollengerechte Größe/Tempo (_pref), Katalog-Anker, dann Fit.
|
||||
score = (2.0 * suit) + _pref(role, params, tps) + (0.6 if cat_match else 0.0) + fit_term
|
||||
if not fits:
|
||||
score -= 5.0
|
||||
|
||||
out.append({
|
||||
"name": m["name"], "current_role": m.get("role"),
|
||||
"params_b": round(params, 1), "quant": quant,
|
||||
"fit": fit, "suitable": suitable, "incomplete": incomplete,
|
||||
"score": round(score, 3),
|
||||
"reason": _reason(role, caps, m["name"], fit, fits, incomplete, cat_match),
|
||||
})
|
||||
|
||||
out.sort(key=lambda x: -x["score"])
|
||||
rec = next((o["name"] for o in out if o["suitable"]), None)
|
||||
for o in out:
|
||||
o["recommended"] = (o["name"] == rec)
|
||||
return {"role": role, "recommended": rec, "models": out}
|
||||
@@ -0,0 +1,90 @@
|
||||
"""
|
||||
Lane-Routing für den eingebauten MC2-Gateway (:9001/v1).
|
||||
|
||||
Zwei virtuelle Lanes, die Clients/IDEs auswählen — der Router pickt das echte Modell:
|
||||
- **chat** (= altes `auto`): Alltag → `fast`, schwer/lang → `heavy`.
|
||||
- **coding**: Code-Arbeit → `coder` (Qwen3-Coder-Next); riesiger/architektonischer Kontext → `heavy`;
|
||||
triviale Kurzfrage ohne Code → `fast` (Tempo).
|
||||
|
||||
Regelbasiert, sub-ms, ohne Cloud. Schwellen/Aliases liegen in einer UI-editierbaren Policy
|
||||
(routing_policy.py, hot-reload; Env = Defaults). Lucy läuft NICHT hierüber — die ist der
|
||||
Hermes-Agent (:8642), eigene Ebene.
|
||||
"""
|
||||
|
||||
import re
|
||||
|
||||
from services.routing_policy import load_policy
|
||||
|
||||
# Modell-Aliases & Zeichen-Schwellen liegen jetzt in der UI-editierbaren Policy
|
||||
# (routing_policy.py) und kommen pro Request via load_policy() (hot-reload). Die Env-Vars
|
||||
# sind dort die Defaults. Die Regex-Keyword-Listen unten bleiben bewusst im Code.
|
||||
|
||||
# Virtuelle Lanes, die im Gateway als „Modelle" sichtbar sind.
|
||||
LANES = ["coding", "chat"]
|
||||
LANE_ALIASES = {"auto": "chat"} # Rückwärtskompatibel: model:auto == chat
|
||||
|
||||
_HEAVY_KW = re.compile(
|
||||
r"\b(beweis|prove|theorem|komplex|complex|schwierig|"
|
||||
r"think\s*hard|reason\s*carefully|tief\s*nachdenk|optimi[sz]e|"
|
||||
r"root\s*cause|analy[sz]e\s+deeply|step[-\s]?by[-\s]?step)\b",
|
||||
re.IGNORECASE,
|
||||
)
|
||||
# Coding-spezifische „das ist groß/architektonisch" Signale → heavy statt coder.
|
||||
_CODING_HEAVY_KW = re.compile(
|
||||
r"\b(architekt|architect|system[-\s]?design|refactor\s+the\s+(whole|entire)|"
|
||||
r"ganze[ns]?\s+(architektur|codebase|projekt)|migrat\w+\s+(the\s+)?(whole|entire|gesamte)|"
|
||||
r"entwirf\s+(eine\s+)?architektur|plane?\s+(die\s+)?architektur)\b",
|
||||
re.IGNORECASE,
|
||||
)
|
||||
# Code-Indikatoren — verhindert, dass echte Code-Anfragen als „trivial" auf fast abrutschen.
|
||||
# Bewusst breit (inkl. natürlichsprachiger Coding-Begriffe DE+EN): in der coding-Lane soll im Zweifel
|
||||
# `coder` gewinnen; nur echte Nicht-Code-Kürze ("hallo", "wie spät") rutscht auf fast.
|
||||
_CODE_HINT = re.compile(
|
||||
r"```|\bdef \b|\bclass \b|\bimport \b|\bfunction\b|=>|;\s*$|"
|
||||
r"\.(py|ts|tsx|js|jsx|go|rs|java|cpp|c|rb|php|sql)\b|/src/|traceback|stack\s*trace|"
|
||||
r"\b(funktion|function|bug|fix|fehler|error|exception|implementier\w*|schreib\w*|"
|
||||
r"code\w*|coden|test\w*|klasse|method\w*|methode|refactor\w*|kompil\w*|compile|"
|
||||
r"build|deploy|debug|script|skript|api|endpoint|query|regex|json|yaml|"
|
||||
r"npm|pip|git|docker|terminal|shell|command)\b",
|
||||
re.IGNORECASE | re.MULTILINE,
|
||||
)
|
||||
|
||||
|
||||
def _text_of(body: dict) -> str:
|
||||
msgs = body.get("messages") or []
|
||||
return "\n".join(str(m.get("content") or "") for m in msgs)
|
||||
|
||||
|
||||
def _route_chat(text: str, n: int) -> tuple[str, str]:
|
||||
p = load_policy()
|
||||
if n > p["heavy_chars"]:
|
||||
return p["heavy"], f"langer Kontext ({n} > {p['heavy_chars']} Zeichen)"
|
||||
if _HEAVY_KW.search(text):
|
||||
return p["heavy"], "Komplexitäts-Schlüsselwort erkannt"
|
||||
return p["fast"], "Standard"
|
||||
|
||||
|
||||
def _route_coding(text: str, n: int) -> tuple[str, str]:
|
||||
# Agentisches Coden (OpenCode/RooCode/…) bleibt IMMER beim dedizierten Coder — NIE heavy/fast
|
||||
# (das sind Allzweck-Modelle, schwächer bei Code). Die Qwen-Coder packen 256K–1M Kontext selbst,
|
||||
# langer Repo-Kontext ist bei Agenten der Normalfall und darf NICHT zu heavy umrouten.
|
||||
# (Phase 2b: warme schnelle Coder-Stufe coder_lite als Default + coder als Eskalation.)
|
||||
p = load_policy()
|
||||
if p["coder_lite"] and not _CODING_HEAVY_KW.search(text) and n <= p["coding_escalate_chars"]:
|
||||
return p["coder_lite"], "Coding (schneller Coder)"
|
||||
return p["coder"], "Coding -> starker Coder"
|
||||
|
||||
|
||||
def choose_for_lane(lane: str, body: dict) -> tuple[str, str]:
|
||||
"""Wählt das echte Modell-Alias für eine Lane. Gibt (alias, begründung) zurück."""
|
||||
lane = LANE_ALIASES.get((lane or "chat").lower(), (lane or "chat").lower())
|
||||
text = _text_of(body)
|
||||
n = len(text)
|
||||
if lane == "coding":
|
||||
return _route_coding(text, n)
|
||||
return _route_chat(text, n) # chat + alles Unbekannte
|
||||
|
||||
|
||||
def choose_model(body: dict) -> tuple[str, str]:
|
||||
"""Rückwärtskompatibel: altes `model:auto` == chat-Lane."""
|
||||
return choose_for_lane("chat", body)
|
||||
@@ -0,0 +1,127 @@
|
||||
"""
|
||||
UI-editierbare Routing-Policy für die Gateway-Lanes (coding/chat).
|
||||
|
||||
Persistiert als JSON unter MC_ROUTING_POLICY_PATH (Default MODELS_DIR/mc2-routing.json —
|
||||
gleiche Konvention wie mc2-discover.json). **Hot-reload:** load_policy() liest die Datei nur
|
||||
bei Änderung neu (mtime-Cache) → UI-Edits greifen ohne Dienst-Neustart. Die Env-Vars (bisher
|
||||
einzige Stellschraube in router_logic.py) bleiben als Defaults/Fallback erhalten.
|
||||
|
||||
Bewusst NICHT editierbar (v1): die Regex-Keyword-Listen (heavy/coding-heavy/code-hint) — die
|
||||
bleiben in router_logic.py im Code.
|
||||
"""
|
||||
|
||||
import json
|
||||
import os
|
||||
import threading
|
||||
from pathlib import Path
|
||||
|
||||
from config import MODELS_DIR
|
||||
|
||||
POLICY_PATH = Path(os.environ.get("MC_ROUTING_POLICY_PATH", str(MODELS_DIR / "mc2-routing.json")))
|
||||
|
||||
|
||||
def _env_bool(name: str, default: str) -> bool:
|
||||
return os.environ.get(name, default) not in ("0", "false", "")
|
||||
|
||||
|
||||
# Defaults aus den Env-Vars — Quelle der Wahrheit, solange keine Policy-Datei existiert.
|
||||
DEFAULTS: dict = {
|
||||
"fast": os.environ.get("MC_ROUTE_FAST", "fast"),
|
||||
"heavy": os.environ.get("MC_ROUTE_HEAVY", "heavy"),
|
||||
"coder": os.environ.get("MC_ROUTE_CODER", "coder"),
|
||||
"coder_lite": os.environ.get("MC_ROUTE_CODER_LITE", "").strip(),
|
||||
"heavy_chars": int(os.environ.get("MC_GATEWAY_HEAVY_CHARS", "8000")),
|
||||
"coding_escalate_chars": int(os.environ.get("MC_CODING_ESCALATE_CHARS", "120000")),
|
||||
"fast_no_think": _env_bool("MC_FAST_NO_THINK", "1"),
|
||||
}
|
||||
|
||||
# Feld-Spezifikation für die UI (Typ + Grenzen + Label). Treibt Editor & Validierung.
|
||||
FIELDS: list[dict] = [
|
||||
{"key": "fast", "label": "fast-Alias (chat: Standard)", "type": "str"},
|
||||
{"key": "heavy", "label": "heavy-Alias (chat: lang/komplex)", "type": "str"},
|
||||
{"key": "coder", "label": "coder-Alias (coding: stark / Eskalation)", "type": "str"},
|
||||
{"key": "coder_lite", "label": "coder-lite-Alias (coding: schneller Default; leer = aus)", "type": "str"},
|
||||
{"key": "heavy_chars", "label": "chat → heavy ab N Zeichen", "type": "int", "min": 500, "max": 1_000_000},
|
||||
{"key": "coding_escalate_chars", "label": "coding → starker Coder ab N Zeichen", "type": "int", "min": 1000, "max": 4_000_000},
|
||||
{"key": "fast_no_think", "label": "fast-Spur: Thinking aus (flotte Antworten)", "type": "bool"},
|
||||
]
|
||||
|
||||
_LOCK = threading.Lock()
|
||||
_CACHE: dict = {"mtime": None, "policy": None}
|
||||
|
||||
|
||||
def _read_file() -> dict:
|
||||
try:
|
||||
with open(POLICY_PATH, "r", encoding="utf-8") as f:
|
||||
data = json.load(f)
|
||||
return data if isinstance(data, dict) else {}
|
||||
except (FileNotFoundError, json.JSONDecodeError, OSError):
|
||||
return {}
|
||||
|
||||
|
||||
def _coerce(patch: dict) -> dict:
|
||||
"""Nur bekannte Keys, typ-/bereichsvalidiert. Wirft ValueError bei ungültigen Werten."""
|
||||
spec = {f["key"]: f for f in FIELDS}
|
||||
out: dict = {}
|
||||
for k, v in (patch or {}).items():
|
||||
f = spec.get(k)
|
||||
if not f:
|
||||
continue # unbekannte Keys still verwerfen
|
||||
if f["type"] == "int":
|
||||
iv = int(v)
|
||||
lo, hi = f.get("min", 1), f.get("max", 10**9)
|
||||
if not (lo <= iv <= hi):
|
||||
raise ValueError(f"{k}={iv} außerhalb [{lo}, {hi}]")
|
||||
out[k] = iv
|
||||
elif f["type"] == "bool":
|
||||
out[k] = bool(v)
|
||||
else: # str
|
||||
sv = str(v).strip()
|
||||
if k != "coder_lite" and not sv:
|
||||
raise ValueError(f"{k} darf nicht leer sein")
|
||||
out[k] = sv
|
||||
return out
|
||||
|
||||
|
||||
def _coerce_safe(patch: dict) -> dict:
|
||||
"""Wie _coerce, aber schluckt Fehler — kaputte Datei darf den Betrieb nicht stoppen."""
|
||||
try:
|
||||
return _coerce(patch)
|
||||
except (ValueError, TypeError):
|
||||
return {}
|
||||
|
||||
|
||||
def load_policy() -> dict:
|
||||
"""Aktuelle Policy (Datei über DEFAULTS gemerged). Hot-reload via mtime-Cache, pro Request billig."""
|
||||
try:
|
||||
mtime = POLICY_PATH.stat().st_mtime
|
||||
except OSError:
|
||||
mtime = None
|
||||
with _LOCK:
|
||||
if _CACHE["policy"] is None or _CACHE["mtime"] != mtime:
|
||||
merged = {**DEFAULTS}
|
||||
if mtime is not None:
|
||||
merged.update(_coerce_safe(_read_file()))
|
||||
_CACHE["mtime"] = mtime
|
||||
_CACHE["policy"] = merged
|
||||
return dict(_CACHE["policy"])
|
||||
|
||||
|
||||
def save_policy(patch: dict) -> dict:
|
||||
"""Validiert + persistiert atomar. Gibt die neue, vollständige Policy zurück."""
|
||||
clean = _coerce(patch) # wirft bei ungültigem Input
|
||||
with _LOCK:
|
||||
current = {**DEFAULTS, **_coerce_safe(_read_file()), **clean}
|
||||
POLICY_PATH.parent.mkdir(parents=True, exist_ok=True)
|
||||
tmp = POLICY_PATH.with_suffix(".json.tmp")
|
||||
with open(tmp, "w", encoding="utf-8") as f:
|
||||
json.dump(current, f, ensure_ascii=False, indent=2)
|
||||
os.replace(tmp, POLICY_PATH)
|
||||
_CACHE["mtime"] = None # nächster load_policy() lädt frisch
|
||||
_CACHE["policy"] = None
|
||||
return current
|
||||
|
||||
|
||||
def policy_meta() -> dict:
|
||||
"""Für den UI-Editor: aktuelle Werte + Defaults (für „Zurücksetzen“) + Feld-Spezifikation."""
|
||||
return {"policy": load_policy(), "defaults": dict(DEFAULTS), "fields": FIELDS}
|
||||
@@ -0,0 +1,141 @@
|
||||
"""
|
||||
Health-Wächter der Box (Lucy-Proaktivität, Faden A3).
|
||||
|
||||
Prüft periodisch die Kern-Dienste (Engine, Agent-Hirn, Hermes-Gateway, Mem0,
|
||||
Voice-Sidecar, Platte) und meldet ZUSTANDSWECHSEL in den Melde-Briefkasten
|
||||
(services/announce.py → Lucy spricht es) und via notify.sh (Telegram).
|
||||
|
||||
Flankenerkennung statt Dauerfeuer: Alarm erst nach FAIL_AFTER Fehl-Ticks in
|
||||
Folge (überlebt Neustarts/Update-Fenster), Entwarnung beim ersten grünen Tick
|
||||
nach einem Alarm. Hält ein Problem an, wird frühestens nach REMIND_S erinnert.
|
||||
|
||||
Braucht KEIN sudo, keine neuen Dienste — läuft als asyncio-Task im MC2-Backend
|
||||
(wie der Re-Warm-Wächter). Abschaltbar via MC_SENTRY_ENABLED=0.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
import os
|
||||
import time
|
||||
|
||||
import httpx
|
||||
import psutil
|
||||
|
||||
from config import HERMES_API_URL, LLAMA_SWAP_URL, MEM0_SERVICE_URL, MODELS_DIR, VOICE_SERVICE_URL
|
||||
from services import announce, llamaswap
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
ENABLED = os.environ.get("MC_SENTRY_ENABLED", "1") != "0"
|
||||
INTERVAL = int(os.environ.get("MC_SENTRY_INTERVAL", "120")) # Sekunden zwischen Ticks
|
||||
START_DELAY = int(os.environ.get("MC_SENTRY_START_DELAY", "90")) # Dienste nach Boot setzen lassen
|
||||
FAIL_AFTER = int(os.environ.get("MC_SENTRY_FAIL_AFTER", "3")) # Fehl-Ticks bis Alarm (3×120s = 6 min)
|
||||
REMIND_S = int(os.environ.get("MC_SENTRY_REMIND_S", "21600")) # Erinnerung bei Dauerproblem: 6 h
|
||||
DISK_ALARM_PCT = float(os.environ.get("MC_SENTRY_DISK_PCT", "90"))
|
||||
|
||||
|
||||
def _reach(url: str, path: str = "/health") -> bool:
|
||||
try:
|
||||
with httpx.Client(timeout=5.0) as c:
|
||||
return c.get(f"{url}{path}").status_code < 500
|
||||
except Exception:
|
||||
return False
|
||||
|
||||
|
||||
def _check_engine() -> bool:
|
||||
return llamaswap.engine_reachable()
|
||||
|
||||
|
||||
def _check_brain() -> bool:
|
||||
"""Hirn tot? Verdrängung durch ein ANDERES laufendes Modell (IDE-Last) ist NORMAL —
|
||||
Alarm nur, wenn gar nichts läuft und das Hirn trotz Re-Warm-Wächter kalt bleibt."""
|
||||
st = llamaswap.brain_status()
|
||||
if st.get("ready"):
|
||||
return True
|
||||
return bool(llamaswap.get_running_models()) # anderes Modell aktiv → Verdrängung, kein Defekt
|
||||
|
||||
|
||||
def _check_disk() -> bool:
|
||||
try:
|
||||
return psutil.disk_usage(str(MODELS_DIR) if MODELS_DIR.exists() else os.getcwd()).percent < DISK_ALARM_PCT
|
||||
except Exception:
|
||||
return True # kein Messwert ≠ Alarm
|
||||
|
||||
|
||||
# name → (Checker, Alarm-Text, Entwarnungs-Text) — Texte sind Lucy-sprechbar (kurz, Alltagssprache).
|
||||
CHECKS: dict[str, tuple] = {
|
||||
"engine": (_check_engine,
|
||||
"Die Modell-Engine antwortet nicht mehr. Ohne sie laufen keine KI-Modelle.",
|
||||
"Die Modell-Engine ist wieder da."),
|
||||
"brain": (_check_brain,
|
||||
"Mein Gehirn lädt nicht — ich kann gerade nicht richtig denken. Ein Neustart der Engine könnte helfen.",
|
||||
"Mein Gehirn ist wieder geladen. Alles klar bei mir."),
|
||||
"hermes": (lambda: _reach(HERMES_API_URL),
|
||||
"Der Agent-Dienst ist ausgefallen — Telegram und meine Tools gehen gerade nicht.",
|
||||
"Der Agent-Dienst läuft wieder."),
|
||||
"mem0": (lambda: _reach(MEM0_SERVICE_URL),
|
||||
"Mein Gedächtnis-Dienst ist ausgefallen — ich merke mir vorübergehend nichts Neues.",
|
||||
"Mein Gedächtnis ist wieder online."),
|
||||
"voice": (lambda: _reach(VOICE_SERVICE_URL),
|
||||
"Der Hör-Dienst auf der Box ist ausgefallen — Spracheingabe könnte haken.",
|
||||
"Der Hör-Dienst läuft wieder."),
|
||||
"disk": (_check_disk,
|
||||
f"Die Platte der Box ist zu über {DISK_ALARM_PCT:.0f} Prozent voll. Es wird eng für Modelle und Backups.",
|
||||
"Die Platte hat wieder genug Luft."),
|
||||
}
|
||||
|
||||
|
||||
class _Watch:
|
||||
__slots__ = ("fails", "alerted", "alert_ts")
|
||||
|
||||
def __init__(self) -> None:
|
||||
self.fails = 0 # Fehl-Ticks in Folge
|
||||
self.alerted = False # Alarm ist raus, Entwarnung steht aus
|
||||
self.alert_ts = 0.0 # Zeitpunkt des letzten Alarms (für REMIND_S)
|
||||
|
||||
|
||||
_watches: dict[str, _Watch] = {name: _Watch() for name in CHECKS}
|
||||
|
||||
|
||||
def _notify_telegram(subject: str, text: str) -> None:
|
||||
"""Telegram-Direktweg (gemeinsamer Helper in announce.py); Briefkasten-Eintrag
|
||||
legt der Wächter selbst ab."""
|
||||
announce.notify_telegram(subject, text + " (Diese Meldung kam auch an Lucy.)")
|
||||
|
||||
|
||||
def _tick() -> None:
|
||||
now = time.time()
|
||||
for name, (check, fail_msg, ok_msg) in CHECKS.items():
|
||||
w = _watches[name]
|
||||
try:
|
||||
ok = bool(check())
|
||||
except Exception:
|
||||
ok = False
|
||||
if ok:
|
||||
w.fails = 0
|
||||
if w.alerted:
|
||||
w.alerted = False
|
||||
announce.add(ok_msg, subject="[Box wieder ok]", source="sentry")
|
||||
_notify_telegram("[Box wieder ok]", ok_msg)
|
||||
continue
|
||||
w.fails += 1
|
||||
due = (not w.alerted and w.fails >= FAIL_AFTER) or (w.alerted and now - w.alert_ts >= REMIND_S)
|
||||
if due:
|
||||
prefix = "" if not w.alerted else "Immer noch: "
|
||||
w.alerted = True
|
||||
w.alert_ts = now
|
||||
announce.add(prefix + fail_msg, subject="[Box-Problem]", source="sentry")
|
||||
_notify_telegram("[Box-Problem]", prefix + fail_msg)
|
||||
log.warning("sentry: %s ALARM (%s Fehl-Ticks)", name, w.fails)
|
||||
|
||||
|
||||
async def sentry_loop() -> None:
|
||||
"""Endlos-Schleife (Hintergrund-Task im MC2-Lifespan)."""
|
||||
await asyncio.sleep(START_DELAY)
|
||||
log.info("sentry: Health-Wächter aktiv (Intervall %ss, Alarm nach %s Fehl-Ticks)", INTERVAL, FAIL_AFTER)
|
||||
while True:
|
||||
try:
|
||||
await asyncio.to_thread(_tick)
|
||||
except Exception:
|
||||
log.debug("sentry: Tick fehlgeschlagen", exc_info=True)
|
||||
await asyncio.sleep(INTERVAL)
|
||||
@@ -0,0 +1,35 @@
|
||||
"""
|
||||
Vertrauenswürdige Quellen + Kategorien für die automatische Modell-Entdeckung.
|
||||
Rollen = die EINE Quelle der Wahrheit, identisch zu den llama-swap-Serving-Rollen
|
||||
und der UI: fast · heavy · coder · vision · hermes (Agent-Hirn) · scout.
|
||||
"""
|
||||
|
||||
# HF-Orgs, die zuverlässig aktuelle, hochwertige GGUF-Quants veröffentlichen.
|
||||
TRUSTED_AUTHORS = ["unsloth", "bartowski", "ggml-org", "lmstudio-community"]
|
||||
|
||||
# Kanonische Rollen — eine Quelle der Wahrheit (deckt sich mit llamaswap.ROLE_IDS,
|
||||
# maintenance.ROLE_MAP, frontend ModelBadges.ROLES + Discover.ROLE_METADATA).
|
||||
# `hermes` = Lucys Agent-Hirn (warm + ko-resident); UI-Label „Hirn".
|
||||
ROLE_IDS = ["fast", "heavy", "coder", "vision", "hermes", "scout"]
|
||||
|
||||
# Kategorien (Reihenfolge = Anzeige + Zuordnungs-Priorität). Ein Modell wird der
|
||||
# ERSTEN Kategorie zugeordnet, deren Stichwort im Repo-Namen vorkommt; sonst „scout".
|
||||
# Die `role` ist zugleich der Alias-Vorschlag und EINE der 5 kanonischen Rollen.
|
||||
CATEGORIES = [
|
||||
{"role": "vision", "title": "Bilder verstehen", "icon": "eye",
|
||||
"kw": ["-vl-", "-vl", "vision", "llava", "multimodal", "-mm-", "pixtral"]},
|
||||
{"role": "coder", "title": "Coden & Programmieren", "icon": "code",
|
||||
"kw": ["coder", "-code-", "code-", "codestral", "starcoder"]},
|
||||
{"role": "hermes", "title": "Lucys Hirn (Agent)", "icon": "brain-circuit",
|
||||
"kw": ["hermes"]}, # Agent-Hirn: Hermes-Familie am robustesten (natives Tool-Calling).
|
||||
{"role": "heavy", "title": "Schweres Reasoning", "icon": "brain",
|
||||
"kw": ["reasoning", "-think", "thinking", "gpt-oss", "deepseek-r", "-r1", "qwq",
|
||||
"-70b", "-72b", "-120b", "-123b", "-235b", "-405b", "-a10b", "-a22b"]},
|
||||
{"role": "fast", "title": "Schnelles Alltags-Hirn", "icon": "zap",
|
||||
"kw": ["-a3b", "-a1", "-a2", "-30b", "-32b", "-14b", "-8b", "-7b", "-4b", "-moe"]},
|
||||
{"role": "scout", "title": "Multimodal-Allrounder", "icon": "compass",
|
||||
"kw": []}, # Fallback: instruct/chat-Modelle, die in keine Spezialrolle fallen
|
||||
]
|
||||
|
||||
# Repo-Namensteile, die bei der Entdeckung übersprungen werden (Roh-/Spezialformate).
|
||||
SKIP_TOKENS = ["-base", "-bnb-", "-gptq", "-awq", "-fp8", "draft", "tokenizer"]
|
||||
@@ -0,0 +1,195 @@
|
||||
"""
|
||||
System/OS-Metriken für die Box (Bosgame / Strix Halo).
|
||||
|
||||
CPU/RAM/Disk via psutil (plattformübergreifend). GPU-Auslastung/VRAM/Temperatur
|
||||
via sysfs (amdgpu) — nur Linux; auf anderen Plattformen None (amd-smi fehlt auf
|
||||
der Box, daher sysfs). Verschachtelte Struktur wie v1 (cpu.percent, ram.used Bytes).
|
||||
"""
|
||||
|
||||
import glob
|
||||
import os
|
||||
import subprocess
|
||||
|
||||
import psutil
|
||||
|
||||
from config import MODELS_DIR
|
||||
|
||||
|
||||
def _read_int(path: str) -> int | None:
|
||||
try:
|
||||
with open(path) as f:
|
||||
return int(f.read().strip())
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
|
||||
def _gpu_sysfs() -> dict | None:
|
||||
"""AMD-GPU-Auslastung + Speicher via sysfs (Linux). Findet die Basis-Card
|
||||
dynamisch (Strix Halo ist oft card1, nicht card0) und überspringt die
|
||||
Connector-Verzeichnisse (card1-DP-1 …). Strix Halo nutzt Unified Memory →
|
||||
GTT ist der eigentliche große Pool; VRAM ist nur der kleine Carve-out."""
|
||||
for dev in sorted(glob.glob("/sys/class/drm/card*/device")):
|
||||
card = dev.split("/")[-2] # z.B. "card1" oder "card1-DP-1"
|
||||
if "-" in card: # Connector-Dir → kein GPU-Device
|
||||
continue
|
||||
busy = _read_int(f"{dev}/gpu_busy_percent")
|
||||
if busy is None:
|
||||
continue
|
||||
return {
|
||||
"busy_percent": busy,
|
||||
"vram_used": _read_int(f"{dev}/mem_info_vram_used"),
|
||||
"vram_total": _read_int(f"{dev}/mem_info_vram_total"),
|
||||
"gtt_used": _read_int(f"{dev}/mem_info_gtt_used"),
|
||||
"gtt_total": _read_int(f"{dev}/mem_info_gtt_total"),
|
||||
}
|
||||
return None
|
||||
|
||||
|
||||
def _temps() -> dict | None:
|
||||
"""CPU/GPU-Temperatur via hwmon (Linux). None bei Fehlen."""
|
||||
out: dict = {}
|
||||
for hw in glob.glob("/sys/class/hwmon/hwmon*"):
|
||||
name = ""
|
||||
try:
|
||||
with open(f"{hw}/name") as f:
|
||||
name = f.read().strip()
|
||||
except Exception:
|
||||
continue
|
||||
t = _read_int(f"{hw}/temp1_input")
|
||||
if t is None:
|
||||
continue
|
||||
c = round(t / 1000.0, 1)
|
||||
if name in ("k10temp", "zenpower", "coretemp"):
|
||||
out["cpu"] = c
|
||||
elif name in ("amdgpu", "edge"):
|
||||
out["gpu"] = c
|
||||
return out or None
|
||||
|
||||
|
||||
def get_git_info(path: str) -> dict | None:
|
||||
expanded = os.path.expanduser(path)
|
||||
if not os.path.isdir(expanded) or not os.path.exists(os.path.join(expanded, ".git")):
|
||||
return None
|
||||
try:
|
||||
res = subprocess.run(
|
||||
["git", "log", "-1", "--format=%h|%cd|%s", "--date=short"],
|
||||
cwd=expanded, capture_output=True, text=True, timeout=3
|
||||
)
|
||||
if res.returncode != 0:
|
||||
return None
|
||||
parts = res.stdout.strip().split("|", 2)
|
||||
h = parts[0]
|
||||
d = parts[1]
|
||||
s = parts[2] if len(parts) > 2 else ""
|
||||
|
||||
branch_res = subprocess.run(
|
||||
["git", "rev-parse", "--abbrev-ref", "HEAD"],
|
||||
cwd=expanded, capture_output=True, text=True, timeout=2
|
||||
)
|
||||
branch = branch_res.stdout.strip() if branch_res.returncode == 0 else "unknown"
|
||||
|
||||
status_res = subprocess.run(
|
||||
["git", "status", "--porcelain"],
|
||||
cwd=expanded, capture_output=True, text=True, timeout=2
|
||||
)
|
||||
dirty = bool(status_res.stdout.strip()) if status_res.returncode == 0 else False
|
||||
|
||||
return {
|
||||
"hash": h,
|
||||
"date": d,
|
||||
"subject": s,
|
||||
"branch": branch,
|
||||
"dirty": dirty,
|
||||
"path": expanded
|
||||
}
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
|
||||
def find_hermes_agent_git() -> dict | None:
|
||||
env_path = os.environ.get("MC_HERMES_AGENT_PATH")
|
||||
if env_path:
|
||||
info = get_git_info(env_path)
|
||||
if info:
|
||||
return info
|
||||
|
||||
candidates = [
|
||||
"~/hermes-agent",
|
||||
"~/.hermes/hermes-agent",
|
||||
"~/hermes-webui/hermes-agent",
|
||||
"~/.hermes"
|
||||
]
|
||||
for c in candidates:
|
||||
info = get_git_info(c)
|
||||
if info:
|
||||
return info
|
||||
return None
|
||||
|
||||
|
||||
def get_engine_version() -> dict:
|
||||
engine_path = os.environ.get("MC_ENGINE_PATH", "/opt/llamacpp")
|
||||
git_info = get_git_info(engine_path)
|
||||
if git_info:
|
||||
return {**git_info, "type": "git"}
|
||||
|
||||
candidates = [
|
||||
os.path.join(engine_path, "llama-server"),
|
||||
os.path.join(engine_path, "bin", "llama-server"),
|
||||
"/usr/local/bin/llama-server",
|
||||
"/usr/bin/llama-server",
|
||||
"llama-server"
|
||||
]
|
||||
|
||||
for binary in candidates:
|
||||
if binary != "llama-server" and not os.path.exists(binary):
|
||||
continue
|
||||
try:
|
||||
res = subprocess.run([binary, "--version"], capture_output=True, text=True, timeout=2)
|
||||
output = (res.stdout or "").strip() or (res.stderr or "").strip()
|
||||
if output:
|
||||
lines = output.splitlines()
|
||||
ver = lines[0] if lines else "unknown"
|
||||
return {"version_text": ver, "type": "binary"}
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
return {"type": "unknown"}
|
||||
|
||||
|
||||
_VERSION_CACHE = {"ts": 0.0, "data": {}}
|
||||
|
||||
|
||||
def check_versions_cached() -> dict:
|
||||
import time
|
||||
now = time.time()
|
||||
if now - _VERSION_CACHE["ts"] < 30.0:
|
||||
return _VERSION_CACHE["data"]
|
||||
|
||||
mc2_path = os.path.abspath(os.path.join(os.path.dirname(__file__), "..", ".."))
|
||||
|
||||
data = {
|
||||
"mc2": get_git_info(mc2_path),
|
||||
"engine": get_engine_version(),
|
||||
"hermes_ui": get_git_info("~/hermes-webui"),
|
||||
"hermes_agent": find_hermes_agent_git()
|
||||
}
|
||||
_VERSION_CACHE["ts"] = now
|
||||
_VERSION_CACHE["data"] = data
|
||||
return data
|
||||
|
||||
|
||||
def system_status() -> dict:
|
||||
vm = psutil.virtual_memory()
|
||||
try:
|
||||
du = psutil.disk_usage(str(MODELS_DIR) if MODELS_DIR.exists() else os.getcwd())
|
||||
disk = {"total": du.total, "used": du.used, "percent": du.percent}
|
||||
except Exception:
|
||||
disk = None
|
||||
return {
|
||||
"cpu": {"percent": psutil.cpu_percent(interval=0.1), "cores": psutil.cpu_count()},
|
||||
"ram": {"total": vm.total, "used": vm.used, "percent": vm.percent},
|
||||
"gpu": _gpu_sysfs(),
|
||||
"temp": _temps(),
|
||||
"disk": disk,
|
||||
"versions": check_versions_cached(),
|
||||
}
|
||||
@@ -0,0 +1,99 @@
|
||||
"""Token-Statistik (Verbrauch je Modell) mit gedrosseltem Persistieren.
|
||||
|
||||
Früher wurde bei JEDEM Request die komplette JSON-Datei gelesen und geschrieben
|
||||
(Disk-Thrash). Jetzt: einmaliges Laden in einen In-Memory-Cache, Inkremente laufen
|
||||
gegen den Cache, Persistieren passiert höchstens alle FLUSH_INTERVAL Sekunden sowie
|
||||
beim Prozess-Ende (atexit). Lesen liefert immer den aktuellen (auch ungeflushten) Stand.
|
||||
"""
|
||||
|
||||
import atexit
|
||||
import json
|
||||
import logging
|
||||
import threading
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
from config import HERMES_HOME
|
||||
|
||||
STATS_FILE = HERMES_HOME / "token_stats.json"
|
||||
FLUSH_INTERVAL = 5.0 # Sekunden zwischen Disk-Writes
|
||||
# Baseline (repräsentiert Verbrauch vor dem modellspezifischen Logging).
|
||||
_BASELINE = {"prompt_tokens": 718400, "completion_tokens": 324200, "models": {}}
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
_lock = threading.Lock()
|
||||
_stats: dict | None = None
|
||||
_dirty = False
|
||||
_last_flush = 0.0
|
||||
|
||||
|
||||
def _load_from_disk() -> dict:
|
||||
if not STATS_FILE.exists():
|
||||
return dict(_BASELINE)
|
||||
try:
|
||||
with open(STATS_FILE, "r", encoding="utf-8") as f:
|
||||
data = json.load(f)
|
||||
data.setdefault("prompt_tokens", 0)
|
||||
data.setdefault("completion_tokens", 0)
|
||||
data.setdefault("models", {})
|
||||
return data
|
||||
except (OSError, json.JSONDecodeError):
|
||||
log.warning("token_stats: Laden fehlgeschlagen, nutze Baseline", exc_info=True)
|
||||
return dict(_BASELINE)
|
||||
|
||||
|
||||
def _ensure_loaded() -> dict:
|
||||
global _stats
|
||||
if _stats is None:
|
||||
_stats = _load_from_disk()
|
||||
return _stats
|
||||
|
||||
|
||||
def _write(stats: dict) -> None:
|
||||
try:
|
||||
STATS_FILE.parent.mkdir(parents=True, exist_ok=True)
|
||||
tmp = STATS_FILE.with_suffix(".tmp")
|
||||
with open(tmp, "w", encoding="utf-8") as f:
|
||||
json.dump(stats, f)
|
||||
tmp.replace(STATS_FILE)
|
||||
except OSError:
|
||||
log.warning("token_stats: Schreiben fehlgeschlagen", exc_info=True)
|
||||
|
||||
|
||||
def get_stats() -> dict:
|
||||
"""Aktueller Stand (inkl. noch nicht geflushter Inkremente) als Kopie."""
|
||||
with _lock:
|
||||
return json.loads(json.dumps(_ensure_loaded()))
|
||||
|
||||
|
||||
def increment_tokens(prompt: int, completion: int, model: str | None = None) -> None:
|
||||
"""Tokens im Cache verbuchen; gedrosselt auf Disk persistieren."""
|
||||
global _dirty, _last_flush
|
||||
with _lock:
|
||||
stats = _ensure_loaded()
|
||||
stats["prompt_tokens"] += prompt
|
||||
stats["completion_tokens"] += completion
|
||||
if model:
|
||||
m = stats.setdefault("models", {}).setdefault(
|
||||
model.lower(), {"prompt": 0, "completion": 0})
|
||||
m["prompt"] += prompt
|
||||
m["completion"] += completion
|
||||
_dirty = True
|
||||
now = time.monotonic()
|
||||
if now - _last_flush >= FLUSH_INTERVAL:
|
||||
_write(stats)
|
||||
_dirty = False
|
||||
_last_flush = now
|
||||
|
||||
|
||||
def flush() -> None:
|
||||
"""Ungeschriebene Inkremente sofort persistieren (z.B. beim Shutdown)."""
|
||||
global _dirty
|
||||
with _lock:
|
||||
if _dirty and _stats is not None:
|
||||
_write(_stats)
|
||||
_dirty = False
|
||||
|
||||
|
||||
atexit.register(flush)
|
||||
@@ -0,0 +1,68 @@
|
||||
"""
|
||||
Per-Stage-Latenz-Metriken für die Voice/Lucy-Pipeline.
|
||||
|
||||
Misst die Server-seitige Dauer jeder Stufe (STT, Vision-Beschreibung, Chat-TTFB, TTS) und hält
|
||||
rollende Statistiken (avg/p50/p95/last) im Speicher. Macht aus Latenz-VERMUTUNGEN gemessene Fakten
|
||||
— die eigentliche Voraussetzung, um gezielt zu optimieren (Stufe 5/C2 des Reviews). Anzeige im
|
||||
Frontend-Overhaul (E) analog zur TokenPerformanceCard.
|
||||
|
||||
In-Memory + thread-safe (keine Datei-I/O — Latenz-Telemetrie ist transient, Restart = Reset).
|
||||
"""
|
||||
|
||||
import threading
|
||||
import time
|
||||
from collections import deque
|
||||
|
||||
_LOCK = threading.Lock()
|
||||
_MAX = 200
|
||||
_STAGES: dict[str, deque] = {}
|
||||
|
||||
# Bekannte Stufen (für stabile UI-Reihenfolge); unbekannte werden trotzdem erfasst.
|
||||
STAGES = ("stt", "vision", "chat_ttfb", "tts")
|
||||
|
||||
|
||||
def record_stage(stage: str, ms: float) -> None:
|
||||
"""Eine gemessene Stage-Dauer (ms) verbuchen. No-op bei negativen Werten."""
|
||||
if ms is None or ms < 0:
|
||||
return
|
||||
with _LOCK:
|
||||
dq = _STAGES.get(stage)
|
||||
if dq is None:
|
||||
dq = _STAGES[stage] = deque(maxlen=_MAX)
|
||||
dq.append(float(ms))
|
||||
|
||||
|
||||
class Timer:
|
||||
"""Context-Manager: misst die verstrichene Zeit und verbucht sie auf `stage`.
|
||||
Funktioniert um `await`-Aufrufe herum (enter → await → exit)."""
|
||||
|
||||
def __init__(self, stage: str) -> None:
|
||||
self.stage = stage
|
||||
self._t0 = 0.0
|
||||
|
||||
def __enter__(self) -> "Timer":
|
||||
self._t0 = time.perf_counter()
|
||||
return self
|
||||
|
||||
def __exit__(self, *exc) -> None:
|
||||
record_stage(self.stage, (time.perf_counter() - self._t0) * 1000.0)
|
||||
|
||||
|
||||
def _summary(vals: list[float]) -> dict:
|
||||
if not vals:
|
||||
return {"count": 0}
|
||||
s = sorted(vals)
|
||||
n = len(s)
|
||||
return {
|
||||
"count": n,
|
||||
"avg_ms": round(sum(s) / n, 1),
|
||||
"p50_ms": round(s[n // 2], 1),
|
||||
"p95_ms": round(s[min(n - 1, int(n * 0.95))], 1),
|
||||
"last_ms": round(vals[-1], 1),
|
||||
}
|
||||
|
||||
|
||||
def get_metrics() -> dict:
|
||||
"""Rollende Zusammenfassung je Stufe."""
|
||||
with _LOCK:
|
||||
return {stage: _summary(list(dq)) for stage, dq in _STAGES.items()}
|
||||
@@ -0,0 +1,105 @@
|
||||
"""
|
||||
Hält das Agent-Hirn (Rolle `hermes`) dauerhaft warm.
|
||||
|
||||
Hintergrund: llama-swap ist EIN-Gruppen-resident — lädt ein on-demand-Modell außerhalb
|
||||
der `brains`-Gruppe, wird die ganze Gruppe (inkl. Hirn) verdrängt. `persist: true`
|
||||
verhindert nur Idle-Unload, NICHT die Gruppen-Verdrängung; neu vorgewärmt wird sonst erst
|
||||
beim nächsten llama-swap-(Re)Start. Dieser Wächter schließt die Lücke: ist die Box idle
|
||||
(nichts geladen), pingt er das Hirn vor. Während aktiver Last (irgendetwas geladen) hält
|
||||
er sich raus, verdrängt also nie ein gerade genutztes Modell.
|
||||
|
||||
Abschaltbar/justierbar via Env: MC_REWARM_ENABLED=0, MC_REWARM_INTERVAL, MC_REWARM_MODEL.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
import os
|
||||
|
||||
import httpx
|
||||
|
||||
from config import LLAMA_SWAP_URL
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
ENABLED = os.environ.get("MC_REWARM_ENABLED", "1") != "0"
|
||||
INTERVAL = int(os.environ.get("MC_REWARM_INTERVAL", "90")) # Sekunden zwischen Checks
|
||||
START_DELAY = int(os.environ.get("MC_REWARM_START_DELAY", "25"))
|
||||
NUDGE_GRACE = int(os.environ.get("MC_REWARM_NUDGE_GRACE", "3")) # llama-swap den Reload abschließen lassen
|
||||
|
||||
# Wecksignal für einen sofortigen Vorwärm-Check (statt bis zum nächsten INTERVAL-Tick zu warten).
|
||||
# Wird von write_config() nach einer Config-Änderung gesetzt: llama-swap (-watch-config) lädt die
|
||||
# neue Config und verwirft dabei ALLE Modelle inkl. Hirn — ohne Nudge bliebe es bis zu INTERVAL
|
||||
# Sekunden kalt liegen, bis der nächste Tick oder eine Anfrage es wieder lädt.
|
||||
_loop: asyncio.AbstractEventLoop | None = None
|
||||
_wake: asyncio.Event | None = None
|
||||
|
||||
|
||||
def nudge() -> None:
|
||||
"""Threadsicher: bittet den Wächter, nach einem Config-Reload bald vorzuwärmen. No-op,
|
||||
solange der Wächter (noch) nicht läuft."""
|
||||
if _loop is not None and _wake is not None and not _loop.is_closed():
|
||||
try:
|
||||
_loop.call_soon_threadsafe(_wake.set)
|
||||
except RuntimeError:
|
||||
pass
|
||||
|
||||
|
||||
def _brain_model() -> str:
|
||||
"""Aktives Agent-Hirn = Hermes' model.default (sonst model.model). So wärmt der Wächter
|
||||
immer das WIRKLICH genutzte Hirn (passt sich Hirn-Wechseln an). Override: MC_REWARM_MODEL."""
|
||||
forced = os.environ.get("MC_REWARM_MODEL")
|
||||
if forced:
|
||||
return forced
|
||||
try:
|
||||
from ruamel.yaml import YAML
|
||||
from config import HERMES_HOME
|
||||
p = HERMES_HOME / "config.yaml"
|
||||
if p.exists():
|
||||
with p.open(encoding="utf-8") as f:
|
||||
cfg = YAML().load(f) or {}
|
||||
m = (cfg.get("model") or {}) if isinstance(cfg, dict) else {}
|
||||
v = m.get("default") or m.get("model")
|
||||
if v:
|
||||
return str(v)
|
||||
except Exception:
|
||||
log.debug("rewarm: Hirn-Lookup fehlgeschlagen", exc_info=True)
|
||||
return "fast"
|
||||
|
||||
|
||||
async def _running_empty() -> bool:
|
||||
async with httpx.AsyncClient(timeout=8.0) as c:
|
||||
r = await c.get(f"{LLAMA_SWAP_URL}/running")
|
||||
data = r.json() or {}
|
||||
return not (data.get("running") or [])
|
||||
|
||||
|
||||
async def _warm(model: str) -> None:
|
||||
async with httpx.AsyncClient(timeout=180.0) as c:
|
||||
await c.post(f"{LLAMA_SWAP_URL}/v1/chat/completions", json={
|
||||
"model": model, "max_tokens": 1,
|
||||
"messages": [{"role": "user", "content": "ping"}],
|
||||
})
|
||||
|
||||
|
||||
async def rewarm_loop() -> None:
|
||||
"""Endlos-Schleife (Hintergrund-Task): prüft periodisch (und sofort nach einem Config-
|
||||
Reload-Nudge), wärmt bei Idle vor."""
|
||||
global _loop, _wake
|
||||
_loop = asyncio.get_running_loop()
|
||||
_wake = asyncio.Event()
|
||||
await asyncio.sleep(START_DELAY) # Box/Engine nach MC-Start setzen lassen
|
||||
while True:
|
||||
try:
|
||||
if await _running_empty():
|
||||
model = _brain_model()
|
||||
log.info("rewarm: Box idle → Hirn '%s' wird vorgewärmt", model)
|
||||
await _warm(model)
|
||||
except Exception:
|
||||
log.debug("rewarm: Tick fehlgeschlagen", exc_info=True)
|
||||
# Bis zum nächsten Tick warten ODER sofort auf einen Config-Reload-Nudge reagieren.
|
||||
try:
|
||||
await asyncio.wait_for(_wake.wait(), timeout=INTERVAL)
|
||||
_wake.clear()
|
||||
await asyncio.sleep(NUDGE_GRACE) # llama-swap den Reload abschließen lassen
|
||||
except asyncio.TimeoutError:
|
||||
pass
|
||||
@@ -0,0 +1,3 @@
|
||||
.venv/
|
||||
__pycache__/
|
||||
*.pyc
|
||||
@@ -0,0 +1,207 @@
|
||||
"""
|
||||
Hermes PC Executor — läuft auf dem Windows PC.
|
||||
Empfängt Tool-Befehle vom MCP-Server der AI Box und führt sie lokal aus.
|
||||
"""
|
||||
import base64
|
||||
import hmac
|
||||
import io
|
||||
import os
|
||||
import socket
|
||||
import subprocess
|
||||
import webbrowser
|
||||
from urllib.parse import quote
|
||||
|
||||
import uvicorn
|
||||
from fastapi import Depends, FastAPI, Header, HTTPException
|
||||
|
||||
app = FastAPI(title="Hermes PC Executor")
|
||||
|
||||
MY_PORT = int(os.environ.get("HERMES_PC_PORT", "7777"))
|
||||
# Auf welchem Interface lauschen. Default 0.0.0.0 (LAN), per Env einschränkbar (z.B. die LAN-IP des PCs).
|
||||
MY_HOST = os.environ.get("HERMES_PC_HOST", "0.0.0.0")
|
||||
|
||||
# ── Auth ───────────────────────────────────────────────────────────────────
|
||||
# Shared Secret. OHNE Token sind die gefährlichen Endpunkte (Shell/Eingabe/Öffnen/Screenshot)
|
||||
# fail-closed gesperrt → ein un-konfigurierter Executor ist KEINE offene Remote-Code-Execution mehr.
|
||||
# Der Token muss identisch auf der Box (Hermes-Env PC_EXECUTOR_TOKEN → mcp_pc.py) gesetzt sein.
|
||||
AUTH_TOKEN = os.environ.get("HERMES_PC_TOKEN", "").strip()
|
||||
|
||||
|
||||
def require_auth(authorization: str | None = Header(default=None)) -> None:
|
||||
"""Bearer-Token-Prüfung (konstante Zeit). 503 wenn der Executor ohne Token läuft (fail-closed),
|
||||
401 bei fehlendem/falschem Token."""
|
||||
if not AUTH_TOKEN:
|
||||
raise HTTPException(
|
||||
503,
|
||||
"Executor ohne HERMES_PC_TOKEN gestartet — Steuer-Endpunkte sind aus Sicherheitsgründen "
|
||||
"gesperrt. Setze die Env-Variable HERMES_PC_TOKEN (identisch zur Box) und starte neu.",
|
||||
)
|
||||
expected = f"Bearer {AUTH_TOKEN}"
|
||||
if not authorization or not hmac.compare_digest(authorization, expected):
|
||||
raise HTTPException(401, "Ungültiges oder fehlendes Bearer-Token.")
|
||||
|
||||
|
||||
# ── Shell ──────────────────────────────────────────────────────────────────
|
||||
|
||||
@app.post("/shell", dependencies=[Depends(require_auth)])
|
||||
async def run_shell(req: dict):
|
||||
cmd = req.get("command", "")
|
||||
try:
|
||||
result = subprocess.run(
|
||||
["powershell", "-NoProfile", "-NonInteractive", "-Command", cmd],
|
||||
capture_output=True, text=True, timeout=60,
|
||||
encoding="utf-8", errors="replace",
|
||||
)
|
||||
return {
|
||||
"stdout": result.stdout.strip(),
|
||||
"stderr": result.stderr.strip(),
|
||||
"returncode": result.returncode,
|
||||
}
|
||||
except subprocess.TimeoutExpired:
|
||||
raise HTTPException(408, "Timeout nach 60s")
|
||||
except Exception as e:
|
||||
raise HTTPException(500, str(e))
|
||||
|
||||
|
||||
# ── Screen ─────────────────────────────────────────────────────────────────
|
||||
|
||||
@app.post("/screenshot", dependencies=[Depends(require_auth)])
|
||||
async def take_screenshot():
|
||||
try:
|
||||
import mss
|
||||
from PIL import Image
|
||||
|
||||
with mss.mss() as sct:
|
||||
monitor = sct.monitors[1]
|
||||
raw = sct.grab(monitor)
|
||||
img = Image.frombytes("RGB", raw.size, raw.bgra, "raw", "BGRX")
|
||||
img.thumbnail((1280, 720))
|
||||
buf = io.BytesIO()
|
||||
img.save(buf, format="JPEG", quality=75)
|
||||
return {"screenshot": base64.b64encode(buf.getvalue()).decode()}
|
||||
except Exception as e:
|
||||
raise HTTPException(500, str(e))
|
||||
|
||||
|
||||
# ── Input ──────────────────────────────────────────────────────────────────
|
||||
|
||||
@app.post("/type", dependencies=[Depends(require_auth)])
|
||||
async def type_text(req: dict):
|
||||
text = req.get("text", "")
|
||||
try:
|
||||
import pyautogui
|
||||
pyautogui.write(text, interval=0.03)
|
||||
return {"ok": True}
|
||||
except Exception as e:
|
||||
raise HTTPException(500, str(e))
|
||||
|
||||
|
||||
@app.post("/key", dependencies=[Depends(require_auth)])
|
||||
async def press_keys(req: dict):
|
||||
keys = req.get("keys", "")
|
||||
try:
|
||||
import pyautogui
|
||||
parts = [k.strip() for k in keys.split("+")]
|
||||
pyautogui.hotkey(*parts)
|
||||
return {"ok": True}
|
||||
except Exception as e:
|
||||
raise HTTPException(500, str(e))
|
||||
|
||||
|
||||
# ── Medien & Lautstärke (A3: Lucy steuert den PC per Sprache) ──────────────
|
||||
# Media-Keys wirken auf den AKTIVEN Player (Spotify, YouTube, VLC …) — genau wie
|
||||
# die Tasten auf einer Multimedia-Tastatur. Lautstärke = System-Master (2 % je Schritt).
|
||||
MEDIA_KEYS = {"play_pause": "playpause", "next": "nexttrack", "prev": "prevtrack", "stop": "stop"}
|
||||
VOLUME_KEYS = {"up": "volumeup", "down": "volumedown", "mute": "volumemute"}
|
||||
|
||||
|
||||
@app.post("/media", dependencies=[Depends(require_auth)])
|
||||
async def media_control(req: dict):
|
||||
action = (req.get("action") or "").strip().lower()
|
||||
key = MEDIA_KEYS.get(action)
|
||||
if not key:
|
||||
raise HTTPException(400, f"Unbekannte Aktion '{action}'. Erlaubt: {', '.join(MEDIA_KEYS)}")
|
||||
try:
|
||||
import pyautogui
|
||||
pyautogui.press(key)
|
||||
return {"ok": True, "action": action}
|
||||
except Exception as e:
|
||||
raise HTTPException(500, str(e))
|
||||
|
||||
|
||||
@app.post("/volume", dependencies=[Depends(require_auth)])
|
||||
async def volume_control(req: dict):
|
||||
action = (req.get("action") or "").strip().lower()
|
||||
key = VOLUME_KEYS.get(action)
|
||||
if not key:
|
||||
raise HTTPException(400, f"Unbekannte Aktion '{action}'. Erlaubt: {', '.join(VOLUME_KEYS)}")
|
||||
steps = 1 if action == "mute" else max(1, min(25, int(req.get("steps", 5))))
|
||||
try:
|
||||
import pyautogui
|
||||
pyautogui.press(key, presses=steps, interval=0.02)
|
||||
return {"ok": True, "action": action, "steps": steps}
|
||||
except Exception as e:
|
||||
raise HTTPException(500, str(e))
|
||||
|
||||
|
||||
# ── Apps / Browser ─────────────────────────────────────────────────────────
|
||||
|
||||
@app.post("/open", dependencies=[Depends(require_auth)])
|
||||
async def open_target(req: dict):
|
||||
target = req.get("target", "")
|
||||
try:
|
||||
if target.startswith("http://") or target.startswith("https://"):
|
||||
webbrowser.open(target)
|
||||
else:
|
||||
os.startfile(target)
|
||||
return {"ok": True}
|
||||
except Exception as e:
|
||||
raise HTTPException(500, str(e))
|
||||
|
||||
|
||||
@app.post("/search", dependencies=[Depends(require_auth)])
|
||||
async def search_web(req: dict):
|
||||
query = req.get("query", "")
|
||||
url = f"https://www.google.com/search?q={quote(query)}"
|
||||
webbrowser.open(url)
|
||||
return {"ok": True, "url": url}
|
||||
|
||||
|
||||
# ── Health ─────────────────────────────────────────────────────────────────
|
||||
|
||||
@app.get("/health")
|
||||
async def health():
|
||||
return {"status": "ok", "host": socket.gethostname()}
|
||||
|
||||
|
||||
def _get_local_ip() -> str:
|
||||
try:
|
||||
with socket.socket(socket.AF_INET, socket.SOCK_DGRAM) as s:
|
||||
s.connect(("8.8.8.8", 80))
|
||||
return s.getsockname()[0]
|
||||
except Exception:
|
||||
return socket.gethostbyname(socket.gethostname())
|
||||
|
||||
|
||||
def _ensure_streams() -> None:
|
||||
"""Unter pythonw (kein Konsolenfenster) sind sys.stdout/stderr = None →
|
||||
uvicorns Logging crasht beim Start. Dann auf eine Logdatei umbiegen."""
|
||||
import sys
|
||||
if sys.stdout is not None and sys.stderr is not None:
|
||||
return
|
||||
logdir = os.path.join(
|
||||
os.environ.get("LOCALAPPDATA", os.path.dirname(os.path.abspath(__file__))),
|
||||
"HermesPCExecutor",
|
||||
)
|
||||
os.makedirs(logdir, exist_ok=True)
|
||||
f = open(os.path.join(logdir, "executor.log"), "a", buffering=1, encoding="utf-8")
|
||||
sys.stdout = f
|
||||
sys.stderr = f
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
_ensure_streams()
|
||||
my_ip = _get_local_ip()
|
||||
auth_state = "Token AKTIV" if AUTH_TOKEN else "KEIN Token → Steuer-Endpunkte GESPERRT (fail-closed)"
|
||||
print(f"Hermes PC Executor läuft auf http://{my_ip}:{MY_PORT} (bind {MY_HOST}) — {auth_state}")
|
||||
uvicorn.run(app, host=MY_HOST, port=MY_PORT)
|
||||
@@ -0,0 +1,8 @@
|
||||
@echo off
|
||||
cd /d "%~dp0"
|
||||
echo === Hermes PC Executor Setup ===
|
||||
python -m venv .venv
|
||||
.venv\Scripts\pip install -r requirements.txt -q
|
||||
echo.
|
||||
echo Fertig. Starte mit start.bat
|
||||
pause
|
||||
@@ -0,0 +1,6 @@
|
||||
fastapi>=0.111.0
|
||||
uvicorn>=0.30.0
|
||||
httpx>=0.27.0
|
||||
mss>=9.0.1
|
||||
Pillow>=10.0.0
|
||||
pyautogui>=0.9.54
|
||||
@@ -0,0 +1,9 @@
|
||||
@echo off
|
||||
cd /d "%~dp0"
|
||||
echo === Hermes PC Executor ===
|
||||
echo Port: 7777
|
||||
echo Hermes (AI Box) kann jetzt auf diesen PC zugreifen.
|
||||
echo Fenster offen lassen solange Hermes PC-Zugriff braucht.
|
||||
echo.
|
||||
.venv\Scripts\python executor.py
|
||||
pause
|
||||
@@ -0,0 +1,182 @@
|
||||
#!/usr/bin/env bash
|
||||
# Wöchentlicher Auto-Update-Lauf der Box (Autonomie E2). Läuft als User via
|
||||
# mc2-autoupdate.timer (So 04:30, nach dem 03:30-Backup) oder manuell.
|
||||
#
|
||||
# Prinzip „Haushaltsgerät": Fremd-Software (Router → Engine → Hermes) updatet sich
|
||||
# selbst über die BESTEHENDEN Pfade (update-swap.sh / update-engine.sh / hermes-update-Job).
|
||||
# Jede Ebene: Update-Check → Update → Postcheck. Rot ⇒ Rollback + SELBST-PINNING im
|
||||
# Pin-Register + Telegram-Meldung. Grün ⇒ Telegram „eingespielt X→Y".
|
||||
# Gepinnte Ebenen werden übersprungen, bis der Pin manuell gelöst wird.
|
||||
#
|
||||
# OS-Ebene: bewusst NICHT hier — unattended-upgrades (nur Security) wird in der
|
||||
# einmaligen sudo-Session eingerichtet (siehe deploy/sudoers-mc2-autonomie).
|
||||
# Modelle: NIE automatisch (User entscheidet nach Bench, Etappe E5).
|
||||
#
|
||||
# Router/Engine laufen als root: braucht die NOPASSWD-Regel aus
|
||||
# deploy/sudoers-mc2-autonomie. Fehlt sie, wird die Ebene übersprungen + gemeldet.
|
||||
set -uo pipefail # bewusst KEIN -e: jede Ebene wird kontrolliert abgearbeitet
|
||||
|
||||
API="${MC_API:-http://127.0.0.1:9001}"
|
||||
PINS="${MC_PINS_FILE:-/srv/models/mc2-pins.json}"
|
||||
SRC_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
HERMES_DIR="${HERMES_AGENT_DIR:-$HOME/.hermes/hermes-agent}"
|
||||
BACKUP_DIR="${MC_BACKUP_DIR:-/srv/models/mc2-backups}"
|
||||
JOB_TIMEOUT="${MC_AUTOUPDATE_JOB_TIMEOUT:-1200}" # Sekunden für den Hermes-Job
|
||||
export XDG_RUNTIME_DIR="${XDG_RUNTIME_DIR:-/run/user/$(id -u)}"
|
||||
|
||||
say(){ echo "[autoupdate] $*"; }
|
||||
notify(){ bash "$SRC_DIR/notify.sh" -s "[Box-Update]" "$1" || true; }
|
||||
|
||||
SUMMARY=() # eine Zeile je Ebene für die Abschluss-Meldung
|
||||
|
||||
# ── Pin-Register ─────────────────────────────────────────────────────────────
|
||||
pins_init(){ [ -f "$PINS" ] || echo '{}' > "$PINS"; }
|
||||
is_pinned(){ jq -e --arg k "$1" '.[$k].pinned == true' "$PINS" >/dev/null 2>&1; }
|
||||
pin_info(){ jq -r --arg k "$1" '.[$k] | "\(.version // "?") seit \(.datum // "?") (\(.grund // "?"))"' "$PINS" 2>/dev/null; }
|
||||
set_pin(){ # $1 komponente $2 version(known-good) $3 grund
|
||||
local tmp; tmp="$(mktemp)"
|
||||
jq --arg k "$1" --arg v "$2" --arg g "$3" --arg d "$(date +%F)" \
|
||||
'.[$k] = {pinned: true, version: $v, grund: $g, datum: $d}' "$PINS" > "$tmp" && mv "$tmp" "$PINS"
|
||||
}
|
||||
|
||||
# ── Helfer ───────────────────────────────────────────────────────────────────
|
||||
details(){ curl -sf --max-time 60 "$API/api/maintenance/update-details?kind=$1"; }
|
||||
|
||||
# „darf ich das PASSWORTLOS?" — sudo -n -l <cmd> reicht nicht (Exit 0 auch wenn nur
|
||||
# MIT Passwort erlaubt); deshalb explizit die NOPASSWD-Zeilen der Whitelist greppen.
|
||||
sudo_ok(){ sudo -n -l 2>/dev/null | grep -F "NOPASSWD" | grep -qF "$1"; }
|
||||
|
||||
# Router und Engine teilen denselben Ablauf: root-Skript mit eingebautem
|
||||
# Postcheck+Rollback (Exit 0 grün / 1 zurückgerollt / 2 kaputt).
|
||||
run_root_update(){ # $1 komponente(swap|engine) $2 skript $3 anzeigename $4 installed $5 latest
|
||||
local comp="$1" script="$2" name="$3" installed="$4" latest="$5" rc
|
||||
say "$name: Update $installed → $latest wird eingespielt…"
|
||||
sudo -n /usr/bin/bash "$SRC_DIR/$script"; rc=$?
|
||||
case "$rc" in
|
||||
0) SUMMARY+=("$name: $installed → $latest eingespielt, Stack-Check grün.")
|
||||
notify "$name aktualisiert: $installed → $latest. Stack-Check grün, alles läuft." ;;
|
||||
1) set_pin "$comp" "$installed" "Update auf $latest zerschoss den Stack-Check; Rollback grün"
|
||||
SUMMARY+=("$name: $latest FEHLGESCHLAGEN → zurückgerollt auf $installed + GEPINNT.")
|
||||
notify "$name-Update auf $latest fehlgeschlagen (Stack-Check rot). Automatisch zurückgerollt auf $installed — läuft wieder. Ebene ist jetzt GEPINNT, künftige Läufe überspringen sie." ;;
|
||||
*) SUMMARY+=("$name: KRITISCH — Update UND Rollback fehlgeschlagen!")
|
||||
notify "KRITISCH: $name-Update auf $latest UND Rollback fehlgeschlagen — Stack möglicherweise kaputt. Bitte melden, ich brauche Hilfe (Backup liegt bereit, restore.sh existiert)." ;;
|
||||
esac
|
||||
}
|
||||
|
||||
check_layer_root(){ # $1 komponente $2 skript $3 anzeigename $4 details-kind
|
||||
local comp="$1" script="$2" name="$3" kind="$4" det installed latest
|
||||
if is_pinned "$comp"; then
|
||||
say "$name: GEPINNT ($(pin_info "$comp")) — übersprungen."
|
||||
SUMMARY+=("$name: gepinnt, übersprungen.")
|
||||
return
|
||||
fi
|
||||
det="$(details "$kind")" || { say "$name: Update-Check nicht erreichbar."; SUMMARY+=("$name: Check fehlgeschlagen."); return; }
|
||||
installed="$(jq -r '.installed_build // empty' <<<"$det")"
|
||||
latest="$(jq -r '.latest_build // empty' <<<"$det")"
|
||||
if [ -z "$installed" ] || [ -z "$latest" ]; then
|
||||
say "$name: Versionen nicht ermittelbar (installed='$installed' latest='$latest')."
|
||||
SUMMARY+=("$name: Versions-Check unklar, nichts getan.")
|
||||
return
|
||||
fi
|
||||
if [ "$latest" -le "$installed" ] 2>/dev/null; then
|
||||
say "$name: aktuell ($installed)."
|
||||
SUMMARY+=("$name: aktuell ($installed).")
|
||||
return
|
||||
fi
|
||||
if ! sudo_ok "$script"; then
|
||||
say "$name: Update $installed → $latest verfügbar, aber sudo-Freischaltung fehlt."
|
||||
SUMMARY+=("$name: Update $installed → $latest wartet — sudo-Freischaltung fehlt (einmalig deploy/sudoers-mc2-autonomie installieren).")
|
||||
return
|
||||
fi
|
||||
run_root_update "$comp" "$script" "$name" "$installed" "$latest"
|
||||
}
|
||||
|
||||
# ── Hermes-Agent (User-Space, via MC2-Job + Polling; Rollback machen WIR) ────
|
||||
check_hermes(){
|
||||
local name="Hermes-Agent" det behind old_head resp job_id state t
|
||||
if is_pinned hermes; then
|
||||
say "$name: GEPINNT ($(pin_info hermes)) — übersprungen."
|
||||
SUMMARY+=("$name: gepinnt, übersprungen.")
|
||||
return
|
||||
fi
|
||||
det="$(details hermes)" || { SUMMARY+=("$name: Check fehlgeschlagen."); return; }
|
||||
behind="$(jq -r '.behind // 0' <<<"$det")"
|
||||
if [ "${behind:-0}" -eq 0 ] 2>/dev/null; then
|
||||
say "$name: aktuell."
|
||||
SUMMARY+=("$name: aktuell.")
|
||||
return
|
||||
fi
|
||||
old_head="$(git -C "$HERMES_DIR" rev-parse HEAD 2>/dev/null)"
|
||||
say "$name: $behind Commits hinterher — Update-Job startet (Backup→update→doctor→Restart→Gehirn-Check)…"
|
||||
resp="$(curl -sf --max-time 30 -X POST "$API/api/maintenance/hermes-update")" || { SUMMARY+=("$name: Job-Start fehlgeschlagen."); return; }
|
||||
job_id="$(jq -r '.job_id // empty' <<<"$resp")"
|
||||
if [ -z "$job_id" ]; then
|
||||
say "$name: kein Job gestartet: $resp"
|
||||
SUMMARY+=("$name: Job-Start abgelehnt ($(jq -r '.error // .detail // "unbekannt"' <<<"$resp" 2>/dev/null)).")
|
||||
return
|
||||
fi
|
||||
t=0
|
||||
while [ "$t" -lt "$JOB_TIMEOUT" ]; do
|
||||
state="$(curl -sf --max-time 15 "$API/api/jobs" | jq -r --arg id "$job_id" '.jobs[] | select(.id==$id) | .state' 2>/dev/null)"
|
||||
case "$state" in done|failed|canceled) break ;; esac
|
||||
sleep 10; t=$((t + 10))
|
||||
done
|
||||
if [ "$state" = "done" ]; then
|
||||
SUMMARY+=("$name: $behind Commits eingespielt, Gehirn-Check grün.")
|
||||
notify "Hermes-Agent aktualisiert ($behind Commits). Gehirn-Check (Mem0, Tools, Voice) grün — alles läuft."
|
||||
return
|
||||
fi
|
||||
# Rot/Timeout: ERST Selbstreparatur versuchen (Config-Bruch selbst ziehen, neue Version behalten) —
|
||||
# Autonomie 0g. Nur Config, reversibel, hart gegatet; scheitert es, fällt es sauber in den Rollback.
|
||||
say "$name: Job endete '$state' — versuche Selbstreparatur (Config) vor dem Rollback…"
|
||||
local sr_out sr_rc sugg=""
|
||||
sr_out="$(bash "$SRC_DIR/self-repair.sh" 2>&1)"; sr_rc=$?
|
||||
echo "$sr_out"
|
||||
if [ "$sr_rc" -eq 0 ]; then
|
||||
SUMMARY+=("$name: Config-Bruch nach Update SELBST repariert — neue Version läuft (kein Rollback).")
|
||||
return
|
||||
fi
|
||||
sugg="$(echo "$sr_out" | sed -n 's/^SELFREPAIR_SUGGESTION=//p' | tail -1)"
|
||||
|
||||
# Rollback: Code auf alten Stand, Config aus dem frischen Backup, Neustart, Gehirn-Check.
|
||||
say "$name: Selbstreparatur griff nicht — ROLLBACK auf $old_head…"
|
||||
if [ -n "$old_head" ]; then
|
||||
git -C "$HERMES_DIR" reset --hard "$old_head" >/dev/null 2>&1
|
||||
local bak stage
|
||||
bak="$(ls -1t "$BACKUP_DIR"/mc2-state-*.tar.gz 2>/dev/null | head -1)"
|
||||
if [ -n "$bak" ]; then
|
||||
stage="$(mktemp -d)"
|
||||
tar -xzf "$bak" -C "$stage" ./hermes 2>/dev/null || tar -xzf "$bak" -C "$stage" hermes 2>/dev/null
|
||||
[ -f "$stage/hermes/config.yaml" ] && cp -a "$stage/hermes/config.yaml" "$HOME/.hermes/config.yaml"
|
||||
[ -f "$stage/hermes/.env" ] && cp -a "$stage/hermes/.env" "$HOME/.hermes/.env"
|
||||
rm -rf "$stage"
|
||||
fi
|
||||
systemctl --user restart hermes-gateway; sleep 6
|
||||
if bash "$SRC_DIR/hermes-postcheck.sh" >/dev/null 2>&1; then
|
||||
set_pin hermes "$old_head" "Update ($behind Commits) riss den Gehirn-Check; Rollback grün"
|
||||
SUMMARY+=("$name: Update FEHLGESCHLAGEN → zurückgerollt + GEPINNT auf ${old_head:0:9}.")
|
||||
notify "Hermes-Update fehlgeschlagen (Gehirn-Check rot). Selbstreparatur der Config griff nicht, deshalb automatisch zurückgerollt auf ${old_head:0:9} — läuft wieder. Hermes ist jetzt GEPINNT.${sugg:+
|
||||
|
||||
Wahrscheinliche Ursache/Fix (Box-Diagnose): $sugg}"
|
||||
return
|
||||
fi
|
||||
fi
|
||||
SUMMARY+=("$name: KRITISCH — Update UND Rollback fehlgeschlagen!")
|
||||
notify "KRITISCH: Hermes-Update UND Rollback fehlgeschlagen — Gehirn-Check bleibt rot. Bitte melden. (Voll-Restore: deploy/restore.sh mit dem letzten Backup.)"
|
||||
}
|
||||
|
||||
# ── Lauf ─────────────────────────────────────────────────────────────────────
|
||||
say "Auto-Update-Lauf startet ($(date '+%F %H:%M'))."
|
||||
pins_init
|
||||
if ! curl -sf --max-time 20 "$API/api/health" >/dev/null; then
|
||||
notify "Auto-Update abgebrochen: MC2-API ($API) nicht erreichbar — bitte Box prüfen."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
check_layer_root swap update-swap.sh "Router (llama-swap)" swap
|
||||
check_layer_root engine update-engine.sh "Engine (llama.cpp)" engine
|
||||
check_hermes
|
||||
|
||||
notify "Commander, die Wochenpflege der Box ist durch — kurz für dich:
|
||||
$(printf '• %s\n' "${SUMMARY[@]}")"
|
||||
say "Fertig."
|
||||
@@ -0,0 +1,72 @@
|
||||
#!/usr/bin/env bash
|
||||
# Voll-Zustands-Backup der AI-Box: alles, was NICHT aus Git oder per Re-Download
|
||||
# zurückkommt. Erzeugt EIN Tarball mc2-state-<ts>.tar.gz, behält die letzten N.
|
||||
# Läuft per systemd-Timer (täglich), manuell, oder über den UI-Snapshot-Button.
|
||||
#
|
||||
# Inhalt: mem0 (Chroma + history.db) · ~/.hermes (config.yaml, .env, plugins/) ·
|
||||
# /etc/llama-swap/config.yaml
|
||||
# NICHT enthalten (bewusst): GGUF-Modelle (riesig, neu ladbar), MC2-Code (Git), venvs.
|
||||
#
|
||||
# ACHTUNG: das Tarball enthält ~/.hermes/.env (Secrets) → chmod 600, nicht in Git.
|
||||
#
|
||||
# Off-Box (C12): nach dem lokalen Tarball wird der Backup-Ordner per rsync auf ein
|
||||
# ZWEITES Gerät gespiegelt (Default: Proxmox-Host), damit ein Plattenausfall der Box
|
||||
# nicht auch die Backups mitnimmt. Fehlschlag ist NICHT fatal — das lokale Backup gilt.
|
||||
set -euo pipefail
|
||||
|
||||
MODELS_DIR="${MC_MODELS_DIR:-/srv/models}"
|
||||
DEST_DIR="$MODELS_DIR/mc2-backups"
|
||||
RETAIN="${MC_BACKUP_RETAIN:-14}"
|
||||
# Off-Box-Ziel (leer = deaktiviert). Default: Proxmox-Host, eigener Key mit minimalem Recht.
|
||||
OFFSITE="${MC_BACKUP_OFFSITE:-root@192.168.178.108:/var/lib/vz/mc2-backups}"
|
||||
OFFSITE_KEY="${MC_BACKUP_OFFSITE_KEY:-$HOME/.ssh/mc2_offsite}"
|
||||
MEM0_DIR="${MC_MEM0_DIR:-/srv/models/mem0}"
|
||||
LSWAP="${MC_CONFIG_PATH:-/etc/llama-swap/config.yaml}"
|
||||
HERMES="${HERMES_HOME:-$HOME/.hermes}"
|
||||
TS="$(date +%Y%m%d-%H%M%S)"
|
||||
|
||||
STAGE="$(mktemp -d)"
|
||||
trap 'rm -rf "$STAGE"' EXIT
|
||||
mkdir -p "$STAGE/mem0" "$STAGE/hermes" "$STAGE/llama-swap"
|
||||
|
||||
[ -d "$MEM0_DIR" ] && cp -a "$MEM0_DIR/." "$STAGE/mem0/" || true
|
||||
[ -f "$HERMES/config.yaml" ] && cp -a "$HERMES/config.yaml" "$STAGE/hermes/" || true
|
||||
[ -f "$HERMES/.env" ] && cp -a "$HERMES/.env" "$STAGE/hermes/" || true
|
||||
[ -d "$HERMES/plugins" ] && cp -a "$HERMES/plugins" "$STAGE/hermes/plugins" || true
|
||||
[ -f "$LSWAP" ] && cp -a "$LSWAP" "$STAGE/llama-swap/" || true
|
||||
|
||||
cat > "$STAGE/MANIFEST.txt" <<EOF
|
||||
mc2-state backup
|
||||
created : $TS
|
||||
host : $(hostname)
|
||||
inhalt : mem0 (chroma + history.db), hermes (config.yaml, .env, plugins/), llama-swap (config.yaml)
|
||||
restore : bash ~/mission-control-v2/deploy/restore.sh mc2-state-$TS.tar.gz
|
||||
EOF
|
||||
|
||||
mkdir -p "$DEST_DIR"
|
||||
OUT="$DEST_DIR/mc2-state-$TS.tar.gz"
|
||||
tar -czf "$OUT" -C "$STAGE" .
|
||||
chmod 600 "$OUT" # enthält .env (Secrets)
|
||||
|
||||
# Retention: nur die letzten N behalten.
|
||||
ls -1t "$DEST_DIR"/mc2-state-*.tar.gz 2>/dev/null | tail -n +$((RETAIN + 1)) | xargs -r rm -f
|
||||
|
||||
echo "OK — Backup: $OUT ($(du -h "$OUT" | cut -f1))"
|
||||
|
||||
# --- Off-Box-Spiegel (C12) -------------------------------------------------
|
||||
# Spiegelt die Tarballs 1:1 aufs Zweitgerät; --delete zieht die Retention mit,
|
||||
# sodass die Off-Box-Kopie nicht unbegrenzt wächst. Perms (600) bleiben erhalten.
|
||||
if [ -n "$OFFSITE" ]; then
|
||||
if [ -f "$OFFSITE_KEY" ]; then
|
||||
if rsync -a \
|
||||
-e "ssh -i $OFFSITE_KEY -o BatchMode=yes -o ConnectTimeout=10 -o StrictHostKeyChecking=accept-new" \
|
||||
--include='mc2-state-*.tar.gz' --exclude='*' --delete \
|
||||
"$DEST_DIR/" "$OFFSITE/"; then
|
||||
echo "OK — Off-Box-Kopie gespiegelt → $OFFSITE"
|
||||
else
|
||||
echo "WARN — Off-Box-Sync fehlgeschlagen (lokales Backup ist gültig): $OFFSITE" >&2
|
||||
fi
|
||||
else
|
||||
echo "WARN — Off-Box-Key fehlt ($OFFSITE_KEY) → nur lokales Backup." >&2
|
||||
fi
|
||||
fi
|
||||
@@ -0,0 +1,40 @@
|
||||
#!/usr/bin/env bash
|
||||
# Brain-Config-Matrix (Referenz vom 02.07.2026): misst Qwen3.6 in mehreren Betriebsarten
|
||||
# (MTP n-max / ohne Spec / KV-Quant) auf separatem Port. Für künftige Re-Benchmarks nach
|
||||
# Engine-Updates oder bei neuen Hirn-Kandidaten. Ergebnis-Referenz (02.07., Vulkan, b9843):
|
||||
# mtp3 78,6 t/s · mtp4 63,7 · ohne Spec 62,4 · mtp3+KVq8 78,6 (=gratis) · nospec+KVq8 62,1
|
||||
set -u
|
||||
MODEL="${BRAIN_GGUF:-/srv/models/Qwen3.6-35B-A3B-MTP-GGUF/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf}"
|
||||
BIN="${BENCH_BIN:-/opt/llamacpp-vulkan/llama-server}"
|
||||
PORT="${BENCH_PORT:-5899}"
|
||||
BASE="-c 65536 -ngl 999 -fa on --no-mmap --jinja --parallel 1"
|
||||
|
||||
run() { # $1=name $2=extra-flags
|
||||
echo "=== $1 ==="
|
||||
$BIN -m "$MODEL" --host 127.0.0.1 --port "$PORT" $BASE $2 >/tmp/brain-bench-server.log 2>&1 &
|
||||
local PID=$!
|
||||
for i in $(seq 1 120); do
|
||||
curl -s --max-time 2 "http://127.0.0.1:$PORT/health" | grep -q ok && break
|
||||
sleep 2
|
||||
kill -0 $PID 2>/dev/null || { echo " LOAD-CRASH"; tail -2 /tmp/brain-bench-server.log; return; }
|
||||
done
|
||||
for r in 0 1 2; do
|
||||
curl -s --max-time 300 -X POST "http://127.0.0.1:$PORT/completion" -H "Content-Type: application/json" \
|
||||
-d '{"prompt":"Erklaere in drei kurzen Saetzen, was ein Mixture-of-Experts-Modell ist.","n_predict":150,"temperature":0}' \
|
||||
> /tmp/brain-bench-probe.json
|
||||
[ "$r" = "0" ] && continue # Warmlauf verwerfen
|
||||
python3 - <<'PY'
|
||||
import json
|
||||
t = json.load(open("/tmp/brain-bench-probe.json")).get("timings", {})
|
||||
print(f" tg={t.get('predicted_per_second',0):.1f} t/s"
|
||||
+ (f" · draft {t.get('draft_n_accepted')}/{t.get('draft_n')}" if t.get('draft_n') else ""))
|
||||
PY
|
||||
done
|
||||
kill $PID 2>/dev/null; wait $PID 2>/dev/null; sleep 3
|
||||
}
|
||||
|
||||
run "mtp3 (live-Config)" "--spec-type draft-mtp --spec-draft-n-max 3 -ctk q8_0 -ctv q8_0"
|
||||
run "mtp4" "--spec-type draft-mtp --spec-draft-n-max 4 -ctk q8_0 -ctv q8_0"
|
||||
run "ohne Spec" "-ctk q8_0 -ctv q8_0"
|
||||
run "mtp3 KV-f16" "--spec-type draft-mtp --spec-draft-n-max 3"
|
||||
echo "BENCH_DONE"
|
||||
@@ -0,0 +1,42 @@
|
||||
#!/usr/bin/env bash
|
||||
# Modell-Bench (Autonomie-Plan E5): misst pp/tg eines GGUF auf separatem Port, ohne die
|
||||
# Live-Instanzen anzufassen. Grundlage der Modell-Selbst-Evaluation (Radar-Kandidaten).
|
||||
# Nutzung: bash model-bench.sh <gguf-pfad> [extra llama-server-flags, z.B. --mmproj ...]
|
||||
# Lehren vom 02.07.2026 eingebaut: python via Heredoc (kein f-String-Quoting-Bruch),
|
||||
# Warmlauf vor der Messung, LOAD-CRASH wird gemeldet statt still zu hängen.
|
||||
set -u
|
||||
MODEL="${1:?Nutzung: model-bench.sh <gguf-pfad> [extra-flags]}"
|
||||
shift || true
|
||||
EXTRA="$*"
|
||||
BIN="${BENCH_BIN:-/opt/llamacpp-vulkan/llama-server}"
|
||||
PORT="${BENCH_PORT:-5899}"
|
||||
CTX="${BENCH_CTX:-32768}"
|
||||
|
||||
echo "=== Bench: $(basename "$MODEL") (ctx $CTX${EXTRA:+, $EXTRA}) ==="
|
||||
$BIN -m "$MODEL" --host 127.0.0.1 --port "$PORT" -c "$CTX" -ngl 999 -fa on --no-mmap --jinja $EXTRA \
|
||||
>/tmp/model-bench-server.log 2>&1 &
|
||||
PID=$!
|
||||
trap 'kill $PID 2>/dev/null; wait $PID 2>/dev/null' EXIT
|
||||
|
||||
for i in $(seq 1 240); do
|
||||
curl -s --max-time 2 "http://127.0.0.1:$PORT/health" | grep -q ok && break
|
||||
sleep 2
|
||||
kill -0 $PID 2>/dev/null || { echo "LOAD-CRASH:"; tail -3 /tmp/model-bench-server.log; exit 1; }
|
||||
done
|
||||
|
||||
probe() {
|
||||
curl -s --max-time 300 -X POST "http://127.0.0.1:$PORT/completion" -H "Content-Type: application/json" \
|
||||
-d '{"prompt":"Erklaere in fuenf Saetzen, wie ein Backup-Konzept fuer einen Homeserver aussieht.","n_predict":200,"temperature":0}' \
|
||||
> /tmp/model-bench-probe.json
|
||||
python3 - <<'PY'
|
||||
import json
|
||||
t = json.load(open("/tmp/model-bench-probe.json")).get("timings", {})
|
||||
print(f" pp={t.get('prompt_per_second',0):.0f} t/s · tg={t.get('predicted_per_second',0):.1f} t/s"
|
||||
+ (f" · draft {t.get('draft_n_accepted')}/{t.get('draft_n')}" if t.get('draft_n') else ""))
|
||||
PY
|
||||
}
|
||||
|
||||
probe >/dev/null 2>&1 # Warmlauf (Shader-/Graph-Compile nicht mitmessen)
|
||||
probe
|
||||
probe
|
||||
echo "BENCH_DONE"
|
||||
@@ -0,0 +1,8 @@
|
||||
#!/usr/bin/env bash
|
||||
# Frontend bauen (vor jedem Deploy auf dem Entwickler-PC). Output → frontend/dist,
|
||||
# das committet wird (kein Node-Build auf der Box).
|
||||
set -euo pipefail
|
||||
cd "$(dirname "$0")/../frontend"
|
||||
npm ci
|
||||
npm run build
|
||||
echo "OK — frontend/dist gebaut."
|
||||
@@ -0,0 +1,96 @@
|
||||
#!/usr/bin/env bash
|
||||
# Deploy AUF DER BOX als systemd-USER-Dienst — KEIN sudo, KEIN /opt, KEIN Passwort.
|
||||
# Erstinstallation + Updates in einem. Läuft als User hitonabi.
|
||||
#
|
||||
# Erstinstallation (einmalig):
|
||||
# git clone https://git.tobisniceshomelab.ddnsfree.com/Hitonabi/mission-control-v2 ~/mission-control-v2
|
||||
# bash ~/mission-control-v2/deploy/deploy.sh
|
||||
set -euo pipefail
|
||||
|
||||
SRC="${MC2_SRC:-$HOME/mission-control-v2}"
|
||||
|
||||
cd "$SRC"
|
||||
git fetch -q origin && git reset -q --hard origin/main
|
||||
|
||||
# venv + Abhängigkeiten
|
||||
if [ ! -d "$SRC/backend/.venv" ]; then
|
||||
python3 -m venv "$SRC/backend/.venv"
|
||||
fi
|
||||
"$SRC/backend/.venv/bin/python" -m pip install -q --upgrade pip
|
||||
"$SRC/backend/.venv/bin/python" -m pip install -q -r "$SRC/backend/requirements.txt"
|
||||
|
||||
# Mem0-Sidecar (auto-lernendes Gedächtnis) — eigenes Python-3.12-venv (~/.mem0/venv),
|
||||
# weil mem0+chromadb unter dem 3.14-Backend-venv nicht laufen. mem0ai/chromadb sind dort
|
||||
# bereits installiert; hier nur den HTTP-Server nachziehen + Alt-DB einmalig migrieren.
|
||||
if [ -x "$HOME/.mem0/venv/bin/python" ]; then
|
||||
# ~/.mem0/venv ist uv-managed (kein pip) → uv pip nutzen.
|
||||
UV="$(command -v uv || echo "$HOME/.local/bin/uv")"
|
||||
"$UV" pip install -q --python "$HOME/.mem0/venv/bin/python" -r "$SRC/mem0_service/requirements.txt"
|
||||
( cd "$SRC/mem0_service" && "$HOME/.mem0/venv/bin/python" migrate.py ) || true
|
||||
else
|
||||
echo "WARN: ~/.mem0/venv fehlt — Mem0-Sidecar wird nicht gestartet (siehe Plan A1)."
|
||||
fi
|
||||
|
||||
# Voice-Sidecar (lokales STT + gestuftes TTS für „Mit Hermes reden") — eigenes Python-3.12-venv
|
||||
# (~/.voice/venv), weil torch/chatterbox/faster-whisper nicht ins 3.14-Backend-venv passen.
|
||||
# install.sh ist idempotent (venv + Deps + Piper-Stimmen). Best-effort: schlägt es fehl, läuft
|
||||
# der restliche Stack weiter (der Voice-Tab meldet den Sidecar dann als offline).
|
||||
bash "$SRC/voice_service/install.sh" || echo "WARN: Voice-Sidecar-Setup fehlgeschlagen — Voice-Tab bleibt offline."
|
||||
|
||||
# Hermes-Memory-Provider-Plugin (auto-lernen/Recall via Mem0-Sidecar) nach ~/.hermes/plugins/
|
||||
# spiegeln. Context-only (kein Tool-Loop). Aktivierung in ~/.hermes/config.yaml:
|
||||
# memory.memory_enabled: true + memory.provider: mc2-memory (einmalig, box-lokal).
|
||||
if [ -d "$HOME/.hermes" ]; then
|
||||
mkdir -p "$HOME/.hermes/plugins/mc2-memory"
|
||||
cp "$SRC/hermes/plugins/mc2-memory/__init__.py" "$SRC/hermes/plugins/mc2-memory/plugin.yaml" \
|
||||
"$HOME/.hermes/plugins/mc2-memory/"
|
||||
fi
|
||||
|
||||
# systemd-USER-Units installieren/aktualisieren
|
||||
mkdir -p "$HOME/.config/systemd/user"
|
||||
cp "$SRC/deploy/mission-control-2.service" "$HOME/.config/systemd/user/mission-control-2.service"
|
||||
# Hermes-Terminal (ttyd -> `hermes chat`), in MC2 als Terminal-Seite eingebettet.
|
||||
# Einmalig manuell noetig: `sudo apt install -y ttyd` + `sudo systemctl disable --now ttyd` (apt-Default-Dienst).
|
||||
cp "$SRC/deploy/hermes-terminal.service" "$HOME/.config/systemd/user/hermes-terminal.service"
|
||||
# Mem0-Sidecar-Unit (nur wenn das venv existiert).
|
||||
[ -x "$HOME/.mem0/venv/bin/python" ] && cp "$SRC/deploy/mem0-service.service" "$HOME/.config/systemd/user/mem0-service.service"
|
||||
# Voice-Sidecar-Unit (nur wenn das venv existiert).
|
||||
[ -x "$HOME/.voice/venv/bin/python" ] && cp "$SRC/deploy/voice-service.service" "$HOME/.config/systemd/user/voice-service.service"
|
||||
# Tägliches Zustands-Backup (mem0 + Configs/Secrets) — Timer + oneshot-Service. Siehe docs/BACKUP.md.
|
||||
cp "$SRC/deploy/mc2-backup.service" "$HOME/.config/systemd/user/mc2-backup.service"
|
||||
cp "$SRC/deploy/mc2-backup.timer" "$HOME/.config/systemd/user/mc2-backup.timer"
|
||||
# Wöchentliches Auto-Update (Router/Engine/Hermes, Rollback+Pin+Telegram) — Autonomie E2.
|
||||
cp "$SRC/deploy/mc2-autoupdate.service" "$HOME/.config/systemd/user/mc2-autoupdate.service"
|
||||
cp "$SRC/deploy/mc2-autoupdate.timer" "$HOME/.config/systemd/user/mc2-autoupdate.timer"
|
||||
# Tägliche Selbst-Smoke-Tests (Gateway/Tools/Voice aktiv, Alarm bei Rot) — Autonomie 0d.
|
||||
cp "$SRC/deploy/mc2-selfsmoke.service" "$HOME/.config/systemd/user/mc2-selfsmoke.service"
|
||||
cp "$SRC/deploy/mc2-selfsmoke.timer" "$HOME/.config/systemd/user/mc2-selfsmoke.timer"
|
||||
# Evolution-Radar-Feed (Autonomie E4): Hermes-cron speist ihn als Prompt ein.
|
||||
# Cron-Job selbst wird EINMALIG registriert (hermes cron create, siehe docs/RUNBOOK.md).
|
||||
mkdir -p "$HOME/.hermes/scripts"
|
||||
cp "$SRC/deploy/radar-feed.sh" "$HOME/.hermes/scripts/radar-feed.sh"
|
||||
# Selbstkritik-Cron (Faden 11b): Zwilling des Radars, Blick nach INNEN (Latenzen/Journal/Skills).
|
||||
cp "$SRC/deploy/selbstkritik-feed.sh" "$HOME/.hermes/scripts/selbstkritik-feed.sh"
|
||||
# Werkstatt-Skill (Autonomie E6): Selbstwartungs-Kreislauf für Hermes.
|
||||
mkdir -p "$HOME/.hermes/skills/wartung"
|
||||
cp "$SRC/deploy/skills/wartung/SKILL.md" "$HOME/.hermes/skills/wartung/SKILL.md"
|
||||
systemctl --user daemon-reload
|
||||
systemctl --user enable mission-control-2 >/dev/null 2>&1 || true
|
||||
systemctl --user enable hermes-terminal >/dev/null 2>&1 || true
|
||||
systemctl --user enable mem0-service >/dev/null 2>&1 || true
|
||||
systemctl --user enable voice-service >/dev/null 2>&1 || true
|
||||
systemctl --user enable --now mc2-backup.timer >/dev/null 2>&1 || true
|
||||
systemctl --user enable --now mc2-autoupdate.timer >/dev/null 2>&1 || true
|
||||
systemctl --user enable --now mc2-selfsmoke.timer >/dev/null 2>&1 || true
|
||||
loginctl enable-linger "$USER" >/dev/null 2>&1 || true
|
||||
# Mem0-Sidecar VOR dem Backend (re)starten, damit /api/memory sofort bedient wird.
|
||||
[ -x "$HOME/.mem0/venv/bin/python" ] && systemctl --user restart mem0-service 2>/dev/null || true
|
||||
# Voice-Sidecar (re)starten (best-effort; Erststart lädt das STT-Modell vor).
|
||||
[ -x "$HOME/.voice/venv/bin/python" ] && systemctl --user restart voice-service 2>/dev/null || true
|
||||
systemctl --user restart mission-control-2
|
||||
command -v ttyd >/dev/null 2>&1 && systemctl --user restart hermes-terminal 2>/dev/null || true
|
||||
|
||||
sleep 2
|
||||
echo "--- Health ---"
|
||||
curl -sf http://127.0.0.1:9001/api/health && echo
|
||||
echo "OK — Mission Control 2.0 läuft auf :9001 (User-Dienst, sudo-frei)."
|
||||
@@ -0,0 +1,94 @@
|
||||
#!/usr/bin/env bash
|
||||
# Post-Update-Check nach einem Hermes-Agent-Update: läuft unser geteiltes „Gehirn"
|
||||
# (Mem0 + die mc2-memory-Integration) noch? Exit 0 = alles ok, sonst 1 → der Job wird
|
||||
# im UI rot, damit ein kaputtes Gehirn sofort auffällt.
|
||||
set -uo pipefail
|
||||
|
||||
MC_URL="${MC_URL:-http://127.0.0.1:9001}"
|
||||
MEM0_URL="${MEM0_SERVICE_URL:-http://127.0.0.1:8765}"
|
||||
HERMES="${HERMES_HOME:-$HOME/.hermes}"
|
||||
fail=0
|
||||
|
||||
echo "=== Hermes Post-Update: Gehirn-Check ==="
|
||||
|
||||
if curl -sf -m 5 "$MEM0_URL/health" >/dev/null 2>&1; then
|
||||
echo "PASS · Mem0-Sidecar erreichbar"
|
||||
else
|
||||
echo "FAIL · Mem0-Sidecar NICHT erreichbar"; fail=1
|
||||
fi
|
||||
|
||||
if curl -sf -m 8 "$MC_URL/api/memory" >/dev/null 2>&1; then
|
||||
echo "PASS · /api/memory antwortet"
|
||||
else
|
||||
echo "FAIL · /api/memory antwortet nicht"; fail=1
|
||||
fi
|
||||
|
||||
if grep -qE "^[[:space:]]*provider:[[:space:]]*'?mc2-memory'?" "$HERMES/config.yaml" 2>/dev/null \
|
||||
&& grep -qE "memory_enabled:[[:space:]]*true" "$HERMES/config.yaml" 2>/dev/null; then
|
||||
echo "PASS · memory.provider=mc2-memory aktiv"
|
||||
else
|
||||
echo "FAIL · memory.provider nicht mehr gesetzt (Config vom Update überschrieben?)"; fail=1
|
||||
fi
|
||||
|
||||
if ( cd "$HERMES/hermes-agent" && HERMES_HOME="$HERMES" ./venv/bin/python -c \
|
||||
"import sys; sys.path.insert(0,'.'); from plugins.memory import load_memory_provider; p=load_memory_provider('mc2-memory'); assert p and p.name()=='mc2-memory'" \
|
||||
>/dev/null 2>&1 ); then
|
||||
echo "PASS · mc2-memory-Plugin lädt unter dem neuen Hermes"
|
||||
else
|
||||
echo "FAIL · mc2-memory-Plugin lädt nicht (MemoryProvider-ABC geändert?)"; fail=1
|
||||
fi
|
||||
|
||||
# --- Config-Drift-Wächter (Lehre aus v0.18, 02.07.2026): Hermes fällt bei unbekannten ---
|
||||
# --- Config-Werten STILL auf Defaults zurück (approvals.mode 'auto' -> manual -> alle ---
|
||||
# --- Tools hingen in pending_approval). Solche Warnungen müssen den Job ROT machen. ---
|
||||
if journalctl --user -u hermes-gateway --since "-3 minutes" --no-pager -o cat 2>/dev/null \
|
||||
| grep -iE "Unknown [a-z._]+ '|deprecated .*(setting|option|config)|defaulting to|invalid config" \
|
||||
| grep -v "check_fn" | head -5 | tee /tmp/hermes-config-warnings.txt | grep -q .; then
|
||||
echo "FAIL · Config-Warnungen nach Update (siehe oben) — Config an neue Version anpassen!"; fail=1
|
||||
else
|
||||
echo "PASS · keine Config-Drift-Warnungen im Gateway-Log"
|
||||
fi
|
||||
|
||||
# --- Tool-Smoke: ein harmloser terminal-Befehl durch den echten Agenten. Fängt kaputte ---
|
||||
# --- Approval-/Tool-Ketten (Antwort muss kommen und darf nicht in pending_approval hängen). ---
|
||||
API_KEY="$(grep -E '^API_SERVER_KEY=' "$HERMES/.env" 2>/dev/null | cut -d= -f2- | tr -d '"' | tr -d "'")"
|
||||
if [ -n "$API_KEY" ]; then
|
||||
TOOL_RESP=$(curl -sf -m 90 -X POST "http://127.0.0.1:8642/v1/chat/completions" \
|
||||
-H "Authorization: Bearer $API_KEY" -H "Content-Type: application/json" \
|
||||
-H "X-Hermes-Session-Id: postcheck-tool-smoke" \
|
||||
-d '{"model":"hermes","stream":false,"messages":[{"role":"user","content":"Führe im Terminal exakt den Befehl echo postcheck-ok aus und gib mir nur dessen Ausgabe zurück."}]}' 2>/dev/null)
|
||||
if echo "$TOOL_RESP" | grep -qi "postcheck-ok"; then
|
||||
echo "PASS · Tool-Smoke (terminal via Agent) liefert Ergebnis"
|
||||
elif echo "$TOOL_RESP" | grep -qi "pending_approval"; then
|
||||
echo "FAIL · Tool-Smoke hängt in pending_approval (approvals.mode prüfen!)"; fail=1
|
||||
else
|
||||
echo "FAIL · Tool-Smoke ohne verwertbares Ergebnis: $(echo "$TOOL_RESP" | head -c 200)"; fail=1
|
||||
fi
|
||||
else
|
||||
echo "SKIP · Tool-Smoke (kein API_SERVER_KEY in $HERMES/.env gefunden)"
|
||||
fi
|
||||
|
||||
# --- Voice-Smoke: Lucys Sprech-Pfad streamt. Mit Retry: direkt nach dem Gateway-Neustart ---
|
||||
# --- zahlt der erste Turn den Session-Kaltstart (Prefill) — ein Versuch wäre ein Fehlalarm. ---
|
||||
# WICHTIG: Subshell OHNE pipefail — `head -c 50` schließt die Pipe früh, curl endet mit 23
|
||||
# (SIGPIPE) und pipefail würde den bestandenen Check als FAIL werten (live diagnostiziert).
|
||||
voice_ok=0
|
||||
for try in 1 2 3; do
|
||||
if ( set +o pipefail; curl -sf -m 45 -N -X POST "$MC_URL/api/voice/chat" -H "Content-Type: application/json" \
|
||||
-d '{"text":"Sag nur ok.","session_id":"postcheck-voice-smoke"}' 2>/dev/null | head -c 50 | grep -q "data:" ); then
|
||||
voice_ok=1; break
|
||||
fi
|
||||
sleep 8
|
||||
done
|
||||
if [ "$voice_ok" -eq 1 ]; then
|
||||
echo "PASS · Voice-Smoke (Lucys /api/voice/chat streamt, Versuch $try)"
|
||||
else
|
||||
echo "FAIL · Voice-Smoke: /api/voice/chat streamt nicht (3 Versuche)"; fail=1
|
||||
fi
|
||||
|
||||
if [ "$fail" -eq 0 ]; then
|
||||
echo "=== GEHIRN OK ✓ ==="
|
||||
else
|
||||
echo "=== GEHIRN-CHECK FEHLGESCHLAGEN — bitte prüfen ==="
|
||||
fi
|
||||
exit "$fail"
|
||||
@@ -0,0 +1,21 @@
|
||||
[Unit]
|
||||
Description=Hermes Terminal — ttyd web terminal wrapping `hermes chat` (interaktiver Agent)
|
||||
Documentation=https://github.com/tsl0922/ttyd
|
||||
After=network.target
|
||||
|
||||
[Service]
|
||||
# Exponiert die interaktive Hermes-Agent-CLI (`hermes chat`, voller Agent mit Tools/PC) als
|
||||
# Web-Terminal. Wird in MC2 per iframe eingebettet (Terminal-Seite) — Ersatz fuer AnythingLLM.
|
||||
#
|
||||
# SICHERHEIT: --writable + LAN-Bind ohne Auth = dasselbe Trust-Modell wie das MC2-Dashboard
|
||||
# (vertrautes Heim-LAN, kein Internet-Exposure). Fuer Basic-Auth: am ExecStart
|
||||
# --credential <user>:<pass>
|
||||
# ergaenzen (Browser fragt dann einmalig im iframe nach).
|
||||
Type=simple
|
||||
# --interface eno1 = LAN-Bind (box-spezifisch; eno1 traegt 192.168.178.151).
|
||||
ExecStart=/usr/bin/ttyd --writable --interface eno1 --port 7681 --max-clients 2 --cwd %h/.hermes %h/.hermes/hermes-agent/venv/bin/python -m hermes_cli.main chat
|
||||
Restart=always
|
||||
RestartSec=3
|
||||
|
||||
[Install]
|
||||
WantedBy=default.target
|
||||
@@ -0,0 +1,114 @@
|
||||
# Auto-generiert von provision.sh - per Mission Control erweiterbar
|
||||
healthCheckTimeout: 300
|
||||
globalTTL: 0
|
||||
|
||||
models:
|
||||
Qwen3.6-35B-A3B:
|
||||
# parallel 2 + c 131072: 2 Slots à 65k (Slot 2 = Mem0-Lern-Extraktion, blockiert Voice-Turns
|
||||
# nicht mehr); KV Q8_0 macht das speicherneutral zu vorher (65k f16). Bench 2026-07-02:
|
||||
# Q8_0 kostet 0 t/s, MTP n-max 3 = +26% vs. ohne Spec.
|
||||
cmd: |
|
||||
llama-server -m /srv/models/Qwen3.6-35B-A3B-MTP-GGUF/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf --host 127.0.0.1 --port ${PORT} -c 131072 -ngl 999 -fa on --no-mmap --jinja --parallel 2 -cram 16384 -ctk q8_0 -ctv q8_0 --spec-type draft-mtp --spec-draft-n-max 3
|
||||
ttl: 0
|
||||
aliases:
|
||||
- hermes
|
||||
- fast
|
||||
# context = nutzbarer Kontext PRO REQUEST (c geteilt durch parallel-Slots).
|
||||
capabilities:
|
||||
in: [text]
|
||||
out: [text]
|
||||
tools: true
|
||||
context: 65536
|
||||
Qwen3.5-122B-A10B:
|
||||
cmd: |
|
||||
llama-server -m /srv/models/Qwen3.5-122B-A10B-GGUF/Q4_K_M/Qwen3.5-122B-A10B-Q4_K_M-00001-of-00003.gguf --host 127.0.0.1 --port ${PORT} -c 32768 -ngl 999 -fa on --no-mmap --jinja --cache-reuse 256 -cram 16384
|
||||
ttl: 600
|
||||
aliases:
|
||||
- heavy
|
||||
capabilities:
|
||||
in: [text]
|
||||
out: [text]
|
||||
tools: true
|
||||
context: 32768
|
||||
Qwen3-Coder-Next:
|
||||
cmd: |
|
||||
llama-server -m /srv/models/Qwen3-Coder-Next-GGUF/Qwen3-Coder-Next-Q4_K_M.gguf --host 127.0.0.1 --port ${PORT} -c 131072 -ngl 999 -fa on --no-mmap --jinja --cache-reuse 256 -cram 16384
|
||||
ttl: 600
|
||||
aliases:
|
||||
- coder
|
||||
capabilities:
|
||||
in: [text]
|
||||
out: [text]
|
||||
tools: true
|
||||
context: 131072
|
||||
Qwen3-VL-8B-Instruct:
|
||||
cmd: |
|
||||
llama-server -m /srv/models/Qwen3-VL-8B-Instruct-GGUF/Qwen3VL-8B-Instruct-Q4_K_M.gguf --host 127.0.0.1 --port ${PORT} -c 32768 -ngl 999 -fa on --no-mmap --mmproj /srv/models/Qwen3-VL-8B-Instruct-GGUF/mmproj-Qwen3VL-8B-Instruct-F16.gguf --jinja
|
||||
# ttl 0 (03.07.): Mitglied des Immer-bereit-Sets — ein Selbstauslöser widerspräche dem UI-Versprechen.
|
||||
ttl: 0
|
||||
aliases:
|
||||
- vision
|
||||
capabilities:
|
||||
in: [text, image]
|
||||
out: [text]
|
||||
context: 32768
|
||||
Qwen3-Embedding-0.6B:
|
||||
cmd: |
|
||||
llama-server -m /srv/models/Qwen3-Embedding-0.6B-GGUF/Qwen3-Embedding-0.6B-Q8_0.gguf --host 127.0.0.1 --port ${PORT} --embedding --pooling last -ngl 999 -fa on --no-mmap -c 8192
|
||||
ttl: 0
|
||||
aliases:
|
||||
- embed
|
||||
Qwen3-Coder-30B-A3B-Instruct:
|
||||
cmd: |
|
||||
llama-server -m /srv/models/Qwen3-Coder-30B-A3B-Instruct-GGUF/Qwen3-Coder-30B-A3B-Instruct-Q4_K_M.gguf --host 127.0.0.1 --port ${PORT} -c 131072 -ngl 999 -fa on --no-mmap --jinja --cache-reuse 256 -cram 16384
|
||||
ttl: 300
|
||||
aliases:
|
||||
- coder-lite
|
||||
capabilities:
|
||||
in: [text]
|
||||
out: [text]
|
||||
tools: true
|
||||
context: 131072
|
||||
gpt-oss-120b:
|
||||
# heavy-Kandidat (Bench 02.07.: tg 54-55 t/s vs. Qwen3.5-122B 23,5 t/s = 2,3x, 60 statt
|
||||
# 73 GB). Noch OHNE heavy-Alias — erst Qualitaets-Check (deutsch/Reasoning), dann Entscheid.
|
||||
cmd: |
|
||||
llama-server -m /srv/models/gpt-oss-120b-GGUF/gpt-oss-120b-mxfp4-00001-of-00003.gguf --host 127.0.0.1 --port ${PORT} -c 32768 -ngl 999 -fa on --no-mmap --jinja --cache-reuse 256 -cram 16384
|
||||
ttl: 600
|
||||
capabilities:
|
||||
in: [text]
|
||||
out: [text]
|
||||
tools: true
|
||||
context: 32768
|
||||
Qwen3-VL-30B-A3B-Instruct:
|
||||
# vision-Upgrade-Kandidat (Bench 02.07.: tg 90-92 t/s, MoE 3B aktiv, stark auf GUI-Benchmarks).
|
||||
# Noch OHNE vision-Alias — erst Screenshot-Qualitaets-Check, dann Entscheid (brains-Budget!).
|
||||
cmd: |
|
||||
llama-server -m /srv/models/Qwen3-VL-30B-A3B-Instruct-GGUF/Qwen3-VL-30B-A3B-Instruct-Q4_K_M.gguf --host 127.0.0.1 --port ${PORT} -c 32768 -ngl 999 -fa on --no-mmap --mmproj /srv/models/Qwen3-VL-30B-A3B-Instruct-GGUF/mmproj-F16.gguf --jinja
|
||||
ttl: 300
|
||||
capabilities:
|
||||
in: [text, image]
|
||||
out: [text]
|
||||
context: 32768
|
||||
GLM-4.6V-Flash:
|
||||
cmd: |
|
||||
llama-server -m /srv/models/GLM-4.6V-Flash-GGUF/GLM-4.6V-Flash-Q4_K_M.gguf --host 127.0.0.1 --port ${PORT} -c 131072 -ngl 999 -fa on --no-mmap --mmproj /srv/models/GLM-4.6V-Flash-GGUF/mmproj-BF16.gguf --jinja
|
||||
ttl: 300
|
||||
aliases:
|
||||
- scout
|
||||
capabilities:
|
||||
in: [text, image]
|
||||
out: [text]
|
||||
tools: true
|
||||
context: 131072
|
||||
groups:
|
||||
brains:
|
||||
swap: false
|
||||
# llama-swap versteht NUR `persistent` (Verdrängungsschutz); `persist` ist der
|
||||
# MC2-interne Lese-Key (API/UI) und wird von llama-swap ignoriert. Beide pflegen!
|
||||
persist: true
|
||||
persistent: true
|
||||
members:
|
||||
- Qwen3-Embedding-0.6B
|
||||
- Qwen3-VL-8B-Instruct
|
||||
- Qwen3.6-35B-A3B
|
||||
@@ -0,0 +1,8 @@
|
||||
[Unit]
|
||||
Description=MC2 Auto-Update (Router/Engine/Hermes mit Rollback+Pin, Telegram-Meldung)
|
||||
Documentation=file:%h/mission-control-v2/docs/AUTONOMIE_PLAN.md
|
||||
After=mc2-backup.service
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
ExecStart=/bin/bash %h/mission-control-v2/deploy/autoupdate.sh
|
||||
@@ -0,0 +1,10 @@
|
||||
[Unit]
|
||||
Description=Wöchentliches MC2 Auto-Update (So 04:30, nach dem 03:30-Backup)
|
||||
|
||||
[Timer]
|
||||
OnCalendar=Sun *-*-* 04:30:00
|
||||
Persistent=true
|
||||
RandomizedDelaySec=600
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
@@ -0,0 +1,7 @@
|
||||
[Unit]
|
||||
Description=MC2 Zustands-Backup (mem0 + Hermes-Configs/Secrets + llama-swap config)
|
||||
Documentation=file:%h/mission-control-v2/docs/BACKUP.md
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
ExecStart=/bin/bash %h/mission-control-v2/deploy/backup.sh
|
||||
@@ -0,0 +1,10 @@
|
||||
[Unit]
|
||||
Description=Tägliches MC2 Zustands-Backup
|
||||
|
||||
[Timer]
|
||||
OnCalendar=*-*-* 03:30:00
|
||||
Persistent=true
|
||||
RandomizedDelaySec=300
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
@@ -0,0 +1,7 @@
|
||||
[Unit]
|
||||
Description=MC2 Selbst-Smoke-Tests (Gateway/Tools/Voice aktiv durchspielen, Alarm bei Rot)
|
||||
Documentation=file:%h/mission-control-v2/docs/AUTONOMIE_PLAN.md
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
ExecStart=/bin/bash %h/mission-control-v2/deploy/self-smoke.sh
|
||||
@@ -0,0 +1,10 @@
|
||||
[Unit]
|
||||
Description=Täglicher MC2-Selbst-Smoke-Test (07:15, unabhängig von Updates)
|
||||
|
||||
[Timer]
|
||||
OnCalendar=*-*-* 07:15:00
|
||||
Persistent=true
|
||||
RandomizedDelaySec=300
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
@@ -0,0 +1,26 @@
|
||||
[Unit]
|
||||
Description=MC2 Mem0 Sidecar — auto-lernendes, semantisches Gedächtnis (mem0 + chroma)
|
||||
Documentation=https://github.com/mem0ai/mem0
|
||||
After=network.target
|
||||
|
||||
[Service]
|
||||
# Mem0 + chromadb laufen nur unter Python 3.12 (~/.mem0/venv) — das MC2-Backend (3.14)
|
||||
# kann sie nicht importieren. Darum dieser schlanke Sidecar; MC2 spricht ihn per HTTP an.
|
||||
# Bind 127.0.0.1: nur lokal erreichbar (MC2 proxyt nach außen).
|
||||
Type=simple
|
||||
WorkingDirectory=%h/mission-control-v2/mem0_service
|
||||
Environment=MEM0_PORT=8765
|
||||
Environment=MEM0_CHROMA_PATH=/srv/models/mem0/chroma
|
||||
Environment=MEM0_HISTORY_DB=/srv/models/mem0/history.db
|
||||
Environment=MEM0_EMBED_URL=http://127.0.0.1:8080/v1
|
||||
Environment=MEM0_LLM_URL=http://127.0.0.1:8080/v1
|
||||
Environment=MEM0_EMBED_MODEL=embed
|
||||
Environment=MEM0_LLM_MODEL=fast
|
||||
Environment=MEM0_EMBED_DIMS=1024
|
||||
Environment=TOKENIZERS_PARALLELISM=false
|
||||
ExecStart=%h/.mem0/venv/bin/python -m uvicorn app:app --host 127.0.0.1 --port 8765
|
||||
Restart=always
|
||||
RestartSec=3
|
||||
|
||||
[Install]
|
||||
WantedBy=default.target
|
||||
@@ -0,0 +1,32 @@
|
||||
# systemd-USER-Unit für Mission Control 2.0 (PARALLEL zu v1, Port 9001).
|
||||
# Läuft sudo-frei aus dem Home-Verzeichnis (Nordstern: kein Passwort/sudo).
|
||||
# Ablage: ~/.config/systemd/user/mission-control-2.service ; dann:
|
||||
# systemctl --user daemon-reload
|
||||
# systemctl --user enable --now mission-control-2
|
||||
# loginctl enable-linger hitonabi # läuft auch ohne aktive Session
|
||||
|
||||
[Unit]
|
||||
Description=Mission Control 2.0 (Cockpit)
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
WorkingDirectory=%h/mission-control-v2/backend
|
||||
ExecStart=%h/mission-control-v2/backend/.venv/bin/python -m uvicorn app:app --host 0.0.0.0 --port 9001
|
||||
Environment=MC_PORT=9001
|
||||
Environment=MC_LLAMA_SWAP_URL=http://127.0.0.1:8080
|
||||
Environment=MC_CONFIG_PATH=/etc/llama-swap/config.yaml
|
||||
Environment=MC_MODELS_DIR=/srv/models
|
||||
# Geteiltes Gedächtnis = die bestehende v1-DB (Kontinuität bis/über Cutover).
|
||||
Environment=MC_MEMORY_DB=/srv/models/mission-control-memory.db
|
||||
# KEIN MC_ENGINE_UPDATE_CMD-Override mehr: Das alte /usr/local/bin/update-llamacpp zog den
|
||||
# ROCm-Build nach /opt/llamacpp (totes Rollback-Dir) statt des aktiven Vulkan-Builds → Updates
|
||||
# liefen ins Leere ("DONE", aber nichts passierte). Ohne Override nutzt das Backend den Default
|
||||
# `sudo bash <repo>/deploy/update-engine.sh` (Vulkan, mit Backup/Stack-Check/Auto-Rollback).
|
||||
Restart=on-failure
|
||||
|
||||
RestartSec=3
|
||||
|
||||
[Install]
|
||||
WantedBy=default.target
|
||||
@@ -0,0 +1,59 @@
|
||||
#!/usr/bin/env bash
|
||||
# Meldeweg der Box (Autonomie E1): eine Nachricht an den User schicken.
|
||||
# Primär: hermes send → Telegram (nutzt Gateway-Credentials, kein LLM nötig).
|
||||
# Fallback: Logfile + wall (falls Telegram/hermes nicht erreichbar).
|
||||
#
|
||||
# Nutzung: notify.sh "Nachricht" (oder via stdin: echo msg | notify.sh)
|
||||
# notify.sh -s "[Update]" "Text" (Betreffzeile voranstellen)
|
||||
# Exit 0 = zugestellt (Telegram ODER Fallback-Log geschrieben).
|
||||
set -uo pipefail
|
||||
|
||||
LOG="${MC_NOTIFY_LOG:-$HOME/mc2-notify.log}"
|
||||
SUBJECT=""
|
||||
|
||||
while getopts ":s:" opt; do
|
||||
case "$opt" in
|
||||
s) SUBJECT="$OPTARG" ;;
|
||||
*) echo "usage: notify.sh [-s subject] <nachricht>" >&2; exit 2 ;;
|
||||
esac
|
||||
done
|
||||
shift $((OPTIND - 1))
|
||||
|
||||
MSG="${1:-}"
|
||||
[ -z "$MSG" ] && [ ! -t 0 ] && MSG="$(cat)"
|
||||
if [ -z "$MSG" ]; then
|
||||
echo "usage: notify.sh [-s subject] <nachricht>" >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
TS="$(date '+%Y-%m-%d %H:%M:%S')"
|
||||
|
||||
# Lucy-Briefkasten (Proaktivität A3): Meldung zusätzlich an MC2 spiegeln — die Desktop-Lucy
|
||||
# spricht sie dann von sich aus. Best-effort (darf den Telegram-Weg nie aufhalten).
|
||||
# MC_NOTIFY_NO_ANNOUNCE=1 unterdrückt das (der Health-Wächter legt seine Einträge selbst ab).
|
||||
if [ "${MC_NOTIFY_NO_ANNOUNCE:-0}" != "1" ]; then
|
||||
curl -sf -m 3 -X POST "${MC_ANNOUNCE_URL:-http://127.0.0.1:9001/api/voice/announce}" \
|
||||
-H 'Content-Type: application/json' \
|
||||
--data "$(python3 - "$SUBJECT" "$MSG" <<'PY'
|
||||
import json, sys
|
||||
print(json.dumps({"text": sys.argv[2], "subject": sys.argv[1], "source": "notify"}))
|
||||
PY
|
||||
)" >/dev/null 2>&1 || true
|
||||
fi
|
||||
|
||||
# Primärweg: Telegram via hermes send (Login-Shell-PATH, falls aus Timer/cron aufgerufen).
|
||||
if [ -n "$SUBJECT" ]; then
|
||||
SEND_OUT="$(bash -lc 'hermes send --to telegram --subject "$1" -- "$2"' _ "$SUBJECT" "$MSG" 2>&1)"
|
||||
else
|
||||
SEND_OUT="$(bash -lc 'hermes send --to telegram -- "$1"' _ "$MSG" 2>&1)"
|
||||
fi
|
||||
if [ $? -eq 0 ]; then
|
||||
echo "$TS OK telegram: $MSG" >> "$LOG"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# Fallback: Logfile + wall — Meldung darf nie verloren gehen.
|
||||
echo "$TS FALLBACK (telegram fehlgeschlagen: $SEND_OUT): $MSG" >> "$LOG"
|
||||
printf '%s\n' "MC2-Meldung: ${SUBJECT:+$SUBJECT }$MSG" | wall 2>/dev/null || true
|
||||
echo "WARN: Telegram fehlgeschlagen, in $LOG protokolliert" >&2
|
||||
exit 0
|
||||
@@ -0,0 +1,68 @@
|
||||
#!/usr/bin/env bash
|
||||
# Provisioniert die Inferenz-Engine auf der AI Box (Strix Halo / Ryzen AI MAX+ 395, gfx1151)
|
||||
# auf **Vulkan/RADV** — reproduzierbar. Braucht root (sudo).
|
||||
#
|
||||
# sudo bash ~/mission-control-v2/deploy/provision-engine.sh
|
||||
#
|
||||
# Hintergrund: Auf gfx1151 ist Vulkan/RADV ggü. ROCm/HIP messbar schneller
|
||||
# (Token-Gen +12–22 %, Prefill gleich; auf der Box per llama-bench verifiziert 2026-06-27)
|
||||
# UND einfacher. Der ROCm-Build bleibt unter /opt/llamacpp als Rollback liegen.
|
||||
set -euo pipefail
|
||||
|
||||
VULKAN_DIR=/opt/llamacpp-vulkan
|
||||
WARMUP_DST=/usr/local/bin/llama-swap-warmup.sh
|
||||
DROPIN_DIR=/etc/systemd/system/llama-swap.service.d
|
||||
SRC_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
# Optional gepinnter Build (z.B. b9821). Leer = neuester Release.
|
||||
PIN_BUILD="${MC_ENGINE_BUILD:-}"
|
||||
|
||||
echo "==> 1. RADV-Treiber + Vulkan-Runtime"
|
||||
export DEBIAN_FRONTEND=noninteractive
|
||||
apt-get update -qq
|
||||
apt-get install -y mesa-vulkan-drivers libvulkan1 vulkan-tools curl jq
|
||||
|
||||
echo "==> 2. llama.cpp Vulkan-Build nach $VULKAN_DIR"
|
||||
if [ -n "$PIN_BUILD" ]; then
|
||||
TAG="$PIN_BUILD"
|
||||
else
|
||||
TAG="$(curl -s https://api.github.com/repos/ggml-org/llama.cpp/releases/latest | jq -r .tag_name)"
|
||||
fi
|
||||
ASSET="llama-${TAG}-bin-ubuntu-vulkan-x64.tar.gz"
|
||||
URL="https://github.com/ggml-org/llama.cpp/releases/download/${TAG}/${ASSET}"
|
||||
TMP="$(mktemp -d)"
|
||||
echo " Lade $URL"
|
||||
curl -sL -o "$TMP/v.tgz" "$URL"
|
||||
mkdir -p "$VULKAN_DIR"
|
||||
tar xzf "$TMP/v.tgz" -C "$TMP"
|
||||
# Tarball entpackt nach .../llama-<tag>/ — Inhalt flach nach $VULKAN_DIR
|
||||
cp -rf "$TMP"/llama-*/. "$VULKAN_DIR"/
|
||||
rm -rf "$TMP"
|
||||
test -x "$VULKAN_DIR/llama-server"
|
||||
|
||||
echo "==> 3. Symlink llama-server -> Vulkan-Build"
|
||||
ln -sfn "$VULKAN_DIR/llama-server" /usr/local/bin/llama-server
|
||||
|
||||
echo "==> 4. systemd Drop-ins für llama-swap (LD_LIBRARY_PATH + Brain-Warmup)"
|
||||
mkdir -p "$DROPIN_DIR"
|
||||
cat > "$DROPIN_DIR/vulkan.conf" <<EOF
|
||||
[Service]
|
||||
Environment=LD_LIBRARY_PATH=$VULKAN_DIR
|
||||
EOF
|
||||
install -m 0755 "$SRC_DIR/warmup.sh" "$WARMUP_DST"
|
||||
cat > "$DROPIN_DIR/warmup.conf" <<EOF
|
||||
[Service]
|
||||
# Nach jedem (Re)Start die brains vorladen. Das Skript detacht sich selbst (blockiert
|
||||
# den Start nicht); '-' macht den Aufruf fehlertolerant (kann llama-swap nie failen lassen).
|
||||
ExecStartPost=-$WARMUP_DST
|
||||
EOF
|
||||
|
||||
echo "==> 5. Reload + Restart"
|
||||
systemctl daemon-reload
|
||||
systemctl restart llama-swap
|
||||
sleep 3
|
||||
systemctl is-active llama-swap
|
||||
|
||||
echo "==> Fertig. Aktive Engine:"
|
||||
readlink -f /usr/local/bin/llama-server
|
||||
vulkaninfo --summary 2>/dev/null | grep -m1 deviceName || true
|
||||
echo "Rollback auf ROCm: ln -sfn /opt/llamacpp/llama-server /usr/local/bin/llama-server && rm $DROPIN_DIR/vulkan.conf && systemctl daemon-reload && systemctl restart llama-swap"
|
||||
@@ -0,0 +1,25 @@
|
||||
#!/usr/bin/env bash
|
||||
# Feed für den Evolution-Radar-Cron (Autonomie E4). Liegt in ~/.hermes/scripts/ (deploy.sh
|
||||
# kopiert ihn); sein stdout wird dem Cron-Agenten als Prompt eingespeist:
|
||||
# versionierter Auftrag (deploy/radar-prompt.md im MC2-Repo) + Live-Inventar der Box.
|
||||
set -uo pipefail
|
||||
|
||||
cat "$HOME/mission-control-v2/deploy/radar-prompt.md"
|
||||
|
||||
echo
|
||||
echo "## LIVE-INVENTAR (automatisch erhoben am $(date '+%d.%m.%Y'))"
|
||||
echo
|
||||
echo "### Update-Status (MC2 /api/maintenance/updates):"
|
||||
curl -sf --max-time 30 http://127.0.0.1:9001/api/maintenance/updates \
|
||||
| jq -c '{os_pakete: .os, engine_update: .engine, router_update: .swap, komponenten: .components}' \
|
||||
2>/dev/null || echo "(MC2-API nicht erreichbar)"
|
||||
echo
|
||||
echo "### Installierte Modelle (llama-swap):"
|
||||
curl -sf --max-time 20 http://127.0.0.1:8080/v1/models | jq -r '.data[].id' 2>/dev/null \
|
||||
|| echo "(llama-swap nicht erreichbar)"
|
||||
echo
|
||||
# llama-server braucht LD_LIBRARY_PATH auf das Vulkan-Verzeichnis für --version.
|
||||
echo "### Versionen: llama.cpp $( LD_LIBRARY_PATH=/opt/llamacpp-vulkan /usr/local/bin/llama-server --version 2>&1 | grep -o 'version: [0-9]*' || true ) · llama-swap $( /usr/local/bin/llama-swap --version 2>/dev/null | head -1 || true ) · hermes-agent $( bash -lc 'hermes --version' 2>/dev/null | head -1 || echo '?' )"
|
||||
echo
|
||||
echo "### Pin-Register (gepinnte Ebenen NICHT zum Update vorschlagen, aber Sprünge melden):"
|
||||
cat /srv/models/mc2-pins.json 2>/dev/null || echo "{}"
|
||||
@@ -0,0 +1,64 @@
|
||||
# Evolution-Radar (monatlicher Auftrag)
|
||||
|
||||
Du bist das Evolution-Radar der AI-Box (AMD Strix Halo, 128 GB, 100 % lokal, Vulkan/RADV).
|
||||
Dein Job: EINMAL im Monat prüfen, ob sich bei unseren Kern-Komponenten oder in deren
|
||||
Kategorien ein echter Sprung ergeben hat — und das als kurzen deutschen Report melden.
|
||||
Unten ist das Live-Inventar der Box eingespeist (Update-Status, installierte Modelle, Pins).
|
||||
|
||||
## Schritt 1 (PFLICHT, vor jeder Recherche): Inventar lesen und respektieren
|
||||
|
||||
Lehre aus Report #1 (02.07.): vier Fehler, eine Ursache — Inventar ignoriert. Deshalb hart:
|
||||
- **Empfiehl NIE ein Update auf eine Version, die laut Live-Inventar schon installiert ist.**
|
||||
(Report #1 empfahl „Hermes v0.18" — die Box WAR auf v0.18.0. `komponenten` im Inventar
|
||||
zeigt das echte git-behind; behind=0 heißt AKTUELL, egal was Release-Notes suggerieren.)
|
||||
- **Bei Komponenten, die wir aktiv nutzen, erst die EIGENE Version/Nutzung feststellen,
|
||||
dann Neuigkeiten bewerten.** Aktiv im Einsatz (Stand Repo):
|
||||
Parakeet via onnx-asr = UNSER STT (25 EU-Sprachen inkl. Deutsch, ~0,45 s — nie als
|
||||
„kein Deutsch-Fokus" abtun) · Silero VAD **v5** via vad-web in Lucy (v6 existiert) ·
|
||||
pocket-tts german_24l (CPU, auf dem PC) · Hirn Qwen3.6-35B (MTP-Draft, immer warm).
|
||||
- **Die Modell-Liste im Inventar enthält auch schon geladene KANDIDATEN** (z. B. gpt-oss-120b,
|
||||
Qwen3-VL-30B) — die liegen bereit und sind gebencht; nicht als „Neuentdeckung" verkaufen.
|
||||
- Jede Versions-Aussage im Report muss gegen das Inventar geprüft sein.
|
||||
|
||||
## Arbeitsweise: Fan-out mit Subagents (WICHTIG)
|
||||
|
||||
Nutze `delegate_task` mit einem `tasks`-Array (role: leaf), um die Recherche in ISOLIERTE
|
||||
Subagents aufzuteilen — pro Task eine Komponenten-Gruppe. Jeder Subagent recherchiert per
|
||||
Web-Suche, FILTERT Marketing-Geschwätz aus und liefert dir nur eine Substanz-Zusammenfassung
|
||||
(max. 5 Zeilen pro Komponente). Du konsolidierst am Ende. Erwähne im Endbericht nur, was
|
||||
Substanz hat — „nichts Nennenswertes" ist ein gültiges und gutes Ergebnis je Gruppe.
|
||||
|
||||
Scheitern Subagents wiederholt (z. B. Kontext-Fehler), recherchiere die betroffenen Gruppen
|
||||
selbst sequenziell — der Report muss IMMER zustande kommen.
|
||||
|
||||
Task-Aufteilung (5 Subagents):
|
||||
1. **Engine/Router:** llama.cpp (Vulkan, gfx1151/Strix Halo relevant!), llama-swap (mostlygeek)
|
||||
2. **Agent/Runtime:** hermes-agent (NousResearch), Electron (Major-Sprünge), three-vrm/@pixiv
|
||||
3. **Voice:** pocket-tts (kyutai), onnx-asr/Parakeet, Silero-VAD, smart-turn — nur Deutsch-taugliches
|
||||
4. **Modell-Kategorien** für 128-GB-Strix-Halo: bessere lokale Coder-Modelle (vs. Qwen3-Coder-Next),
|
||||
Chat/Reasoning-MoE (vs. Qwen3.6-35B/Qwen3.5-122B), Vision (vs. Qwen3-VL), deutsche TTS/STT.
|
||||
Nur GGUF-/llama.cpp-lauffähig, Substanz = Benchmarks/Community-Erfahrung, kein Ankündigungs-Hype.
|
||||
5. **IDE-/Agent-Tools** (Zed, Kilo Code, Claude Code): besonders Tod-/Nachfolger-Meldungen —
|
||||
„Projekt eingestellt/archiviert/Fork übernimmt" ist GENAU die Meldung, die wir fangen müssen
|
||||
(Lehre: Roo Code galt als Empfehlung und war längst eingestellt). Auch: neues dominantes Tool?
|
||||
|
||||
## Bereits ENTSCHIEDEN — nicht wieder vorschlagen (Verdikte der Box)
|
||||
|
||||
- Vulkan/RADV statt ROCm (gemessen schneller auf gfx1151) · llama.cpp+llama-swap gesetzt
|
||||
- Hirn = Qwen3.6-35B-A3B mit MTP-Spec-Draft, immer warm · Mem0 als Memory · Electron-App für Lucy
|
||||
- Pocket-TTS für Lucys Stimme (deutsch, CPU) · KEIN Kyutai-Streaming-STT (kann kein Deutsch)
|
||||
- VERWORFEN: LobeChat/WebUI-Ersatz, Unmute-Vollstack, AnythingLLM, gemma als Hirn
|
||||
- Latenz ist heilig: nichts vorschlagen, was Lucys ~2-s-Sprech-Latenz gefährdet
|
||||
|
||||
## Report-Format (Endantwort = geht direkt als Telegram-Nachricht raus)
|
||||
|
||||
- **Schreibe die Endnachricht in Lucys Stimme an den Commander** (du BIST Lucy, siehe SOUL.md —
|
||||
das Radar ist nur dein Auftrag): erst 1–2 Sätze unmissverständlich, was die Funde für ihn
|
||||
bedeuten und ob er etwas tun muss — danach die knappen Fund-Zeilen.
|
||||
- Deutsch, maximal ~20 Zeilen, Du-Form, kein Markdown-Overkill, nicht technisch —
|
||||
Technik in Alltagssprache, Details nur auf Nachfrage
|
||||
- Struktur: „🛰️ Evolution-Radar <Monat>" → je Fund 1–3 Zeilen: WAS, WARUM relevant für UNS,
|
||||
Einstufung **[lohnt vermutlich]** / **[beobachten]** / ggf. Quelle kurz
|
||||
- Wenn ein Fund ein GEPINNTES Level betrifft (siehe Pins im Inventar): explizit erwähnen
|
||||
- Nichts gefunden? Dann genau das in 2 Zeilen sagen. Keine Pflicht-Funde erfinden.
|
||||
- KEINE Konfig-Änderungen vornehmen, KEINE Downloads starten — nur melden.
|
||||
@@ -0,0 +1,132 @@
|
||||
#!/usr/bin/env bash
|
||||
# Wiederherstellung eines mc2-state-Backups (siehe backup.sh).
|
||||
# Stoppt die Dienste, spielt den Zustand zurück, startet neu — und macht VORHER
|
||||
# automatisch ein Sicherheits-Backup des aktuellen Zustands (Restore ist reversibel).
|
||||
#
|
||||
# Nutzung:
|
||||
# bash restore.sh --list # vorhandene Backups zeigen (lokal + Off-Box)
|
||||
# bash restore.sh --dry-run <datei|latest> # nur anzeigen, was passieren würde
|
||||
# bash restore.sh [--yes] <datei|latest> # wiederherstellen (--yes = ohne Rückfrage)
|
||||
# bash restore.sh --pull-offsite # Off-Box-Kopie (Proxmox) nach lokal holen
|
||||
#
|
||||
# Off-Box (C12): fehlt ein Backup lokal (z. B. Box-Platte neu), wird es automatisch
|
||||
# vom Zweitgerät (Proxmox) geholt. Ziel/Key wie in backup.sh (überschreibbar per Env).
|
||||
set -euo pipefail
|
||||
|
||||
MODELS_DIR="${MC_MODELS_DIR:-/srv/models}"
|
||||
DEST_DIR="$MODELS_DIR/mc2-backups"
|
||||
MEM0_DIR="${MC_MEM0_DIR:-/srv/models/mem0}"
|
||||
LSWAP="${MC_CONFIG_PATH:-/etc/llama-swap/config.yaml}"
|
||||
HERMES="${HERMES_HOME:-$HOME/.hermes}"
|
||||
SRC="${MC2_SRC:-$HOME/mission-control-v2}"
|
||||
SERVICES="mem0-service mission-control-2 hermes-gateway"
|
||||
OFFSITE="${MC_BACKUP_OFFSITE:-root@192.168.178.108:/var/lib/vz/mc2-backups}"
|
||||
OFFSITE_KEY="${MC_BACKUP_OFFSITE_KEY:-$HOME/.ssh/mc2_offsite}"
|
||||
OFFSITE_SSH="ssh -i $OFFSITE_KEY -o BatchMode=yes -o ConnectTimeout=10 -o StrictHostKeyChecking=accept-new"
|
||||
|
||||
# Off-Box-Tarballs (Dateinamen) auflisten — leer, wenn nicht erreichbar/konfiguriert.
|
||||
offsite_ls() {
|
||||
[ -n "$OFFSITE" ] && [ -f "$OFFSITE_KEY" ] || return 0
|
||||
local host="${OFFSITE%%:*}" path="${OFFSITE#*:}"
|
||||
$OFFSITE_SSH "$host" "ls -1 $path/mc2-state-*.tar.gz 2>/dev/null" 2>/dev/null | xargs -r -n1 basename || true
|
||||
}
|
||||
# Ein einzelnes Off-Box-Tarball nach lokal holen. $1 = Dateiname (basename).
|
||||
offsite_fetch() {
|
||||
[ -n "$OFFSITE" ] && [ -f "$OFFSITE_KEY" ] || return 1
|
||||
mkdir -p "$DEST_DIR"
|
||||
echo "→ hole $1 vom Zweitgerät ($OFFSITE)…" >&2
|
||||
rsync -a -e "$OFFSITE_SSH" "$OFFSITE/$1" "$DEST_DIR/"
|
||||
}
|
||||
|
||||
usage() { echo "Nutzung: restore.sh --list | restore.sh --pull-offsite | restore.sh [--dry-run] [--yes] <datei|latest>"; }
|
||||
list_backups() {
|
||||
echo "Verfügbare Backups in $DEST_DIR:"
|
||||
local found=0 f
|
||||
for f in $(ls -1t "$DEST_DIR"/mc2-state-*.tar.gz 2>/dev/null || true); do
|
||||
printf " %s (%s)\n" "$(basename "$f")" "$(du -h "$f" | cut -f1)"
|
||||
found=1
|
||||
done
|
||||
[ "$found" = 1 ] || echo " (keine)"
|
||||
local off; off="$(offsite_ls || true)"
|
||||
if [ -n "$off" ]; then
|
||||
echo "Off-Box ($OFFSITE):"
|
||||
echo "$off" | sed 's/^/ /'
|
||||
fi
|
||||
}
|
||||
|
||||
[ $# -eq 0 ] && { usage; exit 1; }
|
||||
DRY=0; YES=0; TARGET=""
|
||||
for a in "$@"; do
|
||||
case "$a" in
|
||||
--list) list_backups; exit 0 ;;
|
||||
--pull-offsite)
|
||||
mkdir -p "$DEST_DIR"
|
||||
[ -n "$OFFSITE" ] && [ -f "$OFFSITE_KEY" ] || { echo "Off-Box nicht konfiguriert."; exit 1; }
|
||||
echo "→ spiegle Off-Box → lokal ($DEST_DIR)…"
|
||||
rsync -a -e "$OFFSITE_SSH" --include='mc2-state-*.tar.gz' --exclude='*' "$OFFSITE/" "$DEST_DIR/"
|
||||
list_backups; exit 0 ;;
|
||||
--dry-run) DRY=1 ;;
|
||||
--yes) YES=1 ;;
|
||||
-*) usage; exit 1 ;;
|
||||
*) TARGET="$a" ;;
|
||||
esac
|
||||
done
|
||||
[ -z "$TARGET" ] && { usage; exit 1; }
|
||||
# 'latest' bevorzugt lokal; fehlt lokal alles, das jüngste Off-Box-Tarball holen.
|
||||
if [ "$TARGET" = "latest" ]; then
|
||||
TARGET="$(ls -1t "$DEST_DIR"/mc2-state-*.tar.gz 2>/dev/null | head -1 || true)"
|
||||
if [ -z "$TARGET" ]; then
|
||||
latest_off="$(offsite_ls | sort | tail -1 || true)"
|
||||
[ -n "$latest_off" ] && offsite_fetch "$latest_off" && TARGET="$DEST_DIR/$latest_off"
|
||||
fi
|
||||
fi
|
||||
# Konkreten Namen zuerst lokal suchen, sonst vom Zweitgerät holen.
|
||||
if [ ! -f "$TARGET" ]; then
|
||||
base="$(basename "${TARGET:-}")"
|
||||
if [ -f "$DEST_DIR/$base" ]; then
|
||||
TARGET="$DEST_DIR/$base"
|
||||
elif [ -n "$base" ] && offsite_ls | grep -qx "$base"; then
|
||||
offsite_fetch "$base" && TARGET="$DEST_DIR/$base"
|
||||
fi
|
||||
fi
|
||||
[ -f "$TARGET" ] || { echo "Backup nicht gefunden: ${TARGET:-<leer>}"; echo; list_backups; exit 1; }
|
||||
|
||||
echo "Restore-Quelle: $TARGET"
|
||||
STAGE="$(mktemp -d)"; trap 'rm -rf "$STAGE"' EXIT
|
||||
tar -xzf "$TARGET" -C "$STAGE"
|
||||
echo "--- Inhalt ---"; cat "$STAGE/MANIFEST.txt" 2>/dev/null || true; echo "--------------"
|
||||
|
||||
if [ "$DRY" = 1 ]; then
|
||||
echo "[dry-run] würde zurücksetzen:"
|
||||
[ -d "$STAGE/mem0" ] && echo " • mem0 → $MEM0_DIR (wird ersetzt)"
|
||||
[ -f "$STAGE/hermes/config.yaml" ] && echo " • $HERMES/config.yaml"
|
||||
[ -f "$STAGE/hermes/.env" ] && echo " • $HERMES/.env"
|
||||
[ -d "$STAGE/hermes/plugins" ] && echo " • $HERMES/plugins/"
|
||||
[ -f "$STAGE/llama-swap/config.yaml" ] && echo " • $LSWAP"
|
||||
echo " • Dienste neu starten: $SERVICES"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [ "$YES" != 1 ]; then
|
||||
echo "WARNUNG: Das überschreibt den AKTUELLEN Zustand und startet die Dienste neu."
|
||||
read -r -p "Fortfahren? (tippe 'ja'): " ans
|
||||
[ "$ans" = "ja" ] || { echo "Abgebrochen."; exit 1; }
|
||||
fi
|
||||
|
||||
echo "→ Sicherheits-Backup des aktuellen Zustands…"
|
||||
bash "$SRC/deploy/backup.sh" || echo " (Pre-Restore-Backup fehlgeschlagen — fahre fort)"
|
||||
|
||||
echo "→ Dienste stoppen…"
|
||||
systemctl --user stop $SERVICES 2>/dev/null || true
|
||||
|
||||
if [ -d "$STAGE/mem0" ]; then rm -rf "$MEM0_DIR"; mkdir -p "$MEM0_DIR"; cp -a "$STAGE/mem0/." "$MEM0_DIR/"; fi
|
||||
[ -f "$STAGE/hermes/config.yaml" ] && { mkdir -p "$HERMES"; cp -a "$STAGE/hermes/config.yaml" "$HERMES/"; }
|
||||
[ -f "$STAGE/hermes/.env" ] && cp -a "$STAGE/hermes/.env" "$HERMES/"
|
||||
[ -d "$STAGE/hermes/plugins" ] && { mkdir -p "$HERMES/plugins"; cp -a "$STAGE/hermes/plugins/." "$HERMES/plugins/"; }
|
||||
[ -f "$STAGE/llama-swap/config.yaml" ] && cp -a "$STAGE/llama-swap/config.yaml" "$LSWAP" 2>/dev/null || true
|
||||
|
||||
echo "→ Dienste starten…"
|
||||
systemctl --user start $SERVICES 2>/dev/null || true
|
||||
sleep 4
|
||||
printf "→ Health: "; curl -sf http://127.0.0.1:9001/api/health && echo || echo "(Backend nicht erreichbar — Logs prüfen)"
|
||||
echo "OK — Restore abgeschlossen."
|
||||
@@ -0,0 +1,39 @@
|
||||
#!/usr/bin/env bash
|
||||
# Feed für den Selbstkritik-Cron (Faden 11b — Zwilling des Evolution-Radars, Blick nach INNEN).
|
||||
# Liegt in ~/.hermes/scripts/ (deploy.sh kopiert ihn); stdout wird dem Cron-Agenten als
|
||||
# Prompt eingespeist: versionierter Auftrag (deploy/selbstkritik-prompt.md) + Live-Betriebsdaten.
|
||||
set -uo pipefail
|
||||
|
||||
cat "$HOME/mission-control-v2/deploy/selbstkritik-prompt.md"
|
||||
|
||||
echo
|
||||
echo "## LIVE-DATEN (automatisch erhoben am $(date '+%d.%m.%Y %H:%M'))"
|
||||
echo
|
||||
echo "### Voice-Pipeline-Latenzen (rollend, ms — :9001/api/voice/metrics):"
|
||||
curl -sf --max-time 15 http://127.0.0.1:9001/api/voice/metrics | jq -c . 2>/dev/null \
|
||||
|| echo "(nicht erreichbar)"
|
||||
echo
|
||||
echo "### Token-Verbrauch / Cloud-Ersparnis:"
|
||||
curl -sf --max-time 15 http://127.0.0.1:9001/api/system/token-stats | jq -c . 2>/dev/null \
|
||||
|| echo "(nicht erreichbar)"
|
||||
echo
|
||||
echo "### Prompt-Budget (Hinweis: prompt-size ist Allowlist-blind — die Voice-Lane hat real"
|
||||
echo "### deutlich weniger Tools; als Trend-Wächter für den GESAMT-Bestand trotzdem nützlich):"
|
||||
bash -lc "hermes prompt-size" 2>/dev/null | grep -E "total|skills index|Tool schemas" || echo " (Messung fehlgeschlagen)"
|
||||
echo
|
||||
echo "### Gateway-Journal: häufigste Warnungen/Fehler der letzten 30 Tage (Anzahl · Muster):"
|
||||
journalctl --user -u hermes-gateway -p warning --since "30 days ago" --no-pager 2>/dev/null \
|
||||
| grep -oE "(WARNING|ERROR)[^:]*: .{0,90}" | sed "s/[0-9]\{3,\}/N/g" | sort | uniq -c | sort -rn | head -12 \
|
||||
|| echo "(keine oder Journal nicht lesbar)"
|
||||
echo
|
||||
echo "### Tool-Loop-Guardrail-Treffer (letzte 30 Tage):"
|
||||
n=$(journalctl --user -u hermes-gateway --since "30 days ago" --no-pager 2>/dev/null \
|
||||
| grep -ci "loop guard\|hard_stop\|tool loop" || true)
|
||||
echo "${n:-0}"
|
||||
echo
|
||||
echo "### Skill-Bestand & Curator:"
|
||||
bash -lc "hermes curator status" 2>/dev/null | head -14 || echo "(Curator-Status fehlgeschlagen)"
|
||||
echo
|
||||
echo "### Meldeweg-Ausfälle (notify.sh-Fallbacks, letzte 30 Tage):"
|
||||
n=$(grep -c "FALLBACK" "$HOME/mc2-notify.log" 2>/dev/null || true)
|
||||
echo "${n:-0}"
|
||||
@@ -0,0 +1,27 @@
|
||||
# Selbstkritik-Runde (monatlich) — der Blick nach INNEN
|
||||
|
||||
Du bist Lucy und schaust einmal im Monat kritisch auf DICH SELBST und deinen eigenen
|
||||
Betrieb: Wo warst du langsam, wo liefen Tools in Schleifen, wo häufen sich Warnungen,
|
||||
was wird nie benutzt? Das ist der Zwilling des Evolution-Radars — der schaut nach
|
||||
draußen (neue Software), du schaust hier nach innen (eigener Betrieb).
|
||||
|
||||
## Auftrag
|
||||
|
||||
1. Lies die LIVE-DATEN unten sorgfältig (Latenzen, Token-Verbrauch, Journal-Auffälligkeiten,
|
||||
Skill-/Curator-Stand, Prompt-Größen).
|
||||
2. Finde die 1 bis 3 WICHTIGSTEN konkreten Verbesserungen. Jede Idee braucht einen BELEG
|
||||
aus den Daten (Zahl, Log-Zeile, Trend) — keine Bauchgefühle, keine Allgemeinplätze.
|
||||
3. Schicke dem Commander EINE kompakte Nachricht in DEINER Stimme (Lucy: Anrede „Commander",
|
||||
Fazit zuerst, Alltagssprache, Technik-Details knapp dahinter). Je Vorschlag: Was ist
|
||||
auffällig (Beleg) → was schlägst du vor → was bringt es. Wenn ehrlich NICHTS
|
||||
Nennenswertes auffällt: genau das sagen, kurz und zufrieden — KEINE Vorschläge erfinden.
|
||||
|
||||
## Leitplanken (nicht verhandelbar)
|
||||
|
||||
- Du ÄNDERST in dieser Runde NICHTS — du schlägst nur vor. Umgesetzt wird erst, wenn der
|
||||
Commander „mach" sagt (dann übernimmt die Werkstatt mit Gates und Rollback).
|
||||
- Der Hermes-Quellcode ist TABU (fremde Software, Auto-Update-Kanal). Verbesserungen nur
|
||||
über UNSERE Schichten: Config, SOUL.md, Skills, Gedächtnis, Modellwahl, MC2-Code.
|
||||
- Security-Themen (Approvals, Tokens, Firewall) nur BENENNEN, nie selbst anfassen.
|
||||
- Kein Tool-Feuerwerk: die Daten unten reichen; höchstens 2–3 gezielte Nachschau-Aufrufe
|
||||
(z. B. eine Journal-Zeile im Kontext lesen), dann antworten.
|
||||
@@ -0,0 +1,99 @@
|
||||
#!/usr/bin/env bash
|
||||
# Selbstreparatur-Modus (Autonomie 0g, Kür): Macht ein Hermes-Update den Gehirn-Check rot,
|
||||
# WEIL sich Config-Schlüssel geändert haben (v0.18-Klasse: „Unknown key … → stiller Default →
|
||||
# alle Tools hingen"), versucht die Box den Config-Bruch SELBST zu ziehen — bevor zurückgerollt
|
||||
# wird, um die neue Version zu behalten. NUR Config, NIE Code, NIE Security-Schlüssel.
|
||||
#
|
||||
# Sicher durch Kapselung — jeder Schritt ist reversibel und hart gegatet:
|
||||
# Backup → nur EINDEUTIGE (genau 1× vorkommende), nicht-sicherheitsrelevante Keys auskommentieren
|
||||
# → YAML muss gültig bleiben → Gehirn-Check (hermes-postcheck.sh) muss GRÜN werden → sonst zurück.
|
||||
# Schlimmster Fall bei Fehlgriff = exakt der heutige: Rollback. Der Aufrufer (autoupdate.sh) ruft
|
||||
# dies VOR seinem Rollback; bei Rückgabe 2 rollt er wie gehabt zurück + pinnt.
|
||||
#
|
||||
# Rückgabe: 0 = repariert (Gehirn-Check danach grün) · 2 = nichts sicher reparierbar.
|
||||
# Gibt bei Misserfolg eine Zeile „SELFREPAIR_SUGGESTION=<LLM-Diagnose>" auf stdout aus.
|
||||
set -uo pipefail
|
||||
|
||||
SRC_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
HERMES="${HERMES_HOME:-$HOME/.hermes}"
|
||||
CFG="$HERMES/config.yaml"
|
||||
WARN_FILE="${HERMES_CONFIG_WARNINGS:-/tmp/hermes-config-warnings.txt}"
|
||||
LLAMA="${LLAMA_SWAP_URL:-http://127.0.0.1:8080}"
|
||||
VENV_PY="$HERMES/hermes-agent/venv/bin/python"
|
||||
export XDG_RUNTIME_DIR="${XDG_RUNTIME_DIR:-/run/user/$(id -u)}"
|
||||
|
||||
say(){ echo "[self-repair] $*"; }
|
||||
|
||||
# ── 1) Evidenz: Postcheck-Warnungen + frisches Gateway-Journal ────────────────
|
||||
EVID="$(cat "$WARN_FILE" 2>/dev/null)"
|
||||
JOURN="$(journalctl --user -u hermes-gateway --since '-6 minutes' --no-pager -o cat 2>/dev/null \
|
||||
| grep -iE "Unknown [a-z._]+ '|deprecated .*(setting|option|config)|defaulting to|invalid config" \
|
||||
| grep -v check_fn | head -10)"
|
||||
[ -z "$EVID" ] && EVID="$JOURN"
|
||||
if [ -z "$EVID" ]; then
|
||||
say "Keine Config-Bruch-Hinweise gefunden — hier ist (sicher) nichts zu reparieren."
|
||||
exit 2
|
||||
fi
|
||||
say "Config-Hinweise:"; echo "$EVID"
|
||||
|
||||
# LLM-Diagnose/Empfehlung (best-effort) — für die Meldung, egal ob wir reparieren.
|
||||
SUGG="$(curl -sf -m 60 "$LLAMA/v1/chat/completions" -H 'Content-Type: application/json' \
|
||||
-d "$(python3 - "$EVID" <<'PY'
|
||||
import json, sys
|
||||
p = ("Ein Hermes-Agent-Update hat den Selbsttest rot gemacht. Log-Hinweise:\n" + sys.argv[1] +
|
||||
"\n\nNenne auf Deutsch in 1-2 kurzen Sätzen die wahrscheinliche Ursache und die konkrete "
|
||||
"Config-Korrektur in ~/.hermes/config.yaml (Schlüssel → neuer Wert). Keine Einleitung.")
|
||||
print(json.dumps({"model": "coder-lite", "max_tokens": 200, "temperature": 0.1,
|
||||
"messages": [{"role": "user", "content": p}]}))
|
||||
PY
|
||||
)" 2>/dev/null | python3 -c 'import sys,json; print((json.load(sys.stdin).get("choices") or [{}])[0].get("message",{}).get("content","").strip())' 2>/dev/null)"
|
||||
|
||||
# ── 2) Offending Leaf-Keys ziehen ────────────────────────────────────────────
|
||||
KEYS="$(echo "$EVID" | grep -oE "Unknown [a-zA-Z0-9_.]+|'[a-zA-Z0-9_.]+'" | sed -E "s/^Unknown //; s/'//g" | sort -u)"
|
||||
LEAVES="$(for k in $KEYS; do echo "${k##*.}"; done | sort -u | grep -vE '^(true|false|null|)$')"
|
||||
|
||||
SEC_RX='token|secret|password|approval|api_server_key|ufw|sudo|allowlist|credential'
|
||||
BAK="$CFG.selfrepair-bak-$(date +%Y%m%d-%H%M%S)"
|
||||
cp -a "$CFG" "$BAK"
|
||||
|
||||
CHANGED=""; SKIPPED=""
|
||||
for leaf in $LEAVES; do
|
||||
if echo "$leaf" | grep -qiE "$SEC_RX"; then SKIPPED="$SKIPPED $leaf(security-tabu)"; continue; fi
|
||||
# Nur EINDEUTIGE Schlüssel-Zeilen (genau 1×) anfassen — Mehrdeutige sind zu riskant.
|
||||
n="$(grep -cE "^[[:space:]]*${leaf}:" "$CFG" 2>/dev/null || echo 0)"
|
||||
if [ "$n" = "1" ]; then
|
||||
sed -i -E "s|^([[:space:]]*)(${leaf}:.*)$|\1# [selbstrepariert $(date +%F)] \2|" "$CFG"
|
||||
CHANGED="$CHANGED $leaf"; say "auskommentiert: $leaf (unbekannt/veraltet)"
|
||||
else
|
||||
SKIPPED="$SKIPPED $leaf(${n}x)"
|
||||
fi
|
||||
done
|
||||
[ -n "$SKIPPED" ] && say "nicht automatisch angefasst:$SKIPPED"
|
||||
|
||||
if [ -z "$CHANGED" ]; then
|
||||
say "Nichts sicher automatisch reparierbar."
|
||||
cp -a "$BAK" "$CFG"
|
||||
echo "SELFREPAIR_SUGGESTION=$SUGG"
|
||||
exit 2
|
||||
fi
|
||||
|
||||
# ── 3) YAML noch gültig? ──────────────────────────────────────────────────────
|
||||
if [ -x "$VENV_PY" ] && ! "$VENV_PY" -c "import yaml; yaml.safe_load(open('$CFG'))" 2>/dev/null; then
|
||||
say "YAML nach Fix ungültig → zurück zum Backup."
|
||||
cp -a "$BAK" "$CFG"; echo "SELFREPAIR_SUGGESTION=$SUGG"; exit 2
|
||||
fi
|
||||
|
||||
# ── 4) Neustart + Gehirn-Check ────────────────────────────────────────────────
|
||||
systemctl --user restart hermes-gateway; sleep 6
|
||||
if bash "$SRC_DIR/hermes-postcheck.sh" >/dev/null 2>&1; then
|
||||
say "Gehirn-Check nach Selbstreparatur GRÜN."
|
||||
bash "$SRC_DIR/notify.sh" -s "[Box-Update]" \
|
||||
"Commander, nach dem Hermes-Update hakte der Selbsttest wegen geänderter Einstellungen — ich habe das selbst gerichtet (angepasst:${CHANGED// /, }). Der Test ist wieder grün, die neue Version läuft. Ein Rollback war nicht nötig." || true
|
||||
exit 0
|
||||
fi
|
||||
|
||||
say "Auch nach Selbstreparatur rot → zurück zum Backup (der Aufrufer rollt zurück)."
|
||||
cp -a "$BAK" "$CFG"
|
||||
systemctl --user restart hermes-gateway; sleep 6
|
||||
echo "SELFREPAIR_SUGGESTION=$SUGG"
|
||||
exit 2
|
||||
@@ -0,0 +1,86 @@
|
||||
#!/usr/bin/env bash
|
||||
# Tägliche Selbst-Smoke-Tests der Box (Autonomie 0d): spielt die drei Kern-Pfade AKTIV durch —
|
||||
# auch OHNE Update — und meldet NUR bei Rot (kein tägliches Grün-Rauschen). Ergänzt den passiven
|
||||
# Health-Wächter (sentry.py prüft nur Erreichbarkeit) um echte Funktion:
|
||||
# 1) Gateway/Hirn: der Agent liefert überhaupt eine Antwort.
|
||||
# 2) Tools: ein harmloser terminal-Befehl durch den echten Agenten (fängt kaputte Approval-/
|
||||
# Tool-Ketten — Antwort muss kommen und darf nicht in pending_approval hängen).
|
||||
# 3) Voice: Lucys Sprech-Pfad (/api/voice/chat) streamt.
|
||||
# Bei Fehlschlag → notify.sh (Telegram + Briefkasten) in Lucys Stimme. Läuft als User, kein sudo.
|
||||
set -uo pipefail
|
||||
|
||||
SRC_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
MC_URL="${MC_URL:-http://127.0.0.1:9001}"
|
||||
HERMES_URL="${HERMES_API_URL:-http://127.0.0.1:8642}"
|
||||
HERMES="${HERMES_HOME:-$HOME/.hermes}"
|
||||
export XDG_RUNTIME_DIR="${XDG_RUNTIME_DIR:-/run/user/$(id -u)}"
|
||||
|
||||
say(){ echo "[self-smoke] $*"; }
|
||||
FAILED=()
|
||||
|
||||
API_KEY="$(grep -E '^API_SERVER_KEY=' "$HERMES/.env" 2>/dev/null | cut -d= -f2- | tr -d '"' | tr -d "'")"
|
||||
|
||||
# ── 1) Gateway/Hirn: Agent antwortet ────────────────────────────────────────
|
||||
if [ -n "$API_KEY" ]; then
|
||||
RESP=$(curl -sf -m 90 -X POST "$HERMES_URL/v1/chat/completions" \
|
||||
-H "Authorization: Bearer $API_KEY" -H "Content-Type: application/json" \
|
||||
-H "X-Hermes-Session-Id: selfsmoke-gateway" \
|
||||
-d '{"model":"hermes","stream":false,"chat_template_kwargs":{"enable_thinking":false},"messages":[{"role":"user","content":"Antworte NUR mit dem einen Wort: bereit"}]}' 2>/dev/null)
|
||||
if echo "$RESP" | grep -qi "bereit"; then
|
||||
say "PASS · Gateway/Hirn antwortet"
|
||||
else
|
||||
say "FAIL · Gateway/Hirn ohne verwertbare Antwort: $(echo "$RESP" | head -c 160)"
|
||||
FAILED+=("Gehirn/Gateway (keine Antwort)")
|
||||
fi
|
||||
else
|
||||
say "SKIP · Gateway (kein API_SERVER_KEY in $HERMES/.env)"
|
||||
fi
|
||||
|
||||
# ── 2) Tools: terminal via Agent ─────────────────────────────────────────────
|
||||
if [ -n "$API_KEY" ]; then
|
||||
TOOL=$(curl -sf -m 90 -X POST "$HERMES_URL/v1/chat/completions" \
|
||||
-H "Authorization: Bearer $API_KEY" -H "Content-Type: application/json" \
|
||||
-H "X-Hermes-Session-Id: selfsmoke-tool" \
|
||||
-d '{"model":"hermes","stream":false,"messages":[{"role":"user","content":"Führe im Terminal exakt den Befehl echo selfsmoke-ok aus und gib mir nur dessen Ausgabe zurück."}]}' 2>/dev/null)
|
||||
if echo "$TOOL" | grep -qi "selfsmoke-ok"; then
|
||||
say "PASS · Tools (terminal via Agent)"
|
||||
elif echo "$TOOL" | grep -qi "pending_approval"; then
|
||||
say "FAIL · Tools hängen in pending_approval (approvals.mode prüfen!)"
|
||||
FAILED+=("Tools (hängen in Freigabe)")
|
||||
else
|
||||
say "FAIL · Tool-Smoke ohne Ergebnis: $(echo "$TOOL" | head -c 160)"
|
||||
FAILED+=("Tools (kein Ergebnis)")
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── 3) Voice: Lucys Sprech-Stream ────────────────────────────────────────────
|
||||
# Retry: der erste Turn nach Idle zahlt evtl. den Session-Kaltstart (Prefill).
|
||||
# Subshell OHNE pipefail — head -c schließt die Pipe früh (curl endet mit 23/SIGPIPE),
|
||||
# das würde einen bestandenen Check sonst als FAIL werten (Lehre aus hermes-postcheck.sh).
|
||||
voice_ok=0
|
||||
for try in 1 2 3; do
|
||||
if ( set +o pipefail; curl -sf -m 45 -N -X POST "$MC_URL/api/voice/chat" -H "Content-Type: application/json" \
|
||||
-d '{"text":"Sag nur ok.","session_id":"selfsmoke-voice"}' 2>/dev/null | head -c 50 | grep -q "data:" ); then
|
||||
voice_ok=1; break
|
||||
fi
|
||||
sleep 8
|
||||
done
|
||||
if [ "$voice_ok" -eq 1 ]; then
|
||||
say "PASS · Voice (/api/voice/chat streamt, Versuch $try)"
|
||||
else
|
||||
say "FAIL · Voice: /api/voice/chat streamt nicht (3 Versuche)"
|
||||
FAILED+=("Voice (Lucys Sprech-Pfad)")
|
||||
fi
|
||||
|
||||
# ── Meldung nur bei Rot ──────────────────────────────────────────────────────
|
||||
if [ "${#FAILED[@]}" -eq 0 ]; then
|
||||
say "Alle Selbst-Tests grün ($(date '+%F %H:%M'))."
|
||||
exit 0
|
||||
fi
|
||||
|
||||
MSG="Commander, mein täglicher Selbst-Test hat etwas gefunden — das läuft gerade NICHT rund:
|
||||
$(printf '• %s\n' "${FAILED[@]}")
|
||||
Der Rest der Box kann trotzdem stehen; magst du kurz draufschauen?"
|
||||
bash "$SRC_DIR/notify.sh" -s "[Selbst-Test]" "$MSG" || true
|
||||
say "FEHLGESCHLAGEN: ${FAILED[*]}"
|
||||
exit 1
|
||||
@@ -0,0 +1,79 @@
|
||||
---
|
||||
name: wartung
|
||||
description: "Werkstatt-Kreislauf der AI-Box: einen KLEINEN Wartungsauftrag (Config-Migration, Dependency-Bump, Ein-/Zwei-Datei-Patch) eigenständig umsetzen — Branch im Box-Checkout, Patch, separater Reviewer-Subagent, Gate, Telegram-Merge-Vorschlag. Merge/Deploy NIE selbst."
|
||||
version: 1.0.0
|
||||
author: MC2 (Autonomie E6)
|
||||
platforms: [linux]
|
||||
metadata:
|
||||
hermes:
|
||||
tags: [wartung, werkstatt, maintenance, mc2]
|
||||
---
|
||||
|
||||
# Werkstatt — Selbstwartungs-Kreislauf der Box
|
||||
|
||||
Nutze diesen Skill, wenn ein kleiner, klar umrissener Wartungsauftrag für das MC2-Repo
|
||||
(`~/mission-control-v2`) vorliegt — vom Evolution-Radar oder direkt vom User.
|
||||
|
||||
**Scope-Check zuerst:** Klein = Config-Schlüssel-Migration, Dependency-Bump (Lockfile),
|
||||
Patch in 1–2 Dateien. Alles Größere (Architektur, mehrere Module, neue Features):
|
||||
NUR einen Plan liefern (Text im Telegram-Vorschlag), KEINEN Code.
|
||||
|
||||
## Leitplanken (nicht verhandelbar)
|
||||
|
||||
- NIEMALS auf `main` committen. NIEMALS mergen. NIEMALS deployen oder Dienste neu starten.
|
||||
- Security-Config ist TABU: keine Tokens, approvals, ufw, sudoers anfassen.
|
||||
- Die Live-Instanz (`~/mission-control-v2`) bleibt unberührt — gearbeitet wird NUR im Worktree.
|
||||
- Am Ende steht IMMER ein Telegram-Vorschlag; die Entscheidung trifft der User.
|
||||
- Gate rot oder Reviewer dagegen → trotzdem ehrlich melden (Branch bleibt liegen), nichts beschönigen.
|
||||
|
||||
## Ablauf
|
||||
|
||||
1. **Worktree anlegen** (Slug = kurzer Kebab-Case-Name des Auftrags):
|
||||
`cd ~/mission-control-v2 && git fetch origin && git worktree add /tmp/wartung-<slug> -b wartung/<slug> origin/main`
|
||||
2. **Patch** nur im Worktree. Minimal-invasiv, Stil der umgebenden Datei übernehmen
|
||||
(deutsche Kommentare, bestehende Muster).
|
||||
3. **Selbst-Gate** (was zutrifft):
|
||||
- Python geändert → `python3 -m py_compile <dateien>`
|
||||
- Shell geändert → `bash -n <dateien>`
|
||||
- Frontend (`frontend/src/...`) geändert → auf der Box gibt es KEIN Node. Im Vorschlag
|
||||
ausweisen: „Gate eingeschränkt: tsc/Build läuft erst beim Merge auf dem PC."
|
||||
- **Live-Checkout unberührt (PFLICHT, zwei Beweise):**
|
||||
- `git -C ~/mission-control-v2 status --porcelain -uno` **muss LEER sein** — kein einziges
|
||||
geändertes/staged Tracked-File im Live-Checkout. (Nur so ist bewiesen, dass du wirklich
|
||||
ausschließlich im Worktree gearbeitet hast; ein grüner Health-curl allein reicht NICHT,
|
||||
weil der laufende Dienst den alten Code im Speicher hält und eine schmutzige Datei auf
|
||||
der Platte nicht bemerkt — genau diese Lücke ist am 03.07. aufgefallen.)
|
||||
- `curl -sf http://127.0.0.1:9001/api/health` muss grün bleiben.
|
||||
- Ist der `git status` NICHT leer: du hast die Leitplanke verletzt → die fremden Änderungen
|
||||
im Live-Checkout mit `git -C ~/mission-control-v2 checkout -- <datei>` zurücknehmen (deine
|
||||
Arbeit liegt ja sicher im Worktree/Branch) und im Vorschlag ehrlich erwähnen.
|
||||
4. **Reviewer-Subagent** (frischer Kontext, Worker/Reviewer-Muster): `delegate_task` mit
|
||||
role=leaf. Gib ihm den AUFTRAG im Wortlaut + `git diff` des Worktrees.
|
||||
**Kernauftrag an den Reviewer (nicht „sieht plausibel aus", sondern BEWEISEN):**
|
||||
- **Stelle den konkreten Fehlerfall aus dem Auftrag nach.** Nimm eine realistische
|
||||
Beispiel-Eingabe, die den beschriebenen Bug AUSLÖST, und spiele den geänderten Code
|
||||
Schritt für Schritt (oder als kleiner Test) durch: Behebt der Diff DIESEN Fall wirklich?
|
||||
(Am 03.07. hat ein Reviewer einen Patch durchgewunken, der den eigentlichen Fehlerfall
|
||||
gar nicht traf — genau das darf nicht mehr passieren.)
|
||||
- **Suche aktiv nach dem Fall, in dem der Patch versagt** (Randfälle, leere/kurze Eingaben,
|
||||
andere Formate). Findest du einen → Kritik.
|
||||
- Erst danach die Standardfragen: Minimal-invasiv? Nur der beauftragte Scope? Nebenwirkungen?
|
||||
Bei berechtigter Kritik: nachbessern (max. 2 Runden), sonst Kritik in den Vorschlag schreiben.
|
||||
Schreibe das Reviewer-Urteil im Telegram-Vorschlag KONKRET aus („durchgespielt mit Eingabe X →
|
||||
Ergebnis Y"), nicht nur „Reviewer: ok".
|
||||
5. **Commit im Worktree** (Message `Werkstatt: <Auftrag kurz>` + 2–4 Zeilen Was/Warum),
|
||||
dann **Branch zu Gitea pushen:** `git push origin wartung/<slug>` (die Box hat einen
|
||||
eigenen scoped Token, seit 02.07. hinterlegt). Push fehlgeschlagen? Nicht schlimm —
|
||||
Branch bleibt lokal, im Vorschlag erwähnen.
|
||||
6. **Telegram-Vorschlag** über `bash ~/mission-control-v2/deploy/notify.sh -s "[Werkstatt]" "<text>"`:
|
||||
In LUCYS Stimme an den Commander (nicht als anonymer Job): erst 1–2 Sätze, was gemacht wurde
|
||||
und was er jetzt entscheiden soll — dann knapp: geänderte Dateien · Kern des Diffs ·
|
||||
Gate-Ergebnis · Reviewer-Urteil · Branch-Name · Frage „merge oder verwerfen?"
|
||||
7. **Nichts löschen:** Worktree + Branch bleiben liegen, bis der User entschieden hat.
|
||||
|
||||
## Nach dem User-Entscheid (kommt als neuer Auftrag)
|
||||
|
||||
- „verwerfen" → `git worktree remove /tmp/wartung-<slug> --force && git branch -D wartung/<slug>`
|
||||
(+ Remote-Branch löschen, falls gepusht: `git push origin --delete wartung/<slug>`)
|
||||
- „merge" → das Mergen nach main + Deploy macht der PC/Claude (Frontend-Builds gibt es
|
||||
nur dort). Du pushst NIE nach main — auch nicht mit Token.
|
||||
@@ -0,0 +1,80 @@
|
||||
#!/usr/bin/env bash
|
||||
# Stack-Post-Update-Check: läuft NACH jedem Update (OS / Engine / Router) und verifiziert,
|
||||
# dass der komplette Inferenz-Stack noch FUNKTIONIERT — nicht nur "Befehl lief durch".
|
||||
# Exit 0 = alles ok, sonst 1 → der jobengine-Job wird im UI ROT (state=failed).
|
||||
#
|
||||
# Pendant zu hermes-postcheck.sh (prüft das Gehirn/Mem0); dieser prüft Router+Engine+MC2
|
||||
# inkl. einer ECHTEN 1-Token-Inferenz (beweist, dass ein Modell wirklich lädt & generiert).
|
||||
set -uo pipefail
|
||||
|
||||
SWAP_URL="${MC_LLAMA_SWAP_URL:-http://127.0.0.1:8080}"
|
||||
MC_URL="${MC_URL:-http://127.0.0.1:9001}"
|
||||
MEM0_URL="${MEM0_SERVICE_URL:-http://127.0.0.1:8765}"
|
||||
BRAIN="${MC_WARMUP_MODELS:-fast}"; BRAIN="${BRAIN%% *}" # erstes Modell, falls Liste
|
||||
EMBED="${MC_WARMUP_EMBED:-embed}"; EMBED="${EMBED%% *}" # Embedding-Modell (Mem0/Gedächtnis)
|
||||
fail=0
|
||||
|
||||
echo "=== Stack Post-Update: Funktionsprüfung ==="
|
||||
|
||||
# 0. Auf llama-swap warten — ein Engine/Router-Update startet den Dienst neu, der Erststart
|
||||
# (+ erstes Modell-Laden) kann dauern. Max ~120s.
|
||||
swap_up=0
|
||||
for _ in $(seq 1 60); do
|
||||
curl -sf -m 5 "$SWAP_URL/v1/models" >/dev/null 2>&1 && { swap_up=1; break; }
|
||||
sleep 2
|
||||
done
|
||||
|
||||
# 1. Router-Dienst aktiv? (is-active ist read-only → kein sudo nötig)
|
||||
if systemctl is-active --quiet llama-swap; then
|
||||
echo "PASS · llama-swap-Dienst aktiv"
|
||||
else
|
||||
echo "FAIL · llama-swap-Dienst NICHT aktiv"; fail=1
|
||||
fi
|
||||
|
||||
# 2. Engine erreichbar (Modell-Liste)?
|
||||
if [ "$swap_up" -eq 1 ]; then
|
||||
echo "PASS · Engine /v1/models antwortet"
|
||||
else
|
||||
echo "FAIL · Engine /v1/models antwortet nicht (nach 120s)"; fail=1
|
||||
fi
|
||||
|
||||
# 3. Echte Inferenz: lädt ein Modell und generiert es ein Token?
|
||||
resp="$(curl -s -m 240 -X POST "$SWAP_URL/v1/chat/completions" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d "{\"model\":\"$BRAIN\",\"max_tokens\":1,\"messages\":[{\"role\":\"user\",\"content\":\"ping\"}]}" 2>/dev/null)"
|
||||
if printf '%s' "$resp" | grep -q '"choices"'; then
|
||||
echo "PASS · Inferenz auf '$BRAIN' liefert eine Antwort"
|
||||
else
|
||||
echo "FAIL · Inferenz auf '$BRAIN' fehlgeschlagen (Modell lädt/generiert nicht)"; fail=1
|
||||
fi
|
||||
|
||||
# 3b. Embedding-Modell (für Mem0/Gedächtnis) lädt und liefert einen Vektor?
|
||||
eresp="$(curl -s -m 120 -X POST "$SWAP_URL/v1/embeddings" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d "{\"model\":\"$EMBED\",\"input\":\"ping\"}" 2>/dev/null)"
|
||||
if printf '%s' "$eresp" | grep -q '"embedding"'; then
|
||||
echo "PASS · Embedding-Modell '$EMBED' liefert Vektoren"
|
||||
else
|
||||
echo "FAIL · Embedding-Modell '$EMBED' lädt/antwortet nicht (Mem0/Gedächtnis betroffen)"; fail=1
|
||||
fi
|
||||
|
||||
# 4. MC2 selbst gesund (Engine + Gateway erreichbar)?
|
||||
if curl -sf -m 8 "$MC_URL/api/health" 2>/dev/null | grep -qE '"engine_reachable":[[:space:]]*true'; then
|
||||
echo "PASS · MC2 /api/health: engine_reachable=true"
|
||||
else
|
||||
echo "FAIL · MC2 /api/health meldet Engine nicht erreichbar"; fail=1
|
||||
fi
|
||||
|
||||
# 5. Mem0-Sidecar (Gedächtnis) erreichbar?
|
||||
if curl -sf -m 5 "$MEM0_URL/health" >/dev/null 2>&1; then
|
||||
echo "PASS · Mem0-Sidecar erreichbar"
|
||||
else
|
||||
echo "FAIL · Mem0-Sidecar NICHT erreichbar"; fail=1
|
||||
fi
|
||||
|
||||
if [ "$fail" -eq 0 ]; then
|
||||
echo "=== STACK OK ✓ ==="
|
||||
else
|
||||
echo "=== STACK-CHECK FEHLGESCHLAGEN — bitte prüfen ==="
|
||||
fi
|
||||
exit "$fail"
|
||||
@@ -0,0 +1,13 @@
|
||||
# MC2-Autonomie: erlaubt dem Auto-Update-Timer (User hitonabi) die beiden
|
||||
# root-Update-Skripte OHNE Passwort. Einmalig als root installieren:
|
||||
#
|
||||
# sudo install -m 0440 -o root -g root ~/mission-control-v2/deploy/sudoers-mc2-autonomie /etc/sudoers.d/mc2-autonomie
|
||||
# sudo visudo -c
|
||||
#
|
||||
# In derselben einmaligen sudo-Session sinnvoll (siehe AUTONOMIE_PLAN E2 + Faden C14):
|
||||
# sudo apt-get install -y unattended-upgrades && sudo dpkg-reconfigure -plow unattended-upgrades
|
||||
# sudo systemctl disable --now mission-control.service # stale v1-Unit
|
||||
#
|
||||
# Hinweis: die Skripte liegen im User-Checkout (User-schreibbar) — bewusste
|
||||
# Entscheidung für die Single-User-Appliance; die Regel gilt NUR für exakt diese Pfade.
|
||||
hitonabi ALL=(root) NOPASSWD: /usr/bin/bash /home/hitonabi/mission-control-v2/deploy/update-engine.sh, /usr/bin/bash /home/hitonabi/mission-control-v2/deploy/update-swap.sh
|
||||
@@ -0,0 +1,78 @@
|
||||
#!/usr/bin/env bash
|
||||
# Aktualisiert den Vulkan-llama.cpp-Build (ggml-org) auf den neuesten Release.
|
||||
# Wird vom MC2-„Engine Update"-Button via `sudo bash …` als ROOT aufgerufen (kein internes sudo).
|
||||
#
|
||||
# Wichtig: llama-swap wird VOR dem Datei-Austausch gestoppt — sonst ist die laufende llama-server-
|
||||
# Binary „Text file busy" und die mmap'ten .so-Libs dürfen nicht unter dem laufenden Prozess
|
||||
# getauscht werden. Danach Stack-Check (echte Inferenz); bei Fehler Auto-Rollback.
|
||||
#
|
||||
# Exit 0 = neuer Build verifiziert · 1 = fehlgeschlagen, Rollback ok (alter Build läuft)
|
||||
# · 2 = Update UND Rollback fehlgeschlagen (Stack evtl. kaputt — bitte prüfen)
|
||||
set -uo pipefail # bewusst KEIN -e: bei Fehlern kontrolliert zurückrollen statt hart abbrechen
|
||||
|
||||
VULKAN_DIR=/opt/llamacpp-vulkan
|
||||
BAK_DIR=/opt/llamacpp-vulkan.bak # „letzter funktionierender Build" für Auto-Rollback
|
||||
SRC_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
PIN_BUILD="${MC_ENGINE_BUILD:-}" # leer = neuester Release
|
||||
|
||||
die(){ echo "$1" >&2; exit 1; }
|
||||
start_swap(){ systemctl start llama-swap || systemctl restart llama-swap; }
|
||||
|
||||
# Inhalte von $1 nach $VULKAN_DIR spiegeln (Dir vorher leeren → sauberer Austausch).
|
||||
replace_dir(){
|
||||
rm -rf "${VULKAN_DIR:?}" && mkdir -p "$VULKAN_DIR" || return 1
|
||||
cp -a "$1/." "$VULKAN_DIR"/ || return 1
|
||||
ln -sfn "$VULKAN_DIR/llama-server" /usr/local/bin/llama-server
|
||||
[ -x "$VULKAN_DIR/llama-server" ]
|
||||
}
|
||||
|
||||
# 1. Release ermitteln + laden (vor jedem Eingriff am laufenden System)
|
||||
if [ -n "$PIN_BUILD" ]; then
|
||||
TAG="$PIN_BUILD"
|
||||
else
|
||||
TAG="$(curl -s https://api.github.com/repos/ggml-org/llama.cpp/releases/latest | jq -r .tag_name)"
|
||||
fi
|
||||
[ -n "$TAG" ] && [ "$TAG" != "null" ] || die "Konnte neuesten Release-Tag nicht ermitteln"
|
||||
|
||||
ASSET="llama-${TAG}-bin-ubuntu-vulkan-x64.tar.gz"
|
||||
URL="https://github.com/ggml-org/llama.cpp/releases/download/${TAG}/${ASSET}"
|
||||
TMP="$(mktemp -d)"
|
||||
trap 'rm -rf "$TMP"' EXIT
|
||||
|
||||
echo "Lade Engine ${TAG} …"
|
||||
curl -fsSL -o "$TMP/v.tgz" "$URL" || die "Download fehlgeschlagen"
|
||||
tar xzf "$TMP/v.tgz" -C "$TMP" || die "Entpacken fehlgeschlagen"
|
||||
NEWDIR="$(echo "$TMP"/llama-*/)" # Tarball entpackt nach llama-<tag>/
|
||||
[ -x "${NEWDIR}llama-server" ] || die "llama-server im Archiv nicht gefunden (${NEWDIR})"
|
||||
|
||||
# 2. Aktuellen (funktionierenden) Build sichern → Auto-Rollback
|
||||
echo "Sichere aktuellen Build → ${BAK_DIR}"
|
||||
rm -rf "$BAK_DIR"
|
||||
cp -a "$VULKAN_DIR" "$BAK_DIR" || die "Backup fehlgeschlagen"
|
||||
|
||||
# 3. llama-swap stoppen (gibt Binary/Libs frei), austauschen, wieder starten
|
||||
echo "Stoppe llama-swap für den Austausch…"
|
||||
systemctl stop llama-swap
|
||||
if replace_dir "${NEWDIR%/}"; then
|
||||
start_swap
|
||||
echo "Engine auf ${TAG} aktualisiert — verifiziere Stack…"
|
||||
if bash "$SRC_DIR/stack-postcheck.sh"; then
|
||||
echo "Engine-Update ${TAG} erfolgreich verifiziert."
|
||||
exit 0
|
||||
fi
|
||||
echo "!! Stack-Check fehlgeschlagen — ROLLBACK auf vorherigen Build."
|
||||
else
|
||||
echo "!! Datei-Austausch fehlgeschlagen — ROLLBACK auf vorherigen Build."
|
||||
fi
|
||||
|
||||
# 4. Auto-Rollback auf den gesicherten Build
|
||||
systemctl stop llama-swap
|
||||
if replace_dir "$BAK_DIR"; then
|
||||
start_swap
|
||||
if bash "$SRC_DIR/stack-postcheck.sh"; then
|
||||
echo "Rollback erfolgreich: vorheriger Build läuft wieder. (Update ${TAG} NICHT angewendet.)"
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
echo "!!! Rollback fehlgeschlagen — Stack möglicherweise kaputt! Backup liegt unter ${BAK_DIR}."
|
||||
exit 2
|
||||
Executable
+68
@@ -0,0 +1,68 @@
|
||||
#!/usr/bin/env bash
|
||||
# Aktualisiert den llama-swap-Router (mostlygeek/llama-swap) auf den neuesten Release.
|
||||
# Wird vom MC2-„Router Update"-Button via `sudo bash …` als ROOT aufgerufen (kein internes sudo).
|
||||
#
|
||||
# Wichtig: llama-swap wird VOR dem Binary-Tausch gestoppt — sonst ist die laufende Binary
|
||||
# „Text file busy". Danach Stack-Check (echte Inferenz); bei Fehler Auto-Rollback.
|
||||
#
|
||||
# Exit 0 = neue Version verifiziert · 1 = fehlgeschlagen, Rollback ok (alte Version läuft)
|
||||
# · 2 = Update UND Rollback fehlgeschlagen (Stack evtl. kaputt — bitte prüfen)
|
||||
set -uo pipefail # bewusst KEIN -e: bei Fehlern kontrolliert zurückrollen statt hart abbrechen
|
||||
|
||||
SWAP_BIN="${MC_SWAP_BIN:-/usr/local/bin/llama-swap}"
|
||||
BAK="${SWAP_BIN}.bak" # „letzte funktionierende Version" für Auto-Rollback
|
||||
SRC_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
PIN_VER="${MC_SWAP_VERSION:-}" # leer = neuester Release
|
||||
|
||||
die(){ echo "$1" >&2; exit 1; }
|
||||
start_swap(){ systemctl start llama-swap || systemctl restart llama-swap; }
|
||||
|
||||
if [ -n "$PIN_VER" ]; then
|
||||
TAG="$PIN_VER"
|
||||
else
|
||||
TAG="$(curl -s https://api.github.com/repos/mostlygeek/llama-swap/releases/latest | jq -r .tag_name)"
|
||||
fi
|
||||
[ -n "$TAG" ] && [ "$TAG" != "null" ] || die "Konnte neuesten Release-Tag nicht ermitteln"
|
||||
|
||||
# Asset-Name nutzt die nackte Nummer (z.B. v230 -> 230): llama-swap_230_linux_amd64.tar.gz
|
||||
NUM="${TAG#v}"
|
||||
ASSET="llama-swap_${NUM}_linux_amd64.tar.gz"
|
||||
URL="https://github.com/mostlygeek/llama-swap/releases/download/${TAG}/${ASSET}"
|
||||
TMP="$(mktemp -d)"
|
||||
trap 'rm -rf "$TMP"' EXIT
|
||||
|
||||
echo "Lade llama-swap ${TAG} …"
|
||||
curl -fsSL -o "$TMP/s.tgz" "$URL" || die "Download fehlgeschlagen"
|
||||
tar xzf "$TMP/s.tgz" -C "$TMP" || die "Entpacken fehlgeschlagen"
|
||||
NEW="$(find "$TMP" -type f -name llama-swap | head -1)"
|
||||
[ -n "$NEW" ] || die "llama-swap-Binary im Archiv nicht gefunden"
|
||||
|
||||
# Sicherung der aktuellen (funktionierenden) Binary → Auto-Rollback
|
||||
[ -x "$SWAP_BIN" ] && { cp -a "$SWAP_BIN" "$BAK" || die "Backup fehlgeschlagen"; }
|
||||
|
||||
# llama-swap stoppen (gibt die laufende Binary frei), tauschen, wieder starten
|
||||
echo "Stoppe llama-swap für den Binary-Tausch…"
|
||||
systemctl stop llama-swap
|
||||
if install -m 0755 "$NEW" "$SWAP_BIN" && [ -x "$SWAP_BIN" ]; then
|
||||
start_swap
|
||||
echo "Router auf ${TAG} aktualisiert — verifiziere Stack…"
|
||||
if bash "$SRC_DIR/stack-postcheck.sh"; then
|
||||
echo "Router-Update ${TAG} erfolgreich verifiziert."
|
||||
exit 0
|
||||
fi
|
||||
echo "!! Stack-Check fehlgeschlagen — ROLLBACK auf vorherige Version."
|
||||
else
|
||||
echo "!! Binary-Tausch fehlgeschlagen — ROLLBACK auf vorherige Version."
|
||||
fi
|
||||
|
||||
# Auto-Rollback auf die gesicherte Binary
|
||||
systemctl stop llama-swap
|
||||
if [ -x "$BAK" ] && install -m 0755 "$BAK" "$SWAP_BIN"; then
|
||||
start_swap
|
||||
if bash "$SRC_DIR/stack-postcheck.sh"; then
|
||||
echo "Rollback erfolgreich: vorherige llama-swap-Version läuft wieder. (Update ${TAG} NICHT angewendet.)"
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
echo "!!! Rollback fehlgeschlagen — Stack möglicherweise kaputt! Backup liegt unter ${BAK}."
|
||||
exit 2
|
||||
@@ -0,0 +1,23 @@
|
||||
[Unit]
|
||||
Description=MC2 Voice Sidecar — lokales STT (faster-whisper) + gestuftes TTS (Piper/Chatterbox)
|
||||
After=network.target
|
||||
|
||||
[Service]
|
||||
# Eigenes Python-3.12-venv (~/.voice/venv) — torch/chatterbox/faster-whisper passen nicht ins
|
||||
# 3.14-Backend-venv (analog mem0-service). Bind 127.0.0.1: nur lokal; MC2 proxyt nach außen.
|
||||
Type=simple
|
||||
WorkingDirectory=%h/mission-control-v2/voice_service
|
||||
Environment=VOICE_PORT=8650
|
||||
Environment=VOICE_STT_MODEL=medium
|
||||
Environment=VOICE_STT_LANG=de
|
||||
Environment=VOICE_PIPER_DIR=%h/.voice/voices
|
||||
Environment=VOICE_PIPER_DEFAULT=de_DE-thorsten-medium
|
||||
Environment=VOICE_CHATTERBOX_DEVICE=cpu
|
||||
Environment=VOICE_CHATTERBOX_LANG=de
|
||||
Environment=TOKENIZERS_PARALLELISM=false
|
||||
ExecStart=%h/.voice/venv/bin/python -m uvicorn app:app --host 127.0.0.1 --port 8650
|
||||
Restart=always
|
||||
RestartSec=3
|
||||
|
||||
[Install]
|
||||
WantedBy=default.target
|
||||
@@ -0,0 +1,57 @@
|
||||
#!/usr/bin/env bash
|
||||
# Lädt das warme "brains"-Set nach einem llama-swap-(Re)Start vor, damit die ERSTE
|
||||
# Anfrage nicht kalt ist (llama-swap lädt sonst lazy bei Bedarf).
|
||||
# Eingehängt als ExecStartPost=-/usr/local/bin/llama-swap-warmup.sh in llama-swap.
|
||||
#
|
||||
# `fast` = das Agent-Hirn (Qwen3.6-35B-A3B), `vision` = die Augen (Qwen3-VL-8B, ttl 0 seit
|
||||
# 03.07. — gehört zum UI-Versprechen „Immer-bereit-Set"), `embed` = das Embedding-Modell
|
||||
# (Qwen3-Embedding-0.6B, für Mem0/Gedächtnis). Alle in der brains-Gruppe (persistent), werden
|
||||
# aber NICHT automatisch zusammen geladen — jedes Modell einzeln anstoßen, sonst bleibt es kalt.
|
||||
# Embedding läuft im --embedding-Modus → anderer Endpunkt (/v1/embeddings, nicht chat).
|
||||
# Selbst-detachend: blockiert den Service-Start nicht.
|
||||
# NACH ÄNDERUNG: sudo cp deploy/warmup.sh /usr/local/bin/llama-swap-warmup.sh (root-Kopie!)
|
||||
set -u
|
||||
URL="${MC_LLAMA_SWAP_URL:-http://127.0.0.1:8080}"
|
||||
BRAINS="${MC_WARMUP_MODELS:-fast vision}"
|
||||
EMBEDS="${MC_WARMUP_EMBED:-embed}"
|
||||
|
||||
if [ "${1:-}" != "--inner" ]; then
|
||||
setsid "$0" --inner >/dev/null 2>&1 &
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# --- ab hier im entkoppelten Hintergrundprozess ---
|
||||
for _ in $(seq 1 60); do # auf llama-swap warten (max ~120s)
|
||||
curl -sf "$URL/v1/models" >/dev/null 2>&1 && break
|
||||
sleep 2
|
||||
done
|
||||
|
||||
for m in $BRAINS; do # Chat-Hirne über /v1/chat/completions
|
||||
curl -s -m 180 -X POST "$URL/v1/chat/completions" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d "{\"model\":\"$m\",\"max_tokens\":1,\"messages\":[{\"role\":\"user\",\"content\":\"ping\"}]}" \
|
||||
>/dev/null 2>&1 || true
|
||||
done
|
||||
|
||||
for m in $EMBEDS; do # Embedding-Modelle über /v1/embeddings
|
||||
curl -s -m 120 -X POST "$URL/v1/embeddings" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d "{\"model\":\"$m\",\"input\":\"ping\"}" \
|
||||
>/dev/null 2>&1 || true
|
||||
done
|
||||
|
||||
# Agenten-Prompt vorkauen: Hermes' Systemprompt (~70 KB, 33 Tool-Schemas) kostet nach
|
||||
# jedem Modell-(Neu-)Laden sonst 30-40 s Prefill BEIM ERSTEN USER-CALL (gemessen 03.07.:
|
||||
# kalt 37 s, mit gefülltem Prompt-Cache ~6 s). Ein Wegwerf-Turn füllt den -cram-Cache.
|
||||
AGENT_KEY="$(grep -E '^API_SERVER_KEY=' "$HOME/.hermes/.env" 2>/dev/null | cut -d= -f2- | tr -d '"' | tr -d "'")"
|
||||
if [ -n "$AGENT_KEY" ]; then
|
||||
for _ in $(seq 1 45); do # auf hermes-gateway warten (Boot-Reihenfolge)
|
||||
curl -sf -m 2 "http://127.0.0.1:8642/health" >/dev/null 2>&1 && break
|
||||
sleep 2
|
||||
done
|
||||
curl -s -m 120 -X POST "http://127.0.0.1:8642/v1/chat/completions" \
|
||||
-H "Authorization: Bearer $AGENT_KEY" -H 'Content-Type: application/json' \
|
||||
-H 'X-Hermes-Session-Id: warmup-prefill' \
|
||||
-d '{"model":"hermes","stream":false,"messages":[{"role":"user","content":"Sag nur OK."}]}' \
|
||||
>/dev/null 2>&1 || true
|
||||
fi
|
||||
@@ -0,0 +1,75 @@
|
||||
# Kickoff: Voll-Audit MC2 (Front+Backend) + Box-SSH + Optimierungsplan
|
||||
|
||||
> Für eine **frische Claude-Code-Session mit echtem Terminal** (direkter SSH-Zugang zur Box).
|
||||
> Aufgabe vom Nutzer: front/backend komplett scannen (IST-Zustand), via SSH auf die Box gehen
|
||||
> (ausdrücklich erlaubt), und am Ende einen **sauberen, strukturierten Plan** liefern, wie es
|
||||
> weitergeht — inkl. Optimierungen für ALLE LLM-Rollen und einem **neuen `scout`-Modell**.
|
||||
> Erst auditieren/planen — keine destruktiven Änderungen ohne Plan-Freigabe.
|
||||
|
||||
## Zugänge
|
||||
- **Repo (Dev-PC):** `F:\Coding Stuff\mission-control-2` (Windows). Auf der Box: `~/mission-control-v2`.
|
||||
- **Box:** `ssh hitonabi@192.168.178.151` (key-based, funktioniert non-interaktiv). Hostname
|
||||
`tobisniceaiarbeitstier`, AMD Ryzen AI MAX+ 395 (Strix Halo), 122 GB RAM, GTT ~124 GB, Vulkan/RADV.
|
||||
- **Wichtige Box-Pfade:** `/etc/llama-swap/config.yaml` (Engine, `-watch-config`),
|
||||
`~/.hermes/config.yaml` (Hermes-Agent = Lucys Hirn), `/srv/models/` (GGUFs),
|
||||
`/opt/llamacpp-vulkan/llama-server` (Engine v9843).
|
||||
- **Dienste (systemd --user):** `hermes-gateway` (:8642), `hermes-webui` (:8787), llama-swap (:8080),
|
||||
MC2-Gateway (:9001). Live-Logs: `journalctl --user -u hermes-gateway -f`.
|
||||
|
||||
## Architektur (verifiziert 30.06.2026)
|
||||
- **llama-swap** (:8080) lädt GGUFs; Rollen = llama-swap **`aliases`**; Warm-Bleiben über
|
||||
`groups: { brains: { swap:false } }` + `ttl`.
|
||||
- **MC2-Gateway** (:9001/v1, FastAPI) mit virtuellen Lanes `chat`(→fast/heavy) & `coding`(→coder_lite/coder).
|
||||
Code: `backend/services/router_logic.py`, `routing_policy.py`, `roles.py`, `llamaswap.py`, `budget.py`,
|
||||
`fit.py`; Router `backend/routers/{models,routing,voice}.py`. Frontend: `frontend/src/views/ModelsView.tsx`,
|
||||
`components/models/*`, `SystemDrawer.tsx`.
|
||||
- **Hermes-Agent** (:8642) = Lucys Hirn (`~/.hermes/config.yaml: model.default`), Delegation `heavy`.
|
||||
- **Lucy-Client** (`client/lucy-desktop`, Electron) spawnt lokal pocket-tts (:8130); Persona+Umlaut-Logik
|
||||
im Client (`renderer/src/config.ts`, `lib/voice/useVoiceAgent.ts`).
|
||||
- **Features sind eigene Rollen** (hirn-unabhängig): Sehen=`Qwen3-VL-8B`, Hören=Whisper(`stt`),
|
||||
Sprechen=pocket-tts, Embedding=`Qwen3-Embedding-0.6B`(ttl 0, immer warm), Gedächtnis=Mem0/MCP.
|
||||
|
||||
## Modelle / Rollen (Stand jetzt)
|
||||
| Rolle (alias) | Modell | aktiv | ctt/ttl | Notiz |
|
||||
|---|---|---|---|---|
|
||||
| fast | Qwen3.6-35B-A3B | 3B | ttl 0, --parallel 2, mmproj | chat-Lane Default |
|
||||
| heavy | Qwen3.5-122B-A10B | 10B | ttl 600 | Delegation-Ziel |
|
||||
| coder | Qwen3-Coder-Next | — | ttl 600, spec draft-simple (Qwen3-0.6B) | |
|
||||
| coder_lite | Qwen3-Coder-30B-A3B | 3B | ttl 300 | coding-Lane Default |
|
||||
| vision | Qwen3-VL-8B-Instruct | — | ttl 300 | Screen-Beschreibung |
|
||||
| (embedding) | Qwen3-Embedding-0.6B | — | ttl 0 | immer warm |
|
||||
| **brain (NEU, Lucy)** | **gemma-4-26B-A4B-it** | 4B | ttl 0, **noch Alias `scout`** | Hermes default seit 30.06. |
|
||||
|
||||
## Bereits gemacht (30.06.)
|
||||
- Lucys Hirn `fast` → **`gemma-4-26B-A4B-it`** (`~/.hermes/config.yaml model.default`, Backup
|
||||
`~/.hermes/config.yaml.gemma-swap.bak`, hermes-gateway neu gestartet, verifiziert).
|
||||
- Umlaut-Fix: Client-Stammliste erweitert (`useVoiceAgent.ts`) **und** Server-Netz `_fix_umlauts`
|
||||
(`client/lucy-tts/pocket_server.py`). Hermes-System-Prompt fordert echte Umlaute (Client-Persona).
|
||||
- pocket-tts getunt: A1 safetensors-Voice-Cache, A2 Referenz-Cleaning, B3 FP-Gate skip kurze Audios,
|
||||
FAST_FIRST (kurzer 1. Chunk, TTFB ~1,3s), seriell (WORKERS=0) als Default. Details: `docs/LUCY_TTS_PLAN.md`.
|
||||
|
||||
## Offene Probleme / Aufgaben
|
||||
1. **Brain-Rolle fehlt als Konzept.** gemma läuft als Alias `scout`, ist NICHT in `brains: swap:false`
|
||||
→ unter Speicherdruck evakuierbar trotz ttl 0. Detail-Brief: **`docs/CLAUDE_CODE_BRIEF_lucy-brain-role.md`**
|
||||
(Backend+Frontend-Umbau: erstklassige `brain`-Rolle, auto-warm/ko-resident, Hermes-Sync).
|
||||
2. **`scout`-Rolle jetzt vakant** (gemma war scout, ist jetzt Hirn). **Neues scout-Modell suchen:**
|
||||
multimodaler Allrounder, klein/MoE (schnell, KO-Residenz-freundlich), Tools wünschenswert. Tagesaktuell
|
||||
recherchieren + Empfehlung.
|
||||
3. **MTP-Speculative-Decoding für gemma** (1,5–2× Durchsatz, 0 Qualitätsverlust): Engine kann `draft-mtp`
|
||||
✓; MTP-Drafter `gemma-4-26B-A4B-it-assistant` (~0,4B) muss nach `/srv/models/`, gemma-cmd ergänzen
|
||||
(`-md <assistant.gguf> --spec-type draft-mtp --spec-draft-n-max/-n-min`; **alte** `--draft-max/-min`
|
||||
sind entfernt!). MC2-UI/Backend kennt MTP-Drafter nicht (zeigte „Vocab ?") → erweitern. (Im Brain-Brief.)
|
||||
4. **Optimierung ALLER Rollen prüfen:** je Modell ctx/quant/ttl/Warm-Gruppe/Parallelität/Spec-Decoding
|
||||
(MTP für Gemma; Draft-Eignung für Qwen-Modelle) gegen das GTT-/RAM-Budget (~124 GB) und reale t/s
|
||||
auf Strix Halo (bandbreitenlimitiert → MoE bevorzugt; dichte Modelle langsam). KO-Residenz-Set
|
||||
sinnvoll definieren (was muss gleichzeitig warm sein für Lucy + IDE-Coding?).
|
||||
|
||||
## Deliverable
|
||||
Ein **strukturierter Plan** (z.B. `docs/OPTIMIZATION_PLAN.md`): IST-Zustand → konkrete Maßnahmen je Rolle
|
||||
(mit Begründung/Zahlen) → Reihenfolge/Risiko/Revert → offene Entscheidungen. Erst Plan, dann Umsetzung
|
||||
nach Freigabe. Alles reversibel (Config-Backups vor jeder Änderung).
|
||||
|
||||
## Leitplanken
|
||||
- Features (Sehen/Hören/Sprechen/Embedding/Gedächtnis) + IDE-Lanes dürfen NICHT brechen.
|
||||
- Vor jeder Box-Config-Änderung Backup; llama-swap reloadt per `-watch-config`.
|
||||
- Keine destruktiven Aktionen ohne Plan-Freigabe des Nutzers.
|
||||
@@ -0,0 +1,138 @@
|
||||
# AUTONOMIE-PLAN — „Die Box wartet sich selbst" (Kapitel 0)
|
||||
|
||||
**Erstellt:** 02.07.2026, abends — als Kickoff-Paket für die Bau-Session („leg los").
|
||||
**Nordstern (User):** Haushaltsgerät, kein Dev-Projekt. Der User macht KEINE Updates, fasst keine
|
||||
Konsole an. Er bekommt Meldungen von Lucy (Telegram) und sagt höchstens „mach". In 6 Monaten ohne
|
||||
KI-Hilfe darf nichts brachliegen — schlimmster erlaubter Zustand: „selbst-gepinnt und läuft".
|
||||
|
||||
**Kontext-Quellen:** Memory `open-threads.md` (Kapitel 0, Bausteine a–k) · `verdicts.md` ·
|
||||
`project-stack-state.md` · dieser Plan. Review-Historie: docs/REVIEW_2026-07-02.md.
|
||||
|
||||
---
|
||||
|
||||
## Architektur: 3 Säulen + Werkstatt
|
||||
|
||||
1. **Stabilität** — eigene Software (MC2, Lucy) EINGEFROREN + Selbsterhaltung. Gitea = Tresor, kein Update-Kanal.
|
||||
2. **Frische** — Fremd-Software (OS/Engine/Router/Hermes) updatet sich per Timer, mit Gate + Auto-Rollback + Selbst-Pinning.
|
||||
3. **Evolution** — Radar (Hermes-cron + Web-Suche) findet Sprünge/neue Tools; Modelle bencht die Box selbst.
|
||||
4. **Werkstatt** — Box-KIs setzen Wartung/Migrationen selbst um: Telegram-Ping → User „mach" → Branch → Code → Gate → Deploy → Bericht.
|
||||
|
||||
**Leitplanken (nicht verhandelbar):** Merge/Deploy nie ohne grünes Gate · Security-Config
|
||||
(approvals/Tokens/ufw) nie ohne explizites User-Ja · alles reversibel (Backup+Branch+Pin) ·
|
||||
kleine Wartung autonom, große nur vorbereitet+freigegeben.
|
||||
|
||||
---
|
||||
|
||||
## Multi-Agent-Verdikt (recherchiert + entschieden 03.07.2026)
|
||||
|
||||
**JA — chirurgisch an 2 Stellen; NEIN bei allem Echtzeit-nahen.** Hermes-Delegation ist reif:
|
||||
`delegate_task` (goal/context/toolsets, tasks-Array = parallel, role leaf|orchestrator),
|
||||
Config `delegation:` (max_concurrent_children, max_spawn_depth, eigenes delegation.model/provider).
|
||||
|
||||
- **E4 Radar = Fan-out-Muster:** Manager delegiert pro Komponente/Kategorie einen isolierten
|
||||
Leaf-Subagent (eigener 65k-Kontext) → recherchiert, **filtert Marketing-Müll**, liefert nur
|
||||
Substanz-Summary → Manager konsolidiert. Müll erreicht den Haupt-Kontext nie.
|
||||
- **E6 Werkstatt = Worker/Reviewer-Muster:** Worker patcht im Worktree, SEPARATER Reviewer-
|
||||
Subagent (frischer Kontext) prüft Diff gegen Auftrag, dann erst Gate+Telegram. (Verdent-Muster,
|
||||
lokal nachgebaut.)
|
||||
- **NICHT bei Lucy-Voice** (Konsens 2026: Agent-Ketten ≈ 3× Latenz bei Echtzeit — unsere 2 s sind
|
||||
heilig). Lucy bekommt stattdessen Delegation als FEATURE: „recherchier ich im Hintergrund"
|
||||
→ Background-Subagent → Ergebnis via Telegram. **NICHT bei der IDE-Lane** (eigene Agent-Loops).
|
||||
- **Kein Hirn-Tausch:** Parent bleibt Qwen3.6; `delegation.model/provider` → MC2-Gateway.
|
||||
Manager-/Reviewer-Kandidat = **gpt-oss-120b** (Faden B5 entscheidet heavy-Lane UND Manager-Rolle
|
||||
in einem). `max_concurrent_children: 2` — passt exakt auf die 2 Hirn-Slots (keine Swap-Thrashes);
|
||||
`max_spawn_depth: 1` zum Start (flat), Orchestrator erst bei Bedarf.
|
||||
- **Ehrliche Physik:** eine GPU ⇒ Parallelität bringt KEIN Tempo (geteilte Bandbreite), sondern
|
||||
Kontext-Isolation + Qualität. Für nächtliche cron-Jobs genau richtig; deshalb Voice-Ausschluss.
|
||||
|
||||
---
|
||||
|
||||
## Was schon existiert (nutzen, nicht neu bauen!)
|
||||
|
||||
| Baustein | Wo |
|
||||
|---|---|
|
||||
| Update-Pipeline je Ebene mit Backup+Checks | UI-Buttons: OS/Engine (`deploy/update-engine.sh` — hat schon stack-postcheck+Auto-Rollback!), Router (`update-swap.sh`), Hermes (`POST /api/maintenance/hermes-update` → backup→update→doctor→Postcheck) |
|
||||
| Qualitäts-Gate (Abnahmetest für ALLES) | `deploy/hermes-postcheck.sh`: Mem0+Plugin+Config-Drift-Scan+Tool-Smoke+Voice-Smoke, E2E grün. ⚠️ pipefail-Lektion: Stream-Probes brauchen `set +o pipefail`-Subshell |
|
||||
| Tägliches Backup + Restore | `deploy/backup.sh` (Timer 03:30) / `restore.sh` |
|
||||
| Telegram | Hermes-Gateway läuft mit Telegram; CLI `hermes send` (Syntax in Session verifizieren: `hermes send --help`) |
|
||||
| Cron | `hermes cron` (Subcommand existiert) |
|
||||
| Coding-KI | Qwen3-Coder-Next (coder-Lane, 131k ctx); Hermes v0.18 hat „Coding Projects mit git-worktree-Management" |
|
||||
| PC-Fernsteuerung | `client/hermes-pc` Executor (Token `HERMES_PC_TOKEN`) — für Lucy-Builds auf dem Windows-PC |
|
||||
| Bench-Harnesses | `deploy/bench/brain-bench.sh` + `deploy/bench/model-bench.sh` (mit diesem Kickoff ins Repo gelegt) |
|
||||
| Health/Repair | LucyHealthCard + `/api/health` brain-Check; systemd `Restart=on-failure` überall |
|
||||
|
||||
---
|
||||
|
||||
## Etappen (in dieser Reihenfolge bauen)
|
||||
|
||||
### E1 · Nervensystem: Meldeweg (klein, zuerst — alles Weitere nutzt ihn)
|
||||
- `deploy/notify.sh "<Nachricht>"`: primär `hermes send` (Telegram), Fallback: Logfile + `wall`.
|
||||
Syntax von `hermes send` zuerst auf der Box verifizieren (`bash -lc 'hermes send --help'`).
|
||||
- Verifikation: Testnachricht landet beim User auf Telegram.
|
||||
|
||||
### E2 · Frische: Auto-Update-Timer + Rollback + Pin-Register
|
||||
- **Pin-Register:** `/srv/models/mc2-pins.json` — `{komponente: {pinned: bool, version, grund, datum}}`.
|
||||
Gepinnte Ebene wird vom Timer übersprungen; Pin-Meldung ging via E1 raus.
|
||||
- `deploy/autoupdate.sh` (Reihenfolge: swap → engine → hermes; Modelle NICHT auto):
|
||||
je Ebene: Update-Check → wenn Update: bestehenden Pfad nutzen (update-engine.sh / update-swap.sh /
|
||||
hermes-update-Job via API + Job-Polling) → Postcheck rot ⇒ Rollback (existiert je Pfad) + Pin
|
||||
setzen + notify; grün ⇒ notify „eingespielt: X→Y".
|
||||
- systemd-User-Timer `mc2-autoupdate.timer` (wöchentlich So 04:30, nach Backup), Unit nach
|
||||
`~/.config/systemd/user/`, in `deploy/deploy.sh` registrieren (Muster: mc2-backup.timer).
|
||||
- **OS-Ebene:** apt braucht sudo → `unattended-upgrades` (nur Security) in der EINMALIGEN
|
||||
sudo-Session einrichten (zusammen mit Faden C14: stale v1-Unit weg). Bis dahin: OS aus dem Timer raus.
|
||||
- Verifikation: Timer-Dry-Run (`bash autoupdate.sh`), einmal echt durchlaufen lassen, Telegram-Meldungen prüfen; künstlichen Fehlschlag provozieren (Postcheck-FAIL faken) → Rollback+Pin+Meldung.
|
||||
|
||||
### E3 · Stabilität: Lucy-Produktiv-Build (einfrieren)
|
||||
- Im Lucy-Repo (`F:\Coding Stuff\lucy`): electron-builder ergänzen (portable oder NSIS),
|
||||
**asar-Caveat:** `app.getAppPath()` zeigt im Build woanders hin → `LUCY_TTS_DIR` absolut setzen
|
||||
(Installer-Config oder Start-BAT) — siehe Kommentar in `lucy-desktop/src/main/index.ts`.
|
||||
- Verknüpfung auf den Build; `Lucy-Neustart.bat` auf Build umstellen; Dev-Modus bleibt für Wartung.
|
||||
- Verifikation: Build starten → Worker-Pool bereit, Sprech-Roundtrip.
|
||||
- Danach gilt Lucy als EINGEFROREN (Änderungen nur noch via Werkstatt-Kreislauf).
|
||||
|
||||
### E4 · Evolution: Radar
|
||||
- Hermes-cron (monatlich, z. B. 1. des Monats 09:00): Prompt-Job —
|
||||
„Inventar von `/api/maintenance/updates` + `/v1/models` + Versionsliste einlesen; per web_search
|
||||
Changelogs/Neuigkeiten zu: llama.cpp, llama-swap, hermes-agent, pocket-tts, onnx-asr/Parakeet,
|
||||
Silero, smart-turn, Electron, three-vrm + Kategorie-Scan (bessere lokale Coder-/Vision-/TTS-
|
||||
Modelle für Strix Halo 128 GB) + **IDE-/Agent-Tool-Kategorie** (Kilo Code, OpenCode, Claude Code
|
||||
— Lehre 03.07.: Roo Code war Stunden nach der Empfehlung als „eingestellt seit April" entlarvt;
|
||||
genau solche Tod-/Nachfolger-Meldungen muss das Radar fangen); gegen [[verdicts]] halten
|
||||
(Verworfenes nicht wieder vorschlagen);
|
||||
Ergebnis als kurzer deutscher Report → notify.sh".
|
||||
- Ablage des Prompts versioniert: `deploy/radar-prompt.md`, cron ruft ihn.
|
||||
- Verifikation: Cron einmal manuell feuern, Telegram-Report prüfen (Qualität grob checken).
|
||||
|
||||
### E5 · Evolution: Modell-Selbst-Evaluation
|
||||
- `deploy/bench/model-bench.sh <gguf-pfad> [extra-flags]` (liegt bei) → pp/tg-Zahlen.
|
||||
- Workflow (halbautonom v1): Radar meldet Kandidat → User „mach" → Hermes lädt (hf CLI im
|
||||
backend-venv existiert) → bencht → notify mit Vergleichstabelle → User-Entscheid → Config-Eintrag.
|
||||
- Verifikation: einmal mit vorhandenem Modell durchspielen.
|
||||
|
||||
### E6 · Werkstatt: Wartungs-Workflow (der Schachzug — zuletzt, auf E1–E5 aufbauend)
|
||||
- Hermes-Skill/Workflow „wartung": Auftrag (aus Radar oder User) → `git worktree` im Box-Checkout
|
||||
(MC2) bzw. via PC-Executor (Lucy) → Änderung durch coder-Lane → Gate = Postchecks + Build →
|
||||
Branch-Push zu Gitea → notify mit Diff-Zusammenfassung → User „merge"/„verwerfen" via Telegram
|
||||
(Hermes approvals-Mechanik nutzen) → bei merge: Deploy-Pfad + Backup + Postcheck.
|
||||
- Scope v1 BEWUSST klein: Config-Schlüssel-Migrationen, Dependency-Bumps (Lockfile), Ein-Datei-
|
||||
Patches. Architektur-Umbauten: nur Plan+Branch vorbereiten.
|
||||
- Verifikation: einen echten Klein-Fall durchspielen (Kandidat: ServicesCard-Restart-Bug, Faden B9
|
||||
— perfekte Werkstatt-Gesellenprüfung: 1 Datei, klarer Fix, Gate vorhanden).
|
||||
|
||||
---
|
||||
|
||||
## Definition of Done (Kapitel 0)
|
||||
1. Wöchentlicher Auto-Update-Lauf mit Telegram-Meldung, verifiziertem Rollback+Pin-Pfad.
|
||||
2. Lucy läuft als gebaute, eingefrorene App.
|
||||
3. Monatlicher Radar-Report kommt auf Telegram an.
|
||||
4. Ein Modell-Kandidat wurde einmal selbst gebencht + gemeldet.
|
||||
5. Die Werkstatt hat EINEN echten Klein-Fix eigenständig bis zum Merge-Vorschlag gebracht.
|
||||
6. Runbook-Seite existiert (`docs/RUNBOOK.md`, 1 Seite Mensch-Anleitung).
|
||||
|
||||
## Session-Start-Checkliste („leg los")
|
||||
1. `git -C ~/mission-control-v2 log --oneline -1` auf der Box == origin/main? (SSH: hitonabi@192.168.178.151)
|
||||
2. Health: `curl -s http://192.168.178.151:9001/api/health` → brain ready?
|
||||
3. `bash -lc 'hermes send --help'` → Syntax für E1 klären.
|
||||
4. Dann E1 → E6 der Reihe nach; jede Etappe committen + deployen + real verifizieren (Telegram!).
|
||||
Gitea-Push: PowerShell, Auth flatterhaft → einmal Retry.
|
||||
@@ -0,0 +1,62 @@
|
||||
# Backup & Restore
|
||||
|
||||
Sichert den **nicht wiederherstellbaren Zustand** der AI-Box. Code kommt aus Git,
|
||||
Modelle sind neu ladbar — gesichert wird nur, was sonst weg wäre.
|
||||
|
||||
## Was im Backup ist
|
||||
- **Gedächtnis:** `/srv/models/mem0/` (Chroma-Vektoren + `history.db`)
|
||||
- **Hermes:** `~/.hermes/config.yaml`, `~/.hermes/.env` (**Secrets!**), `~/.hermes/plugins/`
|
||||
- **Engine:** `/etc/llama-swap/config.yaml`
|
||||
|
||||
Nicht enthalten (bewusst): GGUF-Modelle, MC2-Code (Git), venvs, systemd-Units (aus `deploy.sh`).
|
||||
|
||||
Jedes Backup ist ein Tarball `mc2-state-<zeitstempel>.tar.gz` unter `/srv/models/mc2-backups/`,
|
||||
`chmod 600` (enthält `.env`). Es werden die letzten **14** behalten.
|
||||
|
||||
## Off-Box-Spiegel (C12) — zweite Kopie WEG von der Box
|
||||
Nach jedem lokalen Backup spiegelt `backup.sh` die Tarballs per `rsync` auf ein **zweites Gerät**
|
||||
(Default: **Proxmox-Host** `root@192.168.178.108:/var/lib/vz/mc2-backups`). Stirbt die NVMe der
|
||||
Box, liegen die Backups noch auf dem Proxmox. Die Retention (14) wird per `--delete` mitgezogen,
|
||||
`chmod 600` bleibt erhalten. Der Off-Box-Sync ist **best-effort**: schlägt er fehl, gilt das
|
||||
lokale Backup trotzdem als erfolgreich (Warnung im Log).
|
||||
|
||||
- **Auth:** eigener SSH-Key `~/.ssh/mc2_offsite` (nur für dieses Backup), im Proxmox in
|
||||
`root/.ssh/authorized_keys` **gehärtet** eingetragen: `from="<Box-IP>"`, kein Port-/Agent-/
|
||||
X11-Forwarding, kein PTY → der Key kann nur rsync-Backup, keinen Voll-Root-Fernzugang.
|
||||
- **Ziel/Key überschreibbar:** `MC_BACKUP_OFFSITE` (leer = Off-Box aus) und `MC_BACKUP_OFFSITE_KEY`.
|
||||
|
||||
## Backup erstellen
|
||||
- **Automatisch:** systemd-Timer `mc2-backup.timer`, täglich ~03:30. Status:
|
||||
`systemctl --user list-timers mc2-backup.timer`
|
||||
- **Manuell (Box):** `bash ~/mission-control-v2/deploy/backup.sh`
|
||||
- **UI:** Wartungs-Drawer → „Snapshot erstellen"
|
||||
|
||||
## Wiederherstellen (Restore)
|
||||
Restore läuft **nur per CLI auf der Box** (bewusst — er stoppt Dienste und überschreibt Configs).
|
||||
Vor dem Zurückspielen macht das Skript automatisch ein Sicherheits-Backup des aktuellen Zustands.
|
||||
|
||||
```bash
|
||||
cd ~/mission-control-v2
|
||||
bash deploy/restore.sh --list # vorhandene Backups anzeigen
|
||||
bash deploy/restore.sh --dry-run latest # zeigen, was passieren würde
|
||||
bash deploy/restore.sh latest # neuestes wiederherstellen (mit Rückfrage)
|
||||
bash deploy/restore.sh mc2-state-YYYYMMDD-HHMMSS.tar.gz # bestimmtes Backup
|
||||
```
|
||||
|
||||
Ablauf: Sicherheits-Backup → Dienste stoppen (`mem0-service`, `mission-control-2`,
|
||||
`hermes-gateway`) → Dateien zurückspielen (mem0 wird **ersetzt**, Configs überschrieben) →
|
||||
Dienste starten → Health-Check.
|
||||
|
||||
### Wenn die Box-Platte tot ist (Off-Box-Restore)
|
||||
`restore.sh` kennt den Off-Box-Spiegel: fehlt ein Backup lokal, holt es sich das Skript
|
||||
automatisch vom Proxmox. Kein manuelles Kopieren nötig.
|
||||
|
||||
```bash
|
||||
bash deploy/restore.sh --list # zeigt lokal UND Off-Box (Proxmox)
|
||||
bash deploy/restore.sh --pull-offsite # ganzen Off-Box-Bestand nach lokal spiegeln
|
||||
bash deploy/restore.sh latest # zieht das jüngste — vom Proxmox, falls lokal leer
|
||||
```
|
||||
|
||||
## Erledigt
|
||||
- **Off-Box-Spiegel (C12):** rsync auf den Proxmox-Host, live E2E verifiziert (2026-07-03).
|
||||
Box ist Bare Metal (kein Proxmox-Gast → kein vzdump), daher aktiver Push statt vzdump-Pull.
|
||||
@@ -0,0 +1,55 @@
|
||||
# Mission Control 2.0 — Bedienung (kurz & klartext)
|
||||
|
||||
**Öffnen:** `http://192.168.178.151:9001` (vom Windows-PC im LAN). Dark/Light-Umschalter oben rechts,
|
||||
**Cmd/Strg+K** springt zu jedem Bereich.
|
||||
|
||||
> Im Alltag fasst du MC kaum an: `model: auto` + Auto-Swap laden Modelle selbst. Du öffnest es, um ein
|
||||
> Modell zu installieren/tauschen, die Auslastung zu prüfen, Gedächtnis zu pflegen oder ein Tool zu verbinden.
|
||||
|
||||
## Die 7 Bereiche
|
||||
- **Zentrale** — Live CPU/RAM/GPU-Auslastung und Updates auf einen Blick.
|
||||
- **Modell-Zentrale** — installierte Modelle + neue finden/laden + Gateway-Routing (auto-Rolle).
|
||||
- **Diagnose** — Live-Dienste, Metriken, detaillierte Auslastung und System-Logs.
|
||||
- **Gedächtnis** — geteilte Fakten/Regeln, die ALLE Tools (Hermes, IDEs) via MCP lesen/schreiben.
|
||||
- **Verbinden** — fertige Konfig-Snippets für deine IDEs (Roo Code, Zed, etc.).
|
||||
- **Hermes** — Agent-Status + „Hermes öffnen".
|
||||
- **Anleitung** — Schritt-für-Schritt Einrichtung für Vibe-Coding auf deinem PC.
|
||||
|
||||
## Modell installieren
|
||||
**Tab „Modelle & Routing" → „Modelle finden":**
|
||||
- **Kuratiert:** auf einer Empfehlungs-Karte „Installieren" klicken (⭐ = beste Wahl je Kategorie).
|
||||
- **Eigenes (HuggingFace):** oben **HF-URL oder `org/repo`** einfügen → „Quants laden" → Quant wählen →
|
||||
„Installieren". Oder die **Suchleiste** nutzen → Treffer anklicken → Quant → Installieren.
|
||||
- Der Download läuft als Job mit **Fortschrittsbalken** oben; llama-swap pflegt das Modell automatisch ein.
|
||||
|
||||
## LLM tauschen (z.B. anderes „fast"-Hirn)
|
||||
„Modelle & Routing" → **„Installiert"**: in der Zeile des Modells im **Rollen-Dropdown**
|
||||
`fast` (bzw. `heavy`/`coder`/`vision`/`scout`) wählen → der Alias wandert auf dieses Modell.
|
||||
- `model: auto` nutzt ab sofort dieses Modell als schnelles/schweres Hirn — für Hermes **und** Vibe Coding.
|
||||
- **Kontext** ändern: auf die Kontext-Zahl (✎) klicken. **Entfernen:** 🗑 (GGUF-Datei bleibt erhalten).
|
||||
- „Auto-Swap" = llama-swap lädt automatisch, was gerade angefragt wird; du musst nichts laden/entladen.
|
||||
|
||||
## IDE verbinden (Vibe Coding am eigenen PC)
|
||||
„Verbinden" → Tool wählen (Roo Code/OpenCode/Zed/Continue) → Snippet kopieren. Zeigt auf
|
||||
`http://192.168.178.151:9001/v1`, Modell **`auto`**. Memory-MCP-Snippet separat einfügen → geteiltes Gedächtnis.
|
||||
|
||||
## Gedächtnis pflegen
|
||||
„Gedächtnis": Fakt/Regel hinzufügen (Kategorie wählen), suchen/filtern, **🧹 Aufräumen** entfernt Dubletten.
|
||||
Das ist die geteilte „Verfassung" für alle Tools.
|
||||
|
||||
## Wartung & Backup
|
||||
„System" → **Wartung & Updates**: Badge (offene OS-Pakete / Engine / Modell-Upgrades), Buttons
|
||||
**OS aktualisieren · Engine aktualisieren · Engine neu starten · Reboot**, Modell-Upgrade-Vorschläge
|
||||
(1-Klick), **Backup jetzt** (Gedächtnis-DB + Configs), Dienste-Health.
|
||||
|
||||
`Engine neu starten`, Logs & Modell-Upgrades laufen sofort (NOPASSWD vorhanden). **OS-Update + Reboot**
|
||||
brauchen einmalig erweiterte sudoers — `sudo visudo`, ergänze:
|
||||
```
|
||||
hitonabi ALL=(root) NOPASSWD: /usr/bin/apt-get, /usr/sbin/reboot
|
||||
```
|
||||
(Engine-Update: `MC_ENGINE_UPDATE_CMD` in der mc2-Unit setzen — Befehl, der /opt/llamacpp aktualisiert.)
|
||||
Danach ist die komplette Wartung klicki-bunti, ohne Passwort.
|
||||
|
||||
## Wenn etwas hakt
|
||||
- Modell antwortet nicht → „System" → Dienste-Health (Engine online?) + Engine-Logs-Link (llama-swap `/ui`).
|
||||
- Hermes langsam/komisch → im Hermes-WebUI **neuen Chat** starten (frische Session); Details: `docs/CUTOVER.md`.
|
||||
@@ -0,0 +1,106 @@
|
||||
# Claude-Code-Auftrag: „Brain"-Rolle für Lucy (Front- + Backend-Umbau)
|
||||
|
||||
> Ziel: Lucys Hirn (aktuell `gemma-4-26B-A4B-it`) als **erstklassige, dauer-warme Rolle** im MC2-Stack
|
||||
> verankern — statt es als `scout` zu führen, das nach `ttl` entladen wird. Features (Sehen/Hören/
|
||||
> Sprechen/Embedding) dürfen NICHT verloren gehen, KO-Residenz + Budget müssen sauber bleiben.
|
||||
|
||||
## Kontext / Architektur (Ist-Zustand, verifiziert 30.06.)
|
||||
- **Engine:** llama-swap (config: `/etc/llama-swap/config.yaml` auf der Box, läuft mit `-watch-config`).
|
||||
Rollen werden als llama-swap **`aliases`** gesetzt; Warm-Bleiben über `groups:` (eine Gruppe
|
||||
`brains: { swap: false }`) + `ttl`. Installierte Modelle: `Qwen3.6-35B-A3B` (fast, ttl 0),
|
||||
`Qwen3.5-122B-A10B` (heavy, ttl 600), `Qwen3-Coder-Next` (coder, ttl 600),
|
||||
`Qwen3-Coder-30B-A3B-Instruct` (coder_lite, ttl 300), `Qwen3-VL-8B-Instruct` (vision, ttl 300),
|
||||
`Qwen3-Embedding-0.6B` (ttl 0 = immer warm), `gemma-4-26B-A4B-it` (ttl 180).
|
||||
- **Gateway:** MC2 builtin auf `:9001/v1` mit virtuellen Lanes `chat` (→ fast/heavy) und
|
||||
`coding` (→ coder_lite/coder). Code: `backend/services/router_logic.py`, `routing_policy.py`.
|
||||
**Lucy nutzt die Lanes NICHT** — sie ist der Hermes-Agent.
|
||||
- **Lucy-Hirn:** Hermes-Agent (`:8642`), Brain = `~/.hermes/config.yaml` → `model.default`
|
||||
(jetzt `gemma-4-26B-A4B-it`), Delegation = `heavy`. Persona/Prompt kommt aus dem **Client**
|
||||
(`client/lucy-desktop/src/renderer/src/config.ts`), nicht aus der Box-Config.
|
||||
- **Hardware:** Strix Halo (Ryzen AI Max+ 395), 122 GB, GTT ~124 GB, bandbreitenlimitiert →
|
||||
MoE mit wenig aktiven Params ist schnell, dichte Modelle langsam.
|
||||
|
||||
## Root-Cause des Bugs
|
||||
1. Kanonische Rollen `ROLE_IDS = {fast, heavy, coder, vision, scout}` — **keine `brain`/Agent-Rolle**.
|
||||
(Kommentar im Code: „kein agent/reasoning mehr".)
|
||||
2. Es gibt keine Verknüpfung „Hirn-Modell ⇒ muss warm bleiben". gemma trägt den Alias `scout`,
|
||||
ist NICHT in der `brains`-Gruppe, `ttl 180` → entlädt im Leerlauf → Lucy lädt kalt nach (~5–10 s).
|
||||
3. `roles.py` referenziert noch eine `hermes`-Rolle, die in `ROLE_IDS` nicht existiert → Inkonsistenz.
|
||||
|
||||
## Aufgabe
|
||||
|
||||
### Backend (`backend/`)
|
||||
1. **Neue erstklassige Rolle `brain` einführen** (oder `hermes` reaktivieren) und in ALLEN
|
||||
Quellen synchronisieren (heute dupliziert!):
|
||||
- `services/llamaswap.py` → `ROLE_IDS`
|
||||
- `services/sources.py` → `ROLE_IDS` + Titel/Icon (analog `scout: {title, icon}`)
|
||||
- `services/maintenance.py` (Modell-Upgrade-Rollenliste)
|
||||
- `services/roles.py` → `_capability_suit`/`_pref`/`_reason` für `brain` (Tools Pflicht,
|
||||
mittlere Größe + MoE bevorzugt, niedrige Latenz). `hermes`-Altlast bereinigen.
|
||||
2. **Rolle `brain` ⇒ automatisch warm + ko-resident.** Beim Zuweisen der `brain`-Rolle
|
||||
(`POST /api/models/{model_id}/role`, `routers/models.py`):
|
||||
- Modell via `set_group("brains", members=[...], swap=False, persist=True)` in die Warm-Gruppe
|
||||
aufnehmen (bestehender Helper in `llamaswap.py`).
|
||||
- `ttl: 0` setzen (oder hoch), vorheriges Brain-Modell aus `brains` entfernen + ttl entspannen.
|
||||
- Embedding (`Qwen3-Embedding-0.6B`, ttl 0) MUSS warm bleiben — nicht verdrängen.
|
||||
3. **Single Source of Truth für „Lucys Hirn":** beim Setzen der `brain`-Rolle auch Hermes'
|
||||
`~/.hermes/config.yaml` → `model.default` auf das Modell setzen + `systemctl --user restart
|
||||
hermes-gateway` (mit Backup der config). So bleibt MC2-UI-Auswahl ↔ Hermes-Brain konsistent.
|
||||
4. **Budget/KO-Residenz-Safety:** Brain(warm) + Embedding(warm) + on-demand Vision + Coder müssen
|
||||
in GTT (~124 GB) passen. `services/budget.py`/`fit.py` nutzen, bei OOM warnen (nicht hart laden).
|
||||
5. **Lanes/IDEs unangetastet lassen:** chat/coding-Routing (`router_logic.py`) bleibt wie es ist;
|
||||
`fast` bleibt für die chat-Lane warm.
|
||||
|
||||
### Frontend (`frontend/src/`)
|
||||
1. **Rollen-Taxonomie** um `brain` erweitern (Badges/Titel/Icon) — Duplikat zur Backend-Liste
|
||||
finden & angleichen (`components/models/*`, `views/ModelsView.tsx`, evtl. `ModelBadges.ROLES`).
|
||||
gemma darf NICHT mehr als `scout` erscheinen.
|
||||
2. **Rollen-Zuweisungs-Modal:** `brain` auswählbar; Warm-/KO-Residenz-Status anzeigen
|
||||
(ist das Brain in `brains`? warm? ttl?). Empfehlung via bestehendem `/api/roles/{role}/recommend`.
|
||||
3. **Warm-Set / GTT-Budget sichtbar machen** (Dashboard/Models): was ist gerade ko-resident,
|
||||
passt es ins Budget? (nutzt `/api/models` `running` + budget-Infos).
|
||||
4. Optional: „Lucy-Hirn"-Selector, der die `brain`-Rolle setzt (treibt Backend #2 + #3).
|
||||
|
||||
### Zusatz: MTP-Speculative-Decoding für Gemma (Durchsatz für agentische Arbeit)
|
||||
Lucy ist ein **voller Agent** (Tools/MCP/PC-Control/Vision/Delegation) — Durchsatz zählt, nicht nur
|
||||
TTFB. Gemma 4 bringt **MTP** (Multi-Token-Prediction) mit → 1,5–2× schneller bei null Qualitätsverlust.
|
||||
- **Engine kann es bereits:** `llama-server` v9843 listet `--spec-type … draft-mtp …`. ✓
|
||||
- **Vocab-kompatibler „Draft" = Gemmas eigener MTP-Kopf** `gemma-4-26B-A4B-it-assistant` (~0,4B,
|
||||
identischer Tokenizer per Konstruktion). Quelle z.B. `unsloth/gemma-4-26B-A4B-it-GGUF` (enthält den
|
||||
MTP-Drafter, PR ggml-org/llama.cpp#23398). Muss nach `/srv/models/...` geladen werden (noch nicht da).
|
||||
- **Flag-Änderung beachten:** `--draft-max/--draft-min` sind ENTFERNT → `--spec-draft-n-max` /
|
||||
`--spec-draft-n-min`. Drafter laden via `-md <assistant.gguf> --spec-type draft-mtp`. Exakte Flags
|
||||
des Builds mit `llama-server --help` gegenprüfen.
|
||||
- **MC2-Bug:** die Spec-Draft-Auswahl (UI + Backend) kennt **nur klassische Drafts** (vergleicht stur
|
||||
`n_vocab`/`pre`) und bot fälschlich nur den Qwen-Draft an („Vocab ?"). **Erweitern:** MTP-Drafter
|
||||
(`gemma4_assistant`-Arch) als gültigen, vocab-kompatiblen Draft erkennen, den passenden `-assistant`-
|
||||
GGUF anbieten, und beim Aktivieren `-md … --spec-type draft-mtp --spec-draft-n-max/-n-min` in die
|
||||
llama-swap-cmd schreiben. Co-Residenz: Drafter ~0,4B → vernachlässigbar.
|
||||
|
||||
## Akzeptanzkriterien
|
||||
- [ ] `gemma-4-26B-A4B-it` wird in der UI als **`brain`** geführt (nicht `scout`).
|
||||
- [ ] Es bleibt **warm**: in `brains: swap:false`, `ttl 0`; überlebt > alter ttl Leerlauf
|
||||
(Verifikation: `curl :8080/running` nach >5 Min Idle zeigt das Modell weiterhin `ready`).
|
||||
- [ ] Lucy antwortet ohne Kalt-Nachladen nach Leerlauf.
|
||||
- [ ] Vision (`Qwen3-VL-8B`), Embedding (`Qwen3-Embedding-0.6B`), Coder bleiben funktionsfähig &
|
||||
ko-resident; **kein OOM**; IDE-Lanes (chat/coding) unverändert.
|
||||
- [ ] Brain-Wechsel in der UI aktualisiert llama-swap (Warm-Gruppe) UND Hermes `model.default`
|
||||
(+ Restart), reversibel (Backup).
|
||||
|
||||
## Sofort-Hotfix (unabhängig vom Umbau — macht Lucy JETZT warm)
|
||||
Auf der Box, bis der saubere Umbau steht:
|
||||
```bash
|
||||
# gemma als immer-warm + in die brains-Gruppe (Backup zuerst!)
|
||||
cp /etc/llama-swap/config.yaml /etc/llama-swap/config.yaml.bak
|
||||
# gemma ttl 180 -> 0 und groups.brains.members um gemma-4-26B-A4B-it ergänzen
|
||||
# (manuell editieren oder via MC2 set_group), dann:
|
||||
# llama-swap reloadt per -watch-config automatisch.
|
||||
```
|
||||
|
||||
## Wichtige Dateien (Einstieg)
|
||||
- Backend: `services/llamaswap.py` (Rollen=aliases, `set_role_alias`, `set_group`, `register_model`),
|
||||
`services/roles.py`, `services/sources.py`, `services/routing_policy.py`,
|
||||
`services/router_logic.py`, `services/budget.py`, `services/fit.py`, `routers/models.py`,
|
||||
`routers/routing.py`.
|
||||
- Frontend: `views/ModelsView.tsx`, `components/models/*`, `components/SystemDrawer.tsx`.
|
||||
- Box (nicht im Repo): `/etc/llama-swap/config.yaml`, `~/.hermes/config.yaml`.
|
||||
@@ -0,0 +1,45 @@
|
||||
# Cutover — Stand & Anleitung
|
||||
|
||||
## Was auf der Box LÄUFT (verifiziert)
|
||||
- **MC2** auf `:9001` (sudo-freier User-Dienst, `~/mission-control-v2`). Update: `deploy/deploy.sh`.
|
||||
- **Modelle/Rollen:** `fast` = Qwen3.6-35B-A3B, `heavy` = Qwen3.5-122B-A10B, `coder` = Qwen3-Coder-30B,
|
||||
`vision` = Qwen3-VL-8B, `scout` = Qwen3-8B, `hermes` = Hermes-4-14B (immer warm, ttl 99999). Alle
|
||||
tool-fähig (`--jinja` wo nötig). Legacy `manager`/`reviewer` entfernt.
|
||||
- **Gateway (eingebaut, `:9001/v1`, OpenAI-kompatibel):** `model: auto` → kurz/Standard = `fast`,
|
||||
lang/komplex = `heavy`. End-to-End verifiziert.
|
||||
- **Hermes:** **eigenes festes Hirn = Hermes-4-14B** (`model.model: hermes`) + **Delegation an `heavy`**.
|
||||
`model:auto` ist NUR für Vibe Coding/IDEs, nicht Hermes. MCP verdrahtet: `mission-control-memory` +
|
||||
`mission-control-stack`. Verifiziert: „bist du da?" → 3s, sauber, kein Thrash.
|
||||
- **Gedächtnis vereinheitlicht:** MC2 nutzt die bestehende DB (`mission-control-memory.db`) — geteilte
|
||||
„Verfassung" für Cockpit, Hermes, IDEs.
|
||||
- **Cockpit-Features:** HF-Link/Suche-Install (W2), Rollen/ctx/löschen-UX (W3), Wartung (W8: OS/Engine-
|
||||
Update, Restart, Reboot, Logs, dyn. Modell-Upgrades), Bedien-Anleitung (BEDIENUNG.md + Hilfe-Link).
|
||||
|
||||
## Thrash-Fix (war der „Hermes ist dumm"-Grund)
|
||||
Ursache war NICHT das Modell/die Session, sondern **kaputte/Cloud-Tools im Toolset** (browser ohne Chrome
|
||||
→ Loop, vision auf Text, natives memory falsch aufgerufen). Global abgeschaltet über
|
||||
`agent.disabled_toolsets` in `~/.hermes/config.yaml` (Achtung: `hermes tools disable` greift nur cli,
|
||||
NICHT den api_server). Natives memory zusätzlich aus (`memory.memory_enabled:false`); geteiltes
|
||||
Gedächtnis bleibt via MCP. Behaltene Tools: web/terminal/file/code_execution/skills/todo/session_search/
|
||||
clarify/delegation/cronjob + 2 MCP.
|
||||
|
||||
## Cutover-Schritte
|
||||
1. v2 läuft bereits auf `:9001` parallel — alles dort testen: `http://192.168.178.151:9001`.
|
||||
2. Vibe-Coding-Tools auf den Gateway zeigen (Verbinden-Tab → `:9001/v1`, `model: auto`).
|
||||
3. **v1 stilllegen** — bereits erledigt: `hermes-dashboard` (:9119) disabled (killte 4 v1-MCP-Zombies).
|
||||
**Noch offen (braucht dein sudo, NOPASSWD deckt nur `restart`):**
|
||||
```
|
||||
sudo systemctl disable --now mission-control # v1-Cockpit :9000 aus
|
||||
```
|
||||
`llama-swap` (System) + `hermes-gateway` (User) bleiben — die nutzt v2 weiter.
|
||||
4. Optional v2 auf den „Haupt"-Port legen — `MC_PORT` in der mc2-Unit.
|
||||
5. Backup vorher: Cockpit → System → „Backup jetzt".
|
||||
|
||||
## Offene Tuning-/Setup-Punkte (kein Blocker)
|
||||
1. **nesquena hermes-webui** (:8787) installieren (Plan Block C) → „Hermes öffnen" zeigt darauf statt :9119.
|
||||
2. **Ko-Residenz** `hermes`+`fast` (swap:false, Plan Block E) — GTT beobachten, bei OOM-Nähe zurück.
|
||||
3. **SSH→Windows** (voller PC-Zugriff): OpenSSH-Server am Windows-PC + Key `id_ed25519_hermes_agent`.
|
||||
4. **sudoers erweitern** (`apt-get`, `reboot`) → OS-Update/Reboot klicki-bunti (Zeilen in BEDIENUNG.md).
|
||||
5. **Delegation an heavy** ist konfiguriert; triggert modell-diskretionär bei echt harten Teilaufgaben.
|
||||
|
||||
> v1 bleibt bis zum `disable` lauffähig — Cutover ist reversibel (`sudo systemctl enable --now mission-control`).
|
||||
@@ -0,0 +1,326 @@
|
||||
# Konzept: Bare-Metal-Wiederaufbau + First-Run-Wizard (VOLLSTÄNDIG)
|
||||
|
||||
> **Zweck:** Plan für den „hard crash"-Fall — die Box (`tobisniceaiarbeitstier`, Ubuntu 26.04,
|
||||
> AMD Ryzen AI MAX+ 395 / gfx1151, 122 GB RAM) muss von **null** wiederherstellbar sein.
|
||||
> **Anspruch:** *ALLES* ist erfasst — Engine, Hermes-Konfig 1:1, 3D-Avatare, Voice/Klonstimme,
|
||||
> Memory, Browser, MCP, Skills, Secrets. Pro Komponente entscheidet der Nutzer im Wizard:
|
||||
> **1:1 zurück** · **neu/Default** · **weglassen**.
|
||||
>
|
||||
> Dieses Dokument ist die **Spezifikation für eine künftige Implementierungs-Session** — es baut
|
||||
> noch nichts. Stand: 2026-06-28.
|
||||
|
||||
---
|
||||
|
||||
## 1. Leitprinzip: 1:1 ODER modular — der Nutzer wählt
|
||||
|
||||
Der Wizard behandelt jede Komponente als **eigene Kachel mit drei Modi**:
|
||||
|
||||
| Modus | Bedeutung |
|
||||
|---|---|
|
||||
| 📦 **1:1 aus Backup** | Exakter alter Zustand wird zurückgespielt (Configs, Daten, Klonstimme, Avatare). |
|
||||
| 🆕 **Neu / Default** | Frische Installation mit sinnvollen Defaults (z.B. entfesseltes Hermes-Profil, Default-Avatar). |
|
||||
| ⏭️ **Weglassen** | Komponente wird (vorerst) nicht installiert. |
|
||||
|
||||
Ein „Alles 1:1"-Knopf wählt überall 📦 (wo ein Backup existiert), sonst 🆕. So bekommt der Nutzer
|
||||
entweder den exakten alten Stand oder kann gezielt entrümpeln.
|
||||
|
||||
---
|
||||
|
||||
## 2. Was es schon gibt (wiederverwenden)
|
||||
|
||||
| Asset | Datei | Deckt ab |
|
||||
|---|---|---|
|
||||
| Code-Deploy | `deploy/deploy.sh` | git pull + venvs + Units + Restart. **Setzt Engine/llama-swap/Dirs/venvs voraus.** |
|
||||
| Zustands-Backup | `deploy/backup.sh` + `mc2-backup.{service,timer}` | **Aktuell** (live): mem0, `~/.hermes/{config.yaml,.env,plugins}`, llama-swap-config. Retention 14. **Für echtes 1:1 zu erweitern** — fertiges Snippet in **Anhang B** (§5). |
|
||||
| Restore | `deploy/restore.sh` | Spielt Tarball zurück (mit Pre-Restore-Sicherung). |
|
||||
| Engine | `deploy/provision-engine.sh` | llama.cpp Vulkan/RADV-Build (root). |
|
||||
| Voice-Setup | `voice_service/install.sh` | venv + Piper-Binary + dt. Piper-Stimmen + Chatterbox (best-effort). |
|
||||
| Updates | `deploy/update-engine.sh`, `update-swap.sh` | Laufende Engine/Router-Updates. |
|
||||
| Postchecks | `deploy/stack-postcheck.sh`, `hermes-postcheck.sh` | Funktionsprüfung → Wizard-Verifikationsschritt. |
|
||||
|
||||
---
|
||||
|
||||
## 3. VOLLSTÄNDIGER Komponenten-Katalog
|
||||
|
||||
> **Scan-verifiziert (2026-06-28)** gegen `backend/config.py`, alle `backend/services/*` & `routers/*`,
|
||||
> `frontend/src/{nav.ts,views,components/voice}`, `voice_service/app.py`, `mem0_service/`, `deploy/*`.
|
||||
> Alle persistenten Schreibpfade, Dienste, Secrets und Env-Vars sind unten erfasst.
|
||||
|
||||
Legende Restore-Quelle: 📦 = aus Backup-Tarball · ⬇️ = Re-Download/Install (Skript) · 🌐 = Git ·
|
||||
🔑 = Secret (Nutzer/Generieren) · 🆕 = im Wizard neu gewählt.
|
||||
„Im Backup?" = deckt der **aktuelle** `backup.sh` es ab.
|
||||
|
||||
### A · OS & System (L0, root/sudo, einmalig)
|
||||
| Komponente | Ort | Quelle | Im Backup? |
|
||||
|---|---|---|---|
|
||||
| Ubuntu 26.04, User `hitonabi`, `enable-linger` | — | ⬇️ manuell | n/a |
|
||||
| System-Pakete: python3.14+venv, git, curl, jq, ttyd, uv | — | ⬇️ apt/curl | nein |
|
||||
| Vulkan-Stack: mesa-vulkan-drivers, libvulkan1, vulkan-tools | — | ⬇️ apt | nein |
|
||||
| Chrome-Libs (Browser): `agent-browser install --with-deps` | — | ⬇️ apt | nein |
|
||||
| Verzeichnis-Layout: `/srv/models/{,mem0,mc2-backups,drafts}`, `/etc/llama-swap/` | — | ⬇️ mkdir+chown | nein |
|
||||
|
||||
### B · Inferenz-Engine + Router (L1, root)
|
||||
| Komponente | Ort | Quelle | Im Backup? |
|
||||
|---|---|---|---|
|
||||
| llama.cpp (Vulkan) → `llama-server` | `/opt/llamacpp-vulkan`, `/usr/local/bin/llama-server` | ⬇️ provision-engine.sh / update-engine.sh (ggml-org Release) | nein (re-build) |
|
||||
| **llama-swap** Binary (Router `:8080`) | `/usr/local/bin/llama-swap` | ⬇️ **bekannt:** `mostlygeek/llama-swap`-Release (Rezept in `update-swap.sh`) | nein (re-download) |
|
||||
| **llama-swap systemd-System-Unit** | `/etc/systemd/system/llama-swap.service` (+ `.d/` Drop-ins) | ⬇️ Bootstrap legt sie an — **kompletter Unit-Inhalt in Anhang A** (von der Box abgegriffen; Drop-ins via provision-engine.sh) | nein |
|
||||
| llama-swap-Config (Modelle/Rollen) | `/etc/llama-swap/config.yaml` | 📦 | **ja** |
|
||||
| Draft-Modelle (Spec-Decoding) | `/srv/models/drafts/` | ⬇️ Re-Download | nein |
|
||||
|
||||
### C · MC2-App (L2, userspace)
|
||||
| Komponente | Ort | Quelle | Im Backup? |
|
||||
|---|---|---|---|
|
||||
| MC2-Code | `~/mission-control-v2` | 🌐 Git (Gitea) | n/a |
|
||||
| Backend-venv (Python 3.14) | `backend/.venv` | ⬇️ deploy.sh | nein (rebuild) |
|
||||
| Frontend (gebaut, inkl. **3D-Avatar `avatar.vrm`**) | `frontend/dist`, `frontend/public/avatar.vrm` | 🌐 Git | n/a |
|
||||
| systemd-User-Dienst `mission-control-2` (`:9001`) | `~/.config/systemd/user/` | ⬇️ deploy.sh | nein |
|
||||
|
||||
### D · 3D-Avatar (Sprechen-Tab)
|
||||
| Komponente | Ort | Quelle | Im Backup? |
|
||||
|---|---|---|---|
|
||||
| Default-Avatar (VRM) | `frontend/public/avatar.vrm` (→ dist) | 🌐 Git | n/a |
|
||||
| Renderer | `frontend/src/components/voice/Avatar3D.tsx` | 🌐 Git | n/a |
|
||||
| **Eigene/zusätzliche Avatare** (falls Nutzer welche ablegt) | **TODO: Ablageort definieren** (z.B. `/srv/models/avatars/` + DB-Verweis) | 📦/🆕 | **nein (Lücke)** |
|
||||
|
||||
> Heute ist der Avatar **fest** (`avatar.vrm`, kommt mit dem Git-Frontend zurück). Wenn künftig
|
||||
> Nutzer-Avatare hochgeladen werden, brauchen sie einen persistenten Ablageort, der ins Backup geht.
|
||||
|
||||
### E · Voice-Sidecar (STT + TTS + Klonstimme)
|
||||
| Komponente | Ort | Quelle | Im Backup? |
|
||||
|---|---|---|---|
|
||||
| Voice-venv (Python 3.12) | `~/.voice/venv` | ⬇️ install.sh | nein (rebuild) |
|
||||
| STT (faster-whisper Modell) | `~/.voice/` (Cache) | ⬇️ install.sh / 1. Start | nein |
|
||||
| Piper-Binary + dt. Stimmen (thorsten/kerstin) | `~/.voice/piper`, `~/.voice/voices` | ⬇️ install.sh | nein |
|
||||
| Chatterbox (Premium-TTS, CPU-torch) | venv | ⬇️ install.sh (best-effort) | nein |
|
||||
| **Klonstimme / Voice-Referenz-Audio** (Nutzer) | `~/.voice/refs/ref.wav` (`VOICE_REFS_DIR`, `/api/voice/reference`) | 📦 | **nein → Anhang B** |
|
||||
| ElevenLabs-Key | `~/.hermes/.env` | 🔑 | ja |
|
||||
| Stimm-/Lautstärke-Wahl | Browser-`localStorage` (pro Gerät) | 🆕 client | n/a |
|
||||
| Dienst `voice-service` (`:8650`) | systemd-User | ⬇️ deploy.sh | nein |
|
||||
|
||||
### F · Mem0 (Langzeitgedächtnis)
|
||||
| Komponente | Ort | Quelle | Im Backup? |
|
||||
|---|---|---|---|
|
||||
| Mem0-venv (Python 3.12, uv) | `~/.mem0/venv` | ⬇️ deploy.sh | nein (rebuild) |
|
||||
| **Gedächtnis-Daten (Chroma + history.db)** | `/srv/models/mem0/` | 📦 | **ja** |
|
||||
| Hermes-Memory-Plugin | `~/.hermes/plugins/mc2-memory/` | 📦/🌐 | ja |
|
||||
| Dienst `mem0-service` (`:8765`) | systemd-User | ⬇️ deploy.sh | nein |
|
||||
|
||||
### G · Hermes-Agent (das Herz)
|
||||
| Komponente | Ort | Quelle | Im Backup? |
|
||||
|---|---|---|---|
|
||||
| Hermes-Code | `~/.hermes/hermes-agent` | 🌐 Git (NousResearch) | nein (re-clone) |
|
||||
| Bundled Node v22 | `~/.hermes/node/bin` | ⬇️ (kommt mit Hermes) | nein |
|
||||
| Hermes-venv | `~/.hermes/hermes-agent/venv` | ⬇️ rebuild | nein |
|
||||
| **`config.yaml` 1:1** (Toolsets, MCP, Engine, Browser, Personalities, alles) | `~/.hermes/config.yaml` | 📦 | **ja** |
|
||||
| **`.env` (Secrets: TELEGRAM_BOT_TOKEN, API_SERVER_KEY, ElevenLabs …)** | `~/.hermes/.env` | 📦/🔑 | **ja** |
|
||||
| Plugins | `~/.hermes/plugins/` | 📦 | ja |
|
||||
| **Skills (eigene)** | `~/.hermes/skills/` (45M; `.hub`-Cache re-downloadbar) | 📦 | **nein → Anhang B (ohne .hub)** |
|
||||
| Sessions/History (optional) | `~/.hermes/sessions/` (856K) | 📦 | **nein → Anhang B** |
|
||||
| Checkpoints (optional) | `~/.hermes/checkpoints/` (existiert nicht) | ⏭️ skip | nein |
|
||||
| **Token-/Ersparnis-Statistik** | `~/.hermes/token_stats.json` | 📦 | **nein → Anhang B** |
|
||||
| Dienst `hermes-gateway` (`:8642`) | systemd-User | ⬇️ | nein |
|
||||
| Hermes-Terminal (ttyd → `hermes chat`, `:7681`) | systemd-User `hermes-terminal` | ⬇️ deploy.sh (+ttyd apt) | nein |
|
||||
|
||||
### H · MCP-Server (4)
|
||||
| Komponente | Ort | Quelle | Im Backup? |
|
||||
|---|---|---|---|
|
||||
| `mission-control-memory`, `-stack` (MC2-venv) | `~/mission-control-v2/mcp/*.py` | 🌐 Git | über config.yaml |
|
||||
| `hermes-pc-control`, `hermes-web-fetch` (Hermes-venv) | dito | 🌐 Git | über config.yaml |
|
||||
| MCP-Verdrahtung | `~/.hermes/config.yaml → mcp_servers` | 📦 | ja |
|
||||
|
||||
### I · Browser-Stack
|
||||
| Komponente | Ort | Quelle | Im Backup? |
|
||||
|---|---|---|---|
|
||||
| agent-browser (npm global) | `~/.local/...`, symlink `~/.hermes/node/bin` | ⬇️ npm i -g | nein |
|
||||
| Chrome (engine) + System-Libs | `~/.agent-browser/browsers/` + apt-Libs | ⬇️ install --with-deps | nein |
|
||||
| lightpanda (optional, leicht) | `~/.local/bin/lightpanda` | ⬇️ Download | nein |
|
||||
| Browser-Verdrahtung (engine=chrome, toolset) | `~/.hermes/config.yaml` | 📦 | ja |
|
||||
|
||||
### J · Modelle (GGUF — der große Brocken)
|
||||
| Komponente | Ort | Quelle | Im Backup? |
|
||||
|---|---|---|---|
|
||||
| GGUF-Modelle (Brain, fast, heavy, coder, vision, embedding …) | `/srv/models/` | ⬇️ Re-Download aus llama-swap-Manifest | **nein (zu groß, bewusst)** |
|
||||
|
||||
### K · Extern (anderer Rechner)
|
||||
| Komponente | Ort | Quelle | Im Backup? |
|
||||
|---|---|---|---|
|
||||
| PC-Executor (Windows) | `client/hermes-pc/` → Scheduled Task | 🌐 Git + Task | n/a (anderer Host) |
|
||||
| Hermes-WebUI (nesquena, optional, `:8787`) | `~/hermes-webui` | ⬇️ git clone | nein |
|
||||
|
||||
---
|
||||
|
||||
## 4. Ziel-Architektur
|
||||
|
||||
### 4.1 `deploy/bootstrap.sh` — orchestrierter From-Zero-Lauf
|
||||
- **Idempotent & resumierbar** (Schritt-Marker in `~/.mc2-bootstrap.state`).
|
||||
- **Getrennt nach sudo-Bedarf**: `bootstrap-root.sh` (L0+L1, bewusst mit sudo) + `bootstrap.sh`
|
||||
(L2+, sudo-frei = Kern ist `deploy.sh`). Passt zum Nordstern „Runtime ohne sudo".
|
||||
- **Mündet in den Wizard**: startet MC2 im First-Run-Modus, gibt Wizard-URL aus.
|
||||
|
||||
Phasen: `0 Vorflug → 1[sudo] System+Dirs → 2[sudo] Engine+llama-swap → 3 Code(MC2+Hermes+Node)
|
||||
→ 4 venvs(backend/mem0/voice/hermes) → 5 Browser → 6 deploy.sh(Units/enable/restart)
|
||||
→ 7 Restore(optional) → 8 Wizard hoch + postcheck`.
|
||||
|
||||
### 4.2 First-Run-Wizard (browserbasiert, von MC2 serviert)
|
||||
MC2 erkennt unkonfigurierten Zustand → Frontend-Route `/setup` statt Dashboard. Schritte:
|
||||
|
||||
1. **Systemcheck** — Live-Ampel je Komponente aus §3 (`GET /api/setup/status` + `…/system/services`).
|
||||
2. **Restore-Quelle wählen** — Backup-Tarball erkennen → globaler Modus „Alles 1:1" / „selektiv" / „frisch".
|
||||
3. **Komponenten-Auswahl (Kernstück)** — pro Katalog-Eintrag aus §3 eine Kachel mit
|
||||
📦/🆕/⏭️ (siehe §1). Zeigt Größe + ob Backup-Daten vorhanden.
|
||||
4. **Netzwerk & Identität** — Box-IP (→ `HERMES_TERMINAL_URL`, `PC_EXECUTOR_URL`), Hostname.
|
||||
5. **Secrets** 🔑 — `TELEGRAM_BOT_TOKEN`, erlaubte User-ID, `API_SERVER_KEY` (oder generieren),
|
||||
ElevenLabs-Key → `~/.hermes/.env` (chmod 600). Bei 📦 vorbefüllt aus Backup.
|
||||
6. **Hermes-Profil** — Brain-Alias + Toolset-Profil (Default = entfesselt: vision/tts/memory/browser
|
||||
an, image_gen/video aus — siehe [[project-hermes-setup]]). Bei 📦 = exakte alte `config.yaml`.
|
||||
7. **Avatar & Voice** — Avatar wählen (Default-VRM oder eigener), Stimme/Klonstimme
|
||||
(📦 Referenz-Audio zurück, oder neu aufnehmen/hochladen).
|
||||
8. **Modelle** — aus llama-swap-Manifest automatisch nachladen (`POST /api/models/install`,
|
||||
Fortschritt `GET /api/jobs`) ODER geführte Discover-Neuauswahl. Reihenfolge: Brain → heavy → Rest.
|
||||
9. **Externe Checkliste** — PC-Executor (Windows-Task), Hermes-WebUI als Haken.
|
||||
10. **Verifikation & Abschluss** — `stack-postcheck.sh` + `hermes-postcheck.sh` → grün/rot-Liste → Dashboard.
|
||||
|
||||
**Technik:** neuer Router `backend/routers/setup.py`
|
||||
(`GET /api/setup/status`, `POST /api/setup/{secrets,network,hermes,components,finish}`).
|
||||
First-Run-Gate: MC2 prüft beim Boot Marker `~/.mc2-setup-done` bzw. Pflicht-Secrets → leitet auf `/setup`.
|
||||
|
||||
---
|
||||
|
||||
## 5. Backup-Scope für echtes 1:1 ERWEITERN (umzusetzen — Snippet in Anhang B)
|
||||
|
||||
Der **aktuelle** `backup.sh` (live auf der Box, unverändert) reicht für „1:1" nicht. Für die Bau-Session
|
||||
liegt das **fertige Erweiterungs-Snippet in Anhang B** — es ergänzt den Tarball um:
|
||||
|
||||
- **Voice-Klonstimme** → `~/.voice/refs/` (`ref.wav`)
|
||||
- **Token-/Ersparnis-Statistik** → `~/.hermes/token_stats.json`
|
||||
- **Hermes-Skills** → `~/.hermes/skills/` (re-downloadbarer `.hub`-Cache per `tar --exclude` ausgelassen)
|
||||
- **Hermes-Sessions/History** → `~/.hermes/sessions/`
|
||||
|
||||
Restore soll skills/sessions **mergen** (frisch geladenen `.hub` nicht überschreiben) und `voice-service`
|
||||
mit neu starten.
|
||||
|
||||
**Offen bleibt:** eigene Avatare (sobald Upload existiert — Ablageort heute undefiniert, siehe §3·D).
|
||||
|
||||
**Bewusst NICHT im Backup (re-downloadbar/rebuildbar):** Piper-Stimmen (install.sh), `.hub`-Skill-Cache,
|
||||
`/srv/models/mc2-discover.json` (Cache), `/srv/models/mc2-memory.db` (Legacy-Migration), alle venvs, GGUF-Modelle.
|
||||
|
||||
→ Aufgabe der Bau-Session: `backup.sh` + `restore.sh` um diese Pfade erweitern (mit klarer
|
||||
„opt-in für große/optionale Teile"-Logik), und den MANIFEST-Inhalt entsprechend.
|
||||
|
||||
---
|
||||
|
||||
## 6. Secrets-Strategie
|
||||
- Single Source: `~/.hermes/.env` (chmod 600), im Tarball (selbst 600).
|
||||
- Rebuild ohne Backup → Wizard-Schritt 5 erzeugt sie. `API_SERVER_KEY` generierbar; Telegram/ElevenLabs liefert Nutzer.
|
||||
- Nie ins Git, nie in Logs, nie in die MC2-DB. **Off-Box-Spiegelung des Tarballs noch offen** ([[mc2-backup-restore]]).
|
||||
|
||||
## 7. Modell-Strategie
|
||||
- **A (empfohlen):** `/etc/llama-swap/config.yaml` (im Backup) listet alle Modelle → Wizard leitet Repos/Quants ab und lädt automatisch (`/api/models/install`). „Ein Klick, lädt über Nacht."
|
||||
- **B:** geführte Discover-Neuauswahl je Rolle. Reihenfolge Brain → heavy → Rest, danach `warmup.sh`.
|
||||
|
||||
---
|
||||
|
||||
## 8. Implementierungs-Reihenfolge (neue Session)
|
||||
1. ✅ **Box-Verifikation erledigt (2026-06-28, nur gelesen — nichts verändert):** llama-swap.service-Inhalt
|
||||
in **Anhang A**; skills/sessions/refs/token_stats existieren (Pfade in §3 verifiziert); Backup-Erweiterung
|
||||
als fertiges Snippet in **Anhang B**. **Keine offenen Box-Fragen mehr** außer §10·3–6. → mit Schritt 2 starten.
|
||||
2. `backup.sh`/`restore.sh` um die 1:1-Lücken erweitern (§5).
|
||||
3. `deploy/bootstrap.sh` + `bootstrap-root.sh` (Phasen, State-Marker); llama-swap-Install skripten.
|
||||
4. `backend/routers/setup.py` + First-Run-Gate.
|
||||
5. Frontend `/setup`-Wizard (Schritte 1–10), Komponenten-Auswahl als Kernstück; bestehende Views
|
||||
(Cockpit/AgentView/Sprechen) wiederverwenden.
|
||||
6. Modell-Restore-aus-Manifest.
|
||||
7. **Abnahme in frischer VM** (§9).
|
||||
|
||||
Vorschlag: 2–4 zuerst (Skelett lauffähig), Wizard danach. Branch + PR.
|
||||
|
||||
## 9. Abnahmekriterien
|
||||
- Frisches Ubuntu 26.04 → `bootstrap-root.sh` + `bootstrap.sh` → laufender Stack, kein Spezialwissen.
|
||||
- Postchecks grün; Telegram, Voice (inkl. Klonstimme bei 📦), 3D-Avatar, Browser funktionieren.
|
||||
- „Alles 1:1" stellt mem0, Hermes-config.yaml, Secrets, Klonstimme, Skills exakt wieder her.
|
||||
- Selektiver Modus: einzelne Komponenten ⏭️ überspringbar, Stack läuft trotzdem.
|
||||
- Idempotenz: zweiter Lauf = No-op.
|
||||
|
||||
## 10. Offene Entscheidungen (vor dem Bau)
|
||||
1. ~~llama-swap systemd-Unit-Inhalt~~ → **geklärt: kompletter Unit in Anhang A** (in der Bau-Session anlegen).
|
||||
2. ~~Voice-Referenz-Audio-Ort~~ → **geklärt: `~/.voice/refs/`** (Backup-Snippet in Anhang B, in der Bau-Session umsetzen).
|
||||
3. **Eigene Avatare**: künftiger Upload-/Ablage-Mechanismus + Backup-Pfad (einziger offener 1:1-Daten-Punkt).
|
||||
4. **Off-Box-Backup-Ziel** (NAS/2. Platte/Cloud) — [[mc2-backup-restore]].
|
||||
5. **Python 3.14** auf frischem Ubuntu beschaffen (deadsnakes?).
|
||||
6. **Bootstrap-sudo-Modell**: getrenntes `bootstrap-root.sh` (empfohlen) vs. interaktives sudo.
|
||||
|
||||
## 11. Referenzen
|
||||
- `backend/config.py`, [[project-mc2-architecture]], [[project-stack-state]]
|
||||
- [[project-hermes-setup]], `docs/HERMES_SETUP.md` (entfesseltes Profil, Browser, PC-Executor)
|
||||
- [[voice-sprechen-feature]] (Voice + 3D-Avatar), [[mem0-memory-architecture]], [[mc2-backup-restore]]
|
||||
- `deploy/{deploy,backup,restore,provision-engine,stack-postcheck,warmup}.sh`, `voice_service/install.sh`
|
||||
|
||||
---
|
||||
|
||||
## Anhang A — `llama-swap.service` (1:1 von der Box abgegriffen, 2026-06-28)
|
||||
|
||||
Base-Unit nach `/etc/systemd/system/llama-swap.service` (root). User/GFX sind boxspezifisch
|
||||
(hitonabi, gfx1151 → HSA 11.5.1) — auf anderer Hardware anpassen. Die `.d/`-Drop-ins
|
||||
(`vulkan.conf`, `warmup.conf`) legt `provision-engine.sh` an.
|
||||
|
||||
```ini
|
||||
[Unit]
|
||||
Description=llama-swap (lokaler LLM Router)
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
User=hitonabi
|
||||
Environment=HSA_OVERRIDE_GFX_VERSION=11.5.1
|
||||
Environment=PATH=/usr/local/bin:/usr/bin:/bin
|
||||
ExecStart=/usr/local/bin/llama-swap --config /etc/llama-swap/config.yaml --listen 0.0.0.0:8080 --watch-config
|
||||
Restart=on-failure
|
||||
RestartSec=3
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
```
|
||||
```ini
|
||||
# /etc/systemd/system/llama-swap.service.d/vulkan.conf
|
||||
[Service]
|
||||
Environment=LD_LIBRARY_PATH=/opt/llamacpp-vulkan
|
||||
```
|
||||
```ini
|
||||
# /etc/systemd/system/llama-swap.service.d/warmup.conf
|
||||
[Service]
|
||||
ExecStartPost=-/usr/local/bin/llama-swap-warmup.sh
|
||||
```
|
||||
|
||||
**Erstinstall-Reihenfolge:** `bash deploy/update-swap.sh` (Binary) → Unit kopieren →
|
||||
`sudo bash deploy/provision-engine.sh` (Engine+Drop-ins) → `sudo systemctl enable --now llama-swap`.
|
||||
|
||||
---
|
||||
|
||||
## Anhang B — Backup-Scope-Erweiterung für echtes 1:1 (fertiges Snippet)
|
||||
|
||||
In `deploy/backup.sh` ergänzen (Variablen oben: `VOICE="${VOICE_HOME:-$HOME/.voice}"`,
|
||||
Stage zusätzlich `"$STAGE/voice"`):
|
||||
|
||||
```bash
|
||||
# 1:1-Erweiterung (scan-verifiziert 2026-06-28):
|
||||
[ -f "$HERMES/token_stats.json" ] && cp -a "$HERMES/token_stats.json" "$STAGE/hermes/" || true
|
||||
[ -d "$HERMES/skills" ] && cp -a "$HERMES/skills" "$STAGE/hermes/skills" || true
|
||||
[ -d "$HERMES/sessions" ] && cp -a "$HERMES/sessions" "$STAGE/hermes/sessions" || true
|
||||
[ -d "$VOICE/refs" ] && cp -a "$VOICE/refs" "$STAGE/voice/refs" || true
|
||||
```
|
||||
tar-Aufruf um den re-downloadbaren Skill-Cache erleichtern:
|
||||
```bash
|
||||
tar --exclude='./hermes/skills/.hub' -czf "$OUT" -C "$STAGE" .
|
||||
```
|
||||
In `deploy/restore.sh` spiegelbildlich (skills/sessions **mergen**, nicht ersetzen) und
|
||||
`voice-service` mit neustarten:
|
||||
```bash
|
||||
VOICE="${VOICE_HOME:-$HOME/.voice}"
|
||||
SERVICES="mem0-service voice-service mission-control-2 hermes-gateway"
|
||||
[ -f "$STAGE/hermes/token_stats.json" ] && cp -a "$STAGE/hermes/token_stats.json" "$HERMES/"
|
||||
[ -d "$STAGE/hermes/skills" ] && { mkdir -p "$HERMES/skills"; cp -a "$STAGE/hermes/skills/." "$HERMES/skills/"; }
|
||||
[ -d "$STAGE/hermes/sessions" ] && { mkdir -p "$HERMES/sessions"; cp -a "$STAGE/hermes/sessions/." "$HERMES/sessions/"; }
|
||||
[ -d "$STAGE/voice/refs" ] && { mkdir -p "$VOICE/refs"; cp -a "$STAGE/voice/refs/." "$VOICE/refs/"; }
|
||||
```
|
||||
@@ -0,0 +1,82 @@
|
||||
# Hermes-Schicht — Box-Runbook
|
||||
|
||||
> Diese Schritte laufen **auf der Bosgame** (`192.168.178.151`, User `hitonabi`).
|
||||
> MC betreibt Hermes nicht — es zeigt nur Status + verlinkt das WebUI. Hier wird die
|
||||
> eigentliche **volle Verdrahtung** gemacht (das war in v1 der „Hermes ist dumm"-Grund).
|
||||
|
||||
## Reihenfolge der Dienste
|
||||
`llama-swap (:8080)` → **builtin Gateway (`:9001/v1`, Teil von MC2)** → `hermes-gateway (:8642)` →
|
||||
`hermes-webui (:8787, nesquena)`
|
||||
|
||||
## 1. Gateway = builtin (kein LiteLLM)
|
||||
LiteLLM scheitert auf Python 3.14 (uvloop/orjson). MC2 bringt einen **eingebauten** OpenAI-kompatiblen
|
||||
Gateway auf `:9001/v1` mit: `model: auto` (kurz→`fast`, komplex→`heavy`) + explizite Aliase
|
||||
(`fast`/`heavy`/`coder`/`vision`/`hermes`). Verifizieren:
|
||||
```bash
|
||||
curl -s http://127.0.0.1:9001/v1/models
|
||||
```
|
||||
|
||||
## 2. Engine: Rollen (llama-swap)
|
||||
Aliase sauber: `fast` (Qwen3.6-35B-A3B), `heavy` (Qwen3.5-122B-A10B), `coder`, `vision`, `scout`,
|
||||
`hermes` (Hermes-4-14B, `ttl 99999` = immer warm). Verwaltung im Cockpit (Modelle & Routing).
|
||||
**Ko-Residenz** (optional): Gruppe `swap:false` für `hermes`+`fast` → beide warm; `heavy`/`vision`
|
||||
on-demand. GTT beobachten (~124 GB Limit).
|
||||
|
||||
## 3. hermes-webui installieren (nesquena, standalone)
|
||||
```bash
|
||||
cd ~ && git clone https://github.com/nesquena/hermes-webui && cd hermes-webui
|
||||
python3 bootstrap.py # erkennt hermes-agent, baut venv, installiert Deps
|
||||
mkdir -p ~/.config/environment.d
|
||||
echo 'HERMES_WEBUI_PASSWORD=<dein-passwort>' > ~/.config/environment.d/hermes-webui.conf
|
||||
# Unit deploy/hermes-webui.service → ~/.config/systemd/user/ (HOST=0.0.0.0, PORT=8787)
|
||||
systemctl --user enable --now hermes-webui
|
||||
loginctl enable-linger hitonabi
|
||||
```
|
||||
Zugriff vom Windows-PC: `http://192.168.178.151:8787` (mit Passwort). MC2-Unit
|
||||
`HERMES_WEBUI_URL=http://192.168.178.151:8787` setzen → „Hermes öffnen" zeigt darauf.
|
||||
**Danach das alte offizielle Dashboard stilllegen** (eine WebUI):
|
||||
`systemctl --user disable --now hermes-dashboard` (:9119).
|
||||
|
||||
## 4. Hermes-Hirn = dediziertes Hermes-4-14B + Delegation (NICHT model:auto)
|
||||
In `~/.hermes/config.yaml`:
|
||||
```yaml
|
||||
model:
|
||||
default: Hermes-4-14B
|
||||
provider: custom
|
||||
base_url: http://127.0.0.1:9001/v1
|
||||
api_key: local
|
||||
model: hermes # Hermes' eigenes Hirn (Alias→Hermes-4-14B), NICHT 'auto'
|
||||
delegation:
|
||||
model: heavy # harte Teilaufgaben → Qwen3.5-122B
|
||||
provider: custom
|
||||
base_url: http://127.0.0.1:9001/v1
|
||||
api_key: local
|
||||
orchestrator_enabled: true
|
||||
subagent_auto_approve: true
|
||||
```
|
||||
`model:auto` bleibt ausschließlich Gateway-Funktion für Vibe Coding/IDEs.
|
||||
|
||||
## 5. Tools/MCP verdrahten — UND kaputte Tools abschalten (Thrash-Fix!)
|
||||
- **Kaputte/Cloud-Tools global abschalten** (sonst Endlos-Loops, siehe Memory `hermes-thrash-rootcause-fix`):
|
||||
```yaml
|
||||
agent:
|
||||
disabled_toolsets: [browser, vision, computer_use, image_gen, tts, video, video_gen, memory]
|
||||
memory:
|
||||
memory_enabled: false
|
||||
user_profile_enabled: false
|
||||
```
|
||||
⚠️ `hermes tools disable <x>` wirkt nur für die cli-Plattform, **nicht den api_server** — nutze
|
||||
`agent.disabled_toolsets` (gilt für ALLE Plattformen).
|
||||
- **Behalten:** terminal/shell, file, code_execution, skills, todo, session_search, clarify,
|
||||
delegation, cronjob, web. `approvals: auto` (rein lokal).
|
||||
- **MCP-Server** (`mcp_servers` in config) — laufen über die v2-venv (hat das `mcp`-Modul nach pip install):
|
||||
- geteiltes Gedächtnis: `~/mission-control-v2/backend/.venv/bin/python ~/mission-control-v2/mcp/mcp_memory.py` (Env `MC_URL=http://127.0.0.1:9001`)
|
||||
- Stack-Management: `~/mission-control-v2/backend/.venv/bin/python ~/mission-control-v2/mcp/mcp_mc.py` (Env `MC_URL=http://127.0.0.1:9001`)
|
||||
- **SSH→Windows-PC** (voller Zugriff): OpenSSH-Server auf Windows aktiv + Key
|
||||
`~/.ssh/id_ed25519_hermes_agent` autorisiert; Hermes nutzt sein terminal-Tool für `ssh TobisPC@<win-ip>`.
|
||||
|
||||
## 6. Verifikation
|
||||
- `curl :8642/v1/chat/completions` „bist du da?" → kurze Antwort in wenigen Sekunden, **kein** Tool-Loop
|
||||
im `journalctl --user -u hermes-gateway`; llama-swap `/running` zeigt `Hermes-4-14B`.
|
||||
- WebUI öffnet vom Windows-PC, Chat antwortet.
|
||||
- Hermes kann via `mcp_mc` Modelle listen/Routing ändern; erreicht (nach §5-SSH) den Windows-PC.
|
||||
@@ -0,0 +1,149 @@
|
||||
# Lucy TTS — Optimierungs- & Strategieplan
|
||||
|
||||
> Ziel: **Lucys Stimme läuft 100 % lokal & on-device** (kein Box-Zwang, keine Cloud).
|
||||
> Primärengine: **Kyutai pocket-tts 2.1.0** (CPU). Strategische Alternative: **F5-TTS** (GPU/ONNX).
|
||||
> Stand: 2026-06-30. Resume-fähig im Stil von `docs/STATUS.md`.
|
||||
|
||||
## Leitentscheidung: Hardware
|
||||
- **pocket-tts ist CPU-gebunden** — Kyutai bestätigt: *kein* GPU-Speedup (Batch 1, 100M Params).
|
||||
→ **RDNA2 vs. RDNA4 ist für pocket irrelevant.** Lucy läuft dort, wo der beste CPU steht.
|
||||
- **GPU zählt nur für F5-TTS** (non-autoregressiv, große Matmuls). Dort gewinnt **RDNA4 (9070 XT) + DirectML**
|
||||
deutlich gegen RDNA2 → falls F5 die Lucy-Stimme wird, läuft sie auf der RDNA4-Maschine.
|
||||
- **Konsequenz:** `oute n_gpu_layers=999` und „pocket auf ROCm/Vulkan" werden **eingestellt**.
|
||||
|
||||
---
|
||||
|
||||
## Phase A — Sofort-Wins (risikoarm, pocket_server.py)
|
||||
|
||||
**A1 — Voice → safetensors exportieren (einmalig)**
|
||||
- Statt `get_state_for_audio_prompt(ref)` bei jedem Boot: einmal `export_model_state(...)` → `lucy_voice.safetensors`.
|
||||
- Server lädt beim Start nur noch die safetensors (liest kvcache, keine Klon-Rechnung) → schnellerer Start.
|
||||
- Fallback behalten: wenn safetensors fehlt → aus `ref.mp3` klonen + direkt exportieren.
|
||||
|
||||
**A2 — Referenz säubern**
|
||||
- Kyutai: Sample-Qualität wird *mitreproduziert*. Sauber entrauschte/normalisierte `ref` → weniger Artefakte.
|
||||
- Schritt in `_prep_ref()` ergänzen (Denoise/Highpass/Lautheit), Ergebnis cachen.
|
||||
|
||||
**A3 — temp-Sweep gegen Kollaps**
|
||||
- Hypothese: `temp=0.9` treibt die best-of-N-Regenerationen. Niedriger testen (0.7–0.85).
|
||||
- Akzeptanz: gleiche/bessere Natürlichkeit bei **messbar weniger Regenerationen** (= weniger Latenz).
|
||||
|
||||
*Gate A:* A1–A3 live, Kollaps-Rate & TTFB protokolliert.
|
||||
|
||||
---
|
||||
|
||||
## Phase B — Benchmark-Harness (datenbasiert tunen)
|
||||
|
||||
**B1 — `bench_lucy.py`** (lokal, CPU)
|
||||
- Misst je Konfig: **RTF**, **TTFB**, **Kollaps-/Regenerations-Rate**, Whisper-Rücktranskription (Verständlichkeit).
|
||||
- Sweep-Achsen: `lsd_decode_steps` (6/8/10/12), `temp`, `noise_clamp`, `quantize` on/off, `frames_after_eos`.
|
||||
- Feste Testsätze (kurz/mittel/lang, wie in `ptts_test.py`).
|
||||
|
||||
**B2 — CPU-Parallelität**
|
||||
- pocket nutzt nur **2 Kerne** → mehrere **Satz-Worker parallel** statt globalem `LOCK`.
|
||||
- Messen: Wall-Clock langer Antworten bei N Workern (1/2/3/4) vs. Qualität/Last.
|
||||
|
||||
*Gate B:* dokumentierte Best-Config (RTF + Kollaps-Rate) als neue Defaults in `pocket_server.py`.
|
||||
|
||||
---
|
||||
|
||||
## Phase C — Runtime-Eval (lokal schneller + Python-3.14-frei)
|
||||
|
||||
**C1 — ONNX-Pfade testen**
|
||||
- Kandidaten: **PocketTTS.cpp** (Single-File C++/ONNX, CLI+HTTP+FFI) und **sherpa-onnx** (Windows, viele Bindings).
|
||||
- Ziel: torch-CPU schlagen **und** das Python-3.14-Packaging-Problem umgehen.
|
||||
- Vergleich gegen Phase-B-Baseline (gleicher Benchmark).
|
||||
|
||||
**C2 — Server-Vertrag prüfen**
|
||||
- Fertige **OpenAI-kompatible Streaming-Server** (teddybear082 / ai-joe-git) gegen den aktuellen Custom-Proxy halten.
|
||||
- Nur übernehmen, wenn Stimm-Wächter (F0 + Fingerabdruck) erhalten/abbildbar bleiben.
|
||||
|
||||
*Gate C:* Entscheidung „torch-CPU behalten" vs. „auf ONNX-Runtime wechseln" (mit Zahlen).
|
||||
|
||||
---
|
||||
|
||||
## Phase D — Strategische Weiche: pocket vs. F5
|
||||
|
||||
**D1 — F5 fair gegen pocket messen**
|
||||
- `lucy-f5/` (F5-TTS DE, ONNX + DirectML, **RDNA4/9070 XT**) ist **non-autoregressiv → strukturell kein
|
||||
Kollaps/Männerstimme/Wiederholung** (genau pockets Schmerz).
|
||||
- Gleicher Benchmark wie Phase B: RTF, TTFB, Natürlichkeit, Stabilität.
|
||||
|
||||
**D2 — Entscheidung**
|
||||
- **pocket gewinnt** (gut genug stabil, CPU, „nur lokal"-Ideal): pocket = Lucy-Stimme, F5 verworfen/geparkt.
|
||||
- **F5 gewinnt** (Stabilität schlägt den Latenz-Overhead der Wächter): F5 = Lucy-Stimme auf RDNA4,
|
||||
**pocket bleibt schneller CPU-Fallback** (z. B. wenn keine GPU verfügbar).
|
||||
|
||||
**Beobachtungsposten (extern):** distilliertes **`german`** (statt `german_24l`) ist noch nicht released
|
||||
(Kyutai fixt Distillations-Datenqualität). Sobald da → größter Einzel-Win (Tempo + Stabilität),
|
||||
dann `german_24l → german` swappen. Release-Feed im Blick behalten.
|
||||
|
||||
---
|
||||
|
||||
## Reihenfolge & Quick-Start
|
||||
1. **A1 + A2 + A3** (heute) — sofort spürbar, kein Risiko.
|
||||
2. **B1 + B2** — Zahlen sammeln, Defaults härten.
|
||||
3. **C1/C2** — Runtime-Wechsel nur wenn Benchmark es trägt.
|
||||
4. **D1/D2** — finale Stimm-Architektur entscheiden.
|
||||
|
||||
## Mess-Ergebnisse & finale Defaults (30.06., 9700X, german_24l)
|
||||
|
||||
Alles auf der lokalen Maschine gemessen (Logs: `sweep_b2.log`, `ttfb_test.log`, `sweep_b2_result.json`).
|
||||
|
||||
**B2-Sweep (6-Satz-Antwort, ~21s Audio):**
|
||||
|
||||
| Konfig | Wall | TTFB | RTF |
|
||||
|---|---|---|---|
|
||||
| seriell (1×alle Kerne) | 15,8s | **3,6s** | 0,74 |
|
||||
| 2w×3t | 12,9s | 6,2s | 0,62 |
|
||||
| 3w×2t | 10,3s | 6,8s | 0,50 |
|
||||
| 4w×2t | 9,7s | 8,2s | **0,44** |
|
||||
|
||||
Erkenntnis: pocket ist **speicherbandbreiten-gebunden**. Der Pool verbessert nur die Gesamt-Wall-Clock
|
||||
(Batch), verschlechtert aber die **TTFB** deutlich. Seriell liefert RTF 0,74 < 1 → generiert schneller
|
||||
als Echtzeit → Streaming spielt **lückenlos** und startet am schnellsten. **Für Lucys Live-Stimme
|
||||
gewinnt seriell.** Pool bleibt Opt-in für Batch (`/tts` ganze Datei): Sweet Spot **4w×2t** / **3w×2t**.
|
||||
|
||||
**B3 (FP-Drift-Gate überspringt kurze Audios) + kurzer erster Chunk:**
|
||||
Der MFCC-Fingerabdruck ist auf <2s Audio unzuverlässig (sim ~0,84 < 0,94) → löste 3× Fehlalarm-
|
||||
Regenerierung aus. B3 prüft den Drift erst ab `LUCY_FP_MIN_S=2.0`s stimmhafter Dauer; der F0-
|
||||
Männerstimmen-Wächter bleibt immer aktiv. Effekt (Drift-Warnungen pro Lauf: ~6 → 1):
|
||||
|
||||
| | TTFB | Wall |
|
||||
|---|---|---|
|
||||
| kurzer 1. Chunk, ohne B3 | 4,71s | 19,9s |
|
||||
| gebündelt (alt) | 3,74s | 15,8s |
|
||||
| **kurzer 1. Chunk, mit B3** | **1,31s** | 16,6s |
|
||||
|
||||
→ Lucy spricht nach **1,3s** statt 3,7s, Wall ~gleich, weiter lückenlos.
|
||||
|
||||
**Finale Defaults in `pocket_server.py`:**
|
||||
`LUCY_WORKERS=0` (seriell) · `LUCY_FAST_FIRST=1` (kurzer 1. Chunk) · `LUCY_FP_MIN_S=2.0` (B3) ·
|
||||
`LUCY_REF_CLEAN=1` (A2) · Voice aus `lucy_voice.safetensors` (A1). Pool-Opt-in für Batch:
|
||||
`LUCY_WORKERS=4 LUCY_WORKER_THREADS=2`.
|
||||
|
||||
**Offen / nächster Hebel:** distilliertes `german`-Modell (statt `german_24l`) abwarten — größter
|
||||
Qualität/Tempo-Sprung. Optional: TTFB weiter drücken über kürzeres erstes Wort / `lsd`-Tuning nur
|
||||
für den ersten Chunk.
|
||||
|
||||
## Umlaute (ae/oe/ue) — Fix (30.06.)
|
||||
|
||||
Symptom: Lucy spricht manchmal „u-e" statt „ü". Ursache: das LLM (Hermes/Qwen) gibt gelegentlich
|
||||
ASCII-Ersatzschreibweisen (ueber/schoen/maerz) aus; pocket liest sie wörtlich. (ß vs ss ist
|
||||
akustisch identisch -> ignoriert.)
|
||||
|
||||
**Wurzel-Fix (empfohlen, auf der Box im Hermes-System-Prompt/Persona ergänzen):**
|
||||
> „Schreibe ausschließlich korrektes Deutsch mit echten Umlauten (ä, ö, ü) und ß. Verwende niemals
|
||||
> die Ersatzschreibweisen ae, oe, ue oder ss anstelle von Umlauten."
|
||||
|
||||
**Schutznetz (in `pocket_server.py`, Default an `LUCY_UMLAUT_FIX=1`):** `_fix_umlauts()` wandelt NUR
|
||||
bekannte deutsche Umlaut-Ganzwörter zurück (Ganzwort + gängige Flexionsendungen, case-erhaltend).
|
||||
Verifiziert (`umlaut_test.py`): konvertiert über/für/größe/natürlich/mögliche/Gespräche; lässt
|
||||
neue/aktuell/Steuer/Quelle/Feuer/Frauen/genau/blaue unverändert. Grenzen: reine Wortliste, deckt
|
||||
nicht jedes seltene Wort — erweiterbar via `LUCY_UMLAUT_EXTRA="wort1,wort2"`. Der Wurzel-Fix bleibt
|
||||
die zuverlässige Lösung.
|
||||
|
||||
## Referenzen
|
||||
- pocket-tts README (GPU-kein-Speedup, export-voice, Sprachen): github.com/kyutai-labs/pocket-tts
|
||||
- Multilingual-Ankündigung (DE = `german_24l`, undistilliert): kyutai.org/blog/2026-05-04-pocket-tts-multilingual
|
||||
- Tech-Report / Performance: kyutai.org/blog/2026-01-13-pocket-tts · deepwiki.com/kyutai-labs/pocket-tts
|
||||
@@ -0,0 +1,300 @@
|
||||
# Optimierungsplan MC2 — Rollen, Warm-Set & Durchsatz (Strix Halo)
|
||||
|
||||
> Audit-Ergebnis + Maßnahmenplan. **Erst Plan, dann Umsetzung nach Freigabe.** Alles reversibel
|
||||
> (Config-Backups vor jeder Box-Änderung). Stand: 2026-06-30.
|
||||
> Quellen: Code-Audit (`backend/`, `frontend/`) + **Live-Box-Verifikation per SSH** (`/running`, voller
|
||||
> `config.yaml`, `free`, `du`, t/s-Probe) + `docs/AUDIT_KICKOFF.md` / `CLAUDE_CODE_BRIEF_lucy-brain-role.md`.
|
||||
|
||||
---
|
||||
|
||||
## 0. TL;DR (Priorität nach echtem Impact)
|
||||
|
||||
1. **✅ ERLEDIGT — Warm-Set korrigiert (W1/W2, 2026-06-30, live verifiziert):** `brains` jetzt = **gemma +
|
||||
embedding + vision** (`swap:false, persist:true`, ~31 GB). **Wichtige Korrektur durch Live-Test:** das
|
||||
ursprünglich vermutete OOM-Risiko existiert NICHT — llama-swap **swappt ganze Gruppen** statt sie zu
|
||||
ko-laden (cross-group Load = Gruppe raus, neues Modell rein; kein Speicher-Überlauf). Der echte Punkt war
|
||||
ein anderer: vorher pinnte `brains` auch `fast` (Chat-Lane-Hirn, kein Nutzen für Lucy → ~28 GB verschenkt),
|
||||
und in der reinen `gemma+embed`-Variante hätte **Lucys Sehen ihr Hirn verdrängt**. Jetzt halten Hirn +
|
||||
Gedächtnis + Sehen zusammen warm (verifiziert: gemma & vision ko-resident, 37 GB, kein Evict). **Reversibel.**
|
||||
2. **MTP-Spec-Decoding für gemma:** 53 t/s → erwartet ~80–100 t/s (1,5–2×, 0 Qualitätsverlust). Drafter-GGUF
|
||||
fehlt auf der Box **und MC2 kennt MTP-Drafter gar nicht** (nur klassisches `draft-simple`). **[mittel]**
|
||||
3. **coder-Spec-Draft verifizieren:** `coder` (Qwen3-Coder-Next) hat `--spec-draft-model Qwen3-0.6B-Q8_0
|
||||
--spec-type draft-simple` **aktiv** — ist dieser Draft wirklich vocab-kompatibel zu Qwen3-Coder-Next?
|
||||
Wenn nein, bremst/bricht Spec still. **[Check]**
|
||||
4. **Neues `scout`-Modell** (multimodaler Allrounder, Tools): HF-verifizierter Primärpick **GLM-4.6V-Flash (9B)**
|
||||
— Q4 ~6 GB + mmproj, natives Tool-Calling, schlägt unser vision-Modell. Alt: Gemma-4-12B. (Qwen3.5-VL-MoE
|
||||
verworfen — kein gepflegtes GGUF.) **[Recherche/Download]**
|
||||
5. **gemma-ctx 128k → 64k:** *kleiner* Hebel (real nur ~2 GB, weil Gemmas KV winzig ist — s. §1.4),
|
||||
trotzdem sinnvoll (freigegeben). Kein Funktions-Risiko für agentische Turns. **[risikoarm]**
|
||||
|
||||
---
|
||||
|
||||
## 1. IST-Zustand (live verifiziert per SSH, 2026-06-30)
|
||||
|
||||
### 1.1 Rollen-Taxonomie — Code ist konsistent ✅
|
||||
Der Brain-Rolle-Umbau aus `CLAUDE_CODE_BRIEF_lucy-brain-role.md` ist **bereits umgesetzt** (Commit `dd99401`).
|
||||
`hermes` (= Lucys Hirn) ist als kanonische Rolle in ALLEN Quellen synchron: `llamaswap.ROLE_IDS`,
|
||||
`sources.ROLE_IDS`+`CATEGORIES`, `roles.py` (`_capability_suit`/`_pref`/`_reason`), Frontend `ModelBadges.ROLES`.
|
||||
Die „hermes-Altlast/Inkonsistenz" aus dem Brief existiert nicht mehr.
|
||||
|
||||
### 1.2 Brain-Warm-Flow — Code ist fertig ✅, Box bestätigt ✅
|
||||
`POST /api/models/{id}/role role=hermes` → `agent.set_agent_brain()`: Alias + `brains:{swap:false,persist:true}`
|
||||
+ ttl 0 (altes Hirn raus/entspannt) + Hermes `model.default` + Gateway-Restart + Budget-Warnung. `brain_status()`
|
||||
prüft per `/running`.
|
||||
**Box bestätigt das Akzeptanzkriterium:** `gemma-4-26B-A4B-it` ist `ready`, `ttl:0`, Alias `hermes`, und in
|
||||
`groups.brains` (`swap:false, persist:true`). **Lucy bleibt warm — kein Kalt-Nachladen.** ✅
|
||||
|
||||
### 1.3 Live-Config aller Rollen (`/etc/llama-swap/config.yaml`, `globalTTL: 0`)
|
||||
|
||||
| Rolle (alias) | Modell | Gewichte (du) | ctx | ttl | Besonderheiten |
|
||||
|---|---|---|---|---|---|
|
||||
| fast | Qwen3.6-35B-A3B | 22 GB | 65536 | 0 | `--parallel 2`, mmproj, `--cache-reuse 256 -cram 16384`, **kein Spec**, **in `brains`** |
|
||||
| heavy | Qwen3.5-122B-A10B | 73 GB | 32768 | 600 | split-GGUF, `--cache-reuse 256 -cram 16384`, kein Spec |
|
||||
| coder | Qwen3-Coder-Next | 46 GB | 131072 | 600 | `--parallel 2`, **Spec aktiv** (Qwen3-0.6B-Q8 / draft-simple) |
|
||||
| coder-lite | Qwen3-Coder-30B-A3B | 18 GB | 131072 | 300 | `--cache-reuse 256 -cram 16384` |
|
||||
| vision | Qwen3-VL-8B-Instruct | 6 GB | 32768 | 300 | mmproj, **in `brains` (gepinnt!)** |
|
||||
| embed | Qwen3-Embedding-0.6B | 1 GB | 8192 | 0 | `--embedding --pooling last`, **in `brains`** |
|
||||
| **hermes** (Hirn) | **gemma-4-26B-A4B-it** | 17 GB (+1,2 mmproj) | **131072** | 0 | mmproj, `--jinja`, **kein Spec**, **in `brains`** |
|
||||
|
||||
`groups.brains = {swap:false, persist:true, members:[embedding, vision, fast, gemma]}`.
|
||||
|
||||
### 1.4 Speicher-Realität (GTT = 124 GB, `amdgpu.gttsize=126976`)
|
||||
**Wichtige Korrektur ggü. der reinen `fit.py`-Schätzung:** mit NUR gemma@131072 geladen meldet die Box
|
||||
**25,9 GB used** → gemma resident ≈ **~22 GB** (≈17 Gewichte + 1,2 mmproj + nur **~4 GB KV**). Gemma-4 nutzt
|
||||
starke Sliding-Window-Attention → KV ist viel kleiner, als `fit.py` (kalibriert an Hermes-14B) vorhersagt.
|
||||
**Folge:** gemma-ctx-Reduktion spart real nur ~2 GB; der echte Druck kommt vom **Warm-Set**, nicht vom Hirn-ctx.
|
||||
|
||||
**llama-swap-Gruppen-Semantik (live verifiziert, korrigiert frühere Annahme):** Eine `swap:false`-Gruppe hält
|
||||
NUR ihre EIGENEN Member ko-resident. Ein Modell in einer ANDEREN Gruppe (heavy/coder = je eigene implizite
|
||||
Gruppe) zu laden **swappt die gesamte aktive Gruppe raus** und lädt das neue Modell allein. `persist:true`
|
||||
verhindert nur das Idle-TTL-Entladen, NICHT das Gruppen-Swapping. **Folge:** es gibt **kein** OOM durch
|
||||
Ko-Laden (heavy lädt allein, 79 GB < 124 ✓) — aber das Hirn ist bei cross-group Loads (heavy/coder) **nicht**
|
||||
geschützt: es wird mitgeswappt und lädt danach kalt nach (~Sekunden). Beweis: heavy laden → `/running` zeigte
|
||||
nur heavy (78,7 GB), gemma weg trotz persist.
|
||||
|
||||
**Konsequenz fürs Design (umgesetzt):** Lucys zusammengehöriges Set (Hirn + Gedächtnis + Sehen) MUSS in DERSELBEN
|
||||
`swap:false`-Gruppe stehen, sonst verdrängt schon Lucys eigenes Sehen ihr Hirn. → `brains = gemma + embed +
|
||||
vision` (~31 GB). heavy/coder/coder-lite/fast = on-demand (swappen die brains-Gruppe beim Laden raus — ok, das
|
||||
sind separate Aktivitäten IDE/Delegation; Lucy lädt danach kurz nach).
|
||||
|
||||
### 1.5 Durchsatz (live gemessen)
|
||||
gemma (Hirn): **~53 t/s** Generierung, ~173 t/s Prompt (warm, kurzer Prompt). Solide Basis für ein 26B-A4B
|
||||
(4B aktiv) auf der bandbreitenlimitierten APU; MTP-Ziel ~80–100 t/s.
|
||||
|
||||
---
|
||||
|
||||
## 2. Verifikations-Status (war §2 „offen" — jetzt aufgelöst)
|
||||
- **V1 (gemma in `brains:{swap:false}`?)** → **JA** ✅ (Akzeptanzkriterium erfüllt).
|
||||
- **V2 (t/s)** → gemma 53 t/s gemessen ✅. fast/heavy/coder noch nicht gemessen (Laden würde Warm-Set stören).
|
||||
- **V3 (Voll-Config)** → komplett gelesen (§1.3) ✅.
|
||||
- **V4 (GTT/Belegung)** → GTT 124 GB, gemma-only 25,9 GB used ✅. Cross-group-Load swappt Gruppen (kein OOM,
|
||||
aber Hirn ungeschützt) — live verifiziert (§1.4). Behoben durch W1/W2.
|
||||
|
||||
- **V5 (coder-Spec-Draft)** → **AUFGELÖST ✅**: Qwen3-0.6B-Q8 ↔ Qwen3-Coder-Next = gleicher `tokens_sha
|
||||
facf459…`, n_vocab 151936, `compatible: True`. coder-Spec ist gültig & aktiv. (Bonus-Opportunität §9.2 F4.)
|
||||
|
||||
---
|
||||
|
||||
## 3. Maßnahmen je Rolle
|
||||
|
||||
### 3.1 ✅ Warm-Set / KO-Residenz — ERLEDIGT (2026-06-30, live verifiziert)
|
||||
| # | Maßnahme | Status / Begründung |
|
||||
|---|---|---|
|
||||
| W1/W2 | **`brains:{swap:false,persist:true}` = gemma + embedding + vision** (~31 GB); fast raus (Chat-Lane-Hirn, kein Lucy-Nutzen, ~28 GB gespart) | ✅ angewendet via `PUT /api/groups` (Backup `config.yaml.bak-20260630-152911`). **Verifiziert:** gemma & vision ko-resident, kein Evict (37 GB used); gemma bleibt `ready`. |
|
||||
| — | heavy/coder/coder-lite/fast = on-demand | Korrekt: kein OOM (Gruppen-Swap), Lucys Set bleibt zusammen warm. Trade-off: nach heavy-Delegation / IDE-Coding lädt Lucy kurz nach (~s). |
|
||||
| Optional | fast wieder pinnen, falls Chat-Lane dauerwarm sein soll | +28 GB Pin, swappt aber bei jedem heavy/coder-Load eh raus → geringer Nutzen. Auf Wunsch nachrüstbar. |
|
||||
|
||||
### 3.2 `hermes` / Lucys Hirn — gemma-4-26B-A4B-it
|
||||
| # | Maßnahme | Begründung / Zahlen | Risiko | Revert |
|
||||
|---|---|---|---|---|
|
||||
| H1 | **ctx 131072 → 65536** | Spart real nur ~2 GB (Gemma-KV winzig), aber konsistent mit `fast` ([[hermes-fast-context-blocker]]) und Agent-Turns brauchen selten >64k. | niedrig | ctx zurück |
|
||||
| H2 | **MTP-Spec-Decoding** (s. §4) | 53 → ~80–100 t/s. Drafter ~0,4 B → Budget vernachlässigbar. | mittel | Spec-Flags raus |
|
||||
| H3 | gemma sollte ggf. `--cache-reuse 256` bekommen (wie fast/heavy/coder) — fehlt aktuell | KV-Reuse über Turns → weniger Prompt-Reprocessing bei Agent-Ketten. | niedrig | Flag raus |
|
||||
|
||||
### 3.3 `fast` — Qwen3.6-35B-A3B (chat-Lane)
|
||||
- ctx 65536, `--parallel 2`, **kein Spec** (gut — vermeidet den qwen2.5-Vocab-Bruch). **Belassen**, aber W2 (raus
|
||||
aus dem harten Pin). Optional Qwen-Spec-Draft nur, wenn vocab-kompatibel zu Qwen3.6 (Check via `gguf_meta`).
|
||||
|
||||
### 3.4 `heavy` — Qwen3.5-122B-A10B
|
||||
- 73 GB, ctx 32768, ttl 600, on-demand. **Belassen.** Profitiert direkt von W1/W2 (kann wieder laden). Kein Spec
|
||||
(kein etablierter kompatibler Qwen3.5-Draft) — erst prüfen, sonst lassen.
|
||||
|
||||
### 3.5 `coder` / `coder-lite`
|
||||
- **coder** (Qwen3-Coder-Next, 46 GB, ctx 131072): **Spec aktiv mit Qwen3-0.6B-Q8** → **V5: Vocab-Kompatibilität
|
||||
verifizieren.** Falls inkompatibel → Draft entfernen (still kein Speed-Gewinn, evtl. Fehlerquelle).
|
||||
- **coder-lite** (Qwen3-Coder-30B-A3B, 18 GB, ctx 131072): coding-Lane-Default, schnell (3B aktiv). **Belassen.**
|
||||
Hinweis: Alias heißt `coder-lite` (Bindestrich) — die Lanes referenzieren `coder_lite`; sicherstellen, dass das
|
||||
Routing (`router_logic.py`/`routing_policy.py`) auf den richtigen Alias zeigt (Konsistenz-Check, kein Box-Risiko).
|
||||
|
||||
### 3.6 `vision` — Qwen3-VL-8B-Instruct
|
||||
- 6 GB, ctx 32768. **Belassen**, aber W2 (verdrängbar statt hart gepinnt). Sehen ist bursty → on-demand/warm-soft reicht.
|
||||
|
||||
### 3.7 `embed` — Qwen3-Embedding-0.6B
|
||||
- 1 GB, ttl 0, **bleibt hart gepinnt** (W1). Unantastbar — Mem0/Gedächtnis hängt dran.
|
||||
|
||||
### 3.8 `scout` — VAKANT → §5
|
||||
|
||||
---
|
||||
|
||||
## 4. MTP-Speculative-Decoding für gemma (Backend+UI-Erweiterung)
|
||||
**Problem:** MC2 kennt nur **klassische** Drafts. `spec_draft_flags()` schreibt `--spec-draft-model <d>
|
||||
--spec-type draft-simple`; `find_compatible_draft()` scannt nur `DRAFTS_DIR` mit stumpfem
|
||||
`gguf_meta.compatible()`. Ein **MTP-Kopf** (`gemma-4-26B-A4B-it-assistant`, Arch `gemma4_assistant`) wird nie
|
||||
angeboten. Engine 9843 kann `--spec-type draft-mtp` ✓; alte `--draft-max/-min` sind **entfernt** →
|
||||
`--spec-draft-n-max/-n-min`.
|
||||
|
||||
**Schritte:**
|
||||
1. **Drafter beschaffen** (Box, fehlt noch) — ⚠️ **Brief-Korrektur (HF-verifiziert 2026-06-30):** der Drafter
|
||||
heißt NICHT `…-assistant`. In `unsloth/gemma-4-26B-A4B-it-GGUF` liegt er als **`mtp-gemma-4-26B-A4B-it.gguf`**
|
||||
(Repo-Root, **0,46 GB**) bzw. **`MTP/gemma-4-26B-A4B-it-Q8_0-MTP.gguf`** (0,46 GB; BF16/F16-Varianten 0,86 GB).
|
||||
Eine Datei nach `/srv/models/gemma-4-26B-A4B-it-GGUF/` laden (Q8-MTP = kleinste, ~0,46 GB).
|
||||
**Flags (laut unsloth MTP/README):** `--model-draft <mtp.gguf> --spec-type draft-mtp --spec-draft-n-max 4`
|
||||
(kein `--spec-draft-n-min` nötig; neuere llama.cpp **auto-entdeckt** die Root-`mtp-*.gguf`). `-md` = Kurzform.
|
||||
2. **Backend (`llamaswap.py`):** MTP-Pfad ergänzen — MTP-Drafter als gültigen, vocab-kompatiblen Draft erkennen
|
||||
(Arch/Name `*-assistant*`); beim Aktivieren `-md <assistant.gguf> --spec-type draft-mtp
|
||||
--spec-draft-n-max N --spec-draft-n-min M` schreiben (statt `--spec-draft-model`/`draft-simple`). `_parse_model`
|
||||
um `draft-mtp`/`-md` erweitern (UI-Status). Flag-Namen mit `llama-server --help` (9843) gegenprüfen.
|
||||
3. **Frontend:** Draft-Modal MTP-Drafter als empfohlene Option für gemma führen („MTP-Kopf, vocab-kompatibel ✓").
|
||||
4. **Verifizieren:** t/s vor/nach (Basis 53 t/s) gegen die Box.
|
||||
|
||||
> Risiko mittel (neue Flag-Logik); voll reversibel.
|
||||
|
||||
---
|
||||
|
||||
## 5. Neues `scout`-Modell (Empfehlung)
|
||||
Profil: multimodaler Allrounder, klein/MoE (schnell auf bandbreitenlimitierter Box), Tool-Calling,
|
||||
KO-Residenz-freundlich, **distinkt** von Hirn (gemma-4) und vision (Qwen3-VL-8B).
|
||||
|
||||
**HF-verifiziert 2026-06-30** (Verfügbarkeit + Dateigrößen real geprüft, nicht geraten):
|
||||
|
||||
| Pick | Modell | Verifizierte Fakten | Hinweis |
|
||||
|---|---|---|---|
|
||||
| **Primär** | **GLM-4.6V-Flash (9B)** | GGUF da: `unsloth/GLM-4.6V-Flash-GGUF` (29 Quants, 12k+ dl), `lmstudio-community`, `MaziyarPanahi`. **Q4_K_M = 6,17 GB + mmproj 1,84 GB ≈ 8 GB.** Natives multimodales Function-Calling, 128k ctx. Benchmarks: **schlägt Qwen3-VL-8B** (= unser aktuelles vision-Modell) in fast allen Kategorien. | **9B dense** (nicht MoE), aber bei 9B/6 GB trotzdem schnell. **on-demand** (nicht pinnen). Könnte perspektivisch sogar `vision` ablösen. |
|
||||
| Alt A | **Gemma-4-12B-it (dense, multimodal)** | GGUF breit verfügbar: `unsloth/gemma-4-12b-it-GGUF` (1,4 Mio dl), `google/…-qat`, `lmstudio`. | dense → auf Bandbreiten-Box etwas langsamer; Familien-Dopplung mit dem Hirn (auch Gemma-4). |
|
||||
| ~~Alt B~~ | ~~Qwen3.5-VL-MoE~~ | **VERWORFEN:** kein gut gepflegtes GGUF — nur EIN Nischen-Repo (`jc-builds/Qwen3.5-9B-VLM-Q4_K_M`, kein Trusted-Author). Meine frühere „familien-treu"-Empfehlung war ungedeckt. | nicht nehmen. |
|
||||
|
||||
**Vorgehen:** GLM-4.6V-Flash (Q4_K_M ~6 GB + mmproj) via „Modelle finden"/`/api/models/install` laden → Rolle
|
||||
`scout`, **on-demand** → Mini-Benchmark (t/s + Vision+Tool-Prompt). Quellen (tagesaktuell, 2026-06):
|
||||
[VentureBeat: GLM-4.6V native tool-calling](https://venturebeat.com/ai/z-ai-debuts-open-source-glm-4-6v-a-native-tool-calling-vision-model-for) ·
|
||||
[zai-org/GLM-4.6V-Flash (HF)](https://huggingface.co/zai-org/GLM-4.6V-Flash) ·
|
||||
[GLM-4.6V Flash 9B Review/Benchmarks](https://binaryverseai.com/glm-4-6v-review-benchmarks-pricing-local-install/) ·
|
||||
[SiliconFlow: schnellste OSS-Multimodal 2026](https://www.siliconflow.com/articles/en/fastest-open-source-multimodal-models).
|
||||
|
||||
---
|
||||
|
||||
## 6. KO-Residenz-Set (umgesetzt, live verifiziert)
|
||||
**Lucys Warm-Set (`brains:{swap:false,persist:true}`):** gemma-Hirn + embedding + vision = **~31 GB.**
|
||||
Diese drei MÜSSEN zusammen in einer Gruppe sein (sonst verdrängt Lucys Sehen ihr eigenes Hirn — Gruppen-Swap).
|
||||
**Rein on-demand (swappen die brains-Gruppe beim Laden raus):** fast, heavy, coder, coder-lite, scout.
|
||||
|
||||
**Verifiziert:** gemma+vision ko-resident (37 GB, kein Evict); heavy-Load swappt die Gruppe (kein OOM, 79 GB).
|
||||
**Hinweis `budget.reserved_gb`:** das Modul nimmt an, fast/vision seien verdrängbar persist-Member, die NEBEN
|
||||
dem Hirn warm bleiben. Real swappt llama-swap aber die ganze Gruppe → die `reserved_gb`-Mathematik ist für
|
||||
cross-group-Loads **zu konservativ** (rechnet Hirn+Rest gleichzeitig, was nie passiert). **Folge-Task (Code):**
|
||||
`budget.reserved_gb` an die echte Gruppen-Swap-Semantik angleichen (nur das brains-Set zählt als resident,
|
||||
on-demand-Modelle laufen allein). Kein Box-Risiko, nur genauere ctx-Empfehlungen.
|
||||
|
||||
---
|
||||
|
||||
## 7. Reihenfolge / Risiko / Revert
|
||||
| Schritt | Aktion | Risiko | Revert |
|
||||
|---|---|---|---|
|
||||
| 0 | ✅ **Box-Backup** `config.yaml.bak-20260630-152911` | – | – |
|
||||
| 1 | ✅ **W1+W2** brains = gemma+embed+vision (via `PUT /api/groups`), verifiziert | niedrig | `set_group` zurück |
|
||||
| 2 | **H1** gemma ctx → 65536; **H3** `--cache-reuse 256` ergänzen | niedrig | ctx/Flag zurück |
|
||||
| 3 | **V5** coder-Draft-Vocab prüfen; inaktive Spec entfernen | niedrig | – |
|
||||
| 4 | **§4** MTP: Drafter laden → Backend/UI-Erweiterung → gemma-cmd → t/s-Vergleich | mittel | Spec-Flags raus |
|
||||
| 5 | **§5** scout recherchieren/installieren/zuweisen (on-demand) | niedrig | Modell löschen |
|
||||
|
||||
---
|
||||
|
||||
## 8. Offene Entscheidungen (für dich)
|
||||
1. **Warm-Set-Fix (W1/W2) jetzt umsetzen?** — behebt die heavy-OOM-Gefahr, mein klare Empfehlung als Erstes.
|
||||
2. **fast/vision verdrängbar** via eigener Gruppe `swap:true` (sie swappen sich gegenseitig) **oder** schlicht
|
||||
`persist:false`? (Detail des Umbaus.)
|
||||
3. **MTP jetzt oder als eigener Block?** (Code-Erweiterung Backend+UI + Drafter-Download.)
|
||||
4. **scout-Pick:** GLM-4.6V (Primär) oder familien-treu Qwen3.5-VL-MoE?
|
||||
5. gemma-ctx **65536** ist gesetzt (deine Wahl) — ok so.
|
||||
|
||||
---
|
||||
|
||||
## 9. Gesamtbild: Engine- (llama.cpp) & Router- (llama-swap) Flags — verifiziert gegen Build 9843
|
||||
> Der Brief deckte das nicht ab. Alle Aussagen hier gegen `llama-server --help` (9843) + Live-Config geprüft.
|
||||
|
||||
### 9.1 Cross-cutting (alle Chat-Modelle): `-ngl 999 -fa 1 --no-mmap -c <ctx> --jinja`
|
||||
- `-ngl 999` ✅ korrekt (volle Offload-Last in GTT auf der APU).
|
||||
- `-fa 1` → funktioniert, ist aber **Legacy-Form**. Aktuell: `-fa [on|off|auto]` (Default `auto`). Auf `-fa on`
|
||||
normalisieren (oder weglassen → auto), sonst Risiko bei künftigem Build. **[niedrig]**
|
||||
- `--no-mmap` ✅ bewusst (volle RAM-Residenz statt file-backed Pages; sinnvoll für warm gehaltene Modelle auf
|
||||
Unified Memory). Behalten.
|
||||
|
||||
### 9.2 Pro-Modell-Befunde & Hebel
|
||||
| # | Befund | Hebel | Prio |
|
||||
|---|---|---|---|
|
||||
| F1 | **gemma (Hirn) fehlt `--cache-reuse 256 -cram 16384`** — der Agent macht die meiste Multi-Turn-/Tool-Arbeit, nutzt aber nur den 8 GB-Default-Cache & **kein KV-Reuse über Turns** | ergänzen → System-Prompt + History-KV werden wiederverwendet, weniger Prompt-Reprocessing, schnellere Folge-Antworten | **hoch (Quick-Win)** |
|
||||
| F2 | **`--parallel 2` (fast/coder) ↔ Kontext:** unified KV ist „enabled if slots **auto**" — mit explizitem `--parallel 2` evtl. AUS → harte ctx-Teilung (fast 32k/Anfrage, coder 65k/Anfrage). `-cram` SOLL unified erzwingen, Help mehrdeutig | **V6:** `n_ctx_per_seq` aus `/props` verifizieren (beim nächsten Load). Dann: `--parallel 1`/auto (volle ctx, wenn keine Nebenläufigkeit) **oder** explizit `-kvu` setzen | mittel |
|
||||
| F3 | **KV-Quant `-ctk q8_0 -ctv q8_0`** — getestet an coder-lite | **EVALUIERT → VERWORFEN (2026-06-30):** kein Speicherdruck (Coder laufen allein, voller GTT), bei kurzem Kontext leichter Overhead (91→88 t/s), kein messbarer Gewinn. Voller-Präzisions-KV = beste Code-Treue. | erledigt |
|
||||
| F4 | **coder-Spec** (Coder-Next) ✅ kompatibel & behalten. Draft auch zu coder-lite kompatibel; heavy INKOMPATIBEL (V7) | **coder-lite + Spec EVALUIERT → VERWORFEN:** 91→**69 t/s (−24 %!)** — Spec schadet dem schnellen 3B-MoE (Draft-Overhead > Gewinn, 54 % Akzeptanz). Bestätigt: Spec nur für große/dichte Modelle (coder), nicht für MoE (fast/coder-lite). heavy kann mangels Vocab-Match ohnehin nicht. | erledigt |
|
||||
| F5 | **fast trägt `--mmproj`** (ist vision-fähig) — bewusst? +~1 GB | wenn die Chat-Lane nie Bilder bekommt: mmproj sparen | niedrig |
|
||||
| F6 | `-b/-ub` (batch/ubatch), `-t` (threads) ungenutzt = Defaults | Prompt-Speed evtl. via `-ub` tunbar — **messen statt raten**, niedrige Prio | niedrig |
|
||||
|
||||
### 9.3 llama-swap (Router-Ebene)
|
||||
- `globalTTL: 0` ✅ ok (Modelle haben überwiegend explizite ttl). `healthCheckTimeout: 300` ✅ reicht auch für
|
||||
heavy (~60 s Load gemessen). Gruppen-Swap-Semantik verstanden (§1.4), brains gefixt (§3.1).
|
||||
- **Systemischer Befund (Code↔Box-Drift):** MC2 `config._DEFAULT_CMD_TEMPLATE` =
|
||||
`llama-server -m {model} --host … --port ${PORT} -c {ctx} -ngl 999 -fa 1 --no-mmap` — kennt **KEINE** der auf
|
||||
der Box manuell ergänzten Optimierungen (`--cache-reuse`, `-cram`, `--parallel`, KV-Quant). **Folge: neu über
|
||||
MC2 installierte Modelle bekommen diese Tunings NICHT automatisch** → die Box driftet von der Code-Quelle weg.
|
||||
**Fix (Code):** sinnvolle Defaults ins Template / `register_model` (rollenabhängig), damit „Modelle finden"
|
||||
schon optimiert registriert. Optional: llama-swap `macros` für wiederholte Flag-Blöcke (Wartbarkeit). **[mittel]**
|
||||
|
||||
### 9.4 Verifikationspunkte — aufgelöst (2026-06-30)
|
||||
- **V6 ✅:** fast-Upstream `/props`: `total_slots:2`, **per-slot n_ctx 32768** → `--parallel 2` teilt den Kontext
|
||||
HART (kein unified-Sharing trotz `-cram`). **Maßnahme F2 angewendet:** fast → `--parallel 1` = **per-slot
|
||||
n_ctx 65536** (verifiziert). coder bleibt `--parallel 2` (IDE-Concurrency, 65k/Req).
|
||||
- **V7 ✅:** Qwen3-0.6B-Draft (`tokens_sha facf459`) → **heavy INKOMPATIBEL** (`a5e1ccff`, kein Spec möglich),
|
||||
**coder-lite KOMPATIBEL** (gleiche sha; aber MoE 3B-aktiv → Spec-Gewinn gering, vor Aktivierung benchmarken).
|
||||
|
||||
### Umsetzungs-Log (Box, 2026-06-30, Backups vorhanden)
|
||||
- ✅ **W1/W2** brains = gemma+embed+vision (`config.yaml.bak-20260630-152911`).
|
||||
- ✅ **gemma F1+H1+fa-on**: ctx 131072→65536, `-fa 1`→`-fa on`, `+ --cache-reuse 256 -cram 16384`
|
||||
(`bak-20260630-154842`). Verifiziert: ready, ~52 t/s, 24,4 GB.
|
||||
- ✅ **F2** fast `--parallel 2`→`1` (voller 65k-Kontext, verifiziert).
|
||||
- ✅ **MTP für gemma** (`bak-20260630-155520`): Drafter `mtp-gemma-4-26B-A4B-it.gguf` (461 MB) geladen,
|
||||
cmd `+ --model-draft … --spec-type draft-mtp --spec-draft-n-max 4`. **Verifiziert: 52 → 70,8 t/s (≈1,36×)**,
|
||||
Draft-Akzeptanz ~50 %, kein Crash, ready.
|
||||
- ✅ **scout** GLM-4.6V-Flash installiert (Q4 6,17 GB + mmproj 1,84 GB), role=scout, ttl 300, on-demand.
|
||||
**Voll verifiziert:** lädt sauber (neue GLM-4.6V-Arch auf Engine 9843), 34,6 t/s; **Vision ✓** (Testbild
|
||||
„blauer Kreis + Zahl 42" → korrekt erkannt: „Form: Kreis, Farbe: blau, Zahl: 42"); reasoniert korrekt.
|
||||
**Thinking-Modell** → No-Think für schnelle Kurzantworten via `chat_template_kwargs:{enable_thinking:false}`
|
||||
ODER `/nothink` im Prompt (beide verifiziert: sofort „Tokio" ohne Reasoning).
|
||||
|
||||
### Code-Änderungen (Dev-Repo F:\, **noch nicht deployt**)
|
||||
- ✅ **MC2 MTP-Draft-Support** (Brief-Deliverable). Verifiziert: py_compile OK, Logik-Test gegen echte GGUF,
|
||||
Frontend `tsc --noEmit` exit 0.
|
||||
- `config.py`: `SPEC_DRAFT_N_MAX` (Default 4).
|
||||
- `services/llamaswap.py`: `_is_mtp_draft`, `_spec_flags_for_draft`, `_sibling_mtp_drafters` (findet MTP-Köpfe
|
||||
NEBEN dem Modell); `find_compatible_draft`/`spec_draft_flags`/`drafts_for`/`set_spec_draft` MTP-bewusst;
|
||||
`_parse_model` erkennt `--model-draft`/`-md`. Backward-kompatibel (klassische Drafts unverändert).
|
||||
- Frontend `api.ts` (`DraftInfo.mtp`), `SpecDraftModal.tsx` (MTP-Badge). → gemmas MTP-Drafter erscheint im
|
||||
Spec-Modal automatisch als kompatibel + aktivierbar, schreibt die korrekten `draft-mtp`-Flags.
|
||||
- ✅ **CMD-Template-Drift** (`register_model`): `--cache-reuse 256 -cram 16384` als Default für alle
|
||||
Template-Modelle; `--parallel 2` nur noch für `coder` (kein ctx-Halbierungs-Footgun mehr); Spec-Auto-Attach
|
||||
MTP-bewusst + für alle Rollen self-guarding. Verifiziert: py_compile OK.
|
||||
- ✅ **`budget.reserved_gb`** an Gruppen-Swap-Semantik angeglichen (`_coresident_members` = `swap:false`-Set;
|
||||
on-demand = läuft allein → reserviert 0, voller GTT; brains-Member = reserviert die übrigen Member).
|
||||
Verifiziert gegen echte Config: hermes→16,2 GB, heavy/coder/scout→0 GB (ondemand-alone). Tote
|
||||
`_persist_members` entfernt.
|
||||
- ✅ **DEPLOYT (2026-06-30):** Commits `f655f09` (Code) + `9842249` (Frontend-`dist`) auf main; Box `git pull`
|
||||
→ `restart mission-control-2`. Verifiziert: `/api/models/drafts` zeigt MTP-Draft (`mtp:true, compatible:true`),
|
||||
gemma-Parse `spec_active:true spec_type:draft-mtp`, FastAPI serviert neues dist, API 200, gemma warm.
|
||||
(Box-`config.yaml`-Tunings waren schon vorher live.)
|
||||
|
||||
---
|
||||
|
||||
## 10. Leitplanken (eingehalten)
|
||||
- Features (Sehen/Hören/Sprechen/Embedding/Gedächtnis) + IDE-Lanes (chat/coding) bleiben funktionsfähig —
|
||||
W1/W2 macht sie sogar robuster (kein OOM mehr).
|
||||
- Vor jeder Box-Config-Änderung Backup; llama-swap reloadt per `-watch-config`.
|
||||
- **Keine destruktiven Aktionen ohne Freigabe dieses Plans.**
|
||||
</content>
|
||||
@@ -0,0 +1,281 @@
|
||||
# Komplett-Review Mission Control 2 + Lucy — 02.07.2026
|
||||
|
||||
**Methodik:** IST-Zustand live erhoben (Box per SSH, lokaler Lucy-PC, Working Tree `lucy-v2` inkl. uncommitted) — nicht aus Docs/Memory übernommen. SOLL-Zustand in 3 Recherche-Runden tagesaktuell (Stand 02.07.2026) ermittelt. Alle früheren Verdikte wurden neu auf den Prüfstand gestellt. Scope: alles außer Security.
|
||||
|
||||
---
|
||||
|
||||
## 1. Management-Summary
|
||||
|
||||
| Bereich | Zustand | Kernaussage |
|
||||
|---|---|---|
|
||||
| LLM-Stack (Box) | ✅ stark, 1 Bug | Vulkan-Pfad richtig, Qwen3.6+MTP läuft; aber chat-Lane kaputt (Autostart-Fund war Fehlalarm) |
|
||||
| Lucy Voice-Latenz | 🔴 3,5–5 s | 2026-Ziel ist 0,5–0,8 s. Schuld sind STT (2,0 s) und VAD-Timer (0,9 s), **nicht** das LLM (TTFT 65 ms) |
|
||||
| Lucy Avatar/App | ✅ sehr ausgereift | VRMA-Mocap, Lippensync v2, MateEngine-Features komplett; Electron 10 Major-Versionen alt |
|
||||
| Backend/Frontend-Code | 🟡 solide | 3 Monolithen, etwas DRY-Schuld, dist/ im Git; keine Rot-Flaggen |
|
||||
| Ökosystem-Aktualität | 🟡 | Hermes 1 Release hinter (v0.18.0 vom 01.07.); pocket-tts, mem0, llama-swap, three-vrm aktuell |
|
||||
| Verdikte | ✅ alle bestätigt | Pocket TTS, Electron, Qwen3.6-Hirn, kein ZeroClaw, Mem0 — alle bestehen die Re-Prüfung |
|
||||
|
||||
**Die zwei größten Hebel:**
|
||||
1. **STT-Tausch** (faster-whisper medium → Parakeet-TDT 0.6B v3): ~2,0 s → ~0,2 s pro Äußerung
|
||||
2. **VAD-Tausch** (RMS-Eigenbau → Silero VAD v6.2): Turn-Ende 900 → ~450 ms bei besserer Genauigkeit
|
||||
|
||||
Zusammen: Voice-to-Voice von ~3,5–5 s auf realistisch **≤1,2 s**, ohne neue Hardware.
|
||||
|
||||
---
|
||||
|
||||
## 2. IST-Zustand (live verifiziert)
|
||||
|
||||
### 2.1 Box (192.168.178.151 — Strix Halo, gfx1151, 122 GB, Vulkan)
|
||||
|
||||
| Komponente | Installiert | Aktuell (02.07.2026) | Bewertung |
|
||||
|---|---|---|---|
|
||||
| llama.cpp | b9843 (86b94708f) | b9859 (01.07.) | ✅ nur ~16 Builds dahinter |
|
||||
| llama-swap | v233 (29.06.) | v233 | ✅ aktuell |
|
||||
| Hermes Agent | **v0.17.0** (19.06.) | **v0.18.0** (01.07.) | 🟡 Update lohnt (3 P0 + 493 P1 gefixt) |
|
||||
| mem0ai / chromadb | 2.0.8 / 1.5.9 | 2.0.8 | ✅ aktuell |
|
||||
| Kernel | 7.0.0-27-generic | — | ✅ |
|
||||
|
||||
**Live-Modell-Lineup** (`/etc/llama-swap/config.yaml`, live gelesen — der Repo-Snapshot `deploy/llama-swap.config.yaml` ist **veraltet** und zeigt noch gemma-4):
|
||||
|
||||
| Alias | Modell | ctx | Besonderheiten |
|
||||
|---|---|---|---|
|
||||
| hermes (Brain) | Qwen3.6-35B-A3B (UD-Q4_K_M, MTP-GGUF) | 65536 | `--spec-type draft-mtp --spec-draft-n-max 3 --parallel 1 -cram 16384`, **kein** cache-reuse |
|
||||
| heavy | Qwen3.5-122B-A10B | 32768 | cache-reuse 256, ttl 600 |
|
||||
| coder | Qwen3-Coder-Next | 131072 | cache-reuse 256, ttl 600 |
|
||||
| coder-lite | Qwen3-Coder-30B-A3B | 131072 | cache-reuse 256, ttl 300 |
|
||||
| vision | Qwen3-VL-8B | 32768 | mmproj, in brains |
|
||||
| embed | Qwen3-Embedding-0.6B | 8192 | ttl 0, in brains |
|
||||
| scout | GLM-4.6V-Flash | 131072 | mmproj, ttl 300 |
|
||||
|
||||
Gruppe `brains` (swap:false, persist:true) = embed + vision + Qwen3.6. Alle Modelle: `-ngl 999 -fa on --no-mmap --jinja`. **Kein Modell nutzt KV-Cache-Quantisierung** (alles F16).
|
||||
|
||||
**Dienste live:** llama-swap, hermes-gateway, hermes-terminal, hermes-webui, mem0-service, voice-service — alle running. Ports 7681/8080/8642/8650/8765/8787/9001 belegt.
|
||||
|
||||
### 2.2 🔴 Live-Bugs (selbst reproduziert)
|
||||
|
||||
**B1 — ~~MC2-Backend ohne Autostart~~ FEHLALARM (korrigiert 02.07., nachverifiziert).**
|
||||
:9001 läuft als **User-Unit** `mission-control-2.service` (`~/.config/systemd/user/`, enabled, `Linger=yes` → reboot-fest). Die erste SSH-Prüfung sah User-Units nicht (fehlendes XDG_RUNTIME_DIR). Einziger echter Fund: Die **stale v1-System-Unit** `mission-control.service` (disabled, /opt/mission-control:9000) liegt noch in /etc/systemd/system und stiftet Verwirrung → bei Gelegenheit mit sudo entfernen (kosmetisch, kein Risiko).
|
||||
|
||||
**B2 — chat-Lane kaputt (live reproduziert).**
|
||||
```
|
||||
POST /v1/chat/completions {"model":"chat",...}
|
||||
→ {"error":"no router for requested model","src":"llama-swap"}
|
||||
```
|
||||
Die Routing-Policy routet `chat → fast ↔ heavy`, aber llama-swap kennt den Alias `fast` nicht mehr (Qwen3.6 hat nur noch den Alias `hermes`). Lucy ist nicht betroffen (sie geht über den Hermes-Agenten :8642), aber jeder Client der chat-Lane bekommt Fehler. Fix: `fast` als zweiten Alias eintragen **oder** Policy auf `hermes` umstellen.
|
||||
|
||||
**B3 — Config-Drift.**
|
||||
`deploy/llama-swap.config.yaml` ist untracked UND veraltet (gemma-4-Ära). Zusätzlich fehlen im Install-Template [`backend/config.py:33`](../backend/config.py) die Produktions-Tunings (cache-reuse, parallel, MTP) → neu installierte Modelle driften von der Box-Realität weg.
|
||||
|
||||
### 2.3 Gemessene Latenzkette Lucy
|
||||
|
||||
Quelle: `/api/voice/metrics` (live) + eigener TTFT-Test gegen llama-swap.
|
||||
|
||||
| Stufe | Messwert | Anteil am Problem |
|
||||
|---|---|---|
|
||||
| VAD-Turn-Ende | **900 ms** Fix-Timer (RMS-Eigenbau, [useVAD.ts](../client/lucy-desktop/src/renderer/src/lib/voice/useVAD.ts)) | groß |
|
||||
| STT (faster-whisper medium int8, Box-CPU) | **avg 2.046 ms · p50 2.092 ms · p95 2.529 ms** (n=13) | **dominant** |
|
||||
| LLM TTFT (Qwen3.6 warm, direkt :8080) | **65 ms** (24 Tokens in 334 ms) | ✅ kein Problem |
|
||||
| TTS erster Ton (pocket, RTF 0,44 + LEAD_IN 0,35 s) | ~0,5–1 s für den ersten Satz | mittel |
|
||||
| **Voice-to-Voice gesamt** | **~3,5–5 s** | Ziel 2026: **0,5–0,8 s** |
|
||||
|
||||
**Fazit: Das teuerste Glied ist NICHT das LLM.** STT + VAD-Timer fressen zusammen ~3 s. Genau dort setzt die Roadmap an.
|
||||
|
||||
### 2.4 Lokaler PC (RX 9070 XT / RDNA4)
|
||||
|
||||
- Lucy-App zum Prüfzeitpunkt nicht gestartet (kein Electron-Prozess, :8130 frei)
|
||||
- `ptts-venv`: pocket-tts **2.1.0** (= neueste Version), torch 2.12.1+cpu (bewusst CPU), onnxruntime 1.27.0, fastapi 0.138.1
|
||||
- GPU-Treiber 32.0.31021.5001
|
||||
- [client/lucy-tts/](../client/lucy-tts/) (3 GB): Produktions-TTS mit Stimmen-Wächter (F0-Floor 160 Hz, MFCC-Fingerprint ≥0,94, Best-of-N), Phase A+B des TTS-Plans abgeschlossen (beste Config: 4 Worker × 2 Threads, RTF 0,44)
|
||||
- [client/lucy-f5/](../client/lucy-f5/) (5,3 GB): F5-TTS-ONNX/DirectML-Experiment — **Phase C/D (Finalentscheid pocket vs. F5) offen**
|
||||
- Lucy-Desktop: Electron **33** (aktuell: **43** vom 30.06. — 2 Jahre Chromium-Rückstand), three-vrm 3.4/animation 3.5.4 (≈ aktuell), TS strict
|
||||
|
||||
### 2.5 Code-Qualität (3 Explore-Agents über das ganze Repo)
|
||||
|
||||
**Backend (4.889 LOC, 11 Router / 28 Services):** insgesamt sauber, Multi-Sidecar-Architektur durchdacht. Findings:
|
||||
- C1: Install-Template ohne Produktions-Tunings ([backend/config.py:33](../backend/config.py)) → Drift
|
||||
- C2: `_describe_images()` blockiert den Voice-Chat bis 120 s synchron ([backend/routers/voice.py](../backend/routers/voice.py))
|
||||
- C3: Thinking-Deaktivierung 3× dupliziert (voice.py, gateway.py, mem0_service/app.py) — DRY
|
||||
- C4: Junk-Fact-Regex im Mem0-Sidecar hartcodiert (Wartungspunkt)
|
||||
- Dead Code: `backend/services/pricing.py`, vermutlich `token_stats.py`
|
||||
- Uncommitted Änderungen (voice.py No-Think, connect.py IDE-only-`coding`, mc2-memory-Guards, Mem0-Junk-Guard, voice_service install.sh) sind **alle konsistent und sinnvoll** — nur eben nicht committet
|
||||
|
||||
**Frontend (React 18/Vite 6/Tailwind 4/TanStack 5):**
|
||||
- Voice-Umzug in die Desktop-App sauber (keine toten Imports/Nav-Reste) ✅
|
||||
- UI-Rework (roleMeta/Health/Expert/WarmSet) vollständig; coder_lite↔coder-lite via ALIAS gelöst
|
||||
- F1: [Cockpit.tsx](../frontend/src/views/models/Cockpit.tsx) (990 Z) + [SystemDrawer.tsx](../frontend/src/components/SystemDrawer.tsx) (934 Z) Monolithen
|
||||
- F2: 🟡 `frontend/dist/`-Asset-Stand auf lucy-v2 inkonsistent (alte gelöscht, neue untracked) — dist-im-Git selbst ist Absicht (Box hat kein Node), avatar.vrm ist via .gitignore korrekt draußen
|
||||
- F4: keine globale Error Boundary
|
||||
- MC3-Prototyp ([frontend/src/mc3/](../frontend/src/mc3/)) non-destruktiv hinter `#mc3`, untracked
|
||||
|
||||
**Lucy-Desktop:**
|
||||
- L1: [useVoiceAgent.ts](../client/lucy-desktop/src/renderer/src/lib/voice/useVoiceAgent.ts) (377 Z) monolithisch (SSE, Chunking, Tools, Perf, Vision in einem Hook)
|
||||
- L2: kein Retry/Backoff beim TTS-Warmup (Crash → 2-min-Spinner)
|
||||
- L3: VAD = reines RMS mit 900-ms-Stille-Regel
|
||||
- L4: Echo-Schutz = 450-ms-Cooldown-Workaround (kein echtes Barge-in)
|
||||
- Neues [voice-core/](../client/lucy-desktop/src/renderer/src/voice-core/)-Adapter-Layer (STT/TTS austauschbar) ist die richtige Architektur ✅
|
||||
- H1: ~23 Dateien uncommitted, darunter die neue voice-core-Architektur → Verlustrisiko
|
||||
|
||||
---
|
||||
|
||||
## 3. SOLL-Zustand (Recherche, Stand 02.07.2026)
|
||||
|
||||
### 3.1 Voice-Latenz — der Stand der Technik
|
||||
|
||||
- 2026-Zielmarke: **500–800 ms voice-to-voice**; Grundprinzip: **Streaming auf jeder Stufe** (STT-Partials sparen 200–400 ms, TTS startet auf Teiltext, VAD+Turn-Taking-Budget 150–300 ms) — [Latenz-Guide](https://futureagi.com/blog/how-to-optimize-voice-agent-latency-2026/), [Latenz-Budget-Design](https://smallest.ai/blog/designing-voice-assistants-stt-llm-tts-tools-and-latency-budget)
|
||||
- **STT:** [Parakeet-TDT 0.6B v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3) — 25 EU-Sprachen inkl. Deutsch, 6,32 % WER (besser als Whisper large-v3 mit 7,44 %), extrem schnell auch auf CPU, keine Silence-Halluzination. Deploy-Wege: [onnx-asr](https://pypi.org/project/onnx-asr/) (unterstützt CPU **und** DirectML → 9070 XT), fertige [OpenAI-kompatible FastAPI-Wrapper](https://github.com/groxaxo/parakeet-tdt-0.6b-v3-fastapi-openai). Kyutais Streaming-STT mit eingebautem semantic VAD ([delayed-streams-modeling](https://github.com/kyutai-labs/delayed-streams-modeling)) wäre architektonisch ideal, ist aber **nur EN/FR** → für Deutsch raus.
|
||||
- **VAD:** [Silero VAD v6](https://github.com/snakers4/silero-vad/releases/tag/v6.0) (v6.2, ONNX): −16 % Fehler auf Noisy-Data vs. v5 — klar besser als der RMS-Eigenbau; erlaubt kürzeres Endpointing (~450 ms) bei weniger Fehlstarts.
|
||||
- **Semantische Turn-Detection:** Smart Turn V2 (leichtgewichtig), [Easy Turn](https://arxiv.org/pdf/2509.23938) (akustisch+linguistisch, 4 Turn-States, open source).
|
||||
- **TTS:** pocket-tts 2.1.0 ist Stand der Technik für CPU-Realtime ([kyutai](https://kyutai.org/blog/2026-01-13-pocket-tts/)). Challenger (CosyVoice2-0.5B, Chatterbox-Turbo) bieten keinen klaren Vorteil. → Einziger offener Entscheid bleibt das hauseigene F5-Experiment (Phase C/D).
|
||||
|
||||
### 3.2 Box maximieren
|
||||
|
||||
- **Vulkan/RADV bleibt der schnellste Pfad auf Strix Halo** (schlägt ROCm/HIP bei pp und tg) — die Box ist richtig aufgestellt ([llm-tracker Strix-Halo](https://llm-tracker.info/_TOORG/Strix-Halo), [Strix-Halo-Guide](https://github.com/hogeheer499-commits/strix-halo-guide)). ROCm 7.2/RDNA4 betrifft nur den lokalen PC.
|
||||
- **Host-Memory-Prompt-Cache:** `--cache-ram` + Prefix-Restore bringt bis **93 % TTFT-Reduktion** bei Agent-Workloads mit wachsender Session ([llama.cpp #20574](https://github.com/ggml-org/llama.cpp/discussions/20574)). `-cram 16384` ist gesetzt; das Zusammenspiel mit dem (am hermes-Alias entfernten) `--cache-reuse` gehört gebencht.
|
||||
- **KV-Cache-Quantisierung ungenutzt:** `-ctk/-ctv q8_0` halbiert den KV-Bedarf (relevant bei coder@131k, hermes@65k) bei praktisch verlustfreier Qualität; [TurboQuant](https://github.com/ggml-org/llama.cpp/discussions/20969) (3-bit, near-lossless ab 13B) steht vor der Integration — beobachten.
|
||||
- **„MoE Speculative Trap":** Ein aktueller [Strix-Halo-Deep-Dive](https://dev.to/agustinsacco/breaking-the-moe-speculative-trap-460-ts-on-amd-strix-halo-446d) zu exakt Qwen3.6-35B-A3B zeigt: Spec-Verify kann bei sparsem MoE pro Verify-Pass fast alle 35B Gewichte anfassen — in seinem Setup war **ohne** Draft (parallel=1, KV Q8_0) 43,1 t/s vs. 17,7 t/s. Das widerspricht der eigenen Box-Messung (+35 % MIT MTP). Konsequenz: **beide Betriebsarten benchen**, auch bei vollem Kontext ([Decode-Drop bis 64 % dokumentiert](https://kmarble.dev/posts/strix-halo-full-context-decode-drops/)).
|
||||
- **llama-swap-Features ungenutzt:** `capabilities` (v225), `-config-dir`-Fragmente (v230), Swap-Matrix/Groups V2 ([Releases](https://github.com/mostlygeek/llama-swap/releases)).
|
||||
- Modell-Landschaft 128 GB für heavy-Kandidaten: gpt-oss-120B, Mistral Small 4 (119B-A6B), Llama 4 Scout — nur mit Bench-Gate.
|
||||
|
||||
### 3.3 Ökosystem
|
||||
|
||||
- **Hermes v0.18.0 „Judgment"** (01.07.): 3 P0 + 493 P1 gefixt, Background-Fan-out für Subagents, `/learn`-Skills, Memory-Graph in der Desktop-App ([Releases](https://github.com/NousResearch/hermes-agent/releases)). **Plugin-API:** `sync_turn()` bekommt optional `messages` (voller Turn-Kontext inkl. Tool-Calls); Legacy-Signatur bleibt → **mc2-memory ist update-kompatibel** und kann später Tool-Ergebnisse mitlernen ([Memory-Provider-Doku](https://hermes-agent.nousresearch.com/docs/developer-guide/memory-provider-plugin)).
|
||||
- **Mem0 v3-Algorithmus** ([Migration](https://docs.mem0.ai/migration/oss-v2-to-v3)): Single-Pass-Extraktion (~2× schneller), Hybrid-Retrieval (semantisch + BM25 + Entity-Graph, ohne externe Graph-DB), +20 LoCoMo / +26 LongMemEval. 2.0.8 installiert, `custom_instructions` schon migriert → Status final verifizieren.
|
||||
- **Electron 43** (30.06.), Chromium 150 / Node 24 — Lucy ist auf 33.
|
||||
- **Vision:** Qwen3-VL ist aktuelle Generation ✅; [Qwen3-VL-30B-A3B](https://github.com/qwenlm/qwen3-vl) (MoE, 3B aktiv, top auf GUI-Benchmarks) wäre der Upgrade-Kandidat für die Bildschirm-Sicht.
|
||||
- **Lokaler GPU-Klon:** [AMD Lemonade](https://runaihome.com/blog/amd-lemonade-local-llm-server-npu-gpu-guide-2026/) (llama.cpp-Vulkan + whisper.cpp + Kokoro, OpenAI-kompatibel, RDNA4-Support) als fertiger Unterbau; llama.cpp-Vulkan ist der [verlässlichste 9070-XT-Pfad](https://digtvbg.com/blog/llama-server-vulkan-rdna4-vllm-rocm-benchmark/).
|
||||
|
||||
### 3.4 Verdikte — Ergebnis der Re-Prüfung
|
||||
|
||||
| Verdikt | Ergebnis | Begründung |
|
||||
|---|---|---|
|
||||
| Pocket TTS bleibt | ✅ **bestätigt** | 2.1.0 aktuell, RTF 0,44 optimiert, Wächter-System; Challenger ohne klaren Vorteil. F5-Entscheid (Phase C/D) bleibt als einziger offener Punkt |
|
||||
| Electron bleibt (nicht Tauri) | ✅ **bestätigt** | Alle nativen Features implementiert; aber Major-Update 33→43 einplanen |
|
||||
| Qwen3.6 als Lucy-Hirn | ✅ **live bestätigt** | Läuft mit MTP, TTFT 65 ms; gemma-4 ist komplett raus |
|
||||
| ZeroClaw bleibt tot | ✅ **bestätigt** | Kein neuer Anlass |
|
||||
| Mem0 als Memory-Layer | ✅ **bestätigt** (neu geprüft) | Zep/Hindsight benchen höher, sind aber schwerer/teils closed; Mem0: beste Integration, v3-Algo, deployt |
|
||||
|
||||
---
|
||||
|
||||
## 4. Roadmap (Aufwand/Nutzen, Voice-Speed-first)
|
||||
|
||||
### P0 — Live-Bugs & Betriebssicherheit (Stunden)
|
||||
| # | Maßnahme | Nutzen |
|
||||
|---|---|---|
|
||||
| 1 | ~~Autostart fixen~~ **erledigt als Fehlalarm** — User-Unit `mission-control-2` mit Linger ist reboot-fest; nur stale v1-Unit bei Gelegenheit löschen (sudo) | Klarheit |
|
||||
| 2 | chat-Lane fixen (`fast`-Alias ergänzen oder Policy auf `hermes`) | Live-Bug, alle chat-Clients betroffen |
|
||||
| 3 | Box-Config → Repo syncen + `deploy/` versionieren; Template-Drift C1 fixen | Quelle der Wahrheit |
|
||||
| 4 | Hermes v0.17.0 → v0.18.0 + Health-Brain-Check | 496 Upstream-Fixes |
|
||||
| 5 | Commit-Hygiene (voice-core, MC3, Lucy-Fixes); `frontend/dist/` in .gitignore | Verlustrisiko weg |
|
||||
|
||||
### P1 — Voice-Latenz: 3,5–5 s → Ziel ≤1,2 s (Tage)
|
||||
| # | Maßnahme | Erwartung |
|
||||
|---|---|---|
|
||||
| 6 | **STT-Tausch**: Parakeet-TDT 0.6B v3 statt faster-whisper medium. Variante A: Box-CPU (onnx-asr im voice_service); Variante B: lokal 9070 XT (DirectML). A/B: Deutsch-WER + Latenz | **~2,0 s → ~0,2 s** |
|
||||
| 7 | **VAD-Tausch**: Silero VAD v6.2 (ONNX) statt RMS in useVAD.ts; Endpointing 900 → ~450 ms | **−450 ms** |
|
||||
| 8 | TTS-TTFA: LEAD_IN 0,35 s prüfen, Erster-Satz-kurz-Regel messen (lucy_perf) | −100–300 ms |
|
||||
| 9 | **Brain-Bench-Matrix** am hermes-Alias: MTP n-max 3↔4↔ohne (Speculative-Trap-These) × ±KV Q8_0 × cache-reuse/cram × -b/-ub — bei Kurz- UND Voll-Kontext | Bestes Setup evidenzbasiert |
|
||||
| 10 | voice_metrics ausbauen: VAD→STT→Agent→First-Audio je Turn | Messbarkeit |
|
||||
| 10b | KV Q8_0 für coder@131k/heavy erwägen | Warm-Set-Reserve |
|
||||
|
||||
### P2 — Lucy v2 Kern & Code-Gesundheit (1–2 Wochen)
|
||||
| # | Maßnahme |
|
||||
|---|---|
|
||||
| 11 | Semantische Turn-Detection (Smart Turn V2 / Easy Turn) auf Parakeet aufsetzen; Barge-in-Vorstufe statt 450-ms-Cooldown |
|
||||
| 12 | F5-TTS Phase C/D abschließen → Finalentscheid pocket vs. F5 (Benchmark-Gate) |
|
||||
| 13 | useVoiceAgent-Refactor in Hooks + TTS-Retry/Backoff; Electron 33→43 |
|
||||
| 14 | Backend C2 (Vision async) + C3 (DRY) + Dead Code; Frontend Cockpit/SystemDrawer-Split + globale Error Boundary |
|
||||
|
||||
### P3 — Strategisch (nach Signal)
|
||||
| # | Maßnahme |
|
||||
|---|---|
|
||||
| 15 | Mem0-v3-Status final verifizieren; Junk-Regex durch v3-Retrieval entschärfen; mc2-memory auf `sync_turn(messages)` heben |
|
||||
| 16 | MC3-Vollbau; llama-swap-Features (capabilities, config-dir, Swap-Matrix) |
|
||||
| 17 | heavy-Bench: Qwen3.5-122B vs. gpt-oss-120B / Mistral Small 4; Vision-Upgrade Qwen3-VL-30B-A3B für Bildschirm-Sicht |
|
||||
| 18 | Lokaler GPU-Klon: AMD Lemonade auf der 9070 XT evaluieren |
|
||||
|
||||
---
|
||||
|
||||
## 5. Findings-Katalog (Referenz)
|
||||
|
||||
| ID | Schwere | Fund | Ort |
|
||||
|---|---|---|---|
|
||||
| B1 | ⚪ (Fehlalarm) | :9001 läuft als User-Unit mit Linger — reboot-fest; nur stale v1-Unit als Kosmetik-Rest | Box, /etc/systemd/system/mission-control.service (v1) |
|
||||
| B2 | 🔴 | chat-Lane → fast-Alias existiert nicht mehr | Box, Routing-Policy ↔ /etc/llama-swap/config.yaml |
|
||||
| B3 | 🟡 | deploy/llama-swap.config.yaml untracked + veraltet | Repo |
|
||||
| C1 | 🟡 | Install-Template ohne Produktions-Tunings | backend/config.py:33 |
|
||||
| C2 | 🟡 | Vision-Beschreibung blockiert Chat bis 120 s | backend/routers/voice.py |
|
||||
| C3 | 🟢 | Thinking-Disable 3× dupliziert | voice.py / gateway.py / mem0_service/app.py |
|
||||
| C4 | 🟢 | Junk-Fact-Regex hartcodiert | mem0_service/app.py |
|
||||
| C5 | ⚪ (Fehlalarm) | pricing.py + token_stats.py sind IN BENUTZUNG (system.py /token_stats-Endpoint, gateway_stream) — kein Dead Code | backend/services/ |
|
||||
| F1 | 🟡 | Monolithen 990/934 Z | frontend/src/views/models/Cockpit.tsx, components/SystemDrawer.tsx |
|
||||
| F2 | 🟡 (korrigiert) | dist/ im Git ist ABSICHT (Box hat kein Node, deploy = git reset --hard; s. deploy/build.sh) — echter Fund war nur der inkonsistente Asset-Stand auf lucy-v2 (behoben durch Rebuild+Commit). Build-on-Box/CI wäre P3-Option | frontend/dist/ |
|
||||
| F4 | 🟢 | Keine globale Error Boundary | frontend/src/App.tsx |
|
||||
| L1 | 🟡 | useVoiceAgent monolithisch (377 Z) | client/lucy-desktop/.../lib/voice/useVoiceAgent.ts |
|
||||
| L2 | 🟡 | Kein TTS-Warmup-Retry | ebd. + main/index.ts |
|
||||
| L3 | 🔴 | RMS-VAD mit 900-ms-Timer | client/lucy-desktop/.../lib/voice/useVAD.ts |
|
||||
| L4 | 🟡 | Echo-Cooldown statt Barge-in | ebd. |
|
||||
| H1 | 🟡 | ~23 Dateien uncommitted (inkl. voice-core, MC3) | lucy-v2 Working Tree |
|
||||
|
||||
---
|
||||
|
||||
## 6. Nachtrag: P0/P1-Umsetzung (02.07.2026)
|
||||
|
||||
**P0 komplett:** chat-Lane gefixt (fast-Alias), Hermes v0.18.0 (Postcheck grün), Config versioniert, Template-Drift behoben, Git aufgeräumt. B1 war Fehlalarm (s.o.).
|
||||
|
||||
**P1-6 STT:** Parakeet-TDT 0.6B v3 (onnx-asr, int8, Box-CPU) als Default — **A/B gemessen: 0,45 s vs. 2,44 s** (whisper-medium) bei identischem Transkript. Live deployt, auch via :9001 verifiziert.
|
||||
|
||||
**P1-7 VAD:** Silero v5 (vad-web 0.0.30) statt RMS — Endpointing 900→450 ms, WAV direkt (kein Opus-Umweg mehr). Build + Asset-Auslieferung verifiziert; Mikro-Test beim User.
|
||||
|
||||
**P1-9 Brain-Bench-Matrix** (Qwen3.6-35B-A3B, separater Port, Vulkan, je ~100 Gen-Token):
|
||||
|
||||
| Config | tg kurz | tg tief (15k) | Draft-Akzeptanz kurz |
|
||||
|---|---|---|---|
|
||||
| MTP n-max 3, KV f16 (live) | **78,6 t/s** | — (Probe-EOS) | 55 % |
|
||||
| MTP n-max 4 | 63,7 | 100,9* | 39 % |
|
||||
| ohne Spec | 62,4 | — (Probe-EOS) | — |
|
||||
| **MTP n-max 3 + KV Q8_0** | **78,6** | 83,7 | 56 % |
|
||||
| ohne Spec + KV Q8_0 | 62,1 | 57,0 | — |
|
||||
|
||||
*\*repetitiver Test-Text → unrealistisch hohe Draft-Akzeptanz (95 %); die Kurz-Spalte ist maßgeblich.*
|
||||
|
||||
**Bench-Verdikte:** (1) MTP n-max 3 bestätigt (+26 % vs. ohne Spec — die „MoE Speculative Trap" gilt auf diesem Vulkan/parallel-1-Setup NICHT; der Artikel testete ROCm + parallel 4). (2) n-max 4 lohnt nicht (Akzeptanz fällt auf 39 %). (3) **KV Q8_0 ist gratis** (identische t/s) und halbiert den KV-Speicher → für hermes in deploy/llama-swap.config.yaml übernommen; Live-Schaltung braucht User-Freigabe. Für coder/heavy vorher cache-reuse×KV-Quant-Verträglichkeit testen (Context-Shift mit quantisiertem KV ist in llama.cpp historisch heikel).
|
||||
|
||||
**P1-10:** `chat_first_content`-Metrik in voice.py — misst die echte Hirn-Latenz (Agent-Overhead + LLM-TTFT bis zum ersten Inhalts-Token) statt nur des SSE-Starts.
|
||||
|
||||
**P2-Nachtrag (02.07.):**
|
||||
- **Profiling:** Folge-Turns 1,3 s, NEUE Sessions 5,6–12 s (Session-Prefill System-Prompt+Tools ~4–5k Tok). Mem0-Search unschuldig (18 ms). **Aber:** Die Mem0-Lern-Extraktion (~2,4 s je Turn, gleiche Qwen3.6-Instanz) blockiert bei `--parallel 1` den nächsten Voice-Turn in der Queue → Fix vorbereitet: `--parallel 2 -c 131072 + KV Q8_0` (2×65k-Slots, speicherneutral zu 1×65k f16). Live-Schaltung = User (deploy/llama-swap.config.yaml ist fertig).
|
||||
- **C2 gefixt:** Bildschirm-Sicht läuft jetzt IM Stream (SSE startet sofort, `hermes.vision.progress`-Event → Lucy kann Warte-Ansage sprechen); Vision-Timeout 120→45 s (MC_VISION_TIMEOUT).
|
||||
- **C3 bewertet:** nur 2 Backend-Stellen mit unterschiedlicher Gating-Logik (dritte im separaten Mem0-venv) → kein Shared-Helper erzwungen, Pattern dokumentiert.
|
||||
- **C5 war Fehlalarm** (s. Findings-Tabelle).
|
||||
- **L2 gefixt:** TTS-Warm-Gate mit Retry/Backoff + sichtbarer Fehlermeldung (vorher: nach 120 s stumm „ready" ohne Stimme).
|
||||
|
||||
## 7. Nachtrag 2: P2/P3-Vollausbau (02.07., zweite Session-Hälfte)
|
||||
|
||||
**Brain-Config LIVE DEPLOYT + verifiziert (02.07., User-Freigabe):** hermes läuft mit `-c 131072 --parallel 2 -ctk/-ctv q8_0` + MTP. Gemessen: **2 gleichzeitige Requests laufen echt parallel** (~1,9 s/2,0 s statt seriell — Mem0-Extraktion blockiert Voice-Turns nicht mehr), tg 66 t/s (leichter parallel-2-Abschlag vs. 78 solo, bewusster Trade), RAM 40/122 GB. **Voice-E2E: neue Session 0,85 s total, first_content 803 ms** (Morgen-Baseline: 5,6–12 s bei neuen Sessions). capabilities werden via /v1/models ausgeliefert.
|
||||
|
||||
| Punkt | Ergebnis |
|
||||
|---|---|
|
||||
| **P2-11 Turn-Detection** | ✅ Smart Turn v3.2 (8 MB ONNX) im Voice-Sidecar (`/turn`, ~110 ms warm) + Semantik-Hold im Client (bei „unfertig" bis 1,8 s auf Fortsetzung warten). **Zwei Upstream-Fallen live diagnostiziert:** Audio muss LINKS gepadded werden (sonst konstant „complete"), und der Output ist P(unfertig) — entgegen der Doku. Verifiziert: fertig=0,74→true, mitten im Wort=0,04→false |
|
||||
| **P2-12 F5 Phase C/D** | ✅ **Verdikt: Pocket bleibt.** F5-ONNX auf 9070 XT/DirectML: RTF 0,59@NFE32 / 0,30@NFE16 — aber nicht streamfähig (TTFA = ganze Satzdauer), NFE16-Qualitätsrisiko, GPU-Dauerbelegung. F5 taugt als Offline-Renderer, nicht für den Dialog-Loop |
|
||||
| **P2-13a Refactor** | ✅ useVoiceAgent 397→305 Z; textPipeline/visionIntent/perf als pure Module |
|
||||
| **P2-13b Electron** | ✅ 33→43 (Chromium 150/Node 24), Boot-Smoke-Test grün (allow-scripts-Falle: install.js manuell) |
|
||||
| **P2-14 Cockpit/SystemDrawer** | ⏸ **bewusst vertagt:** reiner Qualitäts-Refactor von live deployter UI; braucht eine eigene Session mit Browser-Verifikation (MC_API_TARGET-Workflow) statt eines Blind-Splits am Session-Ende |
|
||||
| **P3-15 Mem0 v3** | ✅ verifiziert aktiv (BM25/Entity/Hybrid in 2.0.8); mc2-memory nutzt jetzt `sync_turn(messages)` — lernt Tool-NAMEN mit (Ergebnisse bewusst nicht: Poisoning-Vektor). Deployt, Postcheck grün |
|
||||
| **P3-16 llama-swap** | ✅ `capabilities` (in/out/tools/context) je Modell in deploy-Config; Swap-Matrix aktuell unnötig (brains-Gruppe reicht) |
|
||||
| **P3-17 Kandidaten** | ✅ **Beide gebencht + als testbare Modelle live deployt** (ohne Alias-Wechsel — Qualitäts-Entscheid beim User): (1) **gpt-oss-120B: tg 54–55 t/s vs. heavy (Qwen3.5-122B) 23,5 t/s = 2,3× schneller bei 60 statt 73 GB** → klare Empfehlung für die heavy-Lane nach Deutsch-/Reasoning-Check (`model:"gpt-oss-120b"` am Gateway testen). (2) **Qwen3-VL-30B-A3B: tg 90–92 t/s** → Bildschirm-Sicht-Upgrade nach Screenshot-Check (`MC_VISION_MODEL=Qwen3-VL-30B-A3B-Instruct`); brains-Budget-Frage: 19 vs. 6 GB warm |
|
||||
| **P3-18 Lemonade** | ✅ **Verdikt: nicht als Unterbau.** lemonade-sdk 9.1.4 (Python) installiert + Server lief (OpenAI-API, gute Registry inkl. gpt-oss/GLM) — aber: Python-Edition **offiziell deprecated** (C++-Installer ist der Weg), Server blockiert komplett während Model-Downloads (Health tot >10 min bei 400-MB-Modell), Ryzen-AI-Hybrid auf 9700X unsupported. **Empfehlung für den lokalen GPU-Klon: llama.cpp-Vulkan + llama-swap — derselbe Stack wie die Box** (ein Betriebsmodell, ein Know-how, bewährte Configs). Lemonade-C++ nur, falls Whisper+TTS+LLM aus einer Hand gewünscht |
|
||||
| **MC3-Vollbau** | ⛔ außerhalb des Review-Scopes — eigenes Projekt (Prototyp steht unter #mc3) |
|
||||
|
||||
---
|
||||
|
||||
## 8. Nachtrag 3: Trennung, Update-Playbook, Radikal-Aufräumen (02.07., Abend)
|
||||
|
||||
**Lucy = eigenes Repo:** `F:\Coding Stuff\lucy` ↔ Gitea `Hitonabi/lucy` (privat; Repo via API angelegt). Verzeichnisnamen unverändert (lucy-desktop/ + lucy-tts/) → alle Pfade/BATs funktionieren; Boot aus neuem Pfad verifiziert. MC2 = reiner Box-Stack (+ hermes-pc-Executor). lucy-f5 (Verdikt final) und hermes-voice (Vor-Lucy) gelöscht — Historie bleibt in MC2.
|
||||
|
||||
**Hermes-Update-Playbook LIVE + E2E-verifiziert (7/7 PASS):** Update-Job = Backup → update → **doctor** → Restart → erweiterter Postcheck (**Config-Drift-Scan** im Journal, **Tool-Smoke** via echtem Agenten, **Voice-Smoke** mit Retries). Update-Modal fasst anstehende Commits künftig per fast-Modell zusammen (Breaking Changes zuerst, gecacht). Beim E2E-Test sofort geliefert: doctor fand eine **ausstehende Config-Schema-Migration** (altes flaches `model: hermes` → neues Schema; per `doctor --fix` migriert, Gateway verifiziert) und der Voice-Smoke entlarvte eine **pipefail/SIGPIPE-Falle** im eigenen Check (head schließt Pipe → curl 23 → Fehlalarm; gefixt). Außerdem behoben: `approvals.mode auto→smart` (v0.18-Regression, einfache Fragen 25 s → 2,0 s).
|
||||
|
||||
**Radikal aufgeräumt:** Git: lucy-v2 gemerged + gelöscht (lokal+Gitea), nur noch `main`; alter WIP-Stash gedroppt; stale Worktree entfernt. Lokal: lucy-f5-Assets (~5 GB), lemonade-eval, box_recon/gemma_swap-Scratch gelöscht. Box: 22-GB-Duplikat `Qwen3.6-35B-A3B-GGUF` (non-MTP) + 6 leere Modell-Hüllen gelöscht, v1-Altlasten (mission-control-*.db/json) nach mc2-backups/v1-legacy archiviert, /tmp-Bench-Scratch weg. Modell-Verzeichnis = exakt die 9 referenzierten Modelle + drafts + Backups.
|
||||
|
||||
---
|
||||
|
||||
*Erhoben am 02.07.2026 durch Claude Code (Live-SSH auf Box, lokale PC-Prüfung, 3 Codebase-Agents, 3 Web-Recherche-Runden). Messwerte: /api/voice/metrics (n=13 STT, n=18 chat_ttfb), TTFT-Direktmessung llama-swap :8080.*
|
||||
@@ -0,0 +1,59 @@
|
||||
# RUNBOOK — Die Box in 1 Seite (für Menschen, ohne KI-Hilfe)
|
||||
|
||||
**Grundsatz:** Die Box wartet sich selbst. Du bekommst Telegram-Nachrichten und antwortest
|
||||
höchstens „mach". Dieses Blatt ist NUR für den Fall, dass etwas klemmt.
|
||||
|
||||
## Was die Telegram-Meldungen bedeuten
|
||||
|
||||
| Meldung | Bedeutung | Dein Handgriff |
|
||||
|---|---|---|
|
||||
| „… aktualisiert … grün" | Update eingespielt, alles geprüft | keiner |
|
||||
| „… zurückgerollt … GEPINNT" | Update war schlecht, alte Version läuft wieder | keiner (läuft stabil weiter) |
|
||||
| „KRITISCH: …" | Update UND Rollback kaputt | siehe „Box tot?" unten |
|
||||
| „🛰️ Evolution-Radar …" | monatlicher Chancen-Report | lesen; bei Interesse „mach" antworten |
|
||||
| „[Werkstatt] … merge oder verwerfen?" | Box hat einen Fix vorbereitet | mit „merge" oder „verwerfen" antworten |
|
||||
|
||||
## Box tot / Weboberfläche weg?
|
||||
|
||||
1. **Strom/Netz prüfen**, dann Box **einmal neu starten** (Power-Knopf). Alles startet von selbst
|
||||
(systemd, reboot-fest). 2–3 Minuten warten, dann `http://192.168.178.151:9001` aufrufen.
|
||||
2. Immer noch tot → per SSH (PC, PowerShell): `ssh hitonabi@192.168.178.151`
|
||||
dann: `bash ~/mission-control-v2/deploy/restore.sh` (nimmt automatisch das letzte Backup,
|
||||
liegt in `/srv/models/mc2-backups/`, 14 Tage Vorrat, täglich 03:30 Uhr).
|
||||
3. Totalschaden (neue Platte/Hardware) → `docs/DISASTER_RECOVERY.md` (Bootstrap von Null).
|
||||
|
||||
## Lucy (am PC)
|
||||
|
||||
- Start: Desktop-Verknüpfung **„Lucy"** (startet den eingefrorenen Produktiv-Build).
|
||||
- Hängt? `F:\Coding Stuff\lucy\lucy-desktop\Lucy-Neustart.bat` doppelklicken.
|
||||
- Lucy ist EINGEFROREN — Änderungen macht nur die Werkstatt (Telegram-Vorschlag abwarten).
|
||||
|
||||
## Automatik-Fahrplan (läuft ohne dich)
|
||||
|
||||
- **Täglich 03:30** Backup · **So 04:30** Auto-Update (Router→Engine→Hermes) mit Rollback+Pin
|
||||
- **Monatlich 1., 09:00** Evolution-Radar-Report auf Telegram
|
||||
|
||||
## Pinnwand: eine Ebene ist „GEPINNT" — was heißt das?
|
||||
|
||||
Ein Update hat den Selbsttest gerissen; die Box bleibt bewusst auf der alten Version. Das ist
|
||||
ein STABILER Dauerzustand, kein Fehler. Pin ansehen / lösen (per SSH):
|
||||
|
||||
cat /srv/models/mc2-pins.json
|
||||
jq 'del(.hermes)' /srv/models/mc2-pins.json > /tmp/p && mv /tmp/p /srv/models/mc2-pins.json
|
||||
# (statt .hermes: .engine oder .swap) — nächster So-Lauf versucht das Update erneut
|
||||
|
||||
## Einmalige sudo-Session (steht noch aus — schaltet Engine/Router-Auto-Update frei)
|
||||
|
||||
ssh hitonabi@192.168.178.151
|
||||
sudo install -m 0440 -o root -g root ~/mission-control-v2/deploy/sudoers-mc2-autonomie /etc/sudoers.d/mc2-autonomie && sudo visudo -c
|
||||
sudo apt-get install -y unattended-upgrades && sudo dpkg-reconfigure -plow unattended-upgrades
|
||||
sudo systemctl disable --now mission-control.service # alte v1-Leiche entfernen
|
||||
|
||||
Bis dahin meldet der So-Lauf Engine/Router-Updates nur („wartet — sudo-Freischaltung fehlt").
|
||||
|
||||
## Nützliche Handgriffe (SSH)
|
||||
|
||||
curl -s http://127.0.0.1:9001/api/health # Gesamtzustand (brain ready?)
|
||||
bash ~/mission-control-v2/deploy/autoupdate.sh # Update-Lauf sofort statt Sonntag
|
||||
bash ~/mission-control-v2/deploy/notify.sh "test" # Meldeweg testen (muss auf Telegram ankommen)
|
||||
tail ~/mc2-notify.log # was wurde zuletzt gemeldet
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user