Hitonabi
d887a6c248
Report: Nachtrag 3 — Lucy-Trennung, Update-Playbook E2E gruen, Radikal-Aufraeumen
...
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com >
2026-07-02 16:11:43 +02:00
Hitonabi
2a4eb51bd7
Kandidaten live testbar: gpt-oss-120b (2,3x schneller als heavy) + VL-30B in Config (Review P3-17 abgeschlossen)
...
Bench (Box, Vulkan, 32k ctx): gpt-oss-120b tg 54-55 t/s vs. Qwen3.5-122B 23,5 t/s bei
60 statt 73 GB. Beide OHNE Alias-Wechsel deployt — Qualitaets-Entscheid beim User.
brains-Gruppe nach Reload wieder angewaermt (hermes/embed/vision ready).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com >
2026-07-02 14:52:22 +02:00
Hitonabi
a16287e19c
Report: Live-Deploy verifiziert — parallel-2 wirkt (echte Parallelitaet), neue Session 0,85s/first_content 803ms (vorher 5,6-12s)
...
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com >
2026-07-02 12:08:57 +02:00
Hitonabi
f88c2a95c1
Report: VL-30B-Bench (90-92 t/s, Vision-Upgrade-Kandidat) + gpt-oss-Handoff dokumentiert
...
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com >
2026-07-02 12:01:15 +02:00
Hitonabi
ed55612cfa
Report: Lemonade-Verdikt — nicht als Unterbau; GPU-Klon besser mit llama.cpp-Vulkan + llama-swap (Box-Stack)
...
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com >
2026-07-02 11:59:52 +02:00
Hitonabi
852e723cc8
Report: Nachtrag 2 — P2/P3-Vollausbau dokumentiert (Turn-Detection, F5-Finale, Electron 43, Mem0 v3, Kandidaten, Lemonade)
...
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com >
2026-07-02 11:57:44 +02:00
Hitonabi
68ad291a8f
P2: Vision im Stream (C2), TTS-Warmup-Retry (L2), parallel-2-Config vorbereitet
...
- voice.py: Bildschirm-Sicht laeuft IM SSE-Stream (+hermes.vision.progress-Event fuer
Warte-Ansage), Vision-Timeout 120->45s (MC_VISION_TIMEOUT); deployt + Smoke-Test ok
- useVoiceAgent: Warm-Gate mit Retry/Backoff + sichtbarer Meldung statt stummem Fake-ready
- deploy-Config: hermes -c 131072 --parallel 2 (+KV Q8_0) — Mem0-Extraktion blockiert
Voice-Turns nicht mehr; speicherneutral. Live-Schaltung: User
- Report: P2-Nachtrag; C5 (pricing/token_stats) war Fehlalarm — in Benutzung
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com >
2026-07-02 11:07:00 +02:00
Hitonabi
e08d812598
Brain-Bench + Latenz-Metrik (Review P1-9/P1-10)
...
- Bench-Matrix Qwen3.6 (5 Configs, separater Port): MTP n-max 3 bestaetigt (+26% tg),
n-max 4 lohnt nicht, KV Q8_0 gratis (78,6=78,6 t/s) bei halbem KV-Speicher
- deploy-Config: -ctk/-ctv q8_0 am hermes-Eintrag (Live-Schaltung: User-Freigabe noetig)
- voice.py: chat_first_content-Metrik (echte Hirn-Latenz bis erster Inhalts-Token)
- Report um Umsetzungs-Nachtrag ergaenzt
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com >
2026-07-02 10:56:23 +02:00
Hitonabi
64efe18500
Docs: Komplett-Review 2026-07-02 (IST live verifiziert + Roadmap P0-P3) + Session-Docs
...
- REVIEW_2026-07-02.md: Latenz-Baseline (STT 2s dominant, LLM-TTFT 65ms), chat-Lane-Bug,
SOLL-Recherche (Parakeet v3, Silero VAD v6, KV-Quant, MoE-Spec-Trap), Verdikte bestaetigt
- Audit-/TTS-/ZeroClaw-/DR-Docs von main nachgezogen; launch.json
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com >
2026-07-02 10:30:12 +02:00