Hitonabi
|
e08d812598
|
Brain-Bench + Latenz-Metrik (Review P1-9/P1-10)
- Bench-Matrix Qwen3.6 (5 Configs, separater Port): MTP n-max 3 bestaetigt (+26% tg),
n-max 4 lohnt nicht, KV Q8_0 gratis (78,6=78,6 t/s) bei halbem KV-Speicher
- deploy-Config: -ctk/-ctv q8_0 am hermes-Eintrag (Live-Schaltung: User-Freigabe noetig)
- voice.py: chat_first_content-Metrik (echte Hirn-Latenz bis erster Inhalts-Token)
- Report um Umsetzungs-Nachtrag ergaenzt
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-02 10:56:23 +02:00 |
|
Hitonabi
|
64efe18500
|
Docs: Komplett-Review 2026-07-02 (IST live verifiziert + Roadmap P0-P3) + Session-Docs
- REVIEW_2026-07-02.md: Latenz-Baseline (STT 2s dominant, LLM-TTFT 65ms), chat-Lane-Bug,
SOLL-Recherche (Parakeet v3, Silero VAD v6, KV-Quant, MoE-Spec-Trap), Verdikte bestaetigt
- Audit-/TTS-/ZeroClaw-/DR-Docs von main nachgezogen; launch.json
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-02 10:30:12 +02:00 |
|