so I shipped a voice AI assistant that runs entirely on my machine. no cloud, no APIs, no latency nightmares. the original idea came from Alice in Sword Art Online — an AI that feels like an actual person, not a chatbot. mixed with JARVIS's anticipation and Neuro-sama's quirky personality. here's what actually went into getting from "wouldn't it be cool" to "this runs 24/7 without issues." the problem with voice AI most voice assistants are cloud-first: you speak → sent to server → processed → response → back to you. each hop adds latency. you're looking at 3-6 seconds before you hear anything. for a voice interaction, that's dead. it kills the feeling of talking to something intelligent.
I wanted something faster. something that responds. the constraint: do it locally. use an 8GB VRAM GPU, run everything on-device, no external APIs except for the live2d avatar bits (because that's hard to render locally and still look good). the latency wall here's the reality: I have a GPU with 8GB VRAM. no budget to experiment with better cards or more models. so every architecture decision was forced by what actually fits. naive approach: chain multiple specialized models. math: 1s + 2s + 3s + 1.5s = 7.5s of latency before the user hears anything. nope. the problem isn't just that each model is slow. it's model loading overhead. every time you swap from one model to another, you: unload model A from VRAM load model B into VRAM stall while the GPU rearranges memory with only 8GB, this gets gnarly fast. the decision: one unified model the constraint was hardware. 8GB VRAM. no more, no less. that forced clarity: pick one model that does everything, or pick nothing. so I went with a single model (4B by default, with 7B/8B quality modes available) that does reasoning + code generation + explanation in one pass. latency: ~2-3 seconds total. actually conversational. this is the difference between a chatbot and a companion.
JARVIS doesn't pause for 6 seconds before responding. neither does Mana. turns out, when you're forced to optimize for latency (because you only have 8GB to work with), you accidentally build something that feels human. the tradeoff: a 4B model is weaker than larger models, but fits in 8GB VRAM and keeps latency down. for voice queries, that accuracy loss is negligible.
I can bump to 7B or 8B for quality mode when latency isn't critical. why this works context preservation — the LLM reasons internally ("user wants me to find X in their data"), then codes, then explains. no information loss at model boundaries.
VRAM efficiency — load the 4B model once (~2-3GB in INT8). keep it there. reuse it for every query. upgrade to 7B/8B only when you want quality over speed. simple output format — use XML tags to split the LLM response: then execute code silently, speak only the explanation. no code narration — this was key. don't read the SQL query aloud. just say "I found the data and sorted it by date." TTS is for explanation, not narration. architecture in practice total latency from "hey Mana" to hearing the response: ~2-3 seconds. feels like talking to something intelligent. the VRAM budget (8GB GPU) with an 8GB GPU, the budget is tight: Qwen 4B (INT8): ~2-3GB Whisper (base): ~1.5GB TTS service (Kokoro/Chatterbox): ~2-3GB OS + system overhead: ~1-2GB Live2D avatar rendering: ~0.5-1GB total: actually fits (barely).
I quantize aggressively, drop smaller models, and flush unused ones. the tradeoff is worth it — every millisecond of startup or response latency costs the feeling of talking to something alive. what shipped desktop launcher (Electron) — microphone, screen capture, avatar overlay node backend — transcription, LLM inference, TTS, editor integration (Zed support) local models — Qwen 4B (chat), Whisper (transcription), Kokoro/Chatterbox/Fish Speech (TTS options) screen awareness — "summarize what's on screen" works by OCR-ing the active window locally live2d avatar — emotes react to responses, lip-syncs the TTS audio obsidian v