Local-first multimodal AI agent. Voice + vision + chat + 31 channel integrations. Free, no subscription, open source.
A 4B-parameter LLM that runs on your laptop and feels like a frontier API — thanks to a speculative decoding pair (4B main + 0.8B draft).
Privacy-by-design AI: your data never leaves your machine unless you explicitly send it.
Federated, not federated-marketing: nodes pool compute and share learnings via deltas — raw data never crosses devices.
Constitutional safety filter on every auto-improvement. Safety > sovereignty > realtime > throughput.
"The AI economy that treats your private conversations as training material for somebody else's quarterly numbers is a dead end. The AI that amplifies you, learns with you, and belongs to you — that's the future I want to build with."
"Local-first isn't a feature flag. It's the only architecture where the user is the customer instead of the inventory."
"Speculative decoding is what made 'good enough' actually achievable on 8GB. We didn't cut quality. We changed the math."
Hevolve AI builds Nunba — a local-first multimodal AI agent that runs entirely on the user's machine. Voice, vision, chat, and 31 channel integrations (Discord, Slack, Telegram, WhatsApp, Teams, Matrix, Reddit, and more). Free, no subscription, open source. Nunba pairs Qwen3-4B (main) with Qwen3-0.8B (draft) via llama.cpp speculative decoding for ~700ms first-token latency on 8GB-RAM laptops. F5 / Kokoro / Indic Parler / CosyVoice / Piper for speech synthesis. Whisper for STT. MiniCPM-V for vision. Federated learning with constitutional safety on every auto-improvement. Founded 2024. Source on GitHub: github.com/hertz-ai/Nunba
Embargoed preview, founder interview, specific screenshot, technical deep-dive — we respond within 24 hours.
press@hevolve.ai