Sharing a project I help maintain — Flowcat, an Apache-2.0 native-Rust runtime for real-time voice AI agents (phone + WebRTC). It runs the whole voice pipeline (STT → LLM → TTS, or a single speech-to-speech model) as one self-contained binary you host yourself — in your own VPC or fully air-gapped. No hosted control plane, bring your own keys, and a call’s audio never leaves your infra.
It’s a clean-room counterpart to pipecat’s architecture, but native Rust:
- each pipeline stage is its own tokio task behind a bounded mpsc channel (natural backpressure)
- the hot audio frame is an
Arc<AudioFrame>, so each hop moves a pointer, not a PCM copy - no GIL, so one process uses all cores and holds a flat tail under load (measured: flat ~0.6ms p99 framework overhead from 10 to 2,000 concurrent calls on one box)
- native SIP/RTP, so a single binary can terminate phone calls with no FreeSWITCH
The linked writeup has the full Rust design + a reproducible benchmark vs pipecat (with the honest caveat that it’s framework overhead, not end-to-end voice latency — that’s provider-bound).
Repo: https://github.com/AreevAI/flowcat
Apache-2.0, pre-1.0, built in the open. Disclosure: I help maintain it — happy to talk architecture or the SIP stack.



Putting vibe negativity aside, I checked what the local STT and TTS “support” looked like.
STT is based on rust bindings for
whisper.cpp(2024 called!). And the three local TTS models claim is total bullshit. One is using an external service from a docker image, and the other two are trivial web clients that are not even wired. I also loved the circular reasoning at the start of the three TTS modules. Even an LLM is not something I would have pegged as that bad at logic. But maybe they got that way because of the prompter.So ZERO interesting shit is happening here, vibed or otherwise. Nothing even close to using rust-based ML engines (candle, burn, tract), not even in a completely broken untested way.