Voice Agents

Real-time speech agents. Latency budgets are brutal — interruption handling and turn-taking usually decide whether the result feels usable.

3 projects

PipecatVoice Agents

Assembles the whole speech-to-speech loop — transcription, model, synthesis, interruption handling — as a pipeline of swappable services, which is the tedious part to get right. Latency is the entire game in voice, and turn-taking and barge-in are first-class here rather than bolted on. Running it well means caring about transport and network topology, not only which model you picked.

OfficialBSD-2-Clause14.9k+512
LiveKit AgentsVoice Agents

Builds server-side agents that join LiveKit rooms as programmable participants, with integrated WebRTC transport, dispatch, telephony, turn detection and provider plugins. It is a strong fit when media transport and production session orchestration need to work as one system. Teams that only need a text agent or want a transport-neutral pipeline may find the LiveKit architecture more infrastructure than necessary.

OfficialApache-2.013.3k+180
TEN FrameworkVoice Agents

It covers the parts of a spoken agent that sit outside the model and usually get assembled by hand: detecting when someone is speaking, deciding when a turn has ended, separating speakers, connecting to the phone network, driving a lip-synced avatar, even reaching an embedded device. Read the licence before building on it. Apache 2.0 is qualified by additional conditions from Agora that prohibit hosting the framework on end-user devices, mobile terminals included, and prohibit deploying it in a way that competes with Agora's own offerings.

OfficialSource-available11.1k+22