Real-time voice conversation
Streaming ASR → LLM → TTS with wake word, natural turn taking, and interruption support for responsive conversations.
Streaming
Reasoning
Speaking
Built with Agora Conversational AI
Build connected devices that listen, reason, speak, and act with Agora Conversational AI, flexible model providers, tools, and device operations in one platform.
Agora Conversational AI signal path
A live path from physical input to an intelligent response.
Physical device
Microphone + speaker + sensors
Vibbo
Agent + device orchestration
Agora Conversational AI
Realtime conversation intelligence
Tools & actions
Device + cloud execution
One continuous voice loop
From microphone input to a spoken response and real-world action.
Wake word, streaming audio, and VAD designed for voice input in real rooms.
Agora Conversational AI keeps conversations responsive across changing networks.
Choose the LLM, ASR, TTS, voice, memory, and persona for every agent.
Stream natural speech, handle interruptions, and call device or cloud tools.
A complete real-time voice stack, ready for your product
Capabilities
Vibbo connects streaming speech recognition, your model, speech synthesis, tools, and hardware operations around Agora Conversational AI.
Streaming ASR → LLM → TTS with wake word, natural turn taking, and interruption support for responsive conversations.
Streaming
Reasoning
Speaking
Choose speech recognition, language models, and voices per agent. Keep the hardware integration unchanged.
Streaming speech recognition
Configured for this agent
OpenAI-compatible model
Configured for this agent
Neural voice
Configured for this agent
Register boards, group them by agent, and roll firmware out over the air.
ESP32-S3
Kitchen · firmware 1.6.2
Desk companion
Studio · firmware 1.6.2
Voice speaker
Lab · firmware 1.5.9
Inspect every turn from transcript to tool result, then trace what the agent heard, decided, and said.
Voice stream
Realtime
Turn state
Listening
Tool calls
1
Dim the living room lights a little.
Done — the living room lights are now at 30%.
System architecture
A replaceable hardware-to-cloud path connects the microphone to a live AI agent, while Vibbo keeps every operational control in one place.
MCU + microphone + speaker + display
Wake word, audio codec, connectivity
WebSocket / MQTT / MCP
Agora Conversational AI + ASR + LLM + TTS + tools
Built for hardware
Use one voice platform across products that need to listen, reason, speak, and act in the physical world.
Create expressive desktop companions and ambient devices with persistent personalities, voices, memory, and tools.
Add natural voice control and tool execution to hubs, appliances, panels, and room devices.
Build interactive learning devices that listen, explain, ask follow-up questions, and adapt their voice experience.
Give connected toys and robots safe, configurable characters without embedding the whole AI stack in firmware.
Quickstart
Move from a development board to a managed voice agent without rebuilding the live audio and operations stack.
Use the prebuilt firmware for ESP32-S3, C3, or P4 — or compile your own from source.
The board joins Wi-Fi and shows a six-digit code. Enter it in the console to bind it to your account.
Pick the language model, speech provider, voice, and persona. Changes apply on the next turn.
Wake the device and speak. Every session, tool call, and transcript stays inspectable.
FAQs
Create an agent, connect a device, and bring a production voice experience to your hardware.
Voice AI product updates
Get product updates, hardware integration notes, and new voice AI capabilities.