What is Gandr
Gandr is a text-to-speech and voice-cloning platform priced per concurrent stream, offering flat, unlimited usage without per-character metering. It supports real-time audio, instant voice cloning from about ten seconds of reference audio, and 23 languages on one engine. Gandr is designed for audiobooks, video dubbing, e-learning, podcasts, games, and voice agents, with streaming-first audio measured server side.
How to use Gandr
- Sign in to Gandr to get an API key.
- Use the REST, SSE, or WebSocket APIs to integrate text-to-speech into your application.
- For voice cloning, provide a reference clip of about ten seconds; the clone is created instantly and cached after first use.
- Choose from 23 languages and control expression and prosody.
- Stream audio in real-time with low latency (first audio byte in 146 ms over the open internet).
Features of Gandr
- Unlimited text-to-speech: No per-character or per-minute metering; flat pricing per concurrent stream.
- Instant voice cloning: Clone a voice from a ~10-second reference clip, no training job or per-voice fee.
- 23 languages: One engine supports multiple languages, with the same voice identity across languages.
- Low latency: First audio byte in 146 ms over the open internet, with p50 first audio at 116 ms server-side warm.
- Streaming-first: Audio is streamed in real-time, suitable for conversational and long-form content.
- Expression controls: Adjust prosody and expression on a single engine.
- Watermarked audio: Every clip is invisibly watermarked for provenance and EU AI Act Article 50 transparency.
- Privacy: Nothing you send enters a training set.
- API options: REST, SSE, and WebSocket APIs.
- Integrations: Works with LiveKit, Pipecat, Vapi, Retell, Daily, and Whisper.
- Security: SOC 2 Type 2, HIPAA-verified, GDPR-ready infrastructure.
Use Cases of Gandr
- Audiobook narration: Long-form narration with cloned voices.
- Video dubbing: Localize videos across 23 languages.
- E-learning narration: Generate voiceovers for educational content.
- Podcasts and video voiceover: Produce professional voiceovers.
- Games: Add voice to characters and narratives.
- Voice agents: Build conversational AI with real-time streaming TTS.
- Customer support: Automate support calls with natural-sounding voices.
- Live translation: Provide real-time voice translation.
Pricing
Gandr is priced per concurrent stream, flat and unlimited. The price is $150 per stream per month, with one thing talking at a time, unlimited speech, and no per-character metering. There is a free tier to start, and pricing is discussed on a call for larger volumes (from 500 lines).
FAQ
How does voice cloning work?
Provide a reference clip of about ten seconds; the clone is created instantly and cached after first use, with no training job or per-voice fee.
What languages are supported?
Gandr supports 23 languages, including English, Spanish, French, German, Portuguese, Arabic, Chinese, and Japanese.
What is the latency?
First audio byte is delivered in 146 ms over the open internet, with p50 first audio at 116 ms server-side warm.
Is my data used for training?
No, nothing you send enters a training set.
Is the audio watermarked?
Yes, every clip is invisibly watermarked for provenance and EU AI Act Article 50 transparency.
What integrations are available?
Gandr integrates with LiveKit, Pipecat, Vapi, Retell, Daily, and Whisper.
What are the security certifications?
Gandr runs on SOC 2 Type 2, HIPAA-verified, and GDPR-ready infrastructure.