Speech-driven Avatar
Send speech text, a Ruwana voice, speech rate, emotion and optional performance direction. Speech supports up to 3000 characters.
Ruwana Avatar combines an avatar image with exactly one audio source: generated speech or uploaded finished audio. Platform measures prepared audio server-side, then runs a durable Standard 720p or Pro 1080p request.
Send speech text, a Ruwana voice, speech rate, emotion and optional performance direction. Speech supports up to 3000 characters.
Send your finished audio track instead of speech. TTS-only voice controls are not accepted when uploaded audio is used.
| Tier | Resolution | Rate | TTS |
|---|---|---|---|
| Standard | 720p | $0.11/sec | Included |
| Pro | 1080p | $0.22/sec | Included |
Move from capability discovery to the production surface or implementation guide.
Product overview, controls and pricing.
Speech/audio request rules and implementation examples.
Use movement reference video instead of speech-driven performance.