# Request: High-Quality, Model-Selectable Video, TTS, and Talking Avatar APIs

**URL:** <https://forum.anna.partners/t/request-high-quality-model-selectable-video-tts-and-talking-avatar-apis/300>\
**Category:** Developers\
**Created:** [September 11, 2026, 3:17am UTC](https://forum.anna.partners/t/request-high-quality-model-selectable-video-tts-and-talking-avatar-apis/300 "2026-09-11T03:17:03Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![ELFA](https://yyz1.discourse-cdn.com/flex033/user_avatar/forum.anna.partners/elfa/32/22_2.png) [@ELFA](https://forum.anna.partners/u/ELFA)\
**Post date:** [September 11, 2026, 3:17am UTC](https://forum.anna.partners/t/request-high-quality-model-selectable-video-tts-and-talking-avatar-apis/300/1 "2026-09-11T03:17:03Z")

</div>

## Title

Request: High-Quality, Model-Selectable Video, TTS, and Talking Avatar APIs

## Post

We are building Ad Studio as an Anna App. Text generation, image generation, storage, and job orchestration can already use Anna, but media generation still depends on external services:

- Image editing: fal GPT Image 2
- Video generation: fal Seedance 2.0 / Gemini Omni
- TTS: ElevenLabs Multilingual v2
- Video composition: FFmpeg

Anna already provides `image.generate`, `image.edit`, and `tools.invokeAsync`, so we are mainly requesting the following capabilities.

### 1. Video generation

A unified `video.generate` API supporting:

- Reference images and audio
- Native audio and lip-sync
- 9:16, 16:9, and 1:1 aspect ratios
- 720p and 1080p
- Async status, cancellation, and recovery

The quality should be at least comparable to:

- `bytedance/seedance-2.0/reference-to-video`
- `google/gemini-omni-flash/image-to-video`

### 2. Text-to-speech

Please provide `audio.speech` and `audio.listVoices`.

Quality should be comparable to `ElevenLabs eleven_multilingual_v2`, with Chinese and English support, stable voice IDs, MP3/WAV output, speaking-rate controls, and audio-duration metadata.

### 3. Talking avatars and lip-sync

We would like Anna to support models comparable to:

- `fal-ai/heygen/avatar4/image-to-video`
- `fal-ai/bytedance/omnihuman/v1.5`
- `fal-ai/sync-lipsync/v3/image-to-video`
- For existing-video lip-sync: `fal-ai/kling-video/lipsync/audio-to-video`

Basic mouth-only models such as Wav2Lip or MuseTalk should not be the default implementation, as they often lack natural expressions, body motion, and reliable identity consistency.

### 4. Strict model selection

This is the most important requirement. Please support something like:

```auto
{
  "modelPreferences": {
    "hints": [{"name": "HeyGen Avatar 4"}],
    "strict": true
  },
  "fallbackPolicy": "error"
}

```

If the requested model is unavailable, the API should return an explicit error instead of silently falling back to a lower-quality model.

Responses should also report the actual `provider`, `model`, `usage`, and `cost`.

Our goal is to use Anna’s unified quota, credentials, storage, and auditing while still achieving media quality suitable for commercial advertising.
