Request for Text-to-Speech and Video Generation APIs

Hi Anna Team,

We are building an AI short-video app on the Anna AI Platform. The workflow includes script generation, voice synthesis, video generation, subtitles, and final video composition.

We would like to request two native capabilities:

  • Text-to-Speech API: multiple voices and languages, adjustable speaking styles, common audio formats, and timestamp support for subtitles.
  • Video Generation API: text-to-video, image-to-video, reference-image support, multiple aspect ratios, configurable duration, and asynchronous task status APIs.

Currently, using third-party services requires separate authentication, billing, file management, and error-handling logic. Native Anna APIs would make multimedia app development much simpler and more reliable.

Are these APIs currently planned? We would appreciate any information about the roadmap, pricing, supported models, and commercial usage rights.

Thank you!