AI Audio · text to speech

Text to Speech

Text to Speech は「AI text to speech」向けの LipSyncing AI ページです。ホームの使い回しではなく、スタジオがこの用途のモードで開きます。左メニューの AI Audio に属します。ファイルを上げ、台本か音声を足して生成します。

Text to speech

リップシンクを生成

150 credits · $1.50

Pick a voice from the library, paste a script, then generate. My voices are clones and designs you saved. Public voice AI is the catalog for this engine.

My voices0

No saved voices for this engine yet. They appear here after you clone or design one. You can still pick a public voice below.

Public voice AI10

Staff-curated public catalog for this engine, plus live voices when a key is pasted. Play a sample here. Open Text to Speech when you want a generated track.

Showing the built-in catalog. Staff can edit it on Admin → Voice Library.

Text to speech runs on ElevenLabs, Cartesia, and Chatterbox (RunPod). Pick a voice from My voices or public voice AI. Paste keys on Admin → Models.

See an example

InfiniteTalk example from Wavespeed. Sample generated by this tool’s model. Your clip uses your photo, script, or prompt.
Voice sample ready to clone
Voice sample ready to clone
Script turned into speech
Script turned into speech
Audio attached to a talking avatar next
Audio attached to a talking avatar next

この text to speech ページがある理由

My voices, then public AI

Clones and designed voices you saved sit first. Public catalog voices sit second. Pick one and generate.

Emotion without a second take

Calm, bright, or urgent — match the landing page, not a flat read.

Made to hand off

The file is an input for talking photo, talking avatar, dubbing, and singing (spoken intros).

Script-length honesty

TTS is cheap compared with a crew. The video job still bills the 30-second block plus extras.

使い方

  1. Step 1: record or upload a sample, or pick a voice
    ステップ 1

    Pick a voice from the library

    My voices are clones and designs you already saved. Public voice AI is the catalog.

  2. Step 2: paste the script
    ステップ 2

    Paste the script and generate

    Listen before you spend a talking-video job.

  3. Step 3: generate speech for lip sync
    ステップ 3

    Send to a face tool

    Talking avatar, photo lip sync, or video lip sync.

What this clip costs in credits

Every text to speech job uses the same meter: 150 credits ($1.50) for the first 30 seconds, then 5 credits ($0.05) per extra second. Subscriptions price credits at $0.01; one-time packs are $0.08 each.

DurationcreditsGeneration price
15s150$1.50
30s150$1.50
45s225$2.25
60s300$3.00
90s450$4.50
120s600$6.00

First 30 seconds = 150 credits ($1.50). Plan credits roll over while you stay subscribed.

使われる場所

Avatar reads

Generate audio, then open talking avatar on the same script.

Dubbing drafts

TTS in the target language before you hire a human mixer.

IVR-adjacent help clips

Consistent library voice across 40 macros.

text to speech の質問

関連キーワードページ

それぞれ別の検索に向けています。左メニューで同じクラスターに留まれます。