# TikTok character (audio + text)

> AI video is the slowest, most expensive, most failure-prone part of the recommended stack. Cut it. Use one consistent character image (commission once or generate once) plus ElevenLabs voice and beat-cut typography in CapCut. Ship 3x faster, half the cost.

Source: https://aistack.sh/stack/tiktok-character-audio
Category: Video · Creator
Level: Beginner
Last verified by a human: 2026-05-05

## The tools

### Claude — Character bible + script

- Pricing: Free · $20/mo Pro · $100/mo Max 5x · $200/mo Max 20x · Sonnet API $2/$10 per M tokens (input/output)
- Site: https://claude.ai
- Why this one: Same as the recommended stack. The character lives in the voice and the writing, not the visuals.
- Swap in instead: ChatGPT
- More: https://aistack.sh/tool/claude

### ElevenLabs — Character voice

- Pricing: $6/mo Starter · $22/mo Creator
- Site: https://elevenlabs.io
- Why this one: Even more important here than in the Pika version: the voice IS the character. One marketplace voice or clone, locked.
- More: https://aistack.sh/tool/elevenlabs

### CapCut — Typography + edit

- Pricing: Free · $9.99/mo Standard · $19.99/mo Pro
- Site: https://capcut.com
- Why this one: Animated text on the character image carries the visual. CapCut's keyframe text is the bottleneck-free way to do this.
- Swap in instead: Descript
- More: https://aistack.sh/tool/capcut

## Monthly cost

### $22/mo — 1 ep/day (small)

- Claude: Free
- ElevenLabs: $22
- CapCut: Free

### $42/mo — Daily + 3 shorts (medium)

- Claude: $20
- ElevenLabs: $22
- CapCut: Free

### $210/mo — Multi-character agency (heavy)

- Claude: $80 API
- ElevenLabs: $99
- CapCut: $15
- + misc: $16

## Workflow

### 1. Lock the bible (no visual section) (Claude)

Same prompt as the recommended stack, minus 'visual style'. Replace it with a single character image and a typography lock-up.

**Prompt: Audio-first character bible**

```
Build a character bible for a recurring 60-second character account where the visuals are: one fixed character image, animated typography, and short cutaways. NO AI video generation.

Working idea: {{1-line concept, e.g. "a sarcastic raccoon who reviews tech"}}
Niche: {{e.g. consumer tech, finance for Gen Z, parenting fails}}

Output:

1. **Persona** (120 words): name, age, occupation, where they "live", what they care about, what they hate.
2. **Voice** (100 words): cadence, pet phrases (list 5), what they NEVER say, sample line. THIS IS THE PERSONA — write more here than for video-driven characters.
3. **Visual lock-up** (60 words): the one character image (description, art direction). The font for typography. The 2-color palette. Recurring meme/sticker.
4. **Recurring beats** (5): episode formats this character rotates through.
5. **Hard rules** (list 5): things the character must never do.

Save as project context for every script.
```

### 2. Daily script (Claude)

Same script structure as the Pika version, but every line gets a typography cue instead of a B-ROLL cue.

**Prompt: 60s script with typography cues**

```
Use the character bible from this project.

Today's beat: {{e.g. "rant review of Apple's new Vision Pro update"}}
Source material:
"""
{{paste 1 to 2 articles, tweets, or screenshots-as-text}}
"""

Write a 60-second script (about 150 spoken words) in character voice.

Structure:
- 0:00 to 0:03 — Hook line.
- 0:03 to 0:45 — Body. 3 beats. One concrete fact or quote per beat.
- 0:45 to 0:60 — Payoff line.

Format the output as:
[0:00] Spoken line
[TEXT: what appears on screen — a 2-to-5-word punch from the line, NOT the full line]

Stay under 160 words. No "subscribe for more". No filler.
```

### 3. Voice (ElevenLabs)

ElevenLabs renders the script as one audio file. Same voice, same settings, every episode.

### 4. Typography pass (CapCut)

CapCut: drop voice on track 1, character image as background. Punch text per [TEXT] cue, beat-aligned, with 2 fonts max from the lock-up.

### 5. Ship (CapCut)

Upload at 6pm local. Whole loop ships in 25 minutes once the bible exists.

## What it produced

**Audio-first character account, 4 months**

Same niche as the Pika version (sarcastic tech raccoon). 110k followers vs 90k for the AI-video version, mostly because they shipped 1.7x more episodes (no Pika render bottleneck). Visual drift was a non-problem.

## Pitfalls

- **Visuals get static** — One character image only works if the typography pulls weight. Spend the time on the typography lock-up; it's the difference between 'cheap' and 'identifiable'.
- **Voice fatigue** — Audio-first means the voice is everything. Re-clone every 60 days; pin a sample line for QA. Same as the Pika version, but more load-bearing here.

## Recent changes affecting this stack

- **Scribe v2 Medical speech recognition model released** (2026-09-11, release) — Scribe v2 Medical is now generally available as a specialized batch speech recognition model for medical and clinical audio. It maintains the same billing rate as Scribe v2 and supports keyterm prompting, entity detection, speaker diarization, and no-verbatim mode. Pass scribe_v2_medical as the model_id parameter. Source: https://elevenlabs.io/docs/changelog
- **Agent response attachments in ElevenAgents conversations** (2026-09-07, feature) — agent_response WebSocket events now include an optional attachments field, allowing agents to send files with metadata including URL, filename, and MIME type. Source: https://elevenlabs.io/docs/changelog
- **Twilio outbound calls now support answering machine detection** (2026-09-07, feature) — Outbound calls through Twilio can now detect whether they reach a human or answering machine. Configure detection mode to trigger early detection or wait for voicemail greetings to complete. Results arrive via a new answering_machine_detection webhook event. Source: https://elevenlabs.io/docs/changelog
- **ElevenAgents WebSocket responses now support file attachments** (2026-09-07, feature) — Agent response WebSocket events include optional attachments with URL, name, and optional MIME type. Enables agents to send files and documents to users during conversations. Source: https://elevenlabs.io/docs/changelog
- **Twilio answering machine detection added to ElevenAgents outbound telephony** (2026-09-07, feature) — Outbound telephony configurations now support Twilio answering machine detection with two modes: early human-or-machine verdict or waiting for voicemail greeting to finish. Detection results are delivered via the new answering_machine_detection webhook event. Source: https://elevenlabs.io/docs/changelog

## Other ways to do this

- [TikTok character](https://aistack.sh/stack/tiktok-character): Recurring character, daily uploads — $92/mo

---

Curated by ryan-c on aistack.sh. Last updated 2026-05-05.
