Sep 16, 2026 ElevenLabs stays Live + preplanned Likeness vs generated

Video avatars on ElevenLabs — live interviews and preplanned clips

Keep ElevenLabs as the voice. Pick a face that can do a live two-way interview and a scheduled, scripted video. Stock humans, consented clones, and synthetic faces that never existed. Tavus and HeyGen LiveAvatar in the same matrix as the OSS talking-head stack.

Companion to the orchestration report.

What to actually pick

If the interviewer must look human, live, and you already pay ElevenLabs: HeyGen LiveAvatar in LITE / Avatar-Only mode. ElevenLabs owns audio; LiveAvatar owns the face. Documented by ElevenLabs. PCM 24 kHz. 1 credit = 1 minute in Lite.[6][4][5]

If you want one vendor that is a whole conversational-video product (perception, RAG, Zoom/Meet join) and you are OK not driving the mouth from ElevenLabs by default: Tavus CVI. To keep ElevenLabs, you use Echo / Pipecat’s persona_id=pipecat0 so Tavus eats your TTS audio instead of its own voice.[7][18] That is more glue. CVI overage is ~$0.32–$0.37/min on public Starter/Growth cards — bundled STT+LLM+TTS+face, not face-only.[2]

If you want cheaper live face + still ElevenLabs: Simli (LiveKit plugin) or Anam (~$0.11–$0.16/min on Anam’s page) or Beyond Presence speech-to-video.[13][14][15][16] These are live layers. Preplanned MP4s are a second pipeline.

If the face must be OSS and you own a GPU: one renderer for both modes — LiveTalking / OpenAvatarChat + MuseTalk for live WebRTC; SadTalker / LivePortrait (or the same MuseTalk, offline) for pre-rendered interview invites and recaps. Audio in both cases = ElevenLabs file or stream.[10][11][12][8][9]

Do not clone a real person without a recorded consent path. HeyGen requires a consent clip for custom video LiveAvatars; image avatars are the synthetic-identity loophole vendors themselves document.[4][5] OSS photo-to-talking-head is the same legal surface, just without a vendor saying no.

1. Two products hiding in one sentence

Live Two-way interview

Candidate talks. Agent barges in. Face must track streaming TTS chunks, freeze on interrupt, idle while listening. Transport is WebRTC. Latency budget ~1 s to first frame.

Tavus CVI, HeyGen LiveAvatar, Anam, Simli, Beyond Presence, LiveTalking, OpenAvatarChat, MuseTalk realtime.

Tape Preplanned / scheduled

Script known ahead: invite, briefing, “you’re in 10 minutes,” score recap, rejection. Render an MP4 (or a sequence) from ElevenLabs audio + a still or short driver clip. Quality can be higher because nobody is waiting.

HeyGen Studio (separate from LiveAvatar — avatars are not cross-compatible).[4] Tavus generated-video minutes (separate from CVI).[2] SadTalker, LivePortrait, MuseTalk offline, ElevenLabs Creative talking-video.

Same face both ways?

Vendors: usually two SKUs. HeyGen Digital Twin ≠ LiveAvatar. Tavus Replica can feed CVI and generated video, but live and tape are different meters.[4][2]

OSS: one identity image/video, two inference modes. That is the real roll-your-own win.

2. Who is on camera — three identity classes

This is the split you called out: bolting a face generated from a still / a random virtual human, vs a personal likeness.

ClassHow you get itInterview useRisk
Stock / library human Vendor preset (HeyGen 100+ presets; Tavus stock replicas; Simli face IDs; Anam library).[3][20] Default interviewer. Fast. No consent theater. Looks like “an actor from a SaaS ad.” Fine for screening.
Personal likeness (clone) 2-min talking + listening footage (HeyGen video LiveAvatar) or short video/image-to-replica (Tavus). Consent recording required on HeyGen custom video; Enterprise can use pre-recorded consent.[4][5] Named hiring manager / founder twin. Recruiter who cannot be in 40 rooms. Deepfake + employment law. Need written consent, retention policy, “this is AI” disclosure.
Synthetic / generated identity FLUX / SD portrait that never existed → image-to-avatar (HeyGen image LiveAvatar, Tavus image-to-replica, SadTalker, LivePortrait). No real person.[5][8][9] Brand-consistent interviewer that is not impersonating staff. Rotate looks per role (eng vs sales) without cloning anyone. Uncanny valley; accidental resemblance; some jurisdictions still want disclosure. This is the “going away from personal likeness” move.
Stylized virtual Ready Player Me / MetaHuman / Live2D. Rigged 3D or cartoon.[19] Clearly-not-a-person. Gaming/education vibe. Wrong register for most corporate interviews unless you lean into “this is a bot.”

Practical interview policy: synthetic stock or generated face + ElevenLabs voice that is also not a real employee’s clone, unless you have signed talent. Likeness clones are a product later, not v1.

3. How ElevenLabs actually attaches

ElevenLabs is the mouth audio. The face is a consumer of that audio (or, in Full/CVI mode, a competing TTS you should not double-pay).

Keep:  STT → LLM → ElevenLabs streaming TTS
Add:   TTS audio PCM 24 kHz → avatar renderer → WebRTC video track

Official 11L path:  HeyGen LiveAvatar LITE
  ElevenLabs Agents audio  +  LiveAvatar face
  bills: 11L minutes  AND  1 LiveAvatar credit / min[6]

Pipecat/LiveKit path:  TavusVideoService persona_id=pipecat0
  means “use the bot’s TTS, echo audio, don’t use Tavus voice”[7]

Also works: Simli / Anam / Beyond Presence as a LiveKit participant
  they take the same TTS track[15][16]

Offline tape:  ElevenLabs → wav → SadTalker / LivePortrait / MuseTalk / HeyGen Studio

LiveAvatar Full mode uses ElevenLabs Flash v2.5 internally plus Deepgram/AssemblyAI — that is their ElevenLabs, not necessarily your agent voice, and it costs 2 credits/min.[4][5] For “we use ElevenLabs” as your voice, stay on Lite / Avatar Only.

4. The matrix

Option Live 2-way Preplanned MP4 11L as YOUR tts Stock Likeness Generated still Self-host Complexity
HeyGen LiveAvatar Yes. LITE or Full. First frame <300 ms claimed.[3] No — use HeyGen Studio separately; avatars not cross-compatible.[4] Yes — LITE, documented.[6] 100+ presets[3] 2-min video + consent, or photo[4] Photo → image avatar (no voice clone)[5] No Low if LITE. Medium if you also want Studio twins.
Tavus CVI Yes. Phoenix-4.5, Raven perception, Sparrow dialogue.[18] Yes — generated-video minutes on the same account, different rate.[2] Via Echo / Pipecat pipecat0, not the default.[7] Stock replica library[18] Custom replica trainings (plan-capped)[2] Image-to-replica exists as a product line No Medium. Dual-room with Daily. 30s min + 6s rounding.[2][7]
Anam Yes. Latency-first (~180 ms claimed in roundups).[20][13] Not the product BYO LLM/TTS via LiveKit/Pipecat Yes Custom uploads Custom uploads No Low–medium
Simli Yes. LiveKit plugin.[15] Not the product Yes — audio from your agent Face IDs[15] Upload a face[15] Upload a generated still No Lowest LiveKit glue
Beyond Presence Yes. Speech-to-video API, LiveKit token.[16] Not the product Yes — audio-to-video Library + custom Custom Custom No (EU-hosted) Low if you already have LiveKit
D-ID / Spatius-class Agents / streaming SKUs Classic photo-talk clips BYO audio on some SKUs Yes Photo Photo No Low for tape; live quality varies
LiveTalking OSS Yes. WebRTC / WHEP / RTMP. wav2lip / MuseTalk.[10] Same models, file out Stream ElevenLabs into TTS slot You supply the clip You supply consented footage Animate any still/video you generate Yes Apache-2.0 High. GPU, TURN, session leaks.
OpenAvatarChat OSS Yes. Duplex interrupt, MuseTalk / FlashHead / lite-avatar.[12] Possible via handlers Swap TTS handler → ElevenLabs Bundled demos Your data Your still Yes Apache-2.0 High. Full stack demo, not a plugin.
MuseTalk OSS Realtime infer exists; used inside LiveTalking.[11][10] Excellent offline Audio file/stream in n/a — renderer only Identity = your video Identity = your still+body video Yes MIT code[11] Medium as a library, high in prod
SadTalker OSS No (batch). CVPR 2023, ~14k stars.[8] Yes — the classic one-photo + wav → MP4 ElevenLabs wav in n/a Any portrait Best OSS tape path for generated faces Yes Low for a clip; not live
LivePortrait OSS Video-driven (needs a driver clip), not audio-native.[9] Yes — animate a still from a talking driver Indirect (driver from a talking video) n/a Any portrait Generated still + stock driver video Yes ~19k stars[9] Medium. Great motion, extra audio-lip step.
Open-LLM-VTuber / Live2D OSS Yes, local[19] Screen-record Swap TTS Anime models No Stylized only Yes Low technically, wrong look
MetaHuman / RPM + UE Pixel Streaming Yes if you build visemes from 11L Sequencer / Movie Render Yes — viseme / audio2face Parametric humans (not photoreal photo) Scan optional Randomize sliders = infinite virtual people Engine yes, heavy Highest. Game-studio complexity.

5. Price comparison (public cards, Sep 2026)

Face-only vs bundled. Adding ElevenLabs on Lite/Echo is extra. Numbers move; verify before a PO.

VendorWhat the minute includesPublic $Notes
HeyGen LiveAvatar Lite Face + stream only. Your 11L/LLM separate.[4][6] 1 credit / min. Starter overage $0.10/credit; Essential $0.095; Business $0.09.[4] Plans: $99 / 1,100 cr, $475 / 6,000 cr. Enterprise “down to $0.01/min.”[3] Full mode = 2 credits/min (30 s/credit) and uses their Flash TTS.[4][5] Session caps 20–60 min on self-serve.[3]
HeyGen Studio (tape) Pre-rendered avatar video Different credit system. Avatar III vs IV/V are dollars-per-min on the API sheet (e.g. Twin IV ~$4.83/min API).[17] Do not mix with LiveAvatar slots.[4]
Tavus CVI Bundled live pipeline (LLM, TTS, WebRTC, face).[1][2] Free 25 min. Starter $59 / 100 min, overage $0.37/min, 3 concurrent, 3 replica trains/mo. Growth $397 / 1,250 min, overage $0.32/min, 10 concurrent, 7 trains/mo.[2] 30-second minimum, 6-second rounding. Extra replica train $65 / $40. Generated video $1 / $0.90 per min — not CVI.[2]
Anam Live avatar API Page lists Starter $0.16/min, Professional $0.11/min.[13] Third-party recaps also show $12/mo Starter with 50 included min.[13]
Simli Live STV minutes Free $10 on signup + 50 min/mo top-up; paid volume discounts.[14] Confirm current STV $/min on dashboard — marketing is “pennies vs 5–20¢.”
ElevenLabs Agents Voice (+ LLM extra) See companion page. Still on the bill in Lite. Don’t also pay Full-mode vendor TTS.
OSS GPU Your box CapEx / rent. LiveTalking: wav2lip256 ~60 fps on 3060; MuseTalk ~42 fps 3080 Ti, 72 fps 4090.[10] Wins at high concurrent minutes or residency. Loses at 20 interviews/week.

Worked sketch — 500 live interview-minutes/month, you keep ElevenLabs:

6. Roll-your-own — what people actually run

Live interviewer (OSS)

ElevenLabs stream
    → LiveTalking  (wav2lip for cheap GPU, MuseTalk for nicer mouth)
    or OpenAvatarChat (duplex interrupt already wired)
    → WHIP/WHEP or LiveKit custom track

Identity: drop in a 10-second silent looping body video + face, or a generated still expanded with LivePortrait into a loop, then lip-sync with MuseTalk. That is the synthetic-human factory: text-to-image (no one real) → idle video → live lips.[9][11][10]

Preplanned / scheduled (OSS)

Script → ElevenLabs TTS wav
    → SadTalker (one photo, good enough invite/recap)
    or MuseTalk (need a talking-head video identity)
    or LivePortrait (photo + a driver performance, then mux 11L audio)

SadTalker is still the thing people Colab. It is not conversational. Queue it from cron when a candidate books.[8]

Generated-identity pipeline (the “not a real person” stack)

  1. Generate a face with FLUX/SD + consistency (IP-Adapter / same seed sheet). Not a staff photo.
  2. Optional: LivePortrait + a royalty-free talking driver → idle motion plate.[9]
  3. Live: MuseTalk/LiveTalking on that plate, audio = ElevenLabs.[11][10]
  4. Tape: SadTalker on the still + the same ElevenLabs wav.[8]
  5. Same character sheet for eng-screener vs closer so the brand holds.

Vendors that already do step 3 from a still: HeyGen image LiveAvatar, Tavus image-to-replica, Simli face upload.[5][15] You pay minutes instead of a GPU.

7. Complexity — live vs tape vs identity

LowYou will eat glass
Live LITE face on existing 11L agentHeyGen LITE or Simli pluginPCM rate mismatch, interrupt freeze-frame, TURN for candidates on corp Wi-Fi
Tavus + 11LEcho/persona wiring, dual rooms, 30s minimum on abandoned joins, replica training slots[2][7]
Same face on live and tape at a vendorTavus Replica used twiceHeyGen: two platforms, two trainings, two credit pools[4]
OSS liveOne GPU demoConcurrency = GPU; idle = CPU encode; session thread leaks (LiveTalking has patched this class of bug)[10]
Synthetic identityOne pretty stillTemporal identity drift, teeth/tongue, side views, “is this a real employee?” from candidates
Likeness cloneVendor wizardConsent, union/publicity, deepfake statutes, HR disclosure, takedown

8. Suggested architecture for your interviews

  1. Voice: keep ElevenLabs (Flash for live, higher quality for tape).
  2. Live room: LiveKit or ElevenLabs Agents transport.
  3. Live face v1: HeyGen LiveAvatar Lite with a stock or image-generated avatar — not a manager clone. Or Simli if you want less HeyGen lock-in and already have LiveKit.
  4. Tape v1: ElevenLabs wav → SadTalker (OSS) or HeyGen Studio if you accept a second identity SKU.
  5. Tavus: shortlist if you want perception (does the candidate look confused) and Meet/Zoom joining more than you want to keep 11L as the mouth. Budget CVI at ~$0.32–$0.59 effective / min depending on volume.[2]
  6. OSS later: only when concurrent rooms or data residency pay for a 4090/H100. Then LiveTalking+MuseTalk becomes the live+tape renderer with one generated identity.

Vendor latency and “pennies a minute” claims are theirs. Spatius’s Tavus table was last verified Sep 4, 2026 against tavus.io/pricing.[2] HeyGen credit math is from LiveAvatar FAQ + liveavatar.com.[3][4]

Sources

  1. tavus.io/pricing
  2. Spatius — Tavus pricing 2026
  3. liveavatar.com
  4. HeyGen — LiveAvatar FAQ
  5. HeyGen — LiveAvatar by HeyGen
  6. ElevenLabs — LiveAvatar integration
  7. Pipecat — Tavus Video Avatar
  8. github.com/OpenTalker/SadTalker
  9. github.com/KlingAIResearch/LivePortrait
  10. github.com/lipku/livetalking
  11. github.com/TMElyralab/MuseTalk
  12. github.com/HumanAIGC-Engineering/OpenAvatarChat
  13. anam.ai/pricing
  14. simli.com
  15. LiveKit — Simli plugin
  16. beyondpresence.ai
  17. HeyGen API pricing explained
  18. tavus.io/cvi
  19. Open-LLM-VTuber
  20. Tough Tongue — avatar solutions 2026