Keep ElevenLabs as the voice. Pick a face that can do a live two-way interview and a scheduled, scripted video. Stock humans, consented clones, and synthetic faces that never existed. Tavus and HeyGen LiveAvatar in the same matrix as the OSS talking-head stack.
Companion to the orchestration report.
If the interviewer must look human, live, and you already pay ElevenLabs: HeyGen LiveAvatar in LITE / Avatar-Only mode. ElevenLabs owns audio; LiveAvatar owns the face. Documented by ElevenLabs. PCM 24 kHz. 1 credit = 1 minute in Lite.[6][4][5]
If you want one vendor that is a whole conversational-video product (perception, RAG, Zoom/Meet join) and you are OK not driving the mouth from ElevenLabs by default: Tavus CVI. To keep ElevenLabs, you use Echo / Pipecat’s persona_id=pipecat0 so Tavus eats your TTS audio instead of its own voice.[7][18] That is more glue. CVI overage is ~$0.32–$0.37/min on public Starter/Growth cards — bundled STT+LLM+TTS+face, not face-only.[2]
If you want cheaper live face + still ElevenLabs: Simli (LiveKit plugin) or Anam (~$0.11–$0.16/min on Anam’s page) or Beyond Presence speech-to-video.[13][14][15][16] These are live layers. Preplanned MP4s are a second pipeline.
If the face must be OSS and you own a GPU: one renderer for both modes — LiveTalking / OpenAvatarChat + MuseTalk for live WebRTC; SadTalker / LivePortrait (or the same MuseTalk, offline) for pre-rendered interview invites and recaps. Audio in both cases = ElevenLabs file or stream.[10][11][12][8][9]
Do not clone a real person without a recorded consent path. HeyGen requires a consent clip for custom video LiveAvatars; image avatars are the synthetic-identity loophole vendors themselves document.[4][5] OSS photo-to-talking-head is the same legal surface, just without a vendor saying no.
Candidate talks. Agent barges in. Face must track streaming TTS chunks, freeze on interrupt, idle while listening. Transport is WebRTC. Latency budget ~1 s to first frame.
Tavus CVI, HeyGen LiveAvatar, Anam, Simli, Beyond Presence, LiveTalking, OpenAvatarChat, MuseTalk realtime.
Script known ahead: invite, briefing, “you’re in 10 minutes,” score recap, rejection. Render an MP4 (or a sequence) from ElevenLabs audio + a still or short driver clip. Quality can be higher because nobody is waiting.
HeyGen Studio (separate from LiveAvatar — avatars are not cross-compatible).[4] Tavus generated-video minutes (separate from CVI).[2] SadTalker, LivePortrait, MuseTalk offline, ElevenLabs Creative talking-video.
Vendors: usually two SKUs. HeyGen Digital Twin ≠ LiveAvatar. Tavus Replica can feed CVI and generated video, but live and tape are different meters.[4][2]
OSS: one identity image/video, two inference modes. That is the real roll-your-own win.
This is the split you called out: bolting a face generated from a still / a random virtual human, vs a personal likeness.
| Class | How you get it | Interview use | Risk |
|---|---|---|---|
| Stock / library human | Vendor preset (HeyGen 100+ presets; Tavus stock replicas; Simli face IDs; Anam library).[3][20] | Default interviewer. Fast. No consent theater. | Looks like “an actor from a SaaS ad.” Fine for screening. |
| Personal likeness (clone) | 2-min talking + listening footage (HeyGen video LiveAvatar) or short video/image-to-replica (Tavus). Consent recording required on HeyGen custom video; Enterprise can use pre-recorded consent.[4][5] | Named hiring manager / founder twin. Recruiter who cannot be in 40 rooms. | Deepfake + employment law. Need written consent, retention policy, “this is AI” disclosure. |
| Synthetic / generated identity | FLUX / SD portrait that never existed → image-to-avatar (HeyGen image LiveAvatar, Tavus image-to-replica, SadTalker, LivePortrait). No real person.[5][8][9] | Brand-consistent interviewer that is not impersonating staff. Rotate looks per role (eng vs sales) without cloning anyone. | Uncanny valley; accidental resemblance; some jurisdictions still want disclosure. This is the “going away from personal likeness” move. |
| Stylized virtual | Ready Player Me / MetaHuman / Live2D. Rigged 3D or cartoon.[19] | Clearly-not-a-person. Gaming/education vibe. | Wrong register for most corporate interviews unless you lean into “this is a bot.” |
Practical interview policy: synthetic stock or generated face + ElevenLabs voice that is also not a real employee’s clone, unless you have signed talent. Likeness clones are a product later, not v1.
ElevenLabs is the mouth audio. The face is a consumer of that audio (or, in Full/CVI mode, a competing TTS you should not double-pay).
Keep: STT → LLM → ElevenLabs streaming TTS Add: TTS audio PCM 24 kHz → avatar renderer → WebRTC video track Official 11L path: HeyGen LiveAvatar LITE ElevenLabs Agents audio + LiveAvatar face bills: 11L minutes AND 1 LiveAvatar credit / min[6] Pipecat/LiveKit path: TavusVideoService persona_id=pipecat0 means “use the bot’s TTS, echo audio, don’t use Tavus voice”[7] Also works: Simli / Anam / Beyond Presence as a LiveKit participant they take the same TTS track[15][16] Offline tape: ElevenLabs → wav → SadTalker / LivePortrait / MuseTalk / HeyGen Studio
LiveAvatar Full mode uses ElevenLabs Flash v2.5 internally plus Deepgram/AssemblyAI — that is their ElevenLabs, not necessarily your agent voice, and it costs 2 credits/min.[4][5] For “we use ElevenLabs” as your voice, stay on Lite / Avatar Only.
| Option | Live 2-way | Preplanned MP4 | 11L as YOUR tts | Stock | Likeness | Generated still | Self-host | Complexity |
|---|---|---|---|---|---|---|---|---|
| HeyGen LiveAvatar | Yes. LITE or Full. First frame <300 ms claimed.[3] | No — use HeyGen Studio separately; avatars not cross-compatible.[4] | Yes — LITE, documented.[6] | 100+ presets[3] | 2-min video + consent, or photo[4] | Photo → image avatar (no voice clone)[5] | No | Low if LITE. Medium if you also want Studio twins. |
| Tavus CVI | Yes. Phoenix-4.5, Raven perception, Sparrow dialogue.[18] | Yes — generated-video minutes on the same account, different rate.[2] | Via Echo / Pipecat pipecat0, not the default.[7] |
Stock replica library[18] | Custom replica trainings (plan-capped)[2] | Image-to-replica exists as a product line | No | Medium. Dual-room with Daily. 30s min + 6s rounding.[2][7] |
| Anam | Yes. Latency-first (~180 ms claimed in roundups).[20][13] | Not the product | BYO LLM/TTS via LiveKit/Pipecat | Yes | Custom uploads | Custom uploads | No | Low–medium |
| Simli | Yes. LiveKit plugin.[15] | Not the product | Yes — audio from your agent | Face IDs[15] | Upload a face[15] | Upload a generated still | No | Lowest LiveKit glue |
| Beyond Presence | Yes. Speech-to-video API, LiveKit token.[16] | Not the product | Yes — audio-to-video | Library + custom | Custom | Custom | No (EU-hosted) | Low if you already have LiveKit |
| D-ID / Spatius-class | Agents / streaming SKUs | Classic photo-talk clips | BYO audio on some SKUs | Yes | Photo | Photo | No | Low for tape; live quality varies |
| LiveTalking OSS | Yes. WebRTC / WHEP / RTMP. wav2lip / MuseTalk.[10] | Same models, file out | Stream ElevenLabs into TTS slot | You supply the clip | You supply consented footage | Animate any still/video you generate | Yes Apache-2.0 | High. GPU, TURN, session leaks. |
| OpenAvatarChat OSS | Yes. Duplex interrupt, MuseTalk / FlashHead / lite-avatar.[12] | Possible via handlers | Swap TTS handler → ElevenLabs | Bundled demos | Your data | Your still | Yes Apache-2.0 | High. Full stack demo, not a plugin. |
| MuseTalk OSS | Realtime infer exists; used inside LiveTalking.[11][10] | Excellent offline | Audio file/stream in | n/a — renderer only | Identity = your video | Identity = your still+body video | Yes MIT code[11] | Medium as a library, high in prod |
| SadTalker OSS | No (batch). CVPR 2023, ~14k stars.[8] | Yes — the classic one-photo + wav → MP4 | ElevenLabs wav in | n/a | Any portrait | Best OSS tape path for generated faces | Yes | Low for a clip; not live |
| LivePortrait OSS | Video-driven (needs a driver clip), not audio-native.[9] | Yes — animate a still from a talking driver | Indirect (driver from a talking video) | n/a | Any portrait | Generated still + stock driver video | Yes ~19k stars[9] | Medium. Great motion, extra audio-lip step. |
| Open-LLM-VTuber / Live2D OSS | Yes, local[19] | Screen-record | Swap TTS | Anime models | No | Stylized only | Yes | Low technically, wrong look |
| MetaHuman / RPM + UE Pixel Streaming | Yes if you build visemes from 11L | Sequencer / Movie Render | Yes — viseme / audio2face | Parametric humans (not photoreal photo) | Scan optional | Randomize sliders = infinite virtual people | Engine yes, heavy | Highest. Game-studio complexity. |
Face-only vs bundled. Adding ElevenLabs on Lite/Echo is extra. Numbers move; verify before a PO.
| Vendor | What the minute includes | Public $ | Notes |
|---|---|---|---|
| HeyGen LiveAvatar Lite | Face + stream only. Your 11L/LLM separate.[4][6] | 1 credit / min. Starter overage $0.10/credit; Essential $0.095; Business $0.09.[4] Plans: $99 / 1,100 cr, $475 / 6,000 cr. Enterprise “down to $0.01/min.”[3] | Full mode = 2 credits/min (30 s/credit) and uses their Flash TTS.[4][5] Session caps 20–60 min on self-serve.[3] |
| HeyGen Studio (tape) | Pre-rendered avatar video | Different credit system. Avatar III vs IV/V are dollars-per-min on the API sheet (e.g. Twin IV ~$4.83/min API).[17] | Do not mix with LiveAvatar slots.[4] |
| Tavus CVI | Bundled live pipeline (LLM, TTS, WebRTC, face).[1][2] | Free 25 min. Starter $59 / 100 min, overage $0.37/min, 3 concurrent, 3 replica trains/mo. Growth $397 / 1,250 min, overage $0.32/min, 10 concurrent, 7 trains/mo.[2] | 30-second minimum, 6-second rounding. Extra replica train $65 / $40. Generated video $1 / $0.90 per min — not CVI.[2] |
| Anam | Live avatar API | Page lists Starter $0.16/min, Professional $0.11/min.[13] | Third-party recaps also show $12/mo Starter with 50 included min.[13] |
| Simli | Live STV minutes | Free $10 on signup + 50 min/mo top-up; paid volume discounts.[14] | Confirm current STV $/min on dashboard — marketing is “pennies vs 5–20¢.” |
| ElevenLabs Agents | Voice (+ LLM extra) | See companion page. Still on the bill in Lite. | Don’t also pay Full-mode vendor TTS. |
| OSS GPU | Your box | CapEx / rent. LiveTalking: wav2lip256 ~60 fps on 3060; MuseTalk ~42 fps 3080 Ti, 72 fps 4090.[10] | Wins at high concurrent minutes or residency. Loses at 20 interviews/week. |
Worked sketch — 500 live interview-minutes/month, you keep ElevenLabs:
ElevenLabs stream
→ LiveTalking (wav2lip for cheap GPU, MuseTalk for nicer mouth)
or OpenAvatarChat (duplex interrupt already wired)
→ WHIP/WHEP or LiveKit custom track
Identity: drop in a 10-second silent looping body video + face, or a generated still expanded with LivePortrait into a loop, then lip-sync with MuseTalk. That is the synthetic-human factory: text-to-image (no one real) → idle video → live lips.[9][11][10]
Script → ElevenLabs TTS wav
→ SadTalker (one photo, good enough invite/recap)
or MuseTalk (need a talking-head video identity)
or LivePortrait (photo + a driver performance, then mux 11L audio)
SadTalker is still the thing people Colab. It is not conversational. Queue it from cron when a candidate books.[8]
Vendors that already do step 3 from a still: HeyGen image LiveAvatar, Tavus image-to-replica, Simli face upload.[5][15] You pay minutes instead of a GPU.
| Low | You will eat glass | |
|---|---|---|
| Live LITE face on existing 11L agent | HeyGen LITE or Simli plugin | PCM rate mismatch, interrupt freeze-frame, TURN for candidates on corp Wi-Fi |
| Tavus + 11L | — | Echo/persona wiring, dual rooms, 30s minimum on abandoned joins, replica training slots[2][7] |
| Same face on live and tape at a vendor | Tavus Replica used twice | HeyGen: two platforms, two trainings, two credit pools[4] |
| OSS live | One GPU demo | Concurrency = GPU; idle = CPU encode; session thread leaks (LiveTalking has patched this class of bug)[10] |
| Synthetic identity | One pretty still | Temporal identity drift, teeth/tongue, side views, “is this a real employee?” from candidates |
| Likeness clone | Vendor wizard | Consent, union/publicity, deepfake statutes, HR disclosure, takedown |
Vendor latency and “pennies a minute” claims are theirs. Spatius’s Tavus table was last verified Sep 4, 2026 against tavus.io/pricing.[2] HeyGen credit math is from LiveAvatar FAQ + liveavatar.com.[3][4]