AI News

HeyGen Voice: Which Settings Actually Sound Like You?

If your goal is to make a HeyGen voice sound like its original speaker, the company’s documented starting point is similarity: "max" when creating an instant clone and expressiveness_boost: 0.5 when generating speech. HeyGen Studio uses those defaults. The API defaults are different, and they are the configuration behind the company’s Controlled Voice leaderboard result.

That distinction makes the October 9 HeyGen Voice launch more useful to examine than a simple “number one” headline. The release, timestamped 13:20 Eastern, or 17:20 UTC, introduces HeyGen’s in-house voice model. Developers now have an instant-cloning path alongside professional cloning, which already existed.

Evidence note, October 10, 2026: Kingy checked the announcement, current API documentation and the live Controlled Voice leaderboard. We have not completed a consenting-speaker listening comparison. No naturalness, speaker-identity, pronunciation or latency scores in this article are Kingy measurements.

The two configurations optimize different priorities

HeyGen’s speech documentation explicitly recommends the Studio settings for the closest match to the original speaker. It identifies the API defaults as the configuration that took first place on Artificial Analysis’ Controlled Voice leaderboard.

Documented instant-voice configurations, checked October 10
Configuration Clone creation: similarity Speech: expressiveness_boost Evidence
Studio defaults / identity recommendation max 0.5 HeyGen’s recommended starting point for speaker likeness.
API defaults extra_high 1.0 Configuration HeyGen identifies behind its Controlled Voice result.

Naturalness and identity are separate judgments. A voice can sound relaxed, expressive and convincing while departing from a person’s habitual timing or emphasis. A closer imitation can also reproduce the restrained delivery of a source recording, which a listener might find less engaging. Decide which outcome matters before treating a more animated sample as an improvement.

There is also an experimental detail hidden in the settings. Similarity is chosen when the clone is created; expressiveness is supplied when it speaks. Comparing both documented configurations therefore requires two clones made from the same recording, followed by speech requests with the appropriate settings. Changing expressiveness on one existing clone does not reproduce the full comparison.

What the leaderboard actually measures

The Controlled Voice board uses the same eight cloned voices, four US and four UK, for English speech. At our October 10 check, HeyGen Voice occupied the first row with Elo 1201 ±16 from 1,531 samples. The board displayed a rank range of 1–2; the second model’s interval overlapped HeyGen’s.

Those details support a narrower claim than universal superiority. They establish a strong preference result on that board, at that time, with its voice set and evaluation conditions. They do not establish the best likeness for an arbitrary speaker, best pronunciation of a company glossary, or best performance across every language.

The separate Provider Voice leaderboard compares models using their providers’ own voices. Its rankings answer a different question. A model’s position there should not be substituted for its Controlled Voice position, or vice versa.

For a founder recording a product tutorial, identity may outweigh broad listener preference. For an explainer with no requirement to imitate a particular person, a natural and intelligible performance may matter more. A single overall score conceals that editorial choice.

Instant cloning and professional cloning are different products

The HeyGen Voice model overview labels instant cloning as new. It describes a short-recording route and a professional route trained for an individual speaker. Both ultimately produce speech through the model-backed audio endpoints. Professional voice cloning should not be presented as a capability first invented by this launch.

For instant cloning, only the first three minutes of one recording are used. The API offers similarity choices from medium through max, plus background-noise removal. The voice is usually ready within seconds, according to HeyGen; wait for ACTIVE status before synthesis. To change an instant clone, create another voice rather than retraining it.

The professional API guide requires one to ten recordings totaling at least twenty minutes and a purchased clone slot. It supports retraining. The launch release describes the professional add-on as $99 a month with thirty minutes to three hours of speech. These are different descriptions of the product workflow, so do not use the launch’s recording guidance as the API’s minimum validation rule.

A fair comparison should hold the cloning tier constant. Giving a competitor hours of professional training audio while giving HeyGen a brief instant-clone sample would test two products and two data budgets at once. That can be a useful purchasing comparison if stated openly; it is a poor basis for claiming one model has better identity fidelity.

Use the endpoint for the voice you actually created

Model-backed clones speak through POST /v3/models/audio/tts or its streaming counterpart. HeyGen’s model overview distinguishes these IDs from the separate voice catalog and third-party speech surface. An ID appearing somewhere in a voice library does not prove it belongs to the endpoint a developer is calling.

The completed-speech endpoint returns a WAV; streaming delivers ordered audio parts. For an integration, distinguish time to first playable audio from time to completed audio. A fast first fragment may be useful for conversation, while an editor downloading narration cares about the complete file.

Instant and professional voices also accept different controls. Expressiveness applies to instant voices. Professional voices support seed, speed and pitch controls; a seed is best-effort determinism rather than a guarantee of identical audio. Supplying the wrong control for the voice kind can produce an invalid-parameter error. Copying a professional example into an instant-clone request can therefore fail before voice quality is even evaluated.

Set the language deliberately for a controlled test. HeyGen’s instant guide says regional language tags normalize to the base language and do not select an accent. A regional tag is therefore not an adequate control for checking whether a clone preserves the speaker’s Canadian, Scottish or other regional delivery.

Free creation does not settle the synthesis bill

The launch release describes free access on the platform and API. The current model overview is more specific: an instant voice is free to create during preview, while professional cloning requires a purchased slot. It lists professional speech at 0.6 API credits per generated minute.

That page does not establish a complete promotional dollar price for instant speech. The speech documentation includes insufficient-credit and usage-limit errors. Artificial Analysis displays $30 per million characters for HeyGen Voice, but a third-party leaderboard conversion is not an authenticated quote for your account or proof of a launch promotion.

Keep three questions separate: what it costs to create the voice, what it costs to generate speech, and what it costs to create or export the surrounding video. A zero-priced clone creation step does not answer the other two.

Before a paid trial, capture the account’s applicable rate, credit balance and usage record. Record failed requests and retries as well as completed audio. Promotional API pricing remains unverified in this article; we will not turn a broad free-access announcement into a zero-cost synthesis promise.

The listening test that can answer “does it sound like you?”

Our proposed comparison uses one adult speaker who explicitly consents to cloning, the named services and publication of the resulting samples. Use one clean reference recording for both HeyGen configurations and a competitor’s instant clone. The competitor must accept comparable source material; record its exact model and settings rather than using a brand name as a model identifier.

The speaker should also record the test scripts naturally as a reference. Keep those readings out of the clone-creation sample. This makes it possible to compare identity and delivery on new sentences, rather than rewarding a model for reproducing training material.

Use three identical scripts across every generated condition: a neutral explanation, a passage with contrast and emphasis, and a pronunciation passage containing names, acronyms, dates and numbers. Fix the scripts before hearing any output. For example: “Dr. Nguyen reviewed the API at 9:15. The invoice is $1,204.50, due on October 23.” Agree on the intended readings with the speaker before scoring.

Score separately; no Kingy scores have been assigned
Dimension Scoring question Record
Naturalness Does the delivery sound coherent and human, without distracting synthetic artifacts? Anchored 1–5 listener score, plus artifact notes.
Speaker identity How closely does it resemble the consenting speaker’s reference? Separate anchored 1–5 score from listeners familiar with the reference.
Pronunciation Were the agreed names, acronyms and numbers spoken correctly? Correct targets / total targets; retain each error.

Blind the file labels, randomize order, match playback loudness and use the same headphones. Generate three repetitions per script and condition, retaining every output. Reporting only the best take would overstate how reliably the settings work. Collect the speaker’s own judgment separately from the listener panel.

A one-speaker result is useful for that speaker and those scripts. It cannot resolve performance across accents, languages or recording environments. If one configuration wins naturalness and another wins identity, publish that split rather than averaging it into an invented universal champion.

Kingy’s existing HeyGen review now includes a dated voice update and a link to this guide. Start with HeyGen’s documented identity configuration, then judge new scripts against the real speaker. A leaderboard gives you a candidate to try; the reference recording gives you the comparison that matters.