Hume AI develops expressive voice technology and tools for evaluating how people experience it. Its current homepage emphasizes datasets, expression measurement, Kairos simulations and human feedback, while its developer documentation and pricing still offer Octave text-to-speech and EVI speech-to-speech. These are related capabilities with different buying questions. A voice subscription should not be assumed to include the newer evaluation services.
- 01Two connected jobs. Generation produces speech and conversations; evaluation examines how those experiences work for listeners and participants.
- 02Scientific boundary. Expression measurements describe signals in behaviour, not a direct reading of a person’s internal emotional state.
- 03Commercial boundary. Published subscriptions cover TTS, EVI and Voice features. Evaluation and data offerings direct readers to the research team.
01 / ProductVoice generation alongside a broader evaluation and data offer
Hume's developer introduction describes Octave for expressive text-to-speech and EVI for live voice interaction. At consultation it labels Octave 2 as a preview and says EVI 4-mini is live. The products support different flows: prepared text becoming audio, versus an ongoing spoken exchange in which timing, interruptions and responses all matter.
The current homepage foregrounds four complementary offerings: datasets, Kairos, Expression Measurement and Human Feedback. We treat that as the present public positioning while retaining the documented generation offer. A different homepage emphasis is not evidence that Octave or EVI has shut down, and a live product landing page does not by itself prove that every new API is self-service.
The expression-measurement page distinguishes offline Tagger analysis from real-time Prosody measurement. The practical role is to characterize how speech is expressed and to support evaluation or response design. A voice can sound hurried or warm without that observation proving the speaker's private feelings, intention or reliability.
02 / AudienceVoice teams that need to hear and measure the user experience
Hume is relevant to a team building a spoken assistant, evaluating narration or comparing voice-model releases. Text correctness alone misses whether the delivery is understandable, whether a pause feels awkward or whether an interruption is handled sensibly. A narrowly defined voice task gives a team something concrete to evaluate across those dimensions.
There are different buyers inside that description. A content team may only need expressive narration. An application team may need an ongoing speech-to-speech interface. A model team may already have its generation stack and want external human judgments or training data. Treating all three as one subscription requirement would make both the trial and the budget difficult to interpret.
The ElevenLabs blueprint provides a useful comparison for voice and audio creation. The Retell blueprint is adjacent when the job is operating a voice agent. Hume's evaluation products can also be considered alongside a generation stack, rather than only as a replacement for it. Decide whether you need a different voice, a different agent platform or better evidence about the experience you already have.
03 / WorkflowA proposed evaluation of a spoken appointment-information assistant
Imagine a team with an assistant that explains appointment availability without booking anything. It wants to compare two voice configurations before a wider release. The following is a proposed evaluation design, not a Hume deployment or study conducted by Sequenced. Use fictional schedules and consenting adult participants so the first study does not require customer recordings.
Start with a few explicit situations: a caller asks for the earliest appointment, changes their mind, interrupts the explanation or gives an unclear date. Define the correct information independently of the voice system. Separate factual success from the listener's impression of naturalness; a reassuring response that gives the wrong time must still fail the task.
The Kairos overview describes agent-to-agent simulations and human-to-agent conversations, plus public or private benchmark suites. For this pilot, use simulated conversations to cover repeated variations and reserve human participation for questions where actual listening matters. Agree the accessible integration and study scope with Hume rather than assuming a public landing page supplies a ready-to-call endpoint.
Keep the scenario, transcript and expected answer constant while changing one voice configuration. This makes a preference comparison easier to explain. If the language model, prompt, voice and interruption logic all change together, the result may show a difference without identifying its cause. Record the version of every component so a later regression can be investigated.
The Human Feedback API page describes single-sample ratings, comparisons and live conversations, with questions and participant criteria supplied as part of a study. Ask distinct questions about intelligibility, pacing and appropriate delivery. Avoid one broad satisfaction score that conceals whether participants preferred a pleasant voice despite a failed task.
Expression measurements can add another view of the generated audio, but should not replace listening or task evaluation. Hume's use-case guidance explicitly warns against a one-to-one mapping between expression and emotional experience. In this proposed study, use measurements to locate clips for inspection or compare patterns, not to declare how a participant truly feels.
Inspect disagreement as well as averages. If some listeners find a voice patient and others find it patronizing, replay the relevant examples and review the wording and context. A useful evaluation may reveal that the same configuration fits a short factual response and performs poorly during an interruption. Preserve those distinctions when deciding what to change.
Finish with a release decision tied to the original task. Require correct information, acceptable interruption behaviour and understandable delivery, then use human preferences to guide refinement. Retain a small recurring evaluation set for future changes. The aim is a repeatable way to detect an experience getting worse, not a claim that one score captures empathy or that simulated callers represent every real listener.
04 / PricingPublished voice subscriptions and separately scoped research services
The Hume pricing page, consulted 24 September 2026, displays monthly voice plans including Free at $0, Starter at $3, Creator at $14, Pro at $70, Scale at $200 and Business at $500, with Enterprise quoted. Creator also shows a first-month $7 promotion. The table includes TTS character allowances and EVI minutes, but model toggles and plan-specific limits require care.
For example, the displayed Pro table lists one million TTS characters and 1,200 EVI minutes, with EVI 3 overage at $0.06 per minute. Do not silently apply an EVI 3 row to every other model version or interpret approximate audio minutes as an exact conversion from characters. Confirm the model and quota shown in the account.
The billing guide says one subscription covers TTS, EVI and Voice features, with included usage and overages. Managed external LLMs can add charges; using your own model key leaves its bill with that provider. Kairos, datasets and human-evaluation pages instead offer contact-based access. Their study costs, data licences and integration scope were not published as a numerical tariff in the pages reviewed.
| Route | Commercial basis | Important boundary |
|---|---|---|
| Voice Starter | $3/month displayed | Included usage and model-specific limits apply |
| Voice Creator | $14/month; first month $7 displayed | Promotion is not the continuing monthly rate |
| Voice Pro | $70/month displayed | Published allowances; verify model toggle and overages |
| Enterprise voice | Custom quote | Capacity and controls depend on agreement |
| Kairos / Human Feedback / datasets | Contact research; no numerical tariff verified | Separate evaluation scope or data licence |
Selected Hume voice pricing and billing rules, consulted 24 September 2026. Displayed dollar amounts; newer evaluation services have separate contact-based access.
05 / DistinctionsHuman judgments can complement automated voice evaluation
The useful distinction is the attempt to connect generation, expression research and listener evaluation. A team can ask whether a model has improved on the dimensions its users experience, rather than relying only on a transcript metric. That is valuable when voice quality is part of the product's purpose and when the study is designed to separate meaningful changes from preference noise.
Hume's dataset offering includes conversational, multilingual and expressive speech, with samples and custom collections accessed through its research team. A model developer should match the recording conditions and intended use to the training problem. Buying more audio is not automatically equivalent to obtaining better coverage of the speakers, scenarios or expressions the model lacks.
06 / QuestionsCheck study access, licences and what each measurement means
The Human Feedback page's documentation link resolved to the general generation introduction during this review. That leaves the public implementation contract for the newer service unresolved here. Ask for the actual study API, participant-selection options, export format and result-delivery process before scheduling an evaluation that depends on them.
For voice generation, confirm commercial-use eligibility on the selected plan, the exact model version and any voice-cloning consent requirements. This source review did not make API calls, test expression measurements or commission human ratings. It establishes documented product scope and current published pricing, not that a model is empathetic, a measurement is valid for every setting or the advertised turnaround will hold for a specific study.
07 / DecisionChoose the generation or evaluation job before choosing a plan
Hume AI is most useful when a team can name the voice experience it wants to create or the question it needs an evaluation to answer. Start with a short, representative task and separate factual success from delivery quality. For the newer research services, obtain a concrete scope before relying on access. For generation, evaluate the exact model and plan together so a promising sample can become an operable product.
Evaluate an exact voice model
Compare delivery on your script and confirm commercial-use eligibility for the chosen plan.
Test complete conversations
Use repeated scenarios and human listening to examine correctness, pacing and interruptions.
Scope a focused study
Define participants, questions, integration and result formats before commissioning evaluation or data.
A business worth understanding.
Suggest your business or one you find interesting. Tell us what you want to understand about its product, positioning, design or workflows.
Suggestions are free. Selection and publication stay with the desk.
- developer introductionConsulted
- current homepageConsulted
- expression-measurement pageConsulted
- Kairos overviewConsulted
- Human Feedback API pageConsulted
- use-case guidanceConsulted
- Hume pricing pageConsulted
- billing guideConsulted
- dataset offeringConsulted
