TL;DR
The advertised per-minute rate is the orchestration fee, not the call cost. Vapi charges $0.05 per minute for its own hosting layer and passes speech-to-text, language model, text-to-speech and telephony through at cost. Retell publishes $0.07 to $0.31 per minute. Bland bundles the whole stack into one rate. Best for: understanding why three quotes for the same workload look nothing alike.
Independent worked modelling of a four-minute outbound call puts Vapi with bring-your-own-keys at $0.05 rising to $0.27 per minute all-in, Retell bundled at $0.10 rising to $0.16, and Bland flat-rate at $0.09 rising to $0.13. The cheapest headline is the most expensive fully loaded. Best for: budgeting on real numbers.
Vapi's differentiator is at-cost pass-through with the option to drop provider costs to zero by bringing your own keys, plus no monthly base fee on its self-serve tier and $10 per line per month for concurrency. Best for: teams that want to audit and control every layer.
Bland should not be described as FedRAMP Class C certified. Its published assessment page explicitly states that it is not Class C certified. For a regulated deployment, evaluate the current assurance evidence, deployment scope and required agreements rather than a blanket certification list.
Vapi documents separate HIPAA and Zero Data Retention modes, with configuration constraints described in its HIPAA documentation. Confirm the current commercial terms and the supported retention arrangement for the intended workload.
Contents
- The four cost layers behind every voice minute
- 1. Vapi
- 2. Retell AI
- 3. Bland
- Call control and integrations
- Latency claims and who is making them
- Side by side summary
- Additional tools to consider
- Which one fits which sales motion
- FAQ: AI Voice Platforms
The four cost layers behind every voice minute
Every quote in this category decomposes into the same four layers, and vendors differ mainly in how many they bundle.
Telephony is the carrier minute. Somebody has to originate or terminate the call on the public network, and that cost exists regardless of which orchestration platform sits on top. Vapi's documentation states it provides up to 5 free phone numbers per account for US national use with import available for international or custom cases, and that the telephony minutes themselves are a separate cost with no per-minute price published on its pricing page.
Speech-to-text transcribes the caller in real time. Vapi passes this through at cost with a named default provider and others available.
The language model generates the response. This is usually the largest variable cost on a long call and the one most sensitive to prompt design.
Text-to-speech synthesises the reply. Provider choice here swings both cost and perceived quality more than any other layer.
On top of those sits the orchestration fee, which is what the platform charges for holding the conversation together: turn-taking, interruption handling, tool calls, call control and state.
A voice-agent budget should include orchestration, speech and model services, telephony, numbers and peak concurrency. Reconcile the configuration with provider bills before forecasting at scale.
1. Vapi
Layer: Voice agent orchestration with at-cost provider pass-through
Best for: Teams that want to see, audit and control every component of the stack
Vapi's commercial design is unusually transparent for an orchestration platform, and it is the reason technical teams choose it. It marks up only its own hosting layer at $0.05 per minute and passes speech-to-text, language model, text-to-speech and telephony through at cost, dropping them to zero if you bring your own provider keys. That turns Vapi into a thin, predictable margin on top of provider spend the buyer can independently audit.
The plan structure has two named tiers. Build is usage-based with no monthly base fee at all, includes 60 or more call minutes, passes model costs to the customer, provides 10 call concurrency at $10 per line per month, offers email and Discord community support, and includes custom voices and model access. Scale is an annual contract with a fixed platform fee and committed volume, volume-based per-minute pricing, SOC 2, HIPAA, PCI, SSO and RBAC, data residency, priority model and provider access, a support SLA and a dedicated account team.
The hosting arithmetic is simple enough to do in your head. At $0.05 per minute, 300 call minutes is $15 per month, 4,000 minutes is $200, 10,000 minutes is $500, 50,000 minutes is $2,500 and 150,000 minutes is $7,500, all before models, voices and telephony.
Pricing: include orchestration, model and voice providers, telephony, concurrency and any enterprise or retention requirements. Price the chosen configuration; a base per-minute rate is not the complete cost.
Verdict: The right choice when you want control and auditability and you have the engineering capacity to exercise them. The wrong choice if you need a fixed number for a budget line before you have built anything.
2. Retell AI
Layer: Voice agent platform with bundled per-minute pricing
Best for: Teams that want one number that covers most of the stack
Retell sits between Vapi's at-cost transparency and Bland's fully bundled flat rate, and its published pricing reflects that middle position: a Voice AI pay-as-you-go tier at $0.07 to $0.31 per minute with 20 concurrent calls included and more available on demand, alongside custom enterprise pricing.
The range is the honest part. Retell's base voice engine rate is $0.07 per minute covering the conversation layer including speech-to-text processing and real-time handling, and the spread to $0.31 reflects model and voice choices layered on top. Independent worked modelling puts Retell bundled at $0.10 rising to $0.16 per minute on a four-minute outbound call, which is the tightest range of the three and the most forecastable.
Where Retell is strongest is integration depth into the systems a sales motion actually runs on. The platform documents integrations across a major CRM, an enterprise CRM, two support platforms, two contact-centre platforms, a document platform and custom API stacks. The Salesforce pattern is described precisely and it is the right pattern: OAuth-authenticated API calls the agent makes mid-conversation, with lead lookup, contact updates, opportunity stage changes and case creation happening during the call rather than as a delayed post-call sync. Retell's own framing of why that matters is worth repeating because most vendors gloss it: a connector that posts a transcript to an activity record an hour after the call ends is sufficient for compliance archiving and useless for personalisation.
Documented use cases relevant here are lead qualification for inbound demand and cold calling for outbound prospecting, with the platform positioning a hybrid model where the voice agent takes routine volume and hands off to a person when the conversation requires judgement.
Pricing: Voice AI pay-as-you-go at $0.07 to $0.31 per minute, 20 concurrent calls included with more on demand, enterprise custom. Base voice engine rate $0.07 per minute.
Where it falls short: The $0.07 to $0.31 spread is a factor of more than four, which makes early budgeting imprecise until you have fixed your model and voice choices. Retell is also the least differentiated of the three on positioning: it is neither the most transparent nor the most compliant nor the cheapest fully loaded, and it competes on being adequately good at all three. Concurrency beyond 20 is on demand rather than published.
Verdict: The safest default for a sales team that wants real-time CRM behaviour without owning the provider stack. Fix your model and voice configuration during the pilot, because that decision moves your rate more than the vendor choice does.
3. Bland
Layer: Fully bundled voice platform with enterprise compliance and delivered engineering
Best for: Regulated industries, enterprise procurement and teams without in-house voice engineers
Bland's commercial pitch is the inverse of Vapi's. One per-minute rate covers the language model, speech-to-text, text-to-speech and telephony, with no per-token charges, no per-feature surcharges and no separate vendor invoices. Pricing scales with usage, and enterprise plans are contracted on volume, dedicated infrastructure and compliance requirements. Independent modelling puts the flat rate at $0.09 rising to $0.13 per minute on a four-minute call, the lowest fully loaded figure of the three.
The billing documentation is granular where it matters for outbound specifically. Outbound calls carry a minimum charge of $0.015 per call attempt when using Bland's telephony, failed calls carry the same $0.015 minimum, voicemail is billed as part of standard call time at the plan's per-minute rate, and transfers using Bland-provided numbers are billed at a reduced rate and are free for bring-your-own-telephony customers. Warm transfer billing charges the proxy agent for talk time while active at the plan's per-minute rate. For a high-volume dialing motion, the per-attempt minimum on failed calls is the line item to model, because connect rates on cold mobile data run well below half.
The advanced feature list is gated to higher tiers and reads as an enterprise checklist: warm transfers, live transfers, guardrails described as protected calls, alarm and monitoring, knowledge base gaps, citations, outcomes, custom dialing and custom code extraction. Compliance and security gating covers business associate agreements, SSO, JWT signatures and data residency. Support gating covers a shared Slack channel with the Bland team and forward-deployed engineers.
The delivery model is the underrated differentiator. Bland states most production agents go live in two to six weeks depending on conversation complexity, and that its forward-deployed engineer team builds the customer's first agent end to end, owning the pathway, the integrations and the test loop with the customer's operations team. For a sales organisation with no voice engineering capability, that is the difference between a deployed agent and a stalled project.
Getting started is free: the Start plan gives two credits plus an inbound number valued at $15 per month with no card required.
Pricing: Bundled per-minute rate covering model, speech-to-text, text-to-speech and telephony, independently modelled at $0.09 to $0.13 all-in. Outbound minimum $0.015 per call attempt on Bland telephony; failed calls same. Voicemail billed as call time. Transfers reduced rate on Bland numbers, free on bring-your-own-telephony. Start plan free with two credits and an inbound number valued at $15 per month. Enterprise contracted on volume, infrastructure and compliance.
Where it falls short: Bundling removes the ability to swap a cheaper provider into a layer, which is exactly what Vapi's model exists to enable. Published tier prices are thinner than the feature matrix, so the specific per-minute rate at your volume is a sales conversation. And Bland's own analysis of the category is the most useful caution available on its own pricing model: the published list price for voice AI platforms tells enterprise buyers roughly 30% of their real bill, with the remainder in per-minute add-ons, concurrency overage, compliance surcharges and implementation.
Verdict: The right answer when compliance paperwork, deployment support or an enterprise procurement process is the binding constraint. Also, on independent modelling, the cheapest fully loaded of the three, which is the opposite of what the headline rates suggest.
Call control and integrations
Three call-control capabilities separate a demo from a deployed sales agent, and they are worth checking explicitly rather than assuming.
Transfer. Warm transfer means the agent stays on while a human joins; live transfer means a handoff. Bland documents both as advanced features with distinct billing, including proxy agent talk time during warm transfers. This is the capability that makes a voice agent viable for qualification, because the whole value is getting a qualified prospect to a rep while they are still on the line.
Mid-call tool use. Whether the agent can read and write to your systems while talking rather than after. Retell's Salesforce implementation makes lead lookup, contact updates, opportunity stage changes and case creation happen during the call through OAuth-authenticated API calls, and its documentation is right that a post-call transcript sync serves compliance rather than personalisation.
Latency claims and who is making them
Every vendor in this category publishes a latency figure and almost none publish the measurement conditions, so attribute carefully.
One carrier's technical framing describes streaming raw audio packets directly from the cellular network to the application layer to keep the round-trip loop under 300 milliseconds so conversation flows naturally, and that is a vendor claim about its own relay product. Voice agent platform reviews describe roughly 600 millisecond speech-to-text-to-speech round trips as a general category figure. An independent orchestration benchmark project publishes comparative scores across Vapi, Retell, ElevenLabs and custom configurations, with one published snapshot showing Vapi at 92, Retell at 87, ElevenLabs at 81 and a custom build at 74 across 40 scenarios and four providers. That benchmark is published by a testing vendor with a commercial interest in the category, and the scores are composite rather than pure latency.
The practical guidance: latency is dominated by your model and voice provider choices rather than by the orchestration platform, which is precisely why Vapi's at-cost pass-through model lets you tune it and Bland's bundle does not. If sub-second turn-taking is a requirement, test it with your own prompt and your own provider configuration during the pilot, and do not accept a benchmark run on somebody else's configuration as evidence about yours.
Side by side summary
| Tool | Layer | Best for | Entry price |
|---|---|---|---|
| Vapi Build | Orchestration with at-cost provider pass-through | Engineering teams wanting full control | $0.05 per minute orchestration, no base fee |
| Vapi Scale | Annual enterprise contract with committed volume | Enterprise with data residency needs | Fixed platform fee, quote only |
| Retell AI | Bundled voice platform with deep CRM integration | Sales teams wanting real-time CRM behaviour | $0.07 to $0.31 per minute |
| Bland | Fully bundled platform with delivered engineering | Regulated industries and enterprise procurement | Free Start plan, bundled per-minute thereafter |
Additional tools to consider
Four alternatives are worth including in a pilot, each solving a constraint the three principals do not.
ElevenLabs Agents publishes a free tier with 15 minutes of calls and 4 concurrent calls, Starter at $6 per month with 75 minutes and 6 concurrent calls, Creator at $11 to $22 per month with 275 minutes and 10 concurrent calls, and Pro at $99 per month. Agent rates were reduced in 2026 to $0.08 per minute on Starter, voice-only calls are charged on duration with a 95% discount for silence periods longer than 10 seconds, and language model costs pass through separately.
Twilio Conversation Relay is the build-your-own path, combining fast speech-to-text and text-to-speech with the model of your choice orchestrated through a WebSocket API, with published TwiML reference documentation, configurable text-to-speech providers and automatic language detection, priced at $0.07 per minute on top of standard voice rates.
CloudTalk and similar contact-centre platforms embed voice agents inside an existing dialer rather than alongside one, which suits teams whose reps and agents need to work the same queues and numbers.
Which one fits which sales motion
High-volume outbound cold calling. Model the per-attempt floor before the per-minute rate. Bland's $0.015 minimum per outbound attempt including failed calls is the honest structure, and at a 5% connect rate you are paying twenty attempt minimums per conversation. Bland's flat bundled rate at an independently modelled $0.09 to $0.13 per minute is the strongest total-cost position here, and its answering machine and voicemail handling is documented rather than assumed.
Inbound qualification and routing. Retell, because mid-call CRM reads and writes are the entire value and its documented Salesforce pattern does them during the conversation rather than after.
A custom agent inside an existing stack. Vapi, because at-cost pass-through with bring-your-own-keys means the platform takes $0.05 per minute and nothing else, and every other layer is a vendor relationship you already have or can negotiate independently.
Anything in healthcare, financial services or government. Bland, on certifications, self-hosted and on-premises options and forward-deployed engineering. If you must stay on another platform, price the compliance add-ons explicitly: $2,000 per month for HIPAA and $1,000 for Zero Data Retention on Vapi, and note those two modes are mutually exclusive.
FAQ: AI Voice Platforms
Why do the advertised per-minute rates differ so much from real bills?
The advertised rate may cover only one layer. Add telephony, models, voice processing, numbers, concurrency and relevant platform commitments for the chosen configuration.
Does the platform include phone numbers and telephony minutes?
Partly, and inconsistently. Vapi provides up to 5 free phone numbers per account for US national use with import available for international cases, while telephony minutes are a separate cost with no per-minute price on its pricing page. Bland includes telephony in its bundled rate and offers a reduced transfer rate on its own numbers with transfers free for bring-your-own-telephony customers. Bland's free Start plan includes an inbound number valued at $15 per month. As a floor reference, a major carrier publishes US local numbers at $1.15 per month and outbound US and Canada voice at $0.014 per minute.
How much does HIPAA compliance cost on these platforms?
Confirm the supported deployment, data handling and agreements with each provider. A historical surcharge or general marketing label is not a current quote for your intended workload.
Can these agents transfer a live call to a human rep?
Yes on all three, with different documentation depth. Bland documents warm transfers and live transfers as distinct advanced features with distinct billing, charging the proxy agent for talk time while active during warm transfers at the plan's per-minute rate, and billing transfers at a reduced rate on Bland numbers and free for bring-your-own-telephony customers. Retell positions a hybrid model where the agent handles routine volume and hands off when judgement is needed. Test transfer behaviour specifically during a pilot, because it is the capability that determines whether a qualification agent produces meetings or transcripts.
What happens on voicemail and failed calls?
Specify the expected behaviour for voicemail, no answer, timeouts and incomplete transfers. Confirm billing for each case and preserve failed outcomes in reporting rather than counting every attempted call as completed.
Should we build on a telephony provider directly instead?
Only with engineering capacity and a reason. A major carrier's relay product combines fast speech-to-text and text-to-speech with the model of your choice orchestrated through a WebSocket API at $0.07 per minute on top of standard voice rates, with published TwiML reference documentation and configurable voice providers. That is competitive with the orchestration platforms and gives you complete control. What you give up is call state management, testing tooling, evaluation frameworks and the deployment support that, on Bland, includes forward-deployed engineers building the first agent end to end in a stated two to six weeks.
Call handling also needs a review process. Our guide to conversation intelligence tools covers the adjacent options.
Work with Forma Nôrden
We build signal based outbound systems for B2B companies selling into the enterprise and upper mid market. Voice agents are priced in four layers and quoted in one, which is why so many pilots come in under budget and so many deployments do not. Explore how we work.
For enquiries about this article: partnerships@formanorden.com





