Best APIs for Transcribing and Analyzing Sales Calls

Yananai A. ChiwutaPublished ·15 min readUpdated
Best APIs for Transcribing and Analyzing Sales Calls

TL;DR

  • AssemblyAI is the strongest starting point for recorded sales calls needing speaker-labelled text and selectable analysis. Universal-3.5 Pro is $0.21/audio hour, plus $0.02/hour for diarisation.
  • Deepgram fits developers building both recorded-call processing and live speech applications. Nova-3 monolingual recorded transcription is $0.258/hour with diarisation included; live pricing and add-ons differ.
  • Gladia fits multilingual calls with language switching and included diarisation. Starter is $0.61/hour recorded or $0.75/hour live; distinct stereo channels can double billing.
  • Amazon Transcribe and Google Cloud Speech-to-Text make sense inside their existing clouds. AWS Call Analytics is a different purchase from basic transcription; Google's multi-channel billing and diarisation method/language gates matter.
  • OpenAI's current GPT Transcribe is $0.27/hour for file transcription. Its legacy speaker-labelled model is deprecated, so a new build needing speaker attribution should not depend on that model's future availability.
  • Buy accurate facts and usable ownership, not just fluent summaries. Review time usually costs far more than the speech API.

The output a sales team actually needs

Sales-call transcription has three separate jobs. Recognition turns audio into words. Speaker attribution associates those words with the buyer or representative. Analysis extracts the problem, objection, next step and commercial context. A transcript can be mostly correct while its only important number is wrong; a summary can read naturally while converting the representative's suggestion into the buyer's commitment.

For recorded calls, the useful output is a timed transcript, a reliable speaker or channel map, and extracted facts linked back to the relevant passage. A live agent-assist application additionally needs stable interim results, turn completion and a latency target. Streaming an already recorded file is not the same engineering problem as receiving a live phone call.

An API supplies components for an application. It does not automatically buy a meeting recorder, coaching interface, account permissions or CRM integration. Teams seeking that complete seller product should start with the conversation intelligence guide and Gong, Chorus and Fathom comparison.


Current API and price comparison

This comparison uses public pricing assessed on 1 October 2026. Rates are USD per processed audio hour, before tax, storage, application development and separately selected analysis. Models and billing paths are explicit; language coverage is not a promise that every feature works in every language.

API and selected route Recorded price/output Live price/boundary Best fit
AssemblyAI Universal-3.5 Pro $0.23/hour including $0.02 diarisation add-on; 18-language model Universal-3.6 Pro Realtime $0.45/hour plus $0.12/hour diarisation; $0.57/hour with that add-on Recorded conversations with selectable speaker and understanding features
Deepgram Nova-3 Monolingual $0.258/hour, multilingual $0.312/hour; recorded diarisation included Current monolingual promotion $0.288/hour plus $0.12 diarisation; $0.408/hour; promotion is temporary Recorded and live speech processing in one developer stack
Gladia Starter $0.61/hour, diarisation and 100+ languages included $0.75/hour; automatic language switching; channels billed separately Multilingual calls and a straightforward included-feature base
Amazon Transcribe Current US East pricing example uses $0.006/minute, $0.36/hour, with standard speaker/channel features Current US East streaming example uses $0.01/minute, $0.60/hour AWS-hosted recordings, permissions and operational infrastructure
Google Speech-to-Text V2 Standard $0.96/hour at $0.016/minute; dynamic batch $0.18/hour at $0.003/minute Standard recognition pricing; selected method/model controls features Google Cloud applications; delayed batch processing where suitable
OpenAI GPT Transcribe $0.27/hour at $0.0045/minute; text transcription, not a promised new speaker-label contract GPT Live Transcribe $1.02/hour at $0.017/minute Existing OpenAI applications that can supply speaker identity separately

Price sources: AssemblyAI, Deepgram, Gladia, AWS, Google, OpenAI.

AWS amounts use the current US East (N. Virginia) examples, with regional and feature pricing kept separate. Google's rates refer to V2, rather than the older V1 logged-data rate. OpenAI's row deliberately prices the current general transcription model without claiming it supplies the deprecated diarisation model's output.


The six options

1. AssemblyAI: recorded speech with modular understanding

AssemblyAI offers recorded transcription, speaker labels, timestamps and separate speech-understanding features. Universal-3.5 Pro's current base is $0.21/hour; diarisation adds $0.02/hour. Universal-2 is $0.15/hour and lists ninety-nine languages, while Pro lists eighteen with native code switching. Select the model for the actual language mix, rather than assign Universal-2's coverage to Pro.

The add-on structure lets a buyer price the intended output. Pro keyterms prompting adds $0.05/hour; low-effort summarisation is $0.02/hour, medium-effort $0.07/hour, and sentiment $0.02/hour. Pro plus diarisation, keyterms and medium-effort summarisation therefore totals $0.35/audio hour. Speaker identification, which tries to provide names or roles, is distinct from anonymous diarisation. Current model and add-on matrix.

Choose AssemblyAI for an English or supported-language recorded-call application with a clear feature recipe. It is less compelling when the team needs a language outside the chosen model or expects every analysis feature in the cheapest rate. A timestamped chapter summary can speed review, but the owner still needs the passage behind a proposed deadline or amount.

2. Deepgram: recorded and live speech paths with different costs

Deepgram exposes recorded and streaming transcription, model/language selection, formatting and diarisation. The recorded Nova-3 monolingual rate is $0.0043/minute, or $0.258/hour; multilingual is $0.0052/minute, or $0.312/hour. Recorded diarisation is included. The same assumption is wrong for streaming, where diarisation adds $0.002/minute. Current streaming discounts are labelled promotional.

Audio Intelligence covers functions such as summaries, sentiment, topics and intent. These are analysis outputs, not proof that the customer has approved a deal. Multichannel processing can return separate transcripts for separate source channels, making recorder-supplied speaker identity useful.

Choose Deepgram when developers need speech processing for both offline calls and live features, or an existing integration already works. A batch rate should not be used to cost the live assistant, and a sentiment label should not overwrite opportunity qualification. The current data guide makes request-level model-improvement opt-out explicit; opted-out content is not retained after the response, although metadata remains. That is a more concrete retention choice than assuming the base API stores a permanent call archive.

3. Gladia: multilingual calls with included core features

Gladia's Starter base includes speaker diarisation, word-level timestamps, automatic language detection/switching and 100+ languages at $0.61/hour recorded or $0.75/hour live. Growth's advertised “as low as $0.20/hour” recorded rate requires an upfront commitment, so it is not applied to this guide's pay-as-you-go workload. Plans.

Its summary API offers general, concise and bullet-point summaries. That is convenient for a call application, while the structured extraction of your own sales fields remains a separate design decision. Keep factual passages and inferred conclusions distinct even when one provider supplies both transcript and summary.

The channel documentation supports up to two recorded channels and eight live channels. Distinct recorded channels are billed as two audios; live processing is billed per channel. A fifteen-minute stereo call with one speaker on each channel can therefore cost twice the mono example.

Choose Gladia when multilingual handling and included core capabilities simplify development. Its higher unit rate can still be economical if it reduces corrections. Choose a different route when the required data-retention setting exceeds the selected plan, or when the actual multi-channel cost removes the apparent advantage.

4. Amazon Transcribe: AWS-native processing and a separate analytics purchase

Amazon Transcribe takes S3 recordings or live media and returns transcripts with timing and optional speaker partitioning or channel identification. It supports mono and dual-channel media; the current standard pricing includes up to two channels in the audio duration, rather than charging both separately. The current US East examples use $0.006/minute for batch and $0.01/minute for streaming, superseding the older $0.024/minute example often quoted elsewhere. Input/output documentation, current prices.

Call Analytics is the stronger route when the application needs richer call insights, custom categories and customer/agent analysis. Its current first US East tier is $0.03/minute, or $1.80/hour, with generative summarisation priced additionally. It requires two-channel audio with agent and customer separated; it is not the same product as standard mono transcription. New categories must exist before processing a call. Call Analytics requirements.

Choose standard Transcribe for an AWS application supplying its own sales analysis. Choose Call Analytics when the supported two-party recording and richer insights justify that purchase. A mono group discovery call should not be costed as if it met Call Analytics' input requirement. Existing AWS permissions and storage can make the operational fit more important than a small speech-price difference.

5. Google Cloud Speech-to-Text: cloud integration and delayed batch economics

Google's V2 service provides synchronous, batch and streaming recognition. Chirp 3 adds multilingual transcription and automatic language detection, with a narrower supported language/method set for speaker diarisation. The current guide documents diarisation for BatchRecognize and Recognize, including several English variants and languages such as German, Spanish, French and Japanese. Do not assume the full transcription language list also supports live diarisation. Chirp 3 guide.

Standard V2 recognition costs $0.016/minute in the first published band, or $0.96/hour. Dynamic batch's $0.003/minute, or $0.18/hour, processes at lower urgency and is attractive for delayed back-office work. Each recognised audio channel is billed separately. A dual-channel call can therefore cost $1.92/hour at the standard rate. Billing basis.

Batch processing reads audio from Cloud Storage and returns an operation, with output inline or in storage. Choose Google when the application already uses that cloud and can use the right model and method. Choose another API when immediate multi-speaker live handling is the central requirement or the additional cloud plumbing offers little benefit. Cheap delayed recognition is useful only when the resulting notes can arrive later.

6. OpenAI Audio API: current text transcription, with a legacy speaker-model boundary

The current recommended general file-transcription model is GPT Transcribe, at $0.0045/minute, or $0.27/hour. It supports context, keyword hints and multiple language hints. GPT Live Transcribe is the current low-latency streaming path, at $0.017/minute. A separate text model can extract sales fields from the transcript; its processing cost is additional. GPT Transcribe, file-transcription guide.

The specialised gpt-4o-transcribe-diarize model can return speaker-labelled segments, but OpenAI has announced its shutdown, alongside other legacy transcription models, on 26 February 2027. It is therefore a poor dependency for a new long-lived call-analysis build. The current general model's text price should not be presented as a complete replacement speaker-label service. Deprecation notice.

Choose the current Audio API when the surrounding application already uses OpenAI and the recorder or a separate component supplies speaker identity. Separately transcribing two full-duration channels doubles processed minutes: $0.54 per call hour at the file rate. Choose a current diarisation-capable speech provider when anonymous mono speakers are the main requirement. The file API's twenty-five-megabyte limit also makes compression or careful chunking relevant for long recordings.


A call that should not become a false commitment

Imagine a fictional buyer at 12:40 saying: “We could run a £15,000 pilot in November, subject to security review.” The representative at 13:05 suggests: “I can send a proposal for 15 November.” A weak pipeline writes “£50,000 budget approved; contract starts 15 November” into the CRM.

Three errors have combined: fifteen became fifty, a tentative pilot became an approved contract, and the representative's proposed date became the buyer's promise. Overall word accuracy can be high while the commercial record is unusable.

The useful output keeps the buyer's amount and conditional language attached to 12:40, and the representative's suggestion attached to 13:05. “Pilot interest, security review required” is a defensible summary. “Approved budget” is not. The next-step record can propose sending a pilot outline, with the call owner deciding the actual due date. These are editorial examples, not findings from a comparative product test.

Channel-separated telephony can make the role map clearer. Mono meetings require diarisation and a reliable association of labels with participants; an anonymous “Speaker A” is not automatically the customer. Transfers and overlapping voices deserve particular attention because a stable-looking transcript can retain the wrong label after the participant changes.


A useful audio sample and latency target

A compact evaluation can use twenty authorised two-minute clips: six clean discovery-call excerpts, four narrowband telephone excerpts, four with room noise, four with the team's actual accents/languages, and two with crosstalk or a transfer. Include the original channel structure where available. This is a proposed sample, not a claim that the APIs were run on it.

Label the product names, amounts, dates, objections and commitments manually. Judge transcription errors separately from speaker errors and analysis errors. For example, an otherwise readable transcript should fail the commercial extraction check if it changes £15,000 to £50,000 or attributes the seller's deadline to the buyer. For ordinary notes, minor punctuation has less importance than those fields.

Use different latency targets. A recorded-call process might target a reviewed note within ten minutes of the call ending. A live caption application might target final text within two seconds of the relevant utterance ending. Those are application objectives, not vendor guarantees; live interim text can be revised. Measure the actual endpoint and queue behaviour for the selected recipe instead of using an unrelated marketing speed number.

The forty minutes of sample audio cost only a few dollars across these base APIs. Preparing trustworthy references and comparing the output costs more. The selection should therefore turn on correct critical facts, usable speaker/channel identity and operator effort, rather than the smallest headline number.


The cost of 1,000 recorded calls

Assume 1,000 fifteen-minute calls each month: 15,000 call minutes, or 250 audio hours. The first comparison uses mono audio and the stated recorded base; analysis and operations are added separately.

Route Speech subtotal for 250 hours What the subtotal buys
AssemblyAI Pro + diarisation $57.50 Timed, speaker-labelled transcription; selected analysis extra
Deepgram Nova-3 monolingual $64.50 Recorded transcription with diarisation; intelligence extra
Gladia Starter $152.50 Recorded core transcription and diarisation; chosen analysis assessed separately
AWS standard batch, current US example basis $90 Standard transcription; Call Analytics not included
Google standard V2 $240 Standard recognition; suitable diarisation model/method required
OpenAI GPT Transcribe $67.50 Current text transcription; separate speaker attribution required

For the eligible AWS dual-channel analytics workload, the first-tier Call Analytics speech/analysis base is 250 × $1.80 = $450, before generative summarisation. Google dynamic batch's $0.18/hour would be $45 where its lower-urgency recipe is suitable. These are different output and timing choices, not interchangeable discounts on the same service.

Now assume a separately budgeted $0.02/call extraction step, $30/month storage/application hosting, four maintenance hours at $75/hour, and two minutes of review per call at $60/hour. That adds $20 + $30 + $300 + $2,000 = $2,350/month. The extraction rate is an editorial planning allowance, not a quoted vendor model tariff; built-in summaries do not eliminate review of critical CRM facts.

The resulting AssemblyAI subtotal is $2,407.50. If 900 calls produce accepted records, cost per accepted call is $2.68. Gladia's same mono planning subtotal is $2,502.50, or $2.78 per accepted call. Its $95 speech premium is repaid if it saves just 95 minutes/month at $60/hour, about 5.7 seconds per call over 1,000 calls. That is why correction effort matters more than a few cents per audio hour.

For distinct dual-channel input, Gladia's speech subtotal becomes $305, Google's standard recognition $480, and the OpenAI separate-channel design $135. AWS standard pricing includes the two channels in one duration. Preserve the recording's useful channel separation, but cost the actual processing path rather than comparing mono and stereo as equal bills.


Retention and delivery

Keep the organisation's call archive and the provider's processing retention separate. Deleting the provider transcript does not remove the CRM note, an application log or a cached download. Store a call/job identifier and process completion callbacks once, then retain the transcript and its analysis version according to the team's recording policy.

Provider Practical retention boundary
AssemblyAI Current security page offers account-level zero retention and training opt-out; export required results and delete/expire asynchronous artifacts deliberately
Deepgram mip_opt_out=true excludes content from model improvement and gives zero content retention after response; operational metadata remains
Gladia Paid default audio/transcript retention is three weeks; custom and zero retention are Enterprise options; zero-retention results are delivered through callbacks
AWS Own S3 output stays until removed; default service-bucket transcripts expire with the job at ninety days
Google Batch input/output can live in your Cloud Storage; apply the storage lifecycle alongside the speech-processing design
OpenAI Current data table lists audio transcriptions with no training, abuse-log content retention or application-state retention; a later text-analysis endpoint has its own policy

Sources: AssemblyAI security, Deepgram data, Gladia retention, AWS output, Google batch, OpenAI data controls.

Use recordings the business is authorised to process and give the call owner access to the supporting passage. The signal-based outbound playbook connects reviewed call outcomes to the next account action. A summary should make that action easier to assess, rather than quietly manufacture a commitment.


Which API to choose

Start with AssemblyAI Pro plus diarisation for a new recorded-call application in its supported languages, or Deepgram Nova-3 when recorded and live speech share the same developer operation. Compare correction time on the important fields; their base-price difference is small.

Choose Gladia when language switching and included core features materially improve the team's calls. Keep AWS or Google when existing cloud permissions, storage and deployment make that the simpler maintained system. Select AWS Call Analytics for its eligible agent/customer recording and richer call insights, and Google dynamic batch when the notes can wait.

Choose current OpenAI transcription when speaker identity is supplied separately and the application already uses its text models. A new mono-speaker-labelling build should use a current supported diarisation route, rather than assume the deprecated model will remain a stable long-term dependency.


FAQ

Is diarisation the same as knowing the customer's name?

Diarisation separates voices into labels. Identity requires participant metadata, known channels, speaker references or an identification step. A voice label can be consistent and still mapped to the wrong role, so keep the role map alongside the transcript.

Should we transcribe every call live?

Live processing earns its complexity when captions, accessibility or agent assistance need immediate text. For post-call CRM notes, recorded processing is often simpler and cheaper. Judge live latency separately from recorded completion time, and expect interim results to change.

Can the analysis model infer a buying commitment?

It can propose an interpretation, but a commercial commitment needs a supporting statement. Preserve the speaker, amount, conditions and timestamp. A request for a pilot subject to review should remain conditional in the CRM, even if the generated summary sounds confident.

Why do two-channel recordings change the comparison?

Separate channels preserve role information, but several APIs bill each processed channel. Google's standard rate and Gladia's distinct-channel processing can double, while current AWS standard pricing includes up to two channels in one duration. OpenAI's separate-channel design also processes twice the minutes.

Is the cheapest per-hour model usually the cheapest system?

At the example's volume, two minutes of review per call cost $2,000, far above any listed speech subtotal. A modest reduction in corrections can repay a higher API rate. Include the recorder, analysis, review, storage and maintenance before deciding whether to build an API application or buy a finished conversation-intelligence product.

Yananai A. Chiwuta

Author

Yananai A. Chiwuta

CEO & Co-Founder

Yananai A. Chiwuta is the CEO and Co-Founder of Forma Nôrden, where he builds managed acquisition systems for B2B companies through signal-based outbound and precision paid ad acquisition. He has built and exited two companies, most recently FunnelVision.

Celine Sky-Chiwuta

Article reviewed by

Celine Sky-Chiwuta

Co-Founder & CMO

Celine Sky-Chiwuta is the Co-Founder and CMO of Forma Nôrden, where she shapes the positioning and marketing behind the company’s managed acquisition systems. She previously served as CMO of FunnelVision through its 2025 acquisition.

Related Articles