AssemblyAI

AssemblyAI pricing quotes every speech rate in dollars per audio hour, from $0.15 on Universal-2 to $0.62 on the Dictation API.

Pricing Model:

Pricing Model:

Usage, drawn from a prepaid balance

Usage, drawn from a prepaid balance

Usage, drawn from a prepaid balance

Packaging Model:

Packaging Model:

Freemium, Good / Better / Best (GBB)

Freemium, Good / Better / Best (GBB)

Freemium, Good / Better / Best (GBB)

Credit Model:

Credit Model:

Prepaid, non-expiring. Auto-pay refills to a target you set

Prepaid, non-expiring. Auto-pay refills to a target you set

Prepaid, non-expiring. Auto-pay refills to a target you set

Updated on:

AssemblyAI pricing: what an audio hour really costs

AssemblyAI pricing quotes every speech rate in dollars per audio hour, from $0.15 on Universal-2 to $0.62 on the Dictation API. Rivals quote per minute, so nothing compares until you divide by 60. Do that and the headline rates look cheap. The bill lands elsewhere: add-ons stack additively on every hour, and streaming meters the WebSocket rather than the audio.

Key takeaways

  • Async transcription runs $0.21 an hour on Universal-3.5 Pro and $0.15 on Universal-2, which is $0.0035 and $0.0025 a minute once converted.

  • Streaming bills on how long the socket stays open, not on audio sent, and an unclosed session auto-closes after three hours and bills all three.

  • Voice Agent API costs $4.50 an hour, which is $0.075 a minute, matching Deepgram's Standard Voice Agent tier to the cent.

  • The free tier promises 185 pre-recorded hours, a figure that only works at a rate AssemblyAI stopped selling in early 2026.

AssemblyAI pricing in 2026

Product

Rate

Per minute

Universal-3.5 Pro, async

$0.21/hr

$0.0035

Universal-2, async

$0.15/hr

$0.0025

Universal-3.6 Pro Realtime

$0.45/hr

$0.0075

Universal-Streaming, English or multilingual

$0.15/hr

$0.0025

Sync API

$0.30/hr

$0.0050

Dictation API

$0.62/hr

$0.0103

Voice Agent API

$4.50/hr

$0.0750

What AssemblyAI actually meters

Three different clocks, and the one you get depends on the endpoint rather than the plan.

Pre-recorded meters the duration of the file you submit, pro-rated to the exact second, and failed transcripts aren't charged. Multichannel bills per channel, so a one hour stereo file costs two hours.

Streaming meters the WebSocket session, open to close. Idle time counts. AssemblyAI's own FAQ states that a connection held open for 60 minutes carrying 30 minutes of audio bills 60 minutes, and that sessions left open auto-close at three hours and bill the full three. Concurrent sessions accumulate in parallel, so one call dual-streamed under two session IDs for five minutes bills ten.

Voice Agent meters connected conversation time, billed per second, and the $4.50 covers speech to text, the proprietary agent LLM, text to speech, turn detection and hosting with no per-layer add-on.

Then there's the part the rate table buries. Add-ons bill additively per hour on top of the model. Speaker diarization adds $0.02 async and $0.12 streaming, sentiment $0.02, entity detection $0.08, PII text redaction $0.08 async and $0.12 streaming, topic detection $0.15, medical mode $0.15. A Universal-3.5 Pro job with diarization, sentiment, PII redaction and topic detection costs $0.21 + $0.02 + $0.02 + $0.08 + $0.15 = $0.48 an hour. The add-ons add $0.27, which is 129% on top of the base rate.

How credits work

New accounts get $50 in credits with no card, and the credits don't expire.

They cover pre-recorded, streaming, Voice Agent, Speech Understanding and Guardrails. LLM Gateway is excluded and bills from your balance on the first request.

The pricing FAQ sizes that $50 at "up to 185 hours of pre-recorded transcription and up to 333 hours of streaming". The streaming figure checks out: $50 ÷ $0.15 = 333 hours. The pre-recorded one doesn't match any current model. At Universal-3.5 Pro's $0.21 it buys 238 hours, and at Universal-2's $0.15 it buys 333. It resolves only at $0.27 an hour, the retired SLAM-1 rate from January 2026.

After the credits, you deposit funds and draw them down, and auto-pay recharges the card at a threshold you set. SSO costs $199 per connection per month against the same balance, and adding one forces your auto-pay threshold up to $199 and locks auto-pay on.

What happens when you hit the limit

The balance hits $0 and the API stops, returning Your current account balance is negative. Please top up to continue using the API.

Auto-pay is the whole defence. Without it, a runaway streaming session drains the balance and the next request fails. AssemblyAI names improperly closed streaming sessions as the most common cause of negative balances in its own docs, which is a candid thing to publish.

Concurrency throttles rather than blocks. Free accounts open five new streams a minute, pay-as-you-go starts at 100, and the ceiling rises 10% whenever you use 70% of it, with no stated maximum. Enterprise accounts switch to end-of-month invoicing, which removes the $0 cliff.

How AssemblyAI pricing has changed across all these years

Date

Milestone

Source

1 to 6 Oct 2026

Sync API cut from $0.45 to $0.30 an hour, a 33% reduction, announced nowhere

Vendor

Sep 2026

Dictation API launches at $0.62 an hour, and Sync API at $0.45

Vendor

1 Jul 2026

LLM Gateway in-region endpoints take a 10% surcharge over global routing

Vendor

Apr 2026

Voice Agent API launches at $4.50 an hour, replacing the Speech-to-Speech API

Vendor

3 Mar 2026

Real-time inline diarization ships as a streaming add-on at $0.12 an hour

Vendor

Jan to Feb 2026

Universal-3 Pro replaces SLAM-1 as flagship async, $0.27 down to $0.21

Vendor

10 Jan 2024

Latency work funds a cut to async and real-time rates

Vendor

Flexprice’s Take

AssemblyAI publishes its rates well and its billing basis badly.

The rate card is the best in speech to text. Every add-on carries a number, compliance costs nothing extra, and EU data residency is priced identically to US. Only volume discounts hide behind a sales call.

The unit is where buyers get caught. Quoting per hour against a market that quotes per minute isn't deceptive, but it does mean nobody compares correctly on the first pass. Universal-Streaming at $0.15 an hour is $0.0025 a minute, roughly half Deepgram's $0.0048 promotional Nova-3 rate. Flip to the flagship and Universal-3.6 Pro Realtime at $0.0075 a minute sits 56% above it. Across 100,000 streaming minutes that's $250, $480 and $750 respectively, three very different answers to one question.

Worse, the per-minute rate isn't what you pay. Streaming bills session duration, so the real multiplier is how disciplined your socket handling is, not which model you picked.

Three first-party sources disagree on the rest. The pricing FAQ describes monthly arrears invoicing, the docs describe a prepaid balance, and the agent-facing rate card at /llms/pricing.md omits both September launches.

Best For

Batch transcription where files are finite and add-ons are few.

Watch Out For

Streaming idle time, multichannel doubling, and comparing an hourly rate to a per-minute one.

Manish Choudhary

CEO & Co-founder, Flexprice

Billing audio hours, session seconds and LLM tokens on one invoice?

Flexprice meters all three in a single stream.

Flexprice’s Take

AssemblyAI publishes its rates well and its billing basis badly.

The rate card is the best in speech to text. Every add-on carries a number, compliance costs nothing extra, and EU data residency is priced identically to US. Only volume discounts hide behind a sales call.

The unit is where buyers get caught. Quoting per hour against a market that quotes per minute isn't deceptive, but it does mean nobody compares correctly on the first pass. Universal-Streaming at $0.15 an hour is $0.0025 a minute, roughly half Deepgram's $0.0048 promotional Nova-3 rate. Flip to the flagship and Universal-3.6 Pro Realtime at $0.0075 a minute sits 56% above it. Across 100,000 streaming minutes that's $250, $480 and $750 respectively, three very different answers to one question.

Worse, the per-minute rate isn't what you pay. Streaming bills session duration, so the real multiplier is how disciplined your socket handling is, not which model you picked.

Three first-party sources disagree on the rest. The pricing FAQ describes monthly arrears invoicing, the docs describe a prepaid balance, and the agent-facing rate card at /llms/pricing.md omits both September launches.

Best For

Batch transcription where files are finite and add-ons are few.

Watch Out For

Streaming idle time, multichannel doubling, and comparing an hourly rate to a per-minute one.

Manish Choudhary

CEO & Co-founder, Flexprice

Billing audio hours, session seconds and LLM tokens on one invoice?

Flexprice meters all three in a single stream.

Customer
Sentiment Highlights

"AssemblyAI is quite good, pretty cheap and easy with robust SDK support"

Hacker News user recommending transcription APIs, June 2025

"I somehow still haven't used up my credits after several years lol"

AssemblyAI customer on credit burn, Hacker News, August 2026

Frequently Asked Questions

Frequently Asked Questions

How much does AssemblyAI cost per minute?

Is AssemblyAI cheaper than Deepgram?

Does AssemblyAI charge for silence?

What is AssemblyAI's free tier?

Launch usage-based billing this week, not next quarter

Launch usage-based billing this week, not next quarter

Get Instant Feedback on Your Pricing | Join the Flexprice Community with 500+ Builders on Slack

Join the Flexprice Community on Slack