OpenAI Whisper API Pricing 2026: $0.006/Min, $0.36/Hour
OpenAI Whisper API pricing is $0.006 per minute for whisper-1, which works out to $0.36 per audio hour. At that rate, 100 hours cost $36, 1,000 hours cost $360, and 10,000 hours cost $3,600 before any downstream processing or human review.
That is the answer if you specifically mean Whisper. For a new OpenAI transcription build, however, whisper-1 is no longer the only price worth modelling. OpenAI now lists lower-cost recorded-audio routes as well as more expensive live transcription models. The tables below separate those services so you can price an existing Whisper integration without confusing it with the rest of OpenAI’s speech-to-text stack.
| Whisper-1 audio volume | Calculation | API cost |
|---|---|---|
| 1 minute | 1 x $0.006 | $0.006 |
| 10 minutes | 10 x $0.006 | $0.06 |
| 1 hour | 60 x $0.006 | $0.36 |
| 10 hours | 600 x $0.006 | $3.60 |
| 100 hours | 6,000 x $0.006 | $36 |
| 1,000 hours | 60,000 x $0.006 | $360 |
| 10,000 hours | 600,000 x $0.006 | $3,600 |
Whisper cost formula: audio minutes x $0.006. If you already know the workload in hours, use audio hours x $0.36.
OpenAI transcription pricing: Whisper versus the newer models
OpenAI’s transcription catalogue now has several different price points. Rates below were checked on 24 August 2026 against OpenAI’s official API pricing. OpenAI presents gpt-4o-mini-transcribe and gpt-4o-transcribe with token pricing plus an estimated per-minute cost, while whisper-1, gpt-transcribe and the live transcription models have straightforward duration-based figures.
| OpenAI transcription model | Published or estimated cost/min | Approx. cost/hour | What to budget it for |
|---|---|---|---|
gpt-4o-mini-transcribe | $0.003 estimated | $0.18 | Lowest-cost managed transcription route to test for completed audio |
gpt-transcribe | $0.0045 | $0.27 | Current general-purpose transcription for completed files and batch-style workloads |
whisper-1 | $0.006 | $0.36 | Existing Whisper API integrations and workloads are standardised on Whisper output |
gpt-4o-transcribe | $0.006 estimated | $0.36 | GPT-4o transcription workflows where its output quality justifies the rate |
gpt-live-transcribe | $0.017 | $1.02 | Low-latency live transcription while speech is happening |
gpt-realtime-whisper | $0.017 | $1.02 | Existing real-time Whisper-style integrations |
Speaker diarisation needs its own cost line. gpt-4o-transcribe-diarize is a separate OpenAI model with built-in speaker attribution. OpenAI currently presents its pricing in terms of tokens rather than providing the same simple per-minute estimate shown for the main transcription rows. Do not copy the $0.006/minute gpt-4o-transcribe estimate into a diarisation budget without validating the actual usage pattern.
The price gap looks small at one hour. It becomes material at production volume. At 1,000 audio hours, whisper-1 is $360, gpt-transcribe is $270 and the estimated gpt-4o-mini-transcribe cost is $180. Moving 1,000 hours from Whisper to Mini therefore saves about $180 in base transcription charges. Whether it saves money overall depends on transcript quality and correction time.
Worked Whisper API cost examples for real workloads
Per-minute pricing is easy to quote and surprisingly easy to misread. Production teams usually think in terms of episodes, archives, customer calls, or continuous streams, so those are more useful units for budgeting.
Podcast transcription workload
A weekly 60-minute podcast costs $0.36 per episode with whisper-1. Four one-hour episodes in a typical four-episode month cost $1.44. A full 52-episode year costs $18.72 for the base transcription.
Back catalogues are still cheap enough that the arithmetic can surprise people. Five hundred 45-minute episodes contain 22,500 minutes of audio. At $0.006 per minute, the base Whisper API cost is $135. That figure does not include speaker labelling, transcript editing, chapter generation, summaries or publishing work.
100-hour, 1,000-hour and 10,000-hour archives
| Recorded audio | Mini at $0.003/min | GPT Transcribe at $0.0045/min | Whisper-1 at $0.006/min |
|---|---|---|---|
| 100 hours | $18 | $27 | $36 |
| 1,000 hours | $180 | $270 | $360 |
| 10,000 hours | $1,800 | $2,700 | $3,600 |
This is where a model benchmark becomes more useful than a price sheet. Saving $180 on 1,000 hours seems worthwhile until a cheaper model generates more than $180 in extra review work. As a simple decision threshold, $180 only buys six hours of human correction at an illustrative $30 per hour. A small regression in accuracy across a large archive can quickly consume that saving.
Continuous transcription for 24/7 audio
A single audio channel running continuously for 30 days contains 43,200 minutes. If you chop that stream into completed recordings and send every minute through whisper-1, the base transcription cost is $259.20 per 30-day month. The same duration is about $129.60 at the estimated Mini rate and $194.40 with gpt-transcribe.
True live transcription should be budgeted separately. At $0.017 per minute, 43,200 minutes on gpt-live-transcribe or gpt-realtime-whisper is $734.40 per 30-day month for one continuous audio stream. Four independently transcribed channels would multiply the audio volume by four, so start the cost model with channel-minutes rather than wall-clock minutes.
What Whisper actually charges for: duration, not file size
The published whisper-1 price is based on audio duration. A 60-minute file is therefore a $0.36 base transcription workload, whether the upload is a large WAV or a more compressed audio file of the same length. Compression can reduce transfer time and storage, but it does not turn 60 minutes of speech into fewer billable audio minutes.
The same logic applies to long sections of silence, hold music or dead air. If they remain in the audio you send for duration-priced transcription, budget for that duration. Trimming non-speech before upload can reduce the amount of audio you process, provided the trimming step does not clip quiet speakers or remove context your transcript needs.
Many short files create a different problem. Ten thousand six-second voice notes contain the same total audio duration as 1,000 one-minute files, but the short-file workload creates much more request, queue and retry overhead around the API. Do not confuse that infrastructure overhead with the published transcription rate. Price the audio minutes first, then model the operational cost of moving a very large number of jobs through your own system.
Whisper API versus running Whisper locally
Open-source Whisper can be run on your own hardware, so there is no OpenAI per-minute API charge in that setup. It is not the same as zero-cost transcription. Local processing shifts the bill to hardware, electricity, cloud compute, storage, deployment, monitoring, and engineering time.
| Cost factor | Managed whisper-1 API | Self-hosted Whisper |
|---|---|---|
| OpenAI usage fee | $0.006 per audio minute | No OpenAI per-minute fee |
| Compute | Included in the API rate | You provide CPU, GPU or rented compute |
| Idle capacity | No dedicated transcription server to keep warm | Can dominate cost when expensive hardware is underused |
| Scaling | Mostly an API and quota problem | You own concurrency, queues, workers and failure recovery |
| Privacy and local control | Audio is processed through the managed service | Can remain inside your own environment |
| Best economic fit | Small, variable or operationally simple workloads | High, predictable utilisation or workloads that require local processing |
A recurring pattern among people operating local Whisper stacks is that utilisation matters more than headline GPU speed. A powerful GPU that spends most of the day idle can be more expensive per accepted transcript than an API. The opposite can happen with a large, steady queue that keeps hardware busy. CPU deployments can work too, but throughput may be too low for long recordings or high concurrency, so compare completed audio hours per day rather than hardware specifications in isolation.
Which OpenAI transcription model should you budget for?
| Your workload | Model to price first | Reason |
|---|---|---|
| Existing Whisper integration | whisper-1 | Keep the known $0.36/hour baseline unless migration produces a meaningful saving or output improvement |
| Large archive where the lowest managed rate is the priority | gpt-4o-mini-transcribe | Its current estimated cost is about half the Whisper-1 rate |
| New general recorded-audio feature | gpt-transcribe | Current OpenAI transcription route for completed files and batch-style work at $0.0045/minute |
| Existing GPT-4o transcription workflow | gpt-4o-transcribe | Do not migrate purely on price until the replacement passes your transcript regression set |
| Live captions or live speech-to-text | gpt-live-transcribe | Designed for low-latency transcript updates while speech is happening |
| Local or offline transcription | Self-hosted Whisper | No managed per-minute charge, but you take responsibility for the infrastructure |
If the decision is about accuracy, multilingual performance, hallucination risk or local deployment rather than price, use our OpenAI Whisper review. Keeping quality testing separate from this pricing page prevents the cost answer from disappearing under review material.
The cheapest transcription rate can still produce a higher production bill
The API line item is only one part of an accepted transcript. Four costs regularly change the result after a team moves from a calculator to production.
- Retries and duplicate processing: a timeout in your application does not always prove the upstream transcription failed. Use durable job states and avoid blindly resubmitting audio.
- Speaker labelling: a raw transcript and a diarised meeting transcript are different outputs. Do not assume the
whisper-1rate includes a complete speaker-attribution pipeline. - Downstream AI: summaries, action items, redaction, classification, embeddings and CRM extraction are separate workloads.
- Human correction: names, numbers, overlapping speakers and technical terminology can make a nominally cheaper transcript more expensive to accept.
The practical metric is cost per accepted transcript, not cost per uploaded minute. Build a representative audio set, run the models you are considering, measure correction time and then combine API cost with the labour or downstream processing required to reach usable output.
See more of our reporting in Google Top Stories, AI Overviews and AI Mode.
Common OpenAI Whisper pricing mistakes
- Assuming every OpenAI speech-to-text model costs $0.006/minute. That is the
whisper-1rate, not the whole current transcription catalogue. - Using the Whisper figure for a new build without benchmarking newer models. Mini and GPT Transcribe have lower published or estimated base costs.
- Using a live model for completed recordings. The live rate is substantially higher because it solves a different latency problem.
- Expecting audio compression to reduce transcription cost by itself. If the duration stays the same, the Whisper pricing calculation stays the same.
- Calling self-hosted Whisper free. The API fee disappears, but compute, capacity planning, and operations do not.
- Ignoring correction cost at scale. A few extra seconds of review per recording can erase a small per-minute saving across a large archive.
OpenAI Whisper API pricing FAQs
How much does the OpenAI Whisper API cost per minute in 2026?
whisper-1 costs $0.006 per audio minute. One thousand minutes, therefore, cost $6 in base transcription charges.
How much does one hour of Whisper API transcription cost?
One audio hour costs $0.36 with whisper-1. The calculation is 60 minutes x $0.006.
How much do 100 hours of Whisper cost?
One hundred audio hours contain 6,000 minutes. At $0.006 per minute, the base whisper-1 API cost is $36.
How much do 1,000 hours of Whisper cost?
One thousand audio hours contain 60,000 minutes. At the current whisper-1 rate, that is $360 before retries, downstream processing or manual correction.
Is GPT-4o Mini Transcribe cheaper than Whisper?
Yes, on the current published estimates. OpenAI lists gpt-4o-mini-transcribe at an estimated $0.003 per minute, versus $0.006 per minute for whisper-1. At 1,000 audio hours, the base difference is about $180. Benchmark output quality before moving production traffic solely for that saving.
What is GPT Transcribe pricing?
gpt-transcribe is priced at $0.0045 per minute, equivalent to $0.27 per audio hour. That puts it between Mini and Whisper in terms of base recorded-audio cost.
Does compressing audio make the Whisper API cheaper?
Not if the audio duration stays the same. A smaller file can reduce upload bandwidth and storage, but the published rate is based on audio minutes rather than megabytes.
Does silence increase Whisper transcription cost?
Budget according to the duration you actually send for transcription. Long silence, hold music or dead air still occupies audio time if it remains in the submitted recording. Preprocessing can remove unwanted non-speech, but it should be tested carefully so that quiet speech is not clipped.
Is the Whisper API free?
The managed whisper-1 API has a published usage price of $0.006 per minute. Account credits or promotions can affect what an individual account pays, but they should not be used as the long-term production price. Running open-source Whisper on your own hardware avoids the OpenAI per-minute API fee, while creating your own compute and operational costs.
What should you budget for OpenAI Whisper?
If you specifically need whisper-1, use $0.006 per minute or $0.36 per audio hour as the base budget. That means $36 for 100 hours, $360 for 1,000 hours and $3,600 for 10,000 hours.
For a new OpenAI transcription workload, do not stop at the Whisper number. Price gpt-4o-mini-transcribe for the lowest managed base cost, gpt-transcribe for current general recorded-audio work, and a live transcription model only when the product genuinely needs low-latency output while somebody is speaking. If you want to compare OpenAI against other providers rather than just price the OpenAI stack, use our speech-to-text API pricing comparison.
At production scale, the winning route is the one with the lowest cost per accepted transcript. The model price is easy arithmetic. Correction time, retries, speaker handling and infrastructure are where the real budget diverges.


