• Skip to main content
  • Skip to primary sidebar

DONNY BAARNS

  • HOME
  • DEMOS
  • Audiobooks
  • ABOUT
  • VIDEOS
  • CLIENTS
  • CONTACT
  • Blog
  • FAQ

Uncategorized

What Goes Into Voiceover Pricing? A Real Breakdown

Uncategorized

Voiceover pricing is the first question most buyers ask, and it’s one of the few honest “it depends” answers in the creative services world. The variables that drive a quote aren’t arbitrary — they reflect real differences in scope, usage, and labor. Understanding them before you reach out puts you in a much better position: you’ll know what information to have ready, why two projects that seem similar can come in at very different rates, and what production decisions directly affect your final cost.

Here’s how professional voiceover pricing actually works.


The Four Factors That Drive Every Quote

1. Word Count and Finished Length

The starting point for most voiceover quotes is either the word count of your script or the finished length of the audio. These aren’t the same thing, and which one applies depends on the type of project.

For corporate narration and training videos, pricing is typically based on finished audio length — how many minutes of completed, edited audio you receive. A two-minute corporate narration is priced differently than a fifteen-minute training module, and a forty-five-minute eLearning course is priced differently still.

For explainer videos and short-form digital content, pricing is often per video or per finished piece, with word count as the baseline for the estimate.

For eLearning specifically, per-word rates are common and give buyers a clear, predictable way to estimate costs as the script develops.

The practical implication: when you reach out for a quote, the more precisely you can describe the length of your project, the more accurate the estimate will be. “A few minutes” is not a useful starting point. “Approximately 400 words, expected to run around three minutes” is.


2. Where and How the Voiceover Will Be Used

This is the factor that surprises buyers most, and it’s the one that creates the biggest swings in voiceover pricing.

Voiceover rates aren’t just compensation for the recording session. They also include a usage license — the right to use that recording in a specific context, for a specific period of time, across a specific distribution. Change any of those variables and the rate changes with them.

Here’s what that looks like in practice:

A corporate training video used only inside a company, shown to employees and never distributed publicly, is priced at a significantly lower rate than the same length of audio used in a national television campaign. The performance is the same. The difference is entirely in how many people hear it, through what channels, and for how long.

A product explainer video that lives on a company’s own website has a different usage rate than that same video repurposed as a paid pre-roll ad on YouTube. Adding paid social distribution to an existing buy typically adds cost on top of the base web rate.

Regional use costs less than national use. A one-month license costs less than a one-year license. A video shown internally at a company conference costs less than one that runs on broadcast television.

The most common usage categories in professional voiceover:

Non-broadcast / internal use covers training videos, internal corporate content, and any video that lives only on the client’s own channels and isn’t distributed through paid advertising. This is the lowest usage tier and covers most corporate narration and training work.

Digital visual / web use covers product videos, explainers, and content placed publicly on websites, social platforms, or YouTube without paid promotion. When paid promotion enters the picture, that’s a separate and higher tier.

Online pre-roll and paid social covers ads served to audiences through paid placements: YouTube pre-roll, Facebook and Instagram ads, programmatic digital video. These rates are substantially higher than standard web use because the reach and commercial intent of the placement are greater.

Broadcast — radio and television — carries the highest usage rates, scaled by whether the placement is local, regional, or national, and by the term of use.

The question “where will this be used?” isn’t a bureaucratic formality. It’s how the price gets built.


3. Timing Requirements

Standard voiceover is delivered as a clean read: the talent performs the script, and the producer or editor fits the video to the audio. That’s the straightforward version.

When the video is already locked and the voiceover has to match specific timing — fitting exactly within a defined duration, syncing with particular cuts, or hitting specific moments in the video — the project requires time-synchronized recording. This is a categorically different and more demanding type of work.

Time-sync work requires multiple takes at different pacing adjustments, back-and-forth communication to determine what can be compressed and what can’t, possible script adjustments, and in some cases pickup sessions to patch specific sections. It takes more time, more session hours, and more communication than a standard read.

Professional voice talent charges more for time-synchronized work, meaningfully more. If you’ve built your video first and need the voiceover to match it exactly, expect the quote to reflect that additional scope.

The more efficient and less expensive path is to record the voiceover first and edit the video to match the audio. Producers who work this way consistently get better results for less money. But if the video is locked, the solution exists — it just costs more than a standard session.


4. Additional Services: Music and Editing

Background music and audio finishing are separate from the core voiceover performance, and how they’re handled affects the overall cost and deliverable.

When background music is included in the delivered file, the talent sources it from licensed royalty-free music libraries and provides a mixed, production-ready audio file. This is a value-add for buyers who don’t have their own audio production resources — you receive a finished file, not a raw read that needs further production.

If you have your own audio team or music assets, you can take delivery of the raw voiceover and handle the mix yourself. That keeps the voiceover scope clean and the rate straightforward.

Ask about this upfront. Some buyers assume music is always included. Some assume it never is. Clarifying early prevents surprises in the deliverable.


What This Means When You’re Requesting a Quote

The single most useful thing you can do before requesting a voiceover pricing quote is to have answers to these four questions ready:

How long is the script, or how long should the finished audio be? Word count, estimated duration, or both.

Where will this be used? Internal only, public website, paid social ads, broadcast, or some combination.

Is the video already finished? If so, does the voiceover need to match specific timing?

Do you need a finished, mixed file or a raw read? If music and mixing are needed, say so.

A quote built on clear answers to those questions will be accurate and actionable. A quote built on “I need a voiceover for a video, what does that cost?” will be a range wide enough to be almost useless.


Why Voiceover Pricing Varies Between Similar Projects

Buyers sometimes notice that two projects with similar word counts come in at different prices. This is almost always explained by one or more of the factors above.

A 300-word internal training video and a 300-word national television commercial are the same amount of copy. They’re priced differently because the usage is entirely different. The performance delivered in both sessions may be equally demanding. The license attached to the national spot is worth far more — to the brand, and therefore in the rate.

Similarly, a 500-word explainer video delivered as a standard read is priced differently than a 500-word explainer that has to hit exactly 90 seconds to match an animation that’s already been rendered. The words are the same. The scope of the session is not.

These aren’t arbitrary distinctions. They reflect how professional voice talent has been compensated in the industry for decades, in alignment with standards developed by organizations like the Global Voice Acting Academy, whose rate guide serves as an industry reference across usage categories.


The Right Way to Get a Voiceover Pricing Quote

Send the script. Describe the usage. Mention whether the video is locked. Ask about what’s included in the deliverable.

That’s the information a professional voice talent needs to give you an accurate number. Everything else — turnaround, revision policy, file format, session logistics — can be sorted once the scope is clear.

Voiceover pricing isn’t complicated once you understand what drives it. Most buyers who feel confused about rates are simply missing context about usage. Once that’s established, everything else falls into place quickly.


Donny Baarns is a professional voice talent and sports broadcaster with 13,000+ projects delivered for brands including Budweiser, New Balance, Verizon, Puma, and WebMD. He records broadcast-quality audio with standard 24-hour turnaround from his professional home studio. For project inquiries, contact donny@donnyvoice.com or use the contact form below.

Filed Under: Uncategorized

Voiceover Before Video: What AI Gets Wrong and What It Costs You

Uncategorized

Most production timelines treat voiceover before video as an afterthought. The voiceover gets added as the final layer, something you drop in once everything else is locked.

That order feels logical. It’s, in practice, one of the most expensive mistakes a buyer can make.

This isn’t a minor workflow preference. It’s a structural problem that affects cost, turnaround time, and the quality of the final product. Understanding why requires a look at how professional voiceover actually works, and where the AI tools most buyers are now using to build their productions break down in ways that aren’t immediately obvious.


The Foundational Problem: Video Is Rigid, Voice Is Not

A finished video has a fixed duration. Every cut, every transition, every moment of silence has already been decided. When you bring a voiceover artist in at that stage, you’re asking them to fit a living, breathing performance into a container that was built without them.

Professional voiceover isn’t a recitation. A trained voice artist controls pacing, breath, emphasis, and pause to serve the meaning of the copy and the emotional arc of the piece. Those elements aren’t decorative. They’re what makes the difference between a read that converts and one that sounds like someone reading a script.

When the video is already locked, those tools disappear. The artist isn’t performing anymore. They’re matching. Matching is harder, slower, more technically demanding, and more expensive than performing.


What AI Script Tools Get Wrong About Timing

Here’s where the problem has gotten significantly worse in the last two years.

A growing number of buyers are using ChatGPT, Claude, Gemini, or similar large language models to write their voiceover scripts. That’s understandable. These tools produce clean, professional-sounding copy quickly. The problem is what they can’t do: accurately estimate how long that copy will take a professional voice artist to read aloud.

The standard rule for spoken voiceover is well established. General spoken word delivery runs at approximately 150 words per minute, with a hard practical ceiling of around 180 words per minute at a very fast pace. Slower, more deliberate reads like corporate narration, meditation, or documentary can fall well below 130 words per minute. LLMs know this rule. They’ll tell you they know it if you ask. And yet the scripts they produce almost never conform to it accurately.

The reason isn’t that the model is ignoring the rule. The reason is structural.

Why the Math Never Works Out

Large language models generate text by predicting the next most likely token, which is a word or word-fragment, based on everything that came before it. They don’t have any internal experience of time. When a model writes a 60-second script and estimates its duration, it’s performing arithmetic on a word count after the fact. The model knows the rule the same way it knows any other fact stored in its training data. But knowing a rule and applying it as a real-time constraint during text generation are two entirely different operations.

The model is optimizing for the quality of the writing: coherence, completeness, professional tone. It’s not optimizing for a specific spoken duration. Word count is a byproduct of the generation process, not something the model is actively managing line by line. By the time it finishes writing and checks the math, the script already exists. The model has no mechanism to sense whether 160 words will take a professional 60 seconds or 80 seconds to deliver, because it’s never heard anything. It processes language as text, not as sound.

There’s also a subtler problem that compounds the error. LLMs tend to write copy that’s dense. Well-constructed sentences. Complete thoughts. Layered qualifications. That kind of writing looks efficient on a page but it’s systematically slower to deliver aloud than the casual conversational speech the 150 WPM baseline was built to measure. A sentence with three clauses and a parenthetical aside reads faster silently than it delivers in a professional studio read, where each clause needs a breath and a landing point.

The practical result: a script a language model estimates at 60 seconds will frequently run 70 to 80 seconds in a professional read, sometimes longer depending on the delivery style the project requires. For a meditation piece or a deliberate corporate narration, the gap between estimated and actual runtime can be even larger. And when the video’s already finished, that gap becomes a problem with a price tag attached.


The Time-Sync Premium: Why Fixing This Costs More

When a buyer arrives with a locked video and a script that doesn’t fit it, the work required to resolve that is categorically different from a standard voiceover session.

Standard voiceover: the artist reads the script, delivers a clean and well-paced performance, and the editor fits the video to the audio.

Time-synchronized voiceover: the artist has to record to a specific duration, often to the second. This typically requires:

  1. Multiple takes at different pacing adjustments to find what fits
  2. Back-and-forth communication with the buyer to identify which sections can be compressed and which can’t
  3. Potential re-editing of the script itself to bring total duration into range
  4. In some cases, pickup recordings and patch sessions to match specific cuts in the video

All of that takes more time, more sessions, and more communication. Professional voice artists, when they agree to take on time-synched work at all, charge significantly more for it. Rates for time-synchronized voiceover routinely run 50 to 100 percent higher than standard reads, and that’s before accounting for the additional revision rounds that typically result from trying to fit a performance into a container it wasn’t designed for.

The irony is that buyers who make the video first to save time often spend more of it, and more money, than they would have if they’d simply started with the voiceover.


The Workflow That Actually Works: Voiceover Before Video

The correct production order is: script, voiceover, video.

Write the script. Hire the voice artist. Record the voiceover. Then edit the video to match the audio.

This isn’t a new idea. It’s how broadcast production has worked for decades, because broadcast producers learned early that building video around audio is far easier than the reverse. A video editor can trim a shot by two seconds, extend a hold, or adjust a cut to match a beat in the audio. Those are minor adjustments. Rebuilding an entire performance to match a finished video isn’t.

When you have the voiceover first, you have a performance: pacing, emphasis, and timing that was crafted to serve the script. The video becomes the visual complement to that performance. That’s the intended relationship between the two elements, and it shows in the final product.


If You’ve Already Made the Video

This is the practical reality: many buyers reading this already have a finished or nearly-finished video and a script that may not fit it.

Here’s what to do.

Don’t guess. Don’t ask an AI tool to estimate whether your script will fit your video. As described above, those estimates aren’t reliable for professional voiceover purposes, for reasons that are built into how the technology works.

Get a real timing estimate. Send the script to a professional voice artist before committing to anything. An experienced VO artist can read through your script, tell you roughly where it lands in terms of duration, and flag the sections most likely to cause problems. This takes a few minutes and costs nothing at the inquiry stage.

Be upfront about the video. If the video is locked and can’t be re-edited, say so before getting a quote. This changes the scope of the project, the timeline, and the rate. A professional can work within those constraints, but needs to know about them before quoting.

Look for re-edit flexibility if you can. If your video was built in an AI tool, a standard editor, or even a platform like Canva, there’s often more flexibility than buyers assume. A section that runs five seconds long can frequently be extended with a hold on a visual, a slower transition, or a brief title card. Small adjustments at the edit stage can sometimes eliminate the need for a full time-sync session entirely.


A Note on Scripts Written by AI

If you used an AI tool to write your voiceover script, it’s worth having a professional review it before recording, and not just for timing. Scripts optimized for reading on a page often contain sentence constructions that are difficult to deliver naturally out loud. Run-on clauses, back-loaded qualifiers, and dense technical language that reads fine silently can create stumbling points in a live read.

A good voice artist will flag these issues. A great one will sometimes suggest small rewrites that preserve your meaning while making the copy easier to perform, which makes the final product sound better.


The Short Version

Record the voiceover first. Edit the video second.

If you’re using an AI tool to write your script, don’t trust its runtime estimate for professional voiceover purposes. The 150 words-per-minute rule is real and well established. The problem is that language models apply it as an afterthought rather than as a generative constraint, and they consistently produce copy that’s denser and slower to deliver than their estimates account for.

If your video is already finished and the script doesn’t fit, time-synchronized voiceover is available. But it costs more, takes longer, and requires more back-and-forth than a standard session. The earlier you bring a professional into the process, the less that problem costs you.


Donny Baarns is a professional voice talent and sports broadcaster with 13,000+ projects delivered for brands including Budweiser, New Balance, Verizon, Puma, and WebMD. He records broadcast-quality audio with standard 24-hour turnaround from his professional home studio. For project inquiries, contact donny@donnyvoice.com or use the contact form below.

Filed Under: Uncategorized

Primary Sidebar

©2026 Donny Baarns

The Voice That Makes Your Message Human

donny@donnyvoice.com

818-903-0084