AInora
Pricing ModelsVoice AIVendor SelectionProcurement

Per-Minute vs Per-Outcome AI Voice Pricing: What Each Model Says About the Vendor

JB
Justas ButkusFounder, Ainora
··11 min read

TL;DR

Almost every quote for an AI phone agent takes one of four shapes: per minute of call time, per call or conversation, per seat or flat platform fee, or per outcome. The shape matters more than the number, because it decides what the vendor gains by improving the product and what they lose. Per minute means they earn more when calls run longer. Per call prices attempts, not results. Per seat is indifferent to whether the thing works. Per outcome aligns best and is hardest to define honestly, so the test is whether the definition is published. One question exposes all of it: what happens to your revenue if our call volume halves because the agent got better?

4
Shapes almost every quote takes
60 sec
Documented telephony billing increment
52
UK schemes with a payment-by-results element (NAO, 2015)
3
Questions an outcome definition must answer

If you are trying to work out how much an AI receptionist costs, published prices will not help much. They are quoted in units that do not line up against each other, they move without notice, and two vendors quoting the same figure are often selling different amounts of work. The number is the least stable thing on the page.

The structure underneath it barely moves, and it is far more informative. A pricing model is a statement about what the vendor is optimising, and read that way it tells you what the product will become over the life of the contract. Everything below comes from vendors' own pricing pages, checked in September 2026, because a vendor is the one source that cannot be wrong about its own price list.

How is AI phone agent pricing actually structured?

Four shapes cover nearly the whole market, and each answers "what does the vendor get paid for" differently. That answer is the incentive, and the incentive is what survives after the introductory rate expires and the person who sold it to you has moved on.

Pricing shapeThe vendor is paid forBuyer can forecast?The incentive it creates
Per minute of call timeTime spent talkingPoorly, it tracks customer behaviourLonger conversations raise revenue, so efficiency is a cost to the vendor
Per call or conversationAttempts handledWell, volume is already countedMore calls answered, regardless of what any of them achieved
Per seat or flat platform feeAccess to the softwareVery well, it is a fixed lineRenewal, not results; the labour to operate it sits with the buyer
Per outcomeA defined resultWell, if the definition is soundResults, and a strong incentive to define the result favourably

Vendors are rarely coy about which one they use. Bland's pricing page describes the product as "Pay per minute. One flat rate covers the LLM, speech-to-text, and text-to-speech". Retell publishes "$0.07-$0.31 / min for AI Voice Agents". Vapi lists "$0.05 / min" for its own hosting, with model provider costs charged "At cost" on top. Component-by-component figures for all of these are in our per-minute cost breakdown; this piece is about what the shape implies rather than what the components add up to.

What does per-minute pricing tell you about a vendor's incentives?

It tells you that the vendor's revenue rises when your callers talk for longer. That is the whole of it, and it is enough.

The uncomfortable version of this argument accuses vendors of padding calls. That version is wrong, and it is also unnecessary. Product direction is not a series of ethical decisions, it is a queue: a list of things that could be built, sorted by what they are worth. Under per-minute billing, "resolve this in two turns instead of six" sits on that list as an item that reduces the vendor's revenue if it works. It does not get vetoed. It simply never reaches the top, and no meeting is ever held at which someone says why. Run that queue for two years and the product ends up fast, fluent and slightly long-winded, without any individual having made that choice.

Economists have a formal name for the surrounding problem. Holmstrom and Milgrom's multitask principal-agent analysis, published in the Journal of Law, Economics, and Organization in 1991, modelled what happens when a job has several dimensions and only some of them can be measured: rewarding the measured dimension pulls effort away from the ones nobody is scoring, wherever the two compete for the same effort. Their illustration is incentive pay for teachers based on their students' test scores, and the skills that go untested. The voice equivalent is a supplier paid on minutes whose product is optimised for everything except brevity.

The second consequence lands sooner than the first. Under per-minute billing you cannot budget, because the invoice is a function of how talkative your customers happened to be. A confusing letter, a product recall, a bad week of weather: each arrives later as a bill, and finance reasonably dislikes a line that moves for reasons nobody in the building controls. The model also prices the wrong end of the problem. A missed call has a cost, as the arithmetic on missed calls sets out, and per-minute billing charges you for the answered ones instead.

One variant is worth checking for specifically, because it charges more precisely when you need it most. ElevenLabs publishes its agents platform at $0.080 per additional call minute with "burst pricing" at $0.160 per minute once you exceed the concurrency in your tier. That is transparent, and worth modelling, because a call spike is exactly the event that made you buy an automated line.

What is a billing increment, and why does it move the bill more than the rate?

Underneath every per-minute rate is a rounding rule that almost never appears in a proposal. The billing increment is the block of time a call is rounded up to before the rate is applied, and telephony documents it openly because carriers have argued about it for decades.

Telnyx publishes 60-second increments in its own help centre, spelling out that "if your call is 11 seconds long, you will be charged for 60 seconds" and that a 96-second call is charged as 120 seconds. Twilio's own documentation notes that "by default we round up to the next minute, for a 25 second call duration on the logs the invoice will say 1 minute", which is why the call logs and the invoice disagree.

This matters far more for an automated line than a human one, because an automated line answers a great many very short calls: wrong numbers, hang-ups, opening-hours questions, silent lines, robocalls. Under a 60-second increment an eleven-second wrong number and a fifty-nine-second booking cost the same, so two businesses with identical total talk time can receive very different invoices.

The increment is a point of competition, which is the clearest evidence it is not trivial. ReceptionHQ advertises the opposite policy on its own pricing page: "On ReceptionHQ per-minute plans, you're charged for the exact seconds taken for managing each call. So if a receptionist takes 61 seconds for handling a call, you're billed for 1 minute and 1 second", contrasting that with providers where a 61-second call is billed as two minutes. Ask for the increment in writing before comparing two rates, because without it you are not comparing two rates.

Is per call or per conversation better than per minute?

On forecasting, unambiguously yes. Call volume is something you already measure, it moves slowly, and it does not depend on how chatty a caller is. It also removes the increment problem, because there is no partial unit to round.

The human answering-service market worked this out first, and puts the incentive argument more plainly than any AI company has. Smith.ai answers "why do you bill by the call" on its own pricing page: "Versus by minute? Billing per client call saves you money, and there's no mystery accounting involved. You're trusting us to represent you. You wouldn't rush a client off the phone and we don't either". Ruby sells virtual receptionist plans where, in its own words, "The only difference between our plans? The number of receptionist minutes and chats". Two vendors in one category, two opposite units, each with a coherent story. Both appear in our comparison of AI and human reception costs.

What per-call pricing does not do is price results. A hundred calls that resolved nothing costs exactly what a hundred that filled the diary costs. If the agent takes a message rather than booking, or answers politely and creates no record anywhere, the invoice is identical. That gap is the subject of taking a message versus actually booking, and it is the reason per-call pricing pairs badly with a vendor who has no write access to your system of record. The unit rewards answering. Only the integration turns answering into anything.

What does a per-seat or flat platform fee actually buy?

Predictability, and nothing else guaranteed. A seat licence is the most forecastable line in software and the one least connected to whether the product works. The vendor is paid the same whether the agent resolved every call or none, and the renewal conversation is about adoption rather than outcomes.

It is also the model most likely to hide the real cost, because a seat is usually attached to a product you operate yourself. The prompt writing, the testing, the change management when a policy moves, the person reviewing transcripts on a Monday: none of it is on the pricing page, all of it is on someone's calendar, and it usually lands in a different budget from the one that approved the licence. That is the substance of the in-house versus managed comparison, and it is why the honest comparison is never licence against licence.

The instructive thing about seat pricing now is that large vendors are hedging it. Salesforce's Agentforce page publishes three models side by side: a user licence at $5 per user per month, marked "(Requires Flex Credits)", consumption priced per 100,000 Flex Credits, and a flat "Pay per Conversation" option at $2. Intercom sells a per-seat subscription and prices its AI agent alongside it "From $0.99 per Fin outcome". Microsoft still sells its Copilot as a per-user, per-month add-on. When one vendor offers the same capability under three different units, the unit you pick is a decision about your own incentives, not a billing preference.

Does outcome-based pricing fix the alignment problem?

It aligns the incentive better than anything else on the list, and it replaces the alignment problem with a definition problem, which is where most of the honesty in this category now lives.

As soon as money depends on a result, three things must be settled in writing before the first invoice: what counts as the result, which system it is read from, and who decides when the two sides disagree. A vendor offering outcome pricing without answering those three is selling a slogan, and you can usually tell without asking, because the serious ones publish the answer.

Intercom publishes it in its own help documentation. A resolution counts when, after the agent's last answer, the customer confirms it was satisfactory or leaves without asking for more. A greeting does not count. Most tellingly, the charge can be reversed: "If a conversation is considered resolved (confirmed or assumed), but the customer later returns to the same conversation seeking further assistance (even across billing periods) that resolution will be deducted and not charged". That clause is what an honest outcome definition looks like: it costs the vendor money, and exists to stop them being paid for a result that did not hold.

Zendesk states the unit and the exclusion on its pricing page: "Paying per automated resolution means that you pay only for customer requests that were successfully resolved by the AI agent, without any escalation to a human agent". Sierra states the principle on its product page: "we've pioneered an entirely new business model - outcome-based pricing - where you pay only when the software achieves specific, valuable outcomes", and does not publish the qualifying definition there. That does not rank the three companies. It ranks how much of the contract you can read before getting on a call, which is a real procurement difference.

Outcome payment is not new, and has been audited at national scale. Across six central government departments the UK National Audit Office identified 52 schemes with a payment-by-results element, worth at least fifteen billion pounds and reported that payment by results "is a technically challenging form of contracting, and has attendant costs and risks that government has often underestimated". Two findings transfer to software procurement. The first is substitution: where the real outcome could not be measured, schemes paid for a measurable output instead, the report's example being a malaria programme paid per bed net distributed rather than per infection avoided. The second is selection: "A poorly designed scheme may create perverse incentives for providers, such as welfare-to-work providers prioritising people who are easier to help and 'parking' those who are harder to help." The parked cases on a phone line are the awkward callers and the ambiguous requests, precisely the calls an agent's architecture is tested by, and precisely the ones that belong on a written list of what the agent does not handle rather than inside a payment formula.

Why is the sticker price rarely the cost?

Because the cost is the sticker plus the work still left with you, and a cheaper tier usually moves work rather than removing it. None of this is hidden. It is published, and simply not in the number people copy into a spreadsheet.

Vapi's pricing page is unusually explicit about this: the platform rate covers hosting, and model provider costs are charged "At cost ($0 if you bring your own API key)", with HIPAA listed at $2K/month and Zero Data Retention at $1K/month. Nothing there is misleading. It does mean that a compliance requirement, and a decision about where your call recordings end up, are separate purchases rather than properties of the product. At the other end of the market, Synthflow's pricing page publishes no self-serve tiers and states that "Enterprise contracts start at $30,000 annually", with everything else scoped by sales. A price you cannot see is still information: the vendor has concluded the work varies too much to list.

The comparison that survives contact with reality prices the whole job. Who writes the policy down. Who tests a change before a live caller meets it, the discipline in how an agent gets verified before it speaks. Who owns the integration when your booking system changes a field. Who is accountable at eight in the morning when the diary looks wrong, the failure mode in why AI phone agents book the wrong appointment. Each has a price whether or not it appears on an invoice, and where two quotes for the same workflow are far apart, this is the list to price on both sides before concluding that one of them is cheap. It is the same list that decides whether you build this or buy it, itemised in what it actually costs to build a voice agent in house, and it is what makes the comparison against in-house staffing mean anything.

What should you ask a vendor about their pricing model?

One question does most of the work, and its virtue is that it does not sound like a test:

The question

What happens to your revenue if our call volume halves because the agent got better at resolving things on the first call?

It reads as a boring capacity question rather than a challenge, which is what makes it fair to put in an email rather than save for a call. If the honest answer is that they earn less, nothing improper follows. What you have learned is which improvements run against the vendor interest, and therefore which ones belong in the agreement rather than a roadmap conversation.

Two follow-ups finish it. What is the billing increment, in writing. And if the word outcome appeared anywhere in the answer: where is it defined, which system is it counted from, and who decides a disputed one. A vendor who answers all three in one email is running a business you can model. One who cannot is asking you to price something neither of you has defined. The rest of the shortlist questions are in our vendor evaluation checklist and the evaluation guide.

The shorter version of this argument, structured as a decision rather than an essay, is on our page on pricing models and incentives. Why this site publishes no tier table is explained on the pricing page, and the same criteria above are ones we expect to be held to.

What a useful next conversation looks like

Not a demo, and not a quote against a blank page. A working session: about forty-five minutes on your actual call flow, where we take the policies your front desk already follows, push on them until the edge cases show themselves, and write down what your rules turn out to be. You keep that written version whether or not anything else happens, and it is what makes any pricing conversation, with us or anyone else, about something real.

If something does happen next, it is deliberately small. One workflow, missed calls and after hours, roughly two weeks, and nothing else moves until that one is behaving. If you would rather hear it first, there is a live voice demo, and if you would rather send the flow across in writing, tell us what your calls look like.

Frequently Asked Questions

Published rates are quoted in units that do not compare directly, a per-minute platform rate against a per-conversation rate against an annual enterprise floor, and they move without notice, so the figure on its own is a weak basis for comparison. The structure is more informative and more stable. Almost every quote is per minute of call time, per call or conversation, per seat or flat platform fee, or per outcome, and each shape creates a different incentive for the vendor. Note also that a growing number of vendors publish no price at all: Synthflow states that enterprise contracts start at $30,000 annually and scopes everything else through sales, and Zendesk describes its per-resolution model on its pricing page without publishing a per-resolution figure.

Nothing dishonest, and one structural problem. The vendor is paid more when your calls run longer, so any improvement that shortens a call reduces their revenue. That does not require anyone to pad a call; it only means brevity never reaches the top of the roadmap queue. It also makes your bill unforecastable, because it tracks how talkative your customers were rather than anything you control. Bland, Retell and Vapi all publish per-minute rates on their own pricing pages, which makes the model easy to verify before you commit.

It is the block of time a call is rounded up to before the per-minute rate is applied. Telnyx documents 60-second increments, where an 11-second call is billed as 60 seconds and a 96-second call as 120 seconds. Twilio documents that by default it rounds up to the next minute, so a 25-second call appears on the invoice as one minute. Automated lines answer a lot of very short calls, so the increment can change the bill more than the rate does. ReceptionHQ competes on the opposite policy, advertising that customers are charged for the exact seconds taken on each call.

It aligns the incentive better and it is harder to write honestly. Payment tied to a result requires three things to be settled in advance: what counts as the result, which system it is read from, and who adjudicates a dispute. Intercom publishes its rule, including that a charged resolution is deducted if the customer returns to the same conversation later. Zendesk publishes that an automated resolution excludes anything escalated to a human. The UK National Audit Office, which identified 52 UK government schemes carrying a payment-by-results element across six departments, found that where the real outcome could not be measured, schemes specified a measurable output instead.

Usually because the work varies more than a tier table can express: call volume, concurrency, telephony setup, integrations, security requirements and launch support all move the effort involved. Synthflow says exactly this on its pricing page, listing the factors that scope final pricing. The absence of a public number is itself information. It signals that the vendor is selling an implementation rather than a licence, which is a different purchase with a different risk profile.

It is the most predictable model and the least connected to whether the agent works, because the fee is identical whether every call was resolved or none were. It also tends to sit on a product you operate yourself, so the real cost includes the staff time to write, test and maintain the agent, which never appears on the pricing page. Notably, large vendors are hedging: Salesforce publishes a per-user licence, consumption credits and a per-conversation rate side by side, and Intercom pairs a per-seat subscription with per-outcome pricing for its AI agent.

Ask what happens to their revenue if your call volume halves because the agent got better at resolving things on the first call. It reads as a routine capacity question rather than a challenge, so it is a fair thing to ask in writing. If the answer is that they earn less, you have learned which product improvements run against the vendor interest and therefore need writing into the agreement rather than waiting for. Follow it with a request for the billing increment in writing, and, if outcomes were mentioned, for the outcome definition and the system it is counted from.

JB
Justas Butkus

Founder & CEO, AInora

Building AI digital administrators that replace front-desk overhead for service businesses across Europe. Previously built voice AI systems for dental clinics, hotels, and restaurants.

View all articles

Ready to try AI for your business?

Hear how AInora sounds handling a real business call. Try the live voice demo or book a consultation.