The Short Answer
Measuring AI ROI in a small business takes three things: a 30-day baseline captured before you buy anything, five numbers tracked weekly after launch (answered-call rate, speed-to-lead, bookings per 100 inquiries, hours reclaimed from admin work, and cost per resolved conversation), and one payback formula: setup cost divided by net monthly gain. If you cannot produce those numbers, you do not know your ROI. You know your vendor’s marketing.
This guide gives you the formulas, shows where each input lives in your phone system and CRM, works two complete examples with conservative inputs (a six-operatory dental practice and a two-truck HVAC company), and names the recycled statistics the AI industry keeps citing without a source, so you can pressure-test any pitch. Including ours.
Why Do AI ROI Statistics Contradict Each Other?
Because the headline stats measure different cohorts doing different things. An IDC study commissioned by Microsoft (2024, based on interviews with 4,000+ business leaders) found companies realize an average of $3.70 for every $1 invested in generative AI ($10.30 for top performers), with value typically arriving within about 13 months. Meanwhile MIT NANDA’s 2025 “GenAI Divide” report, built on 52 executive interviews, 153 leader surveys, and 300 analyzed deployments, found that 95% of enterprise gen-AI pilots produce no measurable P&L impact. Both findings are real. They are measuring the winners and the losers of the same game.
RAND Corporation’s 2024 research (interviews with 65 data scientists and engineers) frames the divide more bluntly: more than 80% of AI projects fail, twice the failure rate of ordinary IT projects, and the root causes are organizational rather than technical: misaligned objectives, underestimated data work, and chasing technology instead of a business outcome. MIT’s successful 5% shared one trait: they embedded AI deep inside a specific workflow instead of bolting a generic tool onto the org chart.
Notice what actually separates the cohorts. It is not model quality, industry, or budget. It is whether anyone defined the outcome, measured a starting point, and checked the delta afterward. That is a measurement problem, and measurement problems are solvable with a spreadsheet.
Adoption, meanwhile, keeps compounding around you. Intuit QuickBooks’ AI Impact Report (January 2026, drawing on 34,000+ surveyed owners plus anonymized data from 5.3 million QuickBooks businesses) found 68% of small businesses now use AI regularly (up 42% year over year), with more revenue (41%), lower costs (25%), and shorter workdays (24%) as the top reported benefits. Salesforce’s Small & Medium Business Trends survey (December 2024, 3,350 SMB leaders) found 91% of SMBs using AI say it boosts revenue. Read the verbs carefully: “report,” “say.” Self-reported benefit is sentiment, not measurement. The same Salesforce survey found 83% of growing SMBs had adopted AI versus 55% of declining ones, which proves adoption correlates with growth, not that it caused any of it. The rest of this article is about replacing sentiment with arithmetic.
Quick Definitions Before the Math
Baseline: 30 days of your business’s numbers captured before AI goes live: the “before” photo. Without it, no claim about “after” means anything.
Answered-call rate: the percentage of inbound calls a human or AI actually picks up. Voicemail and calls abandoned in a queue count as missed.
Speed-to-lead: minutes between a lead’s inquiry (form fill, missed call, chat message) and your first substantive response. Measured from CRM timestamps, not memory.
Containment rate: the percentage of conversations an AI resolves end-to-end without a human stepping in. Also called resolution rate.
Cost per resolved conversation: everything you pay for a channel in a month, divided by the number of conversations that channel actually resolved.
Payback period: setup cost divided by net monthly gain: how many months until the project has paid for itself. Distinct from ROI, which is the ongoing return after payback.
What Should You Measure Before Buying AI? The 30-Day Baseline
Before you sign anything, spend 30 days capturing four families of numbers: call flow (total, answered, missed, after-hours), lead response time, booking rate, and admin hours. Every one of them lives in a system you already pay for, and the whole exercise is two to four hours of real work spread across a month. It converts every vendor conversation from “trust us” into “prove it.”
RAND’s 2024 failure research flagged underestimated data work as a leading killer of AI projects. A baseline is the cheapest data work you will ever do, and it doubles as negotiating leverage: a vendor who knows you have measured your own missed-call count cannot quote you fantasy ROI. Here is where each number lives:
| Baseline metric | Where to pull it | What to write down |
|---|---|---|
| Total inbound calls | Phone system admin portal (RingCentral, Weave, Grasshopper) or your carrier’s call log | Calls per month, split weekday vs. weekend |
| Answered vs. missed | Same report: count voicemails and abandoned queue calls as missed | Answered-call rate % |
| After-hours share | Filter the same call log by your business hours | % of calls arriving outside open hours |
| Speed-to-lead | CRM timestamps: lead created → first logged call, text, or email | Median minutes (medians resist one weekend outlier) |
| Booking rate | PMS/FSM reports (Dentrix, Open Dental, ServiceTitan, Housecall Pro, Jobber) | Bookings per 100 inquiries, plus show rate |
| Admin hours | One-week time log by your front desk or dispatcher, multiplied by 4 | Hours per month on phones, scheduling, re-keying data |
If you would rather not do this yourself, this exact baseline is the first deliverable of our AI audit, but nothing in that table requires us. Your phone portal has an export button.
What Are the 5 Numbers That Actually Measure AI ROI?
Five numbers, tracked weekly against the baseline, tell you whether an AI deployment is paying: answered-call rate, speed-to-lead, bookings per 100 inquiries, hours reclaimed from admin work, and cost per resolved conversation. Everything else on a vendor dashboard (sentiment scores, “engagement,” raw conversation counts) is decoration.
| # | The number | Formula | Source of truth |
|---|---|---|---|
| 1 | Answered-call rate | Answered ÷ total inbound calls | Phone system logs |
| 2 | Speed-to-lead | Median minutes, inquiry → first response | CRM timestamps |
| 3 | Booking rate | Bookings ÷ qualified inquiries × 100 | PMS / FSM / calendar |
| 4 | Hours reclaimed | (Admin hours before − after) × loaded hourly cost | Time log + payroll |
| 5 | Cost per resolved conversation | Total channel cost ÷ conversations resolved | AI invoice + call logs |
Number 1: Answered-Call Rate, and What a Missed Call Is Worth
Answered-call rate is answered calls divided by total inbound calls; anything that hits voicemail or dies in a queue counts as missed. It is the first number AI answering moves (typically to effectively 100%, since software takes every concurrent call), which is also why it is the number vendors most like to dress up with folklore.
You have seen the pitch stat: “62% of calls to small businesses go unanswered.” That number traces to a single 2016 study by 411 Locals, an SEO vendor. It is a decade old. Google’s own CallJoy research (2019) found the more conservative “nearly half” of calls to local businesses go unanswered. Neither one is your number. Your number is in your phone portal, and pulling it takes ten minutes.
What a missed call is worth is a formula, not a factoid: missed calls × booking-intent share × close rate × average ticket. A practice missing 125 calls a month, where 30% carry booking intent and a quarter of those would have closed at a $250 average visit, is leaking roughly $2,300 a month, not the $10,000 a rigged vendor calculator will show you. Honest math is still plenty motivating; it just survives due diligence.
And the phone channel is worth measuring because it is where high-stakes buyers still are: Invoca’s 2021 consumer research found roughly 70% of consumers call a business before a high-stakes purchase (67% in healthcare, 60% in home services), and 87% said a phone conversation made them more confident buying. The channel your customers use to judge you is the one most businesses measure least.
Number 2: Speed-to-Lead, the 5-Minute Rule, With Dates Attached
Speed-to-lead is the median minutes between an inquiry and your first substantive response. The canonical finding: Dr. James Oldroyd’s Lead Response Management study (MIT/InsideSales.com, 2007) found companies responding to a web lead within 5 minutes were 21 times more likely to qualify it than companies waiting 30 minutes, and roughly 100 times more likely to make contact at all.
Yes, 2007 (nearly two decades old), and any vendor citing it without the date is committing exactly the sloppiness this article exists to catch. But the direction has been replicated. Harvard Business Review’s 2011 audit of 2,241 U.S. companies found the average lead response time was 42 hours, 23% of companies never responded at all, and firms responding within an hour were about 7x more likely to qualify the lead. The magnitudes are dated; the mechanism, buyers shortlist whoever answers first, has not changed.
Measure it from CRM timestamps (lead created → first outbound activity), use the median so one weekend inquiry does not poison your average, and track business-hours and after-hours separately. AI’s contribution here is structural: an answering layer that picks up in seconds makes your 2 a.m. speed-to-lead equal your 2 p.m. number, which no staffing plan can do.
Number 3: Booked Appointments per 100 Inquiries
Booking rate is bookings divided by qualified inquiries, times 100. Use the per-100 framing instead of raw counts, because inquiry volume swings with seasonality and ad spend: a raw “we booked 20 more appointments” claim means nothing in a month your call volume grew 30%.
Two integrity rules. First, count kept appointments, not just booked ones: an AI that books fast but books badly shows up as a rising no-show rate, and no-shows are a cost, not a return. Track show rate right next to booking rate. Second, define “qualified inquiry” once, in writing, before launch: if spam and wrong numbers count as inquiries after the AI arrives but did not before, your booking rate will look worse while your business does better.
This is also the number that catches a failing deployment early. Answered-call rate hits 100% on day one by construction: that is what the software does. If bookings per 100 inquiries have not moved by week four, the AI is answering calls and then losing them. Pull ten call recordings and find out where.
Number 4: Hours Reclaimed From Admin Work
Hours reclaimed = (admin hours before − admin hours after) × loaded hourly cost. It is the softest of the five numbers, which is exactly why you value it at payroll cost and never let a vendor value it for you.
The opportunity is real. Asana’s 2023 Anatomy of Work index found knowledge workers spend about 58% of the workday on “work about work” (status updates, hunting for information, chasing approvals) rather than skilled work. McKinsey Global Institute’s automation research (2017) estimated about 30% of activities could be automated in roughly 60% of occupations with technology that existed then. In a small business, that looks like your front desk re-typing the same patient details into three systems.
The honesty rule: reclaimed hours only count toward ROI if they are redeployed into something that produces revenue or retention: recall lists, review requests, overdue treatment follow-ups, estimate chasing. If the saved time simply evaporates into a calmer afternoon, that may still be worth having, but strike it from the ROI math. In our own deployments we log where the hours went in the weekly report; if we cannot name the destination, we do not count the hours.
Number 5: Cost per Resolved Conversation
Cost per resolved conversation is your total monthly channel cost divided by conversations resolved end-to-end. It is the number that makes AI’s economics legible against a human benchmark, and the one where dishonest comparisons are easiest, so build both sides carefully.
The human side starts at the U.S. Bureau of Labor Statistics: the median receptionist wage was $17.90 an hour as of May 2024 (roughly $37,000 a year full-time before payroll taxes and benefits), with about 128,500 openings projected per year, a reminder that even at that wage, phone coverage is hard to hire and harder to retain. But a receptionist does far more than answer phones, so do not assign their whole cost to calls. If they spend four hours a day on the phones handling 40 calls, the honest phone-side math is 4 × $17.90 ÷ 40 ≈ $1.80 per handled call, before taxes, benefits, sick days, and the calls that ring while they are checking someone in.
On the AI side, divide your full monthly invoice (platform fee, usage, and any maintenance retainer) by the conversations the AI resolved without human rescue. Resolved is the load-bearing word: at a 60% containment rate you divide by 60% of conversations, not all of them, because the escalated 40% still consumed staff time. We publish the complete cost mechanics (per-minute voice components, chat licensing, build fees, and the hidden line items) in our breakdown of what AI automation actually costs, and you can see how we structure voice AI engagements specifically.
The punchline is not “fire the front desk.” In every deployment we run, the AI takes overflow and after-hours first and humans keep the conversations that need judgment. The comparison exists so you can price coverage (nights, weekends, six concurrent calls during the Monday rush) that you were never going to staff anyway.
How Long Until AI Pays Back? The Formulas and Two Worked Examples
Payback = setup cost ÷ net monthly gain, where net monthly gain = incremental revenue + redeployed-hours value − total monthly AI cost. The IDC study commissioned by Microsoft (2024) found companies typically realize generative-AI value within about 13 months; the deliberately conservative worked examples below land at 2 and 5 months, and the same math goes negative at low volume, which is the part vendor calculators skip.
Formula 1: Missed revenue (what is leaking today): missed calls × booking-intent share × close rate × average ticket.
Formula 2: Net monthly gain: incremental revenue + (redeployed hours × loaded hourly cost) − total monthly AI cost.
Formula 3: Payback period: setup cost ÷ net monthly gain. Under 12 months is a good project; under 6 is a great one; a negative net monthly gain means do not buy.
Worked example 1: a six-operatory dental practice. Every input is shown, every input is deliberately pessimistic, and time savings are excluded on purpose: if the revenue case does not clear on its own, do not let anyone pad it with soft hours.
| Input | Value | Basis |
|---|---|---|
| Inbound calls per month | 500 | Assumption: replace with your phone log |
| Current answered-call rate | 75% | Generous assumption; Google’s CallJoy research (2019) found nearly half of local-business calls go unanswered |
| Missed calls | 125 | Computed |
| Booking-intent share of missed calls | 30% | Conservative assumption: the rest are confirmations, vendors, spam |
| Close rate on recovered calls | 25% | Conservative assumption |
| New appointments booked | 9 / month | Computed (125 × 30% × 25%) |
| Show rate | 80% | Assumption: use your PMS number |
| Kept appointments | 7 / month | Computed |
| Average first-visit value | $250 | Conservative: our published month-one data averaged $290 per AI-booked appointment |
| Incremental revenue | $1,750 / month | Computed |
| AI receptionist cost | $750 / month | Top of our published $350-$750 band (Pro tier) |
| Net monthly gain | $1,000 | Computed: time savings deliberately excluded |
| Setup cost | $2,000 | Top of our published $1,500-$2,000 setup band; use your actual quote |
| Payback | 2 months | $2,000 ÷ $1,000 |
Now the fragility check, because this is what vendor calculators hide: halve the close rate to 12.5% and incremental revenue falls to roughly $900 against $750 of monthly cost. Net gain drops to about $150 a month and payback stretches past a year. The model lives or dies on whether recovered calls actually convert, which is why you audit call recordings in week two instead of admiring a dashboard. For calibration in the other direction: our published month-one data from a Miami dental practice handling 600 calls a month showed 93 AI-booked patients and $27,000 in attributed revenue. Conservative models are floors, not forecasts, but budget on the floor.
Worked example 2: a two-truck HVAC company. Same formulas, home-services inputs.
| Input | Value | Basis |
|---|---|---|
| Inbound calls per month | 150 | Assumption: check the carrier log; owner-answered lines undercount |
| Current answered-call rate | 70% | Assumption: owners on job sites miss midday calls |
| Missed calls | 45 | Computed |
| Booking-intent share of missed calls | 40% | Assumption; Invoca’s 2021 research found 60% of home-services consumers call before buying, so inbound calls skew toward real service requests |
| Close rate on recovered calls | 25% | Conservative assumption |
| Jobs booked | 4 / month | Computed (45 × 40% × 25%, rounded down) |
| Average ticket | $300 | In line with the $304 average in our published month-one HVAC data |
| Incremental revenue | $1,200 / month | Computed |
| AI cost | $750 / month | Same $350-$750 band, top of range |
| Net monthly gain | $450 | Computed: revenue only |
| Setup cost | $2,000 | Top of our published $1,500-$2,000 setup band; use your actual quote |
| Payback | ~5 months | $2,000 ÷ $450 |
Five months is a solid project rather than a spectacular one, and the short payback is mostly a function of a modest setup fee rather than a rich revenue case: net gain here is $450 a month, thin enough that two slow months of call volume erase it. Two things move it. First, time: our real HVAC deployment saved two hours of dispatch and intake time a day; valued at the BLS median receptionist wage of $17.90 an hour, that is roughly $790 a month, lifting net gain to about $1,240 and pulling payback under two months: we leave it out of the headline math because hours only count when they are redeployed, but in a dispatch-heavy business they usually are. Second, emergencies: in HVAC, a single captured after-hours emergency can be a $1,000+ job that covers more than a month of AI cost by itself. Voice-AI vendors report that roughly a third of calls arrive after hours; treat that as directional vendor telemetry rather than research, and then pull your own after-hours share from the call log, because that number decides this purchase more than any statistic in this article.
Which AI ROI Stats Should You Stop Trusting?
Stop trusting any statistic that arrives without a named primary source and a year, which, in AI-automation marketing, is most of them. The five below are the industry’s most recycled numbers, and none survives a request for its origin.
“85% of missed callers never call back.” No verifiable primary source exists: it circulates from vendor blog to vendor blog, mutating between 80% and 85% along the way. Some callers certainly do not call back; nobody has produced the study saying 85% of them. Versions of this one have appeared across the industry’s marketing, and our own industry’s citation hygiene, ours included, has not been spotless, which is part of why this article exists.
“Missed calls cost small businesses $126,000 a year.” Derived vendor math: someone’s assumptions about call volume, close rate, and ticket size multiplied together and laundered into a “finding.” Your leak is computable from your own numbers with Formula 1 above. It will usually be a fraction of six figures, and it will be true.
“62% of calls to small businesses go unanswered.” Real study, wrong decade: it originates with 411 Locals in 2016. Google’s CallJoy research (2019) found “nearly half.” Dated correctly, both are usable context; presented as a fresh 2026 finding, both are folklore.
“35-40% of calls come in after hours.” Sourced only from voice-AI vendors’ own platform telemetry, plausibly biased, since businesses that bought after-hours answering tend to be the ones with after-hours volume. Directionally fine; your phone log is the actual answer.
“Dental practices miss 20-35% of their calls.” Vendor blogs citing vendor blogs. Practice-reported ranges vary enormously with staffing and season. Measure yours.
The two-question pressure test for any vendor (including us): “What is the primary source? What year?” If the answer is another vendor’s blog post, discard the number and note what that tells you about the pitch. Then apply the same skepticism to the respectable stats. IDC’s $3.70-per-$1 figure comes from a study commissioned by Microsoft (a company with an obvious interest in the answer) and is built on self-reported returns. MIT NANDA’s 95% describes enterprise pilots, mostly generic tools bolted onto workflows, not a five-person office deploying one scoped system. Salesforce’s 91% “say it boosts revenue” is perception, not P&L. Treat every published statistic, every one in this article included, as a prior. Your baseline is the verdict.
When Is AI ROI Genuinely Negative?
AI ROI is genuinely negative in three common situations: not enough volume, a broken process upstream, and no internal owner. RAND’s 2024 finding bears repeating: AI projects fail at twice the rate of ordinary IT projects, for organizational reasons. So treat “do not buy yet” as a live option rather than a rhetorical one.
Not enough volume. Run Formula 1 before anything else. As a rule of thumb from our own deal reviews, below roughly 100 inbound calls or 50 web leads a month the arithmetic gets thin: a practice with 100 monthly calls missing a quarter of them recovers maybe $470 a month at the dental inputs above, less than the AI costs. We decline these builds and say why; the right move at that volume is usually fixing hours coverage or call routing first, then revisiting AI when volume grows.
A broken process upstream. Automation multiplies throughput, not quality. If your offer does not convert or your intake asks the wrong questions, an AI that answers every call just produces disappointed callers faster. This is where RAND’s root causes (misaligned objectives, chasing technology over business outcomes) live in small-business form. Fix the close rate before you buy the call capture.
No internal owner. The weekly scorecard needs one named person (owner, office manager, ops lead) who reads the five numbers and listens to a handful of calls. The deployments that drift into MIT’s 95% are not, in our experience, the ones with worse technology; they are the ones where nobody looked at week six’s numbers. If you cannot name the owner, you are not ready to buy, from us or from anyone.
How to Run This Yourself, and When to Call Us
The whole system, compressed: pull a 30-day baseline from your phone system and CRM using the table above; compute your missed-revenue leak with Formula 1; if the leak clears roughly twice a realistic monthly AI cost, pilot one workflow (not three); track the five numbers weekly against the baseline; run the payback formula at day 90 and keep, fix, or kill accordingly.
If you want the baseline done for you, that is literally what our free audit is: we pull your answered-call rate, speed-to-lead, and booking numbers and hand you the same spreadsheet this article describes, whether or not you buy anything from us. For businesses that do deploy, Benian OS puts the five-number scorecard on a live dashboard so the weekly review takes ten minutes instead of an afternoon. Either way, measure first. The $3.70 businesses and the 95% businesses bought similar technology. They did not run similar spreadsheets.
