Where AI belongs in a business that answers the phone
An honest map for a fifteen-person service company: what to deploy, what to leave alone, and the line where a human has to take over.
In a fifteen-person service company there are four places this technology earns its keep and a much longer list of places it will cost you money and credibility. The dividing line is not how clever the task is. It is whether being wrong is recoverable, and in a business that sends people to fix physical equipment, a great deal of it is not.
The rule the evidence gives you
Start with what has actually been measured rather than what is being sold. Across 5,179 customer support agents the average productivity gain from a generative assistant was 14 percent, concentrated as 34 percent among novices and close to nothing among experienced staff. In a separate controlled trial, expert developers working on codebases they knew well were slower with AI than without it.
That gives one clean operating rule. Point this at work where the person doing it is not yet expert, or where the task is drafting and retrieval rather than judgement. Do not point it at your most experienced people doing the thing they are experienced at.
The four that work
Each of these has the same shape: high volume, low variety, and a wrong answer that surfaces immediately and cheaply.
Capturing what came in
Taking a missed call, a voicemail, a form or an email and turning it into a structured record with a name, a number, a location and a description. The highest-value application by a distance, because the alternative is not a slower process, it is no record at all.
Reading the mess
A photograph of a handwritten job sheet, a supplier PDF, a warranty document. Extracting the fields so somebody does not retype them. This is genuinely new: the reason previous attempts at this failed was that they required clean input first.
Drafting the routine reply
The appointment confirmation, the follow-up on an unanswered quote, the response to a review, the reminder that a service is due. A human sends it. The blank page was the bottleneck, not the sending.
Finding what the business already knows
What did we install at this address, when, what did we charge, what went wrong last time. That answer exists somewhere in the records and currently requires the one person who remembers. Retrieval is the least glamorous application and the one that most reduces key-person dependency.
Notice that not one of them makes a decision. Every one produces something a person then uses, checks or sends. That is not timidity, it is where the measured gains actually were: throughput on drafting and retrieval, not autonomy.
The fourth deserves particular attention because it is undersold. In most of these companies the answer to "what did we do at this site in 2019" lives in one person's head plus a filing cabinet. Making that searchable does not save dramatic time on any single query. It removes a dependency, and dependencies are what get priced into the business.
Where the line is
Two lists are only useful for the ten tasks on them, and next quarter somebody will propose an eleventh. So it is worth extracting the test that put each one on its side, because the test generalises and the lists do not. Three questions, and the answers are not independent:
| The task | A wrong answer shows up | Who pays for it | Undoable before anyone acts |
|---|---|---|---|
| Capturing what came in | Immediately, the record is visibly wrong | You | Yes |
| Reading a job sheet | On review, in seconds | You | Yes |
| Drafting a routine reply | Before it is sent | You | Yes |
| Retrieving what you already know | Immediately, against the record | You | Yes |
| Quoting a price | Weeks later, at the invoice | The customer | No, it was said |
| Technical advice on the phone | Possibly never. Possibly at an incident | The customer | No |
| Committing to a time | At the slot you missed | The customer | No |
| Deciding what counts as an emergency | At the flooded plant room | The customer | No |
The columns agree perfectly, which is the point. Every task on the safe list is immediate, yours and reversible. Every task on the forbidden list is delayed, the customer's and final. Nothing sits in between, because the three properties are the same property looked at from three angles: whether the mistake reaches somebody who will act on it before you have a chance to catch it.
So the rule for the eleventh task, whenever it arrives, is that it needs all three. Two out of three is a no, and the tasks that get two out of three are exactly the ones that will be argued for hardest, because they are the ones that look like the safe list and behave like the other one.
Stated as a list, the forbidden side reads like this, and the temptation to cross it always arrives dressed as efficiency.
- Never let it quote a price. Pricing in this trade depends on site conditions, access, history with the customer and what the job actually turns out to be
- Never let it give technical advice to a customer, especially on anything safety-related, and especially not on the phone where nothing is reviewed
- Never let it commit to a time. A promised slot the dispatch cannot honour is worse than no promise
- Never let it be the only thing that answered. A customer who realises they spent four minutes with a machine that could not help has had a worse experience than voicemail
- Never let it write anything a regulator, insurer or court might read: certificates, inspection records, incident reports
- Never let it decide whether a job is an emergency, because that judgement is the difference between a callout tonight and a plant room flooding
The first is the one businesses cross most often and regret fastest. A quoted price is a commitment, customers treat it as one, and honouring a wrong one is expensive while withdrawing it costs the relationship. Somebody who understands the site sets prices.
Every item on that list has the same property: the mistake is discovered by the customer, after you have already promised something.
The fourth is subtler and is where the current wave of voice products will disappoint. Answering a call with something that sounds capable and then cannot help is not neutral, it is worse than an honest voicemail, because it consumed the caller's time and signalled that the company is not there. Handling triage and handover well is a much lower bar than handling the conversation, and it is the bar worth aiming at.
The order to do it in
Sequencing matters more than tool choice, and the correct first step is the least interesting one.
Capture comes first, always, and it is not really an AI project. Make every inbound contact produce a durable record with an owner and a status. That single change is worth more than everything below it, and until it exists nothing else has anywhere to write to.
Second, automatic acknowledgement: something goes back to the caller within seconds saying when a person will respond. Third, extraction, so the paper coming back from vans stops being retyped. Fourth, retrieval over the history. Only after all four would I look at anything that holds a conversation, and then only for triage and booking.
Why that order and not the exciting one
Capture and acknowledgement are cheap, reliable and reversible. Conversation handling is expensive, failure-prone and visible to the customer. Doing the second first means the most fragile component is also the first thing anyone experiences of your company, and the first bad story sets the internal politics for a year.
There is also a practical reason to go in this order. The early steps produce the data the later ones need. A retrieval system over records that were never captured properly returns nothing useful, and the failure gets blamed on the technology rather than on the sequencing.
What this does not fix
Being honest about the ceiling is what makes the rest credible.
None of this puts a technician in a van. If the constraint on the business is that there are not enough qualified people, no amount of this changes the number of jobs that can be completed in a week, and the honest gain is that the people you have spend more of their week on billable work rather than on paperwork. That is real and it is bounded by how much of the week was paperwork in the first place.
It also does not fix a business with a demand problem, a pricing problem, or customers who leave. Those are different layers and they need different work. Installing capture in a business whose real issue is that it charges too little produces a well-documented record of underpriced jobs.
Which is the same conclusion the evidence keeps pointing at. The technology is now genuinely capable of the specific jobs above, and the constraint has moved to whether somebody redesigns the workflow and holds it. That is the argument in the product cannot be disrupted, the operations can, and the reason to distrust any percentage quoted at you is in the automation numbers people quote are not real. The whole map, layer by layer, is on the system page.
The short version
- The test that generalises past the ten examples: does a wrong answer surface immediately, do you bear the cost rather than the customer, and can it be undone before anyone acts. Needs all three. Two out of three is a no, and those are the ones argued for hardest.
- Point it at people early on the learning curve and at drafting and retrieval work. Do not point it at your most experienced technician doing what they are experienced at.
- Four applications earn their keep: capturing inbound contact, extracting fields from messy documents, drafting routine replies, and retrieving what the business already knows.
- None of the four makes a decision. Every one produces something a person checks or sends.
- Hard limits: no quoting, no technical advice, no committing to a time, no being the only thing that answered, nothing a regulator might read, and no deciding what counts as an emergency.
- Order: capture, acknowledgement, extraction, retrieval, and only then anything conversational, and then only for triage.
Questions I get on this
Should a small service business use AI to answer calls?
What should AI never do in a trade business?
What is the highest-value place to start?
Productivity figures from Brynjolfsson, Li and Raymond, "Generative AI at Work", Quarterly Journal of Economics (5,179 customer support agents), and METR's July 2025 randomised trial on experienced open-source developers.
