How to Evaluate an AI Receptionist on a Free Trial · Willison Skip to main content
AI receptionist buyer's guide · 12 min read

How should you evaluate an AI receptionist during a free trial?

Seth Willison ·

You signed up for an AI receptionist trial on Monday. By Wednesday it's live on your line, and by next week you have to decide whether to keep paying for it. That's about seven days to find out whether it can be trusted with your calls, while you're running jobs.

Here's the short version. Judge the trial on the line your customers actually dial, not on the vendor's demo, and write down what counts as a fail before the first call comes in. On the first day it's live, make a handful of test calls aimed at where these systems slip: an address said fast, a second problem after the booking, a job you don't do, a caller who goes quiet, an after-hours emergency. Then hold every real call up against your calendar and your old call log, and score jobs booked right and jobs booked wrong, not calls answered.

Put the cancel date in your calendar on day one, too. Ours takes payment details up front and becomes a paid subscription on its own if you let it run out, so ask any vendor whether theirs does the same.

A word on where we stand: Willison sells an AI receptionist and runs a free trial, which makes us an interested party here. The tests below are written to work on anybody's trial, ours included, and one of the possible answers at the end is to keep what you've got.

The short version, in order

  1. 1. Write your fail list and your cancel date before it goes live.
  2. 2. Pull a baseline from your carrier log. A few weeks of calls, answered and missed.
  3. 3. Run it on your real number, forwarded the way you'd actually keep it.
  4. 4. Make the test calls on the first live day, so there's time to fix a setting and test again.
  5. 5. Read every real call against the calendar and the baseline.
  6. 6. Decide two days early, before the clock decides for you.

Before day one: write the fail list and find the end date

Write the fail list before you've heard a single call. Once a trial's running, every mistake is easy to explain away and every good call feels like proof.

Split it in two. Hard fails are things the product does wrong on its own: telling a caller they're booked when nothing's on the calendar, booking the wrong address after reading it back, quoting a price you never gave it, losing a message, transferring a call outside the hours you set. Settings are things it does wrong because it was told wrong: a job length that doesn't match the work, a missing service, an old phone number. Settings get one fix and one retest. A setting that fails its retest goes on the hard-fail side.

Then find out when the clock starts, signup or go-live, and what happens on the last day. Here are ours, in full. Willison's free trial runs nine days from checkout: the 48-hour setup, then about seven days live. We take payment details at checkout, and the subscription starts automatically when the nine days end unless you cancel. Cancel any time inside those nine days and you aren't charged. It's one trial per business. So set your decision reminder for day seven, not day nine.

Last, pull a baseline: the call log for the two to four weeks before the trial, from your carrier or phone system. Total calls, how many got answered, how many were over in a few seconds, and when they came in. Without it, you'll judge the trial week against your memory, and memory's a bad witness.

Test it on your real line, set up the way you'd run it

A demo line proves the engine. It doesn't prove your setup: your services, your hours, your job lengths, what you charge to show up, and who gets reached when a call can't wait. So point the number customers already dial at the trial, with the forwarding you'd actually keep. If you plan to answer during the day and only send the misses and the nights, run the trial that way, even though there'll be fewer calls to judge.

One check before anything else. Let a call ring out with nobody in the office picking up, and see where it lands. If your old voicemail still grabs it first, the trial is testing nothing. The forwarding options, and what forwarding doesn't move, are covered step by step in what it actually takes to switch your phones over to an AI receptionist.

Then make sure bookings land on the board your office actually works from. A receptionist that books perfectly into a calendar nobody checks has booked nothing. If your techs get their day from dispatch software or a whiteboard, ask whether it writes there, or who re-keys it and how fast.

Seven test calls to make on the first live day

Don't make these yourself if you can help it. You know the script, so you'll call too clean. Hand them to a spouse, a friend or an older relative, calling from their own phones in a truck cab or a noisy kitchen. Give each one the same made-up last name, tell your office and techs that name means a test, and delete those bookings the same day so nobody rolls on one. Warn the testers they may get a confirmation text.

If your office answers during the day, give them the testers' numbers and ask them to let those calls ring through, or make the daytime calls after you close. Otherwise your own front desk catches the test.

1. Say the address fast. Use the caller's own home address, unit number and all, rattled off the way people do on a phone. A pass: it reads the address back and lets the caller fix it before anything gets booked. A fail: it books what it thinks it heard.

2. After it confirms, bring up a second problem. "Oh, and the thermostat upstairs has been acting up too." A pass: the second problem is captured, either added to the visit or passed to you, with no duplicate visit booked for one trip. A fail: it disappears, or two trucks end up booked for one house.

3. Ask for a job you don't do. Pick something next door to your trade, like a water heater at an HVAC-only shop. A pass: it doesn't book it. A fail: a truck on the board for work your crew won't do.

4. Go quiet in the middle of the call. "Hang on, let me go look at the unit," then nothing for a while. A pass: it prompts the caller more than once before giving up. A fail: it hangs up on someone who's standing in the basement.

5. Ask if it's a real person. A pass: a straight answer. A fail: it claims to be a person, or dodges.

6. Ask what it'll cost. A pass: it tells the caller what you told it to, such as a paid diagnostic visit for a repair or a free estimate on a new system, and goes no further. A fail: it makes up a number.

7. Call after hours with an urgent job. Keep it ordinary, like no heat or no cooling, never a gas smell or an alarm going off. Set the night and the time with whoever holds the after-hours number, make the call early in the evening, and have them let one transfer ring out on purpose. A pass: exactly what you set up happens. A text reaches you with enough to call back, or a transfer rings the number you set, only inside the hours you chose, and the caller isn't stranded when nobody picks up.

When the seven calls are done, open the calendar. Every caller told they were booked should be there, at the right time, for the right job. A caller told "you're booked" with nothing on the board is the most expensive fail on this list, because that customer waits at home for a truck that was never sent.

Here's what Willison is built to do on each, so you know what to hold us to. It reads the address back, and confirms a booking only after it's written, with a fixed line the AI can't reword. It books one job per call, so a late second problem becomes a message to you. It books only the services on your list and asks the caller to confirm rather than guess. It nudges a quiet caller more than once. Asked, it says it's a virtual receptionist. It states your fee setup per service and never prices a job.

Escalation is set per shop: a text so you can call back, or a live transfer to the escalation number you choose, inside a window you set. If the check that decides whether a transfer is allowed fails, the call isn't transferred, rather than ringing a sleeping person. If an emergency transfer doesn't connect, it falls back to booking the job and texting you.

It has its own weak spots, and your test calls are how you find them. It can mishear an address, file a job under the wrong type, get interrupted and restate something wrong, or take a message that doesn't reach you. The read-back, the one-question-at-a-time pace and the owner alert are there because those things can happen. And some of what it does is a trade-off you might not want: a caller with two separate jobs gets one booking and a message, and a caller who wants a price on the phone doesn't get one.

Read the calls, not the dashboard

If all you're shown at the end of the week is a count of calls answered, that's the wrong score. In Invoca's 2026 home services benchmarks, 38% of calls answered by a person were leads, and 45% of those leads converted on the call. Most answered calls weren't new-job leads, so a count of calls answered says little about what got booked.

Count your own week instead. Go through the recordings or transcripts and put every call in one of four columns:

  • Booked right. Right job, right address, right slot, right length.
  • Handed off right. It couldn't or shouldn't book, and the details reached you complete.
  • Wrong. Wrong address, wrong job type, wrong slot, or a message that never got to you.
  • Not a job. Wrong numbers, sales calls, robocalls.

Keep quick hangups out of all four, and count them against your baseline instead. The same Invoca report puts the answer rate at 65% for calls over 15 seconds and 73% for calls over 30 seconds, a cut it uses to strip out misdials and quick hangups, so every line gets some misdials. If the trial week's share of calls over in a few seconds runs well above your old log's, count the difference against the trial: those are likely callers who hung up on the greeting.

Willison records and transcribes every call, and the greeting tells every caller so. Whatever you trial, ask any vendor, us included, whether you'll get recordings or transcripts during the trial. If the answer's no, you're grading on the vendor's summary of its own week, and that's a finding in itself.

Then call three to five real callers back and ask one question: how did the call go? It's the only customer reaction you'll get before you decide.

What a week can't tell you

A week won't show you a heat wave or the morning after a storm unless one lands inside it, or anything slower than a week, like a quote that needs a nudge or a month-end report. On a quiet line, your test calls are most of your evidence. That's enough to judge how calls get handled, not what coverage is worth, which is worked through in whether round-the-clock coverage is worth paying for at two night calls a week.

When the answer is to keep what you have

A clean week doesn't mean you should buy. Put the trial next to your baseline. If your own log shows few missed calls, and the answered ones already get booked, the trial is solving a problem you don't have. If a person or an answering service already handles your misses and books them correctly, a clean trial only proves something else can do it too. And if the trial booked nothing your old setup wouldn't have caught, that's your answer.

Each of those is a legitimate reason to cancel inside the trial, even when every test call passed.

When a free trial is the wrong test

If you're really buying coverage for your busy season and the trial would land in a dead month, a quiet week tells you very little, and nobody wants to change how the phone gets answered mid heat wave. The better test may be a paid month started two or three weeks before the rush, sending only the misses and the after-hours calls, so you're back on your old setup before the busy weeks if it fails. It costs a month's fee, and you should count that as the price of a real answer.

Our own trial has limits worth naming. Setup takes two of the nine days. Payment details go in at checkout, and the plan starts on its own if you don't cancel. And there's one trial per business, so if your busy season is six weeks out, it may be worth waiting until it's closer before you start.

Ranked: ways to judge an AI receptionist before you commit, best first

From the most telling to the least:

  1. 1. Your own line, just ahead of the season you're buying for.
  2. 2. A trial on your own line in an ordinary week, plus the seven test calls.
  3. 3. An owner already running it. Worth a call if you can get one, but it's their setup and their callers, not yours.
  4. 4. Your own test calls into the vendor's demo line. The engine, not your setup.
  5. 5. A recorded demo call. Good for the voice and the pace, and it's a call the vendor chose.
  6. 6. Answers on a sales call. It tells you what the vendor says, not what it does.

Day seven: hard fails versus settings

Go back to the list you wrote before day one, and sort the week's problems against it.

A setting problem isn't a reason to walk the first time. Wrong job lengths, a missing service or an outdated phone number would be just as wrong in a person's notes. Fix it, run that test call again, and judge the result. If the vendor entered the setting wrong during setup, note that too: it's a preview of how the next change gets handled.

A hard fail is different. One caller told they were booked with nothing on the board, one address booked wrong after a read-back, one invented price, one lost message or one call transferred outside your hours is enough to stop. If a product does that once in a test week, you've got no reason to think it won't do it again with a real customer.

If the week comes back clean and the baseline says you need it, you've got your own calls, handled the way you'd want. If not, cancel inside the trial, and you're out a week and some setup time.

Before you start a trial with Willison, you can press play on a real call on willisonhq.com, a call to our HVAC receptionist from start to finish. It's a sample HVAC shop, so if you're in another trade, listen for how it handles the call rather than the job.

Frequently asked questions

Should I tell the AI receptionist it's a test call?

No. A caller who announces a test speaks slower and clearer than a real one, which is the opposite of what you want to learn. Warn your own people instead: agree on a made-up last name that marks a test, and delete those bookings the same day.

Can I trial two AI receptionists at the same time?

Not on the same calls. Forwarding sends a line's calls to one place, so two products on one line means splitting the week or the hours, and then each one gets different callers. Running them back to back with the same seven test calls is the fairer comparison.

Do I need a credit card for an AI receptionist free trial?

It depends on the vendor, so ask before you sign up. Willison takes payment details at checkout and starts the subscription automatically when the nine-day trial ends, unless you cancel first, and a cancellation inside the trial is never charged.

What if something goes wrong because my setup information was wrong?

That's a setting, not a verdict on the product, the first time. Correct the hours, services, job lengths or contact number, then repeat the test call that caught it. If it fails again with the right information, treat it as a hard fail.

Want to know if Willison is the right fit for your business?

15 minutes. Tell us how your phone works today, how many calls slip past when you cannot pick up, and what a booked job is worth to you. You leave with a straight yes or no on whether Willison is the right fit, and what it would look like set up for your business.

No pitch, no follow-up unless you want one. Your plan is month-to-month by default: cancel anytime if it's not working for you, no penalty. We work with you to dial the receptionist in for your business.

Written by

Seth Willison

Founder, Willison. Willison builds AI receptionists for trades and restoration companies, so the calls that pay don't get missed.

Free 15-min call See your missed revenue
Book