What we learned auditing 886 conversations of a WhatsApp agent for medical practices

We audited 886 WhatsApp chats for medical practices: offering a specific day and time booked 34%; quoting only the price, 0.8%. Lessons and fixes.

Carlos Betancur Gálvez

By Carlos Betancur Gálvez

AI, digital marketing and medical marketing consultant · btodigital

When the person handling bookings offered the patient a specific day and time, 64 of 189 conversations ended in an appointment: 34%. When they only quoted the price, 2 of 264 did: 0.8%. That was the strongest finding of our audit of 886 WhatsApp conversations for medical practices in Colombia, and it has nothing to do with artificial intelligence.

The conversations come from El mejor DOC, the medical directory I founded, and the agent answering them is Atendio, btodigital’s WhatsApp product. In other words, this is an audit of my own product. That is also why I am publishing what we learned and what we fixed.

If you are a specialist physician running, or considering, a bot on your practice’s WhatsApp, this article gives you three things: where patients are really lost, a reply script you can copy, and five minimum rules any healthcare bot should meet.

What did we audit, and how?

Every conversation with activity between 1 September and 2 October 2026: 886 conversations and 16,562 messages. Patients wrote 7,071 of those messages, the bot 4,232 and the human booking team 5,259.

We ran three passes, because a single AI read-through is not enough to claim anything:

  1. Measurements in code against the database: who wrote, when, how long the reply took, how it ended.
  2. Seven AI auditors read every conversation and logged failures. Then nine independent verifiers re-read samples to confirm or reject each finding and to check whether it still applied after the fixes of 26 September.
  3. A final verification in code: every figure in this article was recounted against the data.

Nothing below names doctors, practices or patients, and there is no clinical data. Working files containing personal data were deleted when the audit ended.

How many patients who asked for an appointment got one?

Of the 886 conversations, 736 asked for an appointment or consultation. About 69 ended in a booking: 56 recorded as such and roughly 13 confirmed in the chat but never recorded. That is 9.4%.

On its own, that number says little. It gets interesting when you split conversations by what the human replied:

What the person repliedConversationsBookedRate
Offered a specific day and time1896434%
Quoted the price, no time offered26420.8%
Neither price nor time18310.5%
No human ever stepped in12700%

A word of caution. This is an association, not an experiment: some patients who were offered a time may already have been more decided. But that alone does not explain 34% versus 0.8%, and it is the only variable in the table the practice fully controls.

In practice: the message with the price followed by “would you like to book?” almost never books. The message with “I have Tuesday at 3 pm or Thursday at 9 am, which works for you?” books one time in three.

Does it matter when patients write and how fast we answer?

A lot. By the day of the first message:

First messageBooking rate
Monday to Friday, business hours10.9%
Saturday6%
Sunday4.3%

The median first human reply was 20 minutes at the fastest practice and 8 hours at the slowest. We also found 80 conversations with clear booking intent, during business hours, where no human ever replied.

The bot answers within seconds, at any hour. But in this service the bot does not book on its own for most practices: it collects details and hands the conversation to a person. If that person takes eight hours, the patient has already booked elsewhere.

What we learned

The audit left these rules, some for the bot and some for the booking team:

  • A greeting is not a goodbye. Auto-close read phrases like “Good evening, I wanted to know the price” or “thanks, I’ll think about it” as an ending: of 55 automatic closures reviewed, 22 came right after a greeting. A polite phrase is not enough to close a conversation.
  • Escalate on intent, not on a name. The “needs attention” flag depended on the patient giving their name, so “I want an appointment on Thursday” without an introduction was not flagged. That explains many of the 80 conversations with no human.
  • The bot only states what is in its sources. The audit found replies with slots, dates and insurance agreements that had not been confirmed. If a detail is not confirmed, the bot says the team will confirm it.
  • One voice per conversation. On 481 occasions the bot and a person answered the same patient message, the bot first and the person about two minutes later, and the bot read the team’s messages as its own. When a person steps in, the bot steps aside.
  • Fewer automated messages, and no health data in them. Follow-ups and recovery messages stacked up: 71 conversations received four to six automated messages in a row with no reply from the patient. And an automated message should not repeat the reason for the visit: in healthcare, that can be a diagnosis.
  • The life-risk check is reviewed regularly. Since 26 September the bot detects messages suggesting suicidal ideation or self-harm, points to the crisis line and alerts the team urgently. The audit found expressions it still missed, along the lines of “I can’t see the point” or “I can’t take it anymore”. That list is never finished.
  • The most used reply was not the one that booked. The price-without-time template appeared in 71% of conversations with human involvement, and it is exactly the one that converts at 0.8%. Templates get reviewed against results, not habit.
  • Payment details go in the same message. When the deposit template ended with “if you’d like to go ahead, we’ll send you the details”, without the account, 0 of 25 moved forward; when it included them, 36 of 53 did. Every extra step costs bookings.
  • Mass re-contact did not revive conversations. A batch of messages sent on 28 September to pick them back up ended 0 for 8.

What did we fix on 3 October?

  • Auto-close no longer triggers on greetings or “I’ll think about it”.
  • The team alert now fires on intent (asking for an appointment, asking about a day), even without a name.
  • The bot pauses for 120 minutes after a human steps in (it used to be 30), so it stops talking over people.
  • The bot now tells the team’s messages apart from its own.
  • Explicit rules against inventing slots, dates, insurance agreements or services.
  • Recovery messages no longer repeat the reason for the visit, send at most one template, and stop if the team wrote in the last 48 hours or the patient said they would think about it.
  • Nine new expressions in the life-risk check.

The code changes passed 464 automated tests before release. What I do not have yet is a before-and-after in bookings: with only a few days of data it cannot be measured, and I would rather say so than invent an improvement.

How can you apply this in your practice?

A reply script you can copy. When a patient asks about price or availability, put everything in one message:

  1. The price, plainly.
  2. Two specific options, with day and time.
  3. The deposit payment details, if you charge one, in the same message.
  4. A closed question: “Which of the two works for you?”

If you do not have slots at hand to offer, that is the problem to solve first, before any bot.

Five rules for a healthcare bot. If your vendor cannot show you each one working, do not let it talk to your patients:

  1. It never speaks for the doctor. It never says the doctor reviewed, prescribed or decided anything. If the patient asks for clinical judgment, it hands over to a person.
  2. It never invents slots, prices or insurance agreements. It only states what is in its sources, and says so when it does not know.
  3. It escalates on intent, not on a name. Anyone asking for an appointment needs a person, whether or not they introduced themselves.
  4. It never closes a conversation on a greeting. A “good evening” is almost always the start, not the end.
  5. It has a life-risk check that runs on every message, at any hour, and alerts a person.

What to measure every week:

MetricWhy
Booking rate over conversations with intentIt is the real outcome, not messages answered
Share of replies with a specific day and timeThe strongest lever we found
Median first human reply, per practiceEight hours is a lost appointment
Conversations with intent and no human at allShould be zero during business hours
Automatic closures and life-risk alertsTo catch bot failures before they cost you

If you are evaluating an AI assistant for your practice, I wrote earlier about what an AI chatbot should do in a medical office and about how to get more appointments for a medical practice. If you want me to look at your case, this is how I approach medical marketing.

Frequently asked questions

Is a WhatsApp bot useful for a medical practice? Yes, to answer within seconds at any hour, collect details and alert the team. But in this audit the biggest difference in bookings came from what the human replied, not from the bot: offering a specific day and time booked 34%, quoting only the price booked 0.8%.

What should I check first on my practice’s WhatsApp? How many of your team’s replies offer a specific day and time, and how long the first human reply takes. You can get both numbers by counting one week of conversations by hand.

Can the bot quote prices and book on its own? It can, if it has confirmed information and real access to the calendar. Without that it is risky: in this audit we found replies with unconfirmed slots and insurance agreements. Better that it says the team will confirm and escalates right away.

How do I stop the bot from speaking for the doctor? With an explicit rule in its instructions and, on top of that, an automatic check that reviews every reply before it is sent and blocks it if it attributes a diagnosis, instruction or decision to the doctor. Instructions alone are not enough.

What if a patient writes something that suggests a risk to their life? The bot should detect it on every message, at any hour, point to crisis lines and alert a person urgently. Review the expressions it detects regularly: in this audit we added nine new ones.

How long until you see results after fixing a bot? I do not know yet for these fixes, and I will not guess. With a few days of data any change in booking rate may be noise. The right move is to measure the same figures as this audit a few weeks later and compare.

Share
Related posts