AI coaching works, but in a narrower band than the category markets itself in. The strongest evidence available is a randomised controlled trial that found a statistically significant improvement in goal attainment and non-significant results on wellbeing, resilience, and stress, which is a real finding rather than a null one, and it points at the specific job AI coaching is good at: structured, goal-directed, repeatable support at a volume no human team can staff. That is the job Vocaliv’s AI Coach is built for, and it is worth being clear about what falls outside it.
Key Takeaways
- The founding evidence for AI coaching is Terblanche et al’s randomised controlled trial of a chatbot coach called Vici, with 75 in the experimental group and 94 controls across six months of use.
- That trial found a statistically significant increase in goal attainment, and non-significant results on resilience, psychological wellbeing, and perceived stress.
- The honest reading is that AI coaching is effective in a narrow application, which is exactly what the researchers concluded, not that it matches human coaching across the board.
- A later head-to-head trial compared accredited human coaches against automated AI coaches with 114 coachees inside a global organisation, so the evidence base is growing but still small.
- For corporate training, the relevant claim is not “as good as a human coach” but “handles the repetitive support load that consumes instructor capacity”, which is a measurable operational outcome rather than a psychological one.
🖥️ Sign In to Access Your Dashboard
What the Research Actually Found
Most vendor claims in this category trace back to one study, so it is worth reading it properly.
In 2022, a team led by Nicky Terblanche published a randomised controlled trial of an AI coaching chatbot called Vici. It was designed as a replication of an earlier human-coach trial. An experimental group of 75 used the chatbot for six months, measured against a control group of 94, with eight measurement points on goal attainment, resilience, psychological wellbeing, and perceived stress.
The result: goal attainment improved significantly. Everything else did not.
| Measure | Result |
| Goal attainment | Statistically significant improvement |
| Resilience | Non-significant |
| Psychological wellbeing | Non-significant |
| Perceived stress | Non-significant |
The researchers’ own conclusion was that AI coaching is effective in a narrow application, and that it could democratise coaching in a cost-effective, scalable way. That is a genuinely positive finding. It is also considerably more specific than “AI coaching works”.
A follow-up trial has since compared accredited human coaches against automated AI coaches head to head, with 114 coachees inside a global organisation, measured across goals, motivation, resilience, and wellbeing using validated psychometrics. The evidence base is expanding, but it remains small enough that anyone quoting a definitive verdict is overreaching.
Checked September 2026. Note that “AI coaching” in this research means life and organisational coaching, which is related to but not identical to AI support inside a training programme.
Where AI Coaching Reliably Works
Reading across the evidence and our own operational data, the pattern is consistent. AI coaching performs where the task is structured and repeatable.
- Goal setting and progress check-ins: The one measure that moved in the trial. Structured, repeatable, and the AI never forgets to follow up.
- Answering questions that have been answered before: In a training context this is the largest category by volume and the least dependent on judgment.
- Availability outside working hours: A learner stuck at 11pm on a Thursday gets an answer instead of dropping the module.
- Consistency across cohorts: The tenth cohort gets the same quality of support as the first, which is rarely true of human delivery.
- Flagging who is struggling: Detecting confusion at scale is a pattern-matching problem, which is what these systems are actually good at.
Where It Does Not Work
This is the part most vendor content omits, and it is the part worth being honest about.
Anything depending on the relationship: Coaching research consistently identifies the working alliance between coach and client as a driver of effectiveness. Whether that transfers to AI is an open research question, not a settled one.
Wellbeing, resilience, and stress: The trial found no significant effect on any of these. If your reason for buying is employee wellbeing, the evidence does not currently support the purchase.
Novel or ambiguous problems: A learner question that has never been asked, or that depends on your organisation’s unwritten context, is where an AI coach should escalate rather than answer.
High-stakes judgment: Career decisions, performance conversations, and anything with a compliance consequence need a human accountable for the outcome.
Motivation that comes from being seen: Part of why cohort programmes outperform self-paced ones is that a human notices when you are absent. An AI noticing is not obviously the same thing, and nobody has demonstrated that it is.
📄 Generate a Free PDF Sample Course in Your Cloned Voice
The Question Corporate Buyers Should Actually Ask
Most training providers evaluating AI coaching are not trying to replicate an executive coaching engagement. They are trying to solve a capacity problem.
A cohort of 40 learners generates roughly 200 questions a week, and most of them repeat questions already answered in a previous cohort. Those questions consume instructor hours regardless of whether they require instructor expertise. There is a hard ceiling on how many trainees one trainer can support, and repetitive support is what sets it.
That reframes the evaluation. The question is not whether an AI coach is as good as a human coach at coaching. It is whether it can absorb the share of support volume that does not need a human, and whether the effect on completion is measurable.
For a provider running two 40-learner cohorts:
| Metric | Before | After |
| Learner questions per cohort per week | ~200 | ~200 |
| Handled without instructor | 0% | 70%+ |
| Instructor support hours per week | 20 | 6 |
| Learner confusion rate | Unmeasured | Under 15%, tracked |
| Completion, 12-week programme | 45% | 60%+ |
These are operational outcomes, and they are the ones you should ask any vendor to evidence on your own content during a trial. A vendor who cannot show you the confusion data is asking you to take the completion claim on faith.
How to Evaluate a Claim in This Category
- Ask what the claim is measuring: Goal attainment, completion, and wellbeing are different outcomes with very different evidence behind them.
- Ask for the sample: A pilot with eight learners is an anecdote.
- Ask what happens on escalation: A system with no escalation path will confidently answer things it should not.
- Test on your own content: Generic demos test the demo, not the product.
- Test in your delivery language: An English-only evaluation tells you nothing about Arabic learner support.
Frequently Asked Questions
Yes, in a narrow application. The strongest available evidence, a randomised controlled trial of an AI coaching chatbot with 75 users over six months, found a statistically significant improvement in goal attainment but non-significant results on resilience, psychological wellbeing, and perceived stress. The researchers concluded AI coaching is effective for structured, goal-directed support and can make coaching more accessible at lower cost.
Not across the board, and the evidence needed to claim otherwise does not exist yet. AI coaching matches human coaching best on structured goal attainment and availability. It has no demonstrated effect on wellbeing or resilience, and the relational element that coaching research identifies as central remains an open question.
No demonstrated effect on wellbeing, resilience, or stress. Weak performance on novel or ambiguous problems that depend on unwritten organisational context. Unsuitable for high-stakes judgment where a human must be accountable. It also needs a working escalation path, or it will answer questions it should hand over.
A general chatbot answers whatever is asked from general knowledge. An AI coach works within a defined programme, tracks progress against goals, follows up unprompted, and escalates what it cannot handle. In a training context it should also be grounded in your course content rather than the open internet.
It can, indirectly, by removing the delay between a learner getting stuck and getting an answer. Most abandonment happens at a point of unresolved confusion, so faster resolution matters. Ask any vendor to evidence this on your own cohorts rather than accepting a benchmark figure.
If a vendor tells you AI coaching is as effective as human coaching, ask which measure they mean. The research supports one of them and is silent on the rest, and knowing which is which is the difference between a purchase that works and one that disappoints.

One thought on “Does AI Coaching Actually Work? What the Evidence Shows and Where It Fails”