Part One

The Observability Gap

Part One: The Evidence
What happens to a sales methodology after the training ends.
Who wrote this, and why you should check it. John Cunningham is the founder of One Click Coaching. He has spent fifty years in sales, sales management, and sales training, and has deployed Sandler, Challenger, SPIN, Gap and MEDDIC. He has a commercial interest in the conclusions of this study, which is a reason to check his sources rather than to trust them.

Why this is Part One

I said I was running a study. I am. It isn't finished, and I am not going to send you a finished-looking document that isn't one.

What follows is the part that is done: the evidence base. Every claim marked by type, every source named, every weakness stated. It took longer than expected, because most of what I started with did not survive being checked.

Part Two is the interviews. Sales leaders, frontline managers, and enablement heads at organisations with an installed methodology. That work is underway. If you want to be in it, the last section tells you how — and there is a four-minute diagnostic before it that will tell you whether you have anything to contribute.

You will get Part Two before it is published. Including the parts that contradict what is written below.

How to read this

Three kinds of evidence appear here, and they are not equal. I have marked each one.

[PEER-REVIEWED] — published, replicated, checkable.

[INDUSTRY DATA] — vendor or analyst research. Directionally useful. Not independent.

[INTERVIEW] — what sales leaders told me directly. Small sample. Stated as such.

I am marking them because most content in this category doesn't. A number without a source is a decoration. There are a lot of decorations in sales enablement, and I have used some of them myself.

One in particular. For two years I repeated the claim that sales training has a half-life of seventy-two hours. It is a good line. It is not a finding. There is no study behind it, and there is no study behind its cousins — 87% forgotten in thirty days, 84% in ninety, 90% with no lasting impact. Every trail leads to a vendor blog citing another vendor blog. No sample size. No methodology. No paper.

I am not repeating it again. What follows is what survives when you take the decorations off.

[PEER-REVIEWED]

1. The problem is not decay. It is design.

You do not need a forgetting percentage to explain why the workshop didn't hold. You need the schedule.

Cepeda, Pashler, Vul, Wixted and Rohrer (2006, Psychological Bulletin) meta-analysed 839 assessments across 317 experiments. The finding is one of the most durable in cognitive psychology: practice distributed across time beats identical practice compressed into a single block. And the effect scales with the interval — the longer you need to retain something, the wider the gaps between exposures must be.

Read that against a two-day sales kickoff.

A methodology you need for the next three years, delivered in a single block, with no scheduled second exposure. That is not a training failure. That is a schedule that was never designed to produce retention in the first place. It was designed to produce a calendar event.

The training is not the weak link. The calendar is.

Where this is weaker than it looks: Cepeda's underlying studies are largely verbal recall under laboratory conditions. Applying the spacing effect to complex interpersonal skill executed under pressure in the field is an inference, not a finding. I believe the inference holds. I cannot prove it holds, and neither can anyone selling you the opposite.
[PEER-REVIEWED] + [INDUSTRY DATA]

2. Coaching works. Coaching skill is the variable — and almost nobody has been taught it.

Let me correct something I have said badly in the past.

I have argued that more coaching is not the answer. That framing is wrong, and it gives cover to every leader looking for a reason not to bother. Coaching is not the problem. Coaching is one of the few interventions in this literature with a defensible evidence base.

Jones, Woods and Guillaume (2016, Journal of Occupational and Organizational Psychology) meta-analysed workplace coaching and found positive effects on organisational outcomes overall — δ = 0.36, with skill-based outcomes at 0.28 and individual-level results considerably higher. A later meta-analysis restricted only to randomised controlled trials — 37 studies, 39 samples, n = 2,528 — put the effect at g = .59, in the moderate range, while noting significant publication bias.

Coaching works. That is settled well enough to build on.

But the Jones meta-analysis explicitly excluded manager-to-subordinate coaching. It measured professional coaches. Which leaves the question every sales leader actually needs answered: does it work when the coach is the frontline manager, and what determines whether it does?

The answer comes from a sales floor.

Dahling, Taylor, Chau and Dwight (2016, Personnel Psychology) tracked 1,246 sales representatives across 136 teams in a pharmaceuticals organisation over a full year, against objective annual sales goal attainment. They separated two things that are usually collapsed together: how often managers coached, and how well they coached — coaching skill assessed independently through a training exercise rather than self-reported.

Coaching skill was directly related to the annual goal attainment of the reps those managers supervised, partially mediated by team-level role clarity.

Then the finding that matters most:

Coaching skill moderated the effect of coaching frequency. Where coaching skill was low, coaching frequency had a negative effect on sales goal attainment.

Not smaller. Negative.

Kluger and DeNisi's feedback intervention theory is the mechanism the authors reach for, and it is the same one that runs underneath the earlier finding: feedback improved performance on average (d = .41 across 607 effect sizes and 23,663 observations), but more than a third of feedback interventions made performance worse, and effectiveness collapsed as attention moved from the task to the person.

So the sentence is not "coach less." The sentence is this:

Coaching frequency is a multiplier. It multiplies whatever skill the manager already has, in whichever direction that skill points.

A skilled coach doing it weekly is the strongest lever in the business. An unskilled coach doing it weekly is subtracting revenue, on schedule, with good intentions.

Which raises the obvious question. How many sales managers have been taught to coach?

[INDUSTRY DATA] — and here the evidence gets thin in a way I want to flag rather than hide. MySalesCoach's State of Sales Coaching research, drawing on more than 3,700 sales professionals, reports that only 34% of sales managers have ever received any training or support to become a more effective coach — roughly two thirds have had none — and that only about one in five leaders has a coach of their own. Separate work attributed to CEB Global puts the figure for new managers generally at 85% receiving no management training before taking the role.

Other vendor sources in this space quote one in five, or one in ten, with no traceable methodology behind either. The spread between 10% and 34% is wide enough that no one should quote a precise number, including me. What survives the spread is the direction: the substantial majority of sales managers have never been trained in the skill that the peer-reviewed evidence says determines whether their coaching helps or harms.

Put the three findings together.

Coaching skill drives sales performance. Coaching frequency without skill drives it downward. And most managers were promoted for selling, handed a team, and asked to coach a methodology they were never taught to coach against.

"We already do weekly one-on-ones" is not a defence. It is a frequency claim, offered in answer to a skill question.

What this means for the argument that follows. If skill is the variable and skill is largely absent, there are two paths. Train every manager to coach — slow, expensive, and it decays like every other workshop, per section 1. Or reduce how much the outcome depends on manager skill in the first place: a defined standard, applied consistently, with the observation and the feedback carried by the system rather than by the manager's judgement in the moment.

I have a commercial interest in the second path. I would rather say that here than have you find it out later. Read section 5 and decide whether the evidence holds regardless of who is holding it.

[INDUSTRY DATA]

3. The manager bottleneck is arithmetic, not attitude.

Ask a sales leader why coaching isn't happening and you get an apology. Ask what a manager's week actually contains and you get a calendar that makes coaching structurally impossible.

McKinsey's work on frontline management puts 30–60% of a manager's time in administration and meetings, and another 10–50% in individual-contributor work — leaving somewhere between 10% and 40% for people management. Coaching is a subset of that remainder. Not the whole of it.

Meanwhile the denominator is growing. Average span of control moved from roughly 10.9 direct reports to 12.1 in a single year.

Run the arithmetic on a manager with twelve reps and eighteen hours of genuine people time. Ninety minutes per rep per week — before the pipeline scrub, before the forecast call, before the escalation, before the hire. What remains for skill coaching against a methodology is fifteen to twenty minutes.

On a week where nothing goes wrong.

Nothing is wrong with these managers. They were promoted for selling, given a span designed by a spreadsheet, and asked to do a developmental job in the margins of an administrative one.

The instruction "coach more" asks them to find time that does not exist.

Run your own version. Your span. Your admin load. Divide what's left. The number you get is the real coaching capacity of your organisation, and it is not the number in your enablement plan.

[INTERVIEW] + [INDUSTRY DATA]

4. Everything you measure arrives after the behaviour.

A sales trainer with twenty-five years in the work put it to me plainly: forgetting is a real problem, and as far as he can tell the industry has no solution for it. That is testimony against interest, from inside the supply side. I am weighting it accordingly, and I will report his reasoning in full in Part Two.

But the more useful thing he did was reframe what your instrument panel actually shows you.

The CRM reports what closed. The pipeline reports what might. Conversation intelligence reports what was said on the one call somebody chose to open.

All three are lagging. Every one of them measures the residue of behaviour that has already happened, on a deal that has already moved, at a point where intervention is no longer available. None of them observe the behaviour itself, at scale, while it is still being executed.

This is the part that is easy to miss, because the raw material is already sitting in your organisation. Every company with a recording platform is capturing the evidence and using it to review deals, not to observe behaviour. The recordings answer "what happened on the Acme call." They are almost never asked "is the methodology showing up across all four hundred calls this month, and where is it breaking, and for whom."

The gap between those two questions is the entire subject of this study.

[INDUSTRY DATA] — and it shows up as a perception gap. Across recent industry surveys, roughly 90% of managers report coaching their teams at least monthly, while only about 62% of reps agree they receive it. These are vendor-sourced, and I have not been able to trace them to a primary methodology, so weigh them as an indication rather than a measurement. But managers and reps are describing two different companies, and only one of those companies exists.

A methodology you cannot observe is a vocabulary. It appears in QBRs. It does not appear in discovery.

[PEER-REVIEWED]

5. Being observed changes what professionals do.

This is the load-bearing claim, so it gets the most careful handling.

The obvious place to reach is the Hawthorne effect, and there is a trap in it. Levitt and List (2011) recovered the original Hawthorne plant illumination data and reanalysed it; the famous patterns did not survive. If you cite the factory lighting story, someone will catch you. It is a second seventy-two hours.

But the folklore being wrong is not the same as the phenomenon being absent, and this is where I under-read the evidence in an earlier draft.

Leonard and Masatu (2010, Journal of Development Economics) studied clinicians in the Arusha region of Tanzania. The measurement problem is the hard part: to know what observation changes, you need data on behaviour when the subject does not know they are being watched. Their solution was patient recall interviews conducted shortly after the visit, reconstructing what the clinician actually did, then compared against the same clinicians under direct observation by trained enumerators.

Observed, the clinicians practised measurably closer to their own best possible standard. Not better than they knew how to be. Closer to what they already knew.

The literature has a name for the distance between those two things: the know-do gap. It is documented across settings — providers who could correctly state the treatment guideline and then did not follow it in the consultation, with knowledge gaps explaining only a fraction of the total shortfall.

Read that against a rep who passed the certification and does not run an up-front contract on a Thursday afternoon call.

The gap is not knowledge. The gap is execution under no observation.

But observation alone is not the intervention. This is the second half, and it is the half that keeps the finding honest.

Michie, Abraham, Whittington, McAteer and Gupta (2009, Health Psychology) meta-regressed 122 evaluations covering 44,747 participants. Of twenty-six classified behaviour change techniques, self-monitoring explained the greatest share of variation in effectiveness between studies. And interventions that paired self-monitoring with at least one other technique derived from control theory — a defined standard, feedback on performance, review against the goal — were significantly more effective than those that did not: 0.42 against 0.26.

Roughly sixty per cent more effect, from the combination.

There is the mechanism, stated plainly.

Observation without a standard is surveillance. A standard without feedback is a poster. Feedback without review is a memo.

The four together are a control loop. A control loop is the only structure in this literature that reliably moves behaviour.

Most sales organisations have installed exactly one of the four.

Two honesty notes. Both studies are healthcare. Neither is a sales study, and I am asking you to accept a structural analogy — professionals, a documented standard, execution unobserved. I think the analogy is strong. It is still an analogy. Second, the Michie result is a meta-regression, which identifies association across studies rather than testing the combination directly. It points hard in one direction. It does not close the question.

The observability test

Six questions. Four minutes. Score yourself honestly — nobody sees this but you.

1. Reinforcement. When is the next scheduled exposure to your methodology, and is it on a calendar right now?
2. Evidence. Think of the last time you knew — not believed — that a rep executed the methodology correctly on a specific call. How did you know?
3. Diagnosis. If I asked you right now which rep is weakest at a named component of your methodology — the up-front contract, the pain funnel, the qualification step — where do you go to find out?
4. Coverage. Last month: how many customer calls happened, and how many were reviewed against the methodology — not against deal status?
5. Detection. If every rep quietly stopped using the methodology next Monday, how long before you found out?
6. Perception. Have you ever asked your managers how often they coach, asked your reps how often they're coached, and compared the two answers?
Answer all six to see your score
Nobody sees this but you.

What I am not yet able to say

I do not know how large this gap is. Nobody does — that is the point of the seventy-two hours problem.

I do not know whether observability changes behaviour in a sales context. I know it changes what clinicians do, and I know self-monitoring paired with a standard and feedback outperforms self-monitoring alone. Nobody has run that study on a sales floor with an installed methodology. Until someone does, section 5 is an inference and I will keep calling it one.

I do not know how durable the effect is. Observation may produce a step change that holds, or a lift that decays the moment the observer becomes routine. The healthcare literature is not settled on this and I will not pretend it is.

I do not know whether teams that measure methodology adherence outperform teams that do not, controlling for the things that actually drive quota attainment. The industry claims this constantly. I have not found a clean study.

I also do not yet have the interviews. Part One is evidence. Part Two is testimony, and testimony takes longer to gather than citations do.

Stating that plainly costs me something. It is also the only reason to trust the rest.

Part Two, and how to be in it

I am running structured interviews with sales leaders, enablement heads, and frontline managers at organisations with an installed methodology and teams of five to twenty-five reps.

Twenty minutes. Six questions. Not a pitch, and I will say so at the start.

Three things I am committing to in advance, because a study that finds only what it went looking for is not a study:

Everyone who takes part gets Part Two before publication, including the parts that contradict Part One.

If you run a team with a methodology you paid for, I want to know whether you can see it working.

Be one of the twenty.

Twenty minutes. Six questions. No pitch. You'll see the aggregate findings before anyone else — including the parts that contradict what's written above.

Book the 20-minute interview →
or

Not ready to book? I'll send you Part Two when it's ready.

Leave your details and you'll get the findings before publication.

No pitch. Your details go to a research list, not a sales list. You'll hear from me once, when Part Two is done.

Sources

Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380.

Kluger, A. N., & DeNisi, A. (1996). The effects of feedback interventions on performance: A historical review, a meta-analysis, and a preliminary feedback intervention theory. Psychological Bulletin, 119(2), 254–284.

Dahling, J. J., Taylor, S. R., Chau, S. L., & Dwight, S. A. (2016). Does coaching matter? A multilevel model linking managerial coaching skill and frequency to sales goal attainment. Personnel Psychology, 69(4), 863–894.

Jones, R. J., Woods, S. A., & Guillaume, Y. R. F. (2016). The effectiveness of workplace coaching: A meta-analysis of learning and performance outcomes from coaching. Journal of Occupational and Organizational Psychology, 89(2), 249–277.

Leonard, K. L., & Masatu, M. C. (2010). Using the Hawthorne effect to examine the gap between a doctor's best possible practice and actual performance. Journal of Development Economics, 93(2), 226–234.

Michie, S., Abraham, C., Whittington, C., McAteer, J., & Gupta, S. (2009). Effective techniques in healthy eating and physical activity interventions: A meta-regression. Health Psychology, 28(6), 690–701.

Levitt, S. D., & List, J. A. (2011). Was there really a Hawthorne effect at the Hawthorne plant? An analysis of the original illumination experiments. American Economic Journal: Applied Economics, 3(1), 224–238.

Gartner (CEB), frontline rep survey (n ≈ 6,000), on the relative impact of coaching quality. Reported via Challenger.

McKinsey & Company, research on frontline manager time allocation and span of control.

MySalesCoach, The State of Sales Coaching (2026), n > 3,700 sales professionals, on coaching training prevalence and coaching frequency versus quota attainment.

Recent industry surveys on sales coaching frequency and rep perception (vendor-sourced, primary methodology not traced).

A note on the last three. Claims attributed to Gartner, McKinsey, and industry surveys are secondary — retrieved through published summaries rather than the primary reports. I am working on primary retrieval for Part Two. Until then, treat sections 1 and 2's peer-reviewed core as load-bearing and the rest as directional.