The Ride-Along Is the Least Representative Call of the Quarter
The short version: Sales managers evaluate reps using observed calls, but people perform differently when they know they are being watched. Researchers call the distance between observed and unobserved performance the know-do gap. Medicine measured it. Aviation redesigned observation to eliminate it. Sales has done neither, and coaches from a small, self-selected, distorted sample.
Two economists went to Tanzania to solve a measurement problem.
They wanted to know how well doctors treated patients. Not how well the doctors could treat patients — that's a test, and doctors pass tests. They wanted to know what actually happened in the room on an ordinary Tuesday when nobody was looking.
The trouble is that finding out requires looking.
Kenneth Leonard and Melkiory Masatu solved it by turning the problem into the instrument. Doctors change their behaviour when an observer walks in. Everyone knows this. It's called the Hawthorne effect, and for a century researchers had treated it as contamination — noise to be minimised, controlled for, apologised for in the limitations section.
Leonard and Masatu measured it instead.
The gap between how a doctor performs while observed and how that same doctor performs unobserved is the answer. It is the distance between best possible practice and actual practice. Researchers call it the know-do gap.
The doctors weren't ignorant. They knew what good care looked like. They demonstrated it, reliably, the moment someone was watching.
They just didn't do it the rest of the time.
What happens when you sit in on a sales call
A manager blocks Thursday morning for ride-alongs. Three calls with a rep who's been trending down.
The rep knows on Monday.
By Thursday he has re-read the discovery framework. He has his questions written on a card next to the monitor. He is not going to skip the up-front contract on the one call his manager is listening to. He is going to be, for ninety minutes, the rep he knows how to be.
The manager watches. Takes notes. Concludes the rep is stronger than the numbers suggest, that the problem is pipeline, or territory, or luck.
The manager has just measured the ceiling and filed it as the baseline.
Everything that follows — the coaching plan, the performance conversation, the decision about who to keep — is built on a number that was never real.
This is not a rep problem
It is worth being precise, because the obvious reading of this is that reps sandbag and can't be trusted, and that isn't what the research says.
The Tanzanian doctors were not lazy or dishonest. They were operating under ordinary conditions — fatigue, volume, no feedback, no consequence for the shortcut, no reward for the extra question. Under those conditions, practice drifts from training. Everywhere. In every profession that has ever been measured this way.
Medicine has measured it.
Aviation did something more interesting.
How aviation solved the observation problem
In the early 1990s, airlines faced a version of the same difficulty. Crew Resource Management training had been rolled out across the industry — the standards for how a flight deck communicates, cross-checks, challenges, and recovers. Pilots were trained in it. Pilots were tested on it.
Nobody knew whether they used it.
Check rides told you nothing, for the obvious reason. A check ride is an examination. The pilot knows it is an examination. What you observe is the pilot's best available performance, produced under conditions that exist nowhere else in his working life.
So the industry built something different, now codified by ICAO as the Line Operations Safety Audit.
A trained observer rides the jumpseat on ordinary revenue flights — not training flights, not check rides. Routine Tuesdays. And the design constraints are the whole point:
Randomly selected flights. Not the ones anyone flagged. Not the ones attached to an incident. Systematically or randomly drawn from the line schedule, so the sample is a sample rather than a collection of interesting cases.
Strict no-jeopardy. Crews are not held accountable for anything observed. Nothing goes in a file. Nothing reaches a supervisor with a name attached. The observation cannot hurt you, so there is nothing to perform for.
The observer never intervenes. He watches. He does not help, correct, or coach in the moment. Short of an actual emergency, he is furniture.
Crews can decline. If a crew says no, the observer takes another flight, no questions asked. Voluntary participation is what makes the no-jeopardy promise credible.
Behaviour is coded against a validated taxonomy, not judged. Threats, errors, undesired states. Defined categories applied identically by calibrated observers, so that what comes back is data rather than opinion.
Then they ran it at more than ten airlines, and found something the training departments did not want to hear.
The actual practice of Crew Resource Management on the line looked substantially different from the version depicted in the classroom.
Not because pilots were incompetent. Because that is what happens to any methodology once it leaves the room where it was taught and meets the ordinary pressures of the job.
What sales has instead
Hold those five design constraints up against a sales ride-along and count the violations.
The flight is not random — the manager picked it, usually because something's wrong.
There is no no-jeopardy provision. The observer sets the rep's compensation and decides whether he keeps the territory.
The observer intervenes constantly, and is often expected to.
The rep cannot decline.
And the behaviour isn't coded against anything. It's judged, against whatever standard the manager holds in his head that morning — which is not the standard he held last month, or the one his colleague holds down the hall.
Five constraints. Sales violates all five, simultaneously, and then treats the resulting impression as evidence.
The sampling problem underneath
Set the observer effect aside and there is a second problem sitting behind it, arguably worse.
A manager with eight reps oversees somewhere north of two hundred conversations a month. He will review a handful. Call it five. Call it ten if he's unusually disciplined.
Which ones?
The ones the rep flagged. The ones attached to deals that went sideways. The ones that came up because someone complained. The ones that happened to be on the calendar when he had a free hour.
Not a random sample. A sample selected by relevance, by drama, by convenience — every one of those a bias, and all of them pointing the same direction.
So the manager's model of his team is built from a small non-random sample of performances that were altered by the act of sampling them.
And then he coaches from that model, with confidence. For years.
What you actually know about your team
Try this honestly.
Take your weakest rep. Write down the three things you believe are wrong with how they sell.
Now write down which specific calls those beliefs came from. Not the general impression — the actual conversations. Date them.
Most managers can name two. Sometimes one. Occasionally the belief traces back to a call from eleven months ago that has hardened into a permanent characteristic.
Then ask the harder version: was that rep aware you were listening?
And when a rep knows they are being watched, it matters where attention lands — on the call, or on the person watching.
The gap is not a training gap
Here is what this research keeps finding, across countries, professions, and decades: the distance between what people know and what they do is large, persistent, and almost entirely invisible to the people responsible for closing it.
More training does not close it. The doctors already knew. The pilots had been through CRM.
More motivation does not close it reliably either.
What closes it is knowing where it is — which requires observation that doesn't distort what it observes, applied to a sample that wasn't chosen for being interesting.
Aviation worked that out thirty years ago and rebuilt its methods around it.
Sales still books Thursday morning and tells the rep on Monday.
Stop measuring the ceiling.
The only call that tells the truth is the one nobody knew you'd hear. Send me one transcript and I will send back a scored report of what actually happened.
Send a Transcript →Common questions
How many sales calls does a manager actually review?
A manager with eight reps oversees roughly two hundred conversations a month and typically reviews five to ten. That is a single-digit percentage, and those calls are not randomly selected — they are chosen because a deal went wrong, a rep flagged them, or the calendar allowed it.
Does the observer effect really change how reps sell?
Yes. Leonard and Masatu demonstrated in Tanzanian clinics that practitioners perform measurably closer to their known best standard while observed and drift from it when unobserved. The effect is well established across professions, and there is no reason sales would be exempt.
What is the know-do gap?
The know-do gap is the distance between what a professional knows how to do and what they actually do under ordinary working conditions. It is a measurement problem before it is a performance problem, because observing someone tends to close the gap temporarily and hide its true size.
How does aviation handle observation better than sales?
Aviation's Line Operations Safety Audit observes randomly selected routine flights under strict no-jeopardy terms, with a non-intervening observer coding behaviour against a validated taxonomy. Sales ride-alongs violate every one of those conditions and are conducted by the person who controls the rep's compensation.
You are not coaching your team. You are coaching your memory of five calls you chose, on days they knew you were coming.
Sources
- Leonard, K. L., & Masatu, M. C. (2010). Using the Hawthorne effect to examine the gap between a doctor's best possible practice and actual performance. Journal of Development Economics, 93(2), 226–234.
- Leonard, K. L., & Masatu, M. C. (2010). Professionalism and the know-do gap: exploring intrinsic motivation among health workers in Tanzania. Health Economics, 19(12), 1461–1477.
- Klinect, J. R., Murray, P., Merritt, A., & Helmreich, R. (2003). Line Operations Safety Audit (LOSA): Definition and operating characteristics. Proceedings of the 12th International Symposium on Aviation Psychology, 663–668.
- International Civil Aviation Organization. Doc 9803 AN/761 — Line Operations Safety Audit (LOSA).