What a thirty-claim sample cannot tell you
Sampling is how most reviews are scoped, and it is why most reviews find the same three things. The expensive findings are not rare enough to be missed by chance — they are in the departments nobody sampled.
The standard operating review samples thirty claims, interviews eight people, reads one quarter of data and extrapolates. It is a defensible method when the thing you are measuring is evenly distributed. Operating loss is not evenly distributed.
Why the sample keeps finding the same things
A thirty-item sample is large enough to characterise a common failure and far too small to find a concentrated one. If four per cent of claims carry a coding error, thirty claims will usually surface it — and that finding will be in the report, because it is easy to find and easy to quantify.
The findings that are worth real money behave differently. A single vendor contract with an auto-renewing escalator. One service being paid for twice under two purchase orders. A credentialing queue that adds eleven days to every new clinician’s start date. Each of those is a single instance, not a rate. Sampling is structurally incapable of finding them, because there is nothing to sample.
This is why reviews commissioned from different firms tend to produce overlapping reports. They are not converging on the truth. They are converging on what a sample can see.
The distribution is the point
In prior engagements, the pattern repeats: the largest single finding is rarely in the department that commissioned the review. Finance asks for a claims review; the money is in contracts. IT asks for an infrastructure assessment; the cost is in a manual approval loop that adds four days to every request.
That is not a criticism of the people who scoped it. It is the natural consequence of scoping a review by department when the loss crosses departments.
What we do instead
We review the operating system rather than a department, and we do it end to end rather than by sample:
- Claims, billing and finance operations, including AR/AP and write-offs
- Invoices, vendors and legal contracts, including terms nobody is enforcing
- IT operations, access and the data workflows underneath them
- Credentialing, marketing and sales handoffs
Full coverage costs more time than a sample. It also produces findings a sample cannot reach, and it removes the argument about whether the sample was representative — because there was no sample.
That is the engagement itself: The Operating Review, one business unit, priced as a single fixed fee and raised on a purchase order like anything else. Multi-site and multi-entity coverage is quoted per engagement, because the number of units and systems drives the effort rather than the headcount.
The check we put on ourselves
Before fieldwork begins we write down three dated predictions about what we expect to find, including one thing we expect is already working well. After the engagement, those predictions are compared against what turned up.
If we are right, it shows the diagnosis came from experience rather than hindsight. If we are wrong, the miss is documented too. Either way the method stays testable, which is more than can be said for a report that only ever describes what it already found.
The predictions are written on every engagement, including the shortest one. If you want to test the method against your own operation before commissioning a full review, the operating diagnostic is the bounded way to do it: a short look across the same areas, enough to establish where the loss is concentrated and whether a full review is worth commissioning at all. If it is not, we say so.