I have seven kids, which means I have spent about two decades living with averages that describe absolutely nobody. The household is doing fine, statistically. Statistically has never once told me which kid needs something on a Tuesday.
I thought about that a lot this summer while I was reading carrier scorecards.
Every carrier will hand you an on-time rate, and most of those numbers are good. What none of them hand you is which of your lanes that number is carrying, and which ones it is covering for. A carrier that visibly stumbles is a problem you can see and staff around. This is about the other version, the one where nothing stumbles, no incident gets opened, and your delivery promise erodes anyway.
Reliability is not a carrier trait
I think the mental model is where this goes wrong. We talk about carriers the way we talk about vendors, as though reliability is a trait each one has. This one is solid. That one has been shaky since spring.
Performance belongs to the carrier, the lane, the week, and the service level together, and those four move independently of one another.
A carrier’s Midwest ground network can hold its transit commitment straight through December while its West Coast lanes run two days long. Both things are true at once, both sit inside the same published number, and both sail through the quarterly review while everyone nods along. Call the aggregate 94% on time. That figure is doing a great deal of work, and some of that work is covering for a lane sitting at 81%.*
Which means the question most teams ask, is this carrier performing, does not really have an answer. The question with an answer is narrower and more annoying. Is this carrier performing on the lanes I actually ship, in the week I am actually shipping them.
Almost nobody is set up to answer the second one.
Shippers already told us this, three answers down the list
I went back through our own data on this, and the answer had been sitting in a spot nobody quoted.
Start with what people said they are worried about. In our 2026 Peak Readiness Index, a survey of 141 verified logistics professionals fielded in June (screened down from 2,004 raw submissions, for anyone about to do the math on the sample), managing multiple carriers and routing decisions came in as the number three peak concern overall at 20%. Among 3PLs, who live inside multi-carrier operations every day, it was the number one concern at 40%. That 3PL subset is small enough to be directional, so read it as a signal. Still, the group closest to this problem ranks it highest, and that is usually a decent sign you are looking at something real.
Then there is the visibility question. We asked the 115 shippers in the survey where visibility is worst in their operation, and everyone quoted the same answer back, including me: 23% named actual all-in shipping costs as their single biggest blind spot.
Sitting directly underneath it were two answers nobody picked up. Twenty-one percent named package delivery and the consumer experience as the area they can see the least. Another 18% named which carriers to use for which shipments.
One caveat before I add those together. We asked about carrier selection and about delivery experience. We did not ask specifically whether teams can measure performance lane by lane, so this does not prove my diagnosis.
But they are separate answers to the same question, which means you can add them. Thirty-nine percent of shippers say their biggest blind spot sits somewhere between choosing a carrier and seeing what actually happened to the package. That is a larger share than cost, and it is a worse kind of blind spot. Cost at least turns up on an invoice eventually. This one does not turn up anywhere.
Two dates, compared by lane and week
The metric is not complicated. It is two dates, compared, across enough shipments for the pattern to be real. The particulars are where it goes sideways, and each wrong particular hides a different kind of failure.
The promise you measure against should be the date you gave your customer, not the date the carrier gave you. Those two diverge more than people expect, and only one of them is the date your support team has to defend on a Tuesday afternoon in December. It is not the carrier’s.
Never cut it by carrier alone. A blended carrier-level number is doing precisely the job an average is built to do, which is make variance disappear. Segment by carrier, lane, and service level instead. The blended figure tells you nothing you can act on; the lane-level figure tells you where to move volume Monday morning.
Timing catches people. Degradation does not arrive as a gentle slope. It shows up around specific dates, so a monthly average will smooth the shape into nothing right when you needed to see it. Recompute weekly once peak starts.
A package marked out for delivery is not a delivered package. The distance between those two events is where a real share of December misses are hiding, so measure at the delivery scan.
Without a volume floor you will spend December chasing noise and moving volume around for no reason. A lane needs enough shipments in the window before a swing means anything. Decide that minimum in September, when you are calm.
| What to compare | The right version | The failure it hides |
|---|---|---|
| On-time against | Your promised date, not the carrier’s | Carrier-promise drift your CX team absorbs |
| Segmented by | Carrier, lane, and service level | A blended rate averaging away one failing lane |
| Recomputed | Weekly through peak | Surge-shaped degradation smoothed flat |
| Measured at | Delivery scan | “Out for delivery” counted as delivered |
| Acted on above | A minimum shipment count per lane | Routing decisions made on noise |
Now the honest part. If you are running a few hundred labels a day across four lanes on one carrier, none of this is worth your Thursday. The monthly view is fine, you would see a real problem anyway, and you have better things to do. This starts paying somewhere north of a few thousand labels a week and several lanes, where a single underperforming origin-destination pair can carry enough volume to matter and still hide inside the total.
Nobody opens the report in December
This is where good intentions go to die.
A weekly carrier performance report that nobody routes against is a report. It gets built in October with real enthusiasm, circulated twice, and by the second week of December everyone is too underwater to open it. I have watched that happen. I have probably caused it once or twice.
We already found that only a third of teams have tested a carrier backup plan. This is the less dramatic cousin of that problem. The plan exists, the report exists, and neither one gets used at the exact moment it would have changed something.
The teams that hold their delivery promise through peak removed the need for discipline, by making the measurement and the routing decision one system instead of two separate acts of willpower in your worst month.
Set the threshold in September and let December trigger it
You set a threshold for a lane before peak, say on-time falling below 85% across a rolling seven days on enough volume to count. When it trips, a predefined share of eligible shipments moves to your second carrier on that lane, and you compare the next seven days. The decision got made in September. December only triggers it.
Doing that by hand, across every lane and service level, in your busiest month, is not a realistic ask of a human being. That is the part software should be carrying.
Luma AI Insights is trained on more than a billion historic shipments across 100+ carriers, enough to separate a real benchmark from one warehouse’s bad week. Its delivery-date estimates run 39% more accurate than carrier estimates. Luma AI Select then acts on it, routing each label to the best-performing or best-value carrier under rules you set, without anyone opening a spreadsheet.
A global recommerce marketplace running about 25,000 labels a day moved on-time delivery from 80% to 83% and cut roughly 273,000 late deliveries a year. What gets me is the list of things they did not do to get there. They added no new carriers. They renegotiated no contracts. They hired nobody.
Three points of on-time delivery sounds like rounding until you multiply it across 25,000 labels a day and count the support tickets that never got opened.
What would tell you?
If one of your carriers holds at 94% overall this December while missing badly on two of your lanes, what in your operation tells you? Not in a January post-mortem, when the number is final and useless. That week, while you can still move volume somewhere else. If the honest answer is a customer complaint, the routing decision is getting made for you by whichever carrier is set as the default.
The average was never going to tell me which kid needed something. It won’t tell you which lane did either.
Run the Peak Readiness Stress Test and find out which of these questions your operation can already answer, while there is still time to change the answer.
* The 94% and 81% figures are an illustration, not EasyPost data.
Find out where your peak plan is thin
The Peak Readiness Stress Test scores your carrier contingency and cost readiness against 141 logistics professionals. Five minutes, no signup required.