← Back to News & Updates

Benchmarking maintenance pricing: how it actually works

Printed maintenance quotes for the same work scope laid side by side on a desk in a dim hangar office with a business jet blurred behind the glass, the comparison an aircraft maintenance pricing benchmark is built from

Suppose we tell you your quote is running twelve percent above market.

The correct response is: says who, based on what, compared against how many, and gathered when. If you did not ask that, you would not be doing your job, and frankly we would trust you less.

So this article is the whole answer. It includes the categories where benchmarking works well, the categories where it does not work at all, and the specific questions you should put to anybody who hands you a market comparison, us very much included.

Why aviation maintenance has no Kelley Blue Book

A used car is a defined object. It has a VIN, a mileage figure, a condition grade, and thousands of comparable transactions every month. That is what makes an index possible.

A maintenance event is not an object. It is a scope of work plus whatever the airframe turns out to be hiding.

Two aircraft of the same make, model, and year, going into the same inspection, can separate by six figures for reasons nobody could have known in advance. Corrosion in a place the last owner never looked. Prior repairs done to a standard you inherit rather than chose. Modification status. How the aircraft was flown, where it was parked, and how carefully the last shop put it back together. You cannot build a price index for a thing whose contents are discovered after the price has been agreed.

This is why the data that does exist in business aviation measures adjacent things. The JSSI Business Aviation Index tracks utilization across roughly two thousand business aircraft and reports average flight hours by region, industry, and cabin type. It is well constructed and genuinely useful, and it tells you how much aircraft are flying, not what their maintenance costs. Conklin and de Decker, now part of JSSI, models operating cost across a very wide range of makes and models on a long horizon. Also useful, and also a model rather than a record of what anybody actually paid.

There is no transaction price index for business aircraft maintenance. Not because the industry has been lazy about it, but because the underlying product resists being indexed.

What a mature version of this looks like

The closest developed analogue is collision repair, and it is worth understanding because it demonstrates both what is achievable and precisely where the achievable stops.

That industry runs on three competing estimating databases: CCC, which uses MOTOR data, along with Mitchell and Audatex, now Solera Qapter. Each publishes labor times and parts prices for defined operations. Insurers and shops negotiate against those numbers every day of the week.

The most instructive part is not the numbers. It is the reference manuals, known in that trade as the p-pages, which document exactly what is and is not included in every published labor time. Experienced estimators will tell you that the core professional skill is not knowing the number. It is knowing the inclusions and exclusions behind it, and those differ between the three systems, which is why an estimator reviewing work prepared in another system has to go read that system's manual.

And even with all that infrastructure, the database publishers mark their own limits in public. Guidance distributed through the Society of Collision Repair Specialists states plainly that the estimating databases are intended as a guide only, because actual time varies with severity, condition, and equipment, and because the professional performing the repair is the one positioned to diagnose the work. Audatex went further on one operation, removing its standard blend labor formula altogether and replacing it with a statement that the operation requires the estimate preparer's judgment and cannot carry a standard allowance.

That is a benchmark provider publicly drawing the boundary of its own product. It is the most honest thing in that industry, and it is exactly the posture that benchmarking aviation maintenance requires.

What can be benchmarked, and what cannot

Specificity matters more than enthusiasm here, so here is the honest map.

WhatHow well it benchmarksWhy
Labor rates by region and disciplineStrong A rate is a published fact, not an estimate. Regional variance is real, measurable, and stable enough to be worth knowing.
Defined scheduled tasks on a known typeStrong A phase inspection on a given model has a published task list. Shops quote it repeatedly, so genuine comparables exist.
Commercial terms: consumables percentage, parts markup, handling fees, overtime policy Strong, and almost never checked These are policies rather than estimates. Directly comparable, and they persist across every event you will ever have.
A specific part number in a stated conditionStrong Same part number, same condition, same tag type is a true like for like comparison.
Turn time on a defined scopePartial Driven by slot availability, staffing, and parts. Comparable within a season, misleading across seasons.
Findings and troubleshooting on your airframeNot benchmarkable Nobody knows what is behind the panel until it is open. Anyone quoting a benchmark for this is guessing confidently.

The row most people skip past is the third one. Commercial terms are the most benchmarkable thing in the entire maintenance relationship and the least frequently benchmarked. A consumables percentage, a parts markup, a handling fee, and an overtime approval policy are not estimates subject to discovery. They are stated positions, they are directly comparable between shops, and unlike any single event, they apply to every event you will ever run with that facility. We have walked through where those four numbers live in how to read a six figure quote and in the five line items worth questioning on every inspection invoice.

Getting those four numbers right once is worth more over five years than winning any individual negotiation.

The last row is the one we want to be loudest about. Findings cannot be benchmarked at the event level, and anybody who tells you otherwise is selling something. What can be tracked is a given shop's historical relationship between quoted base scope and final invoice across many events. That is a genuinely useful number and it is a different claim from pretending to know what is inside your wing.

Where the data comes from

Not all price information is the same quality, and the differences are large enough to matter. Ranked from most reliable to least:

  1. Live competitive bids on an identical written scope, gathered in the same window. This is the strongest evidence available because it is not an opinion about the market. It is the market, responding to a real requirement, with real money attached.
  2. Historical awarded prices on the same defined scope, adjusted for elapsed time. Solid, provided the scope really was the same and the adjustment is disclosed.
  3. Published rate cards and list prices. These tell you what a shop hopes to receive, which is related to but distinct from what a shop accepts.
  4. Surveys. Ask a room full of operators what they paid for their last phase inspection and you will collect recollection, rounding, selection bias, and a durable human tendency to remember whichever number makes the speaker look shrewd. Survey data has its uses. Pricing decisions are not among them.

We build on the first two. Every request for quote that runs through the platform is, by construction, a small controlled experiment: one written scope, several qualified shops, the same week, responses on the same axes. That is not a clever data science trick. It is just what happens when you make people bid on the same thing at the same time, and it is the only mechanism that has ever produced honest price information in any market.

The five questions that make a benchmark honest

Put these to anyone who gives you a market comparison. Put them to us.

1. Comparability

Was the scope identical, and was it identical in writing? If two shops priced meaningfully different work, the comparison is theater. This is the single most common failure and it is usually accidental rather than deliberate.

2. Sample size

How many bids sit behind the number? Four real bids on the same scope is evidence. One is an anecdote wearing a suit. Anybody unwilling to tell you the count is telling you the count.

3. Recency

When were these gathered? In a market where lead times on some categories have moved from four to six weeks out to twenty weeks and beyond, a comparison built on eighteen month old pricing is a history lesson, not a benchmark.

4. Selection

Who was invited, and who declined to bid? A pool of shops competing hard for work in a slow month produces a different picture than the same pool in peak season with full hangars. Non responses carry information, and a benchmark that quietly discards them is flattering itself.

5. Inclusions and exclusions

This is the p-pages problem again, and it is where most comparisons quietly fall apart. Does the number include freight, taxes, consumables, outside services, and tooling? Two totals that treat those differently are not comparable, no matter how similar they look. If nobody can tell you what sits inside the number, the number is decorative.

Ask us all five about any specific figure we give you. If we cannot answer them for that figure, do not act on it.

In range is a result, not a wasted exercise

Most people assume the point of benchmarking is catching a high quote. That is the least frequent outcome and, oddly, the least valuable one.

The most common outcome is that the quote comes back in range. That is not the exercise failing. That is the exercise working, and it is worth understanding why the confirmation has value even when nothing changes.

An unverified quote costs you something even when it is perfectly fair. It costs the background hum of approving a very large number you cannot defend if asked. It costs the meeting where a principal asks whether you shopped it and the honest answer is that you think it is probably fine. It costs the risk that you delay an event, or lean on a shop for a discount they had no room to give, and damage a relationship worth considerably more than the discount you were chasing.

Confidence is a deliverable. A quote confirmed in range converts a low grade anxiety into a decision, and it does that for everyone in the chain at once. The owner stops wondering. The DOM stops defending. The chief of staff has something to put in a file.

The shop gains as well, which people rarely think about. A facility that priced fairly and got questioned anyway has lost something real. A facility that priced fairly and can be shown to have priced fairly has gained a customer who will stop second guessing them, which is worth more to a good shop than the margin on one event.

What we will not tell you

Three things, and they matter as much as anything above.

We will not tell you your findings should have cost less. We were not there, the panels were not off in front of us, and the technician who made that call had information we do not have.

We will not tell you a shop is overcharging when what we actually observed is that a shop is more expensive. Those are genuinely different claims. A higher rate can buy better scheduling, better documentation, cleaner paperwork, a hangar slot you can actually get in March, and a turn time worth considerably more than the difference.

We will not hand you a number with a decimal place it has not earned. A range with a stated sample size is more useful, and more honest, than a point estimate implying a precision that nobody in this industry possesses.

The data most owners never get to see

Here is the structural problem that makes all of this hard for an individual operator, and it has nothing to do with anyone's competence.

A flight department running two significant maintenance events a year gets two observations a year. Even a sharp DOM with two decades in the seat is working from a few dozen data points accumulated across different aircraft, different markets, and different decades. That is not enough to know what normal looks like, and it never will be. Nobody should be expected to build a market picture from a sample that small.

The only way anyone learns what normal looks like is by seeing the same scope quoted many times, by many shops, close together. That is what a request for quote through VHMX produces as a byproduct of doing the thing you actually wanted: same written scope, multiple real bids, current week, presented on the same axes.

Send us a quote and we will tell you where it sits and how confident we are in saying so. Sometimes the honest answer will be that we do not yet have enough comparable bids on your type to tell you anything useful, and when that is the answer, that is what you will hear.

A benchmark you cannot trust is worse than no benchmark, because it feels like knowledge.

VHMX is not a repair station and does not perform maintenance. We do not tell you what work your aircraft needs. We tell you what the market says about the price of the work you have been quoted, and we tell you how much weight that opinion deserves.

Sources and further reading: JSSI Business Aviation Index, methodology and quarterly releases, Jet Support Services Inc. JSSI Conklin and de Decker, Aircraft Cost and Performance Comparison Guide. Society of Collision Repair Specialists, published guidance on estimating database labor times. Autobody News and Repairer Driven News reporting on CCC, Mitchell, and Solera Qapter database reference manuals. Database Enhancement Gateway estimating reference materials.

Share this:Copy linkEmailLinkedIn
See for yourself

Find out where your quote actually sits

Send us your next maintenance quote and we will tell you where it lands against real, current bids on the same scope, and how much weight that comparison deserves.