Benchmarking · Long read

March 18, 2026

14 min read

How to read a benchmarking study you didn’t write.

Most in-house tax teams accept benchmarking studies as black boxes — pages of comparables they don’t have time to scrutinize. This is a  mistake. Here’s a framework for pressure-testing any study in under an hour. 

A tax director at a mid-market multinational recently sent me a benchmarking study prepared by a Big Four firm. Two hundred and forty pages. Comparable searches across three jurisdictions. An interquartile range that supported the company’s existing TP policy almost perfectly. The fee was just under forty thousand dollars.

“Can I trust this?” she asked.

It’s the right question, and the honest answer is that most in-house tax teams have no way to answer it. They receive a study from an external advisor, see that the conclusion supports the position they want, and file it away in the documentation binder. The methodology — the actual analysis that determines whether the study will hold up in audit — goes unread.

This is a mistake. Not because Big Four studies are necessarily wrong, but because an unverified study is functionally equivalent to no study at all when a tax authority starts pushing back. The auditor doesn’t care that you paid forty thousand dollars for it. They care whether the methodology is defensible.

01.The 60-minute framework

Here’s how to pressure-test any benchmarking study in under an hour. The framework has five steps, applied in order. Each step asks one specific question. If a study fails at any step, you have a problem worth surfacing before — not after — the tax authority surfaces it for you.

I’ve used this framework on hundreds of studies over the years, including studies I prepared myself early in my career that, in hindsight, would have failed at step three. The goal isn’t to disqualify studies; it’s to know exactly where the weaknesses live so you can either fix them or document why you’re proceeding anyway.

Step one — Was the method selected, or was it assumed?

Open the study and find the method selection section. Read it carefully. The question you’re asking: does the study explain why this method was chosen, and why competing methods were rejected?

A defensible method selection section names at least three alternative methods, explains the data and functional analysis reasons for rejecting each, and arrives at the chosen method through documented reasoning. A weak section says something like “the Transactional Net Margin Method (TNMM) was selected as the most appropriate method given the nature of the controlled transactions” — and moves on.

“Method selection is the most consequential decision in the study. It’s also the section auditors read first.”

The second formulation is what tax authorities call a conclusory method selection — a stated conclusion without underlying analysis. It will not hold up under scrutiny. If you see it in a study you commissioned, push back on your advisor.

Practice note —

The U.S. Treasury Regulations under §1.482-1 explicitly require the “best method” rule — a documented analysis showing why the chosen method provides the most reliable measure of arm’s length results. The OECD Guidelines (paragraphs 2.1–2.12) apply substantively the same standard. If your study was prepared for any jurisdiction following OECD principles, the method selection analysis is not optional.

Step two — Does the functional analysis match what the business actually does?

Skip to the functional analysis. Read it as if you don’t know the business. Does the description of functions, assets, and risks match what your operations team actually does? Or does it read like generic boilerplate that could describe any distribution entity, any contract manufacturer, any management services arrangement?

This is where most studies are weakest. Functional analyses get drafted from intake forms or three-hour interviews, and they often read like marketing descriptions of the business rather than analytical descriptions of who bears which risks and performs which functions.

A working tool for this step —

The Functional Analysis Workbook

The structured interview template, scoring matrix, and output summary I use to build defensible functional analyses. The unglamorous work that holds up under audit.

See the workbook

The test: pick three functions described in the analysis and ask whether you can name the specific person or team in your company who actually performs them, and where they sit. If you can’t, the analysis is operating at the wrong level of abstraction.

Step three — Was the search strategy documented before the search ran, or after?

This is the question most studies fail. A defensible benchmarking search has its acceptance and rejection criteria written down before the search executes — not after the results come in. The reason matters: it prevents the analyst from choosing screens based on which screens produce the results the client wants to see.

Look for a search strategy memo or methodology section that lists, in this order: the database used, the industry codes searched, the geographic scope, the size screens, the independence criteria, and the qualitative screens. Each criterion should have a brief explanation of why it was applied.

If the search strategy section reads as if it were written to describe what was done after the fact — without explaining why those choices were made before the search — the study has a quiet integrity problem. It may still arrive at the right answer, but the analytical process is compromised.

Step four — Were the rejected comparables actually rejected, or were they removed because the result was inconvenient?

Find the acceptance/rejection matrix. Every study should have one — typically an appendix showing each company considered, the screen that retained or rejected it, and a brief note explaining the reasoning. If there isn’t one, that’s a problem.

For each rejected company, ask: would you reject this comparable if you were defending a different conclusion? If the rejection reasoning would still apply, the rejection is principled. If it wouldn’t — if the company was rejected because its results pulled the interquartile range in the wrong direction — the rejection is post-hoc.

Post-hoc rejections aren’t always wrong, but they need very strong functional or economic justification, and they need to be flagged transparently. A study with hidden post-hoc rejections is the kind that loses in audit.

Step five — Is the interquartile range the conclusion, or a calculation?

The interquartile range is the most-cited number in any benchmarking study and the most-misunderstood. Many studies present the IQR as if the conclusion were a mathematical result — “the arm’s length range is 3.5% to 8.2%” — when in reality the IQR is a statistical summary of a set of comparables whose appropriateness is what was actually argued.

The question to ask: does the study explain why the median or interquartile range is the appropriate benchmark for this specific transaction, given the functional analysis? Or does it assume that landing within the IQR is, by itself, sufficient?

02.Seven red flags that should make you push back

If your study has any of the following, ask your advisor to explain — and document the explanation:

  • The method selection rejects fewer than three alternative methods, or rejects them without specific functional reasoning.
  • The functional analysis describes the business in terms that could apply to any company in the industry.
  • The search strategy section was clearly written after the search, not before.
  • The acceptance/rejection matrix is missing, or contains rejections with no documented reasoning.
  • The interquartile range is presented as a conclusion without explanation of why it’s the appropriate benchmark.
  • The tested party’s result happens to fall almost exactly at the median — too neat a result is sometimes worse than a noisy one.
  • The study contains no discussion of comparability adjustments (working capital, operating expenses, risk), even when industry comparables are imperfect matches.

03.A final word

Benchmarking studies are not infallible documents. They’re analytical arguments dressed in tables and charts, and like any argument, they can be strong or weak. The framework above won’t make you a benchmarking expert in an hour — but it will let you ask the right questions of the people you’ve paid to prepare the work.

The tax director I mentioned at the start of this article ran her forty-thousand-dollar study through the five-step framework. It passed steps one, two, and five. It failed steps three and four — the search strategy had clearly been written backward, and three of the rejected comparables had been rejected for reasons that wouldn’t apply if the rejection had been principled.

She didn’t throw out the study. She went back to her advisor and asked them to document the search strategy more rigorously and to revisit two of the rejections with stronger reasoning. The revised study cost an additional six thousand dollars in advisory time. It also became audit-defensible — which the original was not.

That’s the point of pressure-testing. Not to disqualify the work you’ve paid for. To make sure it’s actually worth what you paid.

T

Tosin Odunuga

Founder, TransferPricingLab

Tosin has spent over a decade in transfer pricing across Big Four advisory, in-house TP at a Fortune 500 multinational, and now as the founder of TransferPricingLab. She writes here on methodology, controversy, and the changing economics of TP advisory.

Other clusters.

01.

Documentation

7 ARTICLES

03.

Methods & methodology

6 ARTICLES

04.

Financial transactions

6 ARTICLES

05.

Controversy & defense

6 ARTICLES

06.

Industry & strategy

6 ARTICLES

Transfer PricingLab

Rigorous templates and advisory for a field that needs both.

Library

All products

Templates

Workbook

Courses

Practice

Documentation

Benchmarking

Controversy

Fractional

Company

About

Contact

LinkedIn

© 2026 TransferPricingLab. All rights reserved.

Privacy · Terms