How to read a benchmarking study you didn’t write.
Most in-house tax teams accept benchmarking studies as black boxes – hundreds of pages of comparables they simply don’t have time to scrutinize. This is a mistake. Here’s a framework for pressure-testing any study in under an hour.
A tax director at a mid-market multinational recently sent me a benchmarking study prepared by an external advisory firm. Two hundred and forty pages. Comparable searches across three jurisdictions. An interquartile range that supported the company’s existing TP policy almost perfectly. The fee was substantial.
“Can I trust this?” she asked.
It is the right question, and the honest answer is that most in-house tax teams have no way to answer it. They receive a study from an external advisor, note that the conclusion supports their desired position, and file it away in the documentation binder. The methodology-the actual analysis that determines whether the study will survive an audit-goes unread.
This is a critical risk. An unverified study is functionally equivalent to no study at all when a tax authority begins asking questions. An auditor’s primary concern is not the premium paid for the report; their focus is entirely on whether the underlying methodology is defensible.
01.The 60-minute framework
Here is how to pressure-test any benchmarking study in under an hour. The framework consists of five sequential steps, each asking one specific question. If a study fails at any point, you have identified a vulnerability that must be addressed before a tax authority surfaces it for you.
My goal here is not to automatically disqualify studies, but to illuminate exactly where the weaknesses live so you can either remediate them or thoroughly document your rationale for proceeding.
Step one: Was the method selected, or was it assumed?
Open the study and locate the method selection section. Read it carefully. The core question: Does the study explicitly explain why this method was chosen, and critically, why competing methods were rejected?
A defensible method selection names at least three alternative methods, provides data-driven and functional reasons for rejecting each, and arrives at the chosen method through documented logic. A weak section relies on a conclusory statement—for example, “the Transactional Net Margin Method (TNMM) was selected as the most appropriate method”—and moves on.
“Method selection is the most consequential decision in the study. It is also the section auditors read first.”
Conclusory method selection will not hold up under regulatory scrutiny. If you see it in a commissioned study, require your advisor to provide the underlying analysis.
Practice note —
The U.S. Treasury Regulations under §1.482-1 explicitly require the “best method” rule – a documented analysis showing why the chosen method provides the most reliable measure of arm’s length results. The OECD Guidelines (paragraphs 2.1–2.12) apply substantively the same standard. In any jurisdiction following OECD principles, method selection analysis is not optional.
Step two: Does the functional analysis reflect operational reality?
Skip to the functional analysis. Read it as if you are entirely unfamiliar with the business. Does the description of functions, assets, and risks match what your operations team actually executes daily? Or does it read like generic boilerplate that could describe any distribution entity or contract manufacturer?
This is where the majority of studies are weakest. Functional analyses drafted from intake forms or brief interviews frequently read like marketing copy rather than analytical assessments of risk allocation and functional performance.
The Functional Analysis Workbook
The structured interview template, scoring matrix, and output summary I use to build defensible functional analyses. The unglamorous work that holds up under audit.
The Test: Pick three functions described in the analysis. Can you name the specific person or team in your company who actually performs them? If you cannot, the analysis is operating at the wrong level of abstraction and must be grounded in operational reality.
Step three: Was the search strategy documented proactively, or retroactively?
This is the standard most studies fail to meet. A defensible benchmarking search establishes its acceptance and rejection criteria before the search is executed—not after the results are generated. This discipline prevents the analyst from reverse-engineering screens to produce a desired outcome.
Look for a search strategy memo or methodology section that systematically lists the following: the database used, industry codes searched, geographic scope, size screens, independence criteria, and qualitative screens. Crucially, each criterion must include a brief explanation of why it was applied.
If the section reads as though it was written to justify the results after the fact, the study’s analytical integrity is compromised. Even if it arrives at a reasonable conclusion, the process itself is difficult to defend.
Step four: Were the rejected comparables removed on principle, or for convenience?
Locate the acceptance/rejection matrix. Every study should include one—typically an appendix detailing each company considered, the specific screen that retained or rejected it, and a brief explanatory note. The absence of this matrix is a significant red flag.
For each rejected company, ask: Would this comparable still be rejected if I were defending a completely different conclusion? If the reasoning holds, the rejection is principled. If it appears the company was removed simply because its data skewed the interquartile range unfavorably, the rejection is post-hoc.
Post-hoc rejections require exceptionally strong functional or economic justification, and they must be flagged transparently. Hidden post-hoc rejections are precisely what tax authorities look for during an audit.
Step five: Is the interquartile range the conclusion, or a calculation?
The interquartile range should be the natural mathematical output of a rigorous screening process, never a predetermined target. When reviewing the final range, ensure that the data points feeding into it directly reflect the principled comparables established in Step 4. If manual adjustments, unexplainable exclusions, or statistical manipulations appear in the final calculation without robust economic justification, the study’s defensibility is structurally flawed.
A benchmarking study is only as valuable as its ability to withstand an audit. By applying this 60-minute framework, corporate tax teams can transform dense reports from unexamined black boxes into transparent, defensible assets. Demand rigor, verify the methodology, and ensure your external advisors are delivering analysis you can confidently stand behind.
02.Seven red flags that require scrutiny
If your commissioned study exhibits any of the following characteristics, require your advisor to explain-and rigorously document-their reasoning:
- The method selection is brief. It rejects fewer than three alternative methods, or dismisses them without specific, data-driven functional reasoning.
- The functional analysis is generic. It describes the business in broad terms that could apply to virtually any company within your industry.
- The search strategy is retroactive. The methodology section reads as a post-hoc justification rather than a proactively established set of criteria.
- The rejection matrix is incomplete. The acceptance/rejection matrix is either missing entirely or contains excluded comparables with no documented economic reasoning.
- The range lacks context. The interquartile range is presented as a standalone conclusion without a transparent explanation of how the benchmark was mathematically constructed.
- The results are suspiciously perfect. The tested party’s result falls exactly at the median. In transfer pricing, an artificially neat result can invite more auditor scrutiny than a noisy but analytically sound one.
- Adjustments are ignored. The study omits any discussion of comparability adjustments (e.g., working capital, operating expenses, or risk), even when the accepted industry comparables are clearly imperfect matches.
03.A final word
Benchmarking studies are not infallible documents. They are analytical arguments supported by data, and like any argument, their strength lies entirely in their foundation. While this framework will not transform you into a transfer pricing economist in an hour, it will equip you to ask the precise, critical questions required to hold your advisors accountable.
The tax director mentioned at the beginning of this article ran her commissioned study through this exact five-step framework. It passed the first two steps, but failed on steps three and four. The search strategy appeared reverse-engineered, and several comparables were rejected based on inconsistent, post-hoc reasoning.
She did not discard the study. Instead, she pushed back. She required her advisors to rigorously document the search strategy and apply a consistent, principled standard to the rejected comparables. The revised study required additional effort to correct, but the final deliverable became genuinely audit-defensible – which the original was not.
That is the true purpose of pressure-testing. The goal is not to arbitrarily disqualify the analysis you commissioned, but to ensure that the final deliverable provides the uncompromising defensibility your corporate tax strategy demands.
Other clusters.
01.
Documentation
7 ARTICLES
03.
Methods & methodology
6 ARTICLES
04.
Financial transactions
6 ARTICLES
05.
Controversy & defense
6 ARTICLES
06.
Industry & strategy
6 ARTICLES