Behind Every Arm's-Length Range: Why Data Coverage Decides Whether a Benchmarking Study Holds Up

9
Min Read
A benchmarking study is decided before the quartiles are calculated. Database coverage, the industry classification used, the screens applied and whether every figure traces back to a filing determine whether the arm's length range is market evidence or just a number that looks like one.
#Interquartile Range (IQR)
#Berry Ratio
#Return on Capital Employed (ROCE)
#Transfer Pricing Documentation
#Comparable Company Analysis
#Database
#Benchmarking
Kakhaber Grubelashvili
on
18.9.26
Tax advisor and auditor at Rödl & Partner Georgia, specializing in tax advisory, audit, and cross-border compliance for internationally active companies.

Most discussions of TNMM benchmarking focus on what happens after the comparable set has been assembled — the interquartile range, the median, the comparison against the tested party's actual result. That part of the process gets the attention because it produces the number everyone is waiting for.

But by the time an analyst is computing quartiles, the outcome of the study has largely already been decided. It was decided earlier, in a set of choices that are easy to overlook: which database was searched, how deep its coverage went, which industry classification was used, how independence and data-quality screens were applied, and whether every figure in the final range can be traced back to a verifiable source.

Get those upstream decisions wrong, and no amount of careful statistics downstream will rescue the analysis. This article looks at the practical database decisions that determine whether a benchmarking study produces genuine market evidence — or just a number that looks like one.

The Coverage Problem: When "No Comparables Found" Isn't True

A common frustration in transfer pricing practice, particularly for entities based in smaller or less liquid markets, is the apparent absence of comparable companies. A practitioner searches a database restricted to listed companies, applies an industry filter, and comes back with three or four results — none of them a close functional match.

The instinctive response is often to widen the geographic search to an entire region or continent, which solves the population-size problem but introduces a new one: the resulting comparables may operate in economic conditions — competitive intensity, regulatory environment, labor costs, market maturity — that look nothing like the tested party's actual market.

The underlying issue is frequently not that comparable companies don't exist. It's that they aren't listed. Most operating companies in the world — including the overwhelming majority of routine distributors, contract manufacturers, and service providers that make the best TNMM comparables — are privately held. A database limited to public-company filings is, by construction, working with a small and often unrepresentative slice of the relevant market.

This is one of the more practical reasons a database with genuinely broad coverage matters. smartZebra indexes standardized financial data across more than 50,000 public companies and over 500,000 private companies globally, which means a search for, say, routine IT services providers in Central or Eastern Europe returns an initial population large enough to survive a rigorous qualitative screening process and still leave a defensible number of accepted comparables — rather than forcing the analyst to choose between an unrepresentative sample and an artificially widened search.

Industry Classification: One Code Rarely Tells the Whole Story

A second practical issue sits one step earlier than the search results themselves: which classification system was used to define the search, and how precisely it maps to what the tested party actually does.

NACE, SIC, and NAICS codes are not interchangeable, and a code that captures the right activity in one system can be too broad — or miss relevant peers entirely — in another. A company classified under a general "computer programming" code, for instance, might include everything from routine contract developers to businesses that own and commercialize proprietary software platforms. Relying on a single classification system, or on a code chosen for convenience rather than functional accuracy, can quietly distort the entire initial population before qualitative review even begins.

A benchmarking platform that supports cross-referencing across classification systems — filtering by NACE and SIC and NAICS together, rather than being locked into one — gives the analyst more control at exactly the stage where a narrow or mismatched code can do the most damage. It doesn't replace the judgment call of deciding which businesses genuinely match the tested party's functional profile, but it does mean that judgment call is being made on a properly assembled starting population rather than one skewed by a classification quirk.

Quantitative Screens: Clearing the Noise Before the Real Work Starts

Not every weak comparable needs to be caught by manual review. A meaningful share of unsuitable candidates can, and should, be filtered out through quantitative screens before the analyst spends time reading business descriptions one by one: companies with persistent losses that suggest a fundamentally different business situation, related-party ownership that undermines independence, or R&D intensity levels inconsistent with a routine functional profile.

This matters for a reason that goes beyond convenience. Qualitative review — genuinely reading each candidate's business description and financial history — is the part of the benchmarking process that carries the most audit weight, and it is also the most time-intensive. An analyst who has to manually work through eighty candidates, twenty of which could have been screened out automatically on independence or loss-making grounds, is spending scarce professional judgment on cases that never had a realistic chance of being accepted. Automated screening applied before the qualitative stage means that judgment gets spent where it actually changes the outcome — on the genuinely borderline candidates where a documented economic reason for acceptance or rejection is what an examiner will actually be testing.

Profit Level Indicators: The Right Ratio for the Right Business

The choice of Profit Level Indicator is not a formality. Operating margin suits a distributor whose profitability is driven by sales; Net Cost Plus or a mark-up on total costs suits a routine service provider or contract manufacturer whose economics are driven by cost recovery; a Berry Ratio can be more meaningful where gross margin, rather than operating margin, is the reliable indicator — for example, for an entity that performs limited marketing functions relative to its cost base; Return on Assets fits a capital-intensive operation whose profitability tracks the assets it deploys.

Because the appropriate PLI depends on the specific functional profile of the tested party — not on a fixed rule — a benchmarking exercise benefits from being able to test more than one indicator against the same comparable set before settling on the one that best reflects the underlying economics. A platform that calculates Operating Margin, Net Cost Plus, Berry Ratio, Return on Assets, Return on Capital Employed, and Gross Margin from the same underlying dataset lets the analyst make that comparison directly, rather than re-running the entire search under a different data source for each ratio under consideration.

Traceability: A Number Without a Source Is Not Evidence

Perhaps the most underrated feature of a good benchmarking database, from an audit-defense perspective, is traceability. An arm's-length range built from figures that cannot be linked back to the original financial statement or filing is, in practice, difficult to defend — a tax authority examiner reviewing the study is entitled to ask where a specific margin figure came from, and "the database said so" is not an adequate answer.

Direct links from each comparable's financial data back to its source annual report turn the benchmarking file from a set of numbers into a set of documented evidence. It also matters for internal quality control: an analyst reviewing a benchmarking study prepared a year earlier — or a colleague picking up the file for the first time — should be able to verify any individual data point without having to re-run the original search from scratch.

From Search Result to Local File: Closing the Loop

The final practical consideration is what happens after the comparable set and the range have been finalized. A benchmarking study that lives only as a set of unstructured spreadsheet exports creates extra work — and extra risk of transcription errors — when it has to be converted into the appendix format expected in a Local File. A platform that can generate the full benchmarking study appendix directly from the accepted comparable set, in a format ready for the documentation package, removes a step where errors typically creep in and keeps the audit trail intact from the initial search criteria through to the final filed document.

Technology Handles the Mechanics — Judgment Still Decides the Outcome

None of this replaces the analytical work described in earlier articles on this blog: defining the tested party precisely, performing a genuine functional and risk analysis, reading business descriptions rather than trusting filters, and documenting every acceptance and rejection with an economic — not a convenient — rationale. A database, however comprehensive, does not make the comparability judgment. It gives the analyst the raw material to make that judgment properly, instead of working with an incomplete or unrepresentative population from the outset.

That combination — broad, standardized data across both public and private companies, flexible industry classification, automated pre-screening, multiple PLI options, full source traceability, and export-ready documentation — is the practical reason a growing number of tax advisors and in-house transfer pricing teams have shifted their benchmarking workflow onto a platform like smartZebra rather than assembling comparable sets manually from scattered filings. It doesn't shorten the qualitative review that a defensible study requires. It makes sure that review is being applied to the right starting population, with every number behind it fully traceable.

Further Reading

Readers who want the full methodology behind this article — the complete eight-step benchmarking workflow, worked case studies with real search criteria and documented rejection logs, and a chapter dedicated specifically to comparability analysis and tax disputes — can find it in The Transfer Pricing Workbook: The Practical Guide to Arm's Length Compliance and Audit Defense, available on Amazon. The book walks through the same database-to-documentation workflow described here, applied to full numerical examples across CUP, RPM, Cost Plus, TNMM, and Profit Split:

The Transfer Pricing Workbook: The Practical Guide to Arm’s Length Compliance and Audit Defense: Grubelashvili, Kakhaber, Lonnemann, Gustav: 9798187061198: Amazon.com: Books

A benchmarking study is only as strong as the population it was built from and the reasoning that narrowed it down. Get the data right, and the statistics take care of themselves. Get it wrong, and no interquartile calculation will fix it.

Related pages

Questions & Answers

Why do database searches return "no comparables found" when comparable companies clearly exist?

Because most searches are run against listed-company filings. The overwhelming majority of routine distributors, contract manufacturers and service providers - the best TNMM comparables - are privately held. A database limited to public filings works with a small and often unrepresentative slice of the market, which is why widening the geography feels like the only fix.

Is it defensible to widen the geographic search to a whole region?

It is a recognised step, but it trades one comparability problem for another: competitive intensity, regulatory environment, labour costs and market maturity may no longer resemble the tested party's market. A deeper initial population in the right geography is preferable to a shallow one stretched across a continent.

Does it matter whether I search by NACE, SIC or NAICS?

Yes. The systems are not interchangeable, and a code that captures the right activity in one can be too broad or simply miss relevant peers in another. Cross-referencing several classification systems keeps a coding quirk from distorting the initial population before qualitative review begins.

Which profit level indicator should I use?

It depends on the tested party's functional profile, not on a fixed rule: operating margin for a distributor driven by sales, net cost plus for a routine service provider or contract manufacturer, the Berry ratio where gross margin is the reliable indicator, return on assets for a capital-intensive operation. Testing several indicators against the same comparable set before choosing is the practical way to make that judgement.

Why does traceability matter so much in an audit?

Because an examiner is entitled to ask where a specific margin figure came from, and "the database said so" is not an adequate answer. Direct links from each comparable's financials back to the source annual report turn a set of numbers into documented evidence - and let a colleague verify any single data point a year later without re-running the search.

Does a better database reduce the amount of qualitative review required?

No. Automated screens remove candidates that never had a realistic chance - persistent loss-makers, related-party ownership, R&D intensity inconsistent with a routine profile - so that professional judgement is spent on the genuinely borderline cases. The qualitative review itself, and the documented economic reasoning behind each acceptance and rejection, is unchanged.

Unlock Your 5 Days 100% Access.

Experience the power of the smartZebra engine risk-free. See how fast you can build a defensible peer group or calculate a compliant WACC.

Full platform access
No hidden fees
No credit card required