Plain English
A benchmarking study is the evidence file behind a transfer price. You describe the tested party, decide which independent companies or deals look similar enough to compare against, screen a database down to a workable shortlist, and calculate what margin or return those comparables actually earned. The output is usually a range of results, and the controlled transaction is judged against that range. Done well, it is defensible years later in an audit; done as a box-ticking exercise, it collapses under the first serious question about why a comparable was kept or rejected.
Technical definition
A structured economic analysis, performed under OECD Transfer Pricing Guidelines Chapter III, that identifies potential comparable uncontrolled transactions or companies, applies quantitative and qualitative screens to arrive at a final comparables set, and computes the relevant financial or price indicator to establish an arm's length range against which the tested party's actual result is measured.
Why it matters
It converts a policy statement into evidence. Tax authorities do not accept assertions that a price is arm's length; they expect a search methodology, a rejection log, and a calculation they can replicate. A weak or undocumented study is the single most common reason transfer pricing positions fail on audit.
How it works in practice
- 01Characterise the tested party and the transaction through the functional analysis.
- 02Choose the most appropriate method and the profit level indicator or price metric it requires.
- 03Define quantitative screens (industry codes, independence, size, loss history) and run them against a database.
- 04Apply qualitative review to remove companies that are not functionally comparable despite passing the quantitative screen.
- 05Apply comparability adjustments where reliable data supports them.
- 06Compute the arm's length range and compare the tested party's actual result to it.
Worked example
Distributor margin search
A UK entity buys finished goods from its French parent and resells to UK wholesalers, bearing limited inventory and credit risk. The analyst searches Orbis for European wholesale distributors, applies independence and loss-making screens, and narrows 1,400 companies to 11 after functional review. Their three-year weighted average operating margins run from 1.8% to 5.4%, interquartile range 2.6%-4.1%. The UK entity earned 3.0% — inside the range, so no adjustment is proposed, and the search strategy and rejection log are retained as documentation.
Common mistakes
- Reusing last year's search without checking whether comparables still qualify.
- Accepting a database screen result without reviewing the underlying business descriptions.
- Selecting a profit level indicator that does not match the tested party's risk profile.
- Cherry-picking comparables that produce a favourable range rather than following the screening criteria consistently.
Audit red flags
- A comparables set with fewer than five or six accepted companies and no explanation for the small sample.
- Search criteria that were tightened only after seeing preliminary results.
- No documented rejection reasons for companies excluded at the qualitative stage.
Documentation & data
Documents to hold
- Search strategy memo: database, date, keywords, industry codes used.
- Full and rejected comparables lists with reasons for rejection.
- Financial data extract for each accepted comparable, with source and date pulled.
- Calculation workbook showing the range and the tested party's position in it.
Data you need
- Segmented financials for the tested party.
- Access to a commercial comparables database covering the relevant market.
- Functional analysis narrative to apply qualitative screens.
Who owns this internally: Typically performed by external advisors or an in-house economics team, reviewed and signed off by group tax.
Jurisdiction notes
- European Union
- Pan-European searches using Orbis are standard; local tax authorities increasingly expect local comparables where available, especially in France and Spain.
- United States
- Section 482 practice favours U.S.-only comparables from databases such as Compustat where a domestic market can be shown to be more reliable.
- Asia-Pacific
- Thin local comparables sets often require regional (Asia-Pacific-wide) searches; India has historically required strict domestic comparables.
Notes by role
Advisors & consultants
The rejection log is what survives audit scrutiny, not the final range. Keep it contemporaneous, not reconstructed after the fact.
Students & job seekers
Interviewers often ask you to walk through a search step by step — practise narrating screens and rejections out loud.
Frequently asked
- How often must a benchmarking study be refreshed?
- OECD guidance and most local rules expect the search itself to be refreshed every three years, with financial data updated annually in between.
- Can internal comparables replace a database search?
- Yes, and they are generally preferred where reliable internal comparable transactions with third parties exist, since they require fewer comparability adjustments.
Sources & status
- Primary source
OECD Transfer Pricing Guidelines, Chapter III
OECD, 2022
- Our interpretation
Typical database search workflow
This glossary, 2026
Reference material only, not advice on a specific fact pattern. Reviewed 2026-06-30.
Careers
How this shows up in the job
Benchmarking is the day-to-day work of junior transfer pricing analysts — fluency here is often what gets you hired.
Careers in transfer pricing