How to Benchmark Your Foundation Against Its Peers
Most foundation benchmarking compares the wrong things to the wrong peers. How to build a defensible comparison set and read the results without fooling yourself.
Every foundation board eventually asks the same question: how do we compare? It is a reasonable question and it is almost always answered badly — usually with a handful of famous foundations that share nothing with the asker except the word "foundation," or with a national average that includes both a $40 billion global institution and a family fund making four grants a year.
The problem is not a shortage of data. Public IRS filings cover the entire e-filing universe, and the underlying facts are free. The problem is peer selection. A benchmark is only as good as the comparison set, and building a defensible one is genuinely hard: it requires deciding what "like us" means, then finding the organizations that satisfy it across a dataset of hundreds of thousands.
This guide covers how to build that set, which metrics survive comparison, which do not, and how to present the result to a board without overclaiming.
Why sector averages tell you almost nothing
US foundation giving totaled roughly $117.15 billion in 2025, around 19% of the $617.20 billion Americans gave in total, according to Giving USA's 2026 report (Indiana University Lilly Family School of Philanthropy). Those are useful headline facts for context. They are useless as a benchmark.
The reason is distribution shape. Grantmaking is extraordinarily concentrated: Candid counts over 86,000 grantmaking entities, of which around 92% are independent foundations, and a small number of them account for a disproportionate share of dollars. Any mean computed across that population is dominated by the top of the distribution and describes almost no one.
Three-quarters of independent foundations have no paid staff at all (Urban Institute). Comparing an unstaffed family foundation's administrative ratio to a national average that includes fully staffed institutions with program teams, evaluation functions, and communications departments produces a number that is arithmetically correct and analytically worthless.
The fix is not a better average. It is a better cohort.
The four axes that define a real peer set
A defensible comparison set matches on the dimensions that actually drive the metrics you plan to compare.
1. Asset size band. The strongest single predictor of nearly every operational ratio. Compare within a band — for example, $10m–$50m in assets — rather than across the whole population. Size drives staffing, which drives administrative expense, which drives payout composition.
2. Geography. A foundation making grants exclusively in one metro area operates in a different competitive and need environment than a national funder. Local grantmakers should be compared to local grantmakers, and ideally to those working in comparable places.
3. Cause mix. Grant sizes, relationship lengths, and application volumes differ enormously between, say, arts funding and human services. A funder whose portfolio is 70% health should be benchmarked against others with substantially overlapping cause profiles.
4. Structure. Independent, family, corporate, community foundation, and operating foundation are meaningfully different vehicles with different constraints. Community foundations in particular carry donor-advised funds, which changes how their filings read.
Get these four right and the comparison starts to mean something. Get them wrong and you are comparing your foundation to a category, not to peers.
Two ways to build the cohort
There are broadly two methods, and they answer slightly different questions.
Rule-based cohorts apply explicit filters: assets between X and Y, headquartered in state Z, at least 40% of grant dollars to NTEE major group N. The advantage is that they are transparent and easy to defend to a board — you can write the definition on one slide. The disadvantage is brittleness at the edges, and that they miss functionally similar funders that fail one filter.
Similarity-based cohorts compare the actual content of grantmaking — the text of grant purposes and grantee missions — and find funders whose portfolios resemble yours regardless of how they are formally categorized. This catches genuine look-alikes that rule-based filters miss, such as a foundation classified under "human services" whose portfolio is really about housing stability. The disadvantage is that the definition is harder to explain.
In practice, the strongest approach is to build a rule-based cohort for board-facing reporting, and use a similarity-based one to sanity-check it. If a similarity method surfaces ten funders your rules excluded, that is worth understanding before you present anything.
Plinth's public dataset applies the similarity approach across the whole e-filing universe — roughly 205,000 grantmakers and 17.9 million grants for fiscal years 2017 to 2025 — so look-alike sets are computed against every filer rather than a curated subset. You can explore any funder's page free without signing up.
Which metrics survive peer comparison
Not every metric is worth benchmarking. Some are robust; some are so sensitive to accounting choices that comparison is misleading.
| Metric | Benchmarks well? | Why |
|---|---|---|
| Total grant dollars | Yes, within size band | Directly reported, hard to distort |
| Number of grants | Yes | Reveals strategy: few large vs. many small |
| Median grant size | Yes | More robust than mean; resistant to one outlier |
| Grantee concentration (top 5 share) | Yes | Strategy signal, comparable across sizes |
| New-grantee rate | Yes | Share of grants to first-time recipients |
| Relationship length | Yes | Median years of support per grantee |
| Geographic spread | Yes | Counties or states reached |
| Payout percentage | With care | Multi-year only; single years are noisy |
| Administrative expense ratio | Poorly | Staffing models differ too much |
| Investment return | Poorly | Depends on asset mix and valuation timing |
| "Impact" | No | Not comparably measured anywhere |
The pattern is that structural metrics — how many, how big, how long, how concentrated — benchmark well, because they are derived from itemized grant records that every filer reports the same way. Financial ratio metrics benchmark worse, because they depend on accounting choices. And outcome metrics do not benchmark at all, because no two foundations measure them the same way.
The metrics worth adding once you have the network
Once your grants sit inside the full national graph, a second class of benchmark becomes available — one that cannot be computed from your own filing alone because it depends on what every other funder did.
- Sole-funder rate. The share of your grantees for whom you are the only institutional funder. Compared against peers, this tells you whether you are unusually load-bearing.
- Co-funder overlap. How many funders share your grantees, and which. A foundation with very low overlap is either finding genuinely undiscovered organizations or operating in isolation — the distinction matters.
- First-money-in rate. How often your grant preceded a grantee's other institutional funding.
- New-organization appetite. Your share of grants to organizations with little or no prior institutional funding history, versus peers.
- Portfolio distinctiveness. How far your cause mix sits from the cohort median.
These are the benchmarks boards find most useful, because unlike payout percentages they describe behavior rather than accounting. They are also the ones no single-foundation report can produce.
Three benchmarking mistakes that survive into board packets
These recur often enough to be worth naming, because each produces a number that looks authoritative and is not.
Benchmarking against aspiration. A $30m regional foundation compares itself to a set of nationally known institutions because those are the foundations its board admires. Every metric then reads as underperformance, and the natural response — grow, professionalize, expand — may be exactly wrong for a funder whose value is local depth.
Comparing a single year to a multi-year cohort median. Your foundation's 2024 figure against a cohort median computed across 2019–2024 is not a comparison. Either both sides are single-year or both are multi-year.
Treating a percentile as a target. Once a board sees "we are in the 38th percentile for grant size," the conversation drifts toward moving the number rather than asking whether the number should move. Percentiles describe position in a distribution; they carry no information about what position is correct for your strategy.
A fourth, subtler failure is worth flagging: survivorship in the cohort itself. If your peer set is built from foundations that filed in every year of your window, you have quietly excluded funders that spent down, merged, or wound up. For most structural metrics this barely matters. For anything about longevity or growth, it biases the comparison upward.
How to present benchmarks without overclaiming
Benchmarking goes wrong at the presentation stage as often as at the analysis stage. A few disciplines help.
Show the cohort definition on the same page as the result. A percentile is meaningless without knowing the population. If your cohort is 34 foundations, say so.
Prefer medians and ranges to averages. In skewed distributions the median is the honest summary. Showing the interquartile range alongside your position is more informative than a single rank.
Date everything. Filings lag 12 to 24 months. A benchmark presented in 2026 is likely describing fiscal 2023 or 2024. Label it.
Do not rank on virtue. "We are 12th of 34 on payout" invites a board to treat a compliance threshold as a scoreboard. Present position, not league tables.
Resist single-metric conclusions. No individual number tells you whether a strategy is working. The value of benchmarking is in the pattern across metrics, and in the questions it prompts.
Separate observation from inference. "Our median grant is $18,000 against a cohort median of $45,000" is an observation. "We are underfunding our grantees" is an inference that may or may not follow — small grants to many organizations is a legitimate strategy.
Turning a benchmark into a decision
The point of benchmarking is not the report. It is the small number of decisions it should inform:
- Grant size policy. If your median grant is far below cohort and your grantees are the same organizations others fund at higher levels, you may be adding administrative burden without proportionate support.
- Concentration. If your top five grantees take 70% of dollars against a cohort norm of 35%, that is either a deliberate strategy worth stating explicitly or a drift worth examining.
- Pipeline. A very low new-grantee rate relative to peers means your portfolio is closed. That may be right for a relationship-based strategy and wrong for a field-building one.
- Coverage. If need indicators in your stated service area diverge from where your dollars land, that is a question for the board — not an indictment, but a question.
- Staffing. Administrative ratios benchmark poorly, but grants-per-staff-member within a size band is a fair capacity signal.
Each of these is a conversation, not a verdict. Benchmarking earns its keep by making the conversation specific.
Where software fits
Most of the difficulty in benchmarking your own side of the comparison is that internal data is rarely in a shape that supports it. If grant purpose, geography, and grantee identity are captured inconsistently across years and spreadsheets, computing your own median grant size is a research project.
Tools like Plinth capture those fields at the point of decision, so internal benchmarks are queries rather than reconstructions. Portfolio insights computes concentration, cause mix, and relationship length across your live portfolio, and impact reporting turns the same records into board-ready output. For the external half of the comparison, Plinth's public dataset is free to search, with a free API tier for teams that want to pull cohort data directly.
Community foundations, which carry additional complexity from donor-advised funds, may find the community foundation guide a useful starting point.
Where the comparison data comes from
Both halves of a benchmark have a source problem worth planning for.
Your own side is usually the harder one, which surprises people. Peer data is public; your data is in your systems. If grantee identity, purpose, geography, and amount have been captured inconsistently across years — or live in spreadsheets that changed structure twice — computing your own median grant size becomes a reconciliation project before it becomes a benchmark. Foundations that have never done this should budget for cleanup first and comparison second.
The peer side comes from public filings, which are complete for e-filers but lagged. Three practical consequences follow. Your cohort will be described by fiscal years 12 to 24 months old, so your own figures should be drawn from the same years rather than from the current one. Foundations that file late will be missing from the most recent year, which slightly biases recent cohorts toward the well-resourced. And any cohort built by asset size uses filed asset values, which move with markets independently of anything strategic.
None of this makes benchmarking unreliable. It makes it a description of the recent past, which is what it should be used as. Benchmarks answer "what shape is our grantmaking, relative to funders like us?" — a question whose answer changes slowly. They do not answer "how are we doing right now."
Frequently asked questions
How many foundations should be in a peer cohort?
Enough that medians are stable and no single member dominates — typically 20 to 60. Below about 15, one unusual foundation distorts the whole comparison. Above a few hundred, the cohort has usually stopped being peers.
Should we benchmark against foundations we know, or against data?
Both, and they should mostly agree. If the funders you consider peers do not appear in a data-derived cohort, that gap is informative — either your mental model or your cohort definition needs revisiting.
Is payout rate a fair benchmark?
Only across multiple years, and only within a size band. Payout is computed against average investment asset value and includes qualifying administrative expenses, so single-year figures move with markets and timing as much as with intent.
Can we benchmark impact?
Not meaningfully across foundations. No two funders define or measure outcomes the same way, and grantee-reported outcomes are not comparable across portfolios. Benchmark structural and behavioral metrics; assess impact internally against your own theory of change.
How current is public benchmark data?
Public filings lag 12 to 24 months behind the fiscal year they describe. This is fine for structural benchmarking, which changes slowly, and poor for anything you need to be current.
Do community foundations benchmark against private foundations?
Generally not directly. Donor-advised funds change how grants appear in filings, and the governance model differs. Compare community foundations to other community foundations of similar asset size and geography.
What if our foundation is genuinely unusual?
Then say so, and benchmark on the axes where comparison still holds. A foundation with an unusual mission can still meaningfully compare grant sizes, relationship lengths, and concentration against similarly sized funders in its region.
Recommended next pages
- What Your Form 990-PF Reveals — the data your benchmark is built from
- Foundation Portfolio Analysis — reading your own portfolio before comparing it
- The 5% Payout Rule Explained — why payout benchmarks need care
- Co-Funder Networks — the network metrics no single filing shows
- Align Grants With Strategy — turning analysis into policy
Last updated: August 2026