In February 2019, Pendo's lead data scientist Suja Thomas published a finding that product teams still quote and still ignore: 80 percent of features in the average software product are rarely or never used.
I have sat in the meetings that produce that number. Nobody plans shelfware. It accumulates one confident opinion at a time: the founder's pet workflow, the integration a big prospect swore would close the deal, the checkbox spotted in a competitor's screenshot. Every one of those decisions felt informed on the day it was made. Almost none was tested against evidence about what the segment actually ships, what buyers verify, or what anyone can prove.
That is the argument of this piece. The 80 percent figure is not an engineering failure. It is a decision-quality failure, and the missing discipline is evidence.
What did Pendo actually measure?
The 2019 Feature Adoption Report aggregated anonymized usage data across 615 Pendo subscriptions, restricted to customers who had used the platform for more than a year, observed over a three-month window. Features in Pendo are tagged by the customer, which matters: these are the parts of the product that teams themselves expected to be used daily, weekly, or monthly. The report then sorted tagged features into four tiers by share of usage volume:
| Usage tier | Definition | Share of features |
|---|---|---|
| Frequent | Generate the top 80% of usage | 12% |
| Moderate | The next 15% of usage | 8% |
| Rare | The last 5% of usage | 56% |
| Never used | No usage recorded at all | 24% |
Then came the price tag. Gartner had forecast public cloud revenue of $175.8 billion for 2018, and publicly traded cloud companies on the Bessemer Emerging Cloud Index spent an average of 21 percent of revenue on R&D, implying roughly $36.9 billion of engineering investment. Eighty percent of that comes to $29.5 billion potentially spent developing features that are rarely or never used. Scaled down, Pendo estimated a $50 million revenue company might be spending $8.4 million on features its customers barely touch.
Treat the dollar figure as directional, not forensic. Pendo sells adoption tooling, the extrapolation is linear, and "up to" is doing real work in that sentence. But the underlying usage distribution came from instrumented products across banking, HR tech, education, logistics, healthcare, and e-commerce, and the report notes it varied only slightly by company size. The shape of the curve is hard to argue with.
Why do opinion-driven roadmaps produce shelfware?
Because human intuition about feature value is worse than anyone in the room believes, mine included.
The strongest data on this comes from Microsoft's experimentation platform team. Reviewing years of controlled experiments, Kohavi, Crook, and Longbotham reported that among well-designed experiments built to improve a key metric, only about one-third succeeded at improving it. These were not random guesses. They were ideas that survived internal review, got specced, got built, and shipped to real users at a company with unusual measurement discipline. Two out of three still failed to move the metric they existed to move. The same paper recounts QualPro's offline testing of 150,000 business improvement ideas over 22 years, which found 75 percent of important business decisions and improvement ideas either had no impact on performance or actively hurt it.
Now remove the experiments. A typical B2B roadmap is prioritized by conviction: the HiPPO (the Microsoft paper's term for the Highest Paid Person's Opinion), the most recent lost deal, the loudest customer, the competitor teardown someone assembled in an afternoon. If ideas that survived Microsoft's vetting failed two times out of three, ideas that survived a heated Tuesday meeting will do worse.
Picture a hypothetical. Sentrix, a mid-market security vendor, loses a deal to CrowdHaven. The seller's write-up says the buyer wanted agentless discovery. Six engineer-months later, Sentrix ships agentless discovery. Nobody checked whether CrowdHaven's agentless claim was substantiated anywhere beyond its own datasheet, whether buyers in the segment verify that capability during procurement, or whether the deal actually turned on price. One anecdote became a roadmap line, and the roadmap line joins next year's 80 percent.
Low usage is not always waste
No, and honesty requires saying so before prescribing anything.
Some features are insurance. A backup restore flow, a breach-notification export, an audit log: measured over a three-month window they will look dormant, and they are still exactly why the deal closed. Some features are procurement keys, present to satisfy a security questionnaire or a compliance regime, exercised once by an auditor and never again. Others run on quarterly or annual cadences that a short observation window mostly misses. Usage analytics grade frequency. They say nothing about criticality.
That cuts both ways, though. If low usage is defensible only when a feature earns its keep some other way (closing deals, passing audits, covering a tail risk), then a team should be able to say, for each dormant feature, which of those jobs it does. Most teams I have worked with cannot. The honest split of the bottom 80 percent is a minority of deliberate insurance and a majority of opinion residue: features built on anecdote that neither users nor buyers ever asked anyone to verify. The problem is not that rarely used features exist. The problem is being unable to tell the two kinds apart.
A roadmap needs evidence when opinion runs out
Evidence about the segment, gathered on a managed schedule and graded honestly.
Internal usage data tells you what customers do with what you already built. It cannot tell you whether the capability a competitor is trumpeting actually exists, whether your segment's buyers verify a claim before paying for it, or whether the gap a seller reported is real. That takes competitive evidence: what rivals ship versus what they merely claim, and what independent sources can substantiate.
OmniAxis enforces that discipline, evidence before engineer-months. The platform grades vendor capability claims by the strength of the independent evidence behind them, and refreshes those grades on a managed schedule, weekly on the top plan, so every profile stays current as the segment moves and carries its last-refresh date. For every tracked vendor, the marketing claim and the evidence score appear side by side. In the hypothetical above, Sentrix would have seen before committing the build that CrowdHaven's agentless-discovery claim graded at the bottom of the scale, with no independent corroboration anywhere. That single screen turns "a competitor has this, so we must too" from a reflex into a testable proposition. Sometimes the claim grades at the top of the scale and the build case is real. Either way the decision rests on graded evidence rather than a seller's memory of a lost deal.
Seven years on from Pendo's report, I see no reason to believe the number has improved, because the inputs to roadmaps have not changed. Teams instrument their own products obsessively and still choose what to build next on anecdote. Flip the discipline: before a feature earns engineer-months, ask what the segment ships, what buyers actually verify, and what the evidence grade of the competing claim is. Build less. Prove more. The alternative is funding the 80 percent, one confident opinion at a time.