5 types of data moats: exhaust generators, currencies, give-to-get consortia, aggregators, and proprietary creators
The AI selloff: aggregators, consortia, and research creators fell ~39-46%
Generators maintain defensibility: exhaust generators rose 17%
Note: Newsletter sign up and here are all past newsletters
TL;DR: Data businesses have historically been valued as if the data itself was the moat. AI has made collecting, cleaning, and summarizing data cheap, and the businesses that sold that work (clearinghouses and research firms) are growing revenue the slowest. The data moats that have held up the best so far generate their own data as a byproduct of something customers have to do anyway, like trading on an exchange. We expect the next generation of durable data businesses to look less like ZoomInfo and more like Flatiron: customers come for the tool, and the data accrues to whoever runs it. For us at Konvoy, we are interested in companies whose data is a byproduct of operating something customers depend on.

Which Data Moats Have Held Up With AI?
ZoomInfo started out in 2007 paying researchers to cold-call company switchboards and confirm who worked where. By 2019 it was running at roughly 90% margins and the brute force work done upfront was one of the most cited examples of why data businesses were near impossible to displace (Pivotal). Today though, ZoomInfo’s stock is down 63% over the past two years alone, and its revenue was flat from 2023 to 2025.
ZoomInfo is the most extreme case, but it is not alone. Over the past two years, information services companies (businesses that sell data and research rather than software) were heavily repriced on the fear that AI agents could just rebuild what they sell. Gartner is down 62%, Thomson Reuters 43%, and Verisk 37%.
That said, not every data business fell. We had a hunch that the split was not random, so we looked at which types of data companies held up and which did not. The businesses that generate their own data by running something (like an exchange) held up, while the ones that aggregate data from other sources took the biggest hits.
We have written before about how proprietary deal flow stopped being a moat for venture firms once AI indexed it, and about where value accrues on top of someone else's platform (The Aftermarket Premium). In this piece we apply the same question to data businesses: we lay out the case for and against data moats, test five prominent types of data moats, and consider where new data businesses get built.
The Case For and Against Data Moats
The debate over whether data is defensible has largely run between two positions:
The case for: data compounds. Data businesses start slowly because building a usable dataset is much harder than building a software product, but once they get going, they accelerate and become very hard to displace. Years of historical data and cleaning turn into an asset a new entrant cannot simply buy (Pivotal, databiz).
The case against: incremental data has diminishing returns and data is easier to replicate today. Each additional data point costs more to collect and adds less value, and most data goes stale. Features like curation, coverage, dashboards, and analytics were never really moats on their own (a16z, The Six Moats of Data Businesses). More recently, LLMs have made collecting and cleaning data orders of magnitude cheaper, which removes the slow start that protected incumbents in the first place (Pivotal).
Data on its own is not a moat, but certain structures around data are, mostly ones where more participants make the data more valuable. Switching costs alone are not a sufficient moat, because AI tools and systems provide a big enough jump in efficiency to justify the pain of migrating (Moats Matter Again, Morningstar). We made a version of this argument earlier this year about AI tools for game development: without network effects or proprietary data, the productivity gains accrue to end users and model providers rather than the tools in between.
Five types of data moat come up most often:
Data currencies: a dataset two parties rely on to transact, like a credit score, a bond rating, or an index (S&P Global, Moody's, FICO, MSCI).
Give-to-get consortia: customers contribute their own data in exchange for access to everyone's (Verisk, CCC Intelligent Solutions, Experian).
Clearinghouses and aggregators: businesses that unify fragmented data from many sources (ZoomInfo, FactSet, RELX, Thomson Reuters).
Proprietary creation: research a company produces itself (Gartner, Morningstar, Forrester).
Testing These Moats In The Market
For each of the five types, we took three or four of its representative public companies and tracked two things over the past two years: how the share price has moved, and how revenue has grown.
Here is how the five types have traded over the past two years:

Source: S&P Capital IQ
From best to worst on share price:
Generators: Cboe is up 35% and CME 20% over two years. ICE is down 6%, but for reasons unrelated to its exchange data. This matches revenue: the group grew ~11% a year from 2023 to 2025, and CME's Q2 market data revenue rose 20% to a record $238m, and Cboe’s market data business grew 10% in 2025.
Data currencies: Moody's and MSCI are each down about 6% and S&P Global 25%. FICO is down 66%, but again, mostly for unrelated reasons. Interestingly, though, revenue here is growing as fast as the generators', at 11% to 15% a year, with FICO's latest twelve months up 24%.
Give-to-get consortia: CCC is down 40%, Experian 39%, and Verisk 37%, largely on fears that customers will build their own analytics in-house. Revenue, however, is still growing 5% to 12%.
Clearinghouses and aggregators: ZoomInfo is down 63%, Thomson Reuters 43%, FactSet 40%, and RELX 30%. This is one of only two groups where revenue growth is in the low-to-mid single digits: ZoomInfo's revenue was flat from 2023 to 2025, and the group overall grew ~5% a year.
Proprietary creation: Gartner is down 62%, Morningstar 40%, and Forrester 36%. Revenue is slowing here too: Gartner's growth has fallen to about 1% over the latest twelve months, and Forrester's revenue has shrunk about 9% a year since 2023.

Source: S&P Capital IQ
Most interestingly, the give-to-get consortia (Verisk, CCC, Experian) fell nearly as far as the clearinghouses and research firms, but their revenue is still growing at the rate of the generators and currencies.
The case against data moats was right about features: curation, coverage, dashboards, and analytics are all things that AI agents can do. The two groups that sell the most features of this nature (clearinghouses and proprietary research) are the only ones where revenue has actually slowed. But the five types did not all fall in unison, so certain data moats are presumed to be more durable than others. A few takeaways on these:
Clearinghouses were overrated. Unifying fragmented data was supposed to be one of the strongest positions in data. Working through a long tail of sources cheaply is now something AI agents do incredibly well.
Currencies are more durable, but not safe. Ratings, scores, and indexes are still growing revenue at double digits, but the market still marked them down, with S&P Global off 25% and Moody's and MSCI each down about 6%. The currencies themselves are not at risk, but the revenue from bundled features, services, and aggregator portions of the business are.
The biggest difference is generate versus aggregate. Verisk and CME both sell gated data that an agent cannot scrape, but Verisk pools data its insurance customers contribute, which leaves it exposed in two places: the analytics it sells on top of the pool are now cheap to rebuild (features), and the largest insurers, whose data makes the pool valuable, need it least. The only way to see CME's order flow is to be CME. Owning your data is not enough on its own, though: Gartner produces its own research, but that is analyst synthesis an LLM now does cheaply. What held up was data generated as a byproduct of something customers must do.
Fast-changing data favors the generator. Data going stale hurts anyone selling a static dataset. For a generator, it is an advantage. Every trade creates new data that only the operator can see, so the faster the data changes, the more valuable that position becomes.
Revenue is weakest for businesses that sell the work itself. That means clearinghouses and research firms, whose core product is the collecting, cleaning, and summarizing of data.
How Will New Data Businesses Get Built?
Generators seem to be the most defensible, but they are also the hardest type to build because you have to create and run the underlying business first. A great example is Flatiron Health, which ran an electronic medical record (EMR) system for oncology practices, essentially as a loss leader to get the rights to the data (The Six Moats of Data Businesses). The playbook has interesting parallels to software’s come for the tools, stay for the network. We made a similar argument in 2024 when we said AppLovin should buy Unity: owning the engine would give AppLovin first-party data from every developer building on it.
Takeaways: Data businesses have historically been valued as if the data itself was the moat. AI has made collecting, cleaning, and summarizing data cheap, and the businesses that sold that work (clearinghouses and research firms) are growing revenue the slowest. The data moats that have held up the best so far generate their own data as a byproduct of something customers have to do anyway, like trading on an exchange. We expect the next generation of durable data businesses to look less like ZoomInfo and more like Flatiron: customers come for the tool, and the data accrues to whoever runs it. For us at Konvoy, we are interested in companies whose data is a byproduct of operating something customers depend on.
Have a great weekend,
Josh
