Real estate listing data in the United States flows through about 500 cooperative MLS systems. These are databases where agents share property information so buyers and other agents can see what is available. That infrastructure depends on trust: the data is accurate because professionals maintain it, and comprehensive because participants agree to share.
According to Patrick Pichette, CEO of RealtyFeed, a significant gray market for that data has developed over decades. It is fed by a distribution model designed in the 1990s that gives MLS organizations no technical way to track where their data goes once it leaves their systems.
The Gray Market’s Origins
The core issue is data replication: the practice of technology vendors signing agreements with MLS systems, then copying entire databases multiple times per day. Once that data leaves the MLS environment through replication, the MLS loses visibility into what happens next.
“We know that a lot of real estate data from MLS systems is being scraped and is being resold on the gray market,” Pichette says. “If you want to see a listing from Tampa Bay or San Diego, you can find that information on many websites, even though they don’t have the right to use it.”
The replication model was a reasonable solution when introduced. It allowed agent websites and brokerage platforms to display listings without making live database calls. But the unintended consequence, according to Pichette, is a distribution chain with no endpoint. A vendor that replicates MLS data can share or resell it to other parties, who can share it further. The originating MLS has no technical mechanism to track where it goes or how it is used.
When listing data circulates through uncontrolled channels, it degrades in quality. Pichette points to markets that lack MLS infrastructure. In parts of Europe, for example, the same property can appear on multiple platforms at different prices, or remain listed long after it has sold. The MLS model was built to prevent that fragmentation, but gray market distribution erodes those guarantees.
Flat-Fee Pricing Problem
The gray market is inseparable from how MLS organizations have historically priced data access. Most charge vendors a flat annual fee regardless of how much data those vendors consume. Pichette argues this creates a direct incentive for data leakage from both ends of the consumption spectrum.
“If you’re charging an annual flat fee, typically you’re aiming for the middle, which means those companies that consume a lot of your data are not paying enough, but those that only need little data are priced out,” Pichette says.
High-volume consumers are effectively subsidized, giving them little reason to stay within authorized channels. Smaller vendors who find the flat fee prohibitive relative to their actual needs sometimes turn to gray market sources instead. The pricing model pushes users toward unauthorized data from both directions.
This dynamic has persisted partly because MLS organizations have lacked the technical infrastructure to measure data consumption in real time. Without usage data, usage-based pricing was not feasible. The flat fee became the default not because it was optimal, but because it was the only administratively practical option.
Why This Matters Now
Pichette says the problem is becoming more urgent as artificial intelligence creates new commercial demand for structured real estate data at scale. AI companies and proptech developers are building tools that require comprehensive, accurate listing data, and they are actively seeking sources. If authorized channels are difficult to access or prohibitively priced, gray market data becomes an attractive alternative.
MLS organizations hold exactly the kind of authoritative, professionally maintained data that AI systems most need. But if that data continues to circulate freely through unauthorized channels, MLS systems lose both the revenue opportunity and the leverage that comes with being the recognized source of record.
“Their data is already on the gray market, so they might as well participate so they can at least shape it and have control over it,” Pichette says. “These companies that might be buying data on the gray market, they would rather get it directly from the source. They would rather log in than break in.”
A Path to Oversight
The alternative to replication is a live API model, where vendors query data on demand instead of storing local copies. Each request is logged, so the MLS retains visibility into who is accessing data and how it’s being used. Access replaces ownership transfer, and that distinction is what proponents argue restores the control that replication removed decades ago.
Pichette points to San Diego MLS as an early adopter of this approach, citing gains in both data-related revenue and oversight after roughly a year of use. He frames it as a matter of timing: as AI-driven demand for structured listing data accelerates, MLS organizations that can offer governed, trackable access will be better positioned than those still relying on unauthorized redistribution to reach the market.
About the Expert: Patrick Pichette is CEO of RealtyFeed, a US-based company building data transport infrastructure for MLS organizations through its MLS Router platform, currently working with more than 120 MLSs and associations.
This article is intended for informational purposes only and does not constitute legal, financial, or investment advice. The views and opinions expressed herein reflect those of the individuals quoted and do not represent an endorsement of any company, product, or service mentioned. Readers should conduct their own due diligence and consult qualified professionals before making any investment decisions.
