← GO BACK

September 8, 2026

Legacy infrastructure, expensive data

Autor:
Bavest
Markets

A major European asset manager estimates that the cost of centrally procured market data rose by approximately 125 percent between 2018 and 2024. The common explanation is that providers are raising their prices. That is an oversimplification. The real reason runs deeper, rooted in an infrastructure whose technical and contractual logic dates back to an era when a person sat in front of a terminal to read quotes.

Executive Summary
  • The high costs are not just a result of list prices, but of an architecture that offloads connectivity, normalization, and license management onto the client.
  • License models count terminals, workstations, and legal entities. This logic was never designed for machine access or agents.
  • Because every established provider uses its own counting method, a clean comparison of quotes is virtually impossible. Without comparability, there is no competition.
  • New technical providers and the European Consolidated Tape are creating substitutes and reference prices for the first time. This is exactly what drives prices.
  • The most effective lever within your own organization is not the discount, but the question of how much integration work a provider leaves for you to do.
The bill is bigger than the license price

In 2024, the global financial industry spent nearly 50 billion US dollars on market data. This figure only covers licenses. What a firm actually pays is spread across three levels.

The first level is the license price. It is stated in the contract and is the only component that gets negotiated. The second level is integration. Feeds arrive in proprietary formats, with their own identifiers, their own time logic, and their own exceptions. Every firm builds the same normalization, the same mapping, and the same reconciliation routine over and over again. This is standard work; it creates no competitive advantage, and in many firms, it is the larger cost block.

The third level is the management of the licenses themselves. Usage classes must be kept technically separate, users counted, records maintained, and audits supported. This, too, ties up personnel without creating a product. Anyone discussing data costs while only considering the first level is negotiating the smaller part of the bill.

Why architecture is the cost driver

The industry's data layer has grown over decades, mostly through acquisitions. The result is a landscape of parallel systems with different data models that look like a single product from the outside. Three consequences of this drive costs.

Delivery is designed for terminals and feeds. Both were designed before data processing in distributed systems became standard. A terminal serves a human. A feed delivers a stream that the recipient must make usable themselves. Both forms shift work to the client.

License logic counts people. Display licenses are billed per user and scale with the number of workstations or end customers. Non-display usage—that is, machine-based processing—forms its own class with its own billing logic, usually as a flat fee per application or instance. Automation, models, and agents do not fit neatly into either counting method. Classification becomes a matter of interpretation and is decided retroactively during an audit.

Metrics are not comparable. User definitions, location concepts, legal entities, usage classes, and redistribution rights differ from provider to provider. Comparing one quote against another means translating two sets of contracts line by line. In most firms, this work does not happen. Consequently, negotiations shift from price-per-performance to a discount on a list price that no one can independently verify.

Why this has persisted for so long

In the market data business of trading venues, margins of around 75 percent are standard. Such margins do not persist in markets with functioning competition. They persist when three conditions coincide: the product is created as a by-product and costs almost nothing to replicate, there is no equivalent alternative source, and switching is expensive.

The last point is the decisive one. Switching costs here are not the price of a new contract, but the renewed integration effort. This is precisely why an architecture that is difficult to connect preserves the pricing power of its operator. In this market, complexity is not an oversight; it is a business model.

What is driving prices now

Two developments are intertwined.

The first is regulatory. In July 2026, ESMA authorized a Consolidated Tape Provider for equities and ETFs, with the operational launch following on October 1, 2026, initially for five years under ESMA supervision. The framework stems from the MiFIR review and provides for separate tapes for equities and ETFs, bonds, and derivatives. The tape does not replace license agreements. However, it creates something that did not exist before: a common reference against which prices and coverage can be verified. To this end, industry associations are calling for uniform licensing standards, public price lists, and royalty-free international identification numbers.

The second is technical and has a faster impact. When the same data arrives via an API that a developer can integrate in a single day, rather than via a feed that requires a full project, switching costs drop. When switching costs drop, an alternative emerges. If an alternative exists, the price becomes negotiable again. In this market, innovation does not lower costs because new providers offer cheaper rates, but because it breaks the lock-in from which pricing power originates.

We are building Bavest for this very reason as a European financial data infrastructure: not as another feed alongside existing ones, but as a replacement for the layer that every firm currently operates itself.

What a modern data infrastructure looks like

Anyone who takes the three cost levels seriously will arrive at a different requirement profile for a data provider. Four points are decisive. We have built Bavest precisely along these four points, so here is what that means in practice for each one.

Access via an API instead of terminal licenses and feeds. The difference is not the transmission method, but the question of who performs the data processing. When data arrives normalized, the very work that every firm currently performs individually is eliminated. At Bavest, access is provided via a REST API and an MCP server with 81 endpoints in under 150 milliseconds. A developer can integrate the first query in a day, not a project.

A single data model for all asset classes. Prices, fundamental data, ETF data, estimates, and macro data with one identifier logic instead of five sources with five symbologies. Mapping between sources is the silent, constant cost block in almost every firm. Bavest delivers these classes from one model, including macro data, instead of distributing them via separate contracts and separate schemas.

A licensing logic that treats machine usage as the standard. One contract instead of individual licenses per source, location, and usage class, with a pricing logic that can be compared with other offers without requiring legal translation. This is exactly how our model is structured: those who automate, run models, or deploy agents do not fall into a special category that is retroactively reclassified during an audit.

Analytics at the same level as your data. Running portfolio analysis where the data resides eliminates the need for an additional environment that requires connection, licensing, and synchronization. With us, QuantOS handles this: allocation, risk, and performance calculations run directly on the same data set that provides the raw data.

The point isn't that such a provider necessarily has the lowest list price. The point is that they lower switching costs. And switching costs are what have protected the pricing power of established providers until now.

What firms should do now
  1. Make the true costs visible. Report licenses, integration efforts, and license management separately. Only then will you see where the money is really going.
  2. Evaluate list price and integration effort together. The license fee remains the hardest point of negotiation. In addition, every request for proposal should include the question of how many person-days are required until productive use, and how much of that is incurred again with every update. Only by combining both figures do you get the actual price.
  3. Separate usage classes clearly, both technically and contractually. Distinguish between display, non-display, and derived data. Any gray area will be interpreted to the user's disadvantage during an audit.
  4. Connect a secondary source productively before the contract expires. An alternative only carries weight in negotiations if it is already technically operational.
  5. Negotiate definitions instead of discounts. A discount on an incorrectly defined metric just preserves the problem for the next contract term.
Frequently asked questions
Why are market data so expensive?

Because the visible license price is only part of the bill. On top of that, there is the connection and normalization of proprietary feeds, as well as the ongoing management of usage classes and compliance reporting. This work is repeated in every firm and creates no competitive advantage.

Why is the tech stack of established market data providers outdated?

It stems from an era when people sat in front of terminals, and it has grown primarily through acquisitions ever since. In practical terms, this means systems from different decades run side-by-side, connected via interfaces rather than being integrated. Data is delivered in proprietary feed protocols that the recipient must normalize themselves. Many holdings are still updated in batch runs rather than continuously. Each asset class brings its own data model and symbology. And to this day, licensing structures count workstations and legal entities, not queries or applications. This architecture was never built for machine access, models, or agents, which is why it is being retrofitted with special classes and add-on contracts. This is why modern projects fail against these providers not because of price, but because of integration.

Why is it impossible to compare offers from different data providers?

Because every provider uses its own licensing metrics, such as user counts, applications, legal entities, or instances. Without a complete translation of both contracts, you end up comparing prices for differently defined services.

What does non-display mean and how is it billed?

Non-display refers to market data that is not displayed to a human user but is processed by machines, for example in valuation, risk management, or order routing. This category is usually billed as a flat fee per application or instance, regardless of the number of users. Display licenses, on the other hand, are billed per user, scale with every additional workstation or end customer, and are usually significantly more expensive.

Is it worth switching from an established provider to a modern data infrastructure?

It is worth it if the new provider takes on the integration work instead of offloading it. If you simply swap one feed for another, the effort is merely shifted, not eliminated. However, if the data arrives normalized via an API, from a single data model, and under a licensing logic that includes machine usage, both switching costs and ongoing operational costs decrease. At Bavest, this is exactly our benchmark: time to the first productive query rather than time to a signed contract.

Conclusion

Market data is not expensive because the data itself is costly. It is expensive because the layer through which it is delivered is technically and contractually outdated, offloading a significant portion of the work onto the customer. This architecture generates high costs while simultaneously protecting the position of those who operate it. Regulation addresses the transparency gap, but it works slowly. Technical competition works faster: as soon as data arrives normalized via modern interfaces with a licensing logic that treats machine usage as the standard, switching costs drop and prices follow. This is the infrastructure we are building at Bavest in Frankfurt.

Sources

blog

More articles