Essays on data, work, and personal growth that help you simplify without flattening what matters.

Where Is Data, Really? Applying Business Strategy Frameworks to a Métier in Crisis

Where Is Data, Really? Applying Business Strategy Frameworks to a Métier in Crisis

Friday, September 4, 2026

Between the "AI will kill BI" hype and major platforms each claiming to offer the definitive, unified approach to enterprise data, I wanted to understand what is actually coming next for the field.

Between the "AI will kill BI" hype and major platforms each claiming to offer the definitive, unified approach to enterprise data, I wanted to understand what is actually coming next for the field. 

A popular answer seems to be that AI itself is the redefinition of data and reporting, and it will simply replace the data and BI function as we know it. I learnt a business strategy framework from Alex Smith that analyses how industries mature over time. I think it is a useful lens to analyse where data stands and what comes next.

***

The data ecosystem is more complex, more expensive, and more specialist-dependent than ever, and one might read this as a sign of a maturing field. I think it is the opposite: a sign that the field is still stuck in an early stage of its development, further from delivering simple, accessible, reliable answers to business questions than the current mood implies.

To make this argument precisely, I have borrowed a framework that describes how industries mature over time. Once the framework is in place, I will use it to argue where data stands today, and what the next step in its journey is likely to look like.

***

The Five Strategic Eras

Alex describes how every industry passes through five distinct strategic eras. The era your industry is in determines what kind of strategy you need to win. Play the wrong strategy for your era and you will lose. Even with good execution, solving the wrong problem rarely brings results.

The five eras are:

Supply: the main challenge is simply answering demand with plentiful stock at a reasonable price. The thing is scarce, and making it available is enough.

Performance: supply has been answered, and the market no longer just wants the thing. It wants the best version of the thing. Vendors compete on features, quality, sophistication. Prices rise. Complexity rises with them.

Democratisation: as an antidote to performance, someone walks in and makes it all simple, accessible, and cheap. No more choosing between the expensive and the inadequate.

Specialisation: once the basics are solved, niche variants emerge for specific audiences and verticals. The category fragments into tailored offerings.

Redefinition: someone questions the premise of the category so fundamentally that they create a new one. The new category is similar enough to feed off the existing market, but different enough to start the whole cycle over again.

The hotel industry is a great illustration of how this works in practice. In the 1950s and 60s, the challenge was simply supply; creating professional hotel groups as people became mobile enough to need them. Then came the performance era: Hilton and Marriott made hotels reliably good, and eventually got locked in a features war that produced luxury brands and ever-expanding amenity lists. Democratisation arrived with budget chains like Premier Inn. These chains provide a simple, decent, and affordable alternative and ended the false choice between the expensive and the grim. Specialisation followed: boutique hotels, design hotels, wellness retreats. Marriott today has close to 40 distinct sub-brands. Then came the moment that changed everything: Airbnb. It didn't fix hotels. It asked whether you needed one. That is redefinition: a new category that competes with the original while starting its own cycle from scratch.

***

Before applying this to data, one clarification is necessary. This framework is built around a simple dynamic: a supplier offering something to a buyer. When we apply it to data, there are two valid lenses, and they produce related but different insights.

The first lens is the data tooling industry: vendors in all layers of the chain, Snowflake, Databricks, Azure, AWS, dbt, Tableau etc. selling platforms and tools to companies and data teams. The buyer is the organisation, and the cycle describes how vendor competition evolves over time.

The second lens is the data function inside a company: the data team acting as an internal supplier, with business stakeholders, like the marketing director, the CFO, the product manager, as the buyers. Here the cycle describes how the capability itself matures within an organisation.

Both lenses will be useful throughout this piece, and I will try to signal which one I am using. What is striking is that they tend to mirror each other: when the tooling industry is in its performance era, so is the internal capability, and for the same underlying reasons.

With those terms defined, the argument can be stated precisely: data, in most settings and organisations, is still in its performance era. Democratisation has not happened. And redefinition; the thing many assume AI is already delivering, is further away than most practitioners assume.

That matters because performance eras are expensive by nature: they reward ever more sophisticated tools and ever more specialised people, without ever making the basic problem, i.e. getting a reliable answer to a business question, simpler to solve. Every added layer of sophistication is sold as progress, but it's also a sign that the category hasn't found its Premier Inn moment yet. Below are my arguments for it, and my sense of what is to come.

***

Data Is in Its Performance Era, and the Evidence Is Everywhere

I believe that data is far from "redefinition" and even from specialisation; it is still in the performance era, because democratisation has not happened. It is tempting to think otherwise. We have vertical tools for finance, healthcare, and retail. We have niche practices for analytics engineering, ML engineering, data governance, and data contracts. Sub-disciplines within sub-disciplines. That sounds like specialisation.

But specialisation only makes sense after democratisation has done its work. Democratisation is not just "more people using data tools", it is the moment when the basic capability becomes so simple, accessible, and affordable that it no longer requires specialists to operate. Like the Premier Inn moment for hotels. The customer does not have to choose between expensive-and-powerful or cheap-and-broken, because someone has finally made the thing work for everyone.

For data, through the tooling lens, democratisation would mean that a business user, not a data engineer, can get a reliable, trustworthy answer to a business question without filing a request, waiting for a pipeline, or understanding what a data model is. That has not happened at scale yet. Tools like Hex and various natural language querying products are gestures in this direction. But the moment that makes the previous generation of tools look laughably over-engineered has not arrived yet.

The telltale signature of a performance era is rising prices and rising complexity. Look around the data ecosystem.

AWS, Azure, Snowflake, Databricks; all locked in a features arms race. The Modern Data Stack promised simplicity but delivered power for specialists. Architecture debates like data mesh versus lakehouse versus data fabric, are not signs of a field finding its niches. These are performance-era religions: ever more sophisticated answers to the same underlying problem, without ever solving to the clients' satisfaction.

Through the internal capability lens, the picture is more segmented and interesting.

Large data-native companies, the ones that have invested consistently over a decade, have pushed data access out to the business. A marketing analyst can answer their own questions. A product manager can run their own funnel analysis without creating a ticket for the data engineering team, so the data team is no longer a bottleneck for routine questions. That is democratisation, inside those organisations. Some are going further: each business domain now owns its own data products, its own definitions, its own metrics. Finance has its own layer, logistics has its own, marketing has its own. That is specialisation.

Mid-market companies, by contrast, are firmly in the performance era. They are navigating vendor complexity, debating which stack to standardise on, struggling to build pipelines that don't break.

Small companies are often still in supply. They just want numbers they can trust.

Vertical SaaS tools like compliance analytics for healthcare, risk platforms for finance etc. represent genuine specialisation, but only for the segments that are already past democratisation and ready for it.

I think that the data ecosystem spans stages two through four (performance through specialisation) simultaneously, depending on which segment you examine. But the centre of gravity, where most of the money flows, most of the hiring happens, and most of the unresolved problems sit, is firmly in the performance era. Which means we are two full stages away from the interesting question of what will the redefinition of data look like.

But let's ask the question anyway, because the most useful place to look is always slightly further ahead than where you currently stand.

***

What Will the Airbnb Moment for Data Look Like?

In the hotel example, Airbnb is the hinge point; the moment when someone stopped trying to make hotels better or cheaper and asked whether the frame was right at all. What is the equivalent moment for data?

I have three scenarios for what redefinition might look like, ordered by how well each one can still answer cross-domain and causal questions, from full capability, to partial, to none.

1. The first scenario is the platform collapse. The modern data stack is fragmented because software was built and sold in silos over the past few decades. Your CRM doesn't speak to your ERP. Your logistics platform stores data in formats your finance tool doesn't recognise. You need a team of data engineers to translate between them constantly. Redefinition, in this reading, means that fragmentation simply ends: one coherent platform that connects to well-known business process tools, and where the data model is given from the start. The reconciliation headache disappears because there is nothing left to reconcile.

SAP tried exactly this in the 1990s. And the idea was right. The problem was that making unification painless required capabilities that didn't exist yet. Integration meant hand-coded connectors and rigid schemas. Companies couldn't conform to SAP's data model, so they customised heavily, and the unified vision broke apart through the customisation layer. SAP became exactly what it was trying to replace: a powerful, expensive, consultant-dependent system. What's different now is that AI might finally provide the invisible integration layer that SAP was missing. Inferring mappings, tolerating schema differences, translating between formats without human intervention. The enabling condition that was missing in 1995 may actually be arriving.

2. The second scenario is more radical: it starts from a different question entirely. What are organisations actually trying to accomplish with all this data infrastructure? Let's argue that they are not after a single source of truth; that is just the means. The end they are trying to achieve is confidence in a decision. And those are different things. If an AI agent can pull correct data from five different systems and tell you "enter this new market, the margins are healthy, and here is why", then the entire data engineering reconciliation problem becomes a historical artefact, like hand-typesetting after the laser printer arrived. You stop needing large data teams because the problem got reframed.

Yes, this scenario implies a degree of uncertainty that feels uncomfortable, to some, even unacceptable. But weigh that against the cost of eliminating this uncertainty, i.e. data infrastructure and large data teams. Many organisations might just tolerate this uncertainty, and invest the capital, otherwise dedicated to data, to improve business processes and their decision-making frameworks.

3. The third scenario is the most radical. What if the data layer becomes ambient and disappears into the application itself? Rather than extracting data from systems to analyse it elsewhere, the intelligence lives inside the operational system. Your supply chain tool doesn't produce data for analysis, it simply tells you what to order next. Your CRM doesn't export pipeline data, it tells you which deal to prioritise. The question "what does the dashboard say?" stops making sense, because the data never leaves the system that generated it. So the oncept of a separate analytical layer becomes obsolete, eliminating the discussions of unifying data or accepting uncertainty.

***

The last two scenarios share a common breaking point: causal and cross-domain questions. Consider this: the marketing campaign was launched, and you want to know what the actual impact was, on retention. Is the churn we are seeing in this cohort related to the pricing change or the product degradation from last quarter?

These questions require data from multiple systems to genuinely talk to each other. No amount of tolerance for uncertainty fixes that. You actually need the joins. And no ambient intelligence inside a single application can answer a question that spans three of them.

The platform collapse, one lakehouse to rule them all, seems like the best solution here. But that, we know, is expensive to set up and maintain. And I think there is an increasingly plausible middle path: selective consolidation. Not one platform to rule everything, but a much lighter and more pragmatic version of reconciliation. You consolidate the two or three data domains that genuinely need to talk to each other for your most critical decisions, say, commercial and product data, and you accept directional, loosely coupled signals everywhere else. You stop chasing the dream of a perfect unified model across all systems, and instead ask: which specific cross-domain questions are worth the engineering cost of creating and maintaining consolidated datasets?

This would represent a genuine cultural shift for data teams, who have historically treated incomplete reconciliation as a problem to be solved rather than a trade-off to be managed. But it may be the most realistic version of redefinition of this entire field. Far from the elegant single-platform vision, and from the radical abandonment of data infrastructure, a more sober renegotiation of what good enough looks like, made possible by tools that can do more of the translation work automatically.

***

So, what do you think? Which stage is data at in your organisation? Your thoughts and comments are always welcome.

No comments yet

Join The Simplicity Stack

The unactionable newsletter. For people tired of doing everything.

Search