I wrote more than 220 posts in six years. Looking back, there are eight I’m really proud of. I believe everyone of you should read those (like now!), they hold true now more then ever, and I simply like them and the thinking that’s contained.
Thing is, 7 of those 8 are from before 2022. And I just spent 2 years writing about AI. I think I got distracted.
To me AI looks increasingly more like just another technology problem with a known trajectory. Data however looks like a problem homo sapiens have been wrestling with for 100,00s of years.
The fact that we have more data isn’t necessarily good news. In my view of the world, it’s actually bad news. The world has become more complex much faster than we’re getting better at understanding it. And data is simply the visible exhaust from that invisible process.
So let’s talk about what I learned about what’s even more important today than when I wrote the initial thoughts down first:
AI is a tool that we’re very comfortable using to get these great entertainment and technology outcomes.” (Elizabeth Stone, Chief Product & Technology Officer at Netflix - What I’m reading here is that Netflix likely doesn’t really have an “AI strategy” either. AI is simply the newest technology they have to use at the cutting edge to keep solving the same underlying business and data problems.)
At least read this if you only have 5 minutes, here’s the thesis: Do make continuous work on your data strategy the core of your company. Don’t make an “AI strategy” the center of your company. The gap between the complexity of the world and our ability to understand it keeps growing, and data strategy is how you fight that gap and gain a serious competitive advantage (because no one else is really doing it).
These are the six things I’ve found to be fundamental for just that:
DO focus on distributing data. DON’T confuse data strategy with collecting, storing, or processing more of it. It’s not called “data warehouse/context/lake/… strategy” for a reason. Get the right information, at the right abstraction, to the right human or machine, at the right moment.
DO build around primitives that preserve optionality. DON’T optimize your whole data system around today’s finished answers. When the business or environment changes, you should be able to produce the newly relevant information fast.
DO treat data strategy as continuous. DON’T write a 6-month roadmap and call it done. The world changes, your company changes, and yesterday’s useful data will suddenly become less useful.
DO design for your company’s unique data snowflake. DON’T expect best practices to solve the important parts. Your combination of data sources × use cases becomes unique very quickly, and the most valuable parts are the least standardized.
DO deliberately invest in the hard, uncertain data opportunities. DON’T optimize only for what is easy and already solved. Companies systematically ignore some of their highest-value data opportunities precisely because they are expensive, uncertain and difficult.
DO find and empower data leaders everywhere. DON’T assume they sit inside the data department. There is no universal “data expert”; the people who create value with data will sit anywhere in the company.
FWIW, I’ve developed a quick prototype of what AI-native blogging actually should look like, if you fancy that, you can also read the article here.
Now, I’m calling these six things “fundamental” pretty deliberately.
I don’t mean they’re best practices. Quite the opposite: after 220 posts, a book, a startup and more than a decade working with data, I’ve become increasingly suspicious that there even are many useful best practices for data strategy.
I mean that each of these seems to fall out of something more basic about the nature of data itself. If I’m right about the underlying mechanism, they should still matter when today’s architectures, tools and AI hype are long gone.
And looking back, these dynamics aren’t new: they were already hiding in my writing six years ago and have kept reappearing ever since. That’s a big part of why I now think they’re fundamental.
So let’s start with the mechanism underneath almost everything else.
(1)The real data problem is distribution. Nothing else.
Let me explain why I think “data” is unsolved and urgent. Data is just a special kind of information. If data has no informational value, it’s noise.
I like to define it through it’s one core capability: Data is information we can copy WITHOUT ANY loss of information.
Now what is the “data problem?” the true data problem? It’s a simple “How do we distribute data to the places that need it.”
Doesn’t sound exciting? Nope. Lots of things might sound more appealing (IMHO usually, because they are EASIER, but they almost never are the bottle neck) like:
How do we generate/capture more data?
How do we process more data?
How do we find more data for AI?
With distribution I don’t mean moving data from A to B, we’re pretty amazing at that (progress here is indeed close to exponential), it’s about
the right piece of information (captured by data)
at the right abstraction
to the right human or machine at the right moment
So why do I think this is a fundamental human problem that’s only getting bigger and bigger?
Because one side of this problem is basically fixed: us, our brain. It’s been pretty static for 100,00 years and can only make so many decisions. (Yes we got agents, but they are not human, this constraint on homo sapiens might only change with BMIs and an upgrade of our species).
The other side has changed A LOT. We went from living in relatively small groups, where most relevant things were close enough to understand, to eight billion people, hundreds of millions of companies, global supply chains, financial markets, software talking to software, sensors, APIs and every stupid object in your house emitting data. The world got exponentially more complex, and data is basically the measurable exhaust of that complexity.
Let’s look at an example of this: Suppose a dude named Sven runs a business called MemeAlchemist, a simple app letting paying users create Memes. When he has 50 users and one feature, Sven can basically understand his whole company by looking at it. Now give him 500,000 users, payments, text refinement, multiple image models, enterprise customers and a marketing team. Suddenly the number of things Sven could need to know explodes. Sven’s brain didn’t get any bigger. So he needs data to compress all of that complexity back down into something he can actually act on.
You might think we humans would deal with this problem. But it turns out, the historical dynamic is rather worrisome: if you look at the few dozen technologies we call “general purpose technologies” we have multiple ones that deal with the problem of distribution information, writing, printing, the internet. For comparison, only a handful of general purpose technologies are directly aimed at another fairly important problem: feeding humanity. To me this says, this is a true problem of our species, and just like feeding humanity, it’s pretty much unsolved and getting bigger.
Worse, while those three technologies offered short jumps in our understanding capabilities, every single one general purpose technology drove up the complexity of this world. One might say that technological progress is also the source of complexity progress. And nothing limited can beat exponential growth.
⇒ complexity keeps compounding, while our ability to distribute the right information mostly improves in jumps.
So what does this mean for data strategy? Mostly this: Lesson 1: data strategy matters. A LOT. It mattered three years ago, it mattered before ChatGPT, and as this gap widens it becomes more important, not less. Amazon, Google, Netflix and Airbnb didn’t become what they are because they had nice data warehouses. Not because they had a great “internet strategy”. They became what they are because they got extremely good at distributing data into products and decisions. And yet inside most companies this fundamental problem is still buried underneath whatever happens to be “urgent” right now. (It also means “small data” is absolutely nuts, want to know why? Ask here.)
(2) Every data piece you use trades usefulness now for optionality later.
Data is basically infinite, and the resources that then do stuff with it are very much limited. Like our brain, or any piece of software. So the key challenge is distributing right data to right targets. BUT and there’s a huge BUT, this is not “set it up once and forget about it” mechanism. In fact the moment you’ve set it up (or a system has set it up) it basically is out of date. Because the way we do this is by abstracting, by reducing 1000 pieces of information into the 1-2 that are relevant to that one use case.
Think about it, we basically reduce a whole person and his life into 10 fields in a CRM. That’s a pretty big reduction mechanism. That data, data in CRMs, documents inside of your company (the “knowledge of your company” or the “knowledge of the world” we train LLMs on), the data warehouse tables, those are all examples of this kind of reduced data, of so called “finished data.”
This is a true paradox (in the sense that it is unsolvable and you’ll have to make a trade off),
to deal with increasing complexity we MUST build a system that reduces data;
But that same system (our brains do the same btw. by forming neural connections!) is automatically not good at handling unexpected stuff (like increasing complexity)
The trade off we thus get: Every useful piece of data is an abstraction. Every abstraction throws information away. Therefore every abstraction trades usefulness now for optionality later.
Which means if data strategy is all about complexity, then all of that is just “more exhaust” data warehouses, architectures like “the data mesh”… What should a data strategy thus take as its primitive building blocks? NOT finished data, that’s already outdated, as are the systems distributing it.
It’s data that’s as close to that complexity as possible of course! (And “luckily”, this is exponentially growing in availability) This in turn means (to the best of my knowledge) at least three things, and most companies ignore at least 2 of those:
“Raw” event-based data (emphasis on event!). Which is extremely hard to define (thank you all you “event sourcing” patterns, and streaming platforms). I have always followed the physics definition of an event: Stuff can have a state “right now” and an event is simply the change of two states. Which means those two kinds of data “should be equal” but really aren’t. So what you want to focus on is the changes of states, the stuff that is happening.
Decentralized data collected as close to the physical world as possible (or digital, if an agent, a software system does it). I still like the word “edge” because it’s NOT inside a centralized system or a core/lake/database. You want to collect data at the edge of your business.
Multi-modal data. Because all data in most companies still is reduced to text, but our world is not at all text based… It’s voice, it’s images, it’s video (and much more senses we don’t yet have any way of getting into computers like odors, tastes,…). It is ok to reduce an image to a textual description in the process of “finishing data” but your data strategy shouldn’t focus on image to text, it should focus on those images in the first place!
Lesson 2 for data strategy: Your data strategy should start with primitives, not answers. Keep the data as close to reality as practical, and finish it only when you know what question you’re trying to answer.
Optionality means to me, the ability to answer a question I didn’t have yet. Here are a couple of examples of optionality in data strategies:
A feature suddenly matters. You shipped something months ago, but nobody tracked the interactions that count. Now you add tracking and wait weeks before you know whether anyone uses it, let alone whether it moved a number.
An API suddenly matters. You’ve always logged which version customers use, but nobody was watching, because nothing had changed for years. Then usage shifts to an inferior version and sits there for a week before anyone notices.
A product behavior changes. You have years of “session_success” data, but the product changed what “success” means several times, and nobody preserved the underlying events and what they meant at the time. You have the data. You can’t re-finish it for today’s question.
(3) Your data frontier can move backward, so data strategy can never be fixed.
There’s a simple rule of thumb: if I double the complexity of something, I might need 4x as much data to understand it just as well as before.
Why? Let’s look at a simple example, and why this produces a “jagged frontier of data” that constantly wiggles back and forth:
t=1: Sven has a simple app that finds good images given a potential joke/text for a meme. Let’s call it the “MemeAlchemist.” People submit text and get images back. By tracking how often people click “generate another image,” Sven can get a pretty good idea of how good his images are.
t=2: Now let’s change the world. People suddenly aren’t happy anymore with just submitting text. They also want to refine their text and get suggestions.
t=3: Fair enough. Sven implements a “refine text” button and starts tracking how often people click it. Now he gets some understanding of how good the texts he generates are.
Now let’s look at what happens to the data:
At t=1, the data point “3 new images generated by user X in session Y” is pretty easy to understand. Image generation isn’t one-shot; apparently the first images weren’t good enough.
At t=2, however, that suddenly doesn’t hold anymore. Sven might now see only “1 new image generated by user X in session Y,” but the same user suddenly has lots of new sessions because they keep changing or refining their text. The exact same data point means less than it did before. The data has become worth less.
At t=3, Sven gets back on track by collecting another kind of data: whether and how often people refine their text.
At t=4, however, that creates the next problem: Sven now needs to understand the combination of both data points. Did someone generate another image because the image was bad? Because the text was bad? Did refining the text improve the images? To get back to the same simple insight, “my app works fine”, he now has to understand a more complex set of data.
At t=5, and that’s the good part, Sven can always choose to collect more data. Why stop there? He could track users across sessions, add an “export” event, track which image they ultimately choose, and suddenly understand much more again.
The important thing is that all of these are independent actions by the world and by the data user, Sven. The world changes. Sven changes what he measures. Both happen at different points in time, both take effort, and every change can make yesterday’s data more, or less, valuable.
That’s why the frontier is jagged: complexity can make existing data worse, while new data can suddenly make the same world understandable again.
So, what does this mean for data strategy? Lesson 3: To me this has lots of implications, but at the core it means that data strategy, unlike all other “strategies” is continuous and “just-in-time.” In a world where the business and the environment keep changing independently of your data, no data strategy (& the accompanying roadmap) should cover more than 2-3 months.
Note: If you feel that’s too short, then you’re likely either focusing on doing the wrong things or are moving too slow.
Examples on JIT data strategies:
All “data mesh” implementations, year long projects, basically failed (to show any ROI) to my knowledge.
Amazon’s initiative to implement minimal pricing on its platform for sellers however? Implemented in months, and with immediate ROI in terms of platform growth.
(4) Every company becomes a data snowflake almost immediately.
Let’s go back to Svens MemeAlchemist, at t=5 he starts to collect the “export” events, and decides that data is cool, so he implements a simple functionality that evaluates those memes and shares the most fun ones on X.
He also uses the export event tracking to further understand he quality of his creation algorithm. And that’s what I like to call the “data snowflake” because his setup already looks pretty unique (at day 2 basically”), doesn’t? But of course it does, if Sven has a good data strategy and smartly uses his data to improve his business, just like his situation, his data, and his use cases, the combination of all of this will be HIGHLY unique, simply because it’s a multiplicative relationship.
As soon as a company does more than one thing (= day 1), it has to deal with:
A diversity of data sources
A diversity of use cases.
It’s a simple “n*m” relationship, meaning the matrix of (data | use cases) becomes pretty big, but it’s the only one that properly describes the value a company gets from its data.
Now while Svens MemeAlchemist might share 1-2 data sources with one other company picked at random, and also likely 1 of the use cases, and even the respective underlying systems (say Google Analytics, HubSpot,….), this multiplicative relationship usually means, that they at most share 80% in numbers, and more like 20-50% when it comes to value.
In other words, there is a reason why Netflix, Airbnb and Amazon put so much engineering work in developing unique systems and are pushing the boundaries of research; It’s because for them the value isn’t in what’s already out there. They need new use cases, new data sources and new systems, and have spent significant portions of their money on it from day 1 (one of the first hires of Netflix was Stan Lanning, their first data scientist and creator of Cinematch, still for DVDs back then).
This means three things
As soon as you start a company, the complexity of the underlying data systems are multiplying.
Only 80% in terms of numbers will likely be covered by “things others have done before”
Only 20-50% in terms of value are covered by “things others have done”
This combination produces significant costs and a very clear bias most companies fall into, after all we are all human right? Since the “dark side” the unknown value stuff, the stuff that would be expensive to develop yourself, to produce, to deal with likely carries the most value but also significant cost, we will downplay the potential value of it (because it comes with a priori uncontrollable risk).
To put it simply: Companies will tend to systematically ignore the most important applications of data to their business, because they are blinded by the potential (!) costs of extracting that value.
What does that mean for you and your data strategy?
Lesson 4: Best practices and tooling in anything touching the data space (and that includes all that fancy new AI shit) cover the generic, almost per definition, not valuable parts of your data strategy.
Lesson 5: Because you’re likely significantly underinvesting in the valuable part of your data strategy, it makes sense to consciously overspent in those areas you have to invent and build new things, constantly.
(5) There is no “data expert.” And “data engineer” is a much stranger title than it sounds.
There’s this thing I noticed in my past writings. Whenever I find an inspiring “data leader” , she doesn’t actually hold any title at all related to “data”. And I think there’s a fundamental lesson hiding inside of that simple fact.
You know what one of the most frustrating things for that said data expert, say a data engineer (let’s call him Tim) inside a company is? It is when Tim finds a software team, or the marketing team working around him. He then feels worthless because he feels he could’ve added value to the company but didn’t get the chance.
IMHO the problem lies in the core of data itself. Data is that which we can copy with zero loss of information (Definition by me).
Since we can copy data cheaply thanks to the internet and IT and more inventions, then it means any piece of data WILL be copied with almost no friction if there is value in doing so (in using that data). And this is phenomenal, because at close to zero cost and lossless replication it means this produces a lot of value!
Which in turn means one thing: People and machines will simply always and everywhere work with the data they need, and get it if they don’t have it (probably in some messy but “it works” kind of way).
So our said data engineer Tim might have a marketing team breaking some good practice rules and not using his shiny data warehouse, because they can do so veeeery cheaply and add a ton of value in turn; even if the data isn’t properly normalized, even if the revenue is off by 10%.
Tim might not be happy. But the thing is, the underlying dynamic means this will always happen, as it will always add value in most situations.
So let’s return to the actual question: Is Tim a data expert? He might be a data warehouse expert, but he cannot and never can be a data expert, because of the simple nature of data. He will lack the expertise in how the marketing team works with data and so many other teams.
Which means to me: There is not a single “data expert” on this planet.
But there are data leaders! That’s the good news (and this newsletter is for you).
A data leader is someone who is all about consistently getting better information into products, processes or decisions, whether “data” is part of the title or not.
A (true) data leader thus can be:
A company founder who believes in the importance of data so much that one of the first people he hires has the job to make data work (AirBnB, Netflix)
It might be a whole data team that decides to build enabling platforms (TVPs) to help every team at the company build faster with data.
It might be an analyst who pushes his decision makers to use more data
It might be the start up founders of a data company that didn’t decide to become a “data and AI” company.
Lesson 6: “Works in data” is a shitty proxy to finding “data leaders.” So it is the prime job of you (hello dear CEO) to identify and empower the data leaders inside of your company, because no one else will do so.
What you can expect from this newsletter in the coming months
So that’s what ThDPTh is going to be about now: true data strategy, and what we should build to make it possible. I don’t think I have anything useful to say about that every Thursday. Expect roughly one piece a month.
Now I have THREE asks of you.
First: read the stuff I still think is worth reading.
Looking back, all of my best writing seems to orbit the same two questions: How should companies actually do data strategy, and what should the people building data products (& founding start ups) do about that? Out of 220+ posts, these are the ones I’d still recommend to basically every data leader:
Data Mesh Applied, what decentralizing data actually looks like in practice. The data mesh may be dead, but the fundamental principles, and the theoretical solution are not.
4 trends that will disrupt your data strategy, the forces underneath almost everything I’ve written since (close to the ideas mentioned here today)
Organizing data teams: where to make the cut, it’s all about how you want to decentralize.
Data as Code, what happens when you apply “real” software-engineering principles to data.
How to become the next $30 billion data company, the economics of data, and why they end in open source, which again is extremely brutal. Do I still believe it’s all about open source? Short answer yes! Long answer: call me, don’t have time to write it up and a lot has happened.
Platforms, dbt, Apache Iceberg in 2 sentences, it’s a short one but thing is, data is and always will be a platform business, and few people seem to understand true platforms.
Building data platforms / team topologies, following up on the one before, there are clear best practices on how to organize teams around platforms, just not really in the data space.
ChatGPT killed the $100k BI analyst role I was about to hire, the one AI piece I still think contains something fundamental; The fundamental piece is the “proxy” nature of the whole data department.
Second: if you had ANY thought or question while reading this, “but what about open source?”, “so should I become a data engineer?”, “what about small data?”, “what about AI strategy?”, go ask it here: Chatting with the appendix/editor (or “how Sven thinks AI-native blogging should actually look like”).
I can only write one article for 2,000+ readers. I’m trying to change that.
I’ve put the additional thoughts, caveats, examples, arguments and rabbit holes I deliberately kept out of this article into that appendix. Go ask your question there and follow whichever branch is useful to you.
Third: call me.
If you’re building, leading or thinking deeply about any of this, or think I’ve got one of these ideas wrong and want to passionately tell me so, reach out on LinkedIn/mail. I want to hear it, especially if you think one of these six lessons is wrong.







