NEWSLETTER

By clicking submit, you agree to share your email address with TFN to receive marketing, updates, and other emails from the site owner. Use the unsubscribe link in the emails to opt out at any time.

The AI companies that will survive the next shakeout all have one thing in common and it’s not the model

Tom Henriksson, General Partner at OpenOcean
Image credits: OpenOcean

In the last 12 months, I have watched several AI startups I know well see their core products become features inside a new model release. Their valuations have not recovered. All three had technically impressive models. None of them had defensible data. That is the pattern I keep seeing – and it is about to accelerate.

The uncomfortable truth for most AI founders right now is this: a model is not a moat. Compute can be rented from AWS. Distribution can be borrowed from an app store. But the data that makes your model genuinely better than anything a competitor can replicate overnight – that is the one thing that cannot, for the most part, be bought.

For AI providers – not to mention their prospective investors – the question is: how can they stay ahead in an ever more saturated market?

As TFN has reported, 2026 is the year the market pivots from AI experimentation to demanding tangible value – and the companies that survive that pivot will be the ones that made themselves structurally irreplaceable.

Some are betting that raw compute will improve LLMs to the point of AGI. Others are building World Models that simulate the physical environment. But whether your vision for AI amounts to narrow efficiency gains, a logistics overhaul, or intelligent thinking machines, one thing remains true: data is the ultimate moat.

Why the model is not enough

AI users do not pay for accuracy metrics. They pay for improved profit margins, lower risk, and the certainty of staying compliant. Proprietary, refined data is what makes those outcomes repeatable and defensible.

Consider what happens when a better model is released. If your product is essentially a wrapper around GPT-4, GPT-5 does not improve your product – it replaces it. But if your product is trained on five years of anonymised clinical decision data from a network of NHS trusts, or ten years of customs clearance decisions from European ports, no new model release changes that position overnight. The data pipeline is what separates a product from a feature.

This also creates a compounding feedback loop that pure-play model companies cannot replicate: better models lead to better outcomes, which generate more usage, which generates more proprietary data, which further improves the model. OpenOcean portfolio companies that have built this flywheel are growing ARR materially faster than peers that optimised for model performance without building the underlying data layer.

The counterargument and why I disagree

The strongest opposing view is that compute is the real moat. Ilya Sutskever’s position – essentially, that sufficiently scaled models will generalise well enough to render proprietary data irrelevant – is not without merit. If GPT-6 can absorb any domain in context, why spend years building a specialised dataset?

My answer is that general models are excellent at general tasks and mediocre at specific ones. A large language model trained on the public internet does not understand the idiosyncratic risk profile of a specific insurance portfolio, the tacit knowledge of a 30-year bond trader, or the regulatory nuance of a cross-border customs declaration.

That knowledge is not on the internet. It lives in the workflows, decisions, and edge cases of specific businesses and capturing it systematically is the work that creates a durable competitive advantage. The companies that are doing this quietly, right now, are the ones worth watching.

The AI-native playbook in practice

The most successful pattern I see is straightforward, but most companies miss it. As Europe’s fastest-growing AI agent companies are demonstrating, the companies that build durable positions tend to start life as vertical SaaS tools – deliberately, because that is where the proprietary dataset gets built. As they accumulate data, they begin offering full-stack services, and that is where the compounding begins.

Companies that acquire AI-native services or try to build one from scratch consistently underperform this path, not because acquisition or greenfield building is wrong in principle, but because they lack the data. I have seen this play out specifically in property management, clinical services, KYC, and customs: all sectors where repetitive, human-centric work creates rich proprietary datasets over time. In KYC, for example, you cannot remove the human from the compliance loop. But a platform that has processed five million KYC decisions has training data that no new entrant can replicate in twelve months.

The early days of the AI boom rewarded companies that added automation to existing processes and called it transformation. That era is ending. The $52 billion AI agents market is being won by companies that redesign workflows entirely, where AI and humans operate in tandem rather than in sequence. The digital coworker model is not a metaphor. It is a specific architectural choice about who holds the data from each interaction.

What this means practically for founders

Three questions I ask every AI founder we evaluate at OpenOcean:

First: Is your data proprietary, or could a competitor with a better model and the same API access replicate your product in six months? If the answer is yes, you have a feature, not a company.

Second: Does your product generate more valuable data with every transaction? The compounding feedback loop: usage generates data, data improves the model, a better model generates more usage, which is the clearest signal of a durable business. If your product does not get better as it scales, it will be commoditised.

Third: Are you capturing tacit knowledge, the institutional experience and edge cases that exist in workflows, not databases? This is the hardest and most valuable form of data to acquire. It cannot be scraped. It has to be earned through deployment.

Data as a north star and a prediction

It is easy to frame data as a defensive moat, a wall you build to keep competitors out. That framing is too static. Data is a living resource. It compounds, decays, and requires active maintenance. Companies that treat it as a one-time asset rather than an ongoing programme will find their advantage eroding faster than they expect as model capabilities improve.

The AI funding boom is real, three-quarters of UK VC raised in Q1 2026 went to AI startups but capital alone does not create defensibility. The companies receiving the largest cheques are not necessarily the ones building the most durable positions. The ones quietly accumulating proprietary datasets in unglamorous verticals often are.

My prediction: by 2028, the AI companies commanding the highest valuations will not be the ones that built the best base models. They will be the ones that made it structurally impossible for a better model to replace them, because the data their product runs on cannot be replicated by any competitor, at any price. The window to build that position is open now. In most verticals, it will not stay open much longer.

Tom Henriksson is General Partner at OpenOcean, a venture capital firm investing in data-driven B2B software companies across Europe and North America. Portfolio companies include DataSnipper, ChannelEngine, and Crayon.

Total
0
Shares
Related Posts
Total
0
Share

This is premium content.

Get full access for £9.99/month and join our insider community.

Get daily funding news briefings in the tech world delivered right to your inbox.

Enter Your Email
join our newsletter. thank you
TFN Banner