- Tension: Brands assumed feeding more data to their AI and CDP vendors was pure upside, without asking what the vendor does with that data once it has it, sitting alongside every other brand’s data in the same system.
- Noise: The industry conversation stayed fixated on “first-party data ownership” as the fix, as if owning the record settles the question, while the value created from patterns across records, and the risk of holding them in one place, sits somewhere past ownership entirely.
- The Direct Message: Owning the record isn’t the same as controlling the value it generates. The vendor aggregating your data alongside everyone else’s, and getting breached in your name, is often keeping more than you realize.
To learn more about our editorial approach, explore The Direct Message methodology.
Ninety-eight percent of marketers say they hit at least one barrier to AI-powered personalization, and data issues are the most common culprit, according to Salesforce’s Tenth Edition State of Marketing report, which surveyed roughly 4,450 marketing decision-makers. That number describes an industry that spent the last two years doing exactly what it was told to do: consolidate, centralize, feed more first-party data into every AI and customer data platform that promised better personalization in return. What almost nobody asked, until recently, is what happens to that data, and the patterns learned from it, once it’s inside somebody else’s system.
The first-party data comfort blanket
The shift toward first-party data was a real and necessary correction. As third-party cookies eroded and privacy regulation tightened, brands were told, correctly, that owning a direct relationship with customer data was safer and more durable than renting audiences from ad networks. Most of the industry took that advice.
But as one recent analysis on MarTech put it plainly, ownership doesn’t automatically translate into understanding. Owning a customer record doesn’t tell you whether the email address is still active, whether the identity behind it has moved on, or whether the profile your CDP is personalizing against reflects who that person is now versus who they were eighteen months ago. First-party data ages the moment it’s collected. The comfort of ownership quietly became a substitute for the harder work of keeping that data current, verified, and connected to a real, active person.
What “owning” the data doesn’t actually buy you
That distinction matters more once AI enters the picture, because AI personalization doesn’t run on raw records. It runs on inference, patterns the model has learned about what a given signal tends to predict. And most of the models doing that inferring aren’t proprietary to any single brand. They’re built, trained, and continuously refined by the vendor, using aggregate signal across every brand’s data flowing through that vendor’s platform.
A brand can hold clear legal title to its customer records while the actual predictive intelligence, the thing driving the personalization the brand is paying for, is a byproduct of a much larger pool the brand contributed to but doesn’t control. The record is the brand’s. The pattern recognition built from thousands of brands’ records combined is the vendor’s. That’s not a conspiracy; it’s how the economics of a shared AI platform work. But it’s rarely spelled out in the pitch deck, and it means the vendor is often the only party in the relationship accumulating compounding value across every customer it serves, while each individual brand resets to zero.
The aggregation vendors don’t advertise
BCG projects that $2 trillion in revenue will shift over the next five years to companies that get personalization right, and its Personalization Index research finds that leaders in the category grow revenue 10 percentage points faster annually than laggards. That framing is usually presented as a call to action for brands: move fast, personalize well, capture the upside. It’s worth reading the same projection a different way. If leaders are pulling away by double digits, laggards are subsidizing that gap, often by feeding the same vendor platforms the exact data that helps competitors get sharper. The brand handing over its data is buying a service and, at the same time, contributing to the training signal that makes the vendor’s product better for the next customer, including a direct competitor.
When the data leaves the building and comes back as a breach
The clearest illustration of where the risk actually concentrates arrived in late 2025, when a group calling itself Scattered LAPSUS$ Hunters claimed to have stolen close to a billion records tied to Salesforce customers, including Qantas, GAP, Fujifilm, and Albertsons. Salesforce’s own platform wasn’t directly compromised; the attackers used social engineering, “vishing,” against employees at customer organizations connected through OAuth integrations. But the practical result was the same regardless of where the technical fault sat: customer data that individual brands believed they owned and controlled was exposed at platform scale, because it had been centralized in one vendor’s infrastructure alongside hundreds of other brands’ data.
That’s the sharpest version of the “who keeps the value” question, restated as “who bears the cost.” When data sits inside a shared vendor environment, the brand carries the reputational and legal fallout of a breach it didn’t cause and often couldn’t have prevented, while the vendor, so long as its own platform wasn’t technically breached, can reasonably claim its systems held. Value concentrates upstream. Risk gets distributed downstream, to whichever brand’s name ends up in the headline.
The clean room compromise, and its limits
Data clean rooms exist precisely because the industry has partly recognized this tension. Nearly two-thirds of companies, 64%, already have one in place, matching first-party datasets against a partner’s without either side seeing the other’s raw records, keeping only aggregated, privacy-safe outputs, according to an IAB report summarized by the CDP Institute. It’s a genuine improvement over blind data-sharing. But the same report cautions that clean rooms are expensive to run properly, averaging $376,000 a year plus six or more dedicated staff at half of deploying companies, and up to two years to fully implement, which means the brands most able to protect the value of their data are, again, the largest ones already positioned to capture the disproportionate gains BCG’s research describes. Mid-market brands are more often stuck choosing between the personalization upside of full data sharing and the cost of infrastructure that would let them share more carefully.
None of this means brands should stop feeding data into AI personalization systems, that ship has sailed and the upside is real. But two years into the buildout, the honest question isn’t whether to keep handing data over. It’s what a brand is actually getting to keep, in pattern, in leverage, in protection, once that data has already left the building. Most contracts never answered that question, because until recently, nobody thought to ask it.