First-Party Data Activation Inside Conversational AI Ad Campaigns
Brands must rethink first-party data for AI ad campaigns.

Activating first-party data inside a conversational AI ad campaign is not a smaller version of what brands already do in search or social. The signal being matched is a live prompt, something a user types right now, not a cookie ID or a profile built from last quarter's clicks. That difference forces a rethink of how owned data connects to a moment of intent, and the brands sorting this out now are the ones likely to be positioned once AI chatbot-native ad spend, still small today, actually scales.
What first-party data contains and why quality determines what AI can do with it
Amperity's guide on the subject states that first-party data covers transaction records, CRM entries, website and app behavior, loyalty activity, and email engagement. Collecting first-party data directly does not make it accurate, permissioned, or usable, which is the part brands skip. That's the part brands skip. Consent mapping, identity resolution, and lineage documentation need to happen before the data does anything useful, not after someone asks why a campaign underperformed.
Zero-party data, meaning preferences a customer states outright, gets treated as gold standard by a lot of marketing teams. It's useful, but a stated preference can change by the time it gets used. Second-party data, which is really just another company's first-party data shared through a direct deal, fills gaps when a brand's own signals run thin, though it comes with less control over how clean that data actually is.
The stakes on data quality go up specifically because of AI, not despite it. LiveRamp CEO Scott Howe said at RampUp 2026 that in an AI world, data is power, and those that control the data will win while those that don't will be captive to the model. Public LLMs largely train on the same public data, so proprietary first-party signals are the actual point of difference between one brand's algorithmically driven targeting and a competitor's. Mey Wong, who heads emerging technology at TurboTax, made a related point at the same event: the opportunity isn't piling up more data, it's figuring out which deterministic signals matter and getting them to work together.
Verified transaction data holds up better than modeled or inferred signals when it comes to training AI systems, simply because it reflects what someone actually did rather than what a model guesses they might do. The IAB's State of Data report found 71% of brands, agencies, and publishers are growing their first-party data sets or plan to, nearly double the rate from two years prior. That tells you the industry has accepted the premise. It doesn't mean anyone's figured out activation inside these new AI surfaces yet.
How the matching signal changes when the ad surface is a live conversation
Search and social match a brand's first-party data against inferred segments or keyword lists built from past behavior. The targeting logic guesses at someone's current state of mind using history as a proxy. A conversational AI surface skips the guessing. A user asking an assistant for the best carry-on luggage for business travel under two hundred dollars has just handed over purchase intent, product category, price ceiling, and use case in a single sentence.
That's intent revealed in the moment, not modeled from six months of browsing data. It changes what the targeting system is actually reacting to.
Current research into how these systems might auction ad placement gives a sense of the mechanics. A genre-based bidding approach has advertisers bid on stable, high-level clusters, things like travel, home improvement, or personal finance, rather than on individual queries a user typed. That keeps computation manageable and limits how much sensitive query data ever needs to move around. Google Research has proposed something further out: a token-level auction where advertisers submit models of their brand voice alongside a bid, and that model shapes the LLM's output as it generates, token by token. It's influence over phrasing itself, not a banner ad or a sponsored link. It's influence over phrasing itself.
A paper called LLM-Auction, presented at ICLR 2026, approaches the problem through a novel auction framework rather than classic ad matching. Its framework emphasizes that the surrounding conversation context matters to predicting placement performance, not just any single query in isolation.
None of this means a brand's CRM data gets matched directly to what a user types. Instead, that data trains the advertiser's own model of which intent clusters are worth bidding on, and that model is what shows up in the auction. The question of whether conversation signals might inform targeting on other surfaces remains live across the industry, and it is a design choice each platform makes differently.
The infrastructure layer: what has to be in place before activation is possible
Identity resolution comes first, full stop. CRM records, app data, loyalty activity, and email engagement usually live in separate systems that don't talk to each other. None of it is usable by an AI system downstream until it's unified into one coherent identity.
David Joosten at GrowthLoop describes the shift toward a data lakehouse or composable CDP pattern in "First-Party Data Activation": data needs to be in the enterprise's own cloud rather than locked inside a vendor's platform, so that agentic AI can operate against a unified, accessible foundation. Bolting AI onto five disconnected systems doesn't work.
Consent and permission mapping has to happen at the level of the individual signal: a brand needs to know exactly which data was collected under which permission and for what allowed use. AI systems making allocation decisions need an audit trail, not just a data feed. Alvaro Palacios, writing in AdExchanger in April 2026, frames it directly: AI decision engines built to optimize outcomes need deterministic identity, clean feedback loops, and governable lineage. First-party data isn't a nice-to-have in that setup, it's structurally required. Palacios compares agentic ad allocation to portfolio management rather than day trading, and offers a line worth keeping: first-party identity isn't fuel, it's the ledger that makes allocation possible.
Data clean rooms solve a separate but related problem. They let a brand match its own signals against a publisher's or platform's data without either side handing over raw records, which matters a great deal for conversational AI surfaces built specifically not to pass user data to advertisers. OpenAI's ad setup inside ChatGPT illustrates how constrained the access model is: a brand shapes targeting through the intent categories and audience signals it configures through its demand-side platform, with no direct visibility into individual user data.
Put together, the checklist looks less like a wishlist and more like a set of prerequisites: identity unified across owned channels, consent and lineage documented at the signal level, audience segments built on behavior (purchase frequency, intent category, engagement) rather than demographics alone, brand voice models ready to submit into auction systems rather than just static creative assets, and an integration path to a DSP or exchange with direct supply relationships into AI publishers.
How platform-specific constraints shape what first-party data can and cannot do on each surface
Each AI surface enforces its own rules, and those rules decide how much of a brand's first-party data can actually do anything.
OpenAI launched ads inside ChatGPT in pilot in February 2026. Product cards appear at the bottom of a response, visually set apart and labeled "Sponsored," limited to logged-in adults on the Free and Go tiers in the US. Plus, Pro, Business, Enterprise, and Education tiers stay ad-free. Criteo signed on as the first technology partner. Targeting runs off the topic of the conversation and, when a user opts into ads personalization, past chats and past ad interactions, though advertisers never see that underlying data themselves. Users can turn personalization off. OpenAI calls one of its guiding rules "Answer Independence," meaning ads aren't allowed to bias the actual answer the model gives. Political topics, health, and mental health queries are excluded from targeting, and accounts flagged as belonging to users under 18 see no ads. For a brand, this means there's no way to pass an audience segment directly into OpenAI's system. Activation happens through a demand-side platform's relationship with OpenAI, and the brand's own data shapes which intent categories get bid on, not who gets reached by identity.
Google's AI Mode has begun showing ads across a growing share of AI results. Shopping ads already run inside AI Mode, and checkout integrations with retail partners have launched within it. AI Mode reportedly draws a substantial number of daily users, and advertisers running AI Max for Search campaigns report improved conversion performance. Google's ecosystem allows more conventional first-party matching through tools like Customer Match and enhanced conversions, making it the closest thing to legacy activation mechanics among the current AI surfaces.
Microsoft Copilot has carried ads since 2023, inherited from Bing Chat, and has since expanded into "Compare & Decide" ad formats, shopping campaigns, and a Copilot Checkout feature launched in January 2026. Microsoft has reported that Copilot ad placements show favorable engagement relative to traditional search formats. Because Microsoft Advertising's existing audience tools, including Customer Match and LinkedIn profile data, carry over into Copilot, it stands as one of the more mature first-party on-ramps available in conversational AI today.
Perplexity took the opposite path. It announced in February 2026 that it was winding down advertising, after testing sponsored follow-up questions through 2025 with select brand partners, citing concerns about user trust. Executives left the door open to a return down the line, but the company is now leaning on subscription revenue instead, even with a sizable number of publishers still in its Publisher Program. Its exit shows how fragile user trust is in this format.
Claude is not currently part of an ad-supported model, so it isn't part of this conversation at all, for anyone mapping the landscape.
Because each platform sets its own data access rules, a brand's ability to activate first-party data shifts from surface to surface. Building a coherent strategy across all of them, without rebuilding an integration from scratch for each one, really only works through a DSP that already holds direct publisher relationships across multiple surfaces.
Where first-party data signals connect to conversational intent in practice
Audience segmentation is the input layer that makes everything downstream possible. CRM and CDP clusters built around purchase frequency, category affinity, and lifecycle stage tell a brand which conversational intent categories are worth bidding on. A travel brand's own data showing that people who search flights around long weekends convert at a higher rate translates directly into a genre bid on travel-planning conversations.
Brand voice models function as the creative half of that equation. In token-level or preference-alignment auction systems, an advertiser submits a model of its own messaging alongside the bid, and that model should be built from what a brand's first-party data already shows resonates with its customers, not from generic ad copy written for no one in particular.
Real-time adjustment is where AI earns its keep operationally. Dynamic creative optimization, as described by StackAdapt, tests headline, message, and offer combinations live and serves whichever performs best. Inside a conversational surface, the "creative" being optimized isn't a banner, it's the shape and tone of the brand's influence on the response itself.
GrowthLoop's Compound Marketing Engine points at where this is heading: agentic AI proposing audiences and journeys straight from a brand's own data cloud, activating across channels, and adjusting continuously based on performance. That architecture extends naturally to AI publisher surfaces once the supply relationships exist to support it.
Retail media offers the clearest proof this model can work at scale. US retail media ad spend is projected by eMarketer to hit $69.33 billion in 2026, up from $58.79 billion in 2025, and it works precisely because identity, inventory, and measurement all sit inside the same closed system. Conversational AI advertising needs that same closed loop, and doesn't have it yet.
The New York Times' BrandMatch AI, launched in 2024, is a useful publisher-side example of what's possible when a company controls both the data and the inventory. It matches advertisers to logged-in readers using first-party data, and after a year in market improved both click-through rates and video completion rates by 30%. Conversational AI hasn't caught up to that. Attribution from a purchase influenced by a conversation remains largely unsolved. Activation in this channel is running ahead of measurement, not alongside it.
Governance and privacy constraints that shape what is permissible, not just possible
Amperity's figures show that nineteen US states now enforce comprehensive privacy laws, and that regulatory floor keeps rising even with no federal law in place. Google's decision in April 2025 to keep third-party cookies in Chrome didn't restore the old tracking environment either. Other browsers still restrict cross-site tracking, and platform-level changes keep chipping away at how much signal is even available to begin with.
The bigger constraint on conversational AI appears in structure rather than regulation. Platforms like OpenAI simply will not pass user data to advertisers, full stop. That means first-party data activation has to happen entirely on the advertiser's side, shaping bid strategy and creative models, rather than through identity matching at the moment an ad shows up.
If an environment can't support auditability, clear consent provenance, and controlled data flows, spend either doesn't clear at all or clears only under conservative limits, as Palacios's AdExchanger piece from April 2026 makes concrete for regulated categories like pharma. eMarketer's 2025 forecast for retail media spend growth backs this up, with US healthcare and pharma projected at 21.3% growth against 11.2% for search overall, suggesting environments that can prove their governance are the ones capturing regulated-category budget.
Data clean rooms remain the mechanism best suited to this constraint, since they let signal matching happen without ever exposing raw records, which fits what conversational AI platforms already require. When compliance becomes a capability rather than a cost center, the advantage shifts toward platforms that can prove accountability while still performing. It's becoming a technical requirement for agentic systems to function, not just a compliance department's talking point.
Brand safety compounds the governance challenge. OpenAI excludes political, health, and mental health topics from ad targeting entirely, so a brand needs to map its own segments against those exclusions ahead of time rather than finding out the hard way mid-campaign. And on the user side, ChatGPT labels every ad "Sponsored," lets people hide or report them, and allows personalization to be switched off. Any first-party activation strategy built for this channel needs to assume users will use those controls, not treat them as an edge case.
What brands need to build now to be ready when conversational AI ad scale arrives
None of this requires waiting for conversational AI ad spend to hit some inflection point before acting. The infrastructure work, identity resolution, consent mapping at the signal level, clean segmentation built on behavior rather than demographics, has to happen regardless of which platform eventually wins the largest share of this spend.
Brands sitting on messy, unresolved identity data will find themselves locked out of the higher-value auction mechanics once genre-based and token-level bidding systems mature past the research stage. The ones with clean, documented, permissioned first-party data will be able to plug into whichever DSP relationships open up next, on whichever surface proves durable.
Perplexity's retreat from advertising is a reminder that user trust in this channel is not guaranteed just because the technology exists. Brands that treat governance, consent clarity, and brand safety exclusions as core to their targeting strategy, not as legal's problem to solve after the fact, will be the ones still standing when the next platform decides to open its own ad surface. The infrastructure is the strategy here. There isn't a separate one.
Sources
- Why 1st-Party Data and AI are the Future of Advertising
- "First-Party Data Activation" Offers Marketing and Technology Leaders a Strategic Playbook for Unlocking AI
- First-Party vs. Third-Party Data: Key Differences | Amperity
- AI Has Already Decided: First-Party Data Will Define Advertising’s Agentic Era
- Ad Insertion in LLM-Generated Responses
- arxiv.org
- openai.com


