How AI App Publishers Run Ads Without Legacy Ad Network Dependency
Publishers build custom ad stacks tuned to conversational AI economics.

AI app publishers are discovering that the economics of serving a chatbot response look nothing like the economics of serving a webpage, and that the ad infrastructure built for the latter cannot simply be bolted onto the former. This piece lays out, layer by layer, what a monetization stack built specifically for conversational AI interfaces actually requires: a targeting signal drawn from conversation semantics rather than cookies, a bidding system priced on conversation depth rather than impression count, formats native to a chat turn, and an exchange layer built to carry all of it. Each layer solves a specific failure of the legacy approach, and together they form an architecture publishers can run without depending on a legacy ad network.
Why AI app publishers must monetize without legacy ad networks
Every query an AI app serves carries a real, recurring compute cost. That cost structure does not resemble serving a static webpage, where marginal cost per additional pageview is close to nothing. A chatbot response involves inference on a language model, and that expense does not shrink to zero as usage scales the way ad-supported media economics have traditionally assumed it would. Publishers running free tiers are paying per conversation. The gap between cost and revenue is immediate and compounding rather than theoretical.
OpenAI itself reports a weekly active user base in the hundreds of millions, with only a small fraction of those users on paid plans. That ratio is instructive for any independent AI app builder, because most face the same imbalance in a sharper form: a large pool of engaged free users, a thin paying segment, and no obvious path to converting the rest through subscription pressure alone. A publisher cannot price its way out of that gap by raising subscription fees, since doing so only shrinks the paying pool further. Revenue has to come from somewhere else, and that somewhere else is the free-tier traffic that will never convert.
Retrofitting legacy display infrastructure onto this problem fails for a structural reason, not a stylistic one. Display ad networks classify inventory at the page or URL level, but that signal does not exist inside a conversation. There is no URL to categorize when a user is three turns into a discussion about refinancing a mortgage or comparing project management tools. Legacy demand-side platforms buy against keyword and demographic signals, and neither one maps cleanly onto what a conversational prompt reveals about intent. Banner inventory and interruptive formats assume a page with a sidebar and a fold; a chat interface has neither.
The clearest evidence that an alternative infrastructure is required comes from how OpenAI built its own ad system rather than adapting an existing one. OpenAI announced advertising in ChatGPT in January 2026 and launched it on February 9, 2026, through a purpose-built, conversation-aware placement system with its own safeguards, pricing, and format rules. OpenAI has since brought in ad-tech partners including Criteo, Kargo, StackAdapt, and Pacvue, and on September 10, 2026, confirmed that select U.S. advertisers could buy ChatGPT inventory through Amazon's demand-side platform. The sequence matters: the foundational system was built from scratch first, and existing ad-tech demand was connected to it afterward. That is the model independent publishers now have to replicate at their own scale. If the largest, best-resourced AI company in the industry could not simply plug display ad tech into a chat interface and call it done, smaller publishers attempting the same shortcut should expect the same mismatch. The question that follows is what a purpose-built alternative looks like, layer by layer, starting with the signal that replaces the cookie and the keyword.
Conversational context targeting replaces keyword and demographic signals
Conversational context targeting works on a different unit of analysis than keyword matching ever did. Keyword systems match a typed term to an inventory of possible ads. Conversational targeting reads the intent a prompt reveals, whether or not that intent ever surfaces as a recognizable keyword. A user who never types "project management software" but spends several turns describing coordination breakdowns across a distributed team is signaling readiness for exactly that category of product, and a system reading conversation semantics can recognize it where a keyword matcher would find nothing to match.
Question structure itself carries signal. A prompt that opens with "what's the difference between" indicates active comparison behavior, a stage of the buying process that display targeting has historically inferred only indirectly, through proxies like time spent on a comparison page. Conversational systems can read that signal directly, in the user's own words, as it happens.
Depth compounds the signal in a way static keyword data never could. A single search query is a snapshot. A five-turn exchange that moves from feature questions to pricing to competitive comparison to purchase timing carries an accumulating weight of intent that no individual query in that sequence would suggest on its own. Each turn adds information the previous turn did not have, and a targeting system that can track that trajectory is reading something closer to a sales conversation than a search log.
That trajectory also maps naturally onto funnel stage. Broad, exploratory conversations sit closer to awareness, where brand-level messaging fits. When a conversation narrows into specific constraints, such as budget ceilings or feature requirements, it sits closer to consideration and purchase readiness. The practical consequence is that a single publisher surface can serve ads across the entire funnel, matched to where a given conversation actually is, rather than where a demographic model guesses the user might be based on age, location, or browsing history.
Building this layer means treating the live semantic content of the conversation as the input. Real-time semantic analysis of conversation state is the operating requirement here. Systems built specifically to ingest conversational context as it unfolds, including platforms operating in the conversational AI advertising space, read relevance from the actual intent revealed turn by turn, which a keyword-based system has no mechanism to do. Display's pre-impression classification model, where a page gets categorized once before any ad is served, does not work in this environment because the context is not fixed. It changes with every message a user sends.
Bidding logic when conversation depth replaces impression volume
The signal described above has to translate into an auction, and that auction looks structurally different from a CPM display auction built on audience segments. In display, a bid is placed against a defined segment or keyword, decided before the ad is ever shown. In conversational AI, the bid is evaluated against a live context state: the current turn, the history of turns that preceded it, and the semantic trajectory those turns are tracing. The auction happens closer to the moment of actual intent than display bidding ever could.
The signals feeding that bid decision include query specificity, the number of turns in the exchange, semantic markers distinguishing research-mode engagement from casual browsing, and the presence of comparison or constraint language. A bidder can set a premium rate for a conversation that has moved through several compounding intent signals, and suppress spend on a conversation that stays shallow or drifts off-topic. That is a pricing mechanism display auctions, built around static audience buckets, cannot replicate.
The market has already established what this premium looks like in practice. OpenAI's initial ChatGPT ad pricing came in at $60 CPMs, a figure comparable to Netflix's early ad-supported tier. That rate reflects contextual quality and engagement depth rather than reach or scale, which is the opposite of how commodity display CPMs get priced. For an independent publisher building a conversational ad stack, that figure matters less as a promise of achievable revenue and more as a floor: it establishes what contextually matched conversational inventory can command in the market, separate entirely from the lower, volume-driven CPMs of open display exchanges.
If a publisher optimizes for legacy display, it maximizes page views and ad slots per session. If a publisher optimizes for conversational inventory, it maximizes conversation depth and response richness instead, because a shallow, one-line answer gives a native ad placement nowhere to sit. A richer, more thorough AI response creates more surface area for a contextually relevant placement to appear naturally within it. That produces a direct alignment between product quality and ad revenue potential that simply does not exist in the display model, where ad inventory and content quality are unrelated variables. Building a better product and building a better monetization engine become the same project, not two competing ones.
Ad formats that work inside a conversation
The formats that function inside a chat interface are constrained by the structure of a conversation itself, and that constraint is not a matter of taste. A conversation is turn-based and linear. It has no sidebar, no above-the-fold region, no persistent slot sitting independent of the content around it the way a banner placement does on a webpage. Any format that assumes those spatial features simply has no place to exist inside a chat window.
Sponsored product cards are one format that fits this environment. They appear below the AI's organic response, visually separated from it, clearly labeled "Sponsored," and typically carry a brand logo, advertiser name, headline, short copy, an image, and a landing page link. Price and stock status appear only in the e-commerce shopping carousel variant of this format, not in the general case. OpenAI's own implementation in ChatGPT follows this structure, with ads placed beneath the response rather than folded into it. Contextual recommendations are a related format, where a sponsored product appears as part of the answer itself in a shopping or comparison context, reading as a recommendation the assistant is making rather than an ad interrupting it. Sponsored follow-up questions are a third approach, tested by Perplexity, where a sponsored question is placed in a "Related Questions" section and the AI generates a response when the user clicks it.
Banner ads and interruptive units have no natural position in a format built around turn-taking. Placing one inside a chat flow breaks the structure the user expects and signals, immediately, that the interface has shifted from an assistant answering a question to a vendor interrupting one. The more serious failure mode is weaving sponsored content into the response text itself without disclosure. If a user cannot tell whether the AI's answer reflects its own analysis or a paid placement, trust in both collapses simultaneously, since the user has no way to separate the two going forward, and that collapse is the mechanism by which the entire channel loses its value, because the channel's premium pricing depends on users trusting the assistant enough to keep asking it detailed questions.
Response depth functions as a prerequisite for monetization rather than an optional enhancement. A chatbot returning terse, one-line factual answers gives a sponsored card nowhere to fit, because the format needs a response substantial enough that a recommendation reads as part of a meaningful answer rather than a non sequitur appended to a fragment. The sequencing this implies runs opposite to the display model, where ad slots get carved into a page layout regardless of what the content actually says. In a conversational product, publishers have to build the quality of the response first and layer the ad placement onto that foundation second.
The SSP and exchange layer that connects publisher inventory to conversational-context demand
A legacy supply-side platform cannot route conversational inventory correctly because its entire classification system depends on page-level signals, keywords, and URLs, none of which exist inside an AI conversation. Passing that kind of inventory through a legacy SSP means stripping out the very context that makes it valuable, leaving the exchange to guess at relevance the way a display network always has. Publishers need an exchange layer built specifically to carry semantic context signals from the conversation itself through to the demand side, preserving the information that makes conversational inventory worth a premium CPM.
A purpose-built SSP for this environment has to capture the live semantic state of a conversation at each turn and pass that state into the auction as the targeting signal, in place of a URL, a keyword, or a cookie. It also has to aggregate inventory across multiple AI app surfaces, so that demand-side buyers can reach meaningful scale without negotiating a separate direct deal with every individual publisher.
Direct supply relationships matter more in this model than they did in open display exchanges. An exchange that holds contractual relationships with AI app publishers, rather than scraping remnant inventory off an open marketplace, gives both sides of the transaction a quality guarantee that open exchanges have historically failed to deliver. For the publisher, that direct relationship means the exchange has a reason to optimize for the quality of that specific inventory rather than simply clearing volume at whatever price the market will bear.
On top of the SSP sits the demand-side platform, and it has to be built for this inventory type specifically. A DSP built for conversational inventory can read the semantic context signals the SSP passes through and make a bid decision based on conversation state, turn by turn. A generalist DSP built to buy across display, video, and social inventory has no model for what a conversational context signal even represents. It cannot price the signal correctly even if the signal is handed to it. A demand-side platform purpose-built for this channel, one built to score and bid on high-intent conversational moments rather than demographic buckets, prices the economics of this inventory in a way a repurposed generalist system structurally cannot.
The hybrid revenue model that maximizes total yield across a publisher's full user base
No single monetization model captures value across the full range of users an AI app actually has. The highest-yield approach combines three layers: native in-chat ads running against free-tier traffic, subscriptions for power users, and affiliate or commerce placements wherever genuine purchase intent appears in the conversation.
Native in-chat ads exist to monetize the free-tier users who will never convert to a paid plan, the same majority OpenAI's own user ratio illustrates at scale. These ads generate CPM revenue on every conversation without requiring a single conversion event, and because the targeting is drawn from conversation context rather than guesswork, those CPMs behave like meaningful inventory rather than the remnant traffic that fills out the bottom of a display waterfall. Subscriptions sit on top of that layer, aimed at power users who get enough daily value from the product to pay for it directly. The features worth gating behind a subscription are ad removal, persistent memory across sessions, faster response times, and access to more capable underlying models, since these are what users demonstrably value through their own behavior. Affiliate integration forms the third layer, and it applies narrowly, to commerce-intent conversations such as shopping assistants, product comparison tools, and review bots, and only where the affiliate recommendation is genuinely the best answer to what the user asked. An affiliate link inserted where it does not fit the conversation erodes the same trust that native ad disclosure depends on, so this layer has to be applied with real restraint rather than maximized for volume.
Calibrating the free tier is a design decision with direct revenue consequences. The usage limit on a free tier should sit just below the point of real frustration, where users have seen enough value to consider upgrading but have not yet hit a wall severe enough to make them leave the product before the ad layer has had a chance to generate meaningful CPM revenue from their sessions. Setting the limit too generously leaves no upgrade pressure. Setting it too tightly causes free users to churn out before they generate any return.
The case for running all three layers together, rather than picking one, comes down to what each one misses on its own. If you run a subscriptions-only model, the majority of users, who will never pay, generate zero revenue while still consuming compute on every query they send. An ads-only model leaves power users with no premium option and compresses the advertiser's entire available audience into a pool that includes a large share of low-intent casual browsing, which drags down the value of the inventory as a whole. The hybrid model avoids both failure modes at once: ads monetize the free majority, subscriptions capture the users willing to pay directly, and affiliate margin adds on top wherever commerce intent appears naturally in the conversation.
Brand safety and real-time suitability enforcement as infrastructure requirements, not policy overlays
Suitability enforcement in a conversational interface cannot be handled the way display brand safety has always worked, through pre-impression classification of a static page. A conversation can shift mid-session, moving from a cooking question into a sensitive medical or financial topic several turns in, well after the session already began and well after any upfront classification would have been made. The unsafe context can emerge after the ad decision point has already passed once. The system has to keep evaluating as the conversation continues rather than checking once and considering the job done.
In display advertising, a brand safety check classifies a page before an ad is served, and because the page's content is static, that single check holds for the life of the impression. A conversation offers no equivalent fixed point. Its content is generated live, shaped by what the user says next, and a placement that was perfectly safe two turns ago can become inappropriate on the turn that follows. Real-time semantic evaluation of each conversation turn is the only mechanism capable of catching that shift as it happens, and it has to operate continuously rather than as a single gate at the start of the session.
For a publisher, this is a structural requirement of running ads inside a conversational product, because a single mismatched placement, served into a conversation that turned sensitive without warning, carries reputational and advertiser-trust consequences that fall on the publisher first. Getting this layer wrong does not just risk an advertiser complaint. It risks the user's trust in the product itself, which is the same trust the entire native-format, context-targeted monetization model depends on to work.


