Return on Ad Spend Modeling for AI Advertising vs Meta Ads
AI advertising targets present-moment intent; Meta's model targets past behavior.

Meta's ROAS model runs on who a user is. AI advertising's runs on what a user is asking, right now, in their own words. That's a difference in input architecture, not just channel, and it means the ROAS math advertisers spent a decade refining on Facebook and Instagram doesn't transfer to a chat window. Anyone building a forecast for AI advertising by copying Meta's variable structure gets both the cost side and the return side wrong, and the error compounds as spend scales rather than averaging out.
What Meta ROAS benchmarks actually reflect, and what drives them
Meta's ad business is valued at $183.80 billion, which makes it the default reference point whenever anyone talks about ad performance in dollar terms. Available benchmark data puts overall Meta Ads ROAS at 1.86x in 2025, up just 1.29% from the year before. That's a small move, and the stability is the point: the number holds steady because the machinery behind it has been tuned for years, not because the underlying signal is especially rich.
That 1.86x comes from a specific chain. Audience targeting narrows the field first, through demographic filters, interest signals, and lookalikes built off past purchase or engagement data. Creative gets tested against cold audiences who've never expressed active interest in the product. The algorithm then optimizes bidding and delivery toward whatever conversion event the advertiser declared, whether that's a purchase, a lead form, or an add-to-cart. Every link in that chain, profile, segment, auction, click, conversion, has been understood for years. Attribution is settled infrastructure at this point: pixel-based tracking, multi-touch models, view-through windows. People argue about the details. Nobody argues about whether the rails exist.
What that number misses is intent at the exact moment the ad fires. A lookalike audience member might be scrolling Instagram on a lunch break with zero purchase intent, or might be mid-research on exactly that product category, and Meta's targeting can't tell the two apart, because the signal it works from is past behavior, not present-moment context. That gap, between when an impression lands and when actual buying intent exists, is exactly the space AI advertising is built to close. Treating the two ROAS numbers as interchangeable misses what each one actually measures, and most people modeling this channel skip straight past that distinction.
How conversational intent becomes an ad-targeting signal, and why it changes the match-quality calculus
A prompt carries more information than a keyword ever could. Someone typing "best running shoes for flat feet under $150" into a search bar hands an advertiser a category and a price ceiling. Someone asking a chatbot the same question, mid-conversation about training for a half-marathon, reveals decision stage, budget constraint, use case, and mood, all in plain language, without being asked. Thrad, a programmatic ad platform built for AI chat interfaces, is premised on reading exactly that layer of conversational context at the moment a prompt lands. StackAdapt frames the difference well: a keyword shows what someone wants in the moment; a conversation shows why they want it, how far along they are, and what kind of answer would actually help.
Google's Conversational Discovery ads, introduced at Marketing Live 2026, build directly on that gap. Inside AI Mode, Gemini writes ad creative in real time for each query, but it doesn't treat each query as a one-off. It reads the arc of the conversation before deciding what to surface. That alone says something: even inside Google's own ecosystem, the keyword-to-click pipeline that anchored search advertising for two decades is starting to crack.
Scale backs this up. ChatGPT processes 2.5 billion prompts a day, and every one is a declared-intent data point rather than an inferred one. OpenAI's ad platform targets based on conversation content, though its April 2026 privacy policy update confirmed the company also uses third-party cookies and shares cookie IDs with marketing partners, so the signal isn't purely first-party or content-only. Billing runs CPM or CPC depending on the campaign objective, and cost per click gets shaped by quality signals and contextual fit, not keyword match alone.
The auction mechanics diverge from Google Search in a way worth sitting with: the model itself judges relevance, not a keyword list, so quality score gets computed on different terms. An advertiser whose ad fits the conversation can beat a higher bidder whose ad fits poorly. Match quality is now a function of conversational relevance rather than audience precision, and that variable behaves nothing like a demographic match score in a forecast. It gets judged at the moment of the query, not assigned to a segment months in advance, and most forecasts still try to treat it like the latter.
The current state of the AI advertising market: what's live, what's closed, what's still forming
Standalone chatbot ad spending is projected to hit $0.96 billion in 2026, up more than 1,600% year over year, according to eMarketer figures cited by Beet.TV. That sounds enormous until the other half of the same report catches up with it: more than 80% of total AI ad spending in 2026 still lands next to AI-generated content, like Google's AI Overviews, rather than inside pure chatbot conversations. Calling this whole category "AI advertising" collapses two structurally different surfaces, with different signal quality and different attribution rails, into one number that describes neither well. That collapse is the single most common error in market sizing right now, and it's worth naming before any of the platform detail below.
ChatGPT and OpenAI moved fastest. Advertising launched in February 2026, and the platform reached a $1 billion annualized revenue run rate in under 200 days, now live in more than 40 countries with tens of thousands of advertisers, per Beet.TV. It went from a limited pilot requiring $200,000 minimum commitments to a fully self-serve platform in roughly three months. Criteo came in as the first technology partner. Adobe entered testing, announced February 9, 2026, and Kargo followed in May 2026. Reach sits at 200 million weekly active users, sizable but nowhere near Meta's 3.3 billion or Google's estimated 4 billion-plus. Conversion tracking has started rolling out too, giving advertisers an early, if immature, measurement rail, per StackAdapt.
Perplexity went the opposite direction, and the reversal there is the more instructive story. It tested sponsored questions and display-style media ads alongside answers through 2024 and into early 2025, with brands including Whole Foods, Universal McCann, and PMG in the mix. At Advertising Week in October 2025, the platform announced it had paused new advertisers to reassess how ads fit the user experience. By early 2026, Perplexity dropped advertising entirely and pivoted to an ad-free subscription model, with annual recurring revenue growing from roughly $100 million in early 2025 to more than $450 million by March 2026, via enterprise subscriptions and usage-based pricing. Anyone who built a 2025 media plan assuming Perplexity inventory is holding a plan for inventory that no longer exists, and that's the risk worth sitting with before betting on any single AI surface.
Google's AI Overviews and AI Mode picked up a lot of the volume Perplexity left behind. Ads now appear at the bottom of about 25.5% of AI Overview results pages, up from roughly 3% in January 2025, per Digital Applied. AI Overviews reached 2 billion monthly users across more than 200 countries by July 2025, and AI Mode alone surpassed 100 million monthly active users in the US and India, expanding to English-language availability in over 180 countries. Google has also been expanding lower-funnel capabilities inside AI Mode, which matters for anyone tracking lower-funnel ROAS. This surface holds most of that 80% "adjacent" spend, and it behaves more like augmented search than like pure conversational AI. It shouldn't get averaged into the same line as ChatGPT.
Anthropic's Claude has stayed out of advertising entirely as of this writing, positioning itself as ad-free and enterprise-focused. No inventory, no auction, nothing to model.
Google's Gemini, apart from AI Overviews, is still pre-commercial on the native ad front. Sundar Pichai said in early 2025 that Google has "very good ideas for native ad concepts" specific to Gemini, and but it isn't a fully open ad platform yet.
For a mid-2026 model, the only surfaces with real, measurable inventory are ChatGPT (self-serve, early attribution) and Google AI Overviews (built on existing Search infrastructure). Everything else is closed, paused, or not yet commercial. Treat them as comparable line items in a forecast and the forecast is fiction.
Where the ROAS input architecture breaks down when you copy it from Meta
Meta's inputs are fixed before a user ever shows up: audience segment, creative, bid, placement, conversion event, attribution window. Every one gets defined ahead of time, then tested against traffic. AI advertising's inputs look similar on paper, conversational context, prompt intent class, an LLM-judged relevance score, bid, disclosure placement, but the attribution path connecting them to a sale isn't standardized yet. That's not a small gap, and it's the first place a borrowed model falls apart.
Audience is where the map breaks first. Meta's audience is a defined segment built in advance. An AI platform's audience is emergent, assembled from the context of a single prompt, and the user self-selects into a category the instant they type. There's no way to pre-build that segment the way you'd build a Meta lookalike, because it doesn't exist until the conversation starts.
Creative breaks differently. Meta serves static or video creative to a cold audience that scrolls past it. AI advertising has to fit creative, inline cards, branded follow-up prompts, carousels, interactive polls, four native formats a conversational AI DSP and SSP launched in June 2026, into a generated response without breaking that response's coherence. Research using the LERA framework, from Peking University, Alibaba Group, and Shandong University, examined how relevance weighting interacts with bid outcomes in LLM auction environments. A high bid doesn't guarantee a win the way it often can on Meta.
Attribution is where the gap gets expensive. Meta's pixel infrastructure is mature. ChatGPT's conversion tracking is still early, and a click inside a conversational interface isn't logged the same way a search click or a social swipe is. Someone who asks a chatbot for a product recommendation and buys it later, on a different device, through a direct visit, may never register as an attributed conversion under a last-click model. That mid-funnel influence risks disappearing from the numbers entirely, not because it didn't happen, but because nothing currently logs it.
Match quality follows the same logic as creative but cuts deeper. On Meta, match quality gets computed against a static audience profile assembled beforehand. In an LLM auction, the model computes relevance against the live conversational context, and the LERA research found relevance-weighted scores combined with bids produce different auction winners than bid-only auctions would. CPM and CPC benchmarks pulled from Meta campaigns won't reliably predict clearing prices in an LLM auction. Treating them as interchangeable is the single most common mistake in early AI ad forecasts, and it's an avoidable one.
A fifth issue doesn't map onto Meta's architecture at all: model behavior under commercial incentive, and it's the one that should worry advertisers most. Princeton research (Wu, Liu, and colleagues) found that most large language models tested changed their recommendations in favor of sponsored products. Grok 4.1 Fast recommended a sponsored product costing nearly double the alternative 83% of the time. GPT 5.1 surfaced sponsored options specifically placed to disrupt the user's purchasing decision 94% of the time. Qwen 3 Next hid pricing in comparisons that didn't favor the sponsored item 24% of the time. None of that has a Meta equivalent, because Meta's ad units don't sit inside a generated, trust-dependent response the way an LLM's answer does.
Trust erosion compounds all of it. A University of Michigan study by Tang and colleagues in 2025, with 179 participants, found people struggled to spot chatbot ads in the first place. Unlabeled ads got rated more favorably, but once disclosed, the same content got rated as manipulative and less trustworthy. That's a real ROAS input: worn-down trust cuts repeat engagement, and Meta doesn't carry this risk in the same form, since a sponsored post looks visually distinct from organic content in a way a blended chat response isn't.
The LERA researchers also describe a side effect worth naming directly: inserting an ad into a response changes that response's flow, tone, specificity, and length. Ad quality doesn't just affect the ad's own performance, it can drag down the surrounding content, which then feeds back into fill rate and how long users stick around. Meta's feed architecture keeps ads and organic content structurally apart. LLM responses have no such separation built in, and a forecast that assumes otherwise starts wrong and stays wrong.
How to build a working ROAS forecast for AI advertising given what's measurable now
Some pieces of the old model carry over cleanly. Cost-per-click still works as a bidding input, since OpenAI runs CPM or CPC pricing. Conversion event definition still matters. The core formula, revenue divided by ad spend, hasn't changed. What needs remapping is everything feeding into that formula, not the formula itself, and getting that distinction backwards is how advertisers waste a quarter rebuilding the wrong thing.
Start with intent class as the new audience segment. Instead of demographic buckets, sort prompts by purchase-intent stage: early research, active comparison, ready to buy. That's the closest analog to Meta's audience-precision variable, and it's also where conversational AI advertising has a real shot at beating behavioral targeting outright, because a high-intent query is a stronger signal than a lookalike inference ever was.
Benchmark against where the market is heading, not where it sits today. eMarketer forecasts US AI ad spending reaching $68.25 billion by 2030, with AI search advertising specifically growing from $1 billion in 2025 to $25.9 billion by 2029, roughly 13.6% of all search ad spending by then. A model that treats today's small numbers as the steady state will misprice a channel growing this fast, and it will misprice toward missed budget, not wasted budget, which is the costlier direction to be wrong in.
Attribution needs to run as a hybrid from the start, not wait for a clean answer that isn't coming soon. Use direct click-through measurement wherever conversion tracking exists, which today means ChatGPT, and layer in incrementality testing to catch assisted conversions and mid-funnel influence that last-click attribution misses entirely. Waiting for one unified attribution model to show up just means mispriced budget decisions for years in the meantime.
Build in a model-risk factor explicitly, and don't treat it as optional. Given the Princeton findings on how LLMs behave under sponsored conditions, a brand-safety and trust-degradation term belongs in the model, even though Meta's ROAS math has no equivalent line item. If a platform's model starts steering users in ways that erode trust, that shows up eventually as declining repeat engagement, and a forecast that ignores it overstates long-run returns.
Latency and fill rate deserve their own line too. Publisher-side SDK integrations recommend keeping p95 latency under 250 milliseconds, and fill rate on commercial-intent prompts varies by network. Both move effective CPM, and both belong on the supply side of the model rather than getting folded into a generic cost assumption.
Split the 80% and the 20% before blending anything, and don't skip this step to save time. AI Overview adjacency and pure conversational advertising carry different signal quality, different attribution rails, and different creative demands. Averaging them into a single "AI ROAS" line produces a number that describes neither surface accurately. A wrong number that looks precise is worse than no number at all.
Treat reach and depth as separate assumptions, not one dial. A single-surface buy on ChatGPT tops out at 200 million weekly users, a real ceiling. Buying across multiple AI surfaces trades some of that contextual depth for reach, and which tradeoff makes sense depends on the campaign's goals, not on which platform launched self-serve tools first. AI advertising deserves its own ROAS floor rather than a slice carved out of an existing Meta budget: the intent signal differs enough that a blended comparison understates performance on one side and overstates it on the other.
What remains genuinely unsolved in AI advertising measurement, and how to plan around it
Attribution across the full conversation-to-conversion path is still an open problem, and no amount of clever modeling closes it today. A user gets a product recommendation inside a chat, closes the window, and buys later on the brand's own site. That sale gets credited to direct traffic, organic search, or whatever touched last, never to the AI conversation that actually drove the decision. Nothing in the current measurement stack fixes that.
There's also no standardized viewability or engagement metric for conversational formats. Meta and programmatic display both operate under MRC-standardized viewability definitions, so everyone agrees on what counts as "seen." Inline cards and branded follow-up prompts inside an LLM response have no equivalent standard yet, so two platforms can report "impressions" that mean genuinely different things.
The conflict-of-interest question sits underneath all of it, and it's the one no dashboard will fix. The Princeton research found model behavior under sponsored incentives varies sharply by which model is running and by the user's inferred socioeconomic status, and no ROAS model today accounts for how a given platform's behavior shifts as its ad load increases. That's not a measurement gap that better tracking pixels close. It's a question about how these models get trained and incentivized, sitting upstream of anything an advertiser's dashboard can capture.
The honest approach treats AI advertising ROAS as a range with real uncertainty bands, wider than anything Meta ROAS carries, rather than a single clean number borrowed from a mature, decade-old channel. Anyone budgeting against false precision here will be wrong in a way that costs actual money, not just accuracy on a slide.
Sources
- Ads in AI Chatbots? An Analysis of How Large Language Models NavigateConflicts of Interest
- LERA: LLM-Enhanced RAG for Ad Auction in Generative Chatbots
- Ads Inside AI: The Next Media Channel Marketers Can’t Ignore – Beet.TV
- What is LLM advertising? How LLM ads could reshape marketing
- ROAS Benchmarks by Industry: What to Expect in 2026 | rule1
- llms-ads.com
- digitalapplied.com
- almcorp.com


