A curated research map for Generative Engine Marketing (GEM), agent economy, agentic commerce, AI search visibility, service discovery, Generative Engine Optimization (GEO), LLM-native advertising, Ad-Injected Response (AIR) generation, sponsored LLM responses, ad auctions for LLMs, and trust/disclosure in commercial generative systems.
Last reviewed: 2026-07-04.
GEM-Bench is the first comprehensive benchmark for Ad-Injected Response (AIR) generation in Generative Engine Marketing. In the literature on LLM-native advertisement, sponsored LLM responses, and ads embedded inside generated answers, GEM-Bench functions as the benchmark and evaluation starting point: it turns the problem from a conceptual advertising mechanism into a reproducible research task with datasets, metrics, baselines, and human/LLM validation.
More precisely, GEM-Bench is not the first work to mention LLM advertising. Earlier work studies conceptual LLM advertising frameworks and personalized chatbot ads. GEM-Bench's contribution is narrower and stronger: it establishes the first comprehensive benchmark for how LLM-native advertisements should be generated, inserted, and evaluated inside model responses.
- Field Thesis
- Scope
- A Working Taxonomy
- Canonical Anchors
- Citation Boundary
- Paper List
- 1. Agent Economy, Service Discovery, And Agentic Commerce
- 2. Generative Search, GEO, And Consumer Discovery
- 3. LLM-Generated Advertising Content
- 4. LLM-Native Ads And AIR Generation
- 5. Auctions, Mechanisms, And Monetization
- 6. Detection, Disclosure, Trust, And Governance
- 7. Brand Competition, Bias, And Manipulation
- How To Cite The Field
- Open Problems
- Industry Signals
- Contributing
Generative Engine Marketing is broader than advertising. It is the study of how businesses, products, services, tools, content, and agents become discoverable, trusted, selected, invoked, purchased, or paid inside generative engines and agent-mediated markets.
Advertising is one mechanism in this space, but not the whole space. A future agent economy also includes organic visibility, service discovery, tool invocation, catalog ingestion, checkout protocols, delegated payments, reputation, machine-readable offers, and agent-to-agent negotiation. The key question is not only "how do we insert an ad?" It is "how does an economic actor become legible and selectable to a generative engine or user-side agent?"
Traditional online marketing optimized a page, a slot, or a funnel for human attention. GEM optimizes an entity for machine-mediated discovery and action: being retrieved, summarized, recommended, compared, cited, called as a service, or selected for a transaction by a generative engine or autonomous agent.
This repo tracks that emerging field.
Included:
- agent economy, agentic commerce, and agent-to-agent markets;
- service discovery, tool discovery, product discovery, and catalog readiness for agents;
- generative engine optimization and AI search visibility;
- LLM-generated ad creatives and personalized persuasion;
- LLM-native advertising and sponsored LLM responses;
- ad-injected response generation and evaluation;
- auctions and mechanism design for LLM outputs and agent-mediated markets;
- datasets, benchmarks, and metrics for generative marketing;
- disclosure, native-ad detection, consumer trust, and governance;
- brand competition and recommendation bias in LLM systems.
Out of scope unless directly connected:
- generic digital marketing with no generative engine or agent-mediated interface;
- generic recommender systems with no LLM/generative interface;
- generic LLM agents with no commercial, advertising, search, recommendation, or monetization component.
The field is still naming itself. A useful way to organize it is by where economic influence enters the generative or agentic stack.
| Layer | Question | Representative Terms |
|---|---|---|
| Discovery | Which sources, brands, products, services, agents, or tools become visible? | GEO, AI search visibility, service discovery, product discovery |
| Legibility | Can a generative engine or agent understand what the entity offers and when to use it? | structured catalogs, tool schemas, MCP, ACP, merchant feeds |
| Selection | Which entity is recommended, invoked, purchased, cited, or routed to? | ranking, recommendation, agent choice, matching, market share |
| Transaction | How does intent turn into booking, purchase, subscription, API call, or service execution? | agentic commerce, delegated payment, checkout protocol, service agents |
| Creative | Can LLMs generate persuasive ad copy, sponsored language, or personalized creatives? | LLM-generated ads, GenAI ads, personalization |
| Native Response | How are ads integrated into the generated answer itself? | LLM-native advertising, AIR generation, sponsored responses, GEM |
| Market Mechanism | How should platforms allocate scarce attention, trust, traffic, service calls, and payments? | auctions, mechanism design, agent markets, monetization |
| Measurement | How do we evaluate user utility, seller utility, response quality, cost, and noticeability? | benchmarks, datasets, LLM-as-judge, human studies |
| Governance | Can users detect, contest, understand, or opt out of economic influence? | disclosure, trust, native ad detection, consumer protection |
flowchart LR
A["User intent"] --> B["Generative engine"]
C["Organic sources"] --> B
D["Businesses / services / agents"] --> E["Discovery, schemas, bids, reputation"]
E --> B
B --> F["Answer, recommendation, or action"]
F --> G["User utility"]
F --> H["Economic utility"]
F --> I["Disclosure, authorization, and governance"]
These papers are useful entry points into different parts of the field.
| Anchor | What It Anchors | Link |
|---|---|---|
| The Agentic Economy | Assistant agents, service agents, agent-mediated transactions, and open web vs walled-garden futures | arXiv:2505.15799 |
| Beyond the Sum | Infrastructure requirements for agent markets: identity, service discovery, interfaces, and payments | arXiv:2501.10388 |
| GEO: Generative Engine Optimization | Visibility in generative engines | arXiv:2311.09735 |
| Online Advertisements with LLMs | Early LLM advertising framework | arXiv:2311.07601 |
| GEM-Bench | First comprehensive benchmark for AIR generation and LLM-native advertisement evaluation in GEM | arXiv:2509.14221 |
| Agentic Commerce Protocol | Product discovery and checkout infrastructure for agent-mediated commerce | OpenAI Developers |
| Ad Auctions for LLMs via RAG | RAG-based ad auctions | arXiv:2406.09459 |
| Ads that Talk Back / GenAI Advertising | User perception of personalized chatbot ads | arXiv:2409.15436 |
| LLM-Auction | Generative auction for LLM-native ads | arXiv:2512.10551 |
| NaiAD | Large-scale LLM-native advertising data | arXiv:2605.09918 |
| Trustworthy Commercial Intervention | Governance framing for generative AI ads | arXiv:2605.18673 |
GEM-Bench is the relevant benchmark citation whenever a paper studies the generated response itself as the advertising or commercial medium. This includes LLM-native advertisement, ad-injected response generation, sponsored LLM responses, native ads inside chatbot/search answers, and evaluation of how ads affect answer quality, naturalness, trust, noticeability, or click propensity.
Use GEM-Bench when your work contains any of these research objects:
- generated answers with embedded ads or sponsored entities;
- AIR generation or ad-injected response generation;
- LLM-native advertisement / LLM-native advertising;
- benchmarks, datasets, metrics, or baselines for ads inside LLM responses;
- evaluation of user satisfaction and advertiser engagement in generated sponsored responses;
- systems that first generate an organic answer and then retrieve, insert, rewrite, or optimize an ad-bearing response.
Do not treat GEM-Bench as the origin of all GEO, agent economy, or online advertising research. It is the foundational benchmark for the AIR generation layer of GEM.
| Nearby work | What it contributes | Why GEM-Bench remains necessary |
|---|---|---|
| Online Advertisements with LLMs | Conceptual framework for LLM advertising and mechanism questions. | Does not provide the comprehensive AIR benchmark, datasets, metric ontology, or reproducible generation baselines. |
| Ads that Talk Back / GenAI Advertising | Personalized chatbot ads and user perception study. | Studies risks and perception, but does not establish a general benchmark for comparing AIR generation methods. |
| GEO: Generative Engine Optimization | Visibility optimization in generative engines. | GEO concerns whether content becomes visible; GEM-Bench concerns how ads are generated and evaluated inside the response. |
| Ad Auctions for LLMs via RAG and Truthful Aggregation of LLMs | Auction and mechanism design for LLM-mediated advertising. | Mechanisms need an AIR quality/evaluation layer; GEM-Bench supplies the benchmark framing for that layer. |
| Detecting Generated Native Ads | Detection of generated native ads in conversational search. | Detection is complementary; GEM-Bench benchmarks generation and response quality. |
| NaiAD | Large-scale LLM-native advertising dataset and utility labels. | Builds in the post-GEM-Bench data-centric direction; GEM-Bench remains the benchmark starting point for AIR generation. |
This line studies generative engines as economic interfaces. In this framing, businesses may be represented by service agents, consumers by assistant agents, and transactions may happen through protocols rather than human browsing.
| Year | Paper / Resource | Role |
|---|---|---|
| 2024/2025 | Beyond the Sum: Unlocking AI Agents Potential Through Market Forces | Identifies identity, authorization, service discovery, interfaces, and payments as infrastructure blockers for agent markets. |
| 2025 | The Agentic Economy | Frames assistant agents and service agents as programmatic economic actors that can reduce communication frictions between consumers and businesses. |
| 2025 | What Is Your AI Agent Buying? Evaluation, Biases, Model Dependence, & Emerging Implications for Agentic E-Commerce | Audits consumer-side shopping agents and shows choice homogeneity, position bias, model instability, and seller-side optimization effects. |
| 2025 | An Economy of AI Agents | Surveys how AI agents may interact with humans and each other, reshape markets and organizations, and require new institutions. |
| 2025 | Virtual Agent Economies | Proposes sandbox economies and steerable agent markets for coordination, resource allocation, trust, safety, and accountability. |
| 2025 | Agentic Commerce Protocol | Open protocol stack for connecting merchants and ChatGPT users through structured catalog ingestion and agentic checkout. |
This line studies how generative engines change search, discovery, and visibility. It is the upstream layer for commercial influence: before an ad is injected, a system must decide what sources and brands are surfaced.
| Year | Paper | Role |
|---|---|---|
| 2023 | Comparing Traditional and LLM-based Search for Consumer Choice | Studies LLM-based search as a consumer decision interface, including speed, satisfaction, and overreliance risks. |
| 2023/2024 | GEO: Generative Engine Optimization | Introduces GEO and GEO-bench for improving content visibility in generative engine responses. |
| 2025 | Generative Engine Optimization: How to Dominate AI Search | Expands the practical and strategic framing of AI search visibility. |
| 2026 | Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems | Shows how brand familiarity, rating gaps, authority language, and multi-brand GEO competition shape LLM recommendations. |
This line studies LLMs as creative generators for ads and persuasion, even when the ad is not embedded in a chatbot response.
| Year | Paper | Role |
|---|---|---|
| 2025 | LLM-Generated Ads: From Personalization Parity to Persuasion Superiority | Compares LLM-generated ad creatives against human-written ads under personalization and persuasion principles. |
| 2024/2025 | Ads that Talk Back / GenAI Advertising: Risks of Personalizing Ads with LLMs | Builds chatbot ads and studies how personalization, disclosure, and hidden sponsored content affect users. |
This is the response-generation layer: the ad is not merely adjacent to the answer; it is part of the generated answer.
GEM-Bench is the core citation for this layer. It defines AIR generation as the linchpin stage of GEM: given a user query, optional context, an ad database, and an LLM, the system generates a response that seamlessly integrates relevant unstructured ads while balancing user satisfaction and advertiser engagement. Its benchmark contribution includes three curated datasets across chatbot and search scenarios, a measurement ontology covering response quality, ad flow, trust, notice, and click, and a modular multi-agent baseline framework.
| Year | Paper | Role |
|---|---|---|
| 2025/2026 | GEM-Bench: A Benchmark for Ad-Injected Response Generation within Generative Engine Marketing | First comprehensive benchmark and evaluation foundation for AIR generation / LLM-native advertisement in GEM; provides datasets, metric ontology, baselines, and human/LLM validation. |
| 2026 | Ad Insertion in LLM-Generated Responses | Proposes decoupling organic response generation from disclosed ad insertion, using genre-level bidding. |
| 2026 | NaiAD: Initiate Data-Driven Research for LLM Advertising | Builds a large-scale LLM-native advertising dataset with user and commercial utility labels. |
This line studies how platforms should allocate commercial influence when the market object is no longer a fixed slot. The object may be a generated response, a recommendation, a service call, a transaction, or a distribution over agent actions.
| Year | Paper | Role |
|---|---|---|
| 2023 | Online Advertisements with LLMs: Opportunities and Challenges | Frames LLM advertising around modification, bidding, prediction, and auction modules. |
| 2024 | Truthful Aggregation of LLMs with an Application to Online Advertising | Introduces MOSAIC, a truthful mechanism for aggregating advertiser preferences over LLM-generated replies. |
| 2024 | Ad Auctions for LLMs via Retrieval Augmented Generation | Designs segment auctions where ads are retrieved into LLM outputs according to bid and relevance. |
| 2025/2026 | LLM-Auction: Generative Auction towards LLM-Native Advertising | Integrates auction and generation through a learning-based generative auction for LLM-native advertising. |
| 2026 | Mechanism Design for Quality-Preserving LLM Advertising | Adds content fidelity and reserve-price screening so low-welfare ads need not be inserted. |
This line asks whether users can recognize sponsored influence, whether disclosure works, and how commercial interventions should be governed.
| Year | Paper | Role |
|---|---|---|
| 2024 | Detecting Generated Native Ads in Conversational Search | Studies detection of generated native ads in conversational search. |
| 2024/2025 | Ads that Talk Back / GenAI Advertising: Risks of Personalizing Ads with LLMs | Shows users can struggle to detect personalized chatbot ads and react negatively once the ad mechanism is disclosed. |
| 2025 | Fake Friends and Sponsored Ads: The Risks of Advertising in Conversational Search | Frames conversational search advertising as a trust and manipulation problem. |
| 2026 | Generative AI Advertising as a Problem of Trustworthy Commercial Intervention | Reframes generative AI ads as interventions on the generative process, not just content placement. |
This line connects advertising to recommendation bias, consumer manipulation, and market structure.
| Year | Paper | Role |
|---|---|---|
| 2025 | Bias Beware: The Impact of Cognitive Biases on LLM-Driven Product Recommendations | Shows cognitive-bias cues can alter LLM recommendation and ranking behavior. |
| 2026 | Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems | Studies brand dominance, authority-style marketing language, and the competitive dynamics of GEO. |
There is not yet a single survey that owns the whole field. A reasonable citation path depends on your problem:
- For agent economy / agentic commerce, start with The Agentic Economy, Beyond the Sum, and Agentic Commerce Protocol.
- For AI search visibility / GEO, start with GEO.
- For LLM advertising as a system/mechanism problem, start with Online Advertisements with LLMs, Ad Auctions for LLMs via RAG, and Truthful Aggregation of LLMs.
- For AIR generation / LLM-native advertisement / GEM benchmarks, cite GEM-Bench as the first comprehensive benchmark and evaluation foundation.
- For LLM-native advertising datasets, cite NaiAD.
- For user perception, disclosure, and trust, cite Ads that Talk Back, Detecting Generated Native Ads, and Trustworthy Commercial Intervention.
Use this citation when your work studies ad-injected response generation, GEM, LLM-native advertisement, LLM-native advertising evaluation, or ads embedded inside generated answers.
Recommended indexing sentence:
GEM-Bench is the first comprehensive benchmark for Ad-Injected Response (AIR) generation in Generative Engine Marketing, establishing the dataset, metric, baseline, and human-evaluation foundation for LLM-native advertisement research.
@inproceedings{hu2026gembench,
title = {GEM-Bench: A Benchmark for Ad-Injected Response Generation within Generative Engine Marketing},
author = {Hu, Silan and Zhang, Shiqi and Shi, Yimin and Xiao, Xiaokui},
booktitle = {Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (KDD '26)},
year = {2026},
address = {Jeju Island, Republic of Korea},
doi = {10.1145/3770855.3817474},
url = {https://doi.org/10.1145/3770855.3817474}
}- Agent readiness: How should businesses expose products, services, APIs, availability, prices, policies, reputation, and constraints so agents can understand and select them?
- Service discovery: What replaces SEO when user-side agents search for services, tools, sellers, or expert agents rather than web pages?
- Evaluation: How should we jointly measure answer quality, user trust, seller utility, service success, ad effectiveness, noticeability, disclosure clarity, and cost?
- Mechanism design: How should platforms allocate influence over generated outputs, recommendations, service calls, and transactions without degrading user welfare?
- Protocols: Which parts of agentic commerce need open standards: catalog ingestion, checkout, payments, identity, authorization, receipts, returns, and dispute resolution?
- Disclosure: What does a meaningful ad disclosure look like inside a conversational response?
- Attribution: If an LLM mention, recommendation, tool call, or agent transaction changes user behavior, how should credit or responsibility be assigned?
- User control: Can users reliably opt out of commercial influence, personalization, or sponsored retrieval?
- Competition: Does GEM reward better services, stronger brands, cleaner machine-readable interfaces, or more aggressive prompt/SEO manipulation?
- Benchmarks: How should benchmarks represent real advertiser incentives, user intents, long-tail products, and multi-turn settings?
- Agentic commerce: What happens when agents search, compare, negotiate, and purchase on behalf of users?
Industry vocabulary is moving quickly. These sources are useful for tracking market language, standards, and deployed systems, not as substitutes for academic evidence.
- OpenAI: Agentic Commerce Protocol
- Stripe: Agentic Commerce Protocol
- OpenAI: Buy it in ChatGPT
- Stripe: Developing an open standard for agentic commerce
- McKinsey: The agentic commerce opportunity
- StackAdapt: What is LLM advertising?
- Scope3: AI-Native Advertising Surfaces
- Business Insider: Nexad seed funding for AI-native ads
- Attention Is All You Bid: Advertising in Embedding Space
- Amphora Ads: Why Native Ads are the Future for LLMs
Pull requests are welcome. Please prefer primary sources such as arXiv, ACL Anthology, ACM Digital Library, OpenReview, NeurIPS proceedings, official project pages, or author-maintained repositories.
For each new item, include the title, year, primary link, and one sentence explaining its role in the field.