A growing share of travelers now put their search for lodging, a travel agency, or an activity directly to a conversational artificial intelligence assistant instead of a traditional search engine. This work examines the extent to which Argentine tourism providers are prepared to be identified and cited by those engines. We evaluated 42 sites belonging to providers from the country's main tourism trade association, spanning travel agencies and tour operators, hotels, tango houses, and a group of smaller categories. The average overall score was 20 out of 100, with no domain reaching the readiness level for generative engines, and 62% of the sample showing zero structured data (schema.org). The weakest criterion was sub-question coverage (13/100), and the least weak, topical specificity (34/100). As a complement, we ran a field experiment: we put to ChatGPT, Gemini, and Perplexity the questions a real traveler would ask, across two of the sample's categories, following the IAB's AI visibility measurement framework. Of the 32 providers covered by that experiment, 78% was never named in a single one of the 12 captures performed. We conclude that the Argentine tourism industry exhibits a structural lag in its readiness for the generative discovery layer, and that this lag has a direct, measurable translation in practice: the vast majority of providers, quite simply, do not exist for an AI a traveler asks for a recommendation.
* Across a sub-sample of 32 providers from two categories, in 12 captures put to ChatGPT, Gemini, and Perplexity. See section 4.4.
For two decades, a website's visibility was determined by its ranking in traditional search engines. The mass adoption of conversational assistants built on large language models has introduced a second discovery layer: a growing share of users now put their queries to a generative engine that answers synthetically, citing or naming a small set of sources rather than offering a list of links. In tourism this is not hypothetical: questions like "which hotel should I book in Bariloche" or "which travel agency should I use for a trip to Mendoza" are exactly the kind of query travelers put to an AI assistant today, ahead of a search engine.
This shift gives rise to a specific discipline, Generative Engine Optimization (GEO), whose purpose is to maximize the probability that a site's content is retrieved and cited by an artificial intelligence engine. Unlike search engine optimization, GEO favors content's ability to be interpreted, chunked, and attributed programmatically.
No systematic diagnosis exists of how prepared the Argentine tourism sector is for this transition. This work aims to fill that gap by measuring a sample of sector providers, and complementing it with an experiment that observes what actually happens when three generative engines are asked the questions a real traveler would ask.
The instrument used evaluates eight criteria, each weighted according to its relative influence. All are grounded in sources that publish their method and their data: peer-reviewed research, official technical standards, and the engines' own official documentation. Every reference is linked at the end of this work:
Data density, direct question-answer relevance, and authority and trust are grounded in Aggarwal et al. (KDD 2024); freshness and updates and content structure, in Chen, Wang, Chen, and Koudas (University of Toronto, EDBT/ICDT 2026 Workshops); technical markup, in the Schema.org / JSON-LD specification, with Dang et al. (University of Nantes, Semantic Web Journal 2025) as a reference for how widely markup is adopted across the web; topical specificity, in Yu, Yang, Ding, and Sato (University of Tokyo, 2026); and sub-question coverage, in Xie et al. (NAACL 2025).
We assembled a sample of 42 domains of commercial tourism providers, drawn from the public member directory of the country's main tourism trade association, spanning a variety of categories (travel agencies and tour operators, hotels, tango houses, transportation, airlines, timeshare). The full list is presented in Annex A.
The analysis was performed with Nostos, a GEO evaluation engine that assigns each domain a score from 0 to 100 on each of the eight criteria described, as well as an overall score and a readiness level across four categories.
| Level | Description | Range |
|---|---|---|
| Level 1 | Ready for AI | 80–100 |
| Level 2 | Targeted adjustments | 55–79 |
| Level 3 | Partial restructuring | 30–54 |
| Level 4 | Reconstruction needed | 0–29 |
Each domain was analyzed in a single measurement during August 2026. For each domain, the instrument examines the home page, a set of additional internal pages, and the site's sitemap, from which it evaluates publicly accessible structural, markup, and update signals. Each criterion's score results from the aggregation of those signals at the domain level, not from a single page.
Nostos incorporates a two-layer detection mechanism to identify content a site builds dynamically through JavaScript, without needing to execute it. The first layer recognizes the footprint of more than twenty known providers of review, social, chat, and booking widgets, by analyzing the external and inline script references present in the document. The second layer applies structural heuristics — empty containers with telltale identifiers and shell patterns typical of single-page application frameworks (React, Vue, Angular, Next.js, Svelte) — as a containment mechanism for providers not catalogued by the first layer. When content is detected under this criterion, the system explicitly reports it as present but not machine-readable, rather than assuming it does not exist.
As a complement to the GEO score, we ran a direct experiment: putting to three generative engines (ChatGPT, Gemini, and Perplexity) the questions a real traveler would ask, and recording whether they name any provider from the sample. The protocol follows the AI visibility measurement framework published by the IAB. Three fixed questions were defined per category — an open recommendation, a concrete purchase-intent query, and a listing request — with fixed conditions for each capture: a new chat with no history or personalization, the geography written into the question itself, the first answer taken without follow-up, the same search mode noted per model, and the date and time recorded. Captures were run manually. Two of the sample's categories were covered — travel agencies and tour operators, with all three questions, and hotels, with the recommendation question — across a total of 32 providers. Per the IAB's own framework, a single capture per question and model is a directional measurement, not a statistically representative one; it is treated as such throughout this work.
The average overall score of the sample was 20 out of 100. No domain reached the readiness level for generative engines, and only one (2.4%) reached the targeted-adjustments level. The highest score observed was 69, and the lowest, 0.
Criterion averages, ordered from most to least deficient, were: sub-question coverage 13/100, direct question-answer relevance 14/100, data density 16/100, technical markup 18/100, content structure 19/100, freshness and updates 21/100, authority and trust 29/100, topical specificity 34/100.
a) Technical markup absent. 62% of domains (26 of 42) showed zero technical markup (0/100). None of the 42 sites exceeded 72/100: not even the best cases mark up their content in full.
b) Severity of the problems. Of the 294 issues detected across the 42 analyses, 42.2% were classified as critical and 41.8% as important — 84% of the total does not correspond to cosmetic adjustments but to structural shortcomings requiring substantive intervention.
c) Differences by category. The smaller categories (transportation, airlines, and timeshare) were grouped into a single block, to avoid an average calculated over one or two sites individually identifying a specific provider's score. With that caveat, and over small subgroups, hotels register the highest average (26/100), and travel agencies and tour operators, the lowest (17/100) — with tango houses and the smaller-categories block in between.
Across the 32 providers covered by the experiment (22 travel agencies and tour operators, with all three questions — 9 captures —, and 10 hotels, with the recommendation question — 3 captures —, 12 captures in total), 7 (21.9%) received at least one mention in at least one capture, and 25 (78.1%) were never named, in any question or by any model.
The relationship between a provider's own GEO score and its probability of being mentioned was not clear in either direction: the average score of the 7 mentioned providers (22/100) was neither clearly higher nor clearly lower than that of the 25 not mentioned (19/100) or of the full set covered by the experiment (19/100) — a small difference, and one drawn from a sub-sample of just 7 cases, that does not support the claim that optimizing a site today changes the probability of being named. In at least one case, a single provider was the only one mentioned by all three models on the same question, despite holding one of the lowest scores in its category — a sign that, when a mention does occur, it may owe as much to brand recognition or track record built outside the website as to the quality of the site itself. When Perplexity mentioned a provider from the sample, in most cases it did so by citing a third-party source (an official registry, a travel guide, a review site) rather than citing the provider's own site directly.
The most telling result of this work is not the average score itself, but what that score predicts: when the real question a traveler would ask is simulated, nearly eight in ten providers from the sample covered by the experiment do not show up. The GEO score stops being an abstract metric and starts describing, with reasonable fidelity, who exists and who does not exist for a prospective customer who today delegates that search to an AI.
The breakdown by criterion offers a clue as to why. The study's least deficient criterion — topical specificity — and the second least deficient — authority and trust — describe attributes many tourism providers genuinely possess: track record, certifications, years in operation, a well-delimited knowledge domain. But the most deficient criterion of all — sub-question coverage, with 95% of the sample below 30 — describes something different: whether the site explicitly and structurally anticipates and answers the concrete questions someone asks themselves before deciding. A provider can have genuine authority and not have a single one of those answers written in a way an AI can extract. That gap — between what a provider is and what its website communicates in a machine-readable way — is, in practical terms, the same gap that separates the 7 mentioned providers from the 25 invisible ones.
From a strategic standpoint, a widespread lag configures a window for early differentiation: in a scenario where practically no provider is ready, early adoption of GEO practices allows capturing visibility in a low-competition space.
This study has limitations worth making explicit; none of them invalidate its results, but rather delimit the scope of its conclusions. First, the sample is drawn from the directory of a single national trade association, so its results are not directly extrapolable to the universe of tourism providers outside that registry. Second, the measurement instrument is in-house (see 3.4); while its criteria are grounded in published research and official technical standards, and its application is consistent and reproducible, it does not constitute an industry-consensus standard. Third, this is a cross-sectional measurement at a single point in time, which does not capture temporal evolution; the figures describe the state of the sector in August 2026. Fourth, the mention experiment covered only two of the sample's categories (travel agencies and hotels, the latter only with the recommendation question) across 32 of the 42 providers, with a single capture per question and model — per the IAB framework used, this makes it a directional measurement, not a statistically representative one, useful to illustrate a pattern rather than to quantify it with precision. The comparative value of the GEO score measurement does not depend on these limitations, since all domains were evaluated with the same instrument, the same criteria, and in the same period.
The Argentine tourism industry exhibits a systematic lag in its readiness to be identified and cited by generative artificial intelligence engines. With an average score of 20 out of 100, no domain at the readiness level, and 62% of the sample showing zero structured data, the sector shows a cross-cutting weakness in the structure and markup of its content. The mention experiment translates that diagnosis into practice: nearly eight in ten providers covered are never named when the real question a traveler would ask is simulated, without the GEO score alone being enough to explain who does appear. This diagnosis, rather than a call-out, delimits an opportunity: the AI discovery layer is, in Argentine tourism, still largely unexplored.
Listed below, in alphabetical order, are the 42 domains that make up the analyzed sample, without each one's individual score. This criterion contributes methodological transparency — it enables independent reproduction of the measurement — without exposing any provider to a nominal ranking of scores.
← See all of Nostos's sectoral studies
Nostos analyzes any website using the same eight criteria from this study, free and in under a minute.
Analyze my site →