Nostos GEO
GEO Study · Nostos

Argentine E-Commerce Visibility Before Generative Artificial Intelligence Engines

Fabián Torres — Nostos  ·  Buenos Aires, Argentina  ·  August 2026
Abstract

The rise of generative artificial intelligence engines (ChatGPT, Perplexity, Google AI Overviews, among others) as a discovery layer for information introduces a visibility criterion distinct from the one governing traditional search engine ranking. This work examines the extent to which Argentina's leading e-commerce sites are prepared to be identified and cited by these engines. We analyzed 39 domains belonging to companies affiliated with the country's main e-commerce trade association, evaluated using a GEO analysis instrument that scores six criteria grounded in recent academic literature. Results show an average overall score of 35 out of 100, with no domain reaching the level considered "ready" for these engines. The weakest criterion is content structure (average 27/100), and the least weak is freshness of updates (45/100), configuring a paradox: sites update frequently but do not structure their content in a way that allows an artificial intelligence to extract and cite it. 23% of the sample shows zero technical markup. We conclude that top-tier Argentine e-commerce exhibits a systematic lag in its preparation for the generative discovery layer, representing both a risk of lost visibility and a window for early differentiation.

Average overall score
35/100
Domains at Level 1
0/39
Zero technical markup
23%
Highest score
62

1. Introduction

For two decades, a website's visibility was determined by its ranking in traditional search engines. The mass adoption of conversational assistants built on large language models has introduced a second discovery layer: a growing share of users pose their queries to a generative engine that responds synthetically, citing a small set of sources instead of offering a list of links. Under this scheme, competition for attention is no longer settled solely by position in a listing, but by the probability of being the source the model selects to construct its answer.

This shift gives rise to a specific discipline, Generative Engine Optimization (GEO), whose purpose is to maximize the probability that a site's content is retrieved and cited by an artificial intelligence engine. Unlike search engine optimization, GEO prioritizes the content's capacity to be interpreted, fragmented, and attributed programmatically.

There is evidence that this shift reduces traffic derived from traditional search and concentrates citations among a limited number of domains. However, no systematic diagnosis exists of how prepared Argentine e-commerce is for this transition. This work seeks to fill that gap by measuring a sample of representative domains from the sector.

2. Evaluation Framework

The instrument used evaluates six criteria. Each rests on sources that publish their method and their data: peer-reviewed research, official technical standards, and the engines' own official documentation. Every reference is linked at the end of this work:

Criteria 1 through 3 are grounded in Aggarwal et al. (Princeton, KDD 2024); criterion 4 in the Schema.org / JSON-LD specification, with Dang et al. (University of Nantes, Semantic Web Journal 2025) as a reference for how widely the markup is adopted across the web; and criteria 5 and 6 in Chen, Wang, Chen, and Koudas (University of Toronto, Workshops of the EDBT/ICDT 2026 Joint Conference).

3. Methodology

3.1. Sample

A sample of 39 domains from companies within the Argentine e-commerce ecosystem was assembled, drawn from the public member directory of the country's main sectoral trade association. Selection sought variety across sectors (apparel, electronics, general retail, tourism, logistics, payment methods, and e-commerce infrastructure) and scale (from large operators to specialized merchants). The full list of domains is presented in Appendix A.

3.2. Instrument

The analysis was performed with Nostos, a GEO evaluation engine that assigns each domain a score from 0 to 100 on each of the six criteria described, as well as an overall score and a readiness level across four categories.

LevelDescriptionRange
Level 1Ready for AI80–100
Level 2Targeted adjustments55–79
Level 3Partial restructuring30–54
Level 4Reconstruction needed0–29

3.3. Procedure

Each domain was analyzed in a single measurement during July 2026. 40 analyses were run across 39 domains; one domain was measured twice as a consistency check and is counted once in the averages. For each domain, the instrument examines the home page, a set of additional internal pages, and the site's sitemap, from which it evaluates structural, markup, and freshness signals that are publicly accessible. Each criterion's score results from aggregating those signals at the domain level, not from a single page.

3.4. Conflict of Interest Statement

This survey was conducted with Nostos, an in-house GEO measurement instrument, conceived from the experience of an Argentine e-commerce operator with the initial purpose of evaluating its own visibility before generative artificial intelligence engines. This work extends that measurement to the sector as a whole. The instrument is also offered to third parties. This condition as an interested party is declared in the interest of transparency; the methodology and scoring criteria are made explicit so that results can be independently verified by third parties.

3.5. Reading Dynamic Content

Nostos incorporates a two-layer detection mechanism to identify content that a site builds dynamically via JavaScript, without needing to execute it. The first layer recognizes the fingerprint of more than twenty known providers of review, social media, chat, and booking widgets, by analyzing references to external and inline scripts present in the document. The second layer applies structural heuristics —empty containers with characteristic identifiers and shell patterns typical of single-page application frameworks (React, Vue, Angular, Next.js, Svelte)— as a containment mechanism for providers not catalogued in the first layer. When content is detected under this criterion, the system explicitly reports it as present but not machine-readable, rather than assuming it does not exist.

4. Results

4.1. Overall Score

The sample's average overall score was 35 out of 100. No domain in the sample reached the readiness level for generative engines. The highest score observed was 62, and the lowest was 12.

39 domains
  • Level 4 · Reconstruction (0–29) 1744%
  • Level 3 · Restructuring (30–54) 1949%
  • Level 2 · Adjustments (55–79) 38%
  • Level 1 · Ready for AI (80–100) 00%
Fig. 1 — Distribution of the 39 domains by readiness level.
051015 2 15 8 7 6 1 10–1920–2930–39 40–4950–5960–69 Overall score
Fig. 2 — Dispersion of the 39 overall scores. No domain exceeds 62; the "Ready for AI" threshold (80) remains empty.

4.2. Breakdown by Criterion

Average scores by criterion, ordered from most to least deficient, were: content structure 27/100, data density 31/100, technical markup 32/100, authority and trust 37/100, thematic specificity 43/100, freshness and updates 45/100.

0255075100 Content structureData density Technical markupAuthority and trust Thematic specificityFreshness and updates 27 31 32 37 43 45
Fig. 3 — Average by GEO criterion, from weakest to strongest.
0%25%50%75%100% Content structureTechnical markup Authority and trustData density Freshness and updatesThematic specificity 62%49%46% 36%36%15%
Fig. 4 — Prevalence of severe deficit: percentage of domains scoring below 30 on each criterion.

4.3. Specific Findings

a) Absent technical markup. 23% of domains (9 of 39) showed zero technical markup (0/100), and 40% (16 of 39) scored 5 or below. Nearly half the sample therefore lacks the structured data an generative engine requires to interpret and attribute its content.

None (0)Near-none (1–5)Low (6–40) Medium (41–60)Acceptable (61–78) 9 · 23% 7 · 18% 7 · 18% 9 · 23% 8 · 21%
Fig. 5 — Technical markup by band. None of the 39 domains exceeds 78: not even the best mark up completely.
Domains ordered by technical markup (highest → lowest)
Fig. 6 — Technical markup gap within the sample. Domains scoring 78 coexist with domains at absolute zero: the gap is not one of resources, but of decision.

b) The activity paradox. The highest-scoring criterion was freshness of updates (45), and the lowest-scoring was content structure (27). Sites remain active —updating prices, stock, and catalog frequently— but do not organize their content so that an artificial intelligence can extract and cite it. This is activity without visibility: constant motion that does not translate into presence before generative engines.

0255075100 Freshness and updates → 0255075100 Content structure →
Fig. 7 — Freshness × structure scatter plot. Each point is a domain; if activity translated into visibility, the cloud would rise diagonally. It does not.

c) Severity of problems. Of the issues detected across the 39 analyses, approximately 40% were classified as critical and 42% as important — 82% of the total does not correspond to cosmetic adjustments but to structural shortcomings requiring substantive intervention.

82% crit. + import.
  • Critical 40%
  • Important 42%
  • Minor 18%
Fig. 8 — Severity of problems detected across the 39 analyses.

d) Differences by sector. Although each subgroup's sample size is small, a trend emerges: e-commerce platform and infrastructure domains —the same providers of the technology the rest of the sector uses— register the highest average (56/100), while apparel and sporting goods, and tourism, register the lowest (28/100). The lag is not necessarily concentrated in smaller-scale sectors, but in those furthest from the ecosystem's own technical infrastructure.

0255075100 Platforms and infrastructure (5)Telecommunications (2) Retail and home (8)Marketplace (2) Electronics (3)Logistics (4) Tourism (3)Apparel and sporting goods (9)Other (3) 56 47 39 37 31 29 28 28 19
Fig. 9 — Average score by sector (trend; small subgroup samples).

5. Discussion

The observed lag is not concentrated among smaller-scale operators. The sample includes top-tier companies with in-house technical teams, whose scores fall within the same deficient range as the rest. This suggests that the gap does not stem from a lack of resources, but from the absence of a still-emerging practice: deliberate preparation of content for the generative discovery layer is not yet part of the sector's development and digital marketing routines.

The paradox between freshness and structure is especially illustrative. High freshness indicates that sites have mature catalog-update processes; however, that updating is limited to transactional data (price, availability) and is not accompanied by structured informational content —guides, answers, comparative data— which is the type of material generative engines retrieve and cite. The sector is operationally active but communicationally invisible to artificial intelligence.

From a strategic standpoint, a widespread lag configures a window for early differentiation: in a scenario where virtually no player is prepared, early adoption of GEO practices allows one to capture visibility in a low-competition space.

6. Limitations

This study has limitations worth making explicit; none of them invalidate its results, but rather delimit the scope of its conclusions. First, the sample, though diverse, is skewed toward large-scale operators and ecosystem infrastructure, so its results are not directly extrapolable to the universe of small and medium-sized businesses with their own storefronts. Second, the measurement instrument is in-house (see 3.4); while its criteria are grounded in published research and official technical standards, and its application is consistent and reproducible, it does not constitute an industry-consensus standard. Third, this is a cross-sectional measurement at a single point in time, which does not capture temporal evolution; the figures describe the state of the sector as of July 2026. The measurement's comparative value does not depend on these limitations, since all domains were evaluated with the same instrument, the same criteria, and in the same period, which makes the differences observed among them and the aggregate averages internally consistent.

7. Conclusion

Top-tier Argentine e-commerce exhibits a systematic lag in its preparation to be identified and cited by generative artificial intelligence engines. With an average score of 35 out of 100 and no domain at the readiness level, the sector shows a cross-cutting weakness in the structure and markup of its content, only slightly offset by a high update frequency that, on its own, is not enough for generative visibility. This diagnosis, rather than a criticism, delimits an opportunity: the AI discovery layer remains, in the Argentine market, largely unexplored.

References

Appendix A — Sample Composition

Listed below, in alphabetical order, are the 39 domains comprising the sample, without each one's individual score. This criterion provides methodological transparency —enabling independent reproduction of the measurement— without exposing any company to a nominal score ranking. Some corporate groups appear with more than one domain, corresponding to distinct sites analyzed independently.

How does your own site score?

Nostos analyzes any website using the same six criteria from this study, free and in under a minute.

Analyze my site →