Nostos GEO
GEO Study · Nostos

Visibility of the Argentine Software Industry Before Generative Artificial Intelligence Engines

Fabián Torres — Nostos  ·  Buenos Aires, Argentina  ·  August 2026
Abstract

The rise of generative artificial intelligence engines (ChatGPT, Gemini, Perplexity, among others) as an information discovery layer introduces a visibility criterion distinct from the one governing traditional search engine ranking. This work examines the extent to which Argentine software development companies are prepared to be identified and cited by these engines, and complements that measurement with a second instrument: a mention test that queries three conversational models directly. We surveyed the entirety of the members of Argentina's main software trade association whose website is operational — 33 of the 42 companies listed in its public directory — evaluated using a GEO analysis instrument that scores eight criteria grounded in recent academic literature. Results show an average overall score of 31 out of 100, with no domain reaching the level considered "ready" for these engines; 55% of the companies fall into the level requiring reconstruction. The weakest criterion is sub-question coverage (22/100), and the least weak is topical specificity (47/100); both are, moreover, strongly correlated with each other, indicating that weaknesses accumulate within the same sites rather than offsetting one another across criteria. 39% of companies show zero technical markup. The mention test, run on the same 33 companies using three fixed questions put to ChatGPT, Gemini, and Perplexity, recorded no mention of any of them in any of the nine queries performed. The Argentine software industry consequently exhibits a systematic lag in its readiness for the generative discovery layer, with an immediate practical consequence: none of the censused companies appears today when a prospective buyer asks a conversational AI for providers in the sector.

Average overall score
31/100
Companies at Level 1
0/33
Zero technical markup
39%
Mentioned by AI
0/33

1. Introduction

For two decades, a website's visibility was determined by its ranking in traditional search engines. The mass adoption of conversational assistants built on large language models has introduced a second discovery layer: a growing share of users now put their queries to a generative engine that answers synthetically, citing or naming a small set of sources rather than offering a list of links. Under this scheme, competition for attention is no longer decided solely by position in a listing, but by the probability of being the source — or the company — that the model selects to build its answer.

This shift gives rise to a specific discipline, Generative Engine Optimization (GEO), whose purpose is to maximize the probability that a site's content is retrieved and cited by an artificial intelligence engine. Unlike search engine optimization, GEO favors content's ability to be interpreted, chunked, and attributed programmatically.

For an industry like software development, whose clients increasingly ask a conversational assistant for recommendations directly ("which software companies would you recommend in Argentina?"), the question is not only whether a company's site is well built, but whether that company exists, in practice, for the model answering it. This work addresses both dimensions: the technical readiness of content, measured with a GEO instrument, and the observable outcome of that readiness, measured through direct queries to three conversational engines.

2. Evaluation framework

The instrument used evaluates eight criteria, each weighted according to its relative influence on the probability of being cited by a generative engine. All are grounded in sources that publish their method and their data: peer-reviewed research, official technical standards, and the engines' own official documentation. Every reference is linked at the end of this work:

Data density, direct question-answer relevance, and authority and trust are grounded in Aggarwal et al. (KDD 2024); freshness and updates and content structure, in Chen, Wang, Chen, and Koudas (University of Toronto, EDBT/ICDT 2026); sub-question coverage, in Xie et al. (NAACL 2025); technical markup, in the Schema.org / JSON-LD specification, with Dang et al. (University of Nantes, Semantic Web Journal 2025) as a reference for how widely the markup is adopted across the web; and topical specificity, in Yu, Yang, Ding, and Sato (University of Tokyo, 2026).

3. Methodology

3.1. Population

The public member directory of Argentina's main software trade association lists 42 companies. Of these, 33 have an operational, linked website; the remaining 9 either have no site of their own or their published link does not resolve. We therefore analyzed the entire population with an available site — not a random sample of that population — which amounts to a census of the measurable universe within the directory. The full list of the 33 censused companies is presented in Annex A.

3.2. Instrument

The analysis was performed with Nostos, a GEO evaluation engine that assigns each domain a score from 0 to 100 on each of the eight criteria described, as well as an overall score and a readiness level across four categories.

LevelDescriptionRange
Level 1Ready for AI80–100
Level 2Targeted adjustments55–79
Level 3Partial restructuring30–54
Level 4Reconstruction needed0–29

3.3. Procedure

Each of the 33 companies was analyzed in a single measurement during August 2026. For each domain, the instrument examines the home page and a set of additional internal pages, from which it evaluates publicly accessible structural, markup, and update signals. Each criterion's score results from the aggregation of those signals at the domain level, not from a single page.

3.4. Conflict-of-interest statement

This survey was conducted with Nostos, an in-house GEO measurement instrument, conceived from the experience of an Argentine e-commerce operator with the initial purpose of evaluating its own visibility before generative artificial intelligence engines. This work extends that measurement to a new sector. The instrument is also offered to third parties. This condition as an interested party is declared in the interest of transparency; the methodology and scoring criteria are made explicit so that results can be independently verified by third parties.

3.5. Reading dynamic content

Nostos incorporates a two-layer detection mechanism to identify content a site builds dynamically through JavaScript, without needing to execute it. The first layer recognizes the footprint of more than twenty known providers of review, social, chat, and booking widgets, by analyzing the external and inline script references present in the document. The second layer applies structural heuristics — empty containers with telltale identifiers and shell patterns typical of single-page application frameworks (React, Vue, Angular, Next.js, Svelte) — as a containment mechanism for providers not catalogued by the first layer. When content is detected under this criterion, the system explicitly reports it as present but not machine-readable, rather than assuming it does not exist.

3.6. Mention test

As a second measurement dimension, each of the 33 censused companies was subjected to a mention test: three fixed questions, worded identically to ChatGPT, Gemini, and Perplexity, that reproduce the way a real buyer would ask an AI for a provider recommendation. The three questions come from fixed templates applied to the sector's field (custom software development):

A · Recommendation
What are the best software development companies in Argentina?
B · Purchase intent
I want to buy a custom management system in Argentina. Which software companies would you recommend?
C · Listing request
Name Argentine software development companies.

Each capture was performed under fixed conditions: a new chat, with no history or personalization; the geography ("in Argentina") written into the question itself, never inferred from location; no follow-up questions, taking the first complete answer; the same search mode noted for each model; and the date and time recorded per capture. For each of the nine question-model combinations, we recorded whether any of the 33 censused companies' names appeared in the text of the answer. The mention test protocol is aligned with the IAB framework, Measuring Visibility in the AI Era (August 2026); under that framework, this survey's measurement level is directional: a single capture per question and model, without repetition for statistical significance.

4. Results: GEO score

4.1. Overall score

The average overall score of the 33 companies was 31 out of 100. None reached the readiness level for generative engines. The highest score observed was 67, and the lowest, 1.

33 companies
  • Level 4 · Reconstruction (0–29) 1855%
  • Level 3 · Restructuring (30–54) 1133%
  • Level 2 · Adjustments (55–79) 412%
  • Level 1 · Ready for AI (80–100) 00%
Fig. 1 — Distribution of the 33 companies by readiness level.
02468 5 8 5 3 3 5 4 0–910–1920–29 30–3940–4950–5960–69 Overall score
Fig. 2 — Distribution of the 33 overall scores. No company exceeds 67; the "Ready for AI" threshold (80) remains empty.

4.2. Breakdown by criterion

Criterion averages, ordered from most to least deficient, were: sub-question coverage 22/100, direct question-answer relevance 24/100, data density 28/100, content structure 32/100, freshness and updates 33/100, technical markup 33/100, authority and trust 36/100, topical specificity 47/100.

0255075100 Sub-question coverage Direct question-answer relevance Data density Content structure Freshness and updates Technical markup Authority and trust Topical specificity 22 24 28 32 33 33 36 47
Fig. 3 — Average score by GEO criterion, weakest to strongest.
0%25%50%75%100% Direct question-answer relevance Sub-question coverage Freshness and updates Data density Technical markup Authority and trust Content structure Topical specificity 64%64%58%55% 55%42%39%24%
Fig. 4 — Prevalence of severe deficiency: percentage of companies scoring below 30 on each criterion.

4.3. Specific findings

a) Technical markup absent. 39% of companies (13 of 33) showed zero technical markup (0/100), and 45% (15 of 33) scored 5 or lower. Nearly half the sector therefore lacks the structured data a generative engine needs to interpret and attribute its content.

Zero (0)Near-zero (1–5)Low (6–40) Medium (41–60)Acceptable (61–78)High (79+) 13 · 39% 2 · 6% 3 · 9% 6 · 18% 5 · 15% 4 · 12%
Fig. 5 — Technical markup by band. A small group scores above 78 while 13 companies sit at absolute zero.
Companies ordered by technical markup (highest → lowest)
Fig. 6 — Technical markup gap within the sector. Companies scoring 88 coexist with thirteen at absolute zero: the gap is not one of resources, but of decision.

b) Deficits accumulate, they do not offset each other. The sector's two weakest criteria — direct question-answer relevance and sub-question coverage — are strongly correlated with each other (r=0.96). No company that fails to answer a buyer's direct question makes up for it by covering their follow-up questions better: both weaknesses appear together, in the same companies, rather than being spread evenly across the sample.

0255075100 Direct question-answer relevance → 0255075100 Sub-question coverage →
Fig. 7 — Scatter of direct relevance × sub-question coverage. Each point is a company; the cloud follows a tight diagonal (r=0.96), with no case of a company strong on one criterion and weak on the other.

c) Severity of the problems. Of the issues detected across the 33 analyses, 38% were classified as critical and 44% as important — 82% of the total does not correspond to cosmetic adjustments but to structural shortcomings requiring substantive intervention.

82% crit. + import.
  • Critical 38%
  • Important 44%
  • Minor 18%
Fig. 8 — Severity of the problems detected across the 33 analyses (230 issues in total).

d) Differences by sub-sector. Although each subgroup is small, a trend emerges: software architecture and consulting, the other specialized services group, and artificial intelligence register the highest average (40/100), while companies grouped under digital experience, UX, and mobile development register the lowest (16/100). The spread between sub-sectors (24 points) is narrower than the one observed in Argentine e-commerce in an earlier Nostos GEO study (37 points), suggesting a more homogeneous lag within the software industry: no sub-sector stands out substantially from the rest.

0255075100 Software architecture and consulting (4) Other specialized services (6) Artificial Intelligence (2) Management software, ERP, and licensing (4) IT services and support (6) Custom software development (8) Digital experience / UX / Mobile (3) 40 40 40 37 30 22 16
Fig. 9 — Average score by grouped sub-sector (trend; small subgroup samples).

5. Mention test: cross-check against conversational engines

Of the 33 companies censused and analyzed with Nostos, none was mentioned by ChatGPT, Gemini, or Perplexity in any of the nine queries performed (three questions across each of the three models). The result is uniform: zero mentions in each of the nine question-model combinations, without exception.

ModelA · RecommendationB · Purchase intentC · Listing request
ChatGPT0 / 330 / 330 / 33
Gemini0 / 330 / 330 / 33
Perplexity0 / 330 / 330 / 33

Censused companies mentioned, per capture. Nine captures performed in August 2026.

Faced with the same questions, the three models did name software providers in Argentina: larger-scale companies with a different corporate profile from the censused sample appeared recurrently in the answers. None of them belongs to the set of 33 companies analyzed. The absence of matches is therefore not a result of the models refraining from naming providers in the sector, but of them naming a different set of companies than the one censused.

6. Discussion

The GEO score and the mention test result are consistent with each other: a sector with an average score of 31 out of 100, with no company ready for generative engines, fails to be mentioned when queried directly. The structural weakness detected by the GEO instrument translates, in practice, into invisibility before conversational models. This does not imply a strict causal relationship — models also weigh signals external to the site itself, such as brand reputation or press coverage — but it does confirm that a content's technical readiness and its actual appearance in an AI answer are not independent phenomena.

The correlation between direct relevance and sub-question coverage (finding b) further suggests that the gap is not distributed evenly: there is no group within the censused sample that offsets a weakness with a strength in the neighboring criterion. A company that fails to answer a buyer's direct question also fails to anticipate their follow-up questions. This is a compound deficit, not strengths and weaknesses that balance each other out.

The observed lag is not concentrated among smaller-scale companies. The censused population includes companies with an established track record and regional presence, whose scores fall within the same deficient range as the rest. This suggests the gap is not due to a lack of resources, but to the absence of a practice that is still nascent: deliberately preparing content for the generative discovery layer is not yet part of the sector's development and digital marketing routines. From a strategic standpoint, a widespread lag configures a window for early differentiation: in a scenario where no censused company is ready or mentioned, early adoption of GEO practices allows capturing visibility in a space of zero competition.

7. Limitations

This study has limitations worth making explicit; none of them invalidate its results, but rather delimit the scope of its conclusions. First, the censused population corresponds to companies with an operational website within the directory of a single trade association; its results are not directly extrapolable to the universe of Argentine software developers that do not belong to that association. Second, the GEO measurement instrument is in-house (see 3.4); while its criteria are grounded in published research and official technical standards, and its application is consistent and reproducible, it does not constitute an industry-consensus standard. Third, this is a cross-sectional measurement at a single point in time, which does not capture temporal evolution; the figures describe the state of the sector as of August 2026. Fourth, the mention test took a single answer per question and model, without repetition: per the IAB framework cited in 3.6, this yields a directional measurement level, not a statistically significant estimate of mention probability over time. The measurement's comparative value does not depend on these limitations, since all companies were evaluated with the same instrument, the same criteria, and in the same period, which makes the differences observed among them and the aggregate averages internally consistent.

8. Conclusion

The Argentine software industry exhibits a systematic lag in its readiness to be identified and cited by generative artificial intelligence engines. With an average score of 31 out of 100 and no company at the readiness level, the sector shows a cross-cutting weakness in covering a real buyer's questions and in the technical markup of its content. That structural weakness has an observable direct correlation: none of the 33 censused companies was mentioned by ChatGPT, Gemini, or Perplexity across the nine queries performed, while the same models did name providers from the sector outside the census. This diagnosis, rather than a call-out, delimits an opportunity: the AI discovery layer is, within the Argentine software industry, not only largely unexplored but currently empty of the companies that make it up.

References

Annex A — Census composition

Listed below, in alphabetical order, are the 33 censused companies, without each one's individual score. This criterion contributes methodological transparency — it enables independent reproduction of the measurement — without exposing any company to a nominal ranking of scores.

← See all of Nostos's sectoral studies

How does your own site score?

Nostos analyzes any website using the same eight criteria from this study, free and in under a minute.

Analyze my site →