Can AI crawlers access Vietnamese business websites? Evidence from 100 company records
Can AI crawlers access Vietnamese business websites? A 100-record audit of crawler policies, access gaps, entity signals and verification steps.
Published · Updated
Quick answer
AI crawlers could access some, but not all, of the audited Vietnamese business websites. In 100 company records across 98 domains, 65 had both readable homepages and available robots files; five had robots access without readable homepages, and six had the reverse. These are historical technical observations—not measured ChatGPT recommendations.
Author
Sam
LUMESEO
SEO & GEO strategy and research
This LUMESEO study examines website access rules and visible entity signals. Its technical observations are kept separate from claims about ChatGPT recommendations.
Research scope
A reproducible convenience sample of 100 company records across 98 stored domains, selected from Wikidata. Original data cutoff: 2026-09-10 17:55 UTC+7. Reanalysis: 2026-09-19. This is not a representative national sample or a measurement of recommendation probability.
Who this is for: Business owners and SEO teams diagnosing crawler access, robots rules and website evidence.
How many Vietnamese business websites could AI crawlers access?
| Observed signal | Count | Correct denominator / interpretation |
|---|---|---|
| Readable homepage HTML | 71 | 100 sampled entity records |
| Available robots.txt | 70 | 100 sampled entity records; unavailable is not blocked |
| OAI-SearchBot policy | 6 / 58 / 1 / 5 / 30 | Allowed / partial / blocked / unspecified / unavailable; total 100 |
| Organization schema | 27 | 27/71 readable homepages; 27/100 whole sample |
| sameAs detected | 26 | 26/71 readable homepages; 26/100 whole sample |
| Sitemap available or declared | 63 | 100 sampled entity records; not a count of indexed sites |
Dataset version: 2026-09-10-r2. This English edition does not add a new measurement window.
Download the public row ledger →What is new in this analysis—and why the denominator matters
The useful question is not simply whether an AI bot is allowed. It is which observable failure stops a buyer-relevant page from being usable, and what evidence would show that the repair worked. We reanalysed the published September 10 ledger on September 19 to connect robots availability, homepage readability and entity signals. No websites were recrawled for this update; the observations are historical, not a statement of their present configuration.
The ledger contains 100 distinct Wikidata QIDs and 100 stored website URLs, but only 98 distinct stored domain strings. Three entity records point to mgallery.accor.com. Consequently, our percentages are entity-record weighted: a shared domain can contribute more than once. The original shorthand “100-site audit” must not be read as 100 independent domain-level observations. We retain the original rows rather than silently deleting two and changing every denominator.
The sampling rule selected records ordered by QID, not by a random draw, industry quota, company size or export relevance. Coverage depends on having a Wikidata business entry, a Vietnam country statement and an official website. Businesses missing those records are outside the frame. This is evidence about this cohort, not an estimate of all Vietnamese businesses, and it cannot support an industry league table.
All derived tables below use the original stored flags. We did not revalidate response bodies, rerun the robots parser or confirm that a readable page was usable by an authenticated search crawler. The new information is the relationship between recorded observations, with those limitations intact.
Finding 1: a readable robots file and a readable website are different checks
Of the 70 records with robots.txt returning 200, 65 also had a homepage marked readable and five did not. Of the 30 without robots availability, six nevertheless had a readable homepage. In other words, a robots-only dashboard would miss two different investigation paths: five records where policy could be inspected but content could not, and six where content was accessible but policy was unresolved.
The first group calls for comparing the content request with the policy request: final destination, response code, challenge page and request context. The second calls for investigating the policy endpoint separately. Neither group justifies announcing that ChatGPT is blocked. In this dataset, availability and content readability are separate observations—not interchangeable definitions of AI visibility.
“Not readable” also needs decomposition. The stored homepage statuses are 71 responses with 200, nine with 403 and 20 recorded as 0. Zero is not a real HTTP response code and does not identify a specific cause. Without retained error details, we cannot tell whether each zero reflects DNS, TLS, timeout or another collection failure. A 403 does not by itself reveal whether a WAF, an application rule or another layer made the decision.
For an owner, the acceptance test is more useful than an overall score: can the intended service or research page return the expected content under the relevant request conditions? A homepage result alone cannot answer that for a product catalogue, a booking page or an English service page.
Robots availability × homepage readability
| Stored robots observation | Readable homepage | Not readable | Total |
|---|---|---|---|
| HTTP 200 | 65 | 5 | 70 |
| Not available under original check | 6 | 24 | 30 |
| Total | 71 | 29 | 100 |
Unit: entity-linked website record. Readability uses the original detector flag; this table does not validate crawler-specific delivery.
Finding 2: search access and training preferences diverge in the same records
The original policy categories sum to 100 for each subject. For OAI-SearchBot they are six allowed, 58 partial, one blocked, five unspecified and 30 unavailable. Among the 58 partial records, 53 had readable homepages and five did not. This reinforces why “partial” cannot be turned into either a pass for every page or a site-wide rejection.
The more useful paired comparison is inside the GPTBot-blocked group. Of 14 such records, 13 were classified partial for OAI-SearchBot and one blocked for OAI-SearchBot. These are paired row observations, not a subtraction of unrelated headline totals. They show different recorded preferences for the two subjects; they do not show that any of the 13 actually received search traffic. Repeated-domain records remain in this comparison.
OpenAI documents OAI-SearchBot as search-related, GPTBot as training-related, and ChatGPT-User as user-triggered fetching. Their controls serve different purposes. A business should therefore decide its search and training preferences separately rather than treating every OpenAI-labelled request as one traffic channel. See the official crawler documentation linked with this section.
The operational follow-up is to inspect the rule that applies to the exact URL you want discovered. A partial classification may reflect restrictions on paths irrelevant to that page. Conversely, a readable homepage does not prove a restricted service path is available. Never remove protective restrictions blindly to improve an audit score.
Google-Extended needs a different interpretation again: Google describes it as a robots product token without its own separate HTTP request User-Agent, not a crawler whose visits can be counted under that name. Its setting does not control inclusion in Google Search. The original policy table records this token's rules, not observed Google-Extended requests.
Paired policy observations within GPTBot-blocked records
| GPTBot policy | OAI-SearchBot policy | Records |
|---|---|---|
| Blocked | Partial | 13 |
| Blocked | Blocked | 1 |
| Total GPTBot-blocked group | All observed search categories | 14 |
Original snapshot categories, not a fresh robots evaluation. “Partial” does not establish permission for an individual landing URL.
Finding 3: entity markup gaps overlap—but are not identical
Among the 71 readable homepages, 27 contained detected Organization markup and 26 contained at least one sameAs URL. Reporting those numbers separately hides the overlap: 22 had both signals, five Organization only, four sameAs only and 40 neither. Thus 31/71, or 43.7%, had at least one of these detected signals, while 22/71, or 31.0%, had both.
The 40/71 records with neither signal are not 40 untrustworthy companies. We measured detected markup on the sampled page, not business legitimacy, whole-site markup coverage or whether independent sources know the company. Similarly, the four sameAs-only records do not prove a broken Organization schema: the detector does not establish which schema node owns each link.
For a Vietnamese exporter, the useful repair is an identity consistency check. Does the visible page connect the trading name to the legal entity, location and relevant offering? Do linked profiles belong to that same entity? Are claims such as manufacturing capacity, certifications or target markets supported by records the buyer can verify? Those questions can reveal a missing or contradictory fact even when the markup validator returns no error.
Treat structured data as a representation of supported facts, not a substitute for them. We did not measure whether adding either field changes recommendation probability, and this cross-section provides no conversion uplift estimate.
Identity-signal overlap on the 71 readable homepages
| Organization detected | sameAs detected | Records | Share of readable rows |
|---|---|---|---|
| Yes | Yes | 22 | 31.0% |
| Yes | No | 5 | 7.0% |
| No | Yes | 4 | 5.6% |
| No | No | 40 | 56.3% |
22 + 5 = 27 Organization observations; 22 + 4 = 26 sameAs observations; all four groups total 71. Rounded percentages may not sum to exactly 100%.
Finding 4: an HTTP 200 is not proof that an AI support file is valid
Seventeen records returned 200 at llms.txt, but only seven passed the original plaintext-or-Markdown format flag. Ten successful HTTP responses therefore did not pass that format check: 10/17, or 58.8%, of the 200 group. This is a concrete reason not to report file availability from status alone.
We cannot label those ten responses as particular errors without re-inspecting their bodies. They could involve unexpected content or other detector failures; this reanalysis does not decide which. Equally, the seven passing records are not seven editorially reviewed, useful files. Format acceptance is a narrower result.
For a proposed support file, first inspect the body and content type, confirm that referenced URLs resolve, and compare its claims with the visible pages. Do not spend a sprint polishing this file while the service page is unreadable or its identity is contradictory. Google's AI-search guidance does not require special AI text files or a special schema. That guidance is about Google, not a promise about every AI product.
This dataset contains no outcome variable linking llms.txt to citations or referral visits. It cannot establish whether the file helps, harms or has no effect on traffic.
llms.txt: response success versus recorded format acceptance
| Observation | Records | Denominator |
|---|---|---|
| HTTP 200 | 17 | 100 sampled records |
| Passed original format flag | 7 | 17 HTTP-200 records |
| HTTP 200 but failed format flag | 10 | 17 HTTP-200 records |
Reuses the original format check. No manual body review or causal traffic analysis was added.
An evidence ladder: what each test can actually establish
A technical audit becomes misleading when a successful early test is presented as proof of a later outcome. A robots policy, a server response, a cited answer, a referred session and a qualified enquiry are five different kinds of evidence. This study covers policy observations and basic page signals only.
For reporting, attach each claim to its evidence object. “The page is technically readable in this check” needs the response and content check. “ChatGPT cited this page” needs a retained answer and citation URL. “The citation produced a visit” needs attribution evidence for an entry session. “That visit produced a lead” needs the conversion record. Do not fill a missing step with a model's assessment of how good the page looks.
Claim-to-evidence map for a business GEO audit
| Claim | Evidence to retain | What it does not establish |
|---|---|---|
| Policy permits the target path | Dated robots file, applicable group and exact path evaluation | Successful fetch or recommendation |
| Content is retrievable | Request context, final URL, status and expected content | Selection in an AI answer |
| Page was cited | Saved answer, test settings, timestamp and citation destination | A user clicked |
| AI-attributed visit occurred | Entry-session referrer or disclosed tagged attribution | Unique person, recommendation wording or sale |
| Business value occurred | Qualified enquiry or transaction tied to the visit | That all future AI traffic will convert |
This is a proposed evidence framework, not five outcomes measured in the 100-row cohort.
What a useful verification package should contain
Start with a URL inventory rather than a bot list. Choose the homepage, one commercial page, one evidence page and the important language variants. Record the intended canonical URL and which buyer question each page is supposed to answer. These are suggested test targets, not additional pages checked in our original audit.
For each target, save the timestamp, requested URL, redirect destination, status, content type, robots group, relevant path rule and a short body excerpt confirming the intended content. Record visible language and whether a challenge or login interrupted access. If an error occurs, save its category instead of writing only status 0. Keep response bodies privately when they might contain sensitive material; publish only necessary, sanitized evidence.
A request with a bot-like User-Agent is a diagnostic simulation, not proof that an official crawler made it. When analysing your own server logs, preserve the request evidence and use the provider's verification information where applicable. Do not bypass security protections or allow arbitrary clients just because they claim to be a search bot.
Repeat a suspected failure under a documented second request context or time. If it changes, report the difference and investigate rather than selecting the favourable result. A reproducible incident record lets engineering distinguish policy intent from transport failure and lets marketing avoid buying content work to solve a delivery problem.
Close the issue only after the exact target passes its acceptance test. Keep the original failed observation, the change made and the retest together. A screenshot of a green dashboard without request context is not an adequate handover.
How to prioritise repairs without inventing a recommendation score
The fastest useful prioritisation asks two questions: does the observed issue affect the page that serves the buyer, and can we reproduce it? A confirmed wrong destination, missing content or unintended restriction on that page deserves investigation before optional formatting enhancements. A policy label whose path impact is unknown first needs validation.
Use separate owners. Engineering can fix delivery, redirects and unintended access restrictions. Content and business teams must validate identity, offers and evidence. Analytics must define citations, entry sessions and enquiries separately. One combined “AI readiness” number obscures those responsibilities and suggests precision this dataset cannot support.
Do not interpret the legacy technical_access_score_0_to_8 field as a ChatGPT ranking score. Its component weights are not calibrated against recommendations or revenue. In this article we use observable fields and explicit denominators instead of publishing a leaderboard of the sampled businesses.
Proposed repair queue and closure criteria
| Observed issue | Next action | Closure evidence |
|---|---|---|
| Target response absent, denied or wrong | Reproduce and locate the delivery failure | Expected target content under recorded conditions |
| robots classification partial or unresolved | Evaluate the target path and intended policy | Saved applicable rule plus path-specific result |
| Identity facts missing or inconsistent | Verify entity, offering and owned profiles | Visible supported facts and matching markup |
| No known AI visits despite readable content | Run a separate fixed-question study | Saved positive and negative answers, not a promised uplift |
| Optional support file fails format check | Inspect body and references | Valid intended content; not a recommendation claim |
Editorial workflow guidance. No repair experiment or before/after uplift was measured here.
Three Vietnam-market scenarios—and the different evidence each needs
Exporter scenario, illustrative: a buyer asks for a Vietnamese supplier with a specific process and shipping market. A readable homepage may say little about the actual product, minimum order quantity, documented quality process or route to an enquiry. Audit the page owning that buying decision and the evidence behind its claims. This dataset did not evaluate exporters as a subgroup or measure those buyer prompts.
Hospitality scenario, illustrative: several properties can share a chain domain. Our repeated-domain finding shows why entity and host counts should be separated. A property-specific URL, location and booking context may matter more to the buyer than a generic chain homepage. Do not treat three entity rows on the same host as three independent crawler-policy decisions.
Local-service scenario, illustrative: a Vietnamese page can be technically readable while failing an English visitor's task. Check the actual English service page, visible service area, contact path and substantiated credentials. Translating the homepage navigation alone does not demonstrate that the decision content is available in English.
These scenarios turn the observed audit gaps into testable business questions. They are not case studies of the sampled companies, and we have not assigned any company a failure or success based on these hypothetical prompts.
A follow-up study that could test visibility rather than just access
A defensible next experiment would freeze a list of target pages and buyer questions before making changes. Include questions with no brand mention, specify language and market, record product or model and whether search was enabled, and save all answers—including those that do not cite the site. Keep access fixes separate from major content changes where practical.
Measure citation share over the fixed question set, unique cited URLs, attributed entry sessions and qualified enquiries as separate series. Decide the observation schedule in advance and repeat tests, because one answer is not a stable estimate. Keep a comparable unchanged group where feasible, while recording page age, demand and other releases that could explain differences.
If a page becomes readable and later receives a citation, report the sequence. Without an adequate comparison design, do not conclude that the access repair caused the citation. This current audit establishes neither an expected time to first citation nor a budget-to-lead conversion rate.
For immediate business use, request two deliverables from an agency: a reproducible technical issue ledger and a separate buyer-evidence plan. The first should show what is broken and how to verify the repair. The second should explain what information a buyer cannot yet substantiate. Paying for both as distinct deliverables is easier to evaluate than paying for a guaranteed AI recommendation.
Reproduce the new tables and preserve the historical baseline
The original R2 dataset remains unchanged. A separate derived file publishes the cross-tab counts, source SHA-256, source version, calculation rules and repeated-domain QIDs. This avoids presenting a September 19 reanalysis as a September 19 crawl. The original checked_at remains September 10, 2026 at 17:55 UTC+7.
To reproduce the access matrix, split rows by robots_status === 200 and homepage_readable_html === true. To reproduce the identity matrix, first retain only the 71 readable rows, then split on organization_schema_present === true and same_as_url_count > 0. To reproduce the policy comparison, filter GPTBot.policy === blocked before counting OAI-SearchBot categories.
To reproduce the support-file result, count llms_txt_status === 200 and compare that group with the llms_txt_valid_plaintext_or_markdown flag. To reproduce the domain count, count distinct stored domain strings—not QIDs and not website URLs. No new registrable-domain consolidation was performed.
Editorial update, September 19: added these cross-tabulations, clarified 100 entity records versus 98 domains, expanded the repair and verification framework, and retained the historical collection date. No new traffic, citation or recommendation outcome is claimed.
Frequently asked questions
Does this mean most Vietnamese companies block ChatGPT?
No. The convenience sample is not nationally representative. Only one record was classified as fully blocked for OAI-SearchBot; partial and unavailable are different categories, and the study did not test recommendation outcomes.
Can a company block GPTBot without applying the same policy to search?
Yes, the two controls serve different purposes. In this snapshot, 13 of the 14 GPTBot-blocked records had a partial, rather than blocked, OAI-SearchBot classification. The exact target path still needs evaluation.
Why does the article now say 98 domains rather than 100 sites?
The original ledger has 100 entity records and website URLs, but three records share mgallery.accor.com. Counting distinct stored domain strings gives 98. The published tables retain the 100-row denominator and disclose that rows are not independent domains.
Is an unavailable robots file evidence that the whole site is inaccessible?
No. Six of the 30 records without robots availability still had readable homepages in the original check. The two endpoints must be investigated separately.
Does adding Organization or sameAs guarantee recommendations?
No. The study records detected markup, not recommendation outcomes. Verify the visible company facts and the ownership of linked profiles; markup cannot create independent proof of a business's capabilities.
Are these September 19 crawl results?
No. The observations were collected on September 10. September 19 is the date of the derived analysis and editorial expansion, not a new crawl.
Sources and further reading
Apply the method to your market
Define your evidence and measurement plan before scaling production. We do not guarantee rankings or AI recommendations.