ARTICLES
Korean SEO and GEO research from original data and official documentation
DATAPREPREP articles publish only first-party data we collected ourselves and statements verified in official documentation from search engines and AI companies. Most claims about search and AI answers in Korea trace back to no source at all.
Search engine optimization (SEO)
SEO is the work of getting pages into the indexes of Google, Naver and Daum so they appear in search results. This area covers official search engine documentation and webmaster tool data (Google Search Console, Naver Search Advisor).
Google Search
Standards set in Google's official documentation: spam policies, AI Overviews controls
Naver & Daum
What Korea's portals set in their official documentation and robots.txt
Official documentation
Naver's webmaster guide mentions AI in 4 of its 55 pages
All 55 pages of Naver Search Advisor's webmaster guide, counted for AI-related words, with the Korean source sentences and our translations.
Official documentation
Naver names web documents and institutional sites as AI Briefing sources for health and public information
8 posts on Naver's official search blog, compared in date order for what Naver wrote about AI Briefing, sources and document types. Korean originals with our translations.
Original data · official documentation
Naver Blog's robots.txt blocks Naver's own crawler and allows posts for Daum's crawler
robots.txt of 10 platforms, from Naver Blog to Tistory and YouTube, evaluated for whether Daum's crawler is allowed to crawl a post URL, and compared with Daum's official help documentation.
Measurement & queries
Interpreting webmaster tool data and real user search queries
Original data
22.3% of Korean search queries were typed without spaces
1,590 Korean search queries split by word count. 22.3% were typed as one unspaced string, and the share varies by industry.
Original data
89.6% of clicks came from queries with 10 or fewer impressions
221 clicks on 15 Korean sites, split by the impression size of the query. 89.6% came from queries with 10 or fewer impressions.
Original data
Search Advisor's top-30 query report accounts for 4.1% of total impressions
For 53 Korean sites, we measured what share of total impressions the top 30 rows of the query report account for. Combined, it is 4.1%.
Original data
The data covers a 60-day reporting period, but the sites were live for 33 days on average
For 53 Korean sites, we subtracted each go-live date from the capture date. Of the 60-day reporting period, the sites were actually live for 33 days on average.
Original data
Korean searchers typed one intent, "good at it", in 11 different forms
We picked 6 search intents and counted the query variations people actually typed for each in Korean queries. The intent with the most forms had 11.
Original data
Impressions took off at 53 sites, and 11 of them held their ground
For 53 Korean sites, we scored how far each impression curve rose and how much of that rise remained at the end, by site and by industry group. The weak side of the record is included as it is.
Generative engine optimization (GEO)
GEO is the work of making pages readable and citable for generative AI answers such as ChatGPT, Claude, Perplexity and Google's AI Overviews. This area covers AI crawler access and platform controls.
How AI answers pick sources
Which sources ChatGPT, Claude, Perplexity and Google's AI features crawl, and how they select what to cite
Official documentation
Google's own AI search guide lists 5 "GEO" tasks you don't need to do
The "what you don't need to do" section of Google Search Central's generative AI guide, quoted sentence by sentence with the date we checked it.
Official documentation
Google, Bing and a controlled test give three different answers on schema and AI citations
Google's documentation, the Bing Webmaster Blog and the Ahrefs schema experiment side by side, showing where the case for adding schema diverges.
Official documentation
4 of 7 AI companies say robots.txt may not apply to their user-triggered fetchers
Crawler documentation published by OpenAI, Anthropic, Perplexity, Google, Microsoft, Apple, and Meta, compared user agent by user agent: what each one does and how far robots.txt reaches, in the companies' own words.
Crawling & platforms
What each platform's robots.txt allows and disallows, crawler by crawler
Original data
3 of 18 Korean shopping, clinic and booking platforms disallow listing pages even for Googlebot
robots.txt and llms.txt of 18 Korean shopping, beauty-clinic and booking platforms, fetched on 2026-09-14 and evaluated for listing pages.
Original data
Sitemap lines in social platforms' robots.txt range from 182 on Facebook to none on LinkedIn, TikTok, Quora, and Reddit
Sitemap lines in the robots.txt of 11 social platforms, counted and grouped by file name, plus a Googlebot check for each platform's URL types.
Original data · official documentation
4 of 14 Korean content platforms block every AI search crawler
robots.txt of 14 Korean blog, cafe and community platforms, fetched on 2026-09-13 and evaluated per crawler for public post URLs.
Browse by search journey
Articles ordered along the buyer's search journey: problem awareness, concept, implementation, vendor comparison, then performance validation.
01
Problem awareness
"We don't show up in search or AI answers"
02
Concept
Precise definitions of terms such as GEO and AI crawlers
03
Implementation
Settings and standards you run into when implementing it yourself
04
Vendor comparison
Whether the work vendors propose is backed by evidence
05
Performance validation
How to interpret and validate performance metrics
Browse by entity
Entities covered by 4 or more articles. Select one to see its articles.
robots.txt9 articles
- 3 of 18 Korean shopping, clinic and booking platforms disallow listing pages even for Googlebot
- 4 of 14 Korean content platforms block every AI search crawler
- 7 of 11 social platforms disallow public posts for every AI search crawler
- 4 of 7 AI companies say robots.txt may not apply to their user-triggered fetchers
- Naver Blog's robots.txt blocks Naver's own crawler and allows posts for Daum's crawler
- LinkedIn disallows 3 of 13 URL types for Googlebot, and the feed post URL is one of them
- Sitemap lines in social platforms' robots.txt range from 182 on Facebook to none on LinkedIn, TikTok, Quora, and Reddit
- Naver's webmaster guide mentions AI in 4 of its 55 pages
- Google's controls over AI answers sit in 4 different places, and robots.txt is only one of them
Narrower entities inside this one
OAI-SearchBot5 articles
- 3 of 18 Korean shopping, clinic and booking platforms disallow listing pages even for Googlebot
- 4 of 14 Korean content platforms block every AI search crawler
- 7 of 11 social platforms disallow public posts for every AI search crawler
- 4 of 7 AI companies say robots.txt may not apply to their user-triggered fetchers
- LinkedIn disallows 3 of 13 URL types for Googlebot, and the feed post URL is one of them
Google-Extended4 articles
- Google's controls over AI answers sit in 4 different places, and robots.txt is only one of them
- 4 of 14 Korean content platforms block every AI search crawler
- 4 of 7 AI companies say robots.txt may not apply to their user-triggered fetchers
- LinkedIn disallows 3 of 13 URL types for Googlebot, and the feed post URL is one of them
Googlebot6 articles
- Google's controls over AI answers sit in 4 different places, and robots.txt is only one of them
- 7 of 11 social platforms disallow public posts for every AI search crawler
- Instagram posts can be indexed by Google only when three account conditions are all met
- 4 of 7 AI companies say robots.txt may not apply to their user-triggered fetchers
- LinkedIn disallows 3 of 13 URL types for Googlebot, and the feed post URL is one of them
- Sitemap lines in social platforms' robots.txt range from 182 on Facebook to none on LinkedIn, TikTok, Quora, and Reddit
LinkedIn5 articles
- LinkedIn disallows 3 of 13 URL types for Googlebot, and the feed post URL is one of them
- 7 of 11 social platforms disallow public posts for every AI search crawler
- Instagram posts can be indexed by Google only when three account conditions are all met
- Sitemap lines in social platforms' robots.txt range from 182 on Facebook to none on LinkedIn, TikTok, Quora, and Reddit
- Search Console can report Google performance for 4 social platforms: Instagram, TikTok, X, and YouTube
Narrower entities inside this one
Instagram · YouTube4 articles
- Instagram posts can be indexed by Google only when three account conditions are all met
- 7 of 11 social platforms disallow public posts for every AI search crawler
- Sitemap lines in social platforms' robots.txt range from 182 on Facebook to none on LinkedIn, TikTok, Quora, and Reddit
- Search Console can report Google performance for 4 social platforms: Instagram, TikTok, X, and YouTube
Long-tail keyword5 articles
- 89.6% of clicks came from queries with 10 or fewer impressions
- 22.3% of Korean search queries were typed without spaces
- Search Advisor's top-30 query report accounts for 4.1% of total impressions
- Korean searchers typed one intent, "good at it", in 11 different forms
- Google's own AI search guide lists 5 "GEO" tasks you don't need to do