Crawling & platforms · Original data · official documentation
4 of 14 Korean content platforms block every AI search crawler
We fetched the robots.txt of 14 Korean content platforms and checked 11 crawlers against each platform's public post URL pattern. 6 platforms allow all three AI search crawlers, 4 block all three, and Naver Cafe also blocks Googlebot.
In Korea, much of what people read about a product or a place is published on Naver Blog, Naver Cafe, Tistory and a few large communities, not on brand websites. Whether an AI search engine can cite that content depends on each platform's robots.txt.
We fetched the robots.txt of 14 of those platforms on 2026-09-13 and checked, crawler by crawler, whether a public post URL is allowed.
Finding 01
The 4 platforms that disallow AI search crawlers are all run by Naver
- Allowed
- Blocked
- No robots.txt (404)
Fetched on 2026-09-13. Hover a cell for the platform and user-agent name.
Blocking all three AI search crawlers: Naver Blog, Naver Cafe, Naver KnowledgeiN, and Naver Influencer. Allowing all three: Daum Cafe, Tistory, Brunch, velog, Ppomppu, and Namuwiki.
Finding 02
Search crawlers are allowed in 86.1% of cases; AI training crawlers in 41.7%
Share of allowed cells across the 12 platforms with a robots.txt, per crawler type.
Finding 03
Naver Cafe blocks Googlebot too
# BOT ACCESS FOR THE PURPOSES OF AI TRAINING AND RETRIEVAL-AUGMENTED GENERATION (RAG) IS STRICTLY PROHIBITED.
Naver Blog's robots.txt even disallows Yeti, Naver's own crawler. robots.txt governs outside crawlers; Naver Search reaches its own blog posts through a path outside this file.
So where should content for Korean audiences live?
- If you want AI answers outside Naver to cite you, publish where those crawlers are allowed.
- Your own domain is the one place where you set robots.txt yourself.
- If you want to appear inside Naver's answers, Naver Blog and Cafe are the core surfaces, and they disallow outside AI search crawlers.
Methodology
- Evaluated with RFC 9309: the group naming the user agent applies if present, otherwise the * group; the longest matching rule wins.
- Paths were tested against each platform's public post URL pattern. Testing only the home page can give a different answer.
- theqoo and Clien returned 404 for robots.txt. RFC 9309 treats that as unrestricted; firewalls outside the file are not captured here.
- Fetched on 2026-09-13 with the user agent MarketResearchBot/0.1, robots.txt files only.
- robots.txt is a set of crawl directives, not an access control. An allowed cell is not a record that the crawler actually fetched a page.