LEARN

The Chinese and English Ecosystems Cite Different Sources

Two root systems growing apart in dark soil, one shallow and widely spread, one deep and tightly bundled, each feeding its own glowing bulb, representing citation source structures that differ between the Chinese and English ecosystems.
The citation source structures of English and Chinese ecosystems are two separate systems that cannot be directly migrated or reused.

IN ONE SENTENCE

The Chinese and English citation mixes are two different sets; porting the English list across means investing where citations do not happen.

English-language citations concentrate in a handful of platform types — communities, encyclopedias, video, trade media, B2B review sites. The Chinese mix is a different set: ByteDance properties, WeChat public accounts, Zhihu, CSDN and Baijiahao. Porting the English list across means investing where citations do not happen.

OUR POSITION

This is not one extra market, it is a second plan. Channels differ, available measurement differs, even crawler identifiers differ — one plan across both usually serves neither well.

01

How the source mixes differ

On the English side, public data (Semrush, June 2025, 150,000+ citations): Reddit 40.1%, Wikipedia 26.3%, YouTube 23.5%.

There is no comparable public dataset for the Chinese side. Practitioner observation generally holds that ByteDance properties carry visible weight for Doubao, WeChat public accounts weigh heavily inside the WeChat ecosystem's AI, Zhihu and developer communities are cited across several platforms, and Baijiahao plus encyclopedias matter for Baidu's AI.

⚠️ Because comparable public research does not exist, judgements on the Chinese side depend more on running your own measurement — which is also why original data is worth more there.

02

How available measurement differs

Most English-side platforms list citations per answer, so citation share can be counted directly.

Most Chinese-side platforms do not, so citation share can only be estimated — which moves the measurement focus onto mention rate and analysis of the answer text itself.

Crawler identifiers differ too: the set to allow for the Chinese ecosystem is not the same, and some platforms publish none.

03

What that means for the plan

Each side needs its own question set, channel plan and measurement definition. Report them separately — blending hides a decline on one side.

Content is not a translation relationship either: the same decision question is phrased differently, weighted differently, and supported by different evidence types on each side.

Data behind this page

Reddit 40.1% · Wikipedia 26.3% · YouTube 23.5%

The three most-cited source types

SourceSemrush, 150,000+ citations across 5,000 keywords,2025-06

under 5%

Ceiling on any single domain's citation share on one platform

SourceEvertune, 200M prompts,2026

Sources

  1. [1]GEO: Generative Engine Optimization.Aggarwal et al., KDD 2024.2024

Updated 2026-08-10