← Back to blog

SEO for Bilingual Websites: hreflang, Canonical and Common Mistakes

7 min read
SEOhreflangMultilingual

Running a website in two languages solves one problem and creates another. It lets you reach both audiences, but it also gives Google two sets of pages that look related, and if you do not tell it exactly how they relate, it has to guess. Sometimes it guesses correctly. Often it decides one language is a duplicate of the other, ranks the wrong version for a query, or splits the authority of your content across versions that should have been working together. This guide covers what actually needs to be configured, in plain terms, using the setup of our own bilingual site as the working example.

What hreflang actually does

The hreflang annotation tells search engines: this page has an equivalent in another language, and here is its address. It does not translate anything and it does not directly boost rankings. What it does is help Google show the right language version to the right searcher and understand that your Persian and English pages are alternates of one piece of content rather than two competing, duplicate pages.

The annotation can live in three places: an HTML link tag in the page head, an HTTP header, or inside your XML sitemap. For most sites, putting it in the sitemap is the most maintainable option, because it lives in one generated file instead of being repeated across every page template.

The three rules that matter most

Every language version needs a self-reference. The Persian page's set of alternates must include a link to itself, not only to the English version. Skipping the self-reference is one of the most common implementation mistakes.

Links must be reciprocal. If the Persian page points to the English page as an alternate, the English page must point back to the Persian page. A one-directional link is treated as invalid by Google and the pair will not be understood as connected.

Include an x-default. This tells Google which version to show a searcher whose language does not match any of your alternates, such as someone searching in French on a site that only has Persian and English. Pointing x-default at your primary language, whichever that is for your business, is the usual choice.

A working sitemap entry for one page, with both language versions declared, looks like this:

{
  url: "https://example.com/fa/services/seo",
  alternates: {
    languages: {
      fa: "https://example.com/fa/services/seo",
      en: "https://example.com/en/services/seo",
      "x-default": "https://example.com/fa/services/seo",
    },
  },
}

Canonical tags and hreflang are not the same job

A canonical tag says "this is the preferred URL for this exact content." hreflang says "here is an equivalent piece of content in another language." Confusing the two causes a specific and damaging mistake: setting the English page's canonical to the Persian URL. That tells Google the English page is not a real, separate page at all, just a duplicate of the Persian one, which can remove the English version from the index entirely. Each language version should have its own self-referencing canonical, and hreflang, not canonical, is the tool that connects them.

URL structure

Each language needs its own crawlable URL. A common and reliable pattern is a language prefix in the path, such as /fa/... and /en/..., which is straightforward to route in a framework like Next.js and easy to reason about in a sitemap. Other valid patterns exist, such as separate subdomains or country-code domains, but for most bilingual business sites a path prefix is the simplest to maintain correctly.

Avoid detecting the visitor's browser language and silently redirecting them without a stable URL for each version. If the URL changes based on who is asking, crawlers cannot reliably index either version, and a visitor cannot bookmark or share a specific language page.

lang and dir matter beyond SEO

Every page should set the HTML lang attribute to match its actual language, and a right-to-left language such as Persian needs dir="rtl" on the relevant root element. These are not hreflang, but they affect accessibility tools, browser behavior and how correctly the page renders, and getting them wrong on a Persian page is a visible, immediate problem for readers even before it becomes an SEO one.

Content should genuinely match, not just exist

hreflang assumes that the linked pages are true equivalents of each other. If your Persian page has three paragraphs and your English page has three sentences, you have technically satisfied the annotation but not its intent, and it will not perform as well as if both versions gave a visitor the same value. This does not mean a literal, word-for-word translation is required; write naturally in each language, but keep the same depth and structure so that neither audience gets a lesser version.

Search behavior differs by language, not just by translation

A literal translation of your best-performing Persian keyword is not necessarily what an English-speaking searcher types. Keyword research should be done separately for each language, because search volume, phrasing and even intent can differ. A term that is common and specific in Persian may be rare or phrased completely differently in English, and treating the English site as a mirror of the Persian one rather than its own audience is a common way bilingual SEO underperforms even when the technical setup is correct.

Where teams usually go wrong

In practice, the failures we see most often are not exotic. A site launches with only one language's canonical and hreflang set up correctly, because the second language was added later without revisiting the sitemap. A redesign changes URL paths for one language but not the other, breaking the reciprocal link silently until someone checks Search Console. Or a translated page is created as a stub, with the intention to "fill it in later," and hreflang links it as an equivalent before it actually is one. Each of these is easy to prevent with a short checklist before launch and a repeat check after any structural change to the site.

Testing your setup

After configuring hreflang, verify it rather than assume it is correct. Google Search Console's Page indexing report will show if a page is treated as a duplicate with a different canonical chosen, which is the clearest sign of a hreflang or canonical mistake. Fetch your sitemap directly and check that each URL lists the languages you expect, including the self-reference. And search for exact phrases from your content in both languages to see which version actually appears for which searcher.

A short checklist

  1. Every page has a self-referencing canonical in its own language.
  2. Every set of alternates includes a link back to itself, not only to the other language.
  3. Every hreflang link is reciprocal in both directions.
  4. An x-default is set to your primary language.
  5. Each language has its own stable, crawlable URL, not a redirect based on browser settings.
  6. lang and dir are set correctly on each page.
  7. Translated content matches in depth and quality, not just in existence.

Getting this right once, in your sitemap generation code, is far more reliable than getting it right by hand on every page. We cover the surrounding technical setup in our guide to SEO for Next.js websites and in our pre-launch technical SEO checklist. If you want your own bilingual setup reviewed, our SEO service includes a check of hreflang, canonicals and language routing as part of a full technical audit.

Related articles

Comments