Why Pages Don't Get Indexed: A Case Study of 32 Pages
A month after launching this website, we opened Google Search Console and looked at the indexing report. The site has 32 pages, 16 in Persian and 16 in English. Google had indexed 24 of them. That is 75 percent, which means 8 pages were sitting outside the index, invisible to anyone searching.
We work on SEO for a living, so it would be tempting to skip this and only publish success stories. But most guides on indexing describe the theory in the abstract, and the theory is much easier to follow with a real example. This article walks through what we found on our own site, how we found it, what we changed and which shortcuts we are avoiding. We will update it once Google has recrawled the changes.
What "not indexed" actually means
Being crawled and being indexed are two separate things. Google can visit a page, read it and still decide not to add it to its index. Google has been clear that it does not index every page it finds, and a new site should expect some pages to wait. The Page indexing report in Search Console groups these pages by reason, and the reason is the most useful piece of information you have. These are the statuses you will see most often.
Discovered, currently not indexed. Google knows the address exists, usually from your sitemap or a link, but has not crawled it yet. On a brand-new site this is often just a matter of time, although it can also mean Google does not expect much value from the section.
Crawled, currently not indexed. Google visited the page and decided not to index it, at least for now. This is the status that most often points to a content problem: pages that are thin, very similar to others or not clearly useful to a searcher.
Duplicate, Google chose a different canonical. Google considers the page a copy of another URL and picked that other URL as the main version. This one points to a duplication or canonical problem.
Alternate page with proper canonical tag. Usually fine. It means the page correctly declares another page as its preferred version.
Excluded by noindex, blocked by robots.txt, page with redirect, soft 404. These are technical states, and they are usually deliberate or the result of a configuration mistake you can fix directly.
The key lesson is to read the reason before you do anything. Two pages that are both "not indexed" can need completely opposite fixes.
Step one: rule out technical blockers
We began with the boring checks, because a technical blocker is the cheapest problem to fix and the most embarrassing to miss. We checked that robots.txt did not block anything important, that no page carried a noindex directive, that every page had a self-referencing canonical, that the sitemap listed only working pages, and that the pages returned normal responses for both Persian and English versions.
The foundation was sound. Nothing was blocked and nothing was accidentally excluded. The audit did turn up smaller problems, and we fixed them anyway. Eleven pages had the brand name repeated twice in the title. Four pages had no social sharing image. The sitemap did not tell Google which Persian page matched which English page. A shared layout also carried a default canonical that would have pointed any future page without its own canonical at the homepage. None of these explained why a page was missing from the index, and it is worth being honest about that. Fixing small technical issues feels productive, but it would not have moved our number.
Step two: read the reason Google gives
The report is where the answer was. In our case, the reason behind the pages that were left out came down to content quality and depth, not to a technical fault. Search Console does not literally write "this page is thin", so the diagnosis is ours to make: it shows the status, and we have to judge the page honestly against it.
Step three: measure your own content honestly
We counted the words on every page. The result was uncomfortable. Our blog articles contained between roughly 130 and 520 words of body text, and most were under 350. The service pages had about 300 to 380 words of text that was unique to each page. Each Persian page also had an English counterpart that said almost the same thing, so we were not offering Google two different pieces of content in the way we thought we were.
Word count is not a ranking factor, and adding words to a page does not make it better. Yet short pages are a symptom. A 250-word article on "technical SEO" competes with pages that cover the same topic in depth, with examples, steps and specifics. Google has almost no reason to include a shallow version in its index when better ones already exist. Our articles were correct, but they were generic, they contained no first-hand detail and they did not answer the follow-up question a reader would have after the first paragraph.
What we changed
The fix was to make each page worth indexing. We did the following.
- We rewrote all seven articles in both languages, so each now runs between roughly 1,000 and 1,500 words. The extra length came from practical steps, reasons, common mistakes and examples, not from padding.
- We added a related-articles section and contextual links between articles and service pages, so that pages support each other and are easy for a crawler to reach.
- We built pages that a business site was missing: a dedicated website support service, a page that lists all services, and an About page.
- We gave every page its own clear title and description, matched to the words people actually search for.
- We added hreflang annotations to the sitemap and fixed the smaller technical issues from step one.
- We added an estimated reading time to each article, because it sets expectations for the reader.
What we are deliberately avoiding
There are several tempting shortcuts, and we are staying away from all of them. We are not submitting every page for indexing over and over. The request tool has a limited quota, and asking for indexing does not change Google's opinion of a page. We are not buying links to push the pages. We are not deleting the weak pages in a panic, and we are not adding noindex to hide the problem. Removing or excluding pages can be a legitimate choice when a page has no purpose, but here the pages covered topics that our audience cares about, so improving them was the better path. And we are not changing the whole site again every few days. Google needs time to recrawl, and constant changes make it impossible to know what worked.
What to expect
After improving a page, the honest answer to "how long?" is that it varies. Google may recrawl within days or take several weeks, and even a much better page is not guaranteed to be indexed. We will measure the same simple ratio each week, indexed pages out of total pages, and record the date of each change so that cause and effect stay clear. When the numbers settle, we will add the result to this article, whether it is good or bad.
A checklist you can use on your own site
- Open the Page indexing report and write down the total, the indexed count and the reason for each excluded group.
- Rule out technical blockers: robots, noindex, canonicals, redirects, sitemap and server errors.
- Inspect two or three affected pages with the URL Inspection tool and note what Google says it saw.
- Count the real words on the affected pages and compare them with the pages that were indexed.
- Ask honestly what a reader gets from each page that they could not get elsewhere.
- Improve the weakest pages first, add internal links to them and make sure they appear in the sitemap.
- Wait a few weeks, then compare the indexed ratio again.
If you want a starting point, our technical SEO checklist before launch covers the technical side in detail, and the free SEO analyzer gives a quick first look at any page. Speed also matters for how efficiently Google crawls your site, which we explain in our guide to site speed and Core Web Vitals.
The takeaway
Our indexing problem was not a bug. It was a quality problem, and it was ours. The pages were technically correct but did not give Google or a reader a strong reason to keep them. If a share of your pages is missing from the index, read the reason first, be honest about the content, and improve the pages before you reach for tricks. If you would like an expert to look at your own site, our SEO service begins with exactly this kind of diagnosis.
Related articles
- Why Site Speed Affects Conversions and Google RankingsHow Core Web Vitals (LCP, INP, CLS) affect conversions and Google rankings, and how to measure and improve them on a business website.
- A Technical SEO Checklist Before You LaunchThe technical SEO items to check before launch: indexing, metadata, canonicals, sitemap, structured data, speed and multilingual setup.
- Next.js or WordPress for Your Business Website?An honest Next.js vs WordPress comparison for business sites: cost, SEO, speed, security, editing, hosting and languages, plus how to decide.
