Skip to content
← All writing
Ecommerce11 July 2026Updated 7 October 20269 min read14 sources

How to decide which faceted navigation pages to index

A worked method for sorting filter combinations into the few that get their own indexable page and the many that crawlers should never see, using search demand, product depth and duplication.

By Mani Bharij, SEO & AI Search Consultant

In this postHover or tap a point
0102030405

Chapter 011 min

Why the default is to keep facets out

Every shop with filters has to decide which filtered pages search engines should index. Many shops never make it on purpose. The platform or a filter app decides for them, and the result is either thousands of crawlable parameter URLs or no filtered pages in the index at all. Both cost money. The first wastes crawling on pages nobody searches for, and the second leaves real searches like "leather corner sofa" to competitors who built a page for them.

I make the decision with three tests and a fixed order of work, shown below with an example shop's filters. The ecommerce SEO guide covers where faceted navigation sits alongside categories, crawl budget and product data.

Chapter 01 / 051 min

Why the default is to keep facets out

Faceted navigation, in Nielsen Norman Group's definition, provides multiple filters, one for each aspect of the content.1 Every aspect multiplies the number of possible URLs, and that is where the cost comes from. Crawlers fetch a very large number of filtered URLs before they can tell those URLs are useless, and the time spent there slows the discovery of new pages that matter.2

So I start from the position that no filtered URL is indexable, and each one that becomes indexable has to make a case. Indexing everything and pruning later is much harder to unwind once thousands of URLs are in the index.

None of this argues for fewer filters on the page. Baymard Institute's 2025 product list benchmark, which scored more than 170 leading ecommerce sites in the US and Europe, rated 58% of desktop sites and 78% of mobile sites "mediocre" or worse on product list usability. One of its core practices is letting shoppers combine several values of the same filter type, and 14% of the sites still did not allow it.3 The filters stay for shoppers. The decision here is which of their URLs search engines get to see.

Chapter 02 / 052 min

The three tests a filtered page has to pass

I run every candidate through three tests, in this order. A combination that fails any one of them stays out of the index.

1. Search demand in its own right

The combination has to match something people type. "Grey velvet sofa" is a search. "Sofas, grey, under £800, 3 seats, sorted by newest" is a browsing session. Three sources help here:

  • Keyword Planner gives estimates of monthly searches for a keyword.4 Treat the figures as rough and compare candidates against each other more than against a fixed target.
  • The Search Console performance report can be filtered by query and by page, with clicks, impressions, CTR and average position for each.5 Filtering queries that contain a filter value, such as "velvet", against the parent category's URL shows whether Google already associates that category with the term.
  • Your own site search logs. Repeated searches for a filter value are a direct signal that shoppers think of it as a thing to shop by.

Keywords with very low search volumes are not discoverable or forecastable in Keyword Planner,4 so a missing figure is not the same as zero demand. Where the tool shows nothing, I lean on Search Console impressions and site search instead.

2. Enough products, all year

A page for "green leather sofas" with two products on it is a weak answer to the search, and it may hold none in a month. Count the products each candidate returns today, then check how that count moves through the year. Use the lowest point, because that is when the page is weakest. My own floor is enough products to fill the first screen of a listing on mobile and desktop. That is a judgement, and a shop with very high-value items can work with fewer.

Depth also decides what happens when the count falls to nothing. The documented approach is to return a 404 when a filter combination has no results, and not to redirect it.2 A 404 fits because it does not say whether the absence is temporary or permanent, while a 410 says it is likely to be permanent,6 which is wrong for a combination that may fill up again next season. For a combination you have chosen to index, I would rather it never reach zero, which is one more reason to set the floor conservatively.

3. No duplication with a page you already have

A filtered page has to show a meaningfully different set of products from every other indexable page. Typical failures:

  • The filter repeats a subcategory. If "Corner sofas" is already a category, the "shape: corner" filter on the main sofa listing produces the same set at a second URL.
  • The filter barely narrows the set. If nearly every leather sofa you stock is brown, "brown leather sofas" is almost the same page as "leather sofas".
  • The filter is a brand, and you stock one brand in that category. The brand page and the category page are then the same list.

When a duplicate turns up, keep the stronger page and point the other at it. Redirects and rel="canonical" are both strong signals for which URL becomes canonical, sitemap inclusion is a weak one, and different canonical methods should not be mixed for the same page.7

Chapter 03 / 052 min

Running an example shop through the method

Take an online furniture shop with a Sofas category. The category has subcategories for corner sofas, sofa beds and armchairs. The listing offers these filters: material, colour, number of seats, shape, brand, width, price, delivery time, "in stock only" and sort order. The demand labels and product counts below are invented for illustration. The reasoning is the part to copy.

FilterDemandDepth at lowest pointDuplicationDecision
Material: leather, velvet, fabricClear for leather and velvet. Fabric is searched mainly as a modifier.Leather 60, velvet 35, fabric 300None for leather or velvet. Fabric is most of the range.Index leather and velvet. Leave fabric unindexed.
Colour: grey, green, blue and othersClear for grey and green. Thin for the rest.Grey 90, green 14, others under 10NoneIndex grey and green. Leave the rest unindexed.
Seats: 2, 3, 4Clear for 2 and 3 seater2 seat 80, 3 seat 110, 4 seat 9NoneIndex 2 and 3 seater.
Shape: corner, chaiseClearCorner 70Corner duplicates the Corner sofas subcategorySend the filter link to the subcategory URL. Do not create a second page.
BrandSome for two brandsOne brand has 5 sofasThe larger brand already has a brand pageLink to the brand page. Do not index the filter.
Width, price, delivery timeSearched as vague modifiers, if at allVariesOverlaps heavily with everythingKeep out of the crawl.
In stock only, sort orderNoneNot applicableSame products in a different state or orderKeep out of the crawl.
Example: single filters on the Sofas category (illustrative figures)

Out of ten filters, six single-filter pages pass: leather, velvet, grey, green, 2 seater and 3 seater. One filter is redirected into an existing subcategory and one into an existing brand page. Everything else stays out.

Two-filter combinations

Only now do I look at pairs, and only pairs built from filters that passed on their own. "Grey velvet sofa" and "leather 3 seater sofa" are the kind of phrase people search, so they go through the same three tests. Depth is where most pairs fail. If the shop stocks a handful of grey velvet sofas at the low point of the year, the pair fails even with clear demand, and the better answer is to make sure the velvet page sorts and shows grey options well.

In my judgement, three-filter combinations almost never pass all three tests. The phrases exist, but the product count collapses. Treat any you are tempted by as an exception that needs a strong argument.

Chapter 04 / 052 min

Building the indexed set so the rest is easy to block

The cleanest setup separates the two groups at the URL level. Give each chosen combination a fixed path, such as /sofas/velvet/ or /sofas/grey-velvet/, with its own title, heading and a line or two of copy, and link to it from the parent category. Keep every other filter on query parameters. A short robots.txt rule set can then block the parameters without any risk of blocking the pages you want.

The faceted navigation documentation includes a robots.txt example that disallows URLs containing specific filter parameters while allowing a chosen pattern.2 That works because, under the Robots Exclusion Protocol, the most specific matching rule wins, meaning the longest one, and an allow rule beats an equivalent disallow.8 URL fragments are the other option: filters applied after a # have no effect on crawling, because Google Search generally does not support fragments in crawling and indexing.2 The fragment is not even sent to the server when the URL is requested. The browser handles it after the page has loaded.9 On a new build I would use fragments for filters that will never be indexed. On an existing site, parameters plus robots.txt is usually the less disruptive change.

Checklist0 of 6 done
Chapter 05 / 051 min

How big a shop needs to worry about this

The demand and duplication tests matter on any shop, because they decide which searches you can rank for. The crawl side matters more as the catalogue grows. Google's crawl budget guide is written for sites with around a million or more pages changing weekly, or ten thousand or more changing daily. On those sites, unneeded URLs should be blocked with robots.txt, because with noindex Google still requests the page before dropping it.14 A small shop with a handful of filters can afford to be less strict about blocking. It still benefits from building the few filtered pages that pass the tests.

Revisit the list at least once a year, and whenever the range changes. A colour that sold well two seasons ago can fall below the depth floor, and a new material can earn a page. Product variants raise a related question about which URLs to index, which I cover in the post on variant URLs, canonicals and Merchant Center.

Questions5 answered

Common questions

01How many products should a filtered page have before I index it?
There is no documented number. My rule is enough products to fill the first screen of the listing at the lowest stock point of the year. A shop selling expensive items can work with fewer, because each product carries more of the page.
02Should I use robots.txt or noindex for filter pages?
Use robots.txt for filter URLs that were never indexed, because it stops the crawling. Noindex does not help with crawling, since Google still requests the page.14 If the URLs are already indexed, use noindex first and add the robots.txt block once they have dropped out, because Google cannot see a noindex on a page it is blocked from crawling.13
03Is canonicalising every filter page to the category enough?
On a small facet space, often yes. A canonical to the unfiltered page may reduce crawling of the filtered versions over time.2 On a large one, the URLs are still crawled, so robots.txt or fragments do more.
04What should happen when an indexed filter combination runs out of products?
Return a 404 for filter combinations with no results.2 For a page you chose to index, it is better to set the depth floor high enough that this rarely happens, and to review the list when the range changes.
05Do filter pages need their own copy?
A unique title and heading are the minimum. A sentence or two that helps a shopper choose within that set is worth adding to the pages with the most demand. Long blocks of keyword text add little.
References14 sources

Sources

  • Platform docs10
  • Regulator2
  • Industry study2
  1. 01Filters vs. Facets: DefinitionsNielsen Norman GroupIndustry study
  2. 02Managing crawling of faceted navigation URLsGoogle Crawling InfrastructurePlatform docs
  3. 03Product List UX Best Practices 2025Baymard InstituteIndustry study
  4. 04Use Keyword PlannerGoogle Ads HelpPlatform docs
  5. 05Performance report (Search results): Overview and basic setupSearch Console HelpPlatform docs
  6. 06RFC 9110: HTTP SemanticsIETFRegulator
  7. 07How to specify a canonical with rel="canonical" and other methodsGoogle Search CentralPlatform docs
  8. 08RFC 9309: Robots Exclusion ProtocolIETFRegulator
  9. 09URI fragmentMDN Web DocsPlatform docs
  10. 10Keeping Content Discoverable with Sitemaps in AI Powered SearchBing Webmaster BlogPlatform docs
  11. 11Overview of OpenAI CrawlersOpenAIPlatform docs
  12. 12Robots.txt Introduction and GuideGoogle Search CentralPlatform docs
  13. 13Block Search Indexing with noindexGoogle Search CentralPlatform docs
  14. 14Crawl Budget ManagementGoogle Crawling InfrastructurePlatform docs

Written by Mani Bharij, SEO & AI Search Consultant in London. More on this subject in the Ecommerce SEO guide.

Keep reading

Questions about any of this?

LinkedIn is the easiest place to reach me. Send a message or connect, and I will reply there.