Every shop with filters has to decide which filtered pages search engines should index. Many shops never make it on purpose. The platform or a filter app decides for them, and the result is either thousands of crawlable parameter URLs or no filtered pages in the index at all. Both cost money. The first wastes crawling on pages nobody searches for, and the second leaves real searches like "leather corner sofa" to competitors who built a page for them.
I make the decision with three tests and a fixed order of work, shown below with an example shop's filters. The ecommerce SEO guide covers where faceted navigation sits alongside categories, crawl budget and product data.
Why the default is to keep facets out
Faceted navigation, in Nielsen Norman Group's definition, provides multiple filters, one for each aspect of the content.1 Every aspect multiplies the number of possible URLs, and that is where the cost comes from. Crawlers fetch a very large number of filtered URLs before they can tell those URLs are useless, and the time spent there slows the discovery of new pages that matter.2
So I start from the position that no filtered URL is indexable, and each one that becomes indexable has to make a case. Indexing everything and pruning later is much harder to unwind once thousands of URLs are in the index.
None of this argues for fewer filters on the page. Baymard Institute's 2025 product list benchmark, which scored more than 170 leading ecommerce sites in the US and Europe, rated 58% of desktop sites and 78% of mobile sites "mediocre" or worse on product list usability. One of its core practices is letting shoppers combine several values of the same filter type, and 14% of the sites still did not allow it.3 The filters stay for shoppers. The decision here is which of their URLs search engines get to see.
The three tests a filtered page has to pass
I run every candidate through three tests, in this order. A combination that fails any one of them stays out of the index.
1. Search demand in its own right
The combination has to match something people type. "Grey velvet sofa" is a search. "Sofas, grey, under £800, 3 seats, sorted by newest" is a browsing session. Three sources help here:
- Keyword Planner gives estimates of monthly searches for a keyword.4 Treat the figures as rough and compare candidates against each other more than against a fixed target.
- The Search Console performance report can be filtered by query and by page, with clicks, impressions, CTR and average position for each.5 Filtering queries that contain a filter value, such as "velvet", against the parent category's URL shows whether Google already associates that category with the term.
- Your own site search logs. Repeated searches for a filter value are a direct signal that shoppers think of it as a thing to shop by.
Keywords with very low search volumes are not discoverable or forecastable in Keyword Planner,4 so a missing figure is not the same as zero demand. Where the tool shows nothing, I lean on Search Console impressions and site search instead.
2. Enough products, all year
A page for "green leather sofas" with two products on it is a weak answer to the search, and it may hold none in a month. Count the products each candidate returns today, then check how that count moves through the year. Use the lowest point, because that is when the page is weakest. My own floor is enough products to fill the first screen of a listing on mobile and desktop. That is a judgement, and a shop with very high-value items can work with fewer.
Depth also decides what happens when the count falls to nothing. The documented approach is to return a 404 when a filter combination has no results, and not to redirect it.2 A 404 fits because it does not say whether the absence is temporary or permanent, while a 410 says it is likely to be permanent,6 which is wrong for a combination that may fill up again next season. For a combination you have chosen to index, I would rather it never reach zero, which is one more reason to set the floor conservatively.
3. No duplication with a page you already have
A filtered page has to show a meaningfully different set of products from every other indexable page. Typical failures:
- The filter repeats a subcategory. If "Corner sofas" is already a category, the "shape: corner" filter on the main sofa listing produces the same set at a second URL.
- The filter barely narrows the set. If nearly every leather sofa you stock is brown, "brown leather sofas" is almost the same page as "leather sofas".
- The filter is a brand, and you stock one brand in that category. The brand page and the category page are then the same list.
When a duplicate turns up, keep the stronger page and point the other at it. Redirects and rel="canonical" are both strong signals for which URL becomes canonical, sitemap inclusion is a weak one, and different canonical methods should not be mixed for the same page.7
Running an example shop through the method
Take an online furniture shop with a Sofas category. The category has subcategories for corner sofas, sofa beds and armchairs. The listing offers these filters: material, colour, number of seats, shape, brand, width, price, delivery time, "in stock only" and sort order. The demand labels and product counts below are invented for illustration. The reasoning is the part to copy.
| Filter | Demand | Depth at lowest point | Duplication | Decision |
|---|---|---|---|---|
| Material: leather, velvet, fabric | Clear for leather and velvet. Fabric is searched mainly as a modifier. | Leather 60, velvet 35, fabric 300 | None for leather or velvet. Fabric is most of the range. | Index leather and velvet. Leave fabric unindexed. |
| Colour: grey, green, blue and others | Clear for grey and green. Thin for the rest. | Grey 90, green 14, others under 10 | None | Index grey and green. Leave the rest unindexed. |
| Seats: 2, 3, 4 | Clear for 2 and 3 seater | 2 seat 80, 3 seat 110, 4 seat 9 | None | Index 2 and 3 seater. |
| Shape: corner, chaise | Clear | Corner 70 | Corner duplicates the Corner sofas subcategory | Send the filter link to the subcategory URL. Do not create a second page. |
| Brand | Some for two brands | One brand has 5 sofas | The larger brand already has a brand page | Link to the brand page. Do not index the filter. |
| Width, price, delivery time | Searched as vague modifiers, if at all | Varies | Overlaps heavily with everything | Keep out of the crawl. |
| In stock only, sort order | None | Not applicable | Same products in a different state or order | Keep out of the crawl. |
Out of ten filters, six single-filter pages pass: leather, velvet, grey, green, 2 seater and 3 seater. One filter is redirected into an existing subcategory and one into an existing brand page. Everything else stays out.
Two-filter combinations
Only now do I look at pairs, and only pairs built from filters that passed on their own. "Grey velvet sofa" and "leather 3 seater sofa" are the kind of phrase people search, so they go through the same three tests. Depth is where most pairs fail. If the shop stocks a handful of grey velvet sofas at the low point of the year, the pair fails even with clear demand, and the better answer is to make sure the velvet page sorts and shows grey options well.
In my judgement, three-filter combinations almost never pass all three tests. The phrases exist, but the product count collapses. Treat any you are tempted by as an exception that needs a strong argument.
Building the indexed set so the rest is easy to block
The cleanest setup separates the two groups at the URL level. Give each chosen combination a fixed path, such as /sofas/velvet/ or /sofas/grey-velvet/, with its own title, heading and a line or two of copy, and link to it from the parent category. Keep every other filter on query parameters. A short robots.txt rule set can then block the parameters without any risk of blocking the pages you want.
The faceted navigation documentation includes a robots.txt example that disallows URLs containing specific filter parameters while allowing a chosen pattern.2 That works because, under the Robots Exclusion Protocol, the most specific matching rule wins, meaning the longest one, and an allow rule beats an equivalent disallow.8 URL fragments are the other option: filters applied after a # have no effect on crawling, because Google Search generally does not support fragments in crawling and indexing.2 The fragment is not even sent to the server when the URL is requested. The browser handles it after the page has loaded.9 On a new build I would use fragments for filters that will never be indexed. On an existing site, parameters plus robots.txt is usually the less disruptive change.
How big a shop needs to worry about this
The demand and duplication tests matter on any shop, because they decide which searches you can rank for. The crawl side matters more as the catalogue grows. Google's crawl budget guide is written for sites with around a million or more pages changing weekly, or ten thousand or more changing daily. On those sites, unneeded URLs should be blocked with robots.txt, because with noindex Google still requests the page before dropping it.14 A small shop with a handful of filters can afford to be less strict about blocking. It still benefits from building the few filtered pages that pass the tests.
Revisit the list at least once a year, and whenever the range changes. A colour that sold well two seasons ago can fall below the depth floor, and a new material can earn a page. Product variants raise a related question about which URLs to index, which I cover in the post on variant URLs, canonicals and Merchant Center.
Common questions
01How many products should a filtered page have before I index it?
02Should I use robots.txt or noindex for filter pages?
03Is canonicalising every filter page to the category enough?
04What should happen when an indexed filter combination runs out of products?
05Do filter pages need their own copy?
Sources
- Platform docs10
- Regulator2
- Industry study2
- 01Filters vs. Facets: DefinitionsNielsen Norman GroupIndustry study
- 03Product List UX Best Practices 2025Baymard InstituteIndustry study
- 04Use Keyword PlannerGoogle Ads HelpPlatform docs
- 05Performance report (Search results): Overview and basic setupSearch Console HelpPlatform docs
- 06RFC 9110: HTTP SemanticsIETFRegulator
- 07How to specify a canonical with rel="canonical" and other methodsGoogle Search CentralPlatform docs
- 08RFC 9309: Robots Exclusion ProtocolIETFRegulator
- 09URI fragmentMDN Web DocsPlatform docs
- 10Keeping Content Discoverable with Sitemaps in AI Powered SearchBing Webmaster BlogPlatform docs
- 11Overview of OpenAI CrawlersOpenAIPlatform docs
- 12Robots.txt Introduction and GuideGoogle Search CentralPlatform docs
- 13Block Search Indexing with noindexGoogle Search CentralPlatform docs
- 14Crawl Budget ManagementGoogle Crawling InfrastructurePlatform docs
Written by Mani Bharij, SEO & AI Search Consultant in London. More on this subject in the Ecommerce SEO guide.