An online shop is one of the easier sites for a search engine to understand and one of the easier sites to break. The pages follow a pattern, so a few templates decide almost everything. The same templates can also generate millions of near-identical URLs from filters, sorts and variants. Most ecommerce SEO work goes into getting the right few thousand pages found and ranked while keeping crawlers away from everything else.
The guide works down the catalogue from categories to products, then covers the URLs that filters and site search create, crawl budget, stock, product data and measurement. Website migrations, international SEO and AI search each have their own pillar.
Why category pages carry most of the non-brand demand
Someone looking for running shoes types "womens trail running shoes", "waterproof hiking boots" or "running shoes for flat feet" long before they type a model number. Those are category-shaped queries. The searcher wants a choice, and the page that answers them best is a well-stocked, well-filtered listing of relevant products.
Product queries exist too, but they skew towards brand and model names. If you sell products that other retailers also stock, you compete with the manufacturer and every other stockist for those terms, often with the same description and the same images. Your category pages are where you have the most room to be different: your range, your selection, your copy and your filters.
In my judgement, on most shops that are not single-product brands, category and subcategory pages are where the bulk of non-brand organic revenue comes from. That is a pattern, so check it against your own data before acting on it, using the measurement approach set out later. Where it holds, categories deserve the most care: their own copy, a sensible product order, and a place in the site structure that makes them easy to reach.
| Page type | Typical query | What the page needs to win |
|---|---|---|
| Top-level category | "womens running shoes" | Strong internal links, a broad and well-sorted range, short useful copy |
| Subcategory | "womens trail running shoes" | Enough products to be a real choice, a clear title and heading, links from the parent |
| Indexable filtered page | "womens trail running shoes wide fit" | Proven search demand, enough matching products, a stable URL and its own title |
| Product page | "brand model 4 womens" | Unique details, stock and price that match the feed, reviews, structured data |
| Guide or buying advice | "how to choose trail running shoes" | Useful content that links down to the categories it recommends |
Building site architecture and internal links around categories
The basic shape is simple: link from menus to category pages, from category pages to subcategories, and from subcategories to all product pages.1 Google generally does not rely on URL structure to work out a site's structure and looks at the links between pages instead.1 So a tidy folder path matters far less than whether the links exist in the HTML.
Category-led architecture
- Running
- Walking boots
- Sandals
- Running
- Trainers
- Formal
- Socks
- Insoles
- Shoe care
- Use real
<a href>links for navigation, and do not rely on JavaScript events on other elements to move between pages.1 Menus that only build links after a click are a common way to hide a whole category tree from crawlers. - Link priority products and categories from places with authority, for example a best-selling product from the home page or from other content such as blog posts.1
- Keep the hierarchy shallow enough that every product is reachable from a category or subcategory listing, and avoid categories that exist only in the menu with no products behind them.
- Add sideways links where they help shoppers: related categories, "shop by" collections, and links from buying guides down into the categories they recommend.
- Decide one primary category per product for breadcrumbs, even if the product appears in several listings.
On category pages, a paragraph that explains what is in the range, how to choose between options, and any sizing or compatibility caveats helps shoppers and gives the page something distinct. Several hundred words of keyword-stuffed text pushed below the product grid helps nobody. If the copy would not survive a merchandiser reading it, cut it.
The listing itself deserves as much attention as the copy. Baymard Institute's 2025 product list benchmark, which scored more than 170 leading ecommerce sites in the US and Europe, rated 58% of desktop sites and 78% of mobile sites "mediocre" or worse on product list usability.2 That is usability research, not a ranking study. But a category page shoppers struggle to filter and scan is a weak answer to a category search, whatever position it holds.
Product pages, variants and duplicate products
Product pages tend to go wrong in two ways. One is thin or borrowed content, such as the manufacturer's description copied by every stockist. The other is duplication inside your own site, where the same product is reachable at several URLs or near-identical products compete with each other.
Making product pages worth indexing
You will not rewrite fifty thousand descriptions, and you do not need to. Prioritise the products that sell, or that people search for by name, and add what a manufacturer feed cannot: sizing advice, compatibility, what is in the box, answers to the questions your customer service team keeps getting, and real reviews. For the long tail, accurate specifications, good images and correct price and stock are usually enough.
Variants: one page or many
A t-shirt in five colours and six sizes can be one page or thirty. Google supports both. Schema.org's ProductGroup type exists for this case: a group of products that vary only in well-described ways, such as size, colour or material.3 Google's variant markup builds on it, using variesBy, hasVariant and a productGroupID (the parent SKU) to group variants together.4 For a single-page setup there must be one canonical URL for the whole group, typically the page with no variant preselected, and that each variant should still be selectable directly through a distinct URL, for example with a query parameter, so Google can crawl and identify it.4
I give a variant its own indexable page only when people search for it specifically and it differs in ways that matter, such as a colour that people name in searches or a model size with different specifications. Sizes of the same garment almost never qualify. I have written up how variant URLs, canonicals and Merchant Center fit together.
Canonicalising duplicate and near-duplicate products
The methods for indicating a canonical URL, in order of strength, are redirects, then rel="canonical", then inclusion in a sitemap.5 Robots.txt is not a canonicalisation tool, and noindex is the wrong way to choose a canonical within a single site because it removes the page from Search entirely.5 If you do not specify a canonical, Google picks the version it considers best.5
- Same product at several paths, for example under two categories: link internally to one URL and point the others at it with
rel="canonical". - Tracking and session parameters: canonicalise to the clean URL and keep the parameters out of internal links.
- Near-identical products, such as a pack of two and a pack of three: these are usually separate products with separate prices, so keep both and make the differences clear on the page.
- The same product listed twice by mistake, or replaced by a new SKU: merge them and redirect the old URL.
Filters are the largest source of crawl waste on most shops. Five filters with ten values each, combined freely and in any order, produce more URLs than any crawler needs. There are two costs: crawlers access a very large number of faceted URLs because they cannot tell in advance which are useful, and time spent on those URLs means less time for new, useful ones.6
Deciding which combinations to index
Keep the indexed set short and chosen on purpose (I have worked an example shop through which faceted pages to index). A filter combination deserves its own indexable page when all of these are true:
Everything that fails the test, which will be most combinations, should be kept out of the crawl. Sort orders, price sliders, "in stock only" toggles and multi-select combinations almost never pass.
Shoppers still need these filters, so keep them on the page and keep only their URLs away from crawlers. Baymard's 2025 benchmark counts letting users combine several values of the same filter type, such as two colours, as a core product list practice, and found 14% of sites still do not allow it.2
Keeping the rest out of the crawl
| Method | What it does | When to use it |
|---|---|---|
| robots.txt disallow | Stops compliant crawlers fetching matching URLs. The rules are a request to crawlers with no access control behind them, and where an allow and a disallow rule both match, the more specific (longer) one wins.7 It is a documented way to block faceted URLs while leaving product pages and unfiltered listings open.6 | Large facet spaces where the crawling itself is the problem |
| URL fragments (#) | Filtering based on URL fragments has no effect on Google's crawling, positive or negative.6 | New builds where filters can be applied client-side |
| rel="nofollow" on filter links | Less effective long term, because every link to the filtered pages has to carry it.6 | Only as a stopgap |
| rel="canonical" to the parent category | Consolidates signals, but the URLs are still crawled | Small facet spaces where crawling is not a concern |
| noindex | Keeps a page out of the index, but Google still has to fetch it to see the tag | Pages you need crawlable for some other reason |
For the facets you do index, use the standard & separator for parameters, keep filters in a consistent order in the URL, and return a 404 when a filter combination has no results.6 Consistent ordering is the one people miss: if ?colour=red&size=8 and ?size=8&colour=red both resolve, you have doubled the space for no benefit.
Internal site search pages
Pages generated by your own search box behave like an unbounded facet: any string a visitor or a bot types becomes a URL. In most cases I would keep them out of the crawl with robots.txt and stop linking to them in templates. If your site search logs show a phrase people use repeatedly with no matching category, the better response is to build that category or a curated collection, give it a proper URL, and link to it. Site search data is one of the best sources of category ideas you have.
Pagination and incremental loading on category pages
A category with four hundred products needs some way of showing all of them. The guidance for Google Search is specific:8
- Link each page to the next one with
<a href>links, and link back to the first page.8 - Give each page its own URL, for example with
?page=n. Google treats paginated URLs as separate pages.8 - Do not use the first page as the canonical for the whole sequence. Each page should canonicalise to itself.8
- Do not use URL fragments for page numbers, because Google ignores what comes after the
#.8 - Google no longer uses
rel="next"andrel="prev", although other search engines may.8 - Keep alternative sort orders and filtered versions of the listing out of the index with noindex or robots.txt.8
"Load more" buttons and infinite scroll can work for shoppers. Nielsen Norman Group's usability guidance treats a Load More button as a way to reduce some of the problems of classic infinite scrolling, such as a footer that can never be reached.9 For search, what matters is that the underlying paginated URLs exist and are linked. If products only appear after a JavaScript action, Google may never see them on the listing. Sitemaps and Merchant Center feeds are the fallbacks for making sure every product can still be found.8
If deep pages of a category are the only route to some products, those products are poorly linked, and better subcategories and a sensible default sort usually do more than any pagination markup.
When crawl budget matters on a large catalogue
Crawl budget gets more attention than it deserves on small shops. Google's own guide is aimed at large sites with a million or more pages changing about once a week, medium or larger sites with ten thousand or more pages changing daily, and sites with a large share of URLs reported as "Discovered - currently not indexed".10 If your pages are generally crawled the same day they are published, you do not need it.10
A shop with two thousand products and a tidy facet setup is very unlikely to have a crawl budget problem. A marketplace or a retailer with hundreds of thousands of SKUs, daily price changes and open faceted navigation very likely does. The signal to look for is new or changed products taking a long time to be crawled, visible in server logs or the Search Console crawl stats report.
What helps on a large catalogue
- Consolidate duplicate content so crawling goes to unique pages.10
- Block URLs you do not want crawled with robots.txt. Do not rely on noindex for this, because Google still requests the page before dropping it.10
- Return a 404 or 410 for permanently removed pages so Google stops requesting them.10
- Eliminate soft 404s, which keep being crawled and waste budget.10
- Avoid infinite scrolling pages that duplicate linked content and differently sorted versions of the same page.10
- Keep sitemaps current with accurate last-modified dates, and make server response times fast enough that crawling can speed up. Bing uses
lastmodto decide which URLs to recrawl or skip, asks that it reflect when the page content really changed, and ignoreschangefreqandpriority.11
Handling out-of-stock, discontinued and seasonal products
Stock changes are where operational habits most often damage search performance. Deleting a product the moment it sells out throws away the page's links and history, and redirecting everything to the home page leaves shoppers and crawlers with nothing useful. Each case needs a different answer.
| Situation | What to do with the URL | Why |
|---|---|---|
| Temporarily out of stock, coming back | Keep it live and indexable with a 200. Mark availability accurately on the page, in structured data and in the feed. Schema.org has values for this, including OutOfStock, BackOrder and PreOrder.12 Offer a back-in-stock alert and links to alternatives. | The page keeps its rankings and links, and merchant listings can show availability.13 |
| Discontinued, with a direct successor | Permanently redirect to the successor product. | A 301 tells clients the product has a new permanent URL,14 and a redirect is a strong signal to Google that the target should be canonical.5 |
| Discontinued, no close replacement, still gets traffic or links | Keep the page as an archive with specifications, a clear "no longer available" message and links to the nearest category. | The page still answers a query. Redirecting to something unrelated does not. |
| Discontinued, no replacement, no value | Return a 404 or 410. | A 410 says the removal is likely to be permanent, while a 404 does not say either way.14 Google stops requesting permanently removed pages that return either.10 |
| Seasonal product or category | Keep the same URL all year. Out of season, show what is coming and link to related ranges. | A page that exists every year builds history. One recreated each season starts again. |
Avoid soft 404s, pages that return 200 but show "this product is unavailable" with nothing else on them. Those keep being crawled.10 If a discontinued page has no value, give it a real 404. If it has value, give it real content.
For seasonal categories, such as Christmas gifts or summer furniture, publish the page well before the season so it can be crawled and linked before demand rises. Reusing the same URL each year is one of the cheapest wins in ecommerce SEO.
Product structured data, merchant listings and Google Merchant Center
Google gets product information from two places: the structured data on your pages and the product feed you submit to Merchant Center. Doing both is the recommended setup.15 Structured data helps Google understand the page and makes it eligible for product rich results. A Merchant Center feed gives more control, and it is mandatory for some surfaces, such as listings in the Shopping tab.15
Merchant listing structured data
Merchant listing markup is Product and Offer markup on pages where the shopper can buy. Only pages where a shopper can purchase the product are eligible for merchant listing experiences, not pages that link to other sites that sell it.13 The required properties are the product name, an image, and an offer with a price.13 Merchant listings can also show price, availability, shipping and returns information.13
The most common problem I see is drift: the page says one price, the markup another and the feed a third, because they are generated by different systems on different schedules. Merchant Center can be set to update its copy of your product data automatically from the website when there is a mismatch, which addresses price and availability going stale between feed updates.15 Other channels expect the same consistency: Microsoft Merchant Center's product specification says the price must match the price shown on the product's webpage.16 Generate all three from the same source data if you can.
Merchant Center free listings
Free listings let your products appear at no cost on Google surfaces, including Search, Maps, Gemini, YouTube, the Shopping tab, Images and Lens.17 To take part you add products to Merchant Center with the required attributes and follow the product policies, and products are not guaranteed to show.17 For a retailer that has never set up Merchant Center, this is usually one of the highest-value tasks available, because it opens surfaces that organic web pages alone do not reach.
Reviews
Product reviews help twice: they add unique, frequently updated text to product pages, and they can make a result eligible for review stars. Google's review snippet rules require the reviewed content you mark up to be readily available to users on the page, and aggregate ratings must include the average rating to display.18 Where an entity controls the reviews about itself, those pages are ineligible for star review features, a rule aimed at businesses reviewing themselves on their own site.18 Product reviews from customers are the usual, eligible case. Do not mark up reviews that appear only inside a widget Google cannot render, or ratings imported from somewhere the shopper cannot see.
What hosted platforms like Shopify let you change
Hosted platforms make some decisions for you. Know which ones before you plan work that the platform will not allow. Shopify is the most common hosted example: the HTTP Archive's 2025 Web Almanac detected it on 25.3% of the ecommerce sites in its mobile crawl, second to the self-hosted WooCommerce on 44.4%.19 Its documentation is specific about the limits.
- Fixed paths. Shopify's help describes
/products,/collectionsand/collections/allas fixed Shopify paths that you cannot redirect.20 Plan your URLs around the platform's structure. - Redirects only from broken URLs. Shopify says a URL redirect works only when the original URL is broken, meaning it shows a 404. If the old URL still loads a page, the redirect will not work.20 For a discontinued product, that means removing or unpublishing it before the redirect takes effect.
- Collection-aware product URLs. Shopify's Liquid
withinfilter generates a product URL in the context of a collection, and Shopify's own documentation warns that this puts the same content on separate URLs and that you should consider the SEO implications.21 Check your theme's internal links and canonical tags, and link to the plain product URL where you can. - robots.txt. Shopify generates a default robots.txt, and you can replace or extend it with a
robots.txt.liquidtemplate. Shopify describes this as an unsupported customisation and warns that incorrect use can result in loss of all traffic.22
Other platforms have their own versions of these limits: fixed URL prefixes, filter apps that generate parameter URLs you cannot control, or themes that render navigation in JavaScript. Check the platform documentation for the specific behaviour before you promise a change, and test the live HTML, not the admin settings.
Measuring ecommerce SEO by revenue and category
Traffic and rankings are inputs. For a shop, the result is revenue, and the useful view is revenue from organic search broken down by the kind of page people landed on and by product category. That tells you whether category work is paying off, which parts of the catalogue are growing, and where stock or pricing problems are hiding behind a healthy total.
Replatforming, selling abroad and AI shopping answers
Replatforming
Moving a shop to a new platform is the single most common way ecommerce sites lose organic traffic. URL formats change, filter behaviour changes, and category copy, redirects and structured data are easily lost in the move. Treat it as a site migration: map every URL that earns traffic, revenue or links, test the redirects before launch, and compare crawls of the old and new sites. The website migrations pillar covers the process in detail. On a shop, give extra attention to the category layer and to faceted URLs that you deliberately indexed, because they are the pages most likely to be dropped or renamed.
Selling in more than one country
International ecommerce adds currency, shipping, tax, language and stock differences to everything above. Decide which markets get their own localised pages and which share a version, and make sure prices, availability and structured data match each market. The international SEO pillar covers site structure, hreflang and localisation.
AI shopping answers
Assistants and AI search features increasingly answer product questions directly. How they choose which products and retailers to show is not fully documented, and it is changing quickly, so treat any confident claim about it with caution, including this one. Google has published that there are no additional requirements to appear in AI Overviews or AI Mode, that a page must be indexed and eligible to show with a snippet, and that no special structured data or AI text file is needed.24 Beyond that, keep Merchant Center and Business Profile information up to date and make sure structured data matches the visible text.24
OpenAI documents two things that matter for shops. Its OAI-SearchBot crawler is what surfaces websites in ChatGPT search, and sites that opt out of it are not shown in ChatGPT search answers, though they can still appear as navigational links.25 It also accepts product feeds for discovery in ChatGPT, with one row per purchasable item or variant and a URL that opens the product page with the variant selected where possible.26 Check that the robots.txt rules you write for facets do not also block the product and category pages you want these crawlers to reach.
In practice that means the same fundamentals: clean product data in a feed, accurate structured data, crawlable category and product pages, and real reviews. I have written up the product data AI shopping features and agents can use. The AI search pillar covers how assistants choose sources more broadly.
Common questions
01Should I write long descriptions on category pages?
02Should faceted navigation pages be indexed?
03What should I do with a product that is out of stock?
04Does my shop have a crawl budget problem?
05Should each product variant have its own URL?
06Do I need Google Merchant Center if I already have product structured data?
07Is there special optimisation for AI shopping results?
Sources
- Platform docs19
- Regulator4
- Industry study3
- 01Help Google understand your ecommerce website structureGoogle Search CentralPlatform docs
- 02Product List UX Best Practices 2025Baymard InstituteIndustry study
- 03ProductGroupSchema.orgRegulator
- 04Product variant structured data (ProductGroup, Product)Google Search CentralPlatform docs
- 05How to specify a canonical URL with rel="canonical" and other methodsGoogle Search CentralPlatform docs
- 07RFC 9309: Robots Exclusion ProtocolIETFRegulator
- 08Pagination, incremental page loading, and their impact on Google SearchGoogle Search CentralPlatform docs
- 09Infinite Scrolling: When to Use It, When to Avoid ItNielsen Norman GroupIndustry study
- 10Crawl Budget ManagementGoogle Crawling InfrastructurePlatform docs
- 11Keeping Content Discoverable with Sitemaps in AI Powered SearchBing Webmaster BlogPlatform docs
- 12ItemAvailabilitySchema.orgRegulator
- 13Merchant listing (Product, Offer) structured dataGoogle Search CentralPlatform docs
- 14RFC 9110: HTTP SemanticsIETFRegulator
- 16Products Resource (Microsoft Merchant Center Content API)Microsoft AdvertisingPlatform docs
- 17Free listings for productsGoogle Merchant Center HelpPlatform docs
- 18Review snippet (Review, AggregateRating) structured dataGoogle Search CentralPlatform docs
- 19Ecommerce: 2025 Web AlmanacHTTP ArchiveIndustry study
- 20Creating and managing URL redirectsShopify Help CenterPlatform docs
- 21Liquid filters: withinShopifyPlatform docs
- 22Editing robots.txt.liquidShopify Help CenterPlatform docs
- 23Measure ecommerceGoogle AnalyticsPlatform docs
- 24AI features and your websiteGoogle Search CentralPlatform docs
- 25Overview of OpenAI CrawlersOpenAIPlatform docs
- 26Products: Agentic CommerceOpenAI DevelopersPlatform docs