Semalt series

Ecommerce at Scale: Crawl Budget, Facets and Feeds with Semalt

What changes at 50,000 URLs: controlling pagination and filters, where crawl budget disappears, and how organic search connects to your marketplace feed.

Updated: 2026-08-11 10 min read 2.130 words
An online store warehouse with products ready for shipping

Key takeaways

  • At scale, SEO stops being about pages and becomes about templates and rules. One template decision affects tens of thousands of URLs.
  • Faceted navigation is where crawl budget dies. Decide which filter combinations may be indexed before the crawler decides for you.
  • Category pages, not product pages, carry most commercial search demand — and they are usually the thinnest part of the site.
  • Marketplace presence and organic search are not alternatives. The mistake is letting the feed and the site tell different stories.

A shop with four hundred products and a shop with forty thousand are not the same discipline. Below a few thousand URLs, you can look at the site page by page and fix what is wrong. Above that, individual pages stop mattering — what matters is the rules that generate them, and a single careless rule can bury the entire catalogue in near-identical URLs that consume crawling and rank for nothing.

This article covers ecommerce SEO at that scale using the Semalt platform: how to configure the crawl so it produces signal, where crawl budget actually goes, how to handle filters and pagination, and how organic work connects to the marketplace feeds that most Greek shops also depend on.

What changes when the catalogue grows

1 ruleA single template condition can create or remove tens of thousands of URLs
5–8%Share of a large catalogue that typically produces almost all organic revenue
3 filtersEnough combinations to generate more URLs than the shop has products

Illustrative proportions from the patterns described below — measure the actual distribution on your own catalogue before acting on it.

Three things break as a catalogue grows. First, crawl coverage becomes a real constraint: search engines will not fetch everything, so what they fetch becomes a strategic question. Second, duplication stops being an editing mistake and becomes a structural output of the system. Third, reporting collapses into averages, because a visibility number covering forty thousand URLs cannot tell you which part of the shop is growing.

Configuring the crawl so it produces signal

Running an unconfigured crawl on a large shop produces a report with a hundred thousand findings, which is the same as producing none. Three settings turn it into something usable.

  1. Sample by template, not by size

    Crawl a few thousand URLs chosen to cover every template — home, category, subcategory, product, filter result, search result, blog, checkout-adjacent pages. Template-level bugs appear just as clearly in a sample as in a full crawl, and the sample finishes in minutes instead of days.

  2. Decide parameter handling first

    List the query parameters the shop uses and classify each one: navigational (should be crawled), filtering (usually should not), tracking (never). Without this the crawler will spend itself enumerating colour and size combinations while never reaching the categories that matter.

    See also: From the Greek Market to International.

  3. Enable rendering only where it is needed

    Many shops render product listings client-side but serve the rest as HTML. Rendering everything is slow and expensive; rendering nothing produces a report claiming your category pages are empty. Check one URL of each template manually before deciding.

  4. Group tracked keywords by category, not alphabetically

    The reporting question at scale is never "how are our keywords doing". It is "which departments are growing and which are declining", and that only works if the keyword set mirrors the commercial structure of the shop.

Faceted navigation: where crawl budget disappears

Filters are the defining technical problem of large ecommerce. A category with a colour filter, a size filter and a price filter generates every combination of those filters as a distinct URL. Three filters with modest option counts can easily produce more URLs than the shop has products, and each of those URLs contains a subset of content that already exists elsewhere.

The decision framework is simpler than the implementation:

Filter typeExampleTreatment
Has real search demand"waterproof running shoes"Indexable landing page with its own copy and title
Useful, no search demandSort by price, items per pageBlocked from crawling; canonical to the clean category
Combination of two demanded facets"black waterproof running shoes"Case by case — index only if the combined term has genuine volume
Three or more facets combinedColour + size + brand + priceNever indexable, no exceptions
Tracking and session parametersutm, gclid, session idsBlocked and canonicalised, always

The check worth running today. Compare the number of URLs the crawler found against the number of products in the catalogue. If the crawl found ten times more URLs than you have products, the filter rules are generating the difference — and search engines are spending your crawl allowance on it instead of on your category pages.

Category pages are the commercial asset

Most shops invest their content effort in product pages and their marketing effort in the homepage, leaving category pages as bare grids of thumbnails. That is backwards. Commercial search demand concentrates on category-level terms — people search for "running shoes" far more than for any individual model — and category pages are usually the thinnest, least differentiated pages on the site.

A category page that competes has four things a bare grid does not: a title and heading that match how customers actually phrase the category, a genuinely useful introduction that answers the questions buyers ask before choosing, internal links to the subcategories and guides that support the decision, and enough structural signals for the page to be understood as a category rather than a search result.

Where to start on a large catalogue. Sort your categories by revenue, take the top ten, and treat them as ten landing pages rather than as navigation. That is usually two weeks of work and it addresses the pages carrying the majority of commercial demand. Doing the same exercise across four hundred categories is a project nobody finishes.

Product pages: duplication, variants and out-of-stock

Three recurring problems, all structural rather than editorial.

Manufacturer descriptions. If your product text is the supplier's text, it exists on every competing shop and on the marketplace listing as well. Rewriting forty thousand descriptions is not realistic; rewriting the two hundred that generate meaningful revenue is, and the difference in outcome is substantial.

Variants. Size and colour variants as separate indexable URLs create large-scale near-duplication. In most catalogues the right pattern is one canonical product page with selectable variants, unless a specific variant has independent search demand.

Discontinued products. Deleting the URL destroys accumulated authority and produces a 404 for anyone arriving from an old link. Keeping the page live with no stock and no alternative is equally bad. The workable pattern is to keep the URL, state the status honestly, and route the visitor to the closest equivalent or the parent category.

At forty thousand URLs, nobody is optimising pages any more. You are writing rules and auditing what the rules produced.The mental shift that separates ecommerce SEO from everything else

Marketplaces and feeds: complement, not competitor

Almost every Greek shop of scale sells through comparison platforms as well as through its own site, and the two channels are often managed by different people who never compare notes. That creates avoidable problems.

What alignment looks like

  • Titles in the feed use the same phrasing as your category and product pages
  • Pricing and availability match what the site shows, always
  • Feed categories map cleanly to site categories
  • Bestsellers on the platform get priority in organic content work

What misalignment produces

  • Marketplace listing outranks your own product page
  • Price mismatches that cost trust and conversions
  • Duplicate-looking listings competing for the same intent
  • Two teams optimising for different keyword sets

The most useful practical habit is to look at platform bestsellers as keyword research. Products that sell well on a comparison platform have demonstrated demand, and that demand exists in organic search too. Prioritising organic work around proven sellers is a far better use of effort than optimising the catalogue in SKU order.

There is a full walkthrough in The First 30 Days with Semalt.

Performance under load, and why lab scores mislead shops

Speed testing on ecommerce has a particular failure mode: the test runs against a category page with a warm cache, no session, no cart, no personalisation and no third-party scripts firing, and reports a comfortable score. Real customers arrive with a session, a populated cart, a consent banner, a chat widget, a review script and a remarketing tag, on a mobile connection, during the promotion that put the server under load in the first place.

Field data resolves the argument. It reflects what actual users experienced over a rolling window, including everyone on an old device on a bad connection during your busiest hour. When lab and field disagree on a shop, field is right, because field is the population that generates revenue.

We cover this in detail in From Rankings to Revenue.

Two structural culprits appear repeatedly on Greek shops. The first is the image weight of category grids: forty product thumbnails served at full resolution turn a fast template into a slow one, and this is almost always fixable with correct sizing and modern formats rather than with a platform migration. The second is third-party script load, where each individually harmless tag compounds into a page that takes seconds to become interactive. Audit the tag list annually; there is usually something in it that nobody has used for a year.

Time this work before the commercial peak, not during it. A shop that fixes its category template in September enters the November promotional period with a faster site and more crawl efficiency. A shop that discovers the problem during Black Friday discovers it as lost revenue.

Reporting at scale

A single visibility number for a large shop is useless — it will be dominated by whichever category has the most tracked terms. Report by department instead, and the picture becomes actionable immediately: three departments growing, one flat, one declining, with the declining one explained by either a competitor, a technical change or a stock problem.

Two supporting metrics are worth including. The share of the catalogue that receives any organic impressions at all, which reveals how much of the shop is effectively invisible. And the ratio of category-page to product-page traffic, which tells you whether your architecture is doing the work or whether every visit is arriving through a single deep page and bouncing.

Where to start if the catalogue is already large

Do not begin with the audit list. Begin with two numbers: how many URLs exist versus how many products exist, and which ten categories generate the most revenue. The first tells you whether you have a filter problem, which is the most expensive structural issue in ecommerce. The second tells you where any content investment should go.

Everything else — variants, thin descriptions, out-of-stock handling, feed alignment — is real work, but it is second-order compared to a crawl that never reaches your commercial pages because it is enumerating colour filters.

If you have never measured that ratio on your own shop, it takes about twenty minutes. Open the dashboard, configure a template-sampled crawl, and compare what it found against your product count.

Frequently asked questions

Should filter pages be indexed?

Only when the filtered term has genuine search demand of its own. "Waterproof running shoes" deserves an indexable landing page with its own title and copy. "Black, size 42, under 80 euro, sorted by price" does not, and combinations of three or more facets should never be indexable. The test is whether a customer would actually type the phrase, not whether the filter is useful in the interface.

What do we do with products that are permanently out of stock?

Keep the URL, state the status clearly, and offer the nearest equivalent or a link back to the parent category. Deleting the page throws away whatever authority and links it accumulated and produces a 404 for anyone arriving from an old reference. Leaving it live with no stock and no alternative is a dead end for the visitor and a weak signal for the site.

Is it worth rewriting supplier product descriptions?

For the products that generate real revenue, yes. For the whole catalogue, almost never — it is a project measured in years that will not finish. Sort products by revenue, rewrite the top segment properly, and accept that the long tail carries manufacturer text. Your category pages are usually a better investment than the tail of the product catalogue anyway.

Does selling on a comparison platform hurt our own SEO?

Not inherently, but misalignment does. Problems arise when prices differ between channels, when the platform listing is more informative than your own page, or when two teams optimise for different keyword sets. Aligned, the platform is a distribution channel and a source of proven demand data. Unaligned, it competes with you for your own brand and product queries.

Try it

Open your Semalt dashboard

Site audits, rank tracking, competitor data and reporting in one place. Sign in and you will be looking at real numbers for your own domain within minutes.

Sign in to Semalt

Or browse the service overview at semalt.com.