A Czech supplier passes an audit in March and publishes the certificate page the same week. A buyer in Denmark runs his shortlist in April and never sees it, because the page was still not in the index. The gap between publishing and being findable is a commercial delay, and it is measurable.

Indexing gets talked about as a technical housekeeping matter, somewhere below page speed in the list of things to worry about. For a company whose enquiries arrive from abroad, it is closer to a logistics problem. A capability you cannot be found for is a capability you do not have, from the point of view of a purchasing department that will spend one afternoon assembling a list of candidates and never revisit it.

Timing · The cost of a delay

The certificate that nobody could find

Sourcing does not run continuously. It runs in bursts, tied to a project, a re-tender, a plant qualification or a failure at an existing supplier. When a burst happens, a small number of people spend a limited amount of time building a candidate list, and everything they find during those hours is the entire universe of options considered. Anything indexed the following week does not exist.

This is why indexing latency costs a supplier real money in a way it does not cost a shop or a news site. A retailer that gets a product page indexed late loses a fraction of a season. A testing laboratory that gets its new accreditation page indexed six weeks late may have missed the only qualification round that customer will run for three years.

  • New accreditations and standards. An audit passed is an asset only once the page describing it can be found by somebody searching for that standard and nothing else.
  • New equipment and new limits. A larger machine, a wider tolerance envelope, a cleanroom class. These are the phrases that change which briefs you qualify for.
  • Case studies with named sectors. Evidence pages are what turn a listing into a shortlist entry, and they are usually published last and linked least.
  • Event and trade-fair pages. A stand at a show in eight weeks needs its page indexed in two, not in ten. The window closes on a fixed date.
How to think about the delay. Treat time-to-index as a lead time like any other. You would not accept an unknown delivery date on a component; there is no reason to accept an unknown one on the page that wins the enquiry.
Structure · The site a crawler sees

Why a capability catalogue is hard to crawl

Websites belonging to manufacturers, laboratories and technical service firms tend to share a shape that search engines handle badly. The shape is not a blog with a hundred articles. It is a matrix.

Take a machining supplier. Processes multiplied by materials multiplied by industries produces a page count in the hundreds before anyone has written a word of news. Add a downloadable datasheet for each, a filterable parts index, and a full Czech translation of everything, and a site that a human perceives as fifteen sections is several thousand addressable URLs to a crawler.

1,000
URLs per day, per account
10,000
URLs per bulk batch
3
levels of sitemap recursion
1,000
sitemaps per job

Crawling is rationed. A search engine allocates a certain amount of attention to a domain based on its authority, its update frequency and how much waste it encountered last time. A site with thousands of near-identical permutation pages spends that ration on pages nobody needed, and the certificate page published on Tuesday sits in a queue behind them.

Waste source

Filter and sort parameters

A parts index with four filters generates combinations indefinitely. Each combination is a distinct address, and each one is crawled at least once before being judged worthless.

  • Often the single largest drain
  • Rarely visible to anyone who has not looked
Waste source

Permutation pages with one line changed

Twelve pages differing only in a material name read as one page repeated twelve times. They compete with each other and dilute the attention available to genuinely distinct pages.

  • Consolidate or differentiate properly
  • Half-measures make the problem worse
Blind spot

Datasheets living only in PDFs

The technical detail a buyer searches for is often inside a document that nothing on the site links to except a download button loaded by script.

  • Reachable in principle, undiscovered in practice
  • Worth a real HTML page alongside
Blind spot

The English half of a bilingual site

The Czech tree usually carries more internal links, more history and more inbound references, so the English tree is reached later and less often.

  • The half that earns export enquiries
  • Also the half that gets crawled last

That last point deserves emphasis because it is specific to bilingual sites and rarely mentioned. Nothing forces a crawler to treat the two language trees equally. If the English section was added later, is linked from fewer places and is updated less often, it will be visited less often, and its new pages will wait longer. The commercial importance of a directory has no bearing on how quickly it gets crawled.

Discovery · Being found before being judged

Three ways a page gets discovered, and one that fails

Before any question of ranking, a page has to be discovered. There are only a few routes, and understanding which one you are relying on tells you how long to expect to wait.

RouteHow it worksTypical delayWhere it fails on a supplier site
Internal linksA crawler already visiting the site follows a link to the new pageDays to weeks, depending on how often the site is crawledNew pages added to a section nothing links to
Sitemap filesThe page is listed in an XML file the search engine reads periodicallyUnpredictable; the file may not be re-read for some timeA stale sitemap generated once and never regenerated
External linksAnother site references the page directlyFast when it happensAlmost never happens for a capability page
Direct submissionThe URL is pushed to the search engine through an APIThe bot visit is usually promptA visit is not an inclusion; see the warning below

The failure that catches most technical sites is the orphan page. Somebody adds a capability page through the content system, links to it from a single campaign email, and never adds it to the navigation. It exists, it is well written, it names a standard that buyers search for, and no crawler has any path to it. Months later it is discovered by accident, or not at all.

A five-minute check. List every URL your content system knows about, then list every URL reachable by clicking from the homepage. The difference is your orphan set. On most bilingual technical sites it is larger than anybody expects, and it is heavily weighted towards the English tree.
Tooling · What the hub actually does

The Indexing Hub, in plain terms

The indexing section of the unified Semalt panel exists to shorten the discovery step rather than to replace it. It has two working parts, and they solve different problems.

Indexing Hub · Part one

URL tracker and bulk submission

For a known list of addresses you want a bot to visit soon.

1,000 URLs per day · per account
  • A daily budget, not an unlimited queue. One thousand URLs a day per account, which is a planning constraint rather than a limitation to work around.
  • Batches of up to ten thousand. A bulk submission accepts up to 10,000 URLs at once, drawn down against the daily budget rather than processed instantly.
  • A log per address. Bot visit with a timestamp, status, and error detail where there is one, alongside live counters for submitted, found and failed URLs.
1,000
daily submission budget
10,000
maximum batch size
2
bots reached via IndexNow
Indexing Hub · Part two

Sitemap submission and parsing

For a catalogue whose full extent nobody has ever listed by hand.

up to 1,000 sitemaps per job
  • File upload or URL. Either hand over the file or point the job at the address where your system generates it.
  • Recursive parsing three levels deep. Index files pointing to index files pointing to sitemaps are followed, which is exactly how large catalogue sites are usually structured.
  • Queued, not simultaneous. Two jobs run at once with up to twenty waiting, so a full catalogue re-submission is planned work rather than a button pressed in a panic.
3
levels parsed recursively
2 / 20
concurrent jobs, queue depth
1,000
sitemaps per job

Submission itself runs over the IndexNow API, which notifies the participating bots directly instead of waiting for them to come round. That shortens the notification step. It does not shorten, and cannot shorten, the evaluation step that follows.

Limits · The part people misread

Submitting a URL is not the same as being indexed

Say this out loud before you plan anything. Submitting a URL is not the same as being indexed. Submission tells a search engine that an address exists and is worth a look. Whether the page is then added to the index, kept in it, and shown for any query at all remains entirely the search engine's decision. No tool from any vendor changes that, and anyone who tells you otherwise is selling something they cannot deliver.

The distinction matters because the two outcomes look identical in a submission log. A URL can be submitted successfully, visited by a bot within hours, recorded with a clean status, and still never appear in a single result page. That is not a failure of the submission. It is the search engine deciding the page was not worth including — usually because it duplicates something else, because it is too thin to be useful, or because the domain has not earned the room.

  • What submission genuinely fixes. Discovery latency. A page nothing links to and no sitemap lists can be brought to a bot's attention in hours rather than months.
  • What it does not fix. Thin pages, duplicated permutations, and pages whose content is a translation of a page already in the index. The evaluation step catches all three.
  • What the log proves. That a bot arrived and what it saw at the door. Nothing beyond that, and reading more into it produces bad decisions.
  • How to confirm inclusion. A page that is genuinely indexed starts producing impressions in Search Console. That is the confirmation; the submission counter is not.

This is also why bulk submission of a whole catalogue is usually a poor first move. If four hundred of the five hundred pages are permutations that will be judged duplicative, you have spent your daily budget teaching a search engine that your domain produces low-value pages. The order in which you submit is a signal in itself.

Budget · Spending the ration deliberately

Where a supplier's crawl budget actually goes

Crawl budget is not a number you can look up, but its symptoms are readable. Pages taking weeks to appear, older pages losing freshness, a sitemap listing far more addresses than ever produce impressions: these all point the same way.

SymptomLikely causeFirst action
New pages take four weeks or more to appearAttention consumed by low-value addressesIdentify and close off filter combinations
Sitemap lists far more URLs than earn impressionsPermutation pages judged duplicativeConsolidate the matrix into fewer, fuller pages
English pages lag behind Czech ones consistentlyFewer internal links into the English treeCross-link sections and submit the English tree first
A recently published page has no impressions at allPossibly never indexed rather than badly rankedCheck the submission log, then check impressions
Downloadable specifications never appear in resultsDocuments reachable only through scriptGive each an HTML page that links to it plainly

The remedy is ordinary and unglamorous: reduce the number of addresses that exist, make the ones that remain distinct enough to deserve inclusion, and make sure everything commercially important is reachable by clicking. Submission then works on a site that has something to gain from it. The on-site recommendations in the panel flag structural problems of this kind, and where a change needs a developer rather than an editor it belongs with technical SEO work rather than with content.

An order that works. Commercially important pages first, in small batches, on days when nothing else is being submitted. Catalogue-wide sitemap jobs afterwards, once the structural cleanup is done. The daily budget is a ration, and rations reward planning.
Routine · Making it a habit

A monthly rhythm rather than an emergency

Most technical sites treat indexing as something you investigate when a page fails. It works far better as a short recurring routine attached to whatever else happens monthly.

Every month

Submit what changed

New capability pages, updated certificates, new case studies. A dozen addresses, not a thousand, submitted deliberately and checked a week later against impressions.

  • Small batches read as a maintained site
  • Easy to attribute an outcome
Every quarter

Re-run the sitemap job

Regenerate the sitemap, submit it, and compare the URL count against the number of pages actually earning impressions. A widening gap is the signal to consolidate.

  • Catches sections nobody maintains
  • Catches orphaned language trees

Everything the hub records can leave the panel: the export formats available run to CSV or JSON at up to 10,000 rows and to PDF at up to 250 rows through a configurable report builder. For an indexing log the CSV route is the useful one, because comparing this month's submitted list against the pages producing impressions is a spreadsheet operation, not a dashboard one. Where several domains are involved, site tags act as a global filter, and individual sites can be released to another email address so an external developer sees one property and nothing more.

The project assistant, the Stream feed, takes URL lists in batches and keeps to-dos in the same chronological line as automatic reports and campaign news, with states for active, deferred and discarded. Used properly it turns "check whether the certificate page ever got indexed" from something remembered at the wrong moment into an item with a date on it. More on how the surrounding analytics are read sits in the English articles section.

Questions · Straight answers

Common questions

How long should I wait before deciding a page was not indexed?

For a site that is crawled regularly, two to three weeks after submission is a reasonable point to conclude that a page has been seen and not included. Confirm it through impressions rather than through the submission log, since the log only records that a bot arrived.

Should I submit the Czech and English versions of a page together?

Submitting both is fine and normal. The thing to watch is that they are genuinely different pages rather than one page rendered twice, because two language versions of thin content are still thin content. If the English version was written from a specification rather than translated word for word, it will stand on its own.

Is the daily limit of a thousand URLs a problem for a large catalogue?

Rarely, once you stop trying to submit everything. A catalogue of ten thousand addresses almost never contains ten thousand pages worth indexing, and working out which few hundred do is more valuable than the submission itself. For genuine bulk work, the sitemap route handles scale better than the URL tracker.

Our sitemap is generated automatically. Do we still need to submit it?

Automatic generation solves accuracy, not attention. The file may still be re-read infrequently, and an explicit submission after a significant change is the way to shorten that. Recursive parsing three levels deep also means an index file pointing at section files is handled without you flattening it first.

We publish a page for every trade fair. Is that worth indexing?

Only if the page has a life beyond the event, which most do not. A better pattern is one durable page per recurring event or exhibition series, updated each year, which accumulates history instead of starting from zero. A page with a fixed expiry date is rarely worth the crawl attention it consumes.

Close · What to check today

The half hour that answers the question

Pick the three pages you would most want a foreign buyer to find: the newest certificate, the capability you invested in most recently, and the case study you are proudest of. For each, check whether it is producing impressions at all. Not clicks, not position — impressions, which is the evidence that the page exists in the index and has been shown to somebody.

If a page has none, you have separated two problems that usually get confused. A page with impressions and no clicks is a wording problem. A page with no impressions may never have been included, and that is where submission belongs. Fixing the wording of a page that is not in the index is effort spent on nothing.

The limit worth repeating. Everything described here shortens the wait for a bot to arrive. None of it obliges a search engine to keep the page, and no submission counter, however green, is evidence that it did. The only proof of indexing is a page appearing in results, and the only proof of value is what happens after it does.

If you want to see the actual state of your own catalogue rather than assume it, the fastest route is to submit your sitemap and read what comes back: open the dashboard and start a sitemap job, then compare the parsed URL count against the number of pages earning impressions this quarter. On most technical sites the two numbers are far apart, and the distance between them is the work.