A Czech supplier passes an audit in March and publishes the certificate page the same week. A buyer in Denmark runs his shortlist in April and never sees it, because the page was still not in the index. The gap between publishing and being findable is a commercial delay, and it is measurable.
Indexing gets talked about as a technical housekeeping matter, somewhere below page speed in the list of things to worry about. For a company whose enquiries arrive from abroad, it is closer to a logistics problem. A capability you cannot be found for is a capability you do not have, from the point of view of a purchasing department that will spend one afternoon assembling a list of candidates and never revisit it.
The certificate that nobody could find
Sourcing does not run continuously. It runs in bursts, tied to a project, a re-tender, a plant qualification or a failure at an existing supplier. When a burst happens, a small number of people spend a limited amount of time building a candidate list, and everything they find during those hours is the entire universe of options considered. Anything indexed the following week does not exist.
This is why indexing latency costs a supplier real money in a way it does not cost a shop or a news site. A retailer that gets a product page indexed late loses a fraction of a season. A testing laboratory that gets its new accreditation page indexed six weeks late may have missed the only qualification round that customer will run for three years.
- New accreditations and standards. An audit passed is an asset only once the page describing it can be found by somebody searching for that standard and nothing else.
- New equipment and new limits. A larger machine, a wider tolerance envelope, a cleanroom class. These are the phrases that change which briefs you qualify for.
- Case studies with named sectors. Evidence pages are what turn a listing into a shortlist entry, and they are usually published last and linked least.
- Event and trade-fair pages. A stand at a show in eight weeks needs its page indexed in two, not in ten. The window closes on a fixed date.
Why a capability catalogue is hard to crawl
Websites belonging to manufacturers, laboratories and technical service firms tend to share a shape that search engines handle badly. The shape is not a blog with a hundred articles. It is a matrix.
Take a machining supplier. Processes multiplied by materials multiplied by industries produces a page count in the hundreds before anyone has written a word of news. Add a downloadable datasheet for each, a filterable parts index, and a full Czech translation of everything, and a site that a human perceives as fifteen sections is several thousand addressable URLs to a crawler.
Crawling is rationed. A search engine allocates a certain amount of attention to a domain based on its authority, its update frequency and how much waste it encountered last time. A site with thousands of near-identical permutation pages spends that ration on pages nobody needed, and the certificate page published on Tuesday sits in a queue behind them.
Filter and sort parameters
A parts index with four filters generates combinations indefinitely. Each combination is a distinct address, and each one is crawled at least once before being judged worthless.
- Often the single largest drain
- Rarely visible to anyone who has not looked
Permutation pages with one line changed
Twelve pages differing only in a material name read as one page repeated twelve times. They compete with each other and dilute the attention available to genuinely distinct pages.
- Consolidate or differentiate properly
- Half-measures make the problem worse
Datasheets living only in PDFs
The technical detail a buyer searches for is often inside a document that nothing on the site links to except a download button loaded by script.
- Reachable in principle, undiscovered in practice
- Worth a real HTML page alongside
The English half of a bilingual site
The Czech tree usually carries more internal links, more history and more inbound references, so the English tree is reached later and less often.
- The half that earns export enquiries
- Also the half that gets crawled last
That last point deserves emphasis because it is specific to bilingual sites and rarely mentioned. Nothing forces a crawler to treat the two language trees equally. If the English section was added later, is linked from fewer places and is updated less often, it will be visited less often, and its new pages will wait longer. The commercial importance of a directory has no bearing on how quickly it gets crawled.
Three ways a page gets discovered, and one that fails
Before any question of ranking, a page has to be discovered. There are only a few routes, and understanding which one you are relying on tells you how long to expect to wait.
| Route | How it works | Typical delay | Where it fails on a supplier site |
|---|---|---|---|
| Internal links | A crawler already visiting the site follows a link to the new page | Days to weeks, depending on how often the site is crawled | New pages added to a section nothing links to |
| Sitemap files | The page is listed in an XML file the search engine reads periodically | Unpredictable; the file may not be re-read for some time | A stale sitemap generated once and never regenerated |
| External links | Another site references the page directly | Fast when it happens | Almost never happens for a capability page |
| Direct submission | The URL is pushed to the search engine through an API | The bot visit is usually prompt | A visit is not an inclusion; see the warning below |
The failure that catches most technical sites is the orphan page. Somebody adds a capability page through the content system, links to it from a single campaign email, and never adds it to the navigation. It exists, it is well written, it names a standard that buyers search for, and no crawler has any path to it. Months later it is discovered by accident, or not at all.
The Indexing Hub, in plain terms
The indexing section of the unified Semalt panel exists to shorten the discovery step rather than to replace it. It has two working parts, and they solve different problems.
URL tracker and bulk submission
For a known list of addresses you want a bot to visit soon.
- A daily budget, not an unlimited queue. One thousand URLs a day per account, which is a planning constraint rather than a limitation to work around.
- Batches of up to ten thousand. A bulk submission accepts up to 10,000 URLs at once, drawn down against the daily budget rather than processed instantly.
- A log per address. Bot visit with a timestamp, status, and error detail where there is one, alongside live counters for submitted, found and failed URLs.
Sitemap submission and parsing
For a catalogue whose full extent nobody has ever listed by hand.
- File upload or URL. Either hand over the file or point the job at the address where your system generates it.
- Recursive parsing three levels deep. Index files pointing to index files pointing to sitemaps are followed, which is exactly how large catalogue sites are usually structured.
- Queued, not simultaneous. Two jobs run at once with up to twenty waiting, so a full catalogue re-submission is planned work rather than a button pressed in a panic.
Submission itself runs over the IndexNow API, which notifies the participating bots directly instead of waiting for them to come round. That shortens the notification step. It does not shorten, and cannot shorten, the evaluation step that follows.
Submitting a URL is not the same as being indexed
The distinction matters because the two outcomes look identical in a submission log. A URL can be submitted successfully, visited by a bot within hours, recorded with a clean status, and still never appear in a single result page. That is not a failure of the submission. It is the search engine deciding the page was not worth including — usually because it duplicates something else, because it is too thin to be useful, or because the domain has not earned the room.
- What submission genuinely fixes. Discovery latency. A page nothing links to and no sitemap lists can be brought to a bot's attention in hours rather than months.
- What it does not fix. Thin pages, duplicated permutations, and pages whose content is a translation of a page already in the index. The evaluation step catches all three.
- What the log proves. That a bot arrived and what it saw at the door. Nothing beyond that, and reading more into it produces bad decisions.
- How to confirm inclusion. A page that is genuinely indexed starts producing impressions in Search Console. That is the confirmation; the submission counter is not.
This is also why bulk submission of a whole catalogue is usually a poor first move. If four hundred of the five hundred pages are permutations that will be judged duplicative, you have spent your daily budget teaching a search engine that your domain produces low-value pages. The order in which you submit is a signal in itself.
Where a supplier's crawl budget actually goes
Crawl budget is not a number you can look up, but its symptoms are readable. Pages taking weeks to appear, older pages losing freshness, a sitemap listing far more addresses than ever produce impressions: these all point the same way.
| Symptom | Likely cause | First action |
|---|---|---|
| New pages take four weeks or more to appear | Attention consumed by low-value addresses | Identify and close off filter combinations |
| Sitemap lists far more URLs than earn impressions | Permutation pages judged duplicative | Consolidate the matrix into fewer, fuller pages |
| English pages lag behind Czech ones consistently | Fewer internal links into the English tree | Cross-link sections and submit the English tree first |
| A recently published page has no impressions at all | Possibly never indexed rather than badly ranked | Check the submission log, then check impressions |
| Downloadable specifications never appear in results | Documents reachable only through script | Give each an HTML page that links to it plainly |
The remedy is ordinary and unglamorous: reduce the number of addresses that exist, make the ones that remain distinct enough to deserve inclusion, and make sure everything commercially important is reachable by clicking. Submission then works on a site that has something to gain from it. The on-site recommendations in the panel flag structural problems of this kind, and where a change needs a developer rather than an editor it belongs with technical SEO work rather than with content.
A monthly rhythm rather than an emergency
Most technical sites treat indexing as something you investigate when a page fails. It works far better as a short recurring routine attached to whatever else happens monthly.
Submit what changed
New capability pages, updated certificates, new case studies. A dozen addresses, not a thousand, submitted deliberately and checked a week later against impressions.
- Small batches read as a maintained site
- Easy to attribute an outcome
Re-run the sitemap job
Regenerate the sitemap, submit it, and compare the URL count against the number of pages actually earning impressions. A widening gap is the signal to consolidate.
- Catches sections nobody maintains
- Catches orphaned language trees
Everything the hub records can leave the panel: the export formats available run to CSV or JSON at up to 10,000 rows and to PDF at up to 250 rows through a configurable report builder. For an indexing log the CSV route is the useful one, because comparing this month's submitted list against the pages producing impressions is a spreadsheet operation, not a dashboard one. Where several domains are involved, site tags act as a global filter, and individual sites can be released to another email address so an external developer sees one property and nothing more.
The project assistant, the Stream feed, takes URL lists in batches and keeps to-dos in the same chronological line as automatic reports and campaign news, with states for active, deferred and discarded. Used properly it turns "check whether the certificate page ever got indexed" from something remembered at the wrong moment into an item with a date on it. More on how the surrounding analytics are read sits in the English articles section.
Common questions
How long should I wait before deciding a page was not indexed?
For a site that is crawled regularly, two to three weeks after submission is a reasonable point to conclude that a page has been seen and not included. Confirm it through impressions rather than through the submission log, since the log only records that a bot arrived.
Should I submit the Czech and English versions of a page together?
Submitting both is fine and normal. The thing to watch is that they are genuinely different pages rather than one page rendered twice, because two language versions of thin content are still thin content. If the English version was written from a specification rather than translated word for word, it will stand on its own.
Is the daily limit of a thousand URLs a problem for a large catalogue?
Rarely, once you stop trying to submit everything. A catalogue of ten thousand addresses almost never contains ten thousand pages worth indexing, and working out which few hundred do is more valuable than the submission itself. For genuine bulk work, the sitemap route handles scale better than the URL tracker.
Our sitemap is generated automatically. Do we still need to submit it?
Automatic generation solves accuracy, not attention. The file may still be re-read infrequently, and an explicit submission after a significant change is the way to shorten that. Recursive parsing three levels deep also means an index file pointing at section files is handled without you flattening it first.
We publish a page for every trade fair. Is that worth indexing?
Only if the page has a life beyond the event, which most do not. A better pattern is one durable page per recurring event or exhibition series, updated each year, which accumulates history instead of starting from zero. A page with a fixed expiry date is rarely worth the crawl attention it consumes.
The half hour that answers the question
Pick the three pages you would most want a foreign buyer to find: the newest certificate, the capability you invested in most recently, and the case study you are proudest of. For each, check whether it is producing impressions at all. Not clicks, not position — impressions, which is the evidence that the page exists in the index and has been shown to somebody.
If a page has none, you have separated two problems that usually get confused. A page with impressions and no clicks is a wording problem. A page with no impressions may never have been included, and that is where submission belongs. Fixing the wording of a page that is not in the index is effort spent on nothing.
If you want to see the actual state of your own catalogue rather than assume it, the fastest route is to submit your sitemap and read what comes back: open the dashboard and start a sitemap job, then compare the parsed URL count against the number of pages earning impressions this quarter. On most technical sites the two numbers are far apart, and the distance between them is the work.