Skip to content
traze

New · free tool · no login

Site Scan

Enter your domain. We crawl twenty pages, judge the type, intent and cluster of each one, look for pages getting in each other's way, and check whether AI crawlers can use your site at all.

Try:

No login, no credit card. We only crawl public pages and we respect robots.txt. Nothing from your site is stored.Summary free, full report after your email.

What the check looks at · 6 components

What the Site Scan does in thirty seconds

The scan crawls the way a search engine does: through your sitemap and your internal links, respecting robots.txt. What can be counted, we count. What needs judgement, our decision model judges — and we say which is which, per item.

  1. 01

    Twenty pages, really fetched

    We start at your homepage, read your sitemap and follow internal links up to twenty pages. Every page is really fetched; you see the status, the word count and the title we saw.

  2. 02

    Page type and intent per page

    Is this a product page, an article, a service or an FAQ? And is the visitor after an explanation, a comparison or a purchase? The decision model picks from a fixed list per page, with its certainty attached.

  3. 03

    Clusters from your own titles

    The scan derives candidate subjects from the words your own titles repeat and files every page under the subject it belongs to. That shows where your site leans heavy and where it is thin.

  4. 04

    Pages getting in each other's way

    Two pages in the same cluster, with the same intent and overlapping titles, get an explicit question: do they serve the same search? Only when the answer is yes do they end up in your report.

  5. 05

    AI readability

    Is there an llms.txt? Does robots.txt block AI crawlers like GPTBot, ClaudeBot or PerplexityBot? Do your pages have titles, descriptions and enough text to answer a question?

  6. 06

    Three actions, not a scoreboard

    The scan ends in a short list of what you would do now, in order. Including what it takes to keep watching instead of measuring once.

Why it matters · the short version

Why twenty pages are enough to see something

The problems holding your visibility back are rarely on page 21.

A site with a problem shows that problem in its first twenty pages almost every time. Thin pages come in batches. Two pages about the same subject exist because there is no structure, not because someone happened to write the same thing twice. And a robots.txt that shuts AI crawlers out applies to your whole domain.

So this tool scans broad instead of deep: twenty pages, each judged on what it is and who it is for. That gives you a map of your site as a search engine or an AI assistant sees it — including the places where that map stays blank.

What the scan does not do, it says out loud. Twenty pages are not a full crawl, one moment is not a trend, and a model's judgement is a judgement, not a fact. That is why every judgement carries a certainty and every number says whether it was measured or estimated.

Methodology · what is what

What we decide, measure and estimate

Every number in the result belongs to one of these three. The report itself says which, per item, so you know what to lean on.

Assessed by our decision model

A judgement with a probability behind it: certain, likely or uncertain. Not a fact, but traceable — the same input gives the same judgement, and the certainty is always shown.

Per page the type, the search intent and the cluster; per suspicious pair whether they really fight over the same search; and whether your site is usable as a source.

Measured

What we literally fetched or counted: HTTP status, HTML, robots.txt, sitemaps, feeds and search results. If it is not there, we say so.

The twenty fetched pages with status, title, description and word count, plus robots.txt, llms.txt and the title overlap between pages.

Estimated

Models and rules of thumb: search volumes, click rates, expected return. Good enough to prioritise with, not to budget on.

The AI readiness score: a weighting of the measured checks and that one judgement.

Simulated. When no decision model, provider or cluster is available — or the daily ceiling for the free tools is reached — the tool keeps running on deterministic example data. That is labelled on every item it touches.

Good to know

Questions about Site Scan

Because you want an answer inside thirty seconds and we do not want to hammer your server. We crawl gently, respect robots.txt and stop at twenty pages. Want everything: a free account crawls your whole site and keeps watching.

From check to system

One scan is a photo. Traze is the film.

Traze crawls your whole site, tracks who gets named in AI answers, writes the pages you are missing and publishes them into your CMS — with you in the loop. Start with a free account based on this scan.

One check is a snapshot

Traze is the loop.

Traze researches, writes, publishes and optimises the content that gets you found in Google and in AI answers — automatically, with you in the loop. See pricing or book a demo.