> For the complete documentation index, see [llms.txt](https://help.rankability.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://help.rankability.com/audit/site-auditor-guide.md).

# Using Site Auditor

Run a Rankability site audit, control crawl scope, interpret technical SEO findings and duplicate content, and verify fixes after a recrawl.

Use Site Auditor to crawl a website, build a page inventory, find technical SEO issues, compare internal duplicate content, and verify whether issues changed after a later crawl.

Site Auditor is best for site-wide diagnosis. To evaluate one page against a target keyword and search intent, use [Page Content Auditor](/audit/page-content-auditor-guide.md).

## Before you start

Open the client and choose **Audit → Site Auditor**. Confirm that the client domain is correct, then choose the crawl source:

* **Website** — Follow internal links starting from the website.
* **Sitemap** — Start from a specific XML sitemap URL.
* **URL list** — Upload or paste URLs from a TXT or CSV file.

The available page limits are 100, 250, 500, and 1,000 pages. A crawl that reaches its limit is truncated; it does not prove that the uncrawled portion of the site has no issues.

Turn on **Check external links** when broken outbound links matter to the audit. This adds work and can make the crawl slower.

Use the advanced controls only when the default crawl is too broad or too narrow. You can include subdomains, restrict the crawl to allowed paths, exclude disallowed paths, or ignore URL parameters. Check these settings carefully: an overly restrictive rule can omit important pages from the report.

## Usage and completion

Site crawling is included in full-platform pooled usage. The estimate uses the selected page limit, while website activity reflects successfully crawled pages. Large crawls can have a high on-demand usage impact.

The crawl runs in the background, so you can leave the page while it works. Rankability shows pending, crawling, analyzing, completed, error, or cancelled status and sends an in-app notification when the report is ready.

When the report's findings cannot be loaded, site health and top findings say they are unavailable and offer **Retry findings**. Read that as a loading failure, not a clean audit — a crawl that genuinely found nothing says so instead.

## Review the Pages tab

The **Pages** tab is the crawl inventory. Search, filter, and sort it to isolate a template, status, indexability state, action, redirect, thin page, or deep page. Available columns can include:

* HTTP status, redirect target and redirect chain information;
* indexability, robots directives, and canonical URL;
* title, meta description, H1 and H2 headings;
* word count, crawl depth, linking pages, schema, and page weight;
* AI crawler access derived from robots.txt;
* Google Search Console, GA4, Bing, backlink, and performance data when available; and
* target keyword, its source, confidence, reason, and manual notes.

The inventory opens with a deliberately small set of columns — status code, indexability, title, linking pages, crawl depth, AI bot access, and pageviews — so it stays scannable. Use **Manage columns** to switch any other column on. They are grouped by section, including Core audit, On-page, Technical, Search Console, Analytics, and Performance, and **Reset to default** returns you to the opening set. A column you expected is more often switched off than missing data, so check here before concluding an enrichment source returned nothing. Your choice is remembered in the browser you made it in, and the header stays in place while you scroll a long inventory.

Selecting a page's URL opens its **Crawled page details** panel instead of leaving the report. Use the separate action in that panel when you want to open the live page in a new tab.

Use **Identify target keyword** for indexable pages that do not already have one. Rankability tries connected Search Console data first, then can use an AI fallback. Review the page count and usage impact before a large batch.

You can edit a target keyword manually.

## Review Statistics

The **Statistics** tab summarizes the crawl as seven measures: HTTP status codes, indexability, page crawl depth, incoming internal links, canonicalization, schema markup, and AI crawler access. Each card shows a headline percentage, the distribution behind it, and a plain assessment — **Looks healthy**, **Review**, **Needs attention**, or **Context matters** — with a short reason. Open the information control on a card to read what the measure counts, why it matters, and what to do next.

When a card has affected pages, **View affected pages** opens the Pages tab narrowed to exactly that set, such as pages returning 4xx or 5xx, pages more than three clicks deep, or pages with one or fewer incoming links. The narrowing appears as a labeled chip above the inventory; remove the chip to return to the full list.

Treat these measures as diagnostic starting points rather than scores to maximize. Not every page needs schema markup, a restricted AI crawler may be a deliberate policy, and an excluded page is often excluded on purpose. The crawl depth, internal link, and AI access percentages are calculated only over the pages where that information could be determined.

## Review Duplicate content

The **Duplicate content** tab separates actionable findings from raw measured overlap. The comparison reads each page's main content, so shared navigation, headers, footers, and sidebars stay out of it; when a page has no clear main region, Rankability compares the rest of the page instead. A page becomes an actionable finding only when Rankability has concentrated primary-content evidence against one specific other page. Overlap spread thinly across many pages, or contributed by listing pages, does not qualify. Tiny matches and expected repetition on listing, archive, pagination, and HTML sitemap pages are not labeled as defects, and neither is a shared section template — a repeated call to action or standard block — even when it appears on only part of the site rather than every page.

The summary shows the actionable share of analyzed text and separately reports the raw overlap measured across the crawl. Open **Other measured overlap** when you need to inspect lower-confidence or intentional matches without treating them as recommendations to rewrite or consolidate a page. A row that is not an actionable finding also offers **View raw measurement** for the overlap Rankability retained for reference.

Open a row to see duplicated passages highlighted in context. Select a shared page to compare the two pages side by side before deciding whether to rewrite, consolidate, redirect, or canonicalize. The side-by-side view highlights only text both saved pages actually support, so an older audit with incomplete evidence highlights less rather than over-marking the comparison.

Use **Hide** only after confirming that a result is an intentional or harmless match. Hidden duplicate-content findings can be restored later and remain associated with the URL across recrawls. Hiding a finding changes its review state in Rankability; it does not change the website.

Older audits may not contain duplicate-content analysis. Rerun the audit to generate the newer report.

## Review Findings and verify fixes

The **Findings** tab groups issues by Critical, Warning, Info, and Opportunity. Filter by severity or lifecycle state such as New, Fixed, Unchanged, Worsened, or Improved. A finding previews up to three affected pages; choose **View all affected pages** to inspect the complete paginated list. Each page link opens in a new browser tab. The card also includes the evidence, recommendation, and comparison with the prior crawl when one exists.

Every finding carries **How to fix**, which opens a recommended next step for that issue. The guidance does not assume a particular content management system, so apply it where the affected pages are published and then recrawl to confirm the change. **Ask Serena** in the same panel opens a Serena conversation about that finding, grounded in the saved audit rather than in anything the link carries.

After changing the site, use **Verify now** on a supported deterministic page-level finding. This records whether that specific issue is fixed or still open. Heuristic and site-wide findings may not support a live check.

Rerun the same audit scope when you need a complete new crawl, lifecycle comparison, or verification of unsupported findings. A finding marked **Fixed** in the lifecycle view means the later crawl no longer detected it under the audited conditions; it is not a guarantee of rankings or indexation.

## Export the report

Once a crawl is completed, **Export** on the report header covers the whole audit rather than the tab you are viewing.

**Executive PDF report** is the summary to hand a client or a stakeholder. It covers the crawl scope, what to address first, technical readiness, what changed since the prior crawl, and tracked crawler access. It needs the report's findings, so the item is unavailable while they cannot be loaded.

For the underlying data, choose **Full audit workbook** for one Excel file with crawled pages, findings, and duplicate content on separate sheets, or export **Crawled pages**, **Findings**, or **Duplicate content** on their own as an Excel workbook or a CSV spreadsheet. Duplicate content also offers **Markdown for AI** when you want to hand the report to an AI assistant.

An export contains every reportable row for that report, not the subset left by the search, filters, and sorting you applied on screen. Findings are ordered by priority and include resolved ones. Duplicate-content exports retain actionable and lower-confidence measured overlap with an explicit evidence state and reason, and record which part of each page the comparison read; zero-overlap pages are omitted. Rows you hid remain exported with their hidden state and reason recorded. The file reads the saved crawl and does not recrawl the site, so check the crawl date before putting it in a client deliverable.

## Connected-data limitations

The crawl can complete even when GSC, GA4, Bing, backlink, PageSpeed, or other enrichment data is disconnected or temporarily unavailable. In that case the related columns may be blank. A blank value caused by unavailable data is not zero and is not a confirmed pass.

If enrichment fails during a run, review the visible notice, confirm the connection, and rerun to load the missing data. See [Connecting client Google services](/account-and-settings/connecting-client-google-services.md) and [Connecting Bing Webmaster Tools](/track/connecting-bing-webmaster-tools.md).

## Troubleshooting

If expected pages are missing, check the selected page limit, crawl source, allow and disallow paths, subdomain scope, robots rules, authentication, firewall, and JavaScript rendering. For blocked requests, see [Whitelisting the Rankability crawler](/troubleshooting/whitelisting-rankability-crawler.md).

If a supported crawl repeatedly errors, contact support with the client, project name, domain, time of the run, selected crawl settings, and visible error text.

## Related articles

* [Audit overview](/audit/audit-overview.md)
* [Using Page Content Auditor](/audit/page-content-auditor-guide.md)
* [Usage limits reference](/account-and-settings/credit-costs-reference.md)
* [Troubleshooting common issues](/troubleshooting/troubleshooting-common-issues.md)
