> For the complete documentation index, see [llms.txt](https://help.rankability.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://help.rankability.com/audit/site-auditor-guide.md).

# Using Site Auditor

Use Site Auditor to crawl a website, build a page inventory, find technical SEO issues, compare internal duplicate content, and verify whether issues changed after a later crawl.

Site Auditor is best for site-wide diagnosis. To evaluate one page against a target keyword and search intent, use [Page Content Auditor](/audit/page-content-auditor-guide.md).

## Before you start

Open the client and choose **Audit → Site Auditor**. Confirm that the client domain is correct, then choose the crawl source:

* **Website** — Follow internal links starting from the website.
* **Sitemap** — Start from a specific XML sitemap URL.
* **URL list** — Upload or paste URLs from a TXT or CSV file.

The available page limits are 100, 250, 500, and 1,000 pages. A crawl that reaches its limit is truncated; it does not prove that the uncrawled portion of the site has no issues.

Turn on **Check external links** when broken outbound links matter to the audit. This adds work and can make the crawl slower.

Use the advanced controls only when the default crawl is too broad or too narrow. You can include subdomains, restrict the crawl to allowed paths, exclude disallowed paths, or ignore URL parameters. Check these settings carefully: an overly restrictive rule can omit important pages from the report.

## Credits and completion

Site crawling costs **1 credit per 5 successfully crawled pages**, rounded up. The initial estimate uses the selected page limit, but the completed run is charged from the actual number of pages crawled. A run that errors or is cancelled before successful completion is not charged.

The crawl runs in the background, so you can leave the page while it works. Rankability shows pending, crawling, analyzing, completed, error, or cancelled status and sends an in-app notification when the report is ready.

## Review the Pages tab

The **Pages** tab is the crawl inventory. Search, filter, and sort it to isolate a template, status, indexability state, action, redirect, thin page, or deep page. Available columns can include:

* HTTP status, redirect target and redirect chain information;
* indexability, robots directives, and canonical URL;
* title, meta description, H1 and H2 headings;
* word count, crawl depth, linking pages, schema, and page weight;
* AI crawler access derived from robots.txt;
* Google Search Console, GA4, Bing, backlink, and performance data when available; and
* target keyword, its source, confidence, reason, and manual notes.

The optional detail and enrichment columns start hidden so the default inventory stays scannable. Use **Manage columns** to show or hide them by group: **On-page details**, **Technical**, **Search Console**, **Analytics (GA4)**, and **Core Web Vitals**. A column you expected is more often a hidden group than missing data, so check here before concluding an enrichment source did not return anything. Your choice is remembered in the browser you made it in, and the header stays in place while you scroll a long inventory.

Use **Identify target keyword** for indexable pages that do not already have one. Rankability tries connected Search Console data first at no charge. If that cannot identify a keyword, the AI fallback costs **25 credits per page only when it succeeds**. The confirmation shows a worst-case estimate; the actual charge can be lower because GSC-resolved pages are free.

You can edit a target keyword manually. Use **CSV** to export the current page inventory for triage or assignment.

## Review Duplicate content

The **Duplicate content** tab reports the share of analyzable body text duplicated internally, the affected pages, matched words, and pages that share the text.

Open a row to see duplicated passages highlighted in context. Select a shared page to compare the two pages side by side before deciding whether to rewrite, consolidate, redirect, or canonicalize.

Use **Hide** only after confirming that a result is an intentional or harmless match. Hidden duplicate-content findings can be restored later and remain associated with the URL across recrawls. Hiding a finding changes its review state in Rankability; it does not change the website.

Older audits may not contain duplicate-content analysis. Rerun the audit to generate the newer report.

## Review Findings and verify fixes

The **Findings** tab groups issues by Critical, Warning, Info, and Opportunity. Filter by severity or lifecycle state such as New, Fixed, Unchanged, Worsened, or Improved. Open a finding to inspect the affected URL, evidence, recommendation, and comparison with the prior crawl when one exists.

After changing the site, use **Verify now** on a supported deterministic page-level finding. This live check does not use Rankability credits and records whether that specific issue is fixed or still open. Heuristic and site-wide findings may not support a live check.

Rerun the same audit scope when you need a complete new crawl, lifecycle comparison, or verification of unsupported findings. A finding marked **Fixed** in the lifecycle view means the later crawl no longer detected it under the audited conditions; it is not a guarantee of rankings or indexation.

## Connected-data limitations

The crawl can complete even when GSC, GA4, Bing, backlink, PageSpeed, or other enrichment data is disconnected or temporarily unavailable. In that case the related columns may be blank. A blank value caused by unavailable data is not zero and is not a confirmed pass.

If enrichment fails during a run, review the visible notice, confirm the connection, and rerun to load the missing data. See [Connecting client Google services](/account-and-settings/connecting-client-google-services.md) and [Connecting Bing Webmaster Tools](/track/connecting-bing-webmaster-tools.md).

## Troubleshooting

If expected pages are missing, check the selected page limit, crawl source, allow and disallow paths, subdomain scope, robots rules, authentication, firewall, and JavaScript rendering. For blocked requests, see [Whitelisting the Rankability crawler](/troubleshooting/whitelisting-rankability-crawler.md).

If a supported crawl repeatedly errors, contact support with the client, project name, domain, time of the run, selected crawl settings, and visible error text.

## Related articles

* [Audit overview](/audit/audit-overview.md)
* [Using Page Content Auditor](/audit/page-content-auditor-guide.md)
* [Credit costs reference](/account-and-settings/credit-costs-reference.md)
* [Troubleshooting common issues](/troubleshooting/troubleshooting-common-issues.md)
