Most agency technical SEO audits fail for the same reason: they end as a 90-page PDF that nobody implements. The crawl was fine. The findings were real. But the client's developer got a document instead of tickets, and six months later you're re-auditing the same broken canonical tags.
A good technical SEO audit is a project, not a report. Here's the process we see work across agencies running 8–30 client sites: scope it, crawl it, find the gap between what exists and what Google indexes, prioritize ruthlessly, and ship it as work a developer can actually do.
Before you crawl: scope and access (30 minutes)
Half of all audit overruns happen because the team started crawling before they knew what they were auditing. Lock these down first.
Access checklist
- Google Search Console — all property variants (http, https, www, non-www, subdomains)
- Analytics with at least 13 months of history
- CMS admin or staging access
- Server log files (30 days minimum — ask early, this takes IT a week)
- Any existing redirect maps, robots.txt change history, or past migration docs
Define the boundary
Write one sentence: "We're auditing example.com and shop.example.com, excluding the legacy blog on WordPress, for crawlability, indexability, performance, and structured data." That sentence kills three hours of scope creep later. If the client has a subdomain nobody mentioned with 40,000 thin pages, you want to find it now and quote it separately.
Then timebox. A 5,000-URL site is a 12–16 hour audit. A 500,000-URL ecommerce site is 40–60 hours. If you're not tracking that time against the scope, you won't know your audits are losing money until Q3. Agencies that bill retainers tend to lose the most here — audits feel like "included work" until they've consumed 70% of the month's hours. Watching that in advance is the whole point of proper agency capacity planning.
Step 1: Baseline how Google sees the site
Before your own crawler touches anything, get Google's view. Open Search Console and record:
- Pages report — indexed count vs. not-indexed count, and the top five exclusion reasons
- Sitemaps — submitted vs. discovered vs. indexed, per sitemap file
- Core Web Vitals — URL groups failing on mobile
- Crawl stats (Settings → Crawl stats) — average response time, crawl requests per day, response code breakdown
Write these numbers down. A site with 4,200 submitted URLs and 1,100 indexed has a structural problem worth more than every meta description fix you'll find later. Crawl stats showing an average response time above 600ms tells you performance is suppressing crawl budget before you've run a single test.
Step 2: Run the crawl with the right settings
Default crawler settings produce misleading audits. Configure deliberately:
- User agent: Googlebot Smartphone. Mobile-first indexing is the reality; desktop crawls hide mobile-only issues.
- JavaScript rendering: on, if the site is React, Vue, Next.js, or Shopify with heavy apps. Then run a second crawl with rendering off and diff them. The delta is your JS dependency risk.
- Respect robots.txt: on for the first pass, off for the second. The second crawl shows you what's being blocked that shouldn't be.
- Crawl limit: uncapped if under 100k URLs. Otherwise sample intelligently by directory.
- Connect APIs: GSC, GA4, and PageSpeed. This is the step most people skip, and it's the one that turns a crawl into an audit — you can now sort technical issues by sessions and clicks instead of guessing.
Schedule the crawl outside business hours at 3–5 URLs/second for shared hosting. Taking a client's site down during an audit is a memorable way to lose an account.
Step 3: Find the index gap
This is where the real findings live. You now have three URL lists: what your crawler found, what the XML sitemap claims, and what GSC says is indexed. Compare them.
Orphan pages
URLs in the sitemap or analytics with zero internal links. On a typical ecommerce site we'd expect under 2%. If 18% of product pages are orphaned, that's your headline finding — the category architecture is broken, not the individual pages.
Crawled but not indexed
Usually thin content, duplication, or quality signals. Pull 20 sample URLs and look at them manually. If they're all tag archives or paginated filters, the fix is directive-level, not content-level.
Indexed but shouldn't be
Staging subdomains, ?print= parameters, internal search results, PDF duplicates. Search site:example.com inurl:? and see what falls out. One B2B client we've seen had 31,000 indexed faceted-navigation URLs against 400 real pages — crawl budget was going almost entirely to junk.
Step 4: Audit directives and canonicals
Check for contradictions, which are more common than outright errors:
- Pages with noindex that are also in the XML sitemap
- Canonical tags pointing to redirecting or 404 URLs
- Self-referencing canonicals missing on paginated series
- robots.txt blocking URLs that carry canonical tags (Google can't read the tag if it can't crawl the page)
- Mismatched hreflang return tags on international sites — this breaks silently and is worth checking even on 3-language setups
Step 5: Architecture, internal links, and redirects
Pull crawl depth distribution. On a healthy mid-size site, 80%+ of revenue pages sit within three clicks of the homepage. If your top 50 converting product pages average depth 6, you have an internal linking project, and it's usually higher ROI than anything else in the audit.
Then check redirects: chains longer than two hops, redirect loops, and 302s that should be 301s. Export every internal link pointing at a redirect — fixing the source links is a one-time dev task that recovers crawl efficiency immediately.
Step 6: Performance and rendering
Use field data (CrUX via GSC or PageSpeed Insights), not lab scores. Lab scores make clients panic about numbers that don't affect rankings. Record LCP, INP, and CLS for mobile, grouped by template: homepage, category, product, blog post. Fixing the product template fixes 14,000 URLs; fixing one page fixes one.
For JavaScript sites, test rendering directly: fetch a key page with the URL Inspection tool and check whether the main content, internal links, and canonical appear in the rendered HTML. If your product descriptions only exist post-hydration and the hydration takes 4 seconds, that's a finding worth more than the rest of the audit combined.
Step 7: Structured data, sitemaps, robots.txt
Validate schema against Google's Rich Results Test

Nick Quirk

