You run a crawl on your own site and the tool spits out 400 issues. Red links, missing meta descriptions, "low word count" flags on pages you wrote deliberately short, duplicate titles you swear you fixed last year. You fix ten of them, nothing moves, and you close the tab convinced the whole exercise is theatre.
I've been there. Twice, actually, on the same site. The second time around I stopped treating the audit as a to-do list and started treating it as a diagnostic. That shift is the entire game. An SEO audit to improve site health only pays off when you sequence it correctly — crawl first, decide what "healthy" even means for your site, then fix in the order that protects your ability to be found.
Below is the sequence I use now, with the numbers I hold myself to at each step. It's not a 47-point checklist. It's seven moves, in order, because doing them out of order is how people waste a quarter fixing things nobody was going to find anyway.
Key Takeaways
- Crawl before you fix. You cannot diagnose what you haven't measured.
- Indexation is the bottleneck, not rankings. If a page isn't indexed, on-page work on it is pointless.
- Set numeric thresholds for "healthy" before you start, or you'll fix forever.
- Priority is not the same as severity. A broken link on a money page outranks a broken link in a 2019 blog post.
- Re-audit on a schedule. Regression is normal and silent.
- Site health now includes how your pages render inside AI answer engines, not just Google.
What does an SEO audit actually mean in 2026?
An SEO audit is a measurement pass, not a repair pass. You're collecting evidence about three things: whether search engines can reach your pages, whether they've chosen to store them, and whether the stored versions serve the query they're ranking for.
Most of what you read online frames this as a checklist of categories — technical, content, UX, mobile — which is fine as taxonomy and useless as a workflow. Categories don't tell you what to do Monday morning.
The other thing worth naming: the definition has quietly widened. Being crawlable and indexable gets you into the index. It no longer guarantees you'll be surfaced, because a growing share of discovery now happens inside AI-generated answers that summarise pages rather than link to them. I'll come back to what that means practically, because it changes one step in the sequence.
Technical, content, and AI readability are three different audits
Treat them separately, even if you run them in the same week.
- Technical: reachability, indexation, canonicalisation, rendering, structured data, performance budgets.
- Content: query-to-page match, thin or overlapping pages, internal linking, whether each URL has one clear job.
- AI readability: whether a summarising engine can extract a clean answer from your page — clear claims, stated numbers, unambiguous entities.
The third one is genuinely new and most audit templates still ignore it. I'll get to it. First, the crawl.
Step 1: Run a crawl you trust, and check the crawl budget
Every audit starts with a full crawl of the live site, from a fresh user agent, with JavaScript rendering on if your site depends on it. Two crawlers, ideally, because they disagree usefully.
Here's a figure that surprised me the first time I checked properly: on a mid-sized content site I was working on, roughly one page in five was being crawled but never indexed. Not blocked. Not erroring. Just crawled, evaluated, and dropped. That's a signal, and it's invisible in a standard "errors" report because it isn't an error.
So separate your crawl output into four buckets before you read a single recommendation:
- Discovered and indexed
- Discovered and not indexed
- Blocked from crawling
- Orphaned — reachable only via sitemap, not via any internal link
Bucket 2 is where the leverage usually hides. Bucket 4 is where you'll find pages you'd forgotten existed.
The robots.txt and noindex combo that cost me two weeks
Early on, I inherited a site where a previous developer had added a staging Disallow: / to robots.txt that never got removed on launch. Meanwhile half the templates carried a stray noindex tag. Neither alone would have been fatal. Together, they made diagnosis confusing, because a noindexed page blocked by robots.txt can't be reported as noindexed — the crawler never gets far enough to see the tag.
The lesson stuck: check robots.txt and meta robots before anything else. It takes four minutes and it invalidates or explains half of what you're about to see.
Step 2: Establish indexation health before touching on-page
Pull your indexed-page count and compare it to your published-page count. Then subtract the pages you intentionally keep out of the index: tag archives, internal search results, thank-you pages, filtered variants.
What's left over is your real problem set.
I hold myself to a rough benchmark: if more than 15% of your intended-to-rank pages aren't indexed, that's the entire audit. Don't open the content spreadsheet yet. Find out why they're excluded first, and fix that, because every hour spent optimising a page that search engines have declined to store is an hour you don't get back.
How to tell an indexation problem from a quality problem
The two look identical in a dashboard and have opposite fixes.
- Indexation problem: the URL is blocked, canonicalised elsewhere, noindexed, or the server returns a non-200 status to the crawler.
- Quality problem: the URL is fully accessible and simply wasn't selected — usually because the page duplicates a better one, carries almost no unique content, or answers nothing specific.
The first is a five-minute fix. The second usually means the page shouldn't exist in that form. Merging two weak pages into one strong page has done more for my sites than any amount of keyword editing.
Step 3: Fix canonicalisation before content thinness
Canonical chains are the most common silent problem I find. Page A canonicals to B, B canonicals to C, and C canonicals to a URL that returns a redirect. Search engines eventually resolve it, but you've asked them to do work for no reason, and the signals get muddy in the meantime.
Self-referencing canonicals should be exactly that — self-referencing. Anything else on a page you want ranked deserves scrutiny.
Also audit trailing-slash and case variants. On one site I found four live variants of the same URL across http/https and slash/no-slash, all returning 200, none canonicalising to each other. Four pages competing for one query, none of them winning.
Step 4: Measure Core Web Vitals with a budget, not a grade
"Needs improvement" tells you nothing actionable. Set numeric budgets and measure against them.
My standing targets, mobile-first because that's where the traffic is:
- LCP under 2.5 seconds on the 75th percentile of real user sessions — not lab data, field data.
- INP under 200 milliseconds.
- CLS under 0.1.
Lab tools are for debugging. Field data is for judging. If your lab score is green and your field score isn't, you have a real-user problem that synthetic tests can't see — usually a third-party script loading differently depending on geography or consent state.
| Signal | What it measures | My threshold | Typical culprit |
|---|---|---|---|
| LCP | When the main content is painted | < 2.5s (p75, field) | Hero image size, font loading, server response |
| INP | Responsiveness to interaction | < 200ms | Long JavaScript tasks, heavy event handlers |
| CLS | Visual stability | < 0.1 | Images without dimensions, late-injected banners |
| TTFB | Server response start | < 800ms | Uncached queries, slow origin, no CDN |
The 2 MB of tracking scripts I kept defending
I once argued for six months that our performance problem was the CMS. Then I actually counted the third-party scripts in the head: nine of them, roughly 2 MB combined, three loading synchronously. Removing two redundant analytics tags cut LCP by about 1.1 seconds on mobile. The CMS was never the issue. I was.
Worth doing that count before you blame your stack.
Step 5: Audit content in clusters, not page by page
Reading pages one at a time produces a list of small improvements and no strategy. Group them by the query cluster they serve, then ask one question per cluster: does this group of pages cover the topic completely, or does it repeat the same point five ways?
Repeating the same point five ways is the most common content pattern I see. The fix isn't more words on each page. It's consolidation.
- Merge pages that target the same intent with near-identical angles.
- Keep the URL with the most inbound internal links, redirect the rest.
- Rewrite the survivor so it clearly outperforms what it absorbed.
- Delete anything that exists only because a keyword tool said the volume was there.
Internal linking belongs in this step, not the technical one. Links are how you tell a search engine which page in a cluster is the primary one. If three pages all target the same phrase and none links to the others, you've made the choice impossible for anyone but you.
Step 6: Check whether AI answer engines can read your pages
This step barely existed a few years ago and now it's the one I'd protect if forced to drop something else.
The mechanics are straightforward. A summarising system fetches your page, extracts claims, and decides whether to reproduce them. Pages that get cited tend to share a few properties: a direct answer near the top, specific figures stated plainly, unambiguous named entities, and content that doesn't require executing five JavaScript bundles to appear.
Things that hurt: answers buried under three paragraphs of preamble, vague qualifiers instead of numbers, and key content locked behind client-side rendering that a fetch doesn't trigger. I've watched a page with an excellent written answer go uncited for months, purely because the text rendered client-side and the fetching system saw an empty container.
Practical checks I run now:
- Fetch a handful of key URLs with JavaScript disabled and see what text actually arrives.
- Confirm the first 100 words of each page state the answer, not the context.
- Make sure your structured data matches what's visible on the page. Mismatches get ignored at best.
This isn't a replacement for the rest of the audit. It's a filter applied on top of it.
Step 7: Schedule the re-audit and watch for regression
An audit is a snapshot. Site health decays, and it decays quietly — a deploy drops a canonical tag, a plugin update restores a blocked path, someone adds a fourth variant of your URL pattern.
The cadence I use:
- Weekly: automated crawl, alerts only on new hard errors — 404s on linked pages, sudden indexation drops, robots.txt changes.
- Monthly: indexation count versus published count, Core Web Vitals field data.
- Quarterly: full crawl plus a content cluster review.
Regression is the norm, not the exception. I've had a single deploy undo six weeks of canonical fixes because a template default got restored.
Prioritising when you can't fix everything
You will never clear the list. That's fine. Rank by two axes: how many pages a fix affects, and how close those pages are to something that matters — revenue, signups, or your main hub content.
A sitewide template fix that resolves a canonical problem on 300 pages beats 300 individual meta description edits, every time. And a broken internal link pointing at your highest-converting page beats a broken link in a 2019 archive post, even though a crawler flags them identically.
One warning, learned the hard way: don't batch a hundred URL changes into a single deploy. I did that once, introduced a redirect loop across the whole set, and spent a weekend untangling it. Ship in groups of ten to twenty and verify between each batch.
Questions I get asked about SEO audits
How often should you run a full SEO audit?
Quarterly for most sites, monthly if you ship changes frequently or run an e-commerce catalogue that changes daily. The full crawl is the expensive part, so automate a lighter version weekly and reserve the deep pass for once a quarter. What matters more than frequency is having a baseline you measured before you started changing things, because without it you cannot tell improvement from noise.
What should you fix first after an audit?
Whatever prevents pages from being indexed. Blocked crawling, stray noindex tags, canonical chains, and non-200 status codes come before anything about keywords, titles, or content length. A perfectly optimised page that search engines won't store is worth zero, and this is the single most common sequencing mistake I see. Once indexation is clean, move to sitewide template issues, then page-level content.
Do you need paid tools to audit a site?
No, but you need at least one crawler that renders JavaScript and one source of field performance data. Free options cover the crawl and the vitals reasonably well for a small site. Where paid tools earn their cost is history — being able to see that a page dropped out of the index three weeks ago, and what changed that week. That timeline is hard to reconstruct manually and it's usually where the answer lives.
Can an SEO audit hurt your rankings?
The audit can't. The remediation can, if you're careless. Mass URL changes without redirects, bulk content deletion, and noindex tags left on after a staging push are the three ways I've seen an audit cause damage. Test every change on a small subset of URLs first, verify in the index, then roll it out. Slow is smooth here.
The part most audits get wrong
The report is not the work. Nobody has ever improved a site by exporting a CSV.
What actually moves things is deciding, before you start, what "healthy" means in numbers for your specific site — an indexation ratio you'll accept, a vitals threshold you'll hold, a maximum number of internal clicks to your key pages — and then measuring against that definition on a schedule. Everything else is triage.
And keep the failures. The redirect loop, the staging robots.txt, the two megabytes of scripts I defended for six months. Those are the entries that teach you which signals to check first next time, and they're the reason a second audit on the same site takes a fraction of the time the first one did.