How to Write a Testable Technical SEO Specification for a Website Migration
Build a testable technical SEO specification for a migration, covering templates, crawl rules, canonicals, sitemaps, mobile parity and QA.

A migration specification is useful when a developer can implement it and a QA tester can decide whether it passed. Start with page types, not a general instruction to “preserve SEO.” For each type, record the intended URL, what crawlers should reach, what should be eligible for indexing, and what the rendered page must contain. Then attach an owner and a test.
The URLs below are illustrative, not findings about a live site. Replace them with the migration’s approved patterns before work begins.
Start with a page-type requirements matrix
| Page type and example URL | Crawl and index intent | Canonical, metadata and structured data | Verification method |
|---|---|---|---|
Article listing: https://www.example.com/articles?page=2 |
Crawlable; intended for indexing | Self-referencing canonical; page-specific title and description; no structured data specified | Crawl the URL and inspect its links, canonical and page output |
Product: https://www.example.com/products/widget |
Crawlable; intended for indexing | Self-referencing canonical; product title and description; Product markup only for information visible on the page | Compare rendered content with metadata and markup; validate markup |
Member preview: https://www.example.com/members/preview |
Crawlable, but excluded from indexing with noindex |
Record its canonical decision and metadata explicitly; no structured data specified | Fetch the page and inspect its robots directive |
Internal search: https://www.example.com/search/term |
Crawl blocked under the proposed /search/ rule; do not rely on a page-level noindex here |
Not included in the sitemap; no structured data specified | Inspect robots.txt and test the URL pattern |
A matrix makes an omission visible: “same as the old site” does not specify what a newly generated URL should do. Record the old-to-new URL mapping separately. Redirect destinations and acceptance criteria—including status codes and chain handling—need confirmation against the applicable search-engine and platform documentation before implementation; this matrix does not settle them.
Specify links and rendered output
For paths crawlers need to follow, require ordinary HTML links. The pagination requirement should say that each page links to another page in the series and has its own canonical—not that every page canonicalizes to page one. The pagination guidance explains why access to deeper pages matters for discovering their links.
Here is illustrative output for two pages; each URL’s canonical belongs in that URL’s HTML:
<!-- Output at https://www.example.com/articles?page=1 -->
<head>
<link rel="canonical" href="https://www.example.com/articles?page=1">
</head>
<nav aria-label="Article pages">
<a href="/articles?page=2">Page 2</a>
</nav>
<!-- Output at https://www.example.com/articles?page=2 -->
<head>
<link rel="canonical" href="https://www.example.com/articles?page=2">
</head>
<nav aria-label="Article pages">
<a href="/articles?page=1">Page 1</a>
</nav>
If important content is produced by JavaScript, specify what a requester receives and what the rendered page contains. Secondary rendering guidance recommends server-side or hybrid rendering for important content rather than treating bot-only dynamic rendering as the long-term solution. Make that a testable output requirement instead of assuming a framework choice proves the content is available.
Keep crawl rules separate from index rules
robots.txt controls the proposed crawl restriction; a page-level noindex is an index-exclusion instruction that must be available for a crawler to see. For example, the following illustrative rule blocks the internal-search path, not the member preview:
User-agent: *
Disallow: /search/
The crawlable member preview could instead return this page-level directive:
<head>
<meta name="robots" content="noindex">
</head>
Do not add /members/preview to that Disallow rule while relying on its noindex tag. Check both files and pages together; checking either in isolation can miss the conflict. This distinction follows the crawl and noindex guidance.
Define sitemap, mobile and markup acceptance criteria
Require the CMS to regenerate the XML sitemap as site state changes. Every listed URL must be the chosen canonical, eligible for indexing and return HTTP 200; exclude duplicates and non-canonical variants. The example entry below is acceptable only if the live page meets those conditions. A sitemap entry is not proof of indexing.
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://www.example.com/products/widget</loc>
</url>
</urlset>
These criteria and regeneration requirement are described in the sitemap guidance. If mobile and desktop versions differ, add a comparison test for primary content, headings, structured data, titles, descriptions, robots directives and canonicals. The mobile parity guidance calls for those elements to match.
Specify structured data against visible page content, not an inventory value copied from another system without checking the page. This small illustrative fragment, not tested production code, says the widget is out of stock in both the visible text and JSON-LD:
<h1>Widget</h1>
<p>Availability: Out of stock</p>
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "Product",
"name": "Widget",
"offers": {
"@type": "Offer",
"availability": "https://schema.org/OutOfStock"
}
}
</script>
An “In stock” value in that markup would contradict the page. Validate the implemented markup with Google’s Rich Results Test or Search Console, as recommended in the structured-data guidance. Validation can identify markup problems; it does not promise a rich result.
Turn the matrix into a handoff test
Assign named people in the actual project plan; the roles here are placeholders. Give each requirement a pass condition that QA can observe on representative URLs before release and again after launch.
| Requirement | Owner | Pass condition |
|---|---|---|
| Template URLs, links and canonicals | Development | A site crawl finds the expected links and self-referencing canonicals on sampled listing pages; unexpected response codes and canonical errors are investigated |
| Crawl and index directives | SEO lead and development | robots.txt matches the approved blocked patterns, and sampled noindex pages remain crawlable and expose the directive |
| Generated sitemap | CMS owner | Every listed URL checked is canonical, indexable and HTTP 200; the sitemap changes when eligible site URLs change |
| Mobile parity | QA | Where versions differ, the specified content, headings, metadata, directives, canonicals and structured data match |
| Structured data | Development and QA | Visible values agree with markup, and the implemented page is checked with a validation tool |
A site crawler can check response codes, canonical errors and whether expected links are present; it tests implementation and discoverable paths, not search indexing. After launch, review server logs for search-bot requests to see which URLs were requested and when. Logs establish observed crawling, not whether a page was indexed or ranked. Keep those observations separate from the pass conditions: a correct handoff test says what the site delivered, while search-engine reports describe what happened afterward.




