THL React SPA SEO Remediation
Diagnose and fix the '33 pages canonicalising to homepage' failure mode on React + Vite SPAs hosted on Replit (the exact issue currently affecting techhorizonlabs.com and academy.techhorizonlabs.com). Covers pre-render / SSR / SSG trade-offs, Option-A implementation path, build-time meta + JSON-LD injection, sitemap regeneration, and Googlebot UA validation.
Four-step sequence: diagnose with curl + GSC → choose strategy → implement pre-render → validate with Googlebot UA. Reusable as a client deliverable for other Replit-hosted SPAs.
When to use
Trigger on:
- "Fix the SEO on [site]" / "Improve the SEO"
- "Googlebot sees the homepage" / "all my pages look the same to Google"
- "Ahrefs health score is [low]" / "duplicate content warnings"
- Search Console audit shows "duplicate, Google chose different canonical"
- Any React + Vite SPA hosted on Replit where /about, /pricing, /blog etc. all render to the homepage in curl
The diagnosis
Step 1 — confirm the failure mode
Run three checks before picking a fix:
# 1. What Googlebot sees on an interior page curl -A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" \ https://example.com/interior-page -i | head -60 # 2. Compare to what a browser sees curl https://example.com/interior-page -i | head -60 # 3. Compare both to the homepage curl https://example.com/ -i | head -60
If the Googlebot-UA response is identical to the homepage response (same <title>, same meta description, same body), confirmed: the SPA is serving a single HTML shell for all routes, and the per-route meta tags are only injected client-side after JS runs. Google indexes the pre-JS HTML.
Step 2 — quantify
Pull from Search Console:
- Coverage report — count of "Duplicate, Google chose different canonical" and "Submitted URL not selected as canonical"
- URL Inspection on 3 representative interior pages — check the "Rendered HTML" view. If meta tags populate there but not in the initial HTML, Google's renderer is getting them eventually, but the primary index pass is seeing the shell.
- Ahrefs / Sitebulb (if available) — crawl, look for: thin pages <200 words (because only the shell rendered), duplicate titles, duplicate descriptions
Step 3 — decide the severity
Not every SPA needs a rewrite. Severity tiers:
- Critical — the site depends on organic traffic (commerce, content, lead-gen SEO). Fix now.
- Moderate — the site gets ~20–40% of traffic from organic. Fix in the next 30 days.
- Low — the site is primarily used by logged-in users; organic is ~5–10%. Fix when other priorities clear.
For THL's own sites (techhorizonlabs.com, academy.techhorizonlabs.com) this is Moderate-to-Critical: lead flow benefits materially from organic.
Strategy options
Three implementation paths, ranked by effort and fit for a Replit + Supabase + Vite stack:
Option A — Pre-render critical pages at build time
What: Use a Vite plugin (e.g. vite-plugin-ssr in SSG mode, or vite-plugin-prerender, or a custom Puppeteer script) to render each route to static HTML at build time. The rendered HTML has the real title / meta / content populated. JS still hydrates on the client for interactivity.
Fit for Replit + Supabase: Excellent if the public pages have mostly-static content (homepage, about, pricing, feature pages, blog) and the dynamic bits (logged-in state, personalised content) can hydrate in after first paint.
Effort: 1–2 days for a typical THA-scale site (45 pages).
Tradeoffs:
- Pre-rendered pages are a point-in-time snapshot. Content changes require a rebuild + redeploy.
- Not suitable for content that changes per-request (search results, user dashboards, personalised pages). But those pages usually aren't the ones you want indexed anyway.
When to pick: Default recommendation for THL-style sites. Covers 90% of the SEO impact with least infrastructure change.
Option B — Migrate to SSR via Next.js / Remix / Astro
What: Wholesale move from Vite SPA to a framework that server-renders by default. Pages render on each request with full meta/content.
Fit for Replit + Supabase: Works, but Replit's autoscale deployment isn't Next.js's ideal host. Vercel or Netlify are better Next.js hosts.
Effort: 5–15 days depending on codebase size, auth flow complexity, routing conventions.
Tradeoffs:
- Largest rewrite cost
- Gain: personalised / dynamic content can also be SEO-visible
- Lose: forced architectural discipline (SSR requires server-safe code paths)
When to pick: The site has a substantial amount of per-user or per-request dynamic public content (marketplace, directory, user-generated content) AND the team has budget for the migration.
Option C — Hybrid: keep the Vite SPA, pre-render only the top N pages
What: Same tooling as Option A, but you explicitly whitelist the 5–15 pages that matter for SEO (homepage, pricing, each pillar blog post, each product page). Everything else stays client-rendered.
Fit: Good when only a minority of routes benefit from SEO.
Effort: Half a day.
Tradeoffs: Leaves the long tail unsearchable, but that's often fine.
When to pick: For sites where organic matters but mostly for a handful of landing pages.
Recommendation logic (decision tree)
How many routes benefit from organic?
├── Fewer than 15 critical routes → Option C (hybrid pre-render)
├── 15–100 critical routes, mostly static → Option A (full pre-render)
├── 100+ routes OR per-request dynamic content needs SEO → Option B (SSR migration)
For THL's own sites: Option A is the right call. ~45 pages, mostly static, one codebase, no meaningful per-request personalisation on public routes.
Implementation path for Option A (pre-render)
This is the default path. Concrete steps for a Vite + React + Replit stack:
1. Install and configure a pre-renderer
Options, in order of simplicity:
vite-plugin-ssg— configure routes to pre-render, emit HTML alongside the JS bundlereact-snap— post-build step that boots a headless Chromium, crawls, writes HTML- A custom Puppeteer script invoked after
vite build— most control
Pick based on route complexity. If you use wouter (like THA), the Vite SSG path is cleanest because wouter's static route list can be enumerated.
[Huxley: for THA specifically, confirm whether you want to use vite-plugin-ssg or react-snap. I'll default to vite-plugin-ssg for the implementation — it produces cleaner HTML and is more maintained as of April 2026. Swap if you prefer react-snap.]
2. Enumerate the route list at build time
Wouter doesn't have a static route API, so export a routes manifest. In client/src/App.tsx, extract routes into a separate routes.ts that both App.tsx and the pre-render script can import:
// client/src/routes.ts export const PUBLIC_ROUTES = [ "/", "/about", "/pricing", "/membership", "/audit", "/ai-agents-guide", "/learning-hub", "/wiki", "/prompts", // ...enumerate all public pages ] as const;
Dynamic routes (/workshop/:slug) need either a paginated sitemap fetch (from Supabase at build time) or explicit listing.
3. Inject per-route meta at render time
Your SEO component (client/src/components/SEO.tsx) already does this client-side via react-helmet-async. In the pre-renderer, make sure:
- Helmet's
HelmetProvideris wrapping the server-side render HelmetServerStateis captured afterrenderToStringand its output is spliced into the<head>of the emitted HTML
If using vite-plugin-ssg, this is handled via the onRendered hook. Custom script pattern:
// script/prerender.ts import { renderToString } from "react-dom/server"; import { HelmetProvider } from "react-helmet-async"; import { PUBLIC_ROUTES } from "../client/src/routes"; for (const route of PUBLIC_ROUTES) { const helmetContext = {}; const html = renderToString( <HelmetProvider context={helmetContext}> <App initialRoute={route} /> </HelmetProvider> ); const { helmet } = helmetContext; const fullHtml = ` <!DOCTYPE html> <html> <head> ${helmet.title.toString()} ${helmet.meta.toString()} ${helmet.link.toString()} ${helmet.script.toString()} </head> <body> <div id="root">${html}</div> <script type="module" src="/assets/index.js"></script> </body> </html> `; fs.writeFileSync(`dist/public${route}/index.html`, fullHtml); }
4. Handle dynamic Supabase-loaded content
For pages whose content comes from Supabase at runtime (e.g. a workshop detail page), two options:
- Fetch at build time. The pre-render script reads from Supabase using the service role key, renders each known record to its own static HTML file. New records need a rebuild (set up an on-publish hook, or rebuild nightly).
- Skip from pre-render, mark noindex. For truly dynamic content that shouldn't rank, add
<meta name="robots" content="noindex">and drop the route from the sitemap.
5. Meta tag + JSON-LD — build-time discipline
Per the earlier UI/UX audit, THA's SEO.tsx already emits Organization + Event + Product + FAQPage structured data. The pre-render step just has to serialise these into the static HTML.
Additions worth adding at the same time (easier to do during this migration):
BreadcrumbListon every interior pageVideoObjecton pages with embedded videos (especially the homepage testimonials)Courseschema on workshop detail pages- Per-page unique meta descriptions — audit each page's current description and rewrite any duplicates
6. Sitemap generation
Already exists at server/seoRoutes.ts. Confirm it reflects the pre-rendered route list. Two gotchas:
- Dynamic workshop page fetch limit is currently 50 — raise to 200 or paginate via sitemap-index if the workshop catalog grows
lastmodshould reflect the build time, notnew Date()at request time (prevents cache-churn signals to Google)
7. Validation
After deployment:
# 1. Googlebot sees correct HTML curl -A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" \ https://academy.techhorizonlabs.com/audit -i | grep -A2 "<title>" # 2. No "homepage" content leaking curl -A "Googlebot" https://academy.techhorizonlabs.com/audit | grep -c "[homepage-only string]" # Expect 0 # 3. Structured data validates # Open https://search.google.com/test/rich-results and test 5 representative URLs
Then in Search Console:
- Submit the sitemap
- Request indexing on 3–5 representative URLs
- Wait 2–4 weeks for the coverage report to update
- Watch the "Discovered, not indexed" and "Duplicate, Google chose different canonical" counts drop
GEO (generative engine optimisation) layer
Once the SEO fix is in, GEO is the next lever. The pre-rendering step gives you the foundation — AI crawlers like ChatGPT, Perplexity, Claude web, and Gemini consume static HTML the same way Googlebot does.
GEO-specific additions (do after pre-render is working):
- Answer-format content — lead each page with the question as H1/H2 and the direct answer in the first paragraph. AI search engines lift these.
- Cite sources — any claim with a number should link to the source. This is how AI search engines decide who to cite.
- FAQPage schema on any interior page with Q&A content
- Entity establishment — Organization schema with
sameAspointing to LinkedIn, Crunchbase, and any other authoritative profile - Topical authority — interlink related pages with descriptive anchor text (not "click here" / "learn more")
llms.txt is tempting but currently not confirmed to affect AI search citation rates. Deprioritise.
See the ai-seo skill (if installed) for the full GEO playbook.
Reusable as a client deliverable
This same playbook is what THL runs for client Replit sites. When running it as a paid engagement:
- Discover: the diagnosis step above becomes the "Audit Report" deliverable
- Architect: the strategy choice becomes a short doc + sign-off from the client
- Activate: the implementation is billed against the Partner Tier credit pool
Typical effort for a client site:
- 30 pages or fewer, mostly static: 1–2 days (Option A)
- 30–100 pages with Supabase dynamic content: 3–5 days (Option A with build-time fetch)
- 100+ pages or SSR migration: quote separately
Related skills
seo-audit— the broader SEO audit framework; use that first if the client hasn't already diagnosedthl-partner-tier-proposal-generator— when this becomes a paid client engagementai-seo— the AI-search companion to this skill (GEO specifics)
— huxley
- →Fix the SEO on academy.techhorizonlabs.com.
- →Googlebot sees the homepage for every page on this site.
- →Ahrefs health score is 42 — what's wrong with our React SPA?
Source
official
Author
Tech Horizon Labs
Version
1.0
Complexity
Compatible With
Prerequisites
- React + Vite SPA codebase
- Admin access to DNS / Search Console
- Staging environment for pre-render validation
Best For
Tags
