Next.js App Router SEO: What Crawlers Actually Receive from SSG, ISR and SSR
Key takeaways
For SEO, the rendering mode matters less than what the HTML response contains and which status code it carries. This post looks at App Router caching (including the Next.js 15 and 16 changes), metadata, 404s under streaming, canonicals and sitemaps from a crawler's point of view.
Most SEO discussions about Next.js start with “SSR or SSG?”. From a crawler’s point of view that is the wrong first question. Googlebot does not know or care how your HTML was produced. It sees a status code, headers, and a document. The things that actually go wrong in App Router projects are more specific: a page that silently became dynamic and slow, a stale cached page that still shows a deleted product, a “not found” page served with a 200, a canonical pointing at the wrong host, or metadata that is not in the first HTML chunk for crawlers that do not run JavaScript.
This post goes through those in the order a request hits them.
What the rendering modes mean in the App Router
The App Router has no page-level getStaticProps/getServerSideProps switch. Instead, every route is either prerendered (at build time or on first request, then served from cache) or rendered per request, and Next.js decides that from what the route does:
- Static (SSG-like): no request-specific APIs, and every data source is cacheable. HTML is generated once and served as a file.
- ISR: static, but cached data has a lifetime (
revalidate) or can be invalidated on demand by path or tag. The first request after expiry still gets the old page while a fresh one is generated in the background (stale-while-revalidate). - Dynamic (SSR-like): the route reads
cookies(),headers(),searchParams, or uses uncached data. It renders on every request.
The build output (next build) marks each route as static or dynamic. That table is the first thing I check on any SEO complaint, because a route people assume is static often is not.
The fetch default changed, and old snippets are wrong
In Next.js 13 and 14, fetch in a server component was cached by default (force-cache). Next.js 15 reversed that: fetch is no longer cached by default. A lot of tutorials still say “the default is force-cache”, which leads to confusing results.
With the Next.js 15 default, a route with a plain fetch and no dynamic APIs is still prerendered at build time, so the data is fetched once during next build and baked into the HTML. That looks like caching, but it is really “the route is static”. Add a cookies() call anywhere in the tree and the same fetch now runs on every request.
Being explicit avoids the guesswork:
// app/posts/page.tsx
export default async function PostsPage() {
const res = await fetch('https://api.example.com/posts', {
next: { revalidate: 3600, tags: ['posts'] },
});
const posts: { id: string; title: string }[] = await res.json();
return (
<ul>
{posts.map((p) => (
<li key={p.id}>{p.title}</li>
))}
</ul>
);
}
This route is ISR: cached for up to an hour, and invalidatable by the posts tag.
// app/dashboard/page.tsx
import { cookies } from 'next/headers';
export default async function DashboardPage() {
const session = (await cookies()).get('session')?.value;
const res = await fetch('https://api.example.com/me', {
headers: { Authorization: `Bearer ${session}` },
cache: 'no-store',
});
const user = await res.json();
return <div>{user.name}</div>;
}
This one is dynamic, which is correct: it is per-user and should not be indexed anyway (put it behind auth and add robots: { index: false } in its metadata).
Invalidation: revalidateTag in Next.js 16
On-demand invalidation is what keeps ISR pages from serving stale content to crawlers after an edit. In Next.js 16 the API changed: revalidateTag now takes a second argument, a cache-life profile, and the single-argument form is deprecated.
// app/api/revalidate/route.ts (called by the CMS webhook)
import { revalidateTag } from 'next/cache';
export async function POST(req: Request) {
if (req.headers.get('x-revalidate-secret') !== process.env.REVALIDATE_SECRET) {
return new Response('forbidden', { status: 403 });
}
revalidateTag('posts', 'max'); // mark stale; next request triggers a background refresh
return Response.json({ ok: true });
}
With a profile such as 'max', the tagged data is marked stale and refreshed in the background, so the next visitor (or crawler) may still get the old version once. When a Server Action needs the same user to see their own write immediately, Next.js 16 adds updateTag(tag), which expires the entry right away. It can only be called from Server Actions.
Two things tend to break here. First, tags must match exactly between the fetch and the revalidateTag call, and a typo fails silently. Second, a CDN in front of Next.js (Cloudflare, a custom Fastly config) has its own cache. Revalidating inside Next.js does nothing for HTML the CDN already holds, so the webhook also needs to purge the CDN path, or the CDN TTL must be short enough to not matter.
Route segment config and dynamic params
Segment config can force behavior for a subtree:
// app/admin/layout.tsx
export const dynamic = 'force-dynamic';
dynamic accepts 'auto' | 'error' | 'force-static' | 'force-dynamic'. 'error' is underused: it fails the build if anything in the route would make it dynamic, which turns “this marketing page accidentally became SSR” into a build error instead of a slow page.
For dynamic segments, generateStaticParams decides which paths are prerendered, and dynamicParams decides what happens for the rest:
// app/blog/[slug]/page.tsx
export async function generateStaticParams() {
const slugs = await getAllSlugs();
return slugs.map((slug) => ({ slug }));
}
export const dynamicParams = false; // unknown slugs return 404 instead of rendering on demand
With dynamicParams = true (the default), paths not returned at build time are rendered on first request and then cached. That is the right choice for a large catalog where prerendering everything would make builds too slow. With false, any slug not in the list is a real 404, which is exactly what you want when the list is complete.
If you enable Cache Components (cacheComponents in Next.js 16), caching moves to the 'use cache' directive with cacheLife and cacheTag, and the dynamic, revalidate and fetchCache segment options no longer apply. The SEO concerns stay the same; only where you express them changes.
Status codes: the 404 that returns 200
This is the App Router behavior I see cause the most indexing noise. notFound() normally produces a 404. But if the page streams (it has a loading.tsx, or the lookup happens inside a <Suspense> boundary), the response headers, including the 200 status, have already gone out by the time the component calls notFound(). Next.js cannot change the status any more, so it injects <meta name="robots" content="noindex"> into the streamed HTML instead.
Google honors that meta tag, so the page will not be indexed, but in Search Console it shows up as “Excluded by noindex tag” rather than “Not found (404)”, and other tools that look at the status code see a success. The fix is to do the existence check before anything streams:
// app/blog/[slug]/page.tsx
import { notFound } from 'next/navigation';
export default async function PostPage({ params }: { params: Promise<{ slug: string }> }) {
const { slug } = await params; // params is a Promise since Next.js 15
const post = await getPost(slug);
if (!post) notFound(); // before any Suspense boundary renders
return (
<article>
<h1>{post.title}</h1>
<Body post={post} />
</article>
);
}
Calling notFound() in generateMetadata works as well, since metadata for the same slug usually needs the same lookup. The same logic applies to redirect()/permanentRedirect(): before streaming you get a real 307/308; after streaming starts, Next.js falls back to a client-side redirect via a meta refresh.
Metadata: canonical, alternates and streaming
generateMetadata is where titles, descriptions and canonicals come from. Set metadataBase once in the root layout so relative URLs resolve against the production host rather than whatever host served the request:
// app/layout.tsx
export const metadata = {
metadataBase: new URL('https://www.example.com'),
};
// app/blog/[slug]/page.tsx
export async function generateMetadata({ params }: { params: Promise<{ slug: string }> }) {
const { slug } = await params;
const post = await getPost(slug);
if (!post) notFound();
return {
title: post.title,
description: post.excerpt,
alternates: {
canonical: `/blog/${slug}`,
languages: { en: `/en/blog/${slug}`, ko: `/blog/${slug}` },
},
};
}
A mistake I have made myself: deploying preview environments without metadataBase (or with it read from an env var that was unset on preview), so every preview page declared a canonical on the preview domain. If previews are publicly reachable, that tells Google the preview host is authoritative. Block previews with a noindex header or password protection, and hard-code the production origin in metadataBase.
Since Next.js 15.2, metadata from generateMetadata can be streamed: for normal browsers the page shell is sent first and the <title> and <meta> tags arrive later in the body. Next.js keeps a list of “HTML-limited” user agents (Bingbot, social preview crawlers like facebookexternalhit, Twitterbot, Slackbot, and various Google fetchers such as Google-InspectionTool) that get a blocking render with metadata in <head>. Googlebot itself is treated as a JavaScript-executing crawler and receives the streamed version. If some other crawler you care about misses titles, add its user agent to the htmlLimitedBots option in next.config.
Sitemaps and robots
The App Router generates both from route files:
// app/sitemap.ts
import type { MetadataRoute } from 'next';
export default async function sitemap(): Promise<MetadataRoute.Sitemap> {
const posts = await getAllPosts();
return posts.map((p) => ({
url: `https://www.example.com/blog/${p.slug}`,
lastModified: p.updatedAt, // the real edit date, not new Date()
}));
}
// app/robots.ts
import type { MetadataRoute } from 'next';
export default function robots(): MetadataRoute.Robots {
return {
rules: { userAgent: '*', allow: '/', disallow: ['/admin/', '/api/'] },
sitemap: 'https://www.example.com/sitemap.xml',
};
}
Google has stated that it ignores priority and changefreq and uses lastmod only when it is consistently accurate. Setting every lastModified to the build time is a common way to teach Google that your lastmod is meaningless. For more than 50,000 URLs, generateSitemaps splits the output into multiple files.
One interaction with robots.txt worth knowing: do not Disallow a page you want de-indexed via noindex. If the crawler is blocked from fetching the page, it never sees the noindex and the URL can stay in the index as “Indexed, though blocked by robots.txt”.
A per-route policy instead of a label
Rather than calling a site “SSG” or “SSR”, it is more useful to agree per data source how stale it may be and what invalidates it:
| Content | Rendering | Invalidation | Indexable |
|---|---|---|---|
| Docs, legal, marketing | Static | Redeploy | Yes |
| Blog posts, product pages | ISR with tags | CMS webhook → revalidateTag + CDN purge | Yes |
| Search results, filters | Dynamic | n/a | Usually noindex |
| Account, cart, admin | Dynamic | n/a | No |
The last row matters for correctness, not only SEO. Per-user output belongs on dynamic routes. Forcing such a route static with dynamic = 'force-static' does not make it cacheable per user; it makes cookies() and headers() return empty values, so everyone gets whatever was rendered without a session.