seoskills.sh
Catalog/Technical SEO/SEO Sitemaps

SEO Sitemaps

Audit and generate sitemaps and discovery files — validate XML sitemap presence/size/extensions/lastmod, check robots.txt referencing and sitemap-to-canonical consistency, reconcile orphans against the link graph, and produce repaired sitemap entries plus a robots.txt Sitemap line.

Listed ·Updated

174 installs

Install data from skills.sh

Use this skill

Free
$ npx skills add https://github.com/hainrixz/claude-seo-ai --skill seo-sitemaps
Basic pattern scan of SKILL.md text: no matches (checked October 2, 2026)
Runs shell commandsFetches web pages
Uses shared files

This skill is one module of the author's Claude Code plugin and runs the plugin's shared scripts, which a single-skill install does not copy. For the full audit, install the plugin: run "/plugin marketplace add Hainrixz/claude-seo-ai", then "/plugin install claude-seo-ai@claude-seo-ai". That works in Claude Code only.

View the repository

Repository
hainrixz/claude-seo-ai
GitHub stars
71
License
MIT

About this skill

SEO Sitemaps is a Technical SEO skill for AI agents, published in the seoskills.sh catalog. Reach for it when your work involves crawlability, indexing, site architecture, Core Web Vitals, and log-file analysis. Install it with one command and it runs inside your own agent, so the work happens in your workflow, not a separate SEO tool.

Embed a badge

seoskills.sh listing badge

SKILL.md

seo-sitemaps (M17)

Sitemaps are the discovery contract you hand the crawler — they should list exactly the canonical, indexable URLs and nothing else. Schema rules for related markup: references/schema-tier1.md.

Inputs

Work from the PageSnapshot named in your dispatch envelope: read parsed from <run_dir>/pages/<slug>.json plus <run_dir>/site/sitemaps.json (+ site/sitemap-urls.txt), <run_dir>/site/robots.json, and <run_dir>/crawl.json for orphan reconciliation and per-URL status; Grep pages/<slug>.html for verbatim evidence. Deterministic findings already emitted by audit.mjs are listed in <run_dir>/findings.deterministic.json — do not re-emit those ids; add model-judged findings only. If invoked directly with a URL/path and no snapshot exists, first run node "${CLAUDE_PLUGIN_ROOT}/scripts/snapshot.mjs" <target> --out "${CLAUDE_PLUGIN_DATA}/runs" and use the printed snapshot path.

Audits

Working from the PageSnapshot (parsed_rendered when render.used is not none, else parsed) plus site/sitemaps.json and site/robots.json:

  1. Presence & validity: locate XML sitemap(s) (/sitemap.xml, robots Sitemap: lines, sitemap index); parse as well-formed XML against the sitemaps.org schema.
  2. Size limits: each sitemap <=50,000 URLs and <=50MB uncompressed; if exceeded, expect a sitemap index splitting the set.
  3. Extensions: where relevant, validate image:, video:, and news: namespace entries (correct namespace declared, required child elements present).
  4. lastmod accuracy: <lastmod> is valid ISO 8601 and reflects real last-modified time — not a build-time stamp on every URL (which trains crawlers to ignore it).
  5. robots referencing: at least one absolute Sitemap: line in robots.txt.
  6. Sitemap-to-canonical consistency: no URL in the sitemap is noindex, redirected, 4xx/5xx, or non-canonical (self-referencing canonical only). Cross-check indexability with M-indexability.
  7. Orphan reconciliation: diff sitemap URLs against the internal link graph — flag indexable pages absent from the sitemap and sitemap URLs unreachable by internal links.

Fixes

  • AUTO: generate or repair XML sitemap entries (correct <loc>, accurate <lastmod> from observed last-modified data, valid image:/video: extension children where media exists) and add an absolute Sitemap: line to robots.txt. These are additive/deterministic diffs for fix.
  • PROPOSED: removing or splitting entries (e.g. dropping non-canonical/noindex URLs, sharding into a sitemap index) — drafted, accepted per-item.
  • ADVISORY: changing site-wide lastmod strategy or canonical decisions — described, never written by the tool.
  • Never fabricate lastmod times, media URLs, or canonical targets — pull from observed data, ask the user, or leave a clearly-marked TODO placeholder per the schema fixable contract.

Verification

  • From the run (no network): read <run_dir>/site/sitemaps.json and <run_dir>/site/robots.json — the crawl already walked the index, followed gzipped children and recorded the per-file errors.
  • Re-fetch: node "${CLAUDE_PLUGIN_ROOT}/scripts/parse-robots-sitemap.mjs" --url <final_url> [--max-sitemaps 50] [--max-urls 100000] [--no-well-known] — method xml_parse: parses robots.txt + every declared and probed sitemap, checks well-formedness, size caps, namespace/extension validity, and the canonical/noindex consistency assertion. --sitemap <sitemap url> targets one file and --file <robots.txt> [--path /x] parses an offline robots file. This script has no --snapshot mode: sitemaps are site-level artifacts, not page snapshots.
  • When the required data tier (live fetch of sitemap/robots, or the resolved link graph) is unavailable, status is needs_api — never a false pass.

Findings

Emit findings per schema/finding.schema.json. Examples:

  • M17.sitemap.missing — no XML sitemap found at /sitemap.xml or in robots.txt (status fail, severity 3, fixable: auto, axis search, confidence established).
  • M17.robots.no_sitemap_line — sitemap exists but no Sitemap: line in robots.txt (status warn, severity 3, fixable: auto, axis search, confidence established).
  • M17.sitemap.noindex_url — a <loc> in the sitemap points to a noindex/non-canonical URL (status fail, severity 3, fixable: proposed, axis search, confidence established).
  • M17.sitemap.error_url — a <loc> the crawl fetched returned 4xx/5xx (status fail, severity 3, fixable: proposed, axis search, confidence established).
  • M17.sitemap.redirected_url — a <loc> redirects instead of returning the final URL (status warn, severity 3, fixable: proposed, axis search, confidence established).
  • M17.sitemap.missing_indexable_url — an indexable URL the crawl found is listed in no sitemap (status warn, severity 3, fixable: proposed, axis search, confidence directional — a well-linked page is discovered without a sitemap).
  • M17.sitemap.lastmod_identical — every <loc> carries the same <lastmod>, so the file cannot distinguish a changed page from an unchanged one (status warn, severity 2, fixable: advisory, axis search, confidence directional). Per-URL verdicts are only emitted for <loc> entries the crawl actually fetched; a URL that was never visited is left unreported rather than guessed. M17.sitemap.missing is also emitted as needs_api when sitemap discovery never ran in the run (--artifacts none / --no-sitemaps), which is a different fact from "no sitemap exists". Platform-conditional ids. This module also emits 8 ids that fire only when profile.json names the platform (shopify, wordpress, nextjs, nuxt, astro, gatsby, hugo). They are indexed in references/routing.md § Platform-conditional finding ids and specified in references/platforms/<id>.md §10.

Each finding: evidence.observed quotes the page/sitemap verbatim; verification.reproduce is the runnable command above; expected_impact is banded + confidence-tagged (no naked %).

Honesty

  • A sitemap is a discovery aid, not a ranking signal or an indexing guarantee — Google treats <lastmod>, <priority>, and <changefreq> as hints, and <priority>/<changefreq> are largely ignored, so don't promise ranking lift from tuning them (label any such tactic low-magnitude/directional).
  • Submitting a sitemap won't force indexing of low-quality or non-canonical pages; orphan and canonical hygiene matters more than sitemap size.

Not using the CLI? Copy the SKILL.md and paste it straight into ChatGPT, Claude, or any agent.

Install into your agent

$ npx skills add https://github.com/hainrixz/claude-seo-ai --skill seo-sitemaps -a claude-code

More in Technical SEO

Technical SEOaddyosmani/web-quality-skills

SEO

Optimize for search engine visibility and ranking.

48.3K installs
Technical SEOaddyosmani/web-quality-skills

Core Web Vitals

Optimize Core Web Vitals (LCP, INP, CLS) for better page experience using field and lab evidence.

28.4K installs
Technical SEOagricidaniel/claude-seo

SEO Sitemap

Analyze existing XML sitemaps or generate new ones with industry templates.

6.4K installs
Technical SEOhainrixz/claude-seo-ai

SEO Crawlability

Audit and generate robots.txt and general crawl access for a page — verify robots.txt reachability and syntax, detect Disallow rules that block CSS/JS or important content, sanity-check crawl-delay, confirm a Sitemap directive, and assert overall crawl access for Googlebot/Bingbot.

174 installs
Technical SEOhainrixz/claude-seo-ai

SEO International

Audit and generate hreflang annotations for multilingual sites — check reciprocity, BCP-47 validity, self-reference, x-default, hreflang/canonical conflicts, and <html lang> agreement, and emit reciprocal hreflang link sets.

173 installs
Technical SEOseoskills.sh

Crawl-Trap & Facet Detector

Detects URL spaces that grow without limit, such as faceted navigation, calendar loops and session IDs, and recommends robots or canonical containment.

New