seoskills.sh
Catalog/SEO Tool Integration/Site Crawl Data Extractor

Site Crawl Data Extractor

Crawls a site within robots.txt and extracts the same fields from every page (SEO tags, JSON-LD values, patterns, tables) into one CSV or JSON table.

Listed ·Updated

New

Use this skill

Free
$ npx skills add https://seoskills.sh --skill site-crawl-data-extractor
Basic pattern scan of SKILL.md text: no matches (checked October 7, 2026)
Repository
seoskills.sh
GitHub stars
0
License
MIT

About this skill

Site Crawl Data Extractor is an SEO Tool Integration skill for AI agents, published in the seoskills.sh catalog. Reach for it when your work involves ahrefs, Semrush, Screaming Frog, Moz, and Search Console workflows. Install it with one command and it runs inside your own agent, so the work happens in your workflow, not a separate SEO tool.

Embed a badge

seoskills.sh listing badge

SKILL.md

Site Crawl Data Extractor

AGENT ROLE: Autonomous extraction agent. Turn the user's question into field definitions, crawl the right pages, emit the JSON in references/output.schema.json plus the full table file, then answer from the data.

OBJECTIVE

Get the same facts out of many pages into one table, reproducibly, without a paid crawler: a content inventory, product prices from JSON-LD, article dates and authors, or any value a pattern can find.

INPUTS

  • start (REQUIRED unless --urls-file, via --start, repeatable): start URLs. The first one's host is the site crawled.
  • sitemap (OPTIONAL flag): also seed from the site's sitemaps (robots.txt first, then /sitemap.xml).
  • urls_file (OPTIONAL): extract exactly the listed URLs, one per line, without following links.
  • include, exclude (OPTIONAL, repeatable): regular expressions on the URL path, such as --include ^/blog/ or --exclude /tag/.
  • field (OPTIONAL, repeatable): name=kind:argument:
    • regex:PATTERN: the first match (group 1 if present) in the page's visible text;
    • rawregex:PATTERN: the same, in the raw HTML;
    • jsonld:Type.path: a JSON-LD property, for example Product.offers.price or *.datePublished;
    • meta:name: a meta tag's content, by name or property;
    • tag:h2, class:price, id:sku: the text of the first matching element.
  • tables (OPTIONAL flag): up to 5 HTML tables per page, up to 50 rows each (in JSON Lines output).
  • max_pages (OPTIONAL, default 200), delay (OPTIONAL, default 0.5 seconds), text_chars (OPTIONAL, default 600), keep_query (OPTIONAL).
  • out (OPTIONAL): a .csv or .jsonl file for every row. The JSON on stdout shows the first 25.

DATA SOURCES (no keys needed)

The site's pages, robots.txt and sitemaps, fetched as seoskills-crawl-extractor/1.0.

EXPECTED TOOL CALLS

  • Run scripts/crawl_extract.py --start https://example.com/blog/ --include ^/blog/ --field "author=jsonld:Article.author.name" --out articles.csv.
  • One request per page, one after another, at least delay apart.

PROCEDURE (deterministic)

STEP 1: DESIGN the fields from the user's question. Prefer jsonld paths and meta tags, which are stable, over regex on text. Test them on two or three pages with --urls-file before a big crawl. STEP 2: CRAWL breadth-first from the start URLs (and sitemap URLs with --sitemap).

  • Stay on the start host, apply the include and exclude patterns, and skip files and query-string URLs (unless --keep-query).
  • Check every URL against its host's robots.txt.
  • Wait delay seconds between requests, or the robots.txt Crawl-delay when that is longer. STEP 3: EXTRACT each page:
  • the built-in fields: title, meta description, H1, canonical, robots (meta and X-Robots-Tag), lang, published and modified dates (JSON-LD or meta), author, words in the main content, schema types, internal and external links, images, and an excerpt;
  • the custom fields and tables. A page that is not HTML or not 200 is kept as a row with its status. STEP 4: REPORT the fill rate of each field (the share of pages where it was found) so a broken pattern shows at once, and write the table. STEP 5: ANSWER the user's question from the table: counts, lists, gaps or comparisons. Quote values as extracted.

RATE LIMITS & ERROR HANDLING

  • One request at a time, at least 0.5 seconds apart (a larger robots.txt Crawl-delay wins), 20-second timeout per page.
  • A bad --field, a regex that does not compile, a relative start URL or an --out that is not .csv or .jsonl STOPs INPUT_INVALID. An unreadable URL file STOPs FILE_UNREADABLE.
  • A crawl that reaches --max-pages says how many URLs were still queued.

MISSING / INSUFFICIENT DATA

  • The crawler does not run JavaScript. IF every page has under 50 words THEN a note says the site probably renders with JavaScript, and the values must come from a rendering tool or the site's API.
  • Fill rates under 100% mean the pattern is missing or different on some pages. Inspect those rows before trusting totals.
  • Respect each site's terms and robots.txt, and never collect personal data the user has no right to process. Keep crawls to what the question needs.

OUTPUT

One JSON object per references/output.schema.json, the full table in --out, and the answer to the user's question with the numbers from the table.

FILES

  • scripts/crawl_extract.py: robots-aware breadth-first crawler, field extractors (regex, JSON-LD path, meta, tag, class, id, tables) and CSV or JSON Lines writer.
  • references/output.schema.json: output contract.

Not using the CLI? Copy the SKILL.md and paste it straight into ChatGPT, Claude, or any agent.

Install into your agent

$ npx skills add https://seoskills.sh --skill site-crawl-data-extractor -a claude-code

More in SEO Tool Integration

SEO Tool Integrationevery-app/open-seo

SEO Project Setup

Populate a project's shared OpenSEO context — site scope, goals, positioning, competitors, key pages, and preferences — plus MCP checks and Search Console intake.

5.6K installs
SEO Tool Integrationagricidaniel/claude-seo

SEO Google

Google SEO APIs: Search Console (Search Analytics, URL Inspection, Sitemaps), PageSpeed Insights v5, CrUX field data with 25-week history, Indexing API v3, and GA4 organic traffic.

5.5K installs
SEO Tool Integrationagricidaniel/claude-seo

SEO Flow

FLOW framework integration: evidence-led SEO using the Find → Leverage → Optimize → Win loop.

4.5K installs
SEO Tool Integrationseoskills.sh

SEO Metrics API Orchestrator

Fetches authority, backlink, volume and difficulty metrics across Ahrefs, Semrush, Moz and DataForSEO in one pass, within a per-vendor credit budget.

New
SEO Tool Integrationkostja94/marketing-skills

Google Search Console

Analyze Google Search Console data, use the GSC API, or interpret search performance.

1.8K installs
SEO Tool Integrationkostja94/marketing-skills

Core Web Vitals

Optimize Core Web Vitals, fix LCP, INP, or CLS issues.

1K installs