SEO Project Setup
Populate a project's shared OpenSEO context — site scope, goals, positioning, competitors, key pages, and preferences — plus MCP checks and Search Console intake.
Crawls a site within robots.txt and extracts the same fields from every page (SEO tags, JSON-LD values, patterns, tables) into one CSV or JSON table.
Listed ·Updated
$ npx skills add https://seoskills.sh --skill site-crawl-data-extractorSite Crawl Data Extractor is an SEO Tool Integration skill for AI agents, published in the seoskills.sh catalog. Reach for it when your work involves ahrefs, Semrush, Screaming Frog, Moz, and Search Console workflows. Install it with one command and it runs inside your own agent, so the work happens in your workflow, not a separate SEO tool.
AGENT ROLE: Autonomous extraction agent. Turn the user's question into field definitions, crawl the right pages, emit the JSON in references/output.schema.json plus the full table file, then answer from the data.
Get the same facts out of many pages into one table, reproducibly, without a paid crawler: a content inventory, product prices from JSON-LD, article dates and authors, or any value a pattern can find.
start (REQUIRED unless --urls-file, via --start, repeatable): start URLs. The first one's host is the site crawled.sitemap (OPTIONAL flag): also seed from the site's sitemaps (robots.txt first, then /sitemap.xml).urls_file (OPTIONAL): extract exactly the listed URLs, one per line, without following links.include, exclude (OPTIONAL, repeatable): regular expressions on the URL path, such as --include ^/blog/ or --exclude /tag/.field (OPTIONAL, repeatable): name=kind:argument:
regex:PATTERN: the first match (group 1 if present) in the page's visible text;rawregex:PATTERN: the same, in the raw HTML;jsonld:Type.path: a JSON-LD property, for example Product.offers.price or *.datePublished;meta:name: a meta tag's content, by name or property;tag:h2, class:price, id:sku: the text of the first matching element.tables (OPTIONAL flag): up to 5 HTML tables per page, up to 50 rows each (in JSON Lines output).max_pages (OPTIONAL, default 200), delay (OPTIONAL, default 0.5 seconds), text_chars (OPTIONAL, default 600), keep_query (OPTIONAL).out (OPTIONAL): a .csv or .jsonl file for every row. The JSON on stdout shows the first 25.The site's pages, robots.txt and sitemaps, fetched as seoskills-crawl-extractor/1.0.
scripts/crawl_extract.py --start https://example.com/blog/ --include ^/blog/ --field "author=jsonld:Article.author.name" --out articles.csv.delay apart.STEP 1: DESIGN the fields from the user's question. Prefer jsonld paths and meta tags, which are stable, over regex on text. Test them on two or three pages with --urls-file before a big crawl.
STEP 2: CRAWL breadth-first from the start URLs (and sitemap URLs with --sitemap).
--keep-query).delay seconds between requests, or the robots.txt Crawl-delay when that is longer.
STEP 3: EXTRACT each page:X-Robots-Tag), lang, published and modified dates (JSON-LD or meta), author, words in the main content, schema types, internal and external links, images, and an excerpt;Crawl-delay wins), 20-second timeout per page.--field, a regex that does not compile, a relative start URL or an --out that is not .csv or .jsonl STOPs INPUT_INVALID. An unreadable URL file STOPs FILE_UNREADABLE.--max-pages says how many URLs were still queued.One JSON object per references/output.schema.json, the full table in --out, and the answer to the user's question with the numbers from the table.
scripts/crawl_extract.py: robots-aware breadth-first crawler, field extractors (regex, JSON-LD path, meta, tag, class, id, tables) and CSV or JSON Lines writer.references/output.schema.json: output contract.Not using the CLI? Copy the SKILL.md and paste it straight into ChatGPT, Claude, or any agent.
$ npx skills add https://seoskills.sh --skill site-crawl-data-extractor -a claude-codePopulate a project's shared OpenSEO context — site scope, goals, positioning, competitors, key pages, and preferences — plus MCP checks and Search Console intake.
Google SEO APIs: Search Console (Search Analytics, URL Inspection, Sitemaps), PageSpeed Insights v5, CrUX field data with 25-week history, Indexing API v3, and GA4 organic traffic.
FLOW framework integration: evidence-led SEO using the Find → Leverage → Optimize → Win loop.
Fetches authority, backlink, volume and difficulty metrics across Ahrefs, Semrush, Moz and DataForSEO in one pass, within a per-vendor credit budget.
Analyze Google Search Console data, use the GSC API, or interpret search performance.
Optimize Core Web Vitals, fix LCP, INP, or CLS issues.