Skills

ScrapeGraphAI: Build Lead Lists From Public Directories With Claude

7 minute readUpdated October 2026Explore more

TL;DR

ScrapeGraphAI is a free, MIT licensed Python library with over 31,000 stars on GitHub. You give it a web page and a plain-English request, and an AI model reads the page and hands back structured data. It runs on your own computer, and it can use Claude through your API key or a fully local model through Ollama. Below is the setup, a ready script that turns public directory pages into a lead CSV, a Claude Code prompt that builds it for you, and the rules to follow before you email anyone.

Learn Claude Code. Earn income. Only $9.

👉 https://www.skool.com/claudecodeclub

What ScrapeGraphAI is

ScrapeGraphAI/Scrapegraph-ai

31,514 stars on October 3, 2026. MIT license. Latest release v2.3.0 (September 25, 2026).

ScrapeGraphAI describes itself as a web scraping Python library that uses an LLM and graph logic to build scraping pipelines. Its pitch is simple: say which information you want and the library extracts it. You do not write CSS selectors or XPath. You write a sentence like "list every business name, phone and website on this page" and get structured data back.

  • Runs on your own computer. The open-source library runs on your own machine. Your bill is whatever the AI model costs you, and nothing else.
  • Works with Claude. Anthropic is one of its supported model providers, so Claude can be the brain that reads each page.
  • Or fully local. It also supports local models through Ollama, so a run can stay entirely on your machine with no API bill.
  • Several pipelines. SmartScraperGraph reads one page. SmartScraperMultiGraph reads a list of pages. SearchGraph pulls from the top search results. There are also script-generator and audio pipelines.

What you need

  • Python 3.12 or newer (the current release requires it).
  • An Anthropic API key if you want Claude to read the pages. Or Ollama installed if you want a free local model instead.
  • The URLs of the public directory pages you want to read. One URL per results page.
  • Claude Code, if you want it to build and run everything for you (the fastest route).

Fastest route: let Claude Code build it

Open Claude Code in an empty folder and paste this. Swap in your directory, your city and your fields.

promptRead the README at https://github.com/ScrapeGraphAI/Scrapegraph-ai and set up a small lead-list project in this folder.

1. Create a Python 3.12+ virtual environment and install scrapegraphai, langchain-anthropic, and the Playwright browser the README asks for. Tell me each command before you run it.
2. Write leads.py that uses SmartScraperGraph with Claude (model anthropic/claude-haiku-4-5, my key from the ANTHROPIC_API_KEY environment variable) and a Pydantic schema with these fields: name, category, website, phone, public_email, city.
3. Loop over these public directory pages:
   [PASTE PAGE URLS HERE, ONE PER LINE]
4. Only extract business contact details that are printed on the page. No guessing, no personal profiles, no pages behind a login.
5. Remove duplicates by name and website, add a source_url column, and save everything to leads.csv.
6. Before running on all pages, check the site's terms of service and robots.txt and tell me if anything forbids automated collection. Then test on ONE page and show me the first 5 rows.

Claude Code reads the README, installs what it needs, writes the script, and tests it on one page before you run the full list. If the site's rules say no, stop there and pick a different source.

Or do it by hand: install

These are the README's own install steps, plus langchain-anthropic, which LangChain requires when the model provider is Anthropic. Run them inside a virtual environment.

bashpython3 -m venv .venv
source .venv/bin/activate
pip install scrapegraphai langchain-anthropic
playwright install
export ANTHROPIC_API_KEY=sk-ant-your-key-here

The lead-list script

Save this as leads.py, put your page URLs in PAGES, and run python leads.py. It reads each page with Claude, keeps only the fields you asked for, removes duplicates, and writes a clean leads.csv you can open in Google Sheets or Excel.

python# leads.py: turn public directory pages into a clean CSV of business contacts.
# Run inside a virtual environment where scrapegraphai is installed.
import csv
import os
from typing import List, Optional

from pydantic import BaseModel
from scrapegraphai.graphs import SmartScraperGraph

# 1) The public directory pages you want to read (one URL per results page).
PAGES = [
    "https://example-directory.com/plumbers/austin?page=1",
    "https://example-directory.com/plumbers/austin?page=2",
]

# 2) What one row in your sheet looks like.
class Business(BaseModel):
    name: str
    category: Optional[str] = None
    website: Optional[str] = None
    phone: Optional[str] = None
    public_email: Optional[str] = None
    city: Optional[str] = None

class Listing(BaseModel):
    businesses: List[Business]

PROMPT = (
    "List every business shown on this page. For each one return the business name, "
    "category, website, main business phone number, city, and a general business email "
    "only if it is printed on the page. Only use details that are publicly shown on this "
    "page. Do not guess, infer or invent anything. Leave a field empty if it is not shown. "
    "Skip private individuals and personal social media profiles."
)

# 3) The model. This uses Claude through your Anthropic API key.
#    For a fully local run, swap in the Ollama config from the ScrapeGraphAI README.
graph_config = {
    "llm": {
        "model": "anthropic/claude-haiku-4-5",
        "api_key": os.environ["ANTHROPIC_API_KEY"],
        "model_tokens": 200000,
    },
    "headless": True,
    "verbose": False,
}

rows, seen = [], set()
for url in PAGES:
    result = SmartScraperGraph(prompt=PROMPT, source=url, config=graph_config, schema=Listing).run()
    for b in (result or {}).get("businesses", []):
        key = ((b.get("name") or "").strip().lower(), (b.get("website") or "").strip().lower())
        if not key[0] or key in seen:
            continue
        seen.add(key)
        b["source_url"] = url
        rows.append(b)

fields = ["name", "category", "website", "phone", "public_email", "city", "source_url"]
with open("leads.csv", "w", newline="") as fh:
    writer = csv.DictWriter(fh, fieldnames=fields, extrasaction="ignore")
    writer.writeheader()
    writer.writerows(rows)

print(f"Saved {len(rows)} businesses to leads.csv")

Run it fully local with Ollama

Want zero API cost and nothing leaving your machine? Install Ollama, pull a model, and replace the llm block with the one from the README. Local models are slower and less accurate than Claude on messy pages, so spot-check the output.

python# ollama pull llama3.2   (run this in your terminal first)
graph_config = {
    "llm": {
        "model": "ollama/llama3.2",
        "model_tokens": 8192,
        "format": "json",
    },
    "headless": True,
    "verbose": False,
}

Prompts that get cleaner rows

Swap the PROMPT in the script for one of these when you need something more specific.

promptList every business on this page that is a dental clinic. For each one return the clinic name, street address, city, main phone number, website and booking page link if shown. Only use details printed on this page. Leave a field empty if it is not shown. Do not invent anything.
promptThis is a member directory of a local chamber of commerce. For each member business return the business name, industry, website and the general business phone. Ignore named individuals' personal phone numbers and personal emails. Only use details printed on this page.
promptI have a CSV called leads.csv. Clean it: fix capitalization in names, format every phone number the same way, remove rows with no website and no phone, flag any row where the email domain does not match the website domain, and save the result as leads_clean.csv. Show me what you changed.

The rules to follow before you reach out

  • Read the site's terms of service and robots.txt. Many directories forbid automated collection. If they do, do not scrape them.
  • Collect public business contacts only. A company's listed phone, website and general email. Not private individuals, and nothing behind a login.
  • Go slow. Small batches, not thousands of requests at once. You are a guest on someone else's server.
  • Email laws still apply. In the US, the FTC's CAN-SPAM guide says commercial email must not use misleading headers or subject lines, must tell people where you are located, must tell them how to opt out, and must honor opt-outs promptly.
  • Contacting people in the EU or UK? Named contacts count as personal data under GDPR, even at work. Get proper advice before you email them.
  • Telemetry is on by default. The library collects anonymous usage metrics. Set SCRAPEGRAPHAI_TELEMETRY_ENABLED=false to opt out.
CAN-SPAM Act: A Compliance Guide for Business (FTC)

The official US rules for commercial email, in plain language.

Quick start

  1. 1Pick one public directory and check its terms and robots.txt.
  2. 2Paste the Claude Code prompt above, or install by hand and save leads.py.
  3. 3Test on one page and read the first rows yourself.
  4. 4Run the rest, clean the CSV, and follow the email rules before you send anything.

Now you just build lead lists instead of buying them. Inside the Claude Code Club we share the scrapers, prompts and outreach setups we actually run, and help each other get them working. It's nine dollars a month at https://www.skool.com/claudecodeclub/about. Everything on this page works without it.

Common questions

  • Is ScrapeGraphAI free?

    Yes. The open-source library is free and MIT licensed. You only pay for the AI model you plug in, like Claude through your Anthropic API key, or nothing at all with a local model through Ollama. The same team also sells a separate paid cloud API, which you do not need for this guide.

  • Does it really run on my own computer?

    Yes. The open-source library runs on your own machine. With Claude, the page text is sent to Anthropic to be read. With a local Ollama model, the whole run stays on your computer.

  • Does ScrapeGraphAI work with Claude?

    Yes. Anthropic is in its list of supported model providers. Set the model to anthropic/ plus a Claude model name, add model_tokens, and install langchain-anthropic.

  • Is scraping business contacts legal?

    It depends on the site and where you are. Check the site's terms and robots.txt, only collect public business contact details, and follow email laws like CAN-SPAM in the US and GDPR in the EU and UK. When in doubt, get proper advice. This page is not legal advice.

  • Can it export straight to Google Sheets?

    The library returns structured data, not a spreadsheet. The script on this page writes it to a CSV file, which you can import into Google Sheets or open in Excel.

Want the scrapers and outreach setups we actually run?

Get 650+ plug-and-play skills, MCPs & prompts, plus 8,000+ members - $9/mo, cancel anytime.

Join the Club