---
source_url: "https://github.com/scrapfly/python-scrapfly"
title: "GitHub - scrapfly/python-scrapfly: Official Python SDK for the Scrapfly platform: web scraping, screenshots, AI extraction, crawling, and a remote anti-bot browser. Integrates with Scrapy, LlamaIndex, and LangChain. · GitHub"
mirrored_at: 2026-08-13T03:35:36.733Z
host: github.com
cited_in_42a: true
mirror_canonical: "https://index.42a.ai/github.com/scrapfly/python-scrapfly"
---

> **Original source:** https://github.com/scrapfly/python-scrapfly

## Installation

`pip install scrapfly-sdk`

You can also install extra dependencies

-   `pip install "scrapfly-sdk[seepdup]"` for performance improvement
-   `pip install "scrapfly-sdk[concurrency]"` for concurrency out of the box (asyncio / thread)
-   `pip install "scrapfly-sdk[scrapy]"` for scrapy integration
-   `pip install "scrapfly-sdk[webhook-server]"` for have a native webhook server using flask
-   `pip install "scrapfly-sdk[all]"` Everything!

For use of built-in HTML parser (via `ScrapeApiResponse.selector` property) additional requirement of either [parsel](https://pypi.org/project/parsel/) or [scrapy](https://pypi.org/project/Scrapy/) is required.

For reference of usage or examples, please checkout the folder `/examples` in this repository.

This SDK cover the following Scrapfly API endpoints:

-   [Web Scraping API](https://scrapfly.io/docs/onboarding#web-scraping-api)
-   [Extraction API](https://scrapfly.io/docs/onboarding#extraction-api)
-   [Screenshot API](https://scrapfly.io/docs/onboarding#screenshot-api)

## Integrations

Scrapfly Python SDKs are integrated with [LlamaIndex](https://www.llamaindex.ai/) and [LangChain](https://www.langchain.com/). Both framework allows training Large Language Models (LLMs) using augmented context.

This augmented context is approached by training LLMs on top of private or domain-specific data for common use cases:

-   Question-Answering Chatbots (commonly referred to as RAG systems, which stands for "Retrieval-Augmented Generation")
-   Document Understanding and Extraction
-   Autonomous Agents that can perform research and take actions

In the context of web scraping, web page data can be extracted as Text or Markdown using [Scrapfly's format feature](https://scrapfly.io/docs/scrape-api/specification#api_param_format) to train LLMs with the scraped data.

### LlamaIndex

#### Installation

Install `llama-index`, `llama-index-readers-web`, and `scrapfly-sdk` using pip:

pip install llama-index llama-index-readers-web scrapfly-sdk

#### Usage

Scrapfly is available at LlamaIndex as a [data connector](https://docs.llamaindex.ai/en/stable/module_guides/loading/connector/), known as a `Reader`. This reader is used to gather a web page data into a `Document` representation, which can be used with the LLM directly. Below is an example of building a RAG system using LlamaIndex and scraped data. See the [LlamaIndex use cases](https://docs.llamaindex.ai/en/stable/use_cases/) for more.

import os

from llama\_index.readers.web import ScrapflyReader
from llama\_index.core import VectorStoreIndex

\# Initiate ScrapflyReader with your Scrapfly API key
scrapfly\_reader \= ScrapflyReader(
    api\_key\="Your Scrapfly API key",  \# Get your API key from https://www.scrapfly.io/
    ignore\_scrape\_failures\=True,  \# Ignore unprocessable web pages and log their exceptions
)

\# Load documents from URLs as markdown
documents \= scrapfly\_reader.load\_data(
    urls\=\["https://web-scraping.dev/products"\]
)

\# After creating the documents, train them with an LLM
\# LlamaIndex uses OpenAI default, other options can be found at the examples direcotry: 
\# https://docs.llamaindex.ai/en/stable/examples/llm/openai/

\# Add your OpenAI key (a paid subscription must exist) from: https://platform.openai.com/api-keys/
os.environ\['OPENAI\_API\_KEY'\] \= "Your OpenAI Key"
index \= VectorStoreIndex.from\_documents(documents)
query\_engine \= index.as\_query\_engine()

response \= query\_engine.query("What is the flavor of the dark energy potion?")
print(response)
"The flavor of the dark energy potion is bold cherry cola."

The `load_data` function accepts a ScrapeConfig object to use the desired Scrapfly API parameters:

from llama\_index.readers.web import ScrapflyReader

\# Initiate ScrapflyReader with your ScrapFly API key
scrapfly\_reader \= ScrapflyReader(
    api\_key\="Your Scrapfly API key",  \# Get your API key from https://www.scrapfly.io/
    ignore\_scrape\_failures\=True,  \# Ignore unprocessable web pages and log their exceptions
)

scrapfly\_scrape\_config \= {
    "asp": True,  \# Bypass scraping blocking and antibot solutions, like Cloudflare
    "render\_js": True,  \# Enable JavaScript rendering with a cloud headless browser
    "proxy\_pool": "public\_residential\_pool",  \# Select a proxy pool (datacenter or residnetial)
    "country": "us",  \# Select a proxy location
    "auto\_scroll": True,  \# Auto scroll the page
    "js": "",  \# Execute custom JavaScript code by the headless browser
}

\# Load documents from URLs as markdown
documents \= scrapfly\_reader.load\_data(
    urls\=\["https://web-scraping.dev/products"\],
    scrape\_config\=scrapfly\_scrape\_config,  \# Pass the scrape config
    scrape\_format\="markdown",  \# The scrape result format, either \`markdown\`(default) or \`text\`
)

### LangChain

#### Installation

Install `langchain`, `langchain-community`, and `scrapfly-sdk` using pip:

pip install langchain langchain-community scrapfly-sdk

#### Usage

Scrapfly is available at LangChain as a [document loader](https://python.langchain.com/v0.2/docs/concepts/#document-loaders), known as a `Loader`. This reader is used to gather a web page data into `Document` representation, which canbe used with the LLM after a few operations. Below is an example of building a RAG system with LangChain using scraped data, see [LangChain tutorials](https://python.langchain.com/v0.2/docs/tutorials/) for further use cases.

import os

from langchain import hub \# pip install langchainhub
from langchain\_chroma import Chroma \# pip install langchain\_chroma
from langchain\_core.runnables import RunnablePassthrough
from langchain\_core.output\_parsers import StrOutputParser
from langchain\_openai import OpenAIEmbeddings, ChatOpenAI \# pip install langchain\_openai
from langchain\_text\_splitters import RecursiveCharacterTextSplitter \# pip install langchain\_text\_splitters
from langchain\_community.document\_loaders import ScrapflyLoader

scrapfly\_loader \= ScrapflyLoader(
    \["https://web-scraping.dev/products"\],
    api\_key\="Your Scrapfly API key",  \# Get your API key from https://www.scrapfly.io/
    continue\_on\_failure\=True,  \# Ignore unprocessable web pages and log their exceptions
)

\# Load documents from URLs as markdown
documents \= scrapfly\_loader.load()

\# This example uses OpenAI. For more see: https://python.langchain.com/v0.2/docs/integrations/platforms/
os.environ\["OPENAI\_API\_KEY"\] \= "Your OpenAI key"

\# Create a retriever
text\_splitter \= RecursiveCharacterTextSplitter(chunk\_size\=1000, chunk\_overlap\=200)
splits \= text\_splitter.split\_documents(documents)
vectorstore \= Chroma.from\_documents(documents\=splits, embedding\=OpenAIEmbeddings())
retriever \= vectorstore.as\_retriever()

def format\_docs(docs):
    return "\\n\\n".join(doc.page\_content for doc in docs)

model \= ChatOpenAI()
prompt \= hub.pull("rlm/rag-prompt")

rag\_chain \= (
    {"context": retriever | format\_docs, "question": RunnablePassthrough()}
    | prompt
    | model
    | StrOutputParser()
)

response \= rag\_chain.invoke("What is the flavor of the dark energy potion?")
print(response)
"The flavor of the Dark Energy Potion is bold cherry cola."

To use the full Scrapfly features with LangChain, pass a ScrapeConfig object to the `ScrapflyLoader`:

from langchain\_community.document\_loaders import ScrapflyLoader

scrapfly\_scrape\_config \= {
    "asp": True,  \# Bypass scraping blocking and antibot solutions, like Cloudflare
    "render\_js": True,  \# Enable JavaScript rendering with a cloud headless browser
    "proxy\_pool": "public\_residential\_pool",  \# Select a proxy pool (datacenter or residnetial)
    "country": "us",  \# Select a proxy location
    "auto\_scroll": True,  \# Auto scroll the page
    "js": "",  \# Execute custom JavaScript code by the headless browser
}

scrapfly\_loader \= ScrapflyLoader(
    \["https://web-scraping.dev/products"\],
    api\_key\="Your Scrapfly API key",  \# Get your API key from https://www.scrapfly.io/
    continue\_on\_failure\=True,  \# Ignore unprocessable web pages and log their exceptions
    scrape\_config\=scrapfly\_scrape\_config,  \# Pass the scrape\_config object
    scrape\_format\="markdown",  \# The scrape result format, either \`markdown\`(default) or \`text\`
)

\# Load documents from URLs as markdown
documents \= scrapfly\_loader.load()
print(documents)

## Get Your API Key

You can create a free account on [Scrapfly](https://scrapfly.io/register) to get your API Key.

-   [Usage](https://scrapfly.io/docs/sdk/python)
-   [Python API](https://scrapfly.github.io/python-scrapfly/scrapfly)
-   [Open API 3 Spec](https://scrapfly.io/docs/openapi#get-/scrape)
-   [Scrapy Integration](https://scrapfly.io/docs/sdk/scrapy)

## Migration

### Migrate from 0.7.x to 0.8

asyncio-pool dependency has been dropped

`scrapfly.concurrent_scrape` is now an async generator. If the concurrency is `None` or not defined, the max concurrency allowed by your current subscription is used.

    async for result in scrapfly.concurrent\_scrape(concurrency\=10, scrape\_configs\=\[ScrapConfig(...), ...\]):
        print(result)

brotli args is deprecated and will be removed in the next minor. There is not benefit in most of case versus gzip regarding and size and use more CPU.

### What's new

### 0.8.x

-   Better error log
-   Async/Improvement for concurrent scrape with asyncio
-   Scrapy media pipeline are now supported out of the box