Skip to content

About

ScrapegraphAI modified for our use case

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

Β 

History

11 Commits

Folders and files

Repository files navigation

Solstice ScrapeGraph

Solstice ScrapeGraph is a fork of ScrapeGraphAI/Scrapegraph-ai with modifications by Solstice Health. It builds on the original library while tailoring behavior for our use cases.

Features

πŸš€ AI-Powered Extraction: Leverage OpenAI, Mistral, and other LLMs for intelligent data extraction
🌐 Multi-format Support: Handle HTML, JSON, XML, CSV, and PDF documents
🎯 Smart Scraping: Intelligent content parsing and structure recognition
πŸ“· Visual Intelligence: Image-to-text conversion and screenshot analysis
πŸ”„ Multi-graph Architecture: Support for complex data processing workflows
⚑ Async Support: High-performance concurrent processing
πŸ›‘οΈ Robot-friendly: Respects robots.txt and rate limiting

Installation

From Git

pip install git+https://github.com/Solstice-Health/solstice_scrapegraph.git

Development Installation

git clone https://github.com/Solstice-Health/solstice_scrapegraph.git
cd solstice_scrapegraph
pip install -e .[dev]

Quick Start

from scrapegraphai.graphs import SmartScraperGraph

# Configure your LLM
llm_config = {
    "llm_model": "gpt-3.5-turbo",
    "api_key": "your_openai_api_key"
}

# Create a scraper instance
scraper = SmartScraperGraph(
    prompt="Extract all product names and prices",
    source="https://example-ecommerce.com",
    config=llm_config
)

# Run the scraper
result = scraper.run()
print(result)

Available Graphs

  • SmartScraperGraph: Basic AI-powered web scraping
  • JSONScraperGraph: Specialized for JSON data extraction
  • XMLScraperGraph: XML document processing
  • CSVScraperGraph: CSV file analysis
  • DocumentScraperGraph: PDF and document processing
  • OmniScraperGraph: Multi-modal scraping with image analysis
  • SearchGraph: Web search and result extraction
  • CodeGeneratorGraph: Generate scraping code automatically

Supported LLM Providers

  • OpenAI (GPT-3.5, GPT-4)
  • Mistral AI
  • AWS Bedrock
  • Ollama (local models)
  • Azure OpenAI

Key Components

Nodes

  • FetchNode: Web page retrieval and loading
  • ParseNode: HTML/content parsing and cleaning
  • GenerateAnswerNode: LLM-powered data extraction
  • ImageToTextNode: Visual content analysis
  • MergeAnswersNode: Result consolidation

Utilities

  • Output Parsers: Structured data extraction
  • Screenshot Tools: Visual page capture
  • Token Management: Cost optimization
  • HTML Cleanup: Content sanitization

Examples

Extract Structured Data

from scrapegraphai.graphs import SmartScraperGraph
from pydantic import BaseModel

class Product(BaseModel):
    name: str
    price: float
    rating: float

scraper = SmartScraperGraph(
    prompt="Extract product information",
    source="https://shop.example.com",
    config={
        "llm_model": "gpt-4",
        "api_key": "your_key",
        "schema": Product
    }
)

products = scraper.run()

Multi-page Scraping

from scrapegraphai.graphs import SmartScraperMultiGraph

urls = [
    "https://news.example.com/page1",
    "https://news.example.com/page2",
    "https://news.example.com/page3"
]

scraper = SmartScraperMultiGraph(
    prompt="Extract article titles and summaries",
    source=urls,
    config=llm_config
)

results = scraper.run()

Document Processing

from scrapegraphai.graphs import DocumentScraperGraph

scraper = DocumentScraperGraph(
    prompt="Extract key financial metrics",
    source="path/to/financial_report.pdf",
    config=llm_config
)

metrics = scraper.run()

Configuration

Basic Configuration

config = {
    "llm_model": "gpt-3.5-turbo",
    "api_key": "your_openai_api_key",
    "verbose": True,
    "headless": True,
    "max_results": 10
}

Advanced Configuration

config = {
    "llm_model": "gpt-4",
    "api_key": "your_key",
    "temperature": 0.1,
    "max_tokens": 2000,
    "rate_limiter": {
        "requests_per_second": 1,
        "burst_size": 5
    },
    "browser_config": {
        "headless": True,
        "timeout": 30000,
        "user_agent": "custom_agent"
    }
}

Development Setup

# Clone the repository
git clone https://github.com/yourusername/solstice_scrapegraph.git
cd solstice_scrapegraph

# Install in development mode
pip install -e .[dev]

# Install pre-commit hooks
pre-commit install

# Run tests
pytest

# Format code
black .

License

This project is licensed under the MIT License - see the LICENSE file for details.

Support

Attribution

This repository is derived from [ScrapeGraphAI/Scrapegraph-ai] (MIT License).
Original license and copyright notices are preserved.
This fork is not an official distribution and is maintained independently by Solstice Health Team.

Acknowledgments

Built with ❀️ by the Solstice Team, powered by:

About

ScrapegraphAI modified for our use case

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages