Solstice ScrapeGraph is a fork of ScrapeGraphAI/Scrapegraph-ai with modifications by Solstice Health. It builds on the original library while tailoring behavior for our use cases.
π AI-Powered Extraction: Leverage OpenAI, Mistral, and other LLMs for intelligent data extraction
π Multi-format Support: Handle HTML, JSON, XML, CSV, and PDF documents
π― Smart Scraping: Intelligent content parsing and structure recognition
π· Visual Intelligence: Image-to-text conversion and screenshot analysis
π Multi-graph Architecture: Support for complex data processing workflows
β‘ Async Support: High-performance concurrent processing
π‘οΈ Robot-friendly: Respects robots.txt and rate limiting
pip install git+https://github.com/Solstice-Health/solstice_scrapegraph.git
git clone https://github.com/Solstice-Health/solstice_scrapegraph.git
cd solstice_scrapegraph
pip install -e .[dev]from scrapegraphai.graphs import SmartScraperGraph
# Configure your LLM
llm_config = {
"llm_model": "gpt-3.5-turbo",
"api_key": "your_openai_api_key"
}
# Create a scraper instance
scraper = SmartScraperGraph(
prompt="Extract all product names and prices",
source="https://example-ecommerce.com",
config=llm_config
)
# Run the scraper
result = scraper.run()
print(result)- SmartScraperGraph: Basic AI-powered web scraping
- JSONScraperGraph: Specialized for JSON data extraction
- XMLScraperGraph: XML document processing
- CSVScraperGraph: CSV file analysis
- DocumentScraperGraph: PDF and document processing
- OmniScraperGraph: Multi-modal scraping with image analysis
- SearchGraph: Web search and result extraction
- CodeGeneratorGraph: Generate scraping code automatically
- OpenAI (GPT-3.5, GPT-4)
- Mistral AI
- AWS Bedrock
- Ollama (local models)
- Azure OpenAI
- FetchNode: Web page retrieval and loading
- ParseNode: HTML/content parsing and cleaning
- GenerateAnswerNode: LLM-powered data extraction
- ImageToTextNode: Visual content analysis
- MergeAnswersNode: Result consolidation
- Output Parsers: Structured data extraction
- Screenshot Tools: Visual page capture
- Token Management: Cost optimization
- HTML Cleanup: Content sanitization
from scrapegraphai.graphs import SmartScraperGraph
from pydantic import BaseModel
class Product(BaseModel):
name: str
price: float
rating: float
scraper = SmartScraperGraph(
prompt="Extract product information",
source="https://shop.example.com",
config={
"llm_model": "gpt-4",
"api_key": "your_key",
"schema": Product
}
)
products = scraper.run()from scrapegraphai.graphs import SmartScraperMultiGraph
urls = [
"https://news.example.com/page1",
"https://news.example.com/page2",
"https://news.example.com/page3"
]
scraper = SmartScraperMultiGraph(
prompt="Extract article titles and summaries",
source=urls,
config=llm_config
)
results = scraper.run()from scrapegraphai.graphs import DocumentScraperGraph
scraper = DocumentScraperGraph(
prompt="Extract key financial metrics",
source="path/to/financial_report.pdf",
config=llm_config
)
metrics = scraper.run()config = {
"llm_model": "gpt-3.5-turbo",
"api_key": "your_openai_api_key",
"verbose": True,
"headless": True,
"max_results": 10
}config = {
"llm_model": "gpt-4",
"api_key": "your_key",
"temperature": 0.1,
"max_tokens": 2000,
"rate_limiter": {
"requests_per_second": 1,
"burst_size": 5
},
"browser_config": {
"headless": True,
"timeout": 30000,
"user_agent": "custom_agent"
}
}# Clone the repository
git clone https://github.com/yourusername/solstice_scrapegraph.git
cd solstice_scrapegraph
# Install in development mode
pip install -e .[dev]
# Install pre-commit hooks
pre-commit install
# Run tests
pytest
# Format code
black .This project is licensed under the MIT License - see the LICENSE file for details.
- π§ Email: contact@solsticehealth.co
- π¬ Issues: GitHub Issues
- π Documentation: [Coming Soon]
This repository is derived from [ScrapeGraphAI/Scrapegraph-ai] (MIT License).
Original license and copyright notices are preserved.
This fork is not an official distribution and is maintained independently by Solstice Health Team.
Built with β€οΈ by the Solstice Team, powered by:
- LangChain for LLM integration
- Playwright for browser automation
- Pydantic for data validation