A Python-based mini search engine that crawls real web pages using Scrapy, builds a searchable dataset, and returns relevant results based on user queries.
This project is being built from the ground up to understand how search engines work behind the scenes.
The current version uses Scrapy to crawl web pages and collect information such as:
- Quotes
- Authors
- Tags
- Page URLs
The scraped data is stored in a JSON dataset and searched using Python.
Web Pages
│
▼
┌──────────────┐
│ Scrapy │
│ Crawler │
└──────┬───────┘
│
▼
┌──────────────┐
│ dataset.json │
│ 100 records │
└──────┬───────┘
│
▼
┌──────────────┐
│ Python │
│ Search Engine│
└──────┬───────┘
│
▼
Search Results
- Python 3
- Scrapy
- JSON
- Git & GitHub
python-search-engine/
│
├── dataset.json
├── search.py
├── scrapy.cfg
├── README.md
├── .gitignore
│
└── search_engine/
├── __init__.py
├── items.py
├── middlewares.py
├── pipelines.py
├── settings.py
│
└── spiders/
├── __init__.py
└── webcrawler.py
The crawler collects data from quotes.toscrape.com.
Each record contains:
{
"text": "Quote text",
"author": "Author name",
"tags": ["tag1", "tag2"],
"url": "Page URL"
}The crawler currently collects 100 records across multiple pages.
The Python search engine allows users to search the dataset by:
- Quote text
- Author
- Tags
Example:
Search: love
The engine returns matching quotes, authors, tags, and URLs.
This project will be expanded beyond basic keyword matching.
Planned improvements include:
- TF-IDF ranking
- Better text preprocessing
- Stop-word removal
- Stemming and tokenization
- BM25 ranking
- Larger web datasets
- SQLite/PostgreSQL database
- Semantic/vector search
- Search suggestions
- Web-based search interface
- Search result scoring
- Multi-website crawling
The long-term goal is to build a small but functional search engine that demonstrates the core concepts behind modern search systems:
Crawling → Indexing → Ranking → Retrieval
Jesse Adeyemi
Built as a Python learning and portfolio project.