Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🔎 Python Web Search Engine

A Python-based mini search engine that crawls real web pages using Scrapy, builds a searchable dataset, and returns relevant results based on user queries.

🚀 Project Overview

This project is being built from the ground up to understand how search engines work behind the scenes.

The current version uses Scrapy to crawl web pages and collect information such as:

  • Quotes
  • Authors
  • Tags
  • Page URLs

The scraped data is stored in a JSON dataset and searched using Python.

🏗️ Architecture

Web Pages
    │
    ▼
┌──────────────┐
│    Scrapy    │
│    Crawler   │
└──────┬───────┘
       │
       ▼
┌──────────────┐
│ dataset.json │
│ 100 records  │
└──────┬───────┘
       │
       ▼
┌──────────────┐
│    Python    │
│ Search Engine│
└──────┬───────┘
       │
       ▼
 Search Results

🛠️ Technologies

  • Python 3
  • Scrapy
  • JSON
  • Git & GitHub

📂 Project Structure

python-search-engine/
│
├── dataset.json
├── search.py
├── scrapy.cfg
├── README.md
├── .gitignore
│
└── search_engine/
    ├── __init__.py
    ├── items.py
    ├── middlewares.py
    ├── pipelines.py
    ├── settings.py
    │
    └── spiders/
        ├── __init__.py
        └── webcrawler.py

🕷️ Web Crawler

The crawler collects data from quotes.toscrape.com.

Each record contains:

{
    "text": "Quote text",
    "author": "Author name",
    "tags": ["tag1", "tag2"],
    "url": "Page URL"
}

The crawler currently collects 100 records across multiple pages.

🔍 Search

The Python search engine allows users to search the dataset by:

  • Quote text
  • Author
  • Tags

Example:

Search: love

The engine returns matching quotes, authors, tags, and URLs.

📈 Future Improvements

This project will be expanded beyond basic keyword matching.

Planned improvements include:

  • TF-IDF ranking
  • Better text preprocessing
  • Stop-word removal
  • Stemming and tokenization
  • BM25 ranking
  • Larger web datasets
  • SQLite/PostgreSQL database
  • Semantic/vector search
  • Search suggestions
  • Web-based search interface
  • Search result scoring
  • Multi-website crawling

🎯 Goal

The long-term goal is to build a small but functional search engine that demonstrates the core concepts behind modern search systems:

Crawling → Indexing → Ranking → Retrieval

👨‍💻 Author

Jesse Adeyemi

Built as a Python learning and portfolio project.

About

A Python-based search engine built with Scrapy that crawls web pages, creates a searchable dataset, and ranks relevant results.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages