Skip to content

Repository files navigation

Deep Learning & Computer Vision Projects

A curated collection of production-style deep learning projects in PyTorch and TensorFlow/Keras — covering computer vision, generative models, adversarial learning, NLP, and time-series analysis. Every project is a self-contained, modular package with a CLI, configuration, tests where applicable, and its own documentation.

About

A collection of production-ready deep learning projects in PyTorch and TensorFlow — vision transformers, GANs, adversarial attacks, face recognition, image captioning, pose estimation, semantic segmentation, NLP topic modeling, and video analysis. Each project ships as a clean, modular package with CLI entry points, reproducible training, and full documentation.

GitHub Topics

Paste into the repo About → Topics field:

deep-learning, machine-learning, computer-vision, pytorch, tensorflow,
keras, vision-transformer, transformers, gan, generative-models,
adversarial-attacks, image-captioning, pose-estimation, image-segmentation,
face-recognition, object-detection, yolo, nlp, topic-modeling,
sentiment-analysis, time-series, jupyter-notebook, cnn, lstm, opencv

Projects

Every folder is an independent project with its own README.md, code, and requirements. The original Jupyter notebooks are kept untouched as the research record.

Computer Vision

Project What it does Stack
vision-transformer From-scratch ViT & CCT vision transformers trained on CIFAR-10/100 PyTorch
video-YOLO Driving-video analysis: lane detection + YOLOv5 vehicle detection PyTorch, OpenCV
pose estimation Human pose estimation via heatmap regression (LSP) PyTorch
U-NET Binary image segmentation with a pretrained U-Net TensorFlow/Keras
image-color-clustring K-Means color quantization and palette extraction OpenCV, scikit-learn
STN-and-transferLearning Spatial transformer networks + transfer learning (ResNet, DeiT) PyTorch, timm

Generative & Adversarial

Project What it does Stack
GAN GAN & DCGAN trained on Fashion-MNIST / CelebA TensorFlow/Keras + PyTorch
adversarial FGSM & DeepFool adversarial attacks with accuracy analysis PyTorch

Faces & Biometrics

Project What it does Stack
siameseNet Face recognition with a Siamese network + triplet loss (LFW) PyTorch
FER Facial expression recognition across 7 emotions (FER2013) PyTorch

Vision + Language

Project What it does Stack
image_captioning CNN-LSTM image captioning (Inception-v3 + LSTM, Flickr8k) PyTorch

NLP

Project What it does Stack
topic-modeling (NLP) Unsupervised topic discovery in Persian news (embeddings + DBSCAN/community detection) sentence-transformers, scikit-learn
_ML/sentiment_analysis_IMDB IMDB sentiment analysis with TF-IDF + classic classifiers scikit-learn

Time Series

Project What it does Stack
missing-value (time series) Missing-value imputation for the monthly-sunspots series (Conv1D, GRU, BiGRU) TensorFlow/Keras

Video Understanding

Project What it does Stack
hockey-videos Hockey-fight detection with ResNet50 features + classifier TensorFlow/Keras

Training Toolkit

Project What it does Stack
call back Keras callbacks playbook: early stopping, checkpoints, LR scheduling on CIFAR-100 / Fashion-MNIST TensorFlow/Keras
_BASICS PyTorch fundamentals: tensors, CNN, RNN/LSTM, transfer learning, augmentation, custom datasets PyTorch

Repository Layout

.
├── adversarial/                  # FGSM + DeepFool attacks
├── _BASICS/                      # PyTorch fundamentals
├── call back/                    # Keras callback patterns
├── FER/                          # Facial expression recognition
├── GAN/                          # GAN & DCGAN
├── hockey-videos/                # Fight detection in sports video
├── image_captioning/             # CNN-LSTM captioning
├── image-color-clustring/        # Color quantization
├── missing-value(time series)/   # Time-series imputation
├── _ML/sentiment_analysis_IMDB/  # IMDB sentiment analysis
├── pose estimation/              # Heatmap pose estimation
├── siameseNet/                   # Face recognition
├── STN-and-transferLearning/     # Spatial transformers + transfer learning
├── topic-modeling(NLP)/          # Persian news topic clustering
├── U-NET/                        # Binary segmentation
├── video-YOLO/                   # YOLOv5 driving-video analysis
└── vision-transformer/           # ViT / CCT from scratch

Tech Stack

  • Languages: Python 3
  • Deep learning: PyTorch, TensorFlow/Keras, timm
  • Classical ML: scikit-learn, sentence-transformers
  • Vision: OpenCV, albumentations
  • Data & viz: NumPy, Pandas, Matplotlib, Seaborn
  • Ops: Weights & Biases, TensorBoard, Jupyter

Getting Started

Each project is independent. To run one, install its dependencies and follow its README:

cd vision-transformer
pip install -r requirements.txt
python train.py --dataset cifar10 --model cct

Repository-wide environment

pip install torch torchvision tensorflow numpy pandas scikit-learn opencv-python \
            matplotlib seaborn einops timm sentence-transformers wandb

Several projects require external datasets (LFW, Flickr8k, CelebA, FER2013, LSP) — each README documents where to obtain them. Pretrained weights that ship in the repo (e.g. U-NET/binary_segmentation.h5, video-YOLO/yolov5s.pt) are used automatically.

Conventions

  • Every project follows a consistent module layout: config.py, model.py, dataset.py, train.py, evaluate.py, utils.py, requirements.txt, and a README.md.
  • All hyperparameters live in a single config.py.
  • Training/evaluation run through argparse CLI entry points.
  • The original .ipynb files are preserved unmodified as the research record.

License

MIT

About

A curated collection of production-style deep learning projects in PyTorch and TensorFlow — vision transformers, GANs, adversarial attacks, face recognition, image captioning, pose estimation, semantic segmentation, NLP topic modeling, and video analysis. Each project ships as a modular package with CLI entry points, reproducible training.

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages