This project aims to develop and test different lip reading algorithms on words and on sentences, using the GRID Corpus Dataset.
-
Updated
Sep 2, 2021 - Python
This project aims to develop and test different lip reading algorithms on words and on sentences, using the GRID Corpus Dataset.
A Pytorch implementation of the MixMatch algorithm developed by google-research.
Few-shot product image classification using Prototypical Networks (5-10 reference images per class, no training loop) — pretrained MobileNetV3 embeddings + FastAPI, ~40ms inference latency.
This is an image generation tool that implements the generating of images by Diffusers and Kandinsky 2.2 solutions.
🚀 AI-powered end-to-end system for detecting potholes using YOLOv8 and analyzing real-world road conditions from images. This project goes beyond basic object detection by building a practical, production-oriented pipeline that identifies potholes in diverse environments, handles domain shifts, and prepares for real-world deployment.
One tiny INT8 model for edge signals — label-free anomaly detection + neural compression, with a bit-exact C runtime.
An end-to-end NLP pipeline for automated Requirements Engineering. It utilizes a DistilBERT classifier to detect requirements from raw text and a custom transformer-based NER model to extract key entities like Actors, Features, and Quality Attributes.
img2tensor is a high-performance utility to convert images into training-ready tensors for NumPy, PyTorch, and TensorFlow, or stream them directly into TFRecords.
Diabetes risk prediction using XGBoost, Random Forest, Logistic Regression, and a PyTorch 1D CNN with calibration and subgroup evaluation.
This project applies transfer learning using a pretrained AlexNet model to classify FashionMNIST images. The model was fine-tuned and trained on GPU after necessary preprocessing. It achieved 93.16% accuracy on the test set and 95.87% on the training set.
Bidirectional translation toolkit between signed and spoken language — sign video to text/audio and text/audio to synthesized sign video, with modular ML pipeline and REST API.
End-to-end Computer Vision projects covering detection, tracking, segmentation, generation, depth estimation, and 3D face morphing — built with YOLO, MediaPipe, SegFormer, and Diffusion Models.
CNN achieving 99.1% test accuracy — deployed as live Streamlit web app with real-time digit recognition
Flexible and powerful image analysis
introduction to neural networks
Transform words across languages and generate stunning images from text - your all-in-one AI toolkit for multilingual translation and visual creation.
AI Yoga Trainer
To associate your repository with the pytor topic, visit your repo's landing page and select "manage topics."