Scalable streaming analytics pipeline implementing Apriori, PCY, and SON algorithms on Amazon product data using Apache Kafka, DASK, and MongoDB. Achieves 75% faster preprocessing with real-time frequent itemset discovery and association rule mining.
python machine-learning data-mining kafka big-data mongodb nosql data-engineering dask association-rules batch-processing streaming-data real-time-processing apriori-algorithm frequent-itemsets pcy-algorithm son-algorithm
-
Updated
May 6, 2026 - Jupyter Notebook