I’m Konstantin Nikiforov — MD, Molecular Geneticist, and Machine Learning Engineer.
I build end-to-end ML systems: data acquisition → feature engineering → modeling → evaluation → APIs → user-facing applications → reproducible pipelines.
Focus: applied ML, information retrieval & semantic search, recommendation systems, ML platforms, reproducible experimentation, and model evaluation.
Interests: R&D, MLOps, biomedical ML, NLP, CV, recommender systems, interpretable ML, time series, and modern deep learning.
ML / Data Science: Python · SQL · Pandas · NumPy · scikit-learn · CatBoost · XGBoost · LightGBM · Optuna
Retrieval / NLP: Sentence Transformers · SBERT · embeddings · semantic search · hybrid lexical+dense retrieval · Qdrant · Hugging Face
Deep Learning: PyTorch · Hugging Face Transformers
Recommender Systems: ALS / implicit feedback · content-based retrieval · hybrid recommenders · learning-to-rank
MLOps / Backend: FastAPI · Docker / Docker Compose · Airflow · MLflow · experiment tracking · validation pipelines
Data / Storage: PostgreSQL · MySQL · PySpark · Parquet · vector databases
Applications / Visualization: Streamlit · Plotly · Dash · Matplotlib · SHAP · LIME
Time Series: Prophet · TBATS · ETNA · AutoTS · rolling / holdout backtesting
Domains: Tabular ML · Information Retrieval · Recommenders · NLP · Time Series · BioML · Computer Vision
Exploring: RAG · LangChain / LangGraph · GNNs · scientific embeddings & cross-encoders · ESM / protein language models · self-supervised learning · Diffusion · Ray / Dask · Kafka · Kubernetes · observability (Prometheus / Grafana / OpenTelemetry)
Rust · C · C++ · Java
Research intelligence platform for discovering, organizing, and analyzing Machine Learning and AI research.
Built an end-to-end multi-source data pipeline integrating arXiv, OpenAlex, Crossref, Semantic Scholar, and ACL Anthology into a reconciled corpus of ~61K canonical papers.
Implemented:
- metadata ingestion, normalization, reconciliation, and provenance tracking;
- lexical, dense, and hybrid semantic retrieval;
- Sentence Transformer embeddings and Qdrant vector search experiments;
- topic clustering and similar-paper retrieval;
- citation/reference and paper–artifact graph analytics;
- PostgreSQL-backed serving layer;
- FastAPI backend and Streamlit research workspace;
- automated data-quality, regression, and contract validation.
Hybrid recommendation system combining CatBoost + ALS + SBERT, with FastAPI, Streamlit, Docker, and Qdrant.
End-to-end conversion prediction pipeline using a calibrated CatBoost / XGBoost / LightGBM ensemble, FastAPI, Streamlit, Airflow DAGs, and Docker Compose.
Monitoring pipeline with Evidently, SHAP, PSI, and Jensen–Shannon divergence, including configurable alert policies.
Forecasting experiments with ARIMA, TBATS, Prophet, Darts, time-aware validation, and Optuna-tuned baselines.
Bioinformatics project exploring RNA-seq dimensionality reduction and representation learning for survival analysis.
English — B2 · French — B2
Email: konnik1000@gmail.com Telegram: @Konnik1988 GitHub: https://github.com/KonNik88
Machine Learning Engineer with a molecular genetics / biomedical background, building end-to-end ML systems across data engineering, retrieval, modeling, evaluation, APIs, and deployment.
Currently developing ML Research Radar — a multi-source research intelligence platform with semantic search, hybrid retrieval, graph analytics, PostgreSQL, FastAPI, Streamlit, and reproducible data/ML pipelines.
