Data Scientist · AI / ML Engineering
Applied machine learning, retrieval systems, and reproducible experimentation.
I'm a Data Scientist focused on building and evaluating AI systems. My work brings together machine learning, data pipelines, and software engineering, with a particular interest in natural language processing and retrieval-augmented generation (RAG).
I approach projects through clear hypotheses, controlled experiments, and inspectable results. I care about understanding why a model improves, which evidence supports an answer, and how someone else can reproduce the outcome.
- Applied machine learning: data preparation, experimental design, and model evaluation.
- NLP and language models: retrieval-augmented generation and research into model adaptation, grounding, and uncertainty.
- AI engineering: APIs, durable ingestion pipelines, retrieval infrastructure, and workflow orchestration.
- Reproducibility: versioned data, explicit configurations, retrieval traces, and documented experiment artifacts.
A local-first research engineering system, evolved from RAGForge, for computer science and AI experimentation. Its most developed area is an end-to-end RAG workflow that makes documents, retrieval evidence, and query history inspectable.
- Processes versioned documents through Bronze, Silver, and Gold data layers.
- Combines dense and sparse retrieval in Qdrant.
- Streams answers with ranked evidence and preserves query history.
- Includes retrieval tracing and orchestration benchmarks to support analysis of system behavior.
Technologies: Python, FastAPI, PostgreSQL, Qdrant, MinIO, Redis, Airflow, Celery, Next.js, Docker Compose.
Status: RAG capabilities are implemented; a broader framework for experiments across research domains is a planned direction.
An ongoing study of how fine-tuning and retrieval affect language-model answers in a medical question-answering dataset. The central research question is when a model should rely on adaptation, retrieve supporting evidence, or abstain from answering.
The planned evaluation compares four configurations: base Qwen, base Qwen with RAG, QLoRA-adapted Qwen, and QLoRA-adapted Qwen with RAG.
- Completed: data cleaning and group-based splitting designed to prevent data leakage.
- In progress: Qwen input formatting and token analysis.
- Planned: QLoRA training, retrieval-corpus construction, comparative evaluation, uncertainty analysis, and retrieval stress tests.
The study examines performance across common and rare topics, unsupported claims, and model behavior when evidence is irrelevant or contradictory.
| Area | Tools and technologies |
|---|---|
| Machine learning | Python, PyTorch |
| APIs and applications | FastAPI |
| Data and retrieval | PostgreSQL, Qdrant, MinIO, Redis |
| Pipelines and orchestration | Apache Airflow, Celery |
| Development and infrastructure | Docker, Docker Compose, Linux, Git |
I'm open to Data Science opportunities and collaborations in AI/ML engineering, NLP, applied AI, and model evaluation.
Explore my portfolio or contact me at elbaz.ouad1249@gmail.com.