Completely Offline, Privacy-First, Zero-Hallucination Q&A Agent
This project is a completely offline Retrieval-Augmented Generation (RAG) assistant that runs locally on your device without any internet connection. Built with Microsoft Foundry Local, it allows you to query your local documents (PDFs, text files, notes) securely using your device's CPU/NPU resources.
- 🔒 100% Offline & Secure: No data is sent to cloud servers or third-party APIs. All inference happens locally.
- 🧠 Zero Hallucination: Thanks to the RAG architecture and a strict System Prompt, the model does not fabricate answers. If the information is not in the database, it clearly states: "I do not have this information."
- ⚡ High Performance: Optimized inference times averaging ~0.5 - 1.5 seconds per query.
- 🗄️ Lightweight Database: Uses embedded
SQLitefor vector (Embedding) storage and Cosine Similarity calculations without needing an external vector database server. - 🛡️ Error Handling: Robust CLI interface protected against empty or nonsensical user inputs.
The pipeline consists of three main stages:
- Data Ingestion: Documents are chunked, converted to embeddings using
qwen3-embedding-0.6b, and stored in SQLite viaingest.py. - Retrieval: User queries are vectorized and compared against stored vectors using Cosine Similarity to fetch the most relevant context.
- Generation: The retrieved context is fed into the
phi-3.5-miniSmall Language Model (SLM) to generate a grounded response.
- Language: Python 3.11+
- Models:
qwen3-embedding-0.6b(Embedding),phi-3.5-mini(Chat) - Libraries:
foundry-local-sdk,sqlite3,time
git clone https://github.com/BurakHINGE/LocalRagAssistant.git
cd LocalRagAssistant
python3 -m venv venv
source venv/bin/activate # For macOS/Linux
# For Windows: venv\Scripts\activatepip install -r requirements.txtPlace your documents (e.g., .txt) inside the docs/ folder and run the ingestion script:
python3 ingest.pyRun the main app to start chatting via CLI:
python3 app.py- Direct information retrieval: ~0.74 seconds
- Logical deduction queries: ~1.09 seconds
- Out-of-context edge cases (Refusals): ~0.38 seconds
Tamamen Çevrimdışı, Gizlilik Odaklı, Sıfır Halüsinasyon Q&A Asistanı
Bu proje, yerel belgeleriniz (notlar, metin dosyaları) üzerinde çalışan, internet bağlantısı gerektirmeyen ve cihazınızın kaynaklarını (CPU/NPU) kullanarak çalışan bir RAG (Retrieval-Augmented Generation) asistanıdır. Microsoft Foundry Local altyapısı kullanılarak geliştirilmiştir.
- 🔒 %100 Çevrimdışı ve Güvenli: Hiçbir veri bulut sunucularına veya üçüncü parti API'lere gönderilmez.
- 🧠 Sıfır Halüsinasyon (Zero-Hallucination): RAG mimarisi ve katı Sistem Komutu (System Prompt) sayesinde model uydurma yapmaz. Bilgi veritabanında yoksa net bir şekilde cevap veremediğini belirtir.
- ⚡ Ultra Yüksek Performans: Optimizasyonlar sayesinde ortalama yanıt süresi ~0.5 - 1.5 saniye arasındadır.
- 🗄️ Hafif Veritabanı Mimarisi: Vektör (Embedding) depolama ve arama işlemleri için harici bir sunucuya ihtiyaç duymayan gömülü
SQLitekullanır. - 🛡️ Hata Yönetimi (Error Handling): Boş veya anlamsız kullanıcı girdilerine karşı korumalı CLI arayüzü.
Proje üç temel yapıtaşından oluşmaktadır:
- Data Ingestion (Veri Besleme):
ingest.pyaracılığıyla belgeler parçalara ayrılır,qwen3-embedding-0.6bile vektörlere dönüştürülür ve SQLite'a kaydedilir. - Retrieval (Arama): Kullanıcı sorusu vektöre dönüştürülür ve SQLite veritabanındaki kayıtlarla Kosinüs Benzerliği kullanılarak en alakalı bağlam (context) anında çekilir.
- Generation (Üretim): Elde edilen bağlam, küçük ve hızlı bir dil modeli olan
phi-3.5-minimodeline aktarılır ve yerel donanım ivmesiyle cevap üretilir.
git clone https://github.com/BurakHINGE/LocalRagAssistant.git
cd LocalRagAssistant
python3 -m venv venv
source venv/bin/activate # macOS/Linux için
# Windows için: venv\Scripts\activatepip install -r requirements.txtpython3 ingest.pypython3 app.pyMehmet Burak Menteşe Marmara Üniversitesi - Bilgisayar Mühendisliği (İngilizce) Bu proje, Microsoft AI Innovators Summer Program (2026) kapsamında geliştirilmiş bitirme projesidir. Temel amacı bulut sistemlerine bağımlı olmadan, güvenli ve kişiselleştirilmiş yerel yapay zeka asistanlarının gücünü göstermektir.