LocalGPT is a Retrieval-Augmented Generation (RAG) system built using LlamaIndex for managing documents, Ollama for language model inference, and Qdrant for vector-based search and retrieval. This project enables efficient, context-aware Q&A from your local document files using advanced LLMs and vector databases.
- LlamaIndex Integration: Enables document ingestion, chunking, and semantic search for relevant information.
- Ollama for LLM Inference: Use locally hosted Llama models through Ollama API for language generation tasks.
- Qdrant Integration: Fast and efficient vector-based search using Qdrant for document indexing and retrieval.
- Gradio Interface: Simple web interface for uploading documents, interacting with the chatbot, and retrieving answers from your knowledge base.
- Docker (if using Docker-based installation)
- Python 3.8 or later (for manual setup)
- Qdrant: Vector database, preferably running on localhost.
- Ollama: For local Llama-based models.
- Gradio: Web UI for interacting with the chatbot.
- Clone the Repository:
git clone https://github.com/Saurab-Shrestha/LocalGPT.git
cd localgpt- Create a Virtual Environment:
python3 -m venv env
source env/bin/activate- Install Python Dependencies:
pip install -r requirements.txt- Install and Run Qdrant:
docker run -p 6333:6333 qdrant/qdrant
-
Install and Configure Ollama: Install and Configure Ollama: Follow Ollama installation instructions for your operating system.
-
Run the Application:
gradio app.py
Before running the app, you need to configure the settings for the LLM, Qdrant and the other services.
- Edit the Configuration File: Update the
config.pyfile to specify the correct host, port and model setting for you setup. - Model Configuration: Make sure you have your LLM model and embedding running in the Ollama, configured as per you needs.