This project is built for the MindsDB Quest-19: Stress-Test Knowledge Bases challenge. It demonstrates a complete end-to-end pipeline using MindsDB, local Ollama LLMs, and Streamlit UI for:
- 📊 Customer support ticket analysis
- 🎬 Movie-related question answering
- ✅ Accuracy benchmarking
This application performs the following:
- Creates a Knowledge Base (
customer_tickets_kb) - Ingests a CSV dataset (
customer_support_tickets.csv) with metadata - Builds a vector index for semantic search
- Creates an AI model using LLaMA3 via Ollama
- Deploys a retrieval-augmented agent (
ticket_support_agent) - Supports interactive Q&A and summarization of customer issues
- Loads a summary-based knowledge base (
movies_kb) - Creates an agent (
movie_expert_agent) using LLaMA2 - Answers plot, character, and theme-related movie questions
- Measures ingestion time, semantic query latency, and p95/p99 response delays
- Outputs structured
.mdreports underbenchmarks/
Contains 8,469 support tickets with metadata:
- Ticket ID, Subject, Description
- Priority, Status, Type, Channel
Contains:
- Movie title, genre, actors, year, rating
- ✅ Original dataset size: 238,256 rows
- ✅ After removing duplicates: 161,765 rows
- ✅ Saved to
imdb_movies_prepared.csvfor clean ingestion
cd mindsdb-kbpython -m venv venv
source venv/bin/activate # Or venv\Scripts\activate on Windows
pip install -r requirements.txtstreamlit run app.pystreamlit run movie.pypython benchmark.pyEach .md report includes:
| Category | Metric |
|---|---|
| ⏱️ Ingestion Time | Total seconds + time per 1K rows |
| ⚡ Query Latency | Average, p95, p99 response times |
Built with ❤️ using MindsDB, Ollama, Streamlit, and Python.
