Skip to content
syedazeez337Public

About

Local hybrid RAG over the vLLM codebase — BGE-M3 + Qdrant (dense+sparse, RRF) + cross-encoder rerank + Qwen3.5-4B, with a source-grounded evaluation harness.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

raglearn

A hybrid retrieval-augmented generation (RAG) pipeline that answers questions about the vLLM codebase, with an honest, source-grounded evaluation harness. Built to run fully locally on a single 6 GB GPU.

  • Hybrid retrieval — BGE-M3 dense + learned-sparse vectors, fused with RRF in Qdrant
  • Cross-encoder reranking — bge-reranker-v2-m3, with scope-aware dedup
  • Grounded generation — Qwen3.5-4B on llama.cpp; answers cite file paths or refuse
  • Trustworthy eval — 30 questions, references read from source (not the model), LLM-judged, retrieval and generation scored separately

Pipeline architecture

The code is organised so each module maps to exactly one box in the diagram above.

rag/
  config.py                  shared settings: models, collection, endpoints, system prompt
  indexing/                  ── Offline Indexing Pipeline ──
    load.py                  Raw Documents     — load & clean vLLM source
    chunk.py                 Chunking          — AST split (class bodies grouped)
    embed.py                 Embedding Model   — BGE-M3 dense + learned sparse
    index.py                 Vector Database   — Qdrant create + upsert
    pipeline.py              orchestrates load → chunk → embed → index
  retrieval/                 ── Online Retrieval-Generation Pipeline ──
    query_encoder.py         Query Encoder     — embed the query (same BGE-M3)
    search.py                Similarity Search — Qdrant hybrid dense+sparse, RRF fusion
    rerank.py                Re-ranker         — bge-reranker-v2-m3 + scope-aware dedup
    generate.py              LLM Generator     — llama.cpp (Qwen3.5-4B)
    answer.py                Answer + Citations — orchestrates the online pipeline
  server.py                  FastAPI service exposing the online pipeline
  eval/                      evaluation harness (dataset, runner, reference gate, model probe)

scripts/serve_llm.sh         LLM Generator backend (llama.cpp launcher)
docs/                        architecture + eval visualisations
data/                        eval result snapshots (index data is regenerated, not committed)

Prerequisites

This repo is the pipeline code. The large external pieces are not committed — set them up once:

  1. vLLM source (the corpus) into ./vllm:
    git clone --depth 1 https://github.com/vllm-project/vllm.git vllm
  2. llama.cpp (the LLM backend) built into ./llama.cpp:
    git clone https://github.com/ggml-org/llama.cpp.git
    cmake -S llama.cpp -B llama.cpp/build -DGGML_CUDA=ON && cmake --build llama.cpp/build -j
  3. The model — a Qwen3.5-4B GGUF into ./models/Qwen3.5-4B-Q4_K_M.gguf (or set LLAMA_MODEL to your path).
  4. Qdrant via Docker (data persists in a named volume):
    docker run -d --name qdrant -p 6333:6333 -v qdrant_data:/qdrant/storage qdrant/qdrant
  5. Python deps (Python 3.12; CUDA torch — see note in requirements.txt):
    uv venv --python 3.12 && source .venv/bin/activate
    uv pip install -r requirements.txt
  6. Eval key (optional, only for the judge): cp .env.example .env and add your GOOGLE_API_KEY.

Running

# 1. LLM Generator backend (:8080)
scripts/serve_llm.sh

# 2. Build the index once  (load → chunk → embed → upsert)
python -m rag.indexing.pipeline --recreate

# 3. Online service (:8000)
python -m uvicorn rag.server:app --host 127.0.0.1 --port 8000

# 4. Query it
curl -s :8000/query -H 'content-type: application/json' \
  -d '{"question":"What is the default KV cache block_size?","top_k":5}'

Every component is also runnable standalone, e.g. python -m rag.retrieval.search "<query>".

Evaluation

The eval set has 30 questions across six scenarios: A–D answerable (architectural, API-surface, factual, cross-component), E out-of-corpus (must refuse), F false-premise (must reject). A–D references are read from the vLLM source on disk, so answer-correctness is independent of the system's own output. Retrieval (did the ground-truth file get retrieved?) and generation (is the answer right?) are scored separately.

python -m rag.eval.check_references                                  # gate: A-D references filled
JUDGE_MODEL=gemini-2.5-flash python -m rag.eval.run --no-resume      # full 30-question run

Results

Evaluation results

A chunking bug was found where oversized config-class bodies were shattered into micro-chunks that all shared one signature, so retrieval-side dedup made ~15% of the corpus unreachable. Grouping class-body statements and making the dedup key scope/part-aware lifted answer accuracy:

Metric Before After
Answer accuracy (A–D) 0.62 0.85
Scenario D (cross-component) 0.40 0.80
Out-of-corpus refusal (E) 10/10 10/10
False-premise rejection (F) 10/10 10/10

Notably, context-recall stayed flat (0.75 → 0.72) while accuracy rose +0.22 — the bug was never which files were retrieved, but what was inside the chunks and dedup discarding the good ones.

Known upstream warning (not suppressed)

Loading BGE-M3 emits, on Python 3.10+, a DeprecationWarning from sentencepiece: builtin type SwigPyObject has no __module__ attribute. It originates in SWIG-generated bindings and is fixed only by SWIG 4.4 + a sentencepiece rebuild (swig/swig#2881, google/sentencepiece#1150). No released sentencepiece fixes it yet. It is a benign DeprecationWarning (ignored by Python's default filter, so it does not appear in normal runs) and is left visible rather than filtered out.

License

MIT

About

Local hybrid RAG over the vLLM codebase — BGE-M3 + Qdrant (dense+sparse, RRF) + cross-encoder rerank + Qwen3.5-4B, with a source-grounded evaluation harness.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages