HK
MultiDocChat v2 dashboard showing hybrid BM25 and dense retrieval evaluation scorecard
Back to Projects

MultiDocChat v2

Live · Deployed2026

The Problem

Traditional single-retriever RAG systems often miss relevant passages or surface irrelevant ones, especially when dealing with multiple documents and live web pages containing overlapping terminology. Keyword-only search misses semantic meaning, while dense-only retrieval can miss exact terminology, identifiers, and acronyms. Enterprise workflows require systems that combine both search paradigms, verify cross-document factual contradictions, evaluate retrieval metrics with RAGAS, and provide confidence-calibrated attributed answers — all on a 100% free open-source stack requiring zero paid API keys.

Architecture & Approach

The platform implements a multi-tier Hybrid Retrieval architecture. Ingestion processes multi-format files (PDF, DOCX, TXT, MD) and live web URLs (via BeautifulSoup4) into ChromaDB using local sentence-transformers (all-MiniLM-L6-v2) and a pure-Python BM25 Okapi keyword index. User queries and windowed history are condensed before executing min-max normalized weighted fusion (Score = α · Semantic_norm + (1-α) · BM25_norm) with per-source balancing. Generation and synthesis are handled by NVIDIA NIM (nvidia/nemotron-3-super-120b-a12b). Concurrently, a two-stage conflict detection engine runs a heuristic gate followed by targeted LLM verification; a multi-factor confidence scorer assigns 0.0–1.0 ratings with color-coded badges; and an automated RAGAS engine benchmarks Faithfulness, Relevancy, Precision, and Recall. The platform operates across a 6-tab Streamlit UI featuring Plotly analytics, full-text chunk exploration, document diffing, and session report export in Markdown and printable HTML/PDF.

About the Project

A production-grade Document Intelligence and Hybrid RAG platform fusing ChromaDB dense vector search (local all-MiniLM-L6-v2) with a pure-Python BM25 Okapi keyword index via min-max weighted score fusion, achieving 100% retrieval precision@k on multi-document evaluation sets. Features dual ingestion (local multi-format documents + live web scraping), an automated RAGAS evaluation scorecard, two-stage source conflict detection, multi-factor confidence scoring (0.0–1.0 color-coded badges), interactive Plotly analytics dashboard, chunk explorer, contract diff engine with LLM executive summaries, and session report export — all across a 6-tab Streamlit UI. Powered by NVIDIA NIM (nemotron-3-super-120b) with a 100% free open-source stack validated by 34/34 passing automated tests.

Key Metrics

100%

Retrieval Precision@k

34/34 passing

Automated Tests

NVIDIA NIM Nemotron-3

LLM Engine

100% free open-source

Stack Cost

Technologies

PythonLangChainChromaDBNVIDIA NIMStreamlitBM25 OkapiRAGASsentence-transformersPlotlyBeautifulSoup4PyYAML