Offline PDF-based AI chatbot using Ollama + FAISS + LangChain with Streamlit UI and Docker support
332
The Universal Offline AI Chatbot is a privacy-respecting, offline-ready assistant that can chat over any set of PDFs. Itβs ideal for legal, cybersecurity, academic, enterprise, or technical domains.
It uses a locally hosted LLM (mistral:instruct via Ollamaβ ) and semantic search powered by HuggingFace embeddings and FAISS. You get fast, accurate responses, without sending anything to the cloud.
all-MiniLM-L6-v2| Layer | Stack |
|---|---|
| LLM | mistral:instruct via Ollama |
| Embeddings | all-MiniLM-L6-v2 via SentenceTransformers |
| Vector Store | FAISS (in-memory + disk) |
| Framework | LangChain (v0.2+) |
| Language | Python 3.11+ |
| UI | Streamlit |
| Container | Docker |
| CI/CD | GitHub Actions (.github/workflows/python.yml) |
β οΈ HuggingFace Token is required to fetch the embedding model once. It's cached locally afterward.
Example .env:
HF_TOKEN=your_huggingface_token_here
| Chatbot Type | Add These PDFs |
|---|---|
| π¨ββοΈ LawyerBot | Legal, Constitution, HR documents |
| 𧬠ResearchBot | Whitepapers, scientific papers |
| π‘οΈ CyberSecBot | SOC2, GDPR, ISO27001, NIST docs |
| π EdTechBot | Notes, textbooks, question banks |
| π§βπΌ HR/CompanyBot | SOPs, onboarding docs, HR policies |
Universal-Offline-AI-Chatbot/
β
βββ data/ # Place your PDF documents here
β βββ Try.pdf
β
βββ Screenshots/ # UI snapshots
β βββ Loading_Screen.png
β βββ Running_the_Model.png
β
βββ src/ # Modular source code
β βββ chunker.py
β βββ config.py
β βββ embedding.py
β βββ loader.py
β βββ model_loader.py
β βββ prompts.py
β βββ qa_chain.py
β βββ utils.py
β βββ vectorstore.py
β
βββ vectorstore/ # Local FAISS vector index
β βββ db_faiss/
β
βββ Bot.py # CLI script
βββ Bot.ipynb # Jupyter notebook version
βββ main.py # Entry-point (optional)
βββ streamlit_app.py # Frontend UI (Streamlit)
βββ requirements.txt # Python dependencies
βββ setup.ps1 # PowerShell setup script
βββ Dockerfile # Docker image definition
βββ .dockerignore
βββ .env # Contains HF_TOKEN
βββ README.md
βββ LICENSE
.\setup.ps1
This will:
pip install -r requirements.txt
ollama pull mistral:instruct
export HUGGINGFACEHUB_API_TOKEN=your_token # macOS/Linux
set HUGGINGFACEHUB_API_TOKEN=your_token # Windows CMD
python Bot.py
streamlit run streamlit_app.py
.env file containing HF_TOKEN (Hugging Face token)To build and run the chatbot using Docker, follow these steps:
Build the Docker image:
docker build -t ai-chatbot .
Run the container (with volume mount and token):
docker run -p 8501:8501 --env-file .env -v ${PWD}/data:/app/data ai-chatbot
This will:
8501 to local 8501.env for HF_TOKENdata/ folder into the container for access to PDFsAccess the chatbot at http://localhost:8501β
Let me know if you want the actual screenshot file names changed or if youβd like a quick CLI script to generate and store those screenshots automatically during your next run.
# Replace default file(s)
mv your_files/*.pdf ./data/
# Re-run the bot or restart Streamlit
python Bot.py
Automatically re-indexes your new documents using FAISS.
π§ You: What does Article 21 state?
π€ Bot: Article 21 of the Indian Constitution guarantees the protection of life and personal liberty...
Aditya Bhatt
Cybersecurity Specialist | VAPT Expert | OSS Contributor
GitHubβ | Mediumβ
Content type
Image
Digest
sha256:72941a9bbβ¦
Size
3.3 GB
Last updated
over 1 year ago
docker pull adityabhatt3010/ai-chatbot