Sign inSign up

zedd27/ragxollama

By zedd27

•Updated over 2 years ago

Minimal Node.js RESTful API for a RAG system with ollama and chromaDB

Image
Machine learning & AI
Data science
Web servers
0

196

zedd27/ragxollama repository overview

⁠RagXOllama

A minimal Node.js RESTful server for a Retrieval-Augmented Generation (RAG) system using ChromaDB and Ollama!

This repo is intended to be used as a template to easily integrate any LLM with a basic RAG functionality into your MERN (or any other node.js) stack apps natively.

⁠Setup

⁠Method 1 (using Docker)
  1. Add your documents to Docs/ folder.
  2. Simply build the docker containers using
git clone https://github.com/vteam27/RagXOllama
cd RagXOllama
docker compose up 
  1. Wait for the data to be ingested to DB and the LLM model to be downloaded.
  2. Go to http://localhost:3000
⁠Method 2 (build locally)
  1. Ollama⁠ : To serve open source LLMs locally (one time setup)

    ollama serve
    ollama run llama3
    
  2. ChromaDB Backend⁠: Spin up the ChromaDB core

    docker pull chromadb/chroma
    docker run -p 8000:8000 chromadb/chroma
    
  3. Add your PDF documents to Docs/

  4. Install dependencies and run!

    git clone https://github.com/vteam27/RagXOllama
    cd RagXOllama
    npm install
    node app.js
    
  5. Go to http://localhost:3000 to chat using a web demo.

  6. Alternatively, you can run node pipeline.js to interact using terminal, easy development and testing.

⁠Features

  • Chat with private documents securely!
  • Search and query your data easily using Natural Language only!
  • Keep all your LLMs up to date with the latest data.
  • Use any open source LLM of your choice. Browse LLMs⁠

⁠Example

RAG_example

  1. Ingestion: Feed relevant information into the chromaDB vector store collection.
  2. Retrieve: We retrieve relevant (top n) chunks of factual information stored in chromaDB using a algorithm that calculates it's similarity score with the user query.
  3. Augment: We append this context into our prompt.
  4. Generate: We feed this prompt to a 8B llama3 4-bit quantized model served by ollama to generate the desired response.

View the logs.txt file for the full output of my code.

⁠About RAG

With the help of Retrieval-Augmented Generation (RAG) we can achieve:

  • Improved Factual Accuracy: By relying on retrieved information, RAG systems can ensure answers are grounded in real-world data.
  • Domain Specificity: RAG allows you to integrate domain-specific knowledge bases, making the system more knowledgeable in a particular area.
  • Adaptability to New Information: By using external knowledge sources, RAG systems can stay up-to-date with the latest information, even if the LLM itself wasn't specifically trained on it.

image

Tag summary

Content type

Image

Digest

sha256:65dcc1a3c…

Size

217.2 MB

Last updated

over 2 years ago

docker pull zedd27/ragxollama:tinyllama