Sign inSign up

deepak93p/nexlify-data-ingestion-server

By deepak93p

•Updated about 1 year ago

Data Ingestion Service for a Retrieval-Augmented Generation (RAG)-based AI chatbot

Image
Machine learning & AI
0

271

deepak93p/nexlify-data-ingestion-server repository overview

⁠Nexlify Data Ingestion Server

This Docker image provides a production-ready FastAPI service for file upload, text extraction (PDF, HTML, etc.), Google Gemini embedding, Qdrant vector DB ingestion, semantic search with filters, and Confluence page crawling. It includes robust error handling, OCR support for image-based PDFs, and health checks for observability.

⁠Quick Start

To run the container, ensure you have Docker installed. Pull the image (assuming it's published as deepakpant93/nexlify-data-ingestion) and start it with required environment variables:

docker run -d -p 7860:7860 \
  --env QDRANT_HOST=localhost \
  --env QDRANT_PORT=6333 \
  --env GEMINI_API_KEY=your_gemini_api_key \
  deepakpant93/nexlify-data-ingestion

The service will be available at http://localhost:7860. For Confluence integration, add the relevant environment variables.

⁠Environment Variables

Configure the service using these environment variables:

  • QDRANT_HOST: Hostname of the Qdrant instance (default: localhost).
  • QDRANT_PORT: Port for Qdrant (default: 6333).
  • GEMINI_API_KEY: Your Google Gemini API key (required for embeddings).
  • CONFLUENCE_BASE_URL: Base URL for Confluence (e.g., https://your-domain.atlassian.net/wiki).
  • CONFLUENCE_SPACE_KEY: Your Confluence space key.
  • CONFLUENCE_API_USER: Confluence API username (email).
  • CONFLUENCE_API_TOKEN: Confluence API token.

Use a .env file or pass them directly via --env or --env-file.

⁠Usage

⁠Custom Configuration

Mount a custom configuration or override defaults by passing environment variables as shown in Quick Start. For persistent storage or external Qdrant, ensure the host and port are accessible from the container.

⁠Building Locally

Clone the repository and build the image:

git clone [email protected]:DeepakPant93/nexlify.git
cd nexlify/data-ingestion-server
docker build -t nexlify-data-ingestion .
docker run -p 7860:7860 --env-file .env nexlify-data-ingestion

The Dockerfile uses python:3.11-slim as the base, installs dependencies like Tesseract OCR and Poppler, and runs the app with Uvicorn on port 7860.

⁠Features

  • File upload and text extraction for PDFs, HTML, TXT with OCR fallback.
  • Google Gemini embeddings with retry and backoff mechanisms.
  • Qdrant vector database ingestion including metadata like source and filename.
  • Semantic search with optional filters (e.g., by filename or source).
  • Confluence crawler for ingesting pages with HTML cleanup.
  • API endpoints for health checks, embeddings, and administration.
  • Dockerized for easy deployment with extensibility in mind.

⁠Source Code

The source code is available on GitHub⁠.

⁠Support

For issues, open a ticket on the GitHub repository. Contributions are welcome—fork, branch, and submit a PR with tests if possible.

⁠License

This project is licensed under the MIT License⁠. See the repository for details.

Tag summary

Content type

Image

Digest

sha256:8b3755f07…

Size

389.2 MB

Last updated

about 1 year ago

docker pull deepak93p/nexlify-data-ingestion-server