Sign inSign up

eosphorosai/dbgpt

By eosphorosai

•Updated about 1 month ago

dbgpt docker images repo

Image
15

100K+

eosphorosai/dbgpt repository overview

⁠DB-GPT: Revolutionizing Database Interactions with Private LLM Technology

⁠What is DB-GPT?

DB-GPT is an experimental open-source project that uses localized GPT large models to interact with your data and environment. With this solution, you can be assured that there is no risk of data leakage, and your data is 100% private and secure.

⁠Contents

DB-GPT Youtube Video⁠

⁠Demo

Run on an RTX 4090 GPU.

⁠Chat Excel

excel

⁠Chat Plugin

auto_plugin_new

⁠LLM Management

llm_manage

⁠FastChat && vLLM

vllm

⁠Trace

trace_new

⁠Chat Knowledge

kbqa_new

⁠Install

Docker Linux macOS Windows

Usage Tutorial⁠

⁠Features

Currently, we have released multiple key features, which are listed below to demonstrate our current capabilities:

⁠Introduction

The architecture of the entire DB-GPT is shown.

The core capabilities mainly consist of the following parts:

  1. Multi-Models: Support multi-LLMs, such as LLaMA/LLaMA2、CodeLLaMA、ChatGLM, QWen、Vicuna and proxy model ChatGPT、Baichuan、tongyi、wenxin etc
  2. Knowledge-Based QA: You can perform high-quality intelligent Q&A based on local documents such as PDF, word, excel, and other data.
  3. Embedding: Unified data vector storage and indexing, Embed data as vectors and store them in vector databases, providing content similarity search.
  4. Multi-Datasources: Used to connect different modules and data sources to achieve data flow and interaction.
  5. Multi-Agents: Provides Agent and plugin mechanisms, allowing users to customize and enhance the system's behavior.
  6. Privacy & Secure: You can be assured that there is no risk of data leakage, and your data is 100% private and secure.
  7. Text2SQL: We enhance the Text-to-SQL performance by applying Supervised Fine-Tuning (SFT) on large language models
⁠RAG-IN-Action

⁠SubModule

⁠Image

🌐 AutoDL Image⁠

⁠Language Switching
In the .env configuration file, modify the LANGUAGE parameter to switch to different languages. The default is English (Chinese: zh, English: en, other languages to be added later).

⁠Contribution

⁠RoadMap

roadmap

⁠KBQA RAG optimization
  • Multi Documents

    • PDF
    • Excel, CSV
    • Word
    • Text
    • MarkDown
    • Code
    • Images
  • RAG

  • Graph Database

    • Neo4j Graph
    • Nebula Graph
  • Multi-Vector Database

    • Chroma
    • Milvus
    • Weaviate
    • PGVector
    • Elasticsearch
    • ClickHouse
    • Faiss
  • Testing and Evaluation Capability Building

    • Knowledge QA datasets
    • Question collection [easy, medium, hard]:
    • Scoring mechanism
    • Testing and evaluation using Excel + DB datasets
⁠Multi Datasource Support
  • Multi Datasource Support
    • MySQL
    • PostgreSQL
    • Spark
    • DuckDB
    • Sqlite
    • MSSQL
    • ClickHouse
    • Oracle
    • Redis
    • MongoDB
    • HBase
    • Doris
    • DB2
    • Couchbase
    • Elasticsearch
    • OceanBase
    • TiDB
    • StarRocks
⁠Multi-Models And vLLM
⁠Agents market and Plugins
  • multi-agents framework
  • custom plugin development
  • plugin market
  • Integration with CoT
  • Enrich plugin sample library
  • Support for AutoGPT protocol
  • Integration of multi-agents and visualization capabilities, defining LLM+Vis new standards
⁠Cost and Observability
⁠Text2SQL Finetune
  • support llms

    • LLaMA
    • LLaMA-2
    • BLOOM
    • BLOOMZ
    • Falcon
    • Baichuan
    • Baichuan2
    • InternLM
    • Qwen
    • XVERSE
    • ChatGLM2
  • SFT Accuracy

As of October 10, 2023, by fine-tuning an open-source model of 13 billion parameters using this project, the execution accuracy on the Spider evaluation dataset has surpassed that of GPT-4!

nameExecution Accuracyreference
GPT-40.762numbersstation-eval-res⁠
ChatGPT0.728numbersstation-eval-res⁠
CodeLlama-13b-Instruct-hf_lora0.789sft train by our this project,only used spider train dataset ,the same eval way in this project with lora SFT
CodeLlama-13b-Instruct-hf_qlora0.774sft train by our this project,only used spider train dataset ,the same eval way in this project with qlora and nf4,bit4 SFT
wizardcoder0.610text-to-sql-wizardcoder⁠
CodeLlama-13b-Instruct-hf0.556eval in this project default param
llama2_13b_hf_lora_best0.744sft train by our this project,only used spider train dataset ,the same eval way in this project

More Information about Text2SQL finetune⁠

⁠Licence

The MIT License (MIT)

Star History Chart

Tag summary

Content type

Image

Digest

sha256:9bb269844…

Size

7.1 GB

Last updated

about 1 month ago

docker pull eosphorosai/dbgpt