Sign inSign up
P

Thanu Prasad D

Community User

Displaying 1 to 4 of 4 repositories

image

Dockerized Qwen2.5 API with FastAPI, GPU support, and scalable transformer inference.

6m

244

1

image

Dockerized TinyLlama chat API with FastAPI, supports CPU/GPU inference and REST endpoints.

6m

781