Sign inSign up
P

Thanu Prasad D

Community User

Displaying 1 to 4 of 4 repositories

image

Dockerized Qwen2.5 API with FastAPI, GPU support, and scalable transformer inference.

7m

246

1

image

Dockerized TinyLlama chat API with FastAPI, supports CPU/GPU inference and REST endpoints.

7m

788