OpenAI-compatible REST proxy for Amazon Bedrock. AWS Lambda image (Web Adapter).
7.1K
OpenAI-compatible RESTful APIs for Amazon Bedrock
API Gateway Response Streaming Support - You can now deploy with Amazon API Gateway REST API instead of ALB, enabling true response streaming for better latency and cost optimization. See Deployment Options for details.
Latest Models Supported:
It also supports reasoning for Claude 4/4.5 (extended thinking and interleaved thinking) and DeepSeek R1. Check How to Use for more details. You need to first run the Models API to refresh the model list.
Amazon Bedrock offers a wide range of foundation models (such as Claude 3 Opus/Sonnet/Haiku, Llama 2/3, Mistral/Mixtral, etc.) and a broad set of capabilities for you to build generative AI applications. Check the Amazon Bedrock landing page for additional information.
Sometimes, you might have applications developed using OpenAI APIs or SDKs, and you want to experiment with Amazon Bedrock without modifying your codebase. Or you may simply wish to evaluate the capabilities of these foundation models in tools like AutoGen etc. Well, this repository allows you to access Amazon Bedrock models seamlessly through OpenAI APIs and SDKs, enabling you to test these models without code changes.
If you find this GitHub repository useful, please consider giving it a free star ⭐ to show your appreciation and support for the project.
Features:
Please check Usage Guide for more details about how to use the new APIs.
Please make sure you have met below prerequisites:
For more information on how to request model access, please refer to the Amazon Bedrock User Guide (Set Up > Model access)
The following diagram illustrates the reference architecture. It uses Amazon API Gateway response streaming with Lambda for SSE support.

| Option | Pros | Cons | Best For |
|---|---|---|---|
| API Gateway + Lambda | No VPC required, pay-per-request, native streaming support, lower operational overhead | Potential cold starts | Most use cases, cost-sensitive deployments |
| ALB + Fargate | Lowest streaming latency, no cold starts | Higher cost, requires VPC | High-throughput, latency-sensitive workloads |
| Pre-built Docker Hub image | No build step, works with any container runtime (ECS, EKS, Kubernetes, Docker, Podman) | You manage the infra | Self-hosted / non-AWS runtimes, quick trials |
You can also use Lambda Function URL as an alternative, see example
Container images are built by the Publish Docker image to Docker Hub GitHub Actions workflow on every push to main and every v*.*.* tag. Two images are published:
| Image | Dockerfile | Platforms | Use for |
|---|---|---|---|
<your-dockerhub-username>/bedrock-access-gateway | src/Dockerfile_ecs | linux/amd64, linux/arm64 | ECS/Fargate, Kubernetes, docker run, docker-compose |
<your-dockerhub-username>/bedrock-access-gateway-lambda | src/Dockerfile | linux/arm64 | AWS Lambda (ships the Lambda Web Adapter) |
Tags: latest (HEAD of main), semver (1.2.3, 1.2, 1) on release tags, and short commit SHA.
Run locally with the pre-built ECS image:
docker run --rm -p 8000:8080 \
-e API_KEY=my-local-key \
-e AWS_REGION=us-west-2 \
-v ~/.aws:/home/appuser/.aws:ro \
<your-dockerhub-username>/bedrock-access-gateway:latest
Base URL: http://localhost:8000/api/v1. Use any valid AWS credential source — an AWS_PROFILE, explicit AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY env vars, or a mounted ~/.aws dir as shown above.
Enabling the workflow on a fork: add two repository secrets — DOCKERHUB_USERNAME (your Docker Hub namespace) and DOCKERHUB_TOKEN (a Docker Hub personal access token with Read, Write, Delete scope). Trigger with a push, a v*.*.* tag, or manually via Actions → Publish Docker image to Docker Hub → Run workflow.
Staying in sync with upstream: a companion Sync from upstream workflow runs daily (06:00 UTC) and on manual dispatch. It merges aws-samples/bedrock-access-gateway:main into this fork's main (failing loudly on conflict so you can resolve by hand) and pushes the result. A successful sync then chains into the Docker Hub publish workflow via workflow_run, so new upstream commits become new image tags automatically.
Please follow the steps below to deploy the Bedrock Proxy APIs into your AWS account. Only supports regions where Amazon Bedrock is available (such as us-west-2). The deployment will take approximately 10-15 minutes 🕒.
Step 1: Create your own API key in Secrets Manager (MUST)
Note: This step is to use any string (without spaces) you like to create a custom API Key (credential) that will be used to access the proxy API later. This key does not have to match your actual OpenAI key, and you don't need to have an OpenAI API key. please keep the key safe and private.
Open the AWS Management Console and navigate to the AWS Secrets Manager service.
Click on "Store a new secret" button.
In the "Choose secret type" page, select:
Secret type: Other type of secret Key/value pairs:
Click "Next"
In the "Configure secret" page: Secret name: Enter a name (e.g., "BedrockProxyAPIKey") Description: (Optional) Add a description of your secret
Click "Next" and review all your settings and click "Store"
After creation, you'll see your secret in the Secrets Manager console. Make note of the secret ARN.
Step 2: Build and push container images to ECR
Clone this repository:
git clone https://github.com/aws-samples/bedrock-access-gateway.git
cd bedrock-access-gateway
Run the build and push script:
cd scripts
bash ./push-to-ecr.sh
Follow the prompts to configure:
latest)us-east-1)The script will build and push both Lambda and ECS/Fargate images to your ECR repositories.
Important: Copy the image URIs displayed at the end of the script output. You'll need these in the next step.
Step 3: Deploy the CloudFormation stack
Download the CloudFormation template you want to use:
deployment/BedrockProxy.templatedeployment/BedrockProxyFargate.templateSign in to AWS Management Console and navigate to the CloudFormation service in your target region.
Click "Create stack" → "With new resources (standard)".
Upload the template file you downloaded.
On the "Specify stack details" page, provide the following information:
Click "Next".
On the "Configure stack options" page, you can leave the default settings or customize them according to your needs. Click "Next".
On the "Review" page, review all details. Check the "I acknowledge that AWS CloudFormation might create IAM resources" checkbox at the bottom. Click "Submit".
That is it! 🎉 Once deployed, click the CloudFormation stack and go to Outputs tab, you can find the API Base URL from APIBaseUrl, the value should look like http://xxxx.xxx.elb.amazonaws.com/api/v1.
If you encounter any issues, please check the Troubleshooting Guide for more details.
All you need is the API Key and the API Base URL. If you didn't set up your own key following Step 1, the application will fail to start with an error message indicating that the API Key is not configured.
Now, you can try out the proxy APIs. Let's say you want to test Claude 3 Sonnet model (model ID: anthropic.claude-3-sonnet-20240229-v1:0)...
Example API Usage
export OPENAI_API_KEY=<API key>
export OPENAI_BASE_URL=<API base url>
# For older versions
# https://github.com/openai/openai-python/issues/624
export OPENAI_API_BASE=<API base url>
curl $OPENAI_BASE_URL/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-d '{
"model": "anthropic.claude-3-sonnet-20240229-v1:0",
"messages": [
{
"role": "user",
"content": "Hello!"
}
]
}'
Example SDK Usage
from openai import OpenAI
client = OpenAI()
completion = client.chat.completions.create(
model="anthropic.claude-3-sonnet-20240229-v1:0",
messages=[{"role": "user", "content": "Hello!"}],
)
print(completion.choices[0].message.content)
Please check Usage Guide for more details about how to use embedding API, multimodal API and tool call.
This proxy now supports Application Inference Profiles, which allow you to track usage and costs for your model invocations. You can use application inference profiles created in your AWS account for cost tracking and monitoring purposes.
Using Application Inference Profiles:
# Use an application inference profile ARN as the model ID
curl $OPENAI_BASE_URL/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-d '{
"model": "arn:aws:bedrock:us-west-2:123456789012:application-inference-profile/your-profile-id",
"messages": [
{
"role": "user",
"content": "Hello!"
}
]
}'
SDK Usage with Application Inference Profiles:
from openai import OpenAI
client = OpenAI()
completion = client.chat.completions.create(
model="arn:aws:bedrock:us-west-2:123456789012:application-inference-profile/your-profile-id",
messages=[{"role": "user", "content": "Hello!"}],
)
print(completion.choices[0].message.content)
Benefits of Application Inference Profiles:
For more information about creating and managing application inference profiles, see the Amazon Bedrock User Guide.
This proxy now supports Prompt Caching for Claude and Nova models, which can reduce costs by up to 90% and latency by up to 85% for workloads with repeated prompts.
Supported Models:
Enabling Prompt Caching:
You can enable prompt caching in two ways:
ENABLE_PROMPT_CACHING=true
extra_body :Python SDK:
from openai import OpenAI
client = OpenAI()
# Cache system prompts
response = client.chat.completions.create(
model="global.anthropic.claude-haiku-4-5-20251001-v1:0",
messages=[
{"role": "system", "content": "You are an expert assistant with knowledge of..."},
{"role": "user", "content": "Help me with this task"}
],
extra_body={
"prompt_caching": {"system": True}
}
)
# Check cache hit
if response.usage.prompt_tokens_details:
cached_tokens = response.usage.prompt_tokens_details.cached_tokens
print(f"Cached tokens: {cached_tokens}")
cURL:
curl $OPENAI_BASE_URL/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-d '{
"model": "global.anthropic.claude-haiku-4-5-20251001-v1:0",
"messages": [
{"role": "system", "content": "Long system prompt..."},
{"role": "user", "content": "Question"}
],
"extra_body": {
"prompt_caching": {"system": true}
}
}'
Cache Options:
"prompt_caching": {"system": true} - Cache system prompts"prompt_caching": {"messages": true} - Cache user messages"prompt_caching": {"system": true, "messages": true} - Cache bothRequirements:
For more information, see the Amazon Bedrock Prompt Caching Guide.
Make sure you use ChatOpenAI(...) instead of OpenAI(...)
# pip install langchain-openai
import os
from langchain.chains import LLMChain
from langchain.prompts import PromptTemplate
from langchain_openai import ChatOpenAI
chat = ChatOpenAI(
model="anthropic.claude-3-sonnet-20240229-v1:0",
temperature=0,
openai_api_key=os.environ['OPENAI_API_KEY'],
openai_api_base=os.environ['OPENAI_BASE_URL'],
)
template = """Question: {question}
Answer: Let's think step by step."""
prompt = PromptTemplate.from_template(template)
llm_chain = LLMChain(prompt=prompt, llm=chat)
question = "What NFL team won the Super Bowl in the year Justin Beiber was born?"
response = llm_chain.invoke(question)
print(response)
This application does not collect any of your data. Furthermore, it does not log any requests or responses by default.
API Gateway + Lambda uses API Gateway response streaming with Lambda Web Adapter to support SSE streaming without requiring a VPC. This is a cost-effective, serverless option with up to 10 minutes timeout.
ALB + Fargate provides the lowest streaming latency with no cold starts, ideal for high-throughput workloads.
Generally speaking, all regions that Amazon Bedrock supports will also be supported, if not, please raise an issue in Github.
Note that not all models are available in those regions.
You can use the Models API to get/refresh a list of supported models in the current region.
Yes. Three options:
OPENAI_API_KEY=my-local-key docker-compose up --build
src folder:
API_KEY=my-local-key AWS_REGION=us-west-2 uvicorn api.app:app --host 0.0.0.0 --port 8000
The API base url should look like http://localhost:8000/api/v1 in all cases.
Compared with direct AWS SDK calls, the proxy architecture will add some latency. The default API Gateway + Lambda deployment provides good streaming performance with Lambda response streaming.
For lowest latency on streaming responses, consider the ALB + Fargate deployment option which eliminates cold starts and provides consistent performance.
Currently, there is no plan to support SageMaker models. This may change provided there's a demand from customers.
Fine-tuned models and models with Provisioned Throughput are currently not supported. You can clone the repo and make the customization if needed.
To use the latest features, you need follow the deployment guide to redeploy the application. You can upgrade the existing CloudFormation stack to get the latest changes.
See CONTRIBUTING for more information.
This library is licensed under the MIT-0 License. See the LICENSE file.
Content type
Image
Digest
sha256:96fbf42e7…
Size
222.1 MB
Last updated
about 1 month ago
docker pull cliftonzac/bedrock-access-gateway-lambda