Model requests
Goes toAmazon Bedrock, in your account
The prompt, tools and images of each request go to Bedrock in the region you configure. The log records the model, token counts, cache hits and timing, never prompts, answers or keys.
One Rust binary that answers OpenAI Chat Completions, Responses, Completions and Embeddings requests by calling Bedrock with your own AWS account. An OpenAI SDK or an agent needs a base URL and a key, and nothing else changes.
export BGW=http://localhost:8080/api/v1curl -s $BGW/chat/completions \ -H "Authorization: Bearer $API_KEY" --json '{ "model": "us.amazon.nova-2-lite-v1:0", "messages": [{"role": "user", "content": "Reply with exactly: BEDROCK_OK"}]}' | jq -r '.choices[0].message.content'BEDROCK_OKcurl -s $BGW/responses \ -H "Authorization: Bearer $API_KEY" --json '{ "model": "us.amazon.nova-2-lite-v1:0", "input": "Reply with exactly: BEDROCK_OK"}' | jq -r '.output[-1].content[0].text'BEDROCK_OKEverything below works once the gateway has a key for its clients and AWS credentials; GPT-5.x and gpt-oss also need a Bedrock API key. Reasoning across tool calls on Chat Completions and OpenTelemetry export stay off until you turn them on.
Streaming and non-streaming chat with tools, images, JSON output and reasoning effort.
The Responses API that Codex CLI uses, streaming and non-streaming. It keeps no state between requests.
The legacy text-completion route, for editors whose edit prediction still speaks it.
Cohere, Titan and Nova embedding models behind the OpenAI embeddings request.
GET /api/v1/models lists the models and inference profiles your account can call, read from Bedrock and refreshed while the gateway runs.
Sonnet, Opus, Haiku and Fable, through Bedrock model IDs and cross-region inference profiles.
GPT-5.x, GPT-6.x and gpt-oss, under their short names such as gpt-5.6-sol or gpt-6.1-sol.
Amazon Nova, DeepSeek and any model in your account's catalog. A new model is a configuration entry, not a new release.
Cache points are placed on the tools, the system prompt and the messages without any change in the client, on the models that support it.
reasoning_effort becomes the form each model expects, from Claude's adaptive thinking to GPT's reasoning.effort.
Keeps a model's signed reasoning through a Chat Completions tool call and its result. It needs a signing key of your own.
A binary, the Docker image, ECS on Fargate behind a load balancer, or Lambda with a function URL.
The client key can come from SSM Parameter Store or Secrets Manager instead of an environment variable.
A log line per request with the model, tokens, cache hits and timing, never prompts or keys. OpenTelemetry export is a build option.
docker pull sunerpy/bedrock-gateway-rust
Or download the binary for your platform from GitHub Releases, or install it with cargo.
export API_KEY=sk-replace-with-a-private-random-key
Every client sends this key. It is yours to choose and is not an AWS credential.
docker run -p 8080:8080 -e API_KEY -e AWS_REGION=us-east-1 -e AWS_BEARER_TOKEN_BEDROCK sunerpy/bedrock-gateway-rust
AWS_BEARER_TOKEN_BEDROCK is a Bedrock API key. Without it the gateway uses the standard AWS credential chain.
http://localhost:8080/api/v1
That is the base URL for the OpenAI SDKs, Codex CLI and other OpenAI-compatible clients.
Each route takes the request an OpenAI client sends and answers in the same shape, streaming included. Features that exist only on Bedrock, such as the 1-hour prompt cache, go in an extra_body object, which the OpenAI SDKs can add to any request.
| Route | Speaks | Notes |
|---|---|---|
POST /api/v1/chat/completions | OpenAI Chat Completions | Streaming and non-streaming |
POST /api/v1/responses | OpenAI Responses | Stateless |
POST /api/v1/completions | OpenAI Completions | Legacy text completion |
POST /api/v1/embeddings | OpenAI Embeddings | Cohere, Titan and Nova |
GET /api/v1/models | Model list | Read from Bedrock |
GET /api/v1/health | Liveness | No key needed |
Every route except /health asks for your API_KEY, sent as Authorization Bearer. API_ROUTE_PREFIX changes /api/v1.
The OpenAI SDKs take the gateway as their base URL. Codex CLI takes it as a custom model provider that speaks Responses. Editors and agents with an OpenAI-compatible provider take a base URL, a key and a model ID, with no plugin and no patched client.
| Client | API | Set up with |
|---|---|---|
| OpenAI Python and Node SDKs | Chat Completions or Responses | A base URL and an API key |
| Codex CLI | Responses | A model_provider in config.toml |
| Editors and agents | Chat Completions or Completions | An OpenAI-compatible provider |
| curl and scripts | Any route | The Authorization header |
Claude, OpenAI's GPT models on Bedrock, Amazon Nova, DeepSeek and every other model in your catalog are called the same way. Each model's quirks, such as which sampling parameters it accepts or how it reasons, live in a configuration file, so a new model needs an entry there and no new release.
| Family | Model IDs | Served through |
|---|---|---|
| Claude | global.anthropic.claude-sonnet-5-5 us.anthropic.claude-opus-5-5 global.anthropic.claude-fable-5-1 | Bedrock Converse |
| GPT-5.x | gpt-5.4 gpt-5.5 gpt-5.6-sol gpt-5.6-terra gpt-5.6-luna | OpenAI-compatible endpoint |
| GPT-6.x | gpt-6.1-sol gpt-6-sol gpt-6-luna gpt-6-astra | Bedrock Converse |
| gpt-oss | gpt-oss-120b gpt-oss-20b | OpenAI-compatible endpoint |
| Amazon Nova | us.amazon.nova-2-lite-v1:0 us.amazon.nova-pro-v1:0 | Bedrock Converse |
DeepSeek and every other model in your catalog are called the same way. The model list shows the IDs your account can call in the gateway's region.
Every release has a binary for each platform below, a multi-architecture image on Docker Hub and Amazon ECR Public, and the crate on crates.io.
| Platform | Release asset | Container image |
|---|---|---|
| Linux x64Available | x86_64-unknown-linux-musl | linux/amd64 |
| Linux ARM64Available | aarch64-unknown-linux-musl | linux/arm64 |
| macOS IntelAvailable | x86_64-apple-darwin | None |
| macOS Apple SiliconAvailable | aarch64-apple-darwin | None |
| Windows x64Available | x86_64-pc-windows-msvc | None |
The Linux binaries are static, and the image is distroless with the binary and its configuration only.
bedrock-gateway sits between your clients and your AWS account. Apart from the image URLs clients send and an OpenTelemetry collector you configure, it talks to AWS only.
Goes toAmazon Bedrock, in your account
The prompt, tools and images of each request go to Bedrock in the region you configure. The log records the model, token counts, cache hits and timing, never prompts, answers or keys.
Goes toThe Bedrock control plane
At start and then on a schedule, the gateway lists the foundation models and inference profiles your account can call.
Goes toThe host in the URL
When a message carries an http or https image URL, the gateway downloads the image and sends its bytes to Bedrock. A data URL is decoded in the gateway.
Goes toSSM or Secrets Manager, if you use them
When the key comes from Parameter Store or Secrets Manager, the gateway reads it there once at start. Otherwise it stays in the environment.
curl -fsSL https://raw.githubusercontent.com/sunerpy/bedrock-gateway-rust/main/scripts/install.sh | shdocker run -p 8080:8080 -e API_KEY -e AWS_REGION=us-east-1 -e AWS_BEARER_TOKEN_BEDROCK sunerpy/bedrock-gateway-rustgh release download --repo sunerpy/bedrock-gateway-rust --pattern '*-x86_64-unknown-linux-musl.tar.gz'
tar -xzf bedrock-gateway-*-x86_64-unknown-linux-musl.tar.gzcargo install bedrock-gateway-rustThe install guide covers every platform, the image tags and building from source.
Bug reports and feature requests go to GitHub Issues. bedrock-gateway is MIT-0 licensed and is not an AWS product.