Configuration reference
kiro-provider loads configuration from a JSON file, layered with environment variables and (for serve) CLI flags. This document is the complete field reference; see the README for a quick summary.
Precedence
For every field, the effective value is the first one found, in this order:
- CLI flag —
serveonly supports--config,--host,--port,--proxy.loginsupports--config(selects the file, does not override fields) and--help. - Environment variable —
KIRO_PROVIDER_*, listed per field below. - Configuration file — JSON at the resolved config path.
- Schema default — the zod schema default in
src/config/schema.ts.
The config file path defaults to the platform configuration root plus kiro-provider/config.json (see File locations): $XDG_CONFIG_HOME/kiro-provider/config.json or ~/.config/kiro-provider/config.json on Linux/macOS, %APPDATA%\kiro-provider\config.json on Windows. accounts list|import|remove target the provider-owned local authentication store without loading gateway configuration, so accounts import does not accept --config. accounts refresh|relogin load refresh, timeout, region, proxy, and quota_recheck_concurrency settings from the selected config and require auth_source: "local". --version and self-update load neither the config nor the account store; they take a proxy from --proxy, then KIRO_PROVIDER_PROXY_URL, then HTTPS_PROXY/HTTP_PROXY, so a broken config cannot block an upgrade. An explicitly empty --proxy "" selects no proxy instead of falling through to those variables, matching serve. Bun's fetch reads HTTPS_PROXY/HTTP_PROXY on its own, so unset them (env -u HTTPS_PROXY -u HTTP_PROXY kiro-provider self-update) when a fully direct connection is required; --proxy "" alone still suppresses KIRO_PROVIDER_PROXY_URL.
Validation rules
Configuration is validated once at startup; any violation raises a ConfigLoadError naming the offending field (and, for environment values, the variable) and the process exits before binding a port.
- Empty environment variables are unset. A
KIRO_PROVIDER_*variable whose value is empty or whitespace-only is ignored, soKIRO_PROVIDER_PORT=""keeps the config-file value or the default instead of becoming0. This applies to every variable, includingKIRO_PROVIDER_PROXY_URL; to disable a file-configured proxy from the environment, setproxy_urltonullin the file or useserve --proxy "". - Integer variables must be plain decimal integers. Surrounding whitespace and an explicit sign are accepted;
0x1f90,8787.5,1e3, orNaNare rejected with a message such asInvalid environment variable KIRO_PROVIDER_PORT: expected a decimal integer, got "0x1f90". Out-of-range values are reported asport: Number must be less than or equal to 65535 (from KIRO_PROVIDER_PORT). - Unknown config-file keys are rejected. A misspelled key such as
enable_legacy_chat_completionfails withunknown key "enable_legacy_chat_completion" (did you mean "enable_legacy_chat_completions"?)rather than being silently dropped. - Loose file permissions are reported. On POSIX, when the config file is readable or writable by group or others (
mode & 0o077 != 0), startup emits aconfig_file_permissions_loosewarning with the path and current mode because the file usually containsapi_keys. Loading still succeeds; runchmod 600on the file to silence the warning. The check is skipped on Windows. - Every numeric field is a bounded integer. The accepted ranges are listed in the table below; fractional,
NaN, infinite, and out-of-range values are rejected. Millisecond fields are capped at2147483647(see Timeout limits).
Switching models and reasoning effort
New kr2_ tokens use an authenticated v4 envelope and bind replay to the exact Kiro wire model rather than the public model spelling. Base, effort, and thinking aliases of that wire model share one replay identity; changes to reasoning.effort do not invalidate it. v3 tokens retain their original AAD: the server selects the original slug from a bounded registry of known aliases, then authenticates the ciphertext, tenant, assistant output, provenance and TTL. Unknown historical model spellings fail closed.
reasoning_replay_model_switch: "compatible" authenticates a token from a different known model before omitting its incompatible opaque reasoning. It preserves visible assistant messages, historical tool names/IDs/arguments, tool results and authenticated instruction projection. An omitted block grants no account or conversation ownership. Current tool declarations alone still authorize new calls. Responses and Messages return x-kiro-reasoning-model-replay-mode: incompatible-omitted; the separate header can coexist with x-kiro-reasoning-replay-mode: conflict-omitted on newly generated output. reasoning_replay_model_omitted contains only replay_count and ordinary audit metadata. strict, including strict Responses fidelity, returns reasoning_replay_context_mismatch before inference instead.
This applies to provider-authenticated replay, not arbitrary native signatures or a native stored continuation. Cross-wire reasoning reuse is not claimed; even two models in one family omit each other's opaque reasoning until that upstream compatibility has been demonstrated.
Field reference
| Field | Type / default | Environment override | Description |
|---|---|---|---|
host | string (non-empty), default "127.0.0.1" | KIRO_PROVIDER_HOST | HTTP bind address. Leading and trailing whitespace is trimmed; an empty value is rejected. |
port | integer, 0-65535, default 8787 | KIRO_PROVIDER_PORT | HTTP listen port. 0 asks the OS for an ephemeral port (the bound address is printed at startup); serve --port 0 is rejected. Fractional and out-of-range values are rejected, and an empty KIRO_PROVIDER_PORT no longer becomes 0. |
api_keys | string[], required, non-empty after trimming | KIRO_PROVIDER_API_KEYS | Accepted Bearer keys. The environment value is a comma-separated list. An empty or whitespace-only list is rejected and the server refuses to start (fail-closed). |
enable_legacy_chat_completions | boolean, default false | KIRO_PROVIDER_ENABLE_LEGACY_CHAT_COMPLETIONS | Exposes POST /v1/chat/completions. Keep this disabled unless a client cannot use Responses or Anthropic Messages. Environment values accept true, false, 1, 0. |
protocol_projection_mode | "v3-auto" | "safe" | "native-context-safe" | "legacy-user-prefix", default "v3-auto" | KIRO_PROVIDER_PROTOCOL_PROJECTION_MODE | v3-auto uses KiroRuntime native Responses when the request is losslessly supported, then falls back for store:false, max effort, encrypted reasoning, custom/namespace tools, and Codex collaboration items. The other values preserve explicit legacy projection controls. |
responses_fidelity_mode | "compatible" (default) / "strict" | KIRO_PROVIDER_RESPONSES_FIDELITY_MODE | Compatible mode reports enumerated losses in X-Kiro-Compatibility and locally enforces only the bounded single-string-object-v1 metadata profile; strict mode rejects that approximation and other unrepresentable semantics before upstream dispatch. Governs only /v1/responses; the Anthropic Messages output_config.format profile is not gated by this key. |
responses_instruction_lift | "auto" (default) / "off" / "experimental" | KIRO_PROVIDER_RESPONSES_INSTRUCTION_LIFT | Gates native instruction lifting; auto requires complete continuation evidence. Experimental does not bypass storage or account constraints. |
responses_native_tool_bridge | "auto" (default) / "off" / "experimental" | KIRO_PROVIDER_RESPONSES_NATIVE_TOOL_BRIDGE | Gates native namespace/free-form tool adaptation; stored mappings remain readable when disabled. |
session_affinity_mode | "explicit-only" | "legacy-initial-input", default "explicit-only" | KIRO_PROVIDER_SESSION_AFFINITY_MODE | explicit-only never derives a logical session from prompt text. legacy-initial-input temporarily restores the old initial-input fingerprint heuristics without changing model-visible content. |
kiro_prompt_cache_mode | "server-auto" | "explicit-checkpoints" | "off", default "server-auto" | KIRO_PROVIDER_PROMPT_CACHE_MODE | Keeps Kiro's upstream automatic prompt caching by default. explicit-checkpoints projects only catalog-supported message/tool cache points and remains an evidence-gated optimization; off is a diagnostic control. It never changes store, reasoning effort, or model-visible history. |
auth_source | "local", default "local" | KIRO_PROVIDER_AUTH_SOURCE | Authentication authority. Only the provider-owned local store is supported. The former "opencode-shared" value is rejected at startup with a migration message since 0.7.0: copy accounts once with kiro-provider accounts import, then use "local". |
opencode_auth_db_path | string | null, default null | KIRO_PROVIDER_OPENCODE_AUTH_DB_PATH | Deprecated since 0.7.0 and ignored (a warning is logged); scheduled for removal. Point kiro-provider accounts import --from <path> at a non-default OpenCode database instead. |
proxy_url | string | null, default null | KIRO_PROVIDER_PROXY_URL | Optional global HTTP(S) proxy for all upstream egress (model requests, token refresh, quota probes, device-code login). Must be a valid http:// or https:// URL; other schemes (e.g. SOCKS) are rejected. null or an empty string means direct connections. |
default_region | AWS region enum (RegionSchema), default "us-east-1" | KIRO_PROVIDER_DEFAULT_REGION | OIDC region used by login and fallback runtime region for legacy accounts without a profile ARN. A newly discovered profile derives its runtime region from the ARN independently. Must be one of the regions listed in src/kiro/regions.ts (for example us-east-1, eu-west-1, ap-northeast-1); unknown regions are rejected at startup. |
sdk_http_keep_alive | boolean, default false | KIRO_PROVIDER_SDK_HTTP_KEEP_ALIVE | Controls Kiro model-call sockets only. The transport object stays cached in either mode; an SDK client is reused only while its access token is unchanged and is rebuilt immediately after token rotation. false uses fresh direct/proxy SDK sockets; true opts into pooling after deployment-specific validation. Token refresh and device login keep their independent transport policy. |
enforce_single_instance | boolean, default true | KIRO_PROVIDER_ENFORCE_SINGLE_INSTANCE | Acquires one service-process lock before binding the HTTP listener, keeping account/session queues and SDK pools single-owner. Disable only with independent credentials/state or an external serializer. |
instance_lock_path | string | null, default null | KIRO_PROVIDER_INSTANCE_LOCK_PATH | Optional service-lock target. null uses the platform config directory at kiro-provider/service.instance. POSIX mode is 0600; different paths deliberately create independent process domains. |
runtime_endpoint_mode | "kiro-runtime" | "legacy-q", default "kiro-runtime" | KIRO_PROVIDER_RUNTIME_ENDPOINT_MODE | Uses the live-probe-confirmed Kiro runtime endpoint by default. Successful current-runtime streams end with token-usage metadata or valid metering followed by clean EOF. legacy-q is retained for diagnosis/migration only and may not provide either authoritative witness. |
dynamic_model_catalog | boolean, default true | KIRO_PROVIDER_DYNAMIC_MODEL_CATALOG | Discovers models per usable account through Kiro management, routes only to accounts exposing the requested wire model, and uses the checked-in bounded catalog when management is unavailable. |
model_catalog_ttl_ms | integer, 1-2147483647, default 900000 (15 min) | KIRO_PROVIDER_MODEL_CATALOG_TTL_MS | Fresh lifetime of a successful per-account model catalog. |
model_catalog_stale_ttl_ms | integer, 1-2147483647, default 86400000 (24 h) | KIRO_PROVIDER_MODEL_CATALOG_STALE_TTL_MS | Maximum lifetime of a last-known-good account catalog after refresh failures. |
model_catalog_request_timeout_ms | integer, 1-2147483647, default 10000 | KIRO_PROVIDER_MODEL_CATALOG_REQUEST_TIMEOUT_MS | Deadline for one Kiro management model-list request. |
account_selection_strategy | "sticky" | "round-robin" | "lowest-usage", default "lowest-usage" | KIRO_PROVIDER_ACCOUNT_SELECTION_STRATEGY | How the gateway picks an account per request: sticky favors the same account, round-robin cycles, lowest-usage prefers the account with the most remaining quota. |
account_inference_concurrency | integer 1-10, default 10 | KIRO_PROVIDER_ACCOUNT_INFERENCE_CONCURRENCY | Maximum simultaneous inference requests per account in this process, shared by Messages and both Responses transports. Independent branches can share an account; one stateful branch stays ordered. |
rate_limit_max_retries | integer, 0-100, default 3 | KIRO_PROVIDER_RATE_LIMIT_MAX_RETRIES | Maximum retry count shared by pre-acceptance HTTP/transport failures and existing later non-stream recovery. 0 disables those retries; accepted streams are never replayed. |
rate_limit_retry_delay_ms | integer, 1-2147483647, default 5000 | KIRO_PROVIDER_RATE_LIMIT_RETRY_DELAY_MS | Base retry delay in milliseconds before a rate-limit retry. |
quota_recheck_interval_ms | integer, 1-2147483647, default 900000 (15 min) | KIRO_PROVIDER_QUOTA_RECHECK_INTERVAL_MS | Minimum wait before an exhausted account is probed again. If Kiro reports a quota reset time, the probe waits for that reset instead, capped at the larger of this interval and 24 hours. An HTTP 402, a still-exhausted snapshot, or a failed probe advances this timestamp; it does not create a model retry. |
stop_on_overage | boolean, default true | KIRO_PROVIDER_STOP_ON_OVERAGE | Treat an account whose paid overage count exceeds overage_threshold as exhausted for selection (it stays healthy and is re-admitted by the next authoritative usage sync). Set to false to knowingly keep using paid overage. When every eligible account is blocked only by overage, requests fail with 402 paid_overage_blocked. |
overage_threshold | integer, 0-1000000, default 0 | KIRO_PROVIDER_OVERAGE_THRESHOLD | Overage requests tolerated per account before stop_on_overage excludes it. |
quota_recheck_timeout_ms | integer, 1-2147483647, default 10000 | KIRO_PROVIDER_QUOTA_RECHECK_TIMEOUT_MS | Bounds both the request preflight quota-recheck batch and each started account probe. A timed-out probe keeps the account excluded and schedules the next check. |
quota_recheck_concurrency | integer, 1-32, default 4 | KIRO_PROVIDER_QUOTA_RECHECK_CONCURRENCY | Maximum number of due exhausted accounts probed concurrently. Concurrent requests join the same per-account in-flight probe. Also bounds the concurrency of accounts refresh, which uses the same usage prober as the server. |
account_maintenance_enabled | boolean, default true | KIRO_PROVIDER_ACCOUNT_MAINTENANCE_ENABLED | Enables provider-owned background token and usage maintenance. Disable only when an external operator deliberately owns that lifecycle. |
account_maintenance_interval_ms | integer, 1000-2147483647, default 60000 | KIRO_PROVIDER_ACCOUNT_MAINTENANCE_INTERVAL_MS | Interval between background maintenance passes. The first pass is scheduled shortly after startup. |
account_maintenance_timeout_ms | integer, 1000-2147483647, default 120000 | KIRO_PROVIDER_ACCOUNT_MAINTENANCE_TIMEOUT_MS | Absolute deadline for one maintenance pass across all accounts. |
account_maintenance_concurrency | integer, 1-32, default 4 | KIRO_PROVIDER_ACCOUNT_MAINTENANCE_CONCURRENCY | Maximum concurrent proactive access-token refreshes. |
usage_refresh_interval_ms | integer, 1000-2147483647, default 900000 (15 min) | KIRO_PROVIDER_USAGE_REFRESH_INTERVAL_MS | Maximum age of a normal account usage snapshot before background maintenance calls Kiro getUsageLimits. Exhausted accounts continue to use the separate quota-recheck schedule. |
max_request_iterations | integer, 1-1000, default 20 | KIRO_PROVIDER_MAX_REQUEST_ITERATIONS | Global cap on account-switching and retry-loop iterations for a single request. 0 is rejected because it would fail every request. |
request_timeout_ms | integer, 1-2147483647, default 120000 | KIRO_PROVIDER_REQUEST_TIMEOUT_MS | Absolute deadline for a request, in milliseconds. See Timeout limits for the accepted range and a known limitation. |
stream_idle_timeout_ms | integer, 1-2147483647, default 60000 | KIRO_PROVIDER_STREAM_IDLE_TIMEOUT_MS | Maximum idle interval between upstream streaming events before the stream is aborted, in milliseconds. See Timeout limits for the accepted range. |
stream_max_attempts | integer, 1-10, default 3 | KIRO_PROVIDER_STREAM_MAX_ATTEMPTS | Maximum upstream streams for non-stream pre-semantic collection recovery, including an empty-result replacement. Streaming requests stop at their accepted generation and never use this setting to replace it. |
retry_empty_completion | boolean, default true | KIRO_PROVIDER_RETRY_EMPTY_COMPLETION | For non-stream requests only, allow one same-account replacement for a witnessed completion with no text, reasoning or tools, within stream_max_attempts. Accepted empty streams are returned unchanged. |
max_request_body_bytes | integer, 1-2147483647, default 33554432 (32 MiB) | KIRO_PROVIDER_MAX_REQUEST_BODY_BYTES | Maximum accepted request body size (HTTP 413). Also bounds aggregate upstream tool arguments and identities; excess output fails with upstream_tool_arguments_too_large before a completed call. |
max_inflight_requests | integer, 1-10000, default 16 | KIRO_PROVIDER_MAX_INFLIGHT_REQUESTS | Shared request slots covering upload, parsing, queueing, generation, response consumption and cleanup. Overflow returns 503 before dispatch. |
max_inflight_request_body_bytes | integer, 1-2147483647, default 134217728 (128 MiB) | KIRO_PROVIDER_MAX_INFLIGHT_REQUEST_BODY_BYTES | Aggregate body reservation; must be at least max_request_body_bytes. Reserve the per-request maximum before reading, shrink to actual bytes after upload, and release only after complete teardown. |
token_expiry_buffer_ms | integer, 1-2147483647, default 300000 (5 min) | KIRO_PROVIDER_TOKEN_EXPIRY_BUFFER_MS | How long before actual access-token expiry the gateway proactively refreshes. |
session_affinity_ttl_ms | integer, 1-2147483647, default 86400000 (24 h) | KIRO_PROVIDER_SESSION_AFFINITY_TTL_MS | Sliding lifetime of a persisted logical-session binding. A hit extends the expiry; an expired binding is recreated using the normal account strategy. |
session_affinity_max_entries | integer, 1-1000000, default 10000 | KIRO_PROVIDER_SESSION_AFFINITY_MAX_ENTRIES | Maximum persisted session bindings. When over the limit, least-recently-seen entries are removed. |
session_affinity_stall_failover_threshold | integer, 0-100, default 2 | KIRO_PROVIDER_SESSION_AFFINITY_STALL_FAILOVER_THRESHOLD | Consecutive abnormal published-stream terminals on one binding key before the next independent request ignores the stored binding and selects a fresh account and conversation. Counted terminals are the idle timeout, an upstream error, and request_timeout_ms firing while a read of the stream had already been outstanding, with no frame arriving in it, for at least a second and for longer than the stream spent producing — the last case matters when the request deadline is shorter than stream_idle_timeout_ms, because then no idle timeout can ever fire. A client-side cancel never counts, and neither does a deadline that lands while frames are still arriving or between reads, since the stream is pulled by its consumer: a client that stops reading also stops the upstream from being asked for anything, and one frame can carry several output events, so a read can be served from the transformer's buffer without reaching the upstream at all. A terminal the provider caused itself counts as health rather than as a stall: when storing the encrypted reasoning or the output-lineage row fails, the upstream had already delivered a witnessed, fully validated answer, so the client still gets a retryable stream error but the streak is retired exactly as a completion retires it — no other account could have repaired a local keyring or database fault, and leaving the streak armed would re-bind every following request for it. The key is the explicit session-affinity key when the client supplies one, otherwise the history-lineage key, so continuations without a session header fail over too. The accounts the streak was recorded against — not whatever account the stored binding currently names — are held out of that selection unless nothing else can serve the request, so a single-account deployment is still served and a failover that died after rebinding cannot send the next request back to the account that stalled. Deciding the failover does not retire the streak: a request that finds no replacement account leaves it armed, so the next request tries again. 0 disables the failover and keeps a wedged binding sticky. A reasoning replay owner lock always wins, and no committed stream is ever replayed. |
session_affinity_stall_window_ms | integer, 1-2147483647, default 600000 | KIRO_PROVIDER_SESSION_AFFINITY_STALL_WINDOW_MS | Window in which those consecutive failures must land to count toward the threshold, in milliseconds. A longer gap restarts the streak, and so does any healthy answer on that key — a stream that terminates normally or a non-stream completion that is returned to the client — provided that answer proves the stored row healthy. On an explicit session-affinity key it always does, because the binding is rewritten to whatever account served. On the history-lineage key it only does when the answer came from the very account and conversation that row names: the row keys the previous history and is never rewritten, so whenever the request was served elsewhere — a failover, or the bound account merely being unselectable that time — the streak stays armed until the window elapses and a client re-sending the same history keeps failing over. The state is per-process, bounded by session_affinity_max_entries, and cleared by a restart. |
reasoning_replay_key_path | string | null, default null | KIRO_PROVIDER_REASONING_REPLAY_KEY_PATH | Key-file override. null uses the platform config directory and atomically creates reasoning-replay-keys.json when no environment keyring is configured. POSIX mode is forced to 0600. |
reasoning_replay_keys | string[], default [] | KIRO_PROVIDER_REASONING_REPLAY_KEYS | AES-256-GCM keyring. Environment entries are comma-separated key-id:base64url-32-byte-key values; the key ID may be omitted. The first key encrypts new records and later keys only decrypt old records. |
reasoning_replay_token_format | "portable-v2" | "database-v1", default "portable-v2" | KIRO_PROVIDER_REASONING_REPLAY_TOKEN_FORMAT | portable-v2 emits self-contained AEAD replay tokens with an authenticated absolute lifetime and mint provenance; they do not depend on SQLite payload retention. database-v1 is a rollback/diagnostic writer; both formats remain readable. |
reasoning_replay_model_switch | "compatible" | "strict", default "compatible" | KIRO_PROVIDER_REASONING_REPLAY_MODEL_SWITCH | Authenticate historical tokens before omitting only opaque reasoning from a different wire model. Same-model base/effort/thinking aliases share one identity. strict rejects cross-model replay; strict Responses fidelity also rejects it. Visible assistant/tool history is always preserved. |
reasoning_replay_account_failover | "verified" | "strict", default "verified" | KIRO_PROVIDER_REASONING_REPLAY_ACCOUNT_FAILOVER | Allows replay account migration only when the token authenticates the exact mint protocol, model, effective region, profile presence, runtime operation, and replay kind of a proven compatibility cell. Legacy or incomplete provenance stays owner-bound; strict disables all migration. |
reasoning_replay_legacy_account_failover | "strict" | "verified-current-cell", default "strict" | KIRO_PROVIDER_REASONING_REPLAY_LEGACY_ACCOUNT_FAILOVER | Operator-attested failover for authenticated database kr1_ and bounded pre-release kr2_ tokens that lack mint provenance. Requires the verified protocol/model/current-region/profile/runtime/replay-kind cell and global verified mode. Redacted reasoning, Chat hash replay, other models/regions, and safe projection remain owner-bound. Enable only to recover histories from the same deployment after confirming those missing mint dimensions. |
reasoning_replay_ttl_ms | integer, 1-2147483647, default 86400000 (24 h) | KIRO_PROVIDER_REASONING_REPLAY_TTL_MS | Sliding idle lifetime for database-backed kr1_ records and absolute lifetime for newly minted self-contained kr2_ tokens. Pre-release kr2_ envelopes without an authenticated expiry are accepted only during one persisted transition window of this length; legacy migration opt-in does not extend that window. |
reasoning_replay_max_entries | integer, 1-1000000, default 10000 | KIRO_PROVIDER_REASONING_REPLAY_MAX_ENTRIES | Maximum legacy kr1_ records and bounded pre-release kr2_ transition records. Current self-contained kr2_ payloads do not consume the replay table. |
web_search_enabled | boolean, default false | KIRO_PROVIDER_WEB_SEARCH_ENABLED | Allows new provider-executed searches for hosted Responses web_search and Messages web_search_20250305 tools. When false, a request that declares a hosted search tool is rejected with web_search_disabled before any generation, while already authenticated search history in a conversation remains readable. See Web search. |
web_search_max_calls | integer, 1-100, default 20 | KIRO_PROVIDER_WEB_SEARCH_MAX_CALLS | Searches the provider dispatches for one public request across all generation rounds. Further calls are answered with a max_uses_exceeded tool error instead of a new search. |
web_search_timeout_ms | integer, 1-2147483647, default 15000 | KIRO_PROVIDER_WEB_SEARCH_TIMEOUT_MS | Timeout for one InvokeMCP search call, in milliseconds. The remaining public request deadline is applied as well; the shorter one wins. |
web_search_max_result_bytes | integer, 1024-16777216, default 262144 | KIRO_PROVIDER_WEB_SEARCH_MAX_RESULT_BYTES | Largest InvokeMCP search response the provider reads. A larger result fails that call as unavailable; it is never truncated. |
web_search_max_history_bytes | integer, 1024-67108864, default 1048576 | KIRO_PROVIDER_WEB_SEARCH_MAX_HISTORY_BYTES | Total decrypted search snapshot bytes restored from history for one request. A history that needs more is rejected with web_search_history_too_large. |
web_search_replay_ttl_ms | integer, 1-2147483647, default 86400000 (24 h) | KIRO_PROVIDER_WEB_SEARCH_REPLAY_TTL_MS | Lifetime of an encrypted search snapshot. A snapshot referenced by a stored Responses resource lives at least as long as that resource. History that references an expired snapshot is rejected with web_search_replay_expired. |
web_search_max_cache_bytes | integer, 1048576-1099511627776, default 268435456 | KIRO_PROVIDER_WEB_SEARCH_MAX_CACHE_BYTES | Capacity of unexpired encrypted search snapshots. When a new snapshot would not fit, the search is refused with web_search_cache_full before it is dispatched; unexpired snapshots are never evicted to make room. |
effort | "low" | "medium" | "high" | "xhigh" | "max" | null, default null | KIRO_PROVIDER_EFFORT | Optional global reasoning-effort override applied to every request. null leaves effort unset unless the request specifies it. |
auto_effort_mapping | boolean, default true | KIRO_PROVIDER_AUTO_EFFORT_MAPPING | When enabled, the gateway automatically maps model-variant suffixes and request effort. Environment values accept true, false, 1, 0. |
log_level | "debug" | "info" | "warn" | "error", default "info" | KIRO_PROVIDER_LOG_LEVEL | Minimum level of the structured audit log (one JSON object per line on stderr). Levels order debug < info < warn < error; events below the threshold are dropped. warn silences per-request info events such as upstream_affinity_selected. Applied by every command that loads configuration (serve, login, `accounts refresh |
test_upstream_endpoint | string (valid URL), optional, omitted by default | KIRO_PROVIDER_TEST_UPSTREAM | Test-only. Overrides the AWS CodeWhisperer SDK endpoint used for upstream calls. Used by scripts/security-check.sh and isolated tests to point at a non-production endpoint. When set, serve prints a warning to stderr on startup. Do not set this in normal production use. |
Authentication source
auth_source: "local" is the production default. It uses ~/.config/kiro-provider/accounts.db as the sole authentication authority. Populate it either with direct device-code login or a one-time import:
kiro-provider login
# IAM Identity Center:
kiro-provider login --start-url https://example.awsapps.com/start --region us-east-1
# select explicitly when the identity exposes multiple profiles:
kiro-provider login --start-url https://example.awsapps.com/start --region us-east-1 \
--profile-arn arn:aws:codewhisperer:us-east-1:123456789012:profile/PROFILE_ID
# or, after authenticating with OpenCode plus opencode-kiro-auth:
kiro-provider accounts import
# optional non-default source:
kiro-provider accounts import --from /path/to/kiro.db
# overwrite local rows even when they are newer than the source:
kiro-provider accounts import --from /path/to/kiro.db --forcelogin completes the device-code flow, calls Kiro's List-Available-Profiles endpoint with the new access token, persists the selected profileArn, and only then calls the usage endpoint to learn the authenticated email before deriving the account ID. This is a provider-owned flow and has no runtime or installation dependency on Kiro CLI. A unique profile is selected automatically (including a unique start-URL match); use --profile-arn <arn> when multiple profiles remain. Profile discovery fails closed before the database is opened. Without an explicit ARN, the provider queries Kiro's current commercial profile control planes in us-east-1 and eu-central-1; oidcRegion keeps the token issuer region while region comes from the selected profile ARN. If the later usage lookup fails (for example, offline), the row is stored with the placeholder email builder-id@aws.amazon.com, a warning is printed, and a later accounts refresh --all or accounts relogin fills in the real identity. When the identity is verified, logging in again as the same person (same email, start URL, and profile) updates the existing row in place and removes older duplicate rows instead of inserting a second account. Each SSO OIDC request (client registration, device authorization, token poll) has a 30 s deadline, and transient network failures during token polling are retried until the device code expires.
Import copies active account credentials and usage into the provider database; it does not retain a live link, shared lock, or runtime dependency on OpenCode. A source row is skipped when the local copy already has a later access-token expiry or usage sync (kiro-provider has refreshed it since the last import); pass --force to overwrite anyway. accounts import does not read gateway configuration and therefore has no --config option. After import, kiro-provider independently:
- refreshes near-expiry access tokens and persists them before use;
- rebuilds the credential-bound SDK client when the access token changes while preserving the account transport;
- refreshes stale normal usage snapshots in the background;
- excludes exhausted accounts before token refresh or SDK construction;
- probes exhausted accounts only when their persisted recheck time is due and returns them to selection only after an authoritative non-exhausted snapshot;
- marks permanently invalid refresh credentials unhealthy instead of retrying them in model loops;
- deduplicates per-account probes and bounds maintenance concurrency.
The local account store can be operated without OpenCode:
kiro-provider accounts list
kiro-provider accounts list --details
kiro-provider accounts list --json
kiro-provider accounts list --sort availability
kiro-provider accounts list --sort usage --order desc
kiro-provider accounts refresh --all
kiro-provider accounts refresh <id|email> --json
kiro-provider accounts relogin <id|email>
kiro-provider accounts remove <id|email>The default list is an aligned summary. --details and --json expose the stable internal ID needed to disambiguate duplicate emails, but never include access tokens, refresh tokens, or client secrets. Email identifiers are case-insensitive and accepted only when exactly one row matches.
Rows are sorted by email ascending unless --sort <field> says otherwise, and --order asc|desc flips the direction. The supported fields are email (default), id, auth, region, health, availability, usage, overage, last-sync, last-used, token-expires, and generation; LAST_USED and last_used are accepted as aliases. availability ranks accounts from most to least usable, and usage compares the used/limit ratio rather than the raw counter so accounts with different quotas stay comparable. Accounts with no value for the chosen column (an unknown quota, a row never synced) sort last in both directions, and equal keys fall back to email then internal ID, so the order is stable. All three output modes honour the sort.
Manual refresh always calls Kiro's authoritative usage endpoint, including for fresh or currently exhausted rows. It refreshes an access token only when it is near expiry or after one invalid-bearer response. A partial failure produces a non-zero exit code and a per-account result; --json is suitable for monitoring. Background maintenance remains responsible for automatic near-expiry token renewal, normal usage refresh, and periodic quota recovery.
accounts relogin resolves the target before opening device authorization, discovers a profile when the legacy row has none (or validates an explicit --profile-arn), then verifies the authenticated Kiro usage email before writing credentials. It preserves the selected internal account ID so existing session affinity can continue to reference the same account. An account ID already bound to a profile cannot be re-logged into a different profile; add that profile with a fresh login instead. accounts remove prompts by default; --yes is required for non-interactive deletion, which also removes that account's persisted affinity, output-lineage, and reasoning-replay rows.
Run only one authentication owner for an imported rotating refresh token. Continuing to use the same imported account through an independently running OpenCode plugin can race token rotation; re-import only as an intentional operator action.
The former auth_source: "opencode-shared" mode, which read OpenCode's live database and shared its refresh lock, was removed in 0.7.0. A configuration that still selects it fails at startup with a migration message: run kiro-provider accounts import [--from <path>] once, then set auth_source to "local" or delete the key. opencode_auth_db_path is ignored and only logs a deprecation warning until it is removed.
Proxy
proxy_url is the single knob that redirects every kind of upstream traffic through one HTTP(S) proxy:
- Model requests (chat completions).
- Access-token refresh.
- Authoritative quota rechecks and periodic usage refreshes (
getUsageLimits). - Device-code login into the provider-owned local store (
login).
A proxy may be required when a network reaches some model families directly but not others — for example, GPT requests succeed direct while Claude requests need an approved proxy egress and otherwise return HTTP 401/403.
Setting it, in order of precedence for serve:
--proxy <url>(CLI flag,serveonly).KIRO_PROVIDER_PROXY_URL(environment variable).proxy_urlin the config file.
login has no --proxy flag, so device-code login picks up the environment variable or config-file value. One-time import is local SQLite work and does not contact the network.
KIRO_PROVIDER_PROXY_URL=http://proxy.example.com:8080 \
./dist/kiro-provider serve
./dist/kiro-provider serve --proxy https://proxy.example.com:8443Only http:// and https:// schemes are accepted; an invalid or non-HTTP(S) URL fails config validation at startup.
Protocol exposure
POST /v1/responsesis always enabled for OpenAI Responses clients. A particular Codex version is supported only when its standard request stays within the documented verified subset.POST /v1/messagesandPOST /v1/messages/count_tokensare always enabled for Anthropic Messages clients. A particular Claude Code version is supported only when its standard request stays within that subset.POST /v1/chat/completionsreturnslegacy_chat_completions_disabledunlessenable_legacy_chat_completionsis explicitly set totrue.- Authenticated
GET /readyreturns HTTP 200 only when the configured authentication source is readable, at least one active account exists, the provider database is writable, the reasoning keyring is available, and all key IDs used by unexpired replay rows are present. Itsmodel_catalogobject reports whether model metadata currently comes from live, stale, static-fallback, or disabled discovery.
protocol_projection_mode: "v3-auto" is the production default. Ordinary Responses requests use KiroRuntime CreateResponse, including its native instructions field. Request shapes that need store:false, max effort, encrypted reasoning replay, custom/namespace tools, serial-tool compatibility, or Codex collaboration items use the stateless canonical pipeline. Compatible mode also routes the bounded single-string-object-v1 text metadata output profile there for local validation and buffered JSON projection; strict mode rejects it before dispatch. The Anthropic Messages output_config.format variant of that profile is not affected by responses_fidelity_mode. This does not enable arbitrary Structured Outputs.
Explicit safe still applies to the older GenerateAssistantResponse path. Live GPT and Claude probes showed that Kiro accepts a valid required-label additionalContext shape but does not preserve its instruction content or instruction-over-user priority. safe therefore returns unsupported_instruction_projection for instruction roles. The stricter native-context-safe mode uses systemPrompt only when Kiro advertises the private feature; it is not currently enabled for the tested account.
On the stateless path, v3-auto uses the verified native systemPrompt field when the account and instruction shape support it. Otherwise it preserves each instruction's turn boundary through text projection, also used by legacy-user-prefix: a leading or intermediate instruction prefixes the immediately following user/tool turn, or becomes its own user turn before an assistant. A trailing instruction stays on the current user/tool turn without moving its tool results or attachments, or becomes the actual current input after an assistant result. Only original instruction text and \n\n separators are used; no assistant acknowledgement or generic follow-up prompt is fabricated. This fallback preserves timing but cannot promise native role priority. safe and native-context-safe retain their stricter rejection rules.
Explicit legacy-user-prefix emits a content-free startup warning. It does not restore message merging, repeated-content collapse, trailing-character deletion, synthetic tool prose, or any other rewrite. The mode remains deprecated, but it has no fixed removal version: removal requires a protocol-faithful native Kiro instruction channel or completed migration of affected clients.
The exact accepted/rejected API subset is documented in PROTOCOL_COMPATIBILITY.md.
Anthropic POST /v1/messages/count_tokens uses the provider's fallback estimator because Kiro does not expose a standalone tokenizer. Its successful response includes x-kiro-token-count-mode: estimate. OpenAI POST /v1/responses/input_tokens is recognized separately and returns typed HTTP 501 unsupported_endpoint.
Kiro runtime and model catalog
Production requests default to runtime_endpoint_mode: "kiro-runtime". Live A/B capture showed that runtime.<region>.kiro.dev emits authoritative completion witnesses required to distinguish a complete response from a clean but truncated stream: token-usage metadata completes immediately, while valid metering completes only when followed by clean EOF. The old SDK q endpoint may provide neither witness, so legacy-q is an explicit diagnostic/migration option rather than an automatic fallback.
With dynamic_model_catalog: true, the provider calls Kiro's management ListAvailableModels operation using the selected account's current token and the truthful AI_EDITOR origin. Responses are cached per account, concurrent refreshes are deduplicated, and a last-known-good response may be used through model_catalog_stale_ttl_ms. A requested model is sent only to an account whose live/stale catalog contains its exact wire ID. If management is temporarily unreachable and no cached response exists, the checked-in catalog is used as a bounded fallback; unknown models are still rejected before SDK generation.
Session affinity and connection reuse
session_affinity_mode: "explicit-only" is the production default. It does not hash input text, messages, tool arguments, or any other model-visible content to guess whether two requests belong to one conversation. It accepts these explicit sources:
- Responses, in priority order: standard
metadata.zuno_session_id, standardmetadata.kiro_provider_session_id, compatibilityclient_metadata.thread_id|session_id|conversation_id, thenprompt_cache_key. - Chat Completions:
prompt_cache_keyonly. - Anthropic Messages:
x-claude-code-session-id, withx-claude-code-agent-idseparating a subagent's execution branch from its parent and siblings. The agent ID is scoped to the authenticated tenant and family session; it never authorizes reasoning replay. Identity headers are trimmed and bounded to 256 characters. Clients without a valid agent header retain the session-level binding.
With an explicit key, the provider stores only its tenant-isolated hash, the selected account ID, Kiro conversationId, and timestamps—not the original session value or prompt. One execution branch is serialized in-process. Different execution branches can execute concurrently across accounts or within one account's configured capacity. Transport objects are cached per account, and SDK clients are cached only while the account access token is unchanged. A token refresh rebuilds the SDK client against the new immutable credential while preserving the transport.
For requests without a hard replay owner, account selection first restricts the eligible pool to free capacity. The configured strategy and soft affinity break ties within that pool. Selection and reservation occur without an asynchronous gap. When all eligible accounts are busy, the request waits for any of them to become free instead of queuing behind one preselected account. Waiting requests are admitted in arrival order when their eligible capacity is available; a request restricted to a busy owner does not block unrelated requests from using another free account. Messages, stateless Responses, and native Responses share this capacity pool. A busy soft affinity may therefore move to another eligible account and a fresh conversation, trading cache reuse for concurrency. Native continuation and owner-bound reasoning keep their account/region/profile restrictions. account_inference_concurrency defaults to 10, accepts integers from 1 to 10, and is shared by all three inference transports in the process. The least occupied eligible accounts are used first, so idle accounts participate before a busy account receives more work. At the configured limit, requests wait for any eligible slot to be released. Set it to 1 to retain the previous capacity. Native Responses uses the same explicit branch lock as stateless Responses; without an explicit key, a stored continuation serializes on its tenant-scoped previous response or reasoning origin. Identity and historical tool authorization checks remain independent of capacity.
Ten is the upper bound of the tested configuration, not a published Kiro service quota. Real account/model rate limits still apply. An occupied pool queues requests and does not make an in-flight stream portable or replayable.
request_queue_wait reports queue: "session" | "capacity" | "account", duration_ms, and acquired/unavailable/aborted outcome. Capacity admission includes waiting for an eligible free account and chooser overhead; account_selection_completed separately measures the selection step. upstream_attempt_started.preparation_ms measures preparation after obtaining an account lease, only on its first dispatch. upstream_headers_received.wait_ms and upstream_first_frame.wait_ms are measured from that attempt's dispatch. Stream duration is measured against its terminal event. These are separate from pure model inference time.
Accounts with overage_count > 0, or with a positive known limit where used_count >= limit_count, are excluded before refresh and SDK construction. An upstream HTTP 402 marks the account exhausted and excludes it from the current request without retrying it. A 401 or invalid-bearer 403 gets at most one forced refresh per account; if authentication still fails, that account is excluded for the rest of the request and the final response preserves HTTP 401/403 instead of becoming max_request_iterations HTTP 500. A rate-limit, quota, authentication, or unhealthy-account failover rebinds the session to the replacement account and rotates the Kiro conversation ID.
Standard clients that cannot send a stable metadata key still get a safe continuation path when they resend full history. After a completed assistant or tool output, the provider stores only a tenant-isolated fingerprint of that exact output lineage with its account and Kiro conversation. A later request whose latest assistant output matches that lineage reuses both. The first turn, a request without assistant history, or unmatched history starts a fresh Kiro conversation. User text, tool arguments, and initial prompts are never fingerprinted to guess identity.
Account selection and account-scoped SDK/transport object reuse still apply when neither explicit nor history lineage is available. The Kiro SDK's direct/proxy agents use fresh sockets by default; sdk_http_keep_alive: true is an explicit opt-in for deployments that have validated pooled socket behavior.
legacy-initial-input is migration-only. It restores the previous Responses initial-input, Chat user/initial-turn, and Anthropic metadata.user_id/initial-turn heuristics. Startup emits a structured content-free warning. This mode changes only routing affinity; it does not prepend, merge, delete, or otherwise modify model-visible request content.
This maximizes logical-session and SDK-object reuse without tying a session to one physical TCP socket. Even when keep-alive is enabled, the Node/Smithy agent, proxy, remote server, idle timeout, and network can open a new socket. enforce_single_instance: true is therefore the production default: a second provider using the same service lock fails before binding, so account/session queues and socket pools cannot silently split across processes. If this guard is disabled or different lock paths are used, queue serialization becomes per-process; that is safe only with independent credentials/state or an external cross-process serializer.
The gateway mirrors stored OpenAI response objects in its provider-owned SQLite database for 30 days, bounded to 10,000 entries. Tenant-local previous_response_id, retrieve, delete, cancel, and input-items pagination use this mirror. A deleted, expired, unknown, or cross-tenant ID returns response_not_found. Responses conversation objects remain unsupported. Deleting the local mirror blocks gateway continuation but does not prove that Kiro physically deleted its upstream response state.
Encrypted reasoning replay
When Kiro emits signed text (including empty text with a non-empty signature) or redacted reasoning, the default portable-v2 writer returns a self-contained kr2_... AEAD value in Responses reasoning.encrypted_content. The encrypted and authenticated envelope binds tenant, model, complete assistant-output fingerprint, origin account/conversation, mint protocol, effective region, source profile, runtime protocol, upstream operation, issue time, absolute expiry, and key ID. It does not depend on SQLite payload retention, is padded to 1 KiB buckets, and has a 4 MiB fail-closed wire limit. Unsigned text, conflicting signatures, or mixed text/redacted events still produce no token.
database-v1 remains available for rollback and legacy kr1_... tokens remain readable. Their database rows store only token/fingerprint hashes plus AES-256-GCM ciphertext; successful reads are batched, rotate to the active key, and renew the configured idle TTL. Historical kr1_ rows do not authenticate mint protocol/region/profile/operation, so they remain owner-bound by default. Pre-release kr2_ envelopes that lack those fields are also owner-bound by default and accepted only until one persisted compatibility cutoff established when this version first opens the database.
reasoning_replay_account_failover: "verified" permits provenance-authenticated signed reasoning_text only in the following verified cells. All require GenerateAssistantResponse, a profile, and effective region us-east-1:
| Public protocol | Model | Upstream runtime |
|---|---|---|
| Responses | GPT-5.6 Sol | KiroRuntime |
| Responses | GPT-5.6 Sol | CodeWhisperer |
| Anthropic Messages | Claude Sonnet 5 | KiroRuntime |
| Anthropic Messages | Claude Opus 5 | KiroRuntime |
| Anthropic Messages | Claude Opus 5.5 | KiroRuntime; parity, probe owed |
| Anthropic Messages | Claude Fable 5.1 | KiroRuntime; same mint profile only |
Every other cell in that table rests on its own live A→B run. The Claude Opus 5.5 cell does not: it was admitted by explicit maintainer decision on family parity with the Opus 5 cell, because the two models share the same catalog schema and emit the same signed reasoning envelopes. Its dedicated probe is still owed and needs two un-rate-limited accounts in the same region:
bun run scripts/probe-replay-portability.ts --confirm \
--model claude-opus-5.5 --effort maxUntil that run is recorded, treat this one cell as an assumption rather than evidence. If a migrated Opus 5.5 signature is ever rejected upstream, the failure surfaces as an error on that replay; set reasoning_replay_account_failover: "strict", or pin KIROCLAUDE_OPUS_MODEL=claude-opus-5[1m], to fall back to owner-bound replay.
The request protocol and projected runtime operation must match the mint envelope, and the target account must resolve to the same effective region with a profile. The CodeWhisperer cell covers the stateless Responses route, where v3-auto selects legacy-user-prefix for encrypted reasoning history. Redacted reasoning, legacy tokens without the opt-in below, Terra, Luna, other regions, and every unlisted combination remain owner-bound. strict disables all migration.
Fable migration additionally requires the exact authenticated mint profile ARN on the target account. Its signed blocks require an unchanged historical system/tools/message prefix. Migration does not authorize rewriting that prefix, removing signed reasoning, or sharing a cache across accounts.
New replay records authenticate the instruction-projection version. When a pre-fix Fable kr2_ record proves a Messages/KiroRuntime origin, the gateway can preserve the historical forced prefix used by that provider version. Only the signed historical prefix is retained; later system/developer input stays at its original turn. Subsequent records carry the frozen prefix boundary, including when a client removes the oldest thinking block. reasoning_replay_projection_compatibility logs only the model, protocol, and prefix-message count. New sessions do not acquire a synthetic acknowledgement. Both portable and database writers retain projection metadata across restart; an old database token without mint provenance is not used to guess that origin.
An explicitly declared Claude Bash normalization context may also be authenticated in replay records. x-kiro-client-normalization: claude-code-bash-v1 requires a 64-character x-kiro-working-directory-hash (SHA-256 of kiro-provider-working-directory-v1\0 followed by the directory's UTF-8 bytes). It recognizes only a leading literal cd to that exact directory followed by &&; command suffixes, other arguments, and tool identities remain bound. Both portable and database records require the same normalization context when replayed. Untagged records keep their original strict fingerprint contract. For Anthropic Messages requests with this complete normalization context, a user message containing only direct text and image blocks may be split into at most 16 consecutive text/image runs and projected as ordered Kiro user turns. This preserves Claude Code's client-injected text after direct images without flattening or reordering the blocks. Other mixed content and requests without the normalization context retain the strict unsupported_content_block_projection error. This metadata changes neither account eligibility nor current tool authorization.
reasoning_replay_legacy_account_failover: "verified-current-cell" is an explicit recovery switch for database kr1_ tokens and pre-release kr2_ envelopes created by this same deployment. Both must still pass tenant, model, complete-output, owner, content, key, and expiry checks: kr1_ uses its database idle TTL, and pre-release kr2_ uses its persisted transition cutoff. Neither format authenticates the mint protocol/region/profile/operation. Enabling this switch attests that those missing dimensions match the current owner row and request; the gateway cannot reconstruct the original mint provenance. Admission remains limited to the verified runtime/profile cells above except Fable, which requires authenticated mint provenance, with reasoning_replay_account_failover: "verified". Redacted reasoning, Chat hash replay, safe projection, and every unlisted cell remain owner-bound. The default is strict; global strict also disables this recovery switch. Eligible histories can move to a healthy account when their owner becomes unavailable, preserving signed reasoning and complete output. A request that carries verified portable replay prefers its bound origin account for as long as that account is selectable and below account_inference_concurrency, even when another account has a shorter queue; it migrates only when the origin is unavailable (quota exhausted, rate limited, unhealthy, model-ineligible, or quarantined) or already at its concurrency ceiling. The migrated session-affinity binding is committed only after Kiro accepts the migrated request (the first streamed event, or a completed non-streaming collection), so a rejected migration leaves the previous binding untouched. If Kiro rejects the migrated attempt with 400 REQUEST_BODY_INVALID or an invalid reasoning signature before any output, the provider makes exactly one fallback attempt on the origin account and its original conversation when the origin is still selectable; otherwise it returns 400 reasoning_replay_migration_rejected. There is no further retry, and an already accepted stream is never retried on another account.
Restoring an old session does not rewrite its historical kr1_ tokens. With the default portable-v2 writer, subsequent outputs use kr2_; the restored history can contain both formats and multiple verified owner accounts.
A strict owner failure is reported as a typed quota, rate-limit, re-authentication, refresh, health, or model error rather than one generic 503. In particular, reasoning_replay_account_reauthentication_required is distinct from the retryable reasoning_replay_account_refresh_failed.
Key configuration precedence is:
- Non-empty
KIRO_PROVIDER_REASONING_REPLAY_KEYS/reasoning_replay_keys. reasoning_replay_key_path.- The platform default config path.
Example environment keyring (generate keys with a cryptographically secure random source; do not copy this placeholder):
export KIRO_PROVIDER_REASONING_REPLAY_KEYS='2026-08:<base64url-32-byte-key>,2026-07:<old-key>'The first entry is active for encryption. Keep an old key for at least the full configured replay TTL after the last token it minted, and until the persisted pre-release compatibility cutoff has passed. Removing it earlier explicitly revokes those tokens. If an unexpired kr1_ or transition record references a missing key, service construction fails instead of silently breaking sessions. Logs never contain key material, raw replay tokens, signatures, reasoning text, redacted bytes, or request prompt content.
Web search
With web_search_enabled: true the provider executes hosted web search itself through KiroRuntime InvokeMCP web_search, using the selected account's own credentials and profile and an honest kiro-provider/<version> User-Agent. It never starts, reads or borrows the Kiro CLI. Searches are real-time only.
| Surface | Accepted declaration | Rejected before any generation or search |
|---|---|---|
| Responses | web_search or web_search_2025_08_26; external_web_access omitted or true; search_context_size; filters.allowed_domains or filters.blocked_domains | external_web_access: false (cached search), web_search_preview, user_location, search_content_types, return_token_budget, include: ["web_search_call.results"], both filter lists together |
| Messages | web_search_20250305 named web_search; max_uses; allowed_domains or blocked_domains (optional path); allowed_callers omitted or ["direct"] | later tool versions (dynamic filtering, response_inclusion), user_location, code-execution callers, both domain lists together |
Only verified capability cells execute searches: responses and anthropic-messages with gpt-5.6-sol or claude-opus-5.5 on accounts whose profile region is us-east-1. Model aliases and effort suffixes share their model's cell. Other models fail with unsupported_web_search_model; when no usable account is in a verified region the request fails the same way. Forced or named hosted-tool choice stays unsupported; tool_choice auto and none work, and none never runs a search. Fast mode remains a typed rejection (service_tier other than auto/default).
One public request owns one account lease and one Kiro conversation for all of its generations. Kiro receives the backend's own tools/list declaration unchanged, and each search runs once with a timeout of min(web_search_timeout_ms, remaining request time); searches are never retried automatically. A request dispatches at most web_search_max_calls searches (Messages max_uses lowers that) and runs at most web_search_max_calls + 1 generations. Calls over the search budget reach the model as max_uses_exceeded errors.
- Responses always uses the stateless lane. Each executed search is a
web_search_callitem with asearchaction; itssources(the URLs handed to the model) appear only withinclude: ["web_search_call.action.sources"]. A markdown link in the answer whose URL is a retrieved source becomes aurl_citationannotation; offsets count Unicode code points. In a mixed tool group the searches complete first and the function calls are then returned to the client; the next request must answer every one of those calls before any other input item (invalid_tool_historyotherwise). Replay rebuilds such a group as the one turn Kiro produced: its text, searches and function calls in Kiro's order, then all of their results. Reaching the generation limit ends the response withweb_search_iteration_limit.search_context_size: "low"keeps the first 3 filtered sources;mediumandhighkeep all of them (Kiro returns at most 10). A client function namedweb_searchkeeps its name and travels to Kiro under a private alias. - Messages returns
server_tool_use,web_search_tool_resultand text blocks split at cited links withweb_search_result_locationcitations.cited_textis a verbatim excerpt of at most 150 characters within the source's verbatim word limit. Search errors usetoo_many_requests,invalid_tool_input,max_uses_exceeded,query_too_longorunavailable. A group that mixes a search with client tools ends withstop_reason: "tool_use"and runs the search only when the next request returns a result for every client call of the group in a user message that contains nothing buttool_resultblocks; a missing result isinvalid_tool_historyand nothing runs. That response starts with the search result.pause_turnhappens only at a stable checkpoint, before any search is dispatched, when the remaining request time is shorter than one search timeout or the generation limit is reached. Continue it by sending the assistant message back unchanged with the same tools.
Search history is replayed from encrypted, tenant-bound snapshots in the provider's Accounts DB (web_search_replay), sealed with subkeys derived from the reasoning replay keyring. The snapshot, not the client, supplies the exact tool call and result Kiro saw; Messages encrypted_content and encrypted_index are provider envelopes that must be returned unchanged. Unlike Anthropic's self-contained values, they authenticate a snapshot that lives for web_search_replay_ttl_ms (or at least as long as the stored Responses resource that references it), so older history is rejected with web_search_replay_expired. Tampered, foreign-tenant, unknown or unauthenticated history fails with web_search_replay_invalid, web_search_replay_not_found or web_search_replay_key_unavailable. A search that was executing when the process stopped becomes uncertain at the next start and is never run again (web_search_replay_uncertain). A continuation claims all of its pending searches together before any of them runs; when another request already holds one of them, none runs and the request fails with web_search_replay_pending. Each generation keeps its own signed reasoning; the existing conflict rules apply within one generation, and model or effort switches keep the visible search history.
Search history binds its request to the account, region, profile and Kiro conversation that recorded it, whether its searches completed, failed or are still pending. When that owner cannot serve the request it fails with web_search_replay_owner_unavailable instead of moving to another account, and a history whose searches were recorded by different owners is rejected with web_search_replay_invalid. If a stored Responses resource cannot extend its history's snapshots, the request fails with web_search_store_unavailable rather than storing a response that would outlive its history. Restored history, including on /v1/messages/count_tokens, is charged against max_inflight_request_body_bytes. A client that disconnects or cancels the stream aborts the in-flight search and generation; a search whose outcome was not recorded becomes uncertain before the request releases its account lease.
Turning web_search_enabled off only blocks new searches: requests without a hosted declaration still replay existing search history. The snapshot table is additive; an older binary opens the database but cannot serve hosted search history. Before rolling back to a release without web search, remove every web_search_* key from the configuration: older releases reject unknown keys at startup.
File locations
All provider-owned files live under one per-user configuration root:
| Platform | Root | Files |
|---|---|---|
| Linux / macOS | $XDG_CONFIG_HOME or ~/.config | kiro-provider/config.json, kiro-provider/accounts.db, kiro-provider/service.instance, kiro-provider/reasoning-replay-keys.json; OpenCode import source opencode/kiro.db |
| Windows | %APPDATA% or %USERPROFILE%\AppData\Roaming | kiro-provider\config.json, kiro-provider\accounts.db, kiro-provider\service.instance, kiro-provider\reasoning-replay-keys.json; OpenCode import source opencode\kiro.db |
An empty XDG_CONFIG_HOME or APPDATA is treated as unset. Before v0.6 the config file alone was always read from ~/.config/kiro-provider/config.json, even on Windows. On Windows, when %APPDATA%\kiro-provider\config.json does not exist but the legacy ~/.config/kiro-provider/config.json does, the legacy file is still used; move it to %APPDATA% to complete the migration. Use --config <path> to bypass the default lookup entirely.
Timeout limits
request_timeout_ms and stream_idle_timeout_ms both accept integers from 1 through 2147483647 (2³¹−1) milliseconds. Fractional values, 0, negative numbers, NaN, and anything above 2147483647 are rejected at config validation. This bound comes from the JS/Bun setTimeout 32-bit-safe timer range, not from an arbitrary product cap — values above it would silently fire far earlier than configured instead of failing loudly.
request_timeout_ms bounds the gateway's own application-layer cleanup: the pipeline queue lock, the deadline timer, request-scoped idle-timeout lease restoration, and the SDK iterator/reader cleanup attempt are all released deterministically once the deadline fires, regardless of what the client is doing. It does not guarantee that the underlying TCP socket's file descriptor or its outbound Send-Q closes within that same window. On Bun 1.3.14, a client that has paused reading under write backpressure can leave the connection in ESTABLISHED/FIN-WAIT-1 with a nonzero Send-Q even after the gateway has finished its own cleanup — this is a platform-level limitation of Bun's current transport, not something request_timeout_ms can bound on its own.
If you need a hard upper bound on connection lifetime regardless of client read behavior, terminate the connection at a reverse proxy with its own send/write timeout in front of the gateway, or track future Bun releases for stronger transport-level guarantees.
Example config file
{
"host": "127.0.0.1",
"port": 8787,
"api_keys": ["sk-REPLACE-ME"],
"enable_legacy_chat_completions": false,
"protocol_projection_mode": "v3-auto",
"session_affinity_mode": "explicit-only",
"auth_source": "local",
"opencode_auth_db_path": null,
"proxy_url": null,
"default_region": "us-east-1",
"enforce_single_instance": true,
"instance_lock_path": null,
"runtime_endpoint_mode": "kiro-runtime",
"dynamic_model_catalog": true,
"model_catalog_ttl_ms": 900000,
"model_catalog_stale_ttl_ms": 86400000,
"model_catalog_request_timeout_ms": 10000,
"account_selection_strategy": "lowest-usage",
"account_inference_concurrency": 10,
"rate_limit_max_retries": 3,
"sdk_http_keep_alive": false,
"rate_limit_retry_delay_ms": 5000,
"quota_recheck_interval_ms": 900000,
"quota_recheck_timeout_ms": 10000,
"quota_recheck_concurrency": 4,
"account_maintenance_enabled": true,
"account_maintenance_interval_ms": 60000,
"account_maintenance_timeout_ms": 120000,
"account_maintenance_concurrency": 4,
"usage_refresh_interval_ms": 900000,
"max_request_iterations": 20,
"request_timeout_ms": 120000,
"stream_idle_timeout_ms": 60000,
"max_request_body_bytes": 33554432,
"max_inflight_requests": 16,
"max_inflight_request_body_bytes": 134217728,
"token_expiry_buffer_ms": 300000,
"session_affinity_ttl_ms": 86400000,
"session_affinity_max_entries": 10000,
"session_affinity_stall_failover_threshold": 2,
"session_affinity_stall_window_ms": 600000,
"reasoning_replay_key_path": null,
"reasoning_replay_keys": [],
"reasoning_replay_token_format": "portable-v2",
"reasoning_replay_model_switch": "compatible",
"reasoning_replay_account_failover": "verified",
"reasoning_replay_legacy_account_failover": "strict",
"reasoning_replay_ttl_ms": 86400000,
"reasoning_replay_max_entries": 10000,
"web_search_enabled": false,
"web_search_max_calls": 20,
"web_search_timeout_ms": 15000,
"web_search_max_result_bytes": 262144,
"web_search_max_history_bytes": 1048576,
"web_search_replay_ttl_ms": 86400000,
"web_search_max_cache_bytes": 268435456,
"effort": null,
"auto_effort_mapping": true,
"log_level": "info"
}This mirrors config.example.json at the repo root. Replace sk-REPLACE-ME with a private, randomly generated key before deploying; an empty api_keys list is rejected at startup.
Global request admission
Messages, Responses, enabled Chat Completions and Messages count_tokens share request-count and body-byte budgets after authentication and before body reads. Overflow returns 503 with Retry-After: 1, cancels the body and never dispatches to Kiro. Health, readiness and model-catalog requests do not consume this budget.
max_inflight_requests defaults to 16 and max_inflight_request_body_bytes to 128 MiB. Each upload first reserves max_request_body_bytes, regardless of Content-Length. Once uploaded, its reservation shrinks to actual bytes and stays held until both the response and upstream cleanup finish. With the default 32 MiB per-request limit this admits at most 4 unfinished uploads at once; parsed small requests remain bounded by the 16 total slots. Increasing the per-request limit requires an aggregate budget at least as large; inconsistent configuration fails to load.
The per-request default is 32 MiB (33554432 bytes). This is an HTTP JSON-body limit, independent of a model's token context window. Inline base64 images, tool-result history and reasoning envelopes consume body bytes; an image-heavy Codex session can exceed a byte limit while the context indicator is well below 100%. Bun can reject an oversized upload before the application handler runs, returning HTTP 413 and closing the connection. Codex may surface that closure as stream disconnected before completion: error sending request.
Existing JSON files that explicitly set max_request_body_bytes to 10485760 keep their 10 MiB limit after an upgrade. Set it to 33554432 (or set KIRO_PROVIDER_MAX_REQUEST_BODY_BYTES=33554432) and restart the gateway after active requests finish to adopt the new limit. The 128 MiB aggregate budget is unchanged; do not remove body limits to accommodate an ever-growing history.
This is not a JavaScript heap limit: parsing, strings, replay and SDK serialization can amplify resident memory. Size the service from measured bounded-load peaks and limit concurrent builds and full test runs. The separate per-account account_inference_concurrency still applies; admission does not change account eligibility, tenant isolation or signed replay ownership.