Skip to main content

Changing a Model

Changing a model in PR-Agent​

See here for a list of supported models in PR-Agent. The current default model and fallback tier are defined in the configuration file. To use different models, edit these fields in that file:

[config]
model = "..."
fallback_models = ["..."]

To see which of these models actually handled a given PR, enable config.output_run_details (see Additional configurations). To send small pull requests to a cheaper model, see Routing small pull requests to a cheaper model.

For models and environments not from OpenAI, you might need to provide additional keys and other parameters. You can give parameters via a configuration file, or from environment variables.

Model-specific environment variables

See litellm documentation for the environment variables needed per model, as they may vary and change over time. Our documentation per-model may not always be up-to-date with the latest changes. Failing to set the needed keys of a specific model will usually result in litellm not identifying the model type, and failing to utilize it.

Credential isolation boundary

PR-Agent captures request settings and supported provider environment values per handler. Deployment-owned LiteLLM secret managers are outside this request-isolation boundary; their credentials are not snapshotted by PR-Agent. Keep process environment variables and LiteLLM globals stable while requests run. The handler does not isolate arbitrary changes made by embedding applications during a request. Use separate processes when workloads require different deployment-owned credential sources or mutable global authentication and routing state.

Process-wide routing fallbacks such as litellm.api_base, litellm.api_version, litellm.organization, litellm.vertex_project, and litellm.vertex_location can be rejected even when unchanged. Use the corresponding PR-Agent settings (OPENAI.API_BASE, OPENAI.API_VERSION, OPENAI.ORG, VERTEXAI.VERTEX_PROJECT, and VERTEXAI.VERTEX_LOCATION) or supported provider environment variables, and clear the corresponding LiteLLM globals even when they match the intended routing. Use LITELLM.EXTRA_HEADERS instead of litellm.headers.

OpenAI like API​

To use an OpenAI like API, set the following in your .secrets.toml file:

[openai]
api_base = "https://api.openai.com/v1"
api_key = "sk-..."

or use the environment variables (make sure to use double underscores __):

OPENAI__API_BASE=https://api.openai.com/v1
OPENAI__KEY=sk-...

OpenAI Flex Processing​

To reduce costs for non-urgent/background tasks, enable Flex Processing:

[litellm]
extra_body='{"service_tier": "flex"}'

See OpenAI Flex Processing docs for details.

Chat template options​

For OpenAI-compatible endpoints that accept chat_template_kwargs, such as a vLLM deployment serving Qwen, pass a JSON object through the existing global option:

[litellm]
extra_body='{"chat_template_kwargs": {"enable_thinking": false}}'

PR-Agent sends this object in the request body and preserves generated OpenRouter routing, reasoning, and request-attribution fields. The setting applies to every configured model, including fallback models, so use it only when all selected endpoints support the field. Existing service_tier and processing_mode options can be included in the same JSON object.

Azure​

To use Azure, set in your .secrets.toml (working from CLI), or in the GitHub Settings > Secrets and variables (working from GitHub App or GitHub Action):

[openai]
key = "" # your azure api key
api_type = "azure"
api_version = '2023-05-15' # Check Azure documentation for the current API version
api_base = "" # The base URL for your Azure OpenAI resource. e.g. "https://<your resource name>.openai.azure.com"
deployment_id = "" # The deployment name you chose when you deployed the engine

and set in your configuration file:

[config]
model="" # the OpenAI model you've deployed on Azure (e.g. gpt-4o)
fallback_models=["..."]

Azure AD authentication needs the azure extra (pip install "pr-agent[azure]"). To use Azure AD (Entra id) based authentication set in your .secrets.toml (working from CLI), or in the GitHub Settings > Secrets and variables (working from GitHub App or GitHub Action):

[azure_ad]
client_id = "" # Your Azure AD application client ID
client_secret = "" # Your Azure AD application client secret
tenant_id = "" # Your Azure AD tenant ID
api_base = "" # Your Azure OpenAI service base URL (e.g., https://openai.xyz.com/)

The request-local Azure OIDC bridge captures AZURE_CLIENT_ID, AZURE_TENANT_ID, AZURE_AUTHORITY_HOST, AZURE_SCOPE, and competing AZURE_CLIENT_SECRET / AZURE_USERNAME / AZURE_PASSWORD settings when the handler is initialized. It preserves native authentication precedence while binding companion credentials and their cache identity to the captured authority; an absent authority uses the Azure public-cloud default, while an empty authority is rejected for companion credentials. LiteLLM's native assertion and secret-selector resolvers still control updates and source selection; their underlying deployment-owned credential chains are not request-isolated. Keep those sources scoped to the intended workload identity, or use separate processes for distinct identities.

For native Azure SDK routes other than Cloudflare gateways, ordinary AD tokens use the same captured companion settings and SDK client-cache isolation. Azure Responses routes also bind ordinary AD companion selection to the captured settings. Complete companion credentials can also be selected without an initial AD token, including on native Azure AI raw HTTP routes. Captured companion providers retain callable token refresh and native authentication precedence. If SDK initialization or Responses token resolution reaches LiteLLM's implicit credential discovery with enable_azure_ad_token_refresh=True, the request is rejected: that fallback re-reads process-wide identity settings. Configure complete client-secret or username/password companion credentials instead. Managed identity, certificate, and default credential-chain discovery through this fallback are intentionally unsupported; an already cached, isolated SDK client can still be reused without entering discovery.

Passing custom headers to the underlying LLM Model API can be done by setting extra_headers parameter to litellm.

[litellm]
extra_headers='{"projectId": "<authorized projectId >", ...}') #The value of this setting should be a JSON string representing the desired headers, a ValueError is thrown otherwise.

This enables users to pass authorization tokens or API keys, when routing requests through an API management gateway.

Requests that would otherwise inherit non-empty process-wide litellm.headers are rejected, even for non-authentication headers; configure headers through LITELLM.EXTRA_HEADERS instead.

Ollama​

You can run models locally through either VLLM or Ollama

E.g. to use a new model locally via Ollama, set in .secrets.toml or in a configuration file:

[config]
model = "ollama/qwen2.5-coder:32b"
fallback_models=["ollama/qwen2.5-coder:32b"]
custom_model_max_tokens=128000 # set the maximal input tokens for the model
duplicate_examples=true # will duplicate the examples in the prompt, to help the model to generate structured output

[ollama]
api_base = "http://localhost:11434" # or whatever port you're running Ollama on

By default, Ollama uses a context window size of 2048 tokens. In most cases this is not enough to cover pr-agent prompt and pull-request diff. Context window size can be overridden with the OLLAMA_CONTEXT_LENGTH environment variable. For example, to set the default context length to 8K, use: OLLAMA_CONTEXT_LENGTH=8192 ollama serve. More information you can find on the official ollama faq.

Please note that the custom_model_max_tokens setting should be configured in accordance with the OLLAMA_CONTEXT_LENGTH. Failure to do so may result in unexpected model output.

Local models vs commercial models

PR-Agent is compatible with almost any AI model, but analyzing complex code repositories and pull requests requires a model specifically optimized for code analysis.

Commercial models such as GPT-5, Claude Sonnet, and Gemini have demonstrated robust capabilities in generating structured output for code analysis tasks with large input. In contrast, most open-source models currently available (as of January 2025) face challenges with these complex tasks.

Based on our testing, local open-source models are suitable for experimentation and learning purposes (mainly for the ask command), but they are not suitable for production-level code analysis tasks.

Hence, for production workflows and real-world usage, we recommend using commercial models.

Hugging Face​

To use a new model with Hugging Face Inference Endpoints, for example, set:

[config] # in configuration.toml
model = "huggingface/meta-llama/Llama-2-7b-chat-hf"
fallback_models=["huggingface/meta-llama/Llama-2-7b-chat-hf"]
custom_model_max_tokens=... # set the maximal input tokens for the model

[huggingface] # in .secrets.toml
key = ... # your Hugging Face api key
api_base = ... # the base url for your Hugging Face inference endpoint

(you can obtain a Llama2 key from here)

Replicate​

To use Llama2 model with Replicate, for example, set:

[config] # in configuration.toml
model = "replicate/llama-2-70b-chat:2c1608e18606fad2812020dc541930f2d0495ce32eee50074220b87300bc16e1"
fallback_models=["replicate/llama-2-70b-chat:2c1608e18606fad2812020dc541930f2d0495ce32eee50074220b87300bc16e1"]
[replicate] # in .secrets.toml
key = ...

(you can obtain a Llama2 key from here)

Also, review the .secrets_template.toml file for instructions on how to set keys for other models.

Groq​

To use Llama3 model with Groq, for example, set:

[config] # in configuration.toml
model = "llama3-70b-8192"
fallback_models = ["groq/llama3-70b-8192"]
[groq] # in .secrets.toml
key = ... # your Groq api key

(you can obtain a Groq key from here)

SambaNova​

To use MiniMax-M3 model with SambaNova, for example, set:

[config] # in configuration.toml
model = "sambanova/MiniMax-M3"
fallback_models = ["sambanova/MiniMax-M2.7"]
[sambanova] # in .secrets.toml
key = ... # your SambaNova api key

(you can obtain a SambaNova key from here)

xAI​

To use xAI's models with PR-Agent, set:

[config] # in configuration.toml
model = "xai/grok-4.6"
fallback_models = ["xai/grok-4.6"] # or any other model as fallback

[xai] # in .secrets.toml
key = "..." # your xAI API key

You can obtain an xAI API key from xAI's console by creating an account and navigating to the developer settings page.

Grok 4.5 and Grok 4.6 are registered with a 500K token context window (xai/grok-4.5, xai/grok-4.5-latest, xai/grok-build-latest, xai/grok-4.6, openrouter/x-ai/grok-4.5, openrouter/x-ai/grok-4.6). xAI publishes grok-4.5-latest and grok-build-latest aliases for Grok 4.5; Grok 4.6 currently has no published alias.

Grok 4.5 and Grok 4.6 are always-on reasoning models and honor config.reasoning_effort (low, medium, high; "xhigh" is supported on Grok 4.6 and later). PR-Agent sends medium by default; set "high" to restore xAI's native default. This setting is global, so changing it also affects other registered reasoning models. Unsupported values are clamped to the closest accepted level (none/minimal → "low"; "max"/"xhigh" on Grok 4.5 → "high"; "max" on Grok 4.6 → "xhigh").

OpenRouter routes (openrouter/x-ai/grok-4.5, openrouter/x-ai/grok-4.6) apply the same clamping after the OpenRouter effort value is resolved but before the none-versus-budget decision. An explicit openrouter.reasoning_effort overrides the global effort; a positive openrouter.reasoning_max_tokens remains budget-only and suppresses effort, including a clamped "none" value. Routing suffixes such as :nitro require custom_model_max_tokens because token lookup currently uses exact model IDs.

Vertex AI​

Vertex AI needs the google extra (pip install "pr-agent[google]"). To use Google's Vertex AI platform and its associated models (chat-bison/codechat-bison) set:

[config] # in configuration.toml
model = "vertex_ai/codechat-bison"
fallback_models="vertex_ai/codechat-bison"

[vertexai] # in .secrets.toml
vertex_project = "my-google-cloud-project"
vertex_location = ""

Your application default credentials will be used for authentication so there is no need to set explicit credentials in most environments.

If you do want to set explicit credentials, then you can use the GOOGLE_APPLICATION_CREDENTIALS environment variable set to a path to a json credentials file.

Each handler captures the selected Cloud SDK ADC file and its resource project configuration, including CLOUDSDK_CONFIG, the active named configuration, and CLOUDSDK_CORE_PROJECT. Later changes apply to new handlers, not existing ones. The quota project remains separate from the resource project. When no ADC file exists at initialization, the handler retains managed-runtime discovery rather than adopting a file created later.

AWS-backed Vertex workload identity federation captures environment credentials and region per handler. Different captured identities use separate credential caches even when they share the same WIF configuration; identical snapshots can reuse a cache entry. Metadata-backed credentials continue refreshing through the captured metadata source.

Vertex external-account credentials with an executable source are intentionally unsupported, including when GOOGLE_EXTERNAL_ACCOUNT_ALLOW_EXECUTABLES=1. The helper and its cached output can change identity after a handler captures its configuration, so Vertex requests reject this source before using it for authentication. Use a non-executable credential source configured for the intended identity instead; normal credential refresh remains enabled for supported sources.

Google AI Studio​

To use Google AI Studio models, set the relevant models in the configuration section of the configuration file:

[config] # in configuration.toml
model="gemini/gemini-3.8-flash"
fallback_models=["gemini/gemini-3.8-flash"]

[google_ai_studio] # in .secrets.toml
gemini_api_key = "..."

If you don't want to set the API key in the .secrets.toml file, you can set the GOOGLE_AI_STUDIO.GEMINI_API_KEY environment variable.

Anthropic​

To use Anthropic models, set the relevant models in the configuration section of the configuration file:

[config]
model="anthropic/claude-opus-5"
fallback_models=["anthropic/claude-opus-5"]

And also set the api key in the .secrets.toml file:

[anthropic]
KEY = "..."

See litellm documentation for more information about the environment variables required for Anthropic.

Amazon Bedrock​

To use Amazon Bedrock and its foundational models, add the below configuration:

[config] # in configuration.toml
model="bedrock/anthropic.claude-3-5-sonnet-20240620-v1:0"
fallback_models=["bedrock/anthropic.claude-3-5-sonnet-20240620-v1:0"]

[aws]
AWS_ACCESS_KEY_ID="..."
AWS_SECRET_ACCESS_KEY="..."
AWS_REGION_NAME="..."

You can also use the new Meta Llama 4 models available on Amazon Bedrock:

[config] # in configuration.toml
model="bedrock/us.meta.llama4-scout-17b-instruct-v1:0"
fallback_models=["bedrock/us.meta.llama4-maverick-17b-instruct-v1:0"]

Kimi K3 is also available on Amazon Bedrock:

[config] # in configuration.toml
model="bedrock/moonshotai.kimi-k3"
fallback_models=["bedrock/us.moonshotai.kimi-k3"]

Use the bare bedrock/moonshotai.kimi-k3 id where it's directly available, the us. cross-region prefix to route across US regions, or the global. prefix to let Bedrock route across all supported regions. To call it through the Bedrock Converse API instead of the classic runtime, prefix the model id with bedrock/converse/, e.g. bedrock/converse/us.moonshotai.kimi-k3.

Grok 4.3 is available through Amazon Bedrock Mantle rather than the classic Bedrock runtime:

[config] # in configuration.toml
model="bedrock_mantle/xai.grok-4.3"
fallback_models=["bedrock_mantle/xai.grok-4.3"]

Bedrock Mantle uses the same AWS credential sources, but its IAM permissions differ from the classic runtime. See the AWS Mantle inference permissions.

When running PR-Agent on AWS infrastructure (EC2, ECS/Fargate, EKS with IRSA, Lambda, or any self-hosted GitHub Actions runner on AWS), the instance or task already has an IAM role attached. You can use those ambient credentials directly instead of storing long-lived static keys.

Set AWS_USE_IMDS=true in the environment. PR-Agent will resolve credentials via boto3's standard provider chain, which handles all AWS compute contexts transparently:

Compute contextMechanism
EC2 instance with IAM roleIMDSv2 (169.254.169.254)
ECS / Fargate task roleTask metadata endpoint
EKS pod with IRSAWeb identity token + STS
Lambda functionRuntime-injected credentials

Credential discovery runs during handler initialization. Before each eligible SigV4 request, refresh through the same boto3 credentials object runs in a background thread. Non-AWS and bearer-authenticated requests do not trigger this refresh. PR-Agent passes a request-local snapshot to LiteLLM without writing credentials into the process environment. Constructor discovery and credential-file fingerprinting remain synchronous.

AWS calls using this provider chain remain serialized within a handler, including any static-credential retry. Cancelling a request does not stop an already-running boto3 refresh, but its result cannot overwrite the handler's credentials or static-fallback decision. A separate lock serializes SDK refreshes, including cancelled callers' unfinished work. Refresh uses the event loop's shared default executor: blocked operations and workers waiting for the SDK lock after repeated cancellations can delay unrelated executor work and process shutdown. No service-wide worker quota or additional SDK timeout is introduced.

The same opt-in is required for other boto3 provider-chain sources, including AWS_PROFILE and shared credentials files. LiteLLM-specific AWS_PROFILE_NAME and AWS_ROLE_NAME selectors are not supported because they can override request-local credentials; unset them and use AWS_USE_IMDS=true instead.

When upgrading from implicit LiteLLM credential-chain discovery, set AWS_USE_IMDS=true explicitly. Without this opt-in, PR-Agent only uses complete static credentials from settings or environment credentials captured when the handler is initialized; it does not ask LiteLLM to discover a role or profile at request time.

For classic Bedrock, a supported model ARN supplies the request region ahead of environment/settings regions, including during static-credential retries. Converse also uses a captured litellm.model_id ARN, or a region/model path when model_id is absent; Invoke does not derive its region from the separate model_id. Otherwise, without AWS_USE_IMDS=true, set aws.AWS_REGION_NAME, AWS_REGION_NAME, AWS_REGION, or AWS_DEFAULT_REGION. Complete environment credentials alone do not enable region discovery from boto3 profiles or LiteLLM's default region. Static credentials in [aws] still require AWS_REGION_NAME.

For Bedrock Mantle without IMDS or complete static credentials, the region comes from BEDROCK_MANTLE_REGION, AWS_REGION_NAME, aws.AWS_REGION_NAME, or AWS_REGION, then defaults to us-east-1; AWS_DEFAULT_REGION alone does not change that default. A BEDROCK_MANTLE_API_BASE on the standard host https://bedrock-mantle.<region>.api.aws supplies the region and overrides these sources; a custom endpoint host does not. Complete static credentials and opted-in boto3 discovery retain the AWS region policy described above.

Without AWS_USE_IMDS=true, environment authentication follows LiteLLM's AWS_SESSION_TOKEN selection. The legacy AWS_SECURITY_TOKEN alias is recognized only by the opted-in boto3 chain, which prefers it over AWS_SESSION_TOKEN when both are nonempty.

Bedrock bearer tokens must come from handler-captured credentials rather than process-wide LiteLLM fallback, and AWS_BEARER_TOKEN_BEDROCK must be unset for sagemaker_chat and sagemaker_nova routes.

When resolving credentials through profiles with AWS_USE_IMDS=true, PR-Agent does not execute credential_process in the selected profile or its source_profile chain. If such a process is configured, PR-Agent uses complete static credentials from [aws] (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, and AWS_REGION_NAME, plus AWS_SESSION_TOKEN when required) instead; without them, credential resolution fails. This restriction does not apply when using complete environment credentials captured at handler initialization.

Workload token files selected by AWS_WEB_IDENTITY_TOKEN_FILE, profile web_identity_token_file, or AWS_CONTAINER_AUTHORIZATION_TOKEN_FILE must be controlled by the deployment, not switched between tenants. PR-Agent retains native token reload and credential refresh from the selected source; it does not freeze token contents or detect a different identity replacing the contents at the same path. Use separately controlled credential sources and appropriate process/container isolation for distinct workload identities, or supply explicit request-local credentials. Request-local configuration is not an isolation boundary against arbitrary code running in the same process.

Minimal GitHub Actions workflow (no AWS secret keys required):

- uses: the-pr-agent/pr-agent@main
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
AWS_USE_IMDS: "true"
# AWS_REGION_NAME: us-east-1 # optional if the instance metadata provides it
with:
command: review

The IAM role must have bedrock:InvokeModel permission on the target model ARN, for example:

{
"Effect": "Allow",
"Action": "bedrock:InvokeModel",
"Resource": "arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-3-5-sonnet-20240620-v1:0"
}

If you also configure static keys in [aws], they serve as an automatic fallback when ambient credentials cannot be resolved or a SigV4 call using the opted-in provider chain raises an API error other than a rate-limit error. This applies to Bedrock, Bedrock Mantle, and SageMaker routes and preserves the existing fallback behavior for IAM authorization and connection failures, using only the static credentials captured by the handler.

Custom Inference Profiles​

To invoke an application inference profile with a classic bedrock/ model (for cost allocation tags and other configuration settings), set model_id to the profile ARN in your configuration:

[config] # in configuration.toml
model="bedrock/anthropic.claude-3-5-sonnet-20240620-v1:0"
fallback_models=["bedrock/anthropic.claude-3-5-sonnet-20240620-v1:0"]

[aws]
AWS_ACCESS_KEY_ID="..."
AWS_SECRET_ACCESS_KEY="..."
AWS_REGION_NAME="..."

[litellm]
model_id = "your-application-inference-profile-arn"

The litellm.model_id parameter applies only to classic bedrock/ calls made through the bedrock-runtime APIs. It does not apply to bedrock_mantle/; for cost allocation with the Mantle Chat Completions and Responses APIs, use Amazon Bedrock Projects.

The profile is sent only with requests for the model set in config.model. Models in config.fallback_models do not use it, so a fallback is never routed to the primary model's inference profile.

To give a fallback its own application inference profile, list it in litellm.model_ids, keyed by the exact model name. An entry in model_ids takes priority for that model; model_id still applies to config.model when it has no entry. A model that is in neither gets no profile.

[litellm]
model_ids = {"bedrock/anthropic.claude-3-5-sonnet-20240620-v1:0" = "your-primary-profile-arn", "bedrock/qwen.qwen3-235b-a22b-2507-v1:0" = "your-fallback-profile-arn"}
Prompt caching and run cost with an ARN​

Prompt caching and run-cost estimation identify the model by name, so an opaque application inference profile ARN needs extra configuration:

  • Add the ARN to claude_adaptive_thinking_models_override (or claude_extended_thinking_models_override for extended thinking) so PR-Agent treats it as Claude and forwards cache_control_injection_points. See Claude 5 thinking with an application inference profile ARN.
  • Map the ARN to a LiteLLM-priced model id in [litellm] base_models, so run cost is estimated instead of reported as unavailable:
[litellm.base_models]
"bedrock/converse/arn:aws:bedrock:eu-central-1:<account-id>:application-inference-profile/<profile-id>" = "bedrock/anthropic.claude-sonnet-4-5-20250929-v1:0"

Claude 5 thinking with an application inference profile ARN​

Claude Sonnet 5 on Bedrock is invoked through an inference profile rather than a direct foundation-model id. When that profile is an application inference profile, its ARN is an opaque value that carries no model name. Add that ARN to claude_adaptive_thinking_models_override so PR-Agent and LiteLLM both treat it as an adaptive-thinking model:

Use the ARN as the model id and repeat that exact value in the override:

[config] # in configuration.toml
model = "bedrock/converse/arn:aws:bedrock:eu-central-1:<account-id>:application-inference-profile/<profile-id>"
enable_claude_adaptive_thinking = true
claude_adaptive_thinking_models_override = [
"bedrock/converse/arn:aws:bedrock:eu-central-1:<account-id>:application-inference-profile/<profile-id>"
]

The override is additive, so named Claude models in the same fallback chain continue to use built-in detection. PR-Agent also registers each override with LiteLLM, preventing LiteLLM from converting the adaptive payload to the legacy budget_tokens shape that Bedrock rejects.

ARNs only need the override when the suffix is opaque. An ARN that embeds the model family, for example ...:inference-profile/us.anthropic.claude-sonnet-5, normalises to a string the adaptive regex already matches.

To route Bedrock traffic through a VPC interface endpoint instead of the public bedrock-runtime endpoint, set AWS_BEDROCK_RUNTIME_ENDPOINT either as an environment variable or in [aws]:

[aws]
AWS_BEDROCK_RUNTIME_ENDPOINT="https://bedrock-runtime.us-east-1.amazonaws.com"

See litellm documentation for more information about the environment variables required for Amazon Bedrock.

DeepSeek​

To use deepseek-v4 model with DeepSeek, for example, set:

[config] # in configuration.toml
model = "deepseek/deepseek-v4-pro"
fallback_models=["deepseek/deepseek-v4-flash"]

and fill up your key

[deepseek] # in .secrets.toml
key = ...

(you can obtain a deepseek-v4 key from here)

GLM (Z.AI)​

To use GLM models with Z.AI (Zhipu), for example, set:

[config] # in configuration.toml
model = "zai/glm-5.2"
fallback_models=["zai/glm-5.2"]

and fill up your key

[zai] # in .secrets.toml
key = ...

(you can obtain a Z.AI API key from here)

Kimi (Moonshot)​

To use Kimi models with Moonshot, for example, set:

[config] # in configuration.toml
model = "moonshot/kimi-k3"
fallback_models=["moonshot/kimi-k3"]

and fill up your key

[moonshot] # in .secrets.toml
key = ...

(you can obtain a Moonshot API key from here)

If you are on the China endpoint instead, add api_base = "https://api.moonshot.cn/v1" under [moonshot].

Qwen (DashScope)​

To use Qwen models with Alibaba DashScope, for example, set:

[config] # in configuration.toml
model = "dashscope/qwen3.8-max"
fallback_models=["dashscope/qwen3.8-max"]

and fill up your key

[dashscope] # in .secrets.toml
key = ...

(you can obtain a DashScope API key from here)

Xiaomi MiMo​

To use Xiaomi MiMo models, for example, set:

[config] # in configuration.toml
model = "xiaomi_mimo/mimo-v2.5"
fallback_models=["xiaomi_mimo/mimo-v2.5"]

and fill up your key

[xiaomi_mimo] # in .secrets.toml
key = ...

(you can obtain a Xiaomi MiMo API key from here)

DeepInfra​

To use DeepSeek model with DeepInfra, for example, set:

[config] # in configuration.toml
model = "deepinfra/deepseek-ai/DeepSeek-R1-Distill-Llama-70B"
fallback_models = ["deepinfra/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B"]
[deepinfra] # in .secrets.toml
key = ... # your DeepInfra api key

(you can obtain a DeepInfra key from here)

Mistral​

To use models like Mistral or Codestral with Mistral, for example, set:

[config] # in configuration.toml
model = "mistral/mistral-small-latest"
fallback_models = ["mistral/mistral-medium-latest"]
[mistral] # in .secrets.toml
key = "..." # your Mistral api key

(you can obtain a Mistral key from here)

Codestral​

To use Codestral model with Codestral, for example, set:

[config] # in configuration.toml
model = "codestral/codestral-latest"
fallback_models = ["codestral/codestral-2405"]
[codestral] # in .secrets.toml
key = "..." # your Codestral api key

(you can obtain a Codestral key from here)

Databricks​

To use a model hosted on Databricks (e.g. an Azure Databricks serving endpoint), set:

[config] # in configuration.toml
model = "databricks/databricks-claude-sonnet-4"
fallback_models = ["databricks/databricks-claude-sonnet-4"]
[databricks] # in .secrets.toml
api_key = "..." # your Databricks personal access token (PAT)
api_base = "https://adb-xxxx.azuredatabricks.net/serving-endpoints" # your workspace serving-endpoints URL

The model name after the databricks/ prefix is the name of your serving endpoint. See LiteLLM's Databricks provider docs for details.

The configured PAT and endpoint are captured per handler. When both DATABRICKS_CLIENT_ID and DATABRICKS_CLIENT_SECRET are set, LiteLLM performs an OAuth M2M exchange even if a PAT is supplied. An exchange failure can abort the request; after a successful exchange, the supplied PAT replaces the OAuth token in the final Authorization header. Without OAuth M2M credentials or a PAT, optional Databricks SDK authentication remains available. These deployment-owned credentials and profile selection are not request-isolated. Do not change those sources between tenants in a shared process; use request-specific PATs with OAuth M2M credentials unset, or separate processes for distinct workload identities.

Openrouter​

To use model from Openrouter, for example, set:

[config] # in configuration.toml
model="openrouter/anthropic/claude-sonnet-5"
fallback_models=["openrouter/deepseek/deepseek-chat"]
custom_model_max_tokens=20000

[openrouter] # in .secrets.toml or passed an environment variable openrouter__key
key = "..." # your openrouter api key

(you can obtain an Openrouter API key from here)

OpenRouter's router models can be selected directly without setting custom_model_max_tokens:

[config]
model = "openrouter/auto"
fallback_models = ["openrouter/free"]

PR-Agent also registers openrouter/fusion and openrouter/pareto-code. Provider routing, reasoning, and output-cap settings are optional for all four router models; omit them to use OpenRouter's defaults. See OpenRouter's documentation for the Auto, Free, Fusion, and Pareto routers.

Openrouter provider routing, reasoning and output cap​

For openrouter/... models you can optionally restrict which upstream providers Openrouter uses, control reasoning, and cap the completion length. All keys live in the [openrouter] section of configuration.toml. Reasoning-capable models are those litellm's bundled reasoning metadata flags over the model id and its provider-prefixed/xai/-forms, the maintained Grok registry, or config.additional_reasoning_effort_models; they inherit config.reasoning_effort unless an Openrouter-specific effort or token budget is set.

[openrouter]
# Uncomment and adjust the keys you need; unset keys keep Openrouter's defaults.
# provider_only = ["z-ai"] # hard allowlist of upstream providers; empty = default routing
# provider_order = ["z-ai", "novita"] # preferred order instead of an allowlist; ignored when provider_only is set
# allow_fallbacks = true # when provider_order is set, allow routing beyond the list
# reasoning_effort = "low" # override global effort: "none", "minimal", "low", "medium", "high", "xhigh" or "max"
# reasoning_max_tokens = 2048 # explicit budget; ignored only when final effort remains "none"
# max_tokens = 16000 # hard cap on completion tokens for the request

provider_only and reasoning_effort = "none" are useful to pin a specific provider and to bound the cost of reasoning models. Because Openrouter treats effort and token budgets as mutually exclusive, an explicit Openrouter-specific "none" keeps reasoning disabled when the model supports disabling it. Grok 4.5/4.6 and Gemini 3.7/3.8 Flash clamp "none" to "low" before precedence is applied, so a positive budget wins for those models; explicit Openrouter "minimal" remains unchanged for Gemini. Otherwise a positive reasoning_max_tokens value takes precedence over the global effort and other Openrouter-specific values. Invalid Openrouter-specific effort values are warned about and treated as unset, so registered reasoning models fall back to config.reasoning_effort. Openrouter normalizes "max" to "xhigh" in this path for LiteLLM/OpenRouter compatibility. Supported effort values vary by model, and models whose metadata marks reasoning as mandatory reject "none". For Anthropic models using a reasoning budget, set the effective output max_tokens higher than reasoning_max_tokens so the final answer has output headroom. See the Openrouter provider routing and reasoning tokens docs.

OrcaRouter​

OrcaRouter is an OpenAI-compatible AI gateway. It needs no provider-specific code in PR-Agent: the openai/ prefix routes the request to OrcaRouter's base URL through litellm's OpenAI-compatible path, the same way the Neon AI Gateway is handled.

To use a model through OrcaRouter, set:

[config] # in configuration.toml
model = "openai/anthropic/claude-fable-5"
fallback_models = ["openai/auto"]
custom_model_max_tokens = 20000

[openai] # in .secrets.toml
api_base = "https://api.orcarouter.ai/v1"
key = "..." # your OrcaRouter api key

or use the environment variables (make sure to use double underscores __):

OPENAI__API_BASE=https://api.orcarouter.ai/v1
OPENAI__KEY=...

(you can obtain an OrcaRouter API key from here)

Keep the openai/ prefix on the model name, whatever OrcaRouter model ID you use (openai/anthropic/claude-fable-5, openai/auto, ...): the prefix routes the request through litellm's OpenAI-compatible path. A prefixed name is not in the MAX_TOKENS table here, so you also have to set custom_model_max_tokens. OrcaRouter governs routing and guardrails itself, but config.reasoning_effort still reaches it: PR-Agent probes the suffixed model ID against litellm's bundled reasoning metadata (or config.additional_reasoning_effort_models), so an ID such as openai/google/gemini-2.5-pro or openai/o3 sends the configured effort (default "medium") even with nothing set. The example IDs above are not flagged as reasoning-capable and are unaffected.

Neon AI Gateway​

Neon AI Gateway is an OpenAI-compatible inference gateway. Each Neon branch has its own gateway host, so the base URL points at a single branch and not at an account. Neon publishes that host alongside the credential as NEON_AI_GATEWAY_BASE_URL. The value has no path, so append /v1 to reach chat completions.

To use a model served by a Neon branch, set:

[config] # in configuration.toml
model = "openai/gpt-5-mini"
fallback_models = ["openai/gpt-5-mini"]
custom_model_max_tokens = 400000 # the context window Neon publishes for the model

[openai] # in .secrets.toml
api_base = "https://<your-neon-branch-host>/v1"
key = "..." # your Neon AI Gateway credential

or use the environment variables (make sure to use double underscores __):

OPENAI__API_BASE=https://<your-neon-branch-host>/v1
OPENAI__KEY=...

Keep the openai/ prefix on the model name, whichever Neon model ID you use: the prefix routes the request through litellm's OpenAI-compatible path. A prefixed name is not in the MAX_TOKENS table here, so you also have to set custom_model_max_tokens. Take the value from Neon's model catalog.

Create the credential per branch in the Neon Console with the ai_gateway:invoke scope. The credential also works on branches descended from the one it was created on. The gateway is in beta and requires a paid Neon plan. It runs only in AWS US East (Ohio), aws-us-east-2.

Chat completions only

Some model IDs in Neon's catalog are served through the OpenAI Responses API, which Neon exposes under /openai/v1 instead of /v1. The configuration above points at the chat-completions endpoint, so it cannot reach those models. Neon also documents a few models that return message.content as an array of typed blocks rather than a string, and PR-Agent reads the reply as a string.

GitHub Copilot​

Models under an active GitHub Copilot subscription are available through litellm's github_copilot provider, which authenticates as your GitHub identity rather than with a dedicated API key:

[config]
model = "github_copilot/gpt-4o"
fallback_models = ["github_copilot/gpt-4.1"]

The GitHub identity behind the model needs an active Copilot subscription. The token budgets for these models are resolved automatically from litellm's model metadata, so custom_model_max_tokens is not required. However, get_max_tokens clamps the effective window to config.max_model_tokens, which defaults to 32000. To use the input limits recorded for these Copilot routes (e.g., 64000 for gpt-4o, 128000 for gpt-4.1), raise config.max_model_tokens accordingly.

Authentication uses the GitHub Copilot provider flow:

  1. litellm first looks for a pre-seeded GitHub access token in access-token under GITHUB_COPILOT_TOKEN_DIR (default ~/.config/litellm/github_copilot; the filename is overridable with GITHUB_COPILOT_ACCESS_TOKEN_FILE). In a CI runner, write that token before the job runs - for example, mount a secret into the directory or point the variable at a mounted secret directory - and the flow never becomes interactive.
  2. Only when the file is missing or empty does litellm fall back to the interactive device-code flow (POST https://github.com/login/device/code, up to three attempts), which does not suit unattended runners.
  3. The Copilot API key (api-key.json in the same directory) is refreshed automatically against https://api.github.com/copilot_internal/v2/token, using the pre-seeded access token.

Whether Copilot's terms permit this programmatic use is a question for GitHub rather than a guarantee this project can make, so confirm before relying on the route.

Atlas Cloud​

Atlas Cloud is an OpenAI-compatible inference platform. It needs no provider-specific code in PR-Agent: the openai/ prefix routes the request to Atlas's base URL through litellm's OpenAI-compatible path, the same way OrcaRouter and the Neon AI Gateway are handled.

To use a model served by Atlas Cloud, set:

[config] # in configuration.toml
model = "openai/deepseek-ai/deepseek-v4-pro"
fallback_models = ["openai/deepseek-ai/deepseek-v4-pro"]
custom_model_max_tokens = 1048000 # the context window Atlas publishes for the model

[openai] # in .secrets.toml
api_base = "https://api.atlascloud.ai/v1"
key = "..." # your Atlas Cloud api key

or use the environment variables (make sure to use double underscores __):

OPENAI__API_BASE=https://api.atlascloud.ai/v1
OPENAI__KEY=...

(you can obtain an Atlas Cloud API key from the console)

Keep the openai/ prefix on the model name, whichever Atlas model ID you use (openai/deepseek-ai/deepseek-v4-pro, openai/zai-org/glm-5, openai/moonshotai/kimi-k2.6, ...): the prefix routes the request through litellm's OpenAI-compatible path. A prefixed name is not in the MAX_TOKENS table here, so you also have to set custom_model_max_tokens. Take the value from Atlas's model catalog.

Reasoning models need output headroom

Several Atlas models are reasoning models that spend completion tokens on a hidden chain of thought before writing the answer. deepseek-ai/deepseek-v4-pro with max_tokens = 16 returns finish_reason = "length" and an empty message.content — all 16 completion tokens were reasoning tokens. If a tool comes back blank, raise the output budget rather than assuming the request failed. Non-reasoning IDs such as deepseek-ai/DeepSeek-V3.1 are unaffected.

Custom models​

If the relevant model doesn't appear here, you can still use it as a custom model:

  1. Set the model name in the configuration file:
[config]
model="custom_model_name"
fallback_models=["custom_model_name"]
  1. Set the maximal tokens for the model:
[config]
custom_model_max_tokens= ...
  1. Go to litellm documentation, find the model you want to use, and set the relevant environment variables.

  2. Most reasoning models do not support chat-style inputs (system and user messages) or temperature settings. To bypass chat templates and temperature controls, set config.custom_reasoning_model = true in your configuration file.

Dedicated parameters​

OpenAI models​

[config]
reasoning_effort = "medium" # "none", "minimal", "low", "medium", "high", "xhigh", "max"

With the OpenAI models that support reasoning effort (eg: gpt-5.6-terra), you can specify its reasoning effort via config section. The default value is medium. You can change it to any supported value based on your usage. Available values depend on the model and provider. Where litellm marks minimal unsupported for a GPT-5 model, PR-Agent sends low instead.

For a model served through an OpenAI-compatible endpoint that litellm does not recognize as reasoning-capable, add its ID to config.additional_reasoning_effort_models. For known models support is decided by litellm's bundled reasoning metadata plus the maintained Grok registry (Grok ids resolve through their xai/ prefix) with Claude models left out of the metadata path (their reasoning comes from the dedicated extended/adaptive thinking settings; an explicit entry in the list above still applies to them). Config IDs match exactly or through any provider prefix (e.g. "deepseek-v4-flash-0731" matches "openai/deepseek-v4-flash-0731"). When LiteLLM does not recognize the model, PR-Agent sets allowed_openai_params = ["reasoning_effort"] so the parameter reaches the endpoint. Note the default "medium" may be rejected by providers that accept a different subset (e.g. "none"/"low"/"high"/"max"); adding a custom model ID surfaces that provider-side error instead of silently dropping the setting.

For GPT-6 Sol or Luna hosted outside OpenAI, Azure, or OpenRouter under the same model ID, add the ID to this list to explicitly enable reasoning_effort.

To use GPT-6 Sol or GPT-6 Luna:

[config]
model = "gpt-6-sol" # or "gpt-6-luna"
reasoning_effort = "medium" # "none", "low", "medium", "high", "xhigh", "max"

For recognized native GPT-6 Sol/Luna IDs on OpenAI, Azure, and OpenRouter routes, PR-Agent omits temperature. On OpenAI routes, their native none and max reasoning efforts are passed through unchanged. Azure routes preserve none and map max to xhigh for Chat Completions. OpenRouter maps none to disabled reasoning and max to xhigh. The legacy minimal setting is mapped to low on all three routes. Both models have a 1,050,000-token context window, a 922,000-token input ceiling, and support up to 128,000 output tokens. PR-Agent also applies config.max_model_tokens unless a tool bypasses that configured cap, as /help does; the native input ceiling still applies. The existing Chat Completions path is used for PR-Agent's text requests. OpenAI requires the Responses API for built-in tools and function calling with reasoning; Chat Completions function calling is limited to reasoning_effort = "none".

Unrecognized OpenRouter _thinking variants, such as _thinking:batch or _thinking:free, keep their literal IDs without native GPT-6 temperature or effort normalization. Add the full ID to config.additional_reasoning_effort_models to explicitly enable reasoning. For unregistered literal IDs, prompt budgeting requires usable LiteLLM metadata or config.custom_model_max_tokens. For bare Sol/Luna IDs on non-native custom providers, a positive config.custom_model_max_tokens takes precedence over the native registry value.

To use GPT-6 Astra:

[config]
model = "gpt-6-astra"
reasoning_effort = "medium" # "low", "medium", "high", "xhigh", "max"

PR-Agent omits temperature for GPT-6 Astra and maps none or minimal reasoning effort to low. Its 1,050,000-token context window remains subject to config.max_model_tokens. The existing Chat Completions path is used; access depends on your OpenAI account.

Anthropic models​

[config]
enable_claude_extended_thinking = false # Set to true to enable extended thinking feature
extended_thinking_budget_tokens = 2048
extended_thinking_max_output_tokens = 4096

By default, PR-Agent applies the extended-thinking payload only to a built-in list of Claude models (see CLAUDE_EXTENDED_THINKING_MODELS in pr_agent/algo/__init__.py). If you use a newer or custom Claude model that is not in that list, you can override it:

[config]
claude_extended_thinking_models_override = ["anthropic/claude-my-new-model"]

When claude_extended_thinking_models_override is non-empty, it fully replaces the built-in list, so include every model that should receive extended thinking. Leave it empty (the default) to use the built-in defaults.

Only models that accept a thinking budget are supported

PR-Agent enables extended thinking through the manual thinking={"type": "enabled", "budget_tokens": ...} request. Adaptive-only Claude models (e.g. Opus 4.7/4.8, Opus 5/5.5, Sonnet 5, Fable 5, Fable 5.1) reject budget_tokens, so they are intentionally excluded from the built-in defaults. If you add one to claude_extended_thinking_models_override anyway, PR-Agent skips the extended-thinking payload for it and logs a warning rather than sending a request the provider would reject — use enable_claude_adaptive_thinking for those models instead.

Both thinking gates only fire when the model id itself is recognizable: an opaque id such as a Bedrock application inference profile ARN matches neither gate, and PR-Agent logs a warning instead of silently sending nothing. See Claude 5 thinking with an application inference profile ARN.

Output token limit​

[config]
max_output_tokens = 0 # 0 = unset (default)

By default PR-Agent does not send an output token limit on model calls, so the provider's own default applies. On some providers that default is low — for example, AWS Bedrock (Converse API) can cap Claude reasoning models at 4096 output tokens, and since reasoning tokens count against that budget, the visible answer can come back empty or truncated. Set config.max_output_tokens to a positive value (e.g. 16000) to send it to LiteLLM as max_completion_tokens for native OpenAI/Azure GPT-6 models and recognized Astra IDs, including bare Astra IDs on custom providers. OpenRouter Astra uses this parameter only without a routing suffix or with :nitro/:floor. Other routes use max_tokens, and the azure_text and text-completion-openai routes always use max_tokens. Use a value supported by the selected model. GPT-6 Astra, Sol, and Luna support at most 128,000 output tokens, including reasoning tokens. PR-Agent does not automatically clamp this setting to the model's output limit. When Claude extended thinking is enabled, extended_thinking_max_output_tokens takes precedence. For models with small context windows, keep in mind that prompt and completion tokens share the model's context window: size config.max_model_tokens so the packed prompt leaves room for the configured output limit.

Routing small pull requests to a cheaper model​

config.model handles every pull request, however small. With model routing enabled, a pull request under a configured size goes to a cheaper model instead, and config.fallback_models still apply after it:

[model_routing]
enable = true

[[model_routing.rules]]
max_hunks = 3
model = "gpt-5.6-luna"

[[model_routing.rules]]
max_hunks = 15
max_files = 6
model = "gpt-5.6-terra"

Rules are checked in order, and the first one whose limits the pull request fits selects the primary model for the call. A pull request that fits no rule uses config.model. Size is measured by the number of diff hunks (max_hunks) and changed files (max_files) left after the [ignore] rules. Every git provider reports both, and neither depends on a model's tokenizer, so a threshold means the same thing whichever model it selects.

Routing applies only to calls that ask for the regular model: /review, /improve, /generate_labels and /add_docs. Tools that already use config.model_weak (/describe, /ask, /update_changelog) are left alone, and a dedicated config.model_reasoning is still used for self-reflection. With config.output_run_details enabled, the run details show which model a routed pull request ended up on.

Azure deployments

An Azure deployment is tied to one model, so when openai.deployment_id is set each rule also needs its own deployment_id. A rule without one is skipped with a warning and the next rule is tried.