Skip to content

Latest commit

 

History

History
1506 lines (1161 loc) · 109 KB

File metadata and controls

1506 lines (1161 loc) · 109 KB
title Rust ADK
description The Rust ADK (inference-gateway-adk) for building A2A-compatible agents in Rust. Covers the A2AServerBuilder and AgentBuilder fluent builders, the A2AClient and its typed JSON-RPC method helpers, custom function tools (with_function_tool / with_async_function_tool) and custom TaskHandler / StreamableTaskHandler implementations, OIDC bearer-token authentication (OidcJwtVerifier / AuthVerifier), TLS and mutual-TLS termination (PeerCert / ClientCertPrincipal), runtime AgentCardOverrides, the optional MCP client (McpClient / with_mcp_client) exposing the mcp_list_tools and mcp_call_tool selector tools over Streamable HTTP, the artifacts subsystem with filesystem and MinIO backends, and the full A2A_, MCP_, and ARTIFACTS_ environment-variable surface with defaults.

Rust ADK

The Rust ADK (inference-gateway-adk) is the Rust Agent Development Kit for building A2A (Agent-to-Agent) servers. It mirrors the Go ADK handler semantics and consumes the same canonical A2A schema, so agents written against either ADK speak the same wire protocol.

Pre-1.0 status. The Rust ADK is in early development and its public API may change between minor versions. Pin an exact version in production.

This page is a feature reference for the crate: the server and agent builders, the client and its typed JSON-RPC helpers, custom tools and task handlers, authentication, TLS/mTLS, agent-card overrides, the artifacts subsystem, and the full environment-variable surface. The canonical upstream source is the crate README.md.

Installation

Add the crate to your Cargo.toml:

[dependencies]
inference-gateway-adk = "0.4"

Three optional Cargo features extend the defaults:

Feature Pulls in Enables
redis a Redis client RedisStorage, selected when A2A_QUEUE_PROVIDER=redis.
minio the minio crate MinioArtifactStorage, selected when ARTIFACTS_STORAGE_PROVIDER=minio.
telemetry the OpenTelemetry OTLP crates (opentelemetry, opentelemetry-otlp, tracing-opentelemetry) OTLP span export from telemetry::init when enabled. See Telemetry.
inference-gateway-adk = { version = "0.4", features = ["redis", "minio"] }

The crate requires Rust 1.94 or later.

Quick start

Minimal server

The smallest runnable server loads an agent card and opts into the built-in task handlers. A2AServerBuilder::build() validates its inputs: an agent card is required (via with_agent_card or with_agent_card_from_file), and at least one task handler must be configured.

use inference_gateway_adk::A2AServerBuilder;
use tracing::{error, info};

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    tracing_subscriber::fmt().init();

    let server = A2AServerBuilder::new()
        .with_agent_card_from_file(".well-known/agent-card.json", None)
        .with_default_task_handlers()
        .build()
        .await?;

    let addr = "0.0.0.0:8080".parse()?;
    info!("A2A server listening on {addr}");

    if let Err(e) = server.serve(addr).await {
        error!("server stopped: {e}");
    }
    Ok(())
}

With no Agent attached, the default handlers fall back to a built-in echo reply, so the server runs end-to-end without any LLM credentials. The agent card can also be supplied programmatically with with_agent_card(card) instead of loading it from disk.

AI-powered server

Config is a plain serde struct; pick whichever loader you like. The bundled examples use envy with the A2A_ prefix - the same convention adopted by the sibling Go and TypeScript ADKs. With A2A_AGENT_CLIENT_* env vars set, AgentBuilder produces a fully wired LLM agent:

use inference_gateway_adk::{A2AServerBuilder, AgentBuilder, Config};
use inference_gateway_sdk::{
    ChatCompletionTool, ChatCompletionToolType, FunctionObject, FunctionParameters,
};
use serde_json::{Value, json};
use tracing::{error, info};

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    tracing_subscriber::fmt().init();

    // Load A2A_AGENT_CLIENT_PROVIDER, A2A_AGENT_CLIENT_MODEL,
    // A2A_AGENT_CLIENT_API_KEY, A2A_SERVER_PORT, etc. AgentBuilder
    // fails fast at startup if provider/model are missing.
    let config: Config = envy::prefixed("A2A_").from_env()?;

    let tools = vec![ChatCompletionTool {
        type_: ChatCompletionToolType::Function,
        function: FunctionObject {
            name: "get_weather".to_string(),
            description: Some("Get weather information for a city".to_string()),
            parameters: Some(FunctionParameters(
                json!({
                    "type": "object",
                    "properties": {
                        "location": { "type": "string", "description": "City name" }
                    },
                    "required": ["location"]
                })
                .as_object()
                .unwrap()
                .clone(),
            )),
            strict: false,
        },
    }];

    let agent = AgentBuilder::new()
        .with_config(&config.agent_config)
        .with_system_prompt("You are a helpful weather assistant.")
        .with_toolbox(tools)
        .with_function_tool("get_weather".to_string(), |args: Value| {
            let location = args["location"].as_str().unwrap_or("Unknown");
            Ok(json!({ "location": location, "temperature": "22C" }).to_string())
        })
        .build()
        .await?;

    let port = config.server_config.port;
    let server = A2AServerBuilder::new()
        .with_config(config)
        .with_agent(agent)
        .with_agent_card_from_file(".well-known/agent-card.json", None)
        .with_default_task_handlers()
        .build()
        .await?;

    let addr = format!("0.0.0.0:{port}").parse()?;
    info!("AI-powered A2A server running on {addr}");

    if let Err(e) = server.serve(addr).await {
        error!("Server failed to start: {e}");
    }
    Ok(())
}

Switching providers and models

AgentBuilder is provider-agnostic. To swap models, change A2A_AGENT_CLIENT_PROVIDER and A2A_AGENT_CLIENT_MODEL and supply the matching API key - no code edits:

Provider A2A_AGENT_CLIENT_PROVIDER Example A2A_AGENT_CLIENT_MODEL API key env var
OpenAI openai gpt-5-mini OPENAI_API_KEY
DeepSeek deepseek deepseek-v4-flash DEEPSEEK_API_KEY
Anthropic anthropic claude-opus-4-8 ANTHROPIC_API_KEY
Cohere cohere command-a-03-2025 COHERE_API_KEY
Groq groq llama-3.3-70b-versatile GROQ_API_KEY
Cloudflare cloudflare @cf/meta/llama-3.3-70b-instruct-fp8-fast CLOUDFLARE_API_KEY
Ollama Cloud ollama_cloud gpt-oss:120b OLLAMA_CLOUD_API_KEY
Ollama ollama llama3.3 none (gateway reaches Ollama via OLLAMA_API_URL)
llama.cpp llamacpp llama-3.2-3b-instruct LLAMACPP_API_KEY
Google google gemini-3-flash GOOGLE_API_KEY
Mistral mistral mistral-large-3 MISTRAL_API_KEY
MiniMax minimax MiniMax-M3 MINIMAX_API_KEY
Moonshot moonshot kimi-latest MOONSHOT_API_KEY
NVIDIA nvidia nvidia/meta/llama-3.1-8b-instruct NVIDIA_API_KEY
Z-AI zai glm-5.2 ZAI_API_KEY

Set A2A_AGENT_CLIENT_API_KEY to override the per-provider lookup, and A2A_AGENT_CLIENT_BASE_URL to point at the Inference Gateway (recommended - it normalizes provider quirks so the same agent code talks to every provider unchanged) or any other OpenAI-compatible endpoint. The full A2A_AGENT_CLIENT_* surface is in the environment variable reference.

NVIDIA serves the build.nvidia.com NIM catalog (Nemotron, Llama, DeepSeek, Mistral, Qwen) with bearer-token auth at https://integrate.api.nvidia.com/v1. See Supported Providers for the full matrix, auth modes, default URLs, and vision support.

If your agent is going to receive images, A2A_AGENT_CLIENT_MODEL has to name a vision-capable model - see Sending images to an agent.

The server and its builder

A2AServer is the runtime that terminates the A2A JSON-RPC protocol. You never construct it directly - A2AServerBuilder assembles one with a fluent interface, and A2AServer::serve(addr) binds the listener (plaintext, or TLS when Config.tls_config is enabled) and - when artifacts are enabled - spawns the artifacts server and retention loop alongside it.

Method Purpose
with_config(Config) Apply a fully-loaded Config - see the note below for what is actually consumed.
with_agent(Agent) Attach an LLM-backed agent built via AgentBuilder.
with_agent_card(AgentCard) Configure the card served at /.well-known/agent-card.json from an in-memory value.
with_agent_card_from_file(path, Option<AgentCardOverrides>) Load the card from a JSON file, applying optional field overrides.
with_gateway_url(url) Override the Inference Gateway base URL (default http://gateway:8080/v1).
with_storage(Arc<dyn Storage>) Swap the task store (InMemoryStorage default; RedisStorage behind the redis feature).
with_background_task_handler(h) Register a custom SendMessage handler.
with_streaming_task_handler(h) Register a custom SendStreamingMessage handler.
with_default_background_task_handler() Opt into the bundled SendMessage default.
with_default_streaming_task_handler() Opt into the bundled SendStreamingMessage default.
with_default_task_handlers() Opt into both defaults at once.
with_workers(n) Number of background queue workers (defaults to A2A_QUEUE_WORKERS).
with_auth_verifier(Arc<dyn AuthVerifier>) Plug in a custom verifier (overrides A2A_AUTH_ENABLED).
with_artifact_service(Arc<dyn ArtifactService>) Supply a custom artifact service / storage backend.

What with_config consumes. The agent-card URL fallback, TLS, auth, queue, the usage-metadata flag and artifacts. server_config.port only feeds the default advertised URL http(s)://localhost:<port>/a2a - it never binds a listener, which is always the SocketAddr you pass to A2AServer::serve. telemetry_config is read only by telemetry::init, which you call yourself.

Builder validation. build() returns an error unless an agent card is configured and at least one task handler is present. It also cross-checks the card's capabilities.streaming flag: a streaming-enabled card requires a streaming handler, and a streaming-disabled card requires a background handler. with_default_task_handlers() satisfies both.

The default handlers delegate to the registered Agent when one is present (via with_agent); without an agent they return a built-in echo reply, which is what makes the minimal server above work with no LLM wired up.

AgentBuilder

AgentBuilder constructs the OpenAI-compatible Agent that lives inside the server. Seed it with an entire AgentConfig via with_config(&cfg), or set fields individually - per-field setters layered on top of a config override that field only.

use inference_gateway_adk::AgentBuilder;

// Driven entirely by AgentConfig (provider, model, key, ...)
let agent = AgentBuilder::new()
    .with_config(&config.agent_config)
    .with_toolbox(tools)
    .build()
    .await?;

// Or with explicit per-field setters
let agent = AgentBuilder::new()
    .with_provider("deepseek")
    .with_model("deepseek-v4-flash")
    .with_system_prompt("You are a helpful assistant")
    .with_max_chat_completion_iterations(10)
    .build()
    .await?;
Method Purpose
with_config(&AgentConfig) Seed the builder from an AgentConfig. Later setters win per field.
with_provider(s) / with_model(s) LLM provider and model id. Both are required - build() fails fast if either is unset.
with_api_key(s) Provider API key.
with_base_url(s) Override the LLM gateway base URL for the default client.
with_timeout(Duration) / with_max_retries(u32) Bounds every LLM request and each wait for a streamed event, plus the retry budget. Duration::ZERO disables the bound.
with_max_chat_completion_iterations(u32) Cap on chat-completion round trips in the agent tool loop (A2A_AGENT_CLIENT_MAX_CHAT_COMPLETION_ITERATIONS, default 10). with_max_chat_completion is an alias.
with_max_tokens(u32) Token ceiling per completion. Applies to non-streaming completions only - the gateway SDK omits max_tokens from streaming requests.
with_temperature(f64) Sampling temperature (0.0 - 2.0), sent with streaming and non-streaming completions alike. Leave it unset to keep the gateway default.
with_system_prompt(s) System prompt prepended to every conversation.
with_enable_usage_metadata(bool) Override whether terminal tasks carry token usage + execution stats as the usage extension (default from config, true).
with_max_conversation_history(u32) Max conversation-history messages retained (default 20).
with_toolbox(Vec<ChatCompletionTool>) Declare the tool schema advertised to the model.
with_tool_handler(name, h) Register a ToolHandler for a named tool.
with_function_tool(name, closure) Register a synchronous closure tool handler.
with_async_function_tool(name, closure) Register an async closure tool handler.
with_llm_client(C: LLMClient) Replace the default OpenAICompatibleLLMClient with a custom transport.

AgentBuilder::build() fails fast when provider or model are unset (or the provider is unsupported), so a misconfigured server errors out at startup instead of on the first chat request.

Temperature

AgentConfig::temperature is an Option<f64> in the range 0.0 - 2.0, set from A2A_AGENT_CLIENT_TEMPERATURE or AgentBuilder::with_temperature(f64). It is forwarded to streaming and non-streaming chat completions alike, through the gateway SDK's InferenceGatewayClient::with_temperature.

Leaving the variable unset - or setting it to an empty string - keeps the value at None, in which case the ADK sends no temperature at all and the gateway's own default stands. There is no ADK-side default to override.

A2A_AGENT_CLIENT_TEMPERATURE=0.2

The setting was removed in the first release after 0.12.1, because the Rust gateway SDK had no temperature knob and the value was silently dropped. It is back now that inference-gateway-sdk 0.28.0 exposes one, with the same env var and builder names as before.

The release that removed it also made four previously inert settings effective: A2A_AGENT_CLIENT_API_KEY (sent as a bearer token), MAX_TOKENS, TIMEOUT_SECS and MAX_CHAT_COMPLETION_ITERATIONS. An agent that carried unused values for these will now behave according to them, so review them before upgrading.

Custom LLM clients

The default OpenAICompatibleLLMClient wraps the Inference Gateway SDK. To route requests through a different backend - or a mock for tests - implement the LLMClient trait (create_chat_completion + create_streaming_chat_completion, mirroring the Go ADK) and pass it to with_llm_client(...):

use inference_gateway_adk::{AgentBuilder, OpenAICompatibleLLMClient};

// Build the default client explicitly (synchronous; no await)
let llm_client = OpenAICompatibleLLMClient::new(&config.agent_config)?;

let agent = AgentBuilder::new()
    .with_llm_client(llm_client)
    .with_system_prompt("You are a coding assistant.")
    .build()
    .await?;

Custom tools

Tools are declared with the Inference Gateway SDK's ChatCompletionTool / FunctionObject types (the schema the model sees) and backed by a handler keyed on the tool name. Register a synchronous closure with with_function_tool, an async closure with with_async_function_tool, or any ToolHandler implementation with with_tool_handler. When the model emits a tool call, the matching handler runs and its return value is appended to the conversation as a tool message.

use inference_gateway_adk::AgentBuilder;
use inference_gateway_sdk::{
    ChatCompletionTool, ChatCompletionToolType, FunctionObject, FunctionParameters,
};
use serde_json::{Value, json};

let tools = vec![ChatCompletionTool {
    type_: ChatCompletionToolType::Function,
    function: FunctionObject {
        name: "search_web".to_string(),
        description: Some("Search the web for information".to_string()),
        parameters: Some(FunctionParameters(
            json!({
                "type": "object",
                "properties": {
                    "query": { "type": "string" },
                    "limit": { "type": "integer", "default": 5 }
                },
                "required": ["query"]
            })
            .as_object()
            .unwrap()
            .clone(),
        )),
        strict: false,
    },
}];

let agent = AgentBuilder::new()
    .with_config(&config.agent_config)
    .with_system_prompt("You can answer questions and search the web.")
    .with_toolbox(tools)
    .with_function_tool("search_web".to_string(), |args: Value| {
        let query = args["query"].as_str().unwrap_or("");
        Ok(json!({ "query": query, "results": [] }).to_string())
    })
    .build()
    .await?;

Use with_async_function_tool when the handler needs to .await (HTTP calls, database lookups). The closure signature is the same except it returns a future:

let agent = AgentBuilder::new()
    .with_config(&config.agent_config)
    .with_toolbox(tools)
    .with_async_function_tool("search_web".to_string(), |args: Value| async move {
        let query = args["query"].as_str().unwrap_or("").to_string();
        let results = my_async_search(&query).await?;
        Ok(serde_json::to_string(&results)?)
    })
    .build()
    .await?;

See examples/ai-powered/ for a multi-tool walkthrough (weather, math, search).

MCP client

The Rust ADK ships an optional Model Context Protocol (MCP) client that connects an agent to one or more MCP servers over Streamable HTTP, discovers the tools they expose, and lets the LLM invoke them - the same selector-tool design as the Go ADK, so an agent reaches a shared MCP server (or a fleet of them) identically across ADKs. It is disabled by default (MCP_ENABLED=false) and only makes sense with an LLM-backed agent - the selector tools are useless without a model to drive them.

The selector pattern

A single MCP server can expose dozens or hundreds of tools, and loading every schema into the model context would overwhelm the window. So the client does not register MCP tools individually. It registers just two selector tools onto the agent, regardless of how many tools the servers offer:

  • mcp_list_tools - lists discovered tools (server, name, description, input schema); accepts an optional search filter.
  • mcp_call_tool - invokes a tool by name (optionally disambiguated by server) with an arguments object.

A typical turn: the model calls mcp_list_tools to find a relevant tool, reads its input schema, then calls mcp_call_tool to run it. Only the metadata the model actually asks for reaches the context window. The catalog is discovered in the background and refreshed on an interval, so mcp_list_tools returns fast from an in-memory snapshot.

Wiring it up

McpClient::from_config builds the client from an McpConfig, and AgentBuilder::with_mcp_client registers it onto the agent. Like the artifacts subsystem, MCP config loads under its own MCP_ prefix, separate from the A2A_ Config:

use inference_gateway_adk::{AgentBuilder, Config, McpClient, McpConfig};

let config: Config = envy::prefixed("A2A_").from_env()?;
let mcp_config = envy::prefixed("MCP_")
    .from_env::<McpConfig>()
    .unwrap_or_default();

let mut builder = AgentBuilder::new()
    .with_config(&config.agent_config)
    .with_system_prompt("You are a helpful assistant.");

// Attach the client only when MCP is enabled. It connects and refreshes the
// tool catalog in the background, so it never blocks startup when a server is
// still coming up.
if mcp_config.enable {
    builder = builder.with_mcp_client(McpClient::from_config(&mcp_config)?);
}

let agent = builder.build().await?;

Connecting in the background means the agent starts even when an MCP server comes up later or restarts independently: each server retries with exponential backoff (MCP_RETRY_INTERVAL doubling up to MCP_RETRY_MAX_INTERVAL, forever by default with MCP_MAX_RETRIES=0), and a failing server never drops the catalog of the healthy ones.

Limitations

Matching the Go ADK:

  • Transport is Streamable HTTP only; stdio/subprocess MCP servers are not wired.
  • Tool listing fetches a single page - cursor-based pagination is not yet handled.
  • MCP resources and prompts are not exposed; tools only.

The MCP_* tuning knobs and their defaults are in the environment variable reference.

Custom task handlers

The server's two extension points for task execution are the TaskHandler trait (for SendMessage, the background/queue path) and StreamableTaskHandler (for SendStreamingMessage, the SSE path). The defaults wired in by with_default_task_handlers() delegate to the registered Agent; implement either trait to plug in custom logic. A TaskHandler runs on a queue worker while the SendMessage request waits for the task to settle - see Blocking SendMessage.

use async_trait::async_trait;
use inference_gateway_adk::{
    A2AServerBuilder, TaskHandler,
    a2a_types::{Message, Part, Role, Task, TaskState, TaskStatus},
};

#[derive(Debug)]
struct EchoHandler;

#[async_trait]
impl TaskHandler for EchoHandler {
    async fn handle_task(&self, task: Task, message: Option<Message>) -> anyhow::Result<Task> {
        let Some(message) = message else {
            return Ok(task);
        };
        let reply = message
            .parts
            .iter()
            .filter_map(|p| p.text.as_deref())
            .collect::<Vec<_>>()
            .join(" ");

        let mut updated = task.clone();
        updated.status = TaskStatus {
            state: TaskState::TaskStateCompleted,
            ..updated.status
        };
        updated.history.push(Message {
            role: Role::RoleAgent,
            parts: vec![Part { text: Some(reply), ..Default::default() }],
            ..message
        });
        Ok(updated)
    }
}

let server = A2AServerBuilder::new()
    .with_agent_card_from_file(".well-known/agent-card.json", None)
    .with_background_task_handler(EchoHandler)
    .build()
    .await?;

Answering without a task

TaskHandler has a second, optional method: handle_message. The server calls it before it creates or continues any task, and when it returns Some(message) that Message is the SendMessage result - no task is created, nothing is enqueued, and no push notification fires (A2A spec 3.2.2). The default implementation returns None, which keeps the task-based flow above, so existing handlers need no change.

#[async_trait]
impl TaskHandler for GreetHandler {
    async fn handle_task(&self, task: Task, message: Option<Message>) -> anyhow::Result<Task> {
        // ... the regular task flow
    }

    async fn handle_message(&self, message: &Message) -> anyhow::Result<Option<Message>> {
        if !is_greeting(message) {
            return Ok(None); // not ours - fall through to the task flow
        }
        Ok(Some(Message {
            message_id: uuid::Uuid::new_v4().to_string(),
            parts: vec![Part { text: Some("Hello!".to_string()), ..Default::default() }],
            role: Role::RoleAgent,
            task_id: None,
            ..message.clone()
        }))
    }
}

Reach for this when task bookkeeping buys nothing: a health ping, a capability question, a greeting. Anything the caller may want to poll, cancel, or receive artifacts from belongs in a task.

Streaming handlers receive a StreamEmitter and push status updates and artifacts over the SSE stream as work progresses - see Emitting artifacts from a streaming handler for a worked example. For runnable demos, see examples/streaming/ and the TaskStateInputRequired flow in examples/input-required/.

Usage metadata

With enable_usage_metadata on (A2A_AGENT_CLIENT_ENABLE_USAGE_METADATA, default true), the default handlers attach the task's token usage and execution stats to its metadata on terminal states. The ADK publishes them as the usage extension:

  • build() declares the extension in the agent card and the extended card, with required: false.
  • A request activates it with the A2A-Extensions header, and the response echoes the URI. Tasks returned to a request that did not activate it leave the keys out, and push notifications never carry them. The stored task always keeps them.
  • The keys are USAGE_METADATA_KEY and EXECUTION_STATS_METADATA_KEY, both prefixed by USAGE_EXTENSION_URI. Task::without_extension(uri) drops an extension's keys from a task.

A2AClient activates extensions through ClientConfig::extensions, sent as the A2A-Extensions header on every request:

use inference_gateway_adk::{A2AClient, ClientConfig, USAGE_EXTENSION_URI};

let mut config = ClientConfig::new("http://localhost:8080");
config.extensions = vec![USAGE_EXTENSION_URI.to_string()];
let client = A2AClient::with_config(config)?;

A2AClient

A2AClient is the typed client for talking to an A2A server. Construct it with a base URL, or with a ClientConfig for explicit timeout and retry control and the extensions every request activates:

use inference_gateway_adk::A2AClient;

let client = A2AClient::new("http://localhost:8080")?;

// Discovery endpoints (always public, no auth required)
let agent_card = client.get_agent_card().await?;
let health = client.get_health().await?;

JSON-RPC method helpers

The client exposes a typed helper for every method in the A2A specification. Each takes a request struct and returns the matching response struct from inference_gateway_adk::a2a_types. Runnable end-to-end examples - one client binary per method - live in examples/a2a-methods/.

Method A2AClient helper Request type Response type
SendMessage send_message SendMessageRequest SendMessageResponse
SendStreamingMessage send_streaming_message SendMessageRequest SendMessageResponse
GetTask get_task GetTaskRequest Task
ListTasks list_tasks ListTasksRequest ListTasksResponse
CancelTask cancel_task CancelTaskRequest Task
SubscribeToTask resubscribe_task SubscribeToTaskRequest Stream<StreamResponse> (SSE)
CreateTaskPushNotificationConfig set_task_push_notification_config TaskPushNotificationConfig TaskPushNotificationConfig
GetTaskPushNotificationConfig get_task_push_notification_config GetTaskPushNotificationConfigRequest TaskPushNotificationConfig
ListTaskPushNotificationConfigs list_task_push_notification_configs ListTaskPushNotificationConfigsRequest ListTaskPushNotificationConfigsResponse
DeleteTaskPushNotificationConfig delete_task_push_notification_config DeleteTaskPushNotificationConfigRequest serde_json::Value
GetExtendedAgentCard get_authenticated_extended_card GetExtendedAgentCardRequest AgentCard

The method column holds the A2A v1.0.1 wire names, generated from the canonical schema as the A2aMethod enum (A2aMethod::SendMessage, A2aMethod::GetTask, ...). Match on the enum rather than on string literals when you dispatch raw JSON-RPC yourself; the v0.x slash names (message/send, tasks/get, ...) answer -32601 Method not found.

Request params follow the normative proto3 JSON mapping of A2A spec section 1.4, so the server accepts either spelling of every field name - the lowerCamelCase form (pageSize, contextId) or the proto3 form (page_size, context_id) - and ignores params it does not know instead of answering -32602 Invalid params. Hand-rolled clients and clients built against a newer spec revision therefore interoperate without stripping extra fields first. Method names are unaffected: they stay exact-match PascalCase.

POST /a2a accepts application/json, any +json media type, and a missing Content-Type header. Anything else answers a JSON-RPC -32005 (CONTENT_TYPE_NOT_SUPPORTED) envelope with HTTP 200 rather than a bodiless HTTP 415, so a client that sent the wrong header reads the offending type out of error.data like every other A2A error.

A representative SendMessage call, using the typed structs end-to-end:

use inference_gateway_adk::a2a_types::{Message, Part, Role, SendMessageRequest};

let response = client
    .send_message(SendMessageRequest {
        configuration: None,
        message: Message {
            context_id: None,
            extensions: vec![],
            message_id: uuid::Uuid::new_v4().to_string(),
            metadata: None,
            parts: vec![Part {
                data: None,
                file: None,
                metadata: None,
                text: Some("Hello from the typed client".to_string()),
            }],
            reference_task_ids: vec![],
            role: Role::RoleUser,
            task_id: None,
        },
        metadata: None,
        tenant: Some("example".to_string()),
    })
    .await?;

let task = response.task.expect("server returned a task");

Blocking SendMessage

SendMessage is blocking: the server enqueues the task and holds the JSON-RPC response until that task reaches a terminal (TaskStateCompleted, TaskStateFailed, TaskStateCanceled, TaskStateRejected) or an interrupted (TaskStateInputRequired, TaskStateAuthRequired) state, per A2A spec section 3.2.2. The reply therefore carries a finished task, not a TaskStateSubmitted placeholder. A client written against an earlier release, which answered immediately and expected to poll from TaskStateSubmitted, must now pass configuration.returnImmediately: true to keep that behaviour.

SendMessageConfiguration carries the three fields that tune the call:

Field Type Effect
returnImmediately Option<bool> true answers with the freshly submitted task and leaves the client to poll GetTask or subscribe.
historyLength Option<i32> Caps the response to the last n history messages (0 strips history entirely).
taskPushNotificationConfig Option<TaskPushNotificationConfig> Registers a webhook for the task inline, saving a CreateTaskPushNotificationConfig round trip. Leave taskId unset - the server fills it in with the task it just created, and generates an id when none is given.
use inference_gateway_adk::a2a_types::SendMessageConfiguration;

let response = client
    .send_message(SendMessageRequest {
        configuration: Some(SendMessageConfiguration {
            history_length: Some(5),
            return_immediately: Some(true),
            ..Default::default()
        }),
        message,
        metadata: None,
        tenant: Some("example".to_string()),
    })
    .await?;

Long-running agents. A task that takes minutes holds the HTTP request open for minutes, and clients usually hit their own read timeout first. The server stops waiting after 30 seconds and answers with the task's latest known state, which for a slow handler is still TaskStateWorking. For anything slower than that, either set returnImmediately: true and poll, or use SendStreamingMessage so progress arrives as it happens.

Continuing an existing task

A SendMessage whose message.taskId names a stored task appends the message to that task's history, puts it back into TaskStateSubmitted, and re-runs it instead of creating a new task. This is how a client answers a task the agent paused in TaskStateInputRequired:

let paused = response.task.expect("server returned a task");

let answer = client
    .send_message(SendMessageRequest {
        configuration: None,
        message: Message {
            context_id: paused.context_id.clone(),
            extensions: vec![],
            message_id: uuid::Uuid::new_v4().to_string(),
            metadata: None,
            parts: vec![Part {
                text: Some("Berlin".to_string()),
                ..Default::default()
            }],
            reference_task_ids: vec![],
            role: Role::RoleUser,
            task_id: Some(paused.id.clone()),
        },
        metadata: None,
        tenant: None,
    })
    .await?;

Three things are rejected rather than silently creating a task:

  • an unknown taskId answers -32001 (TaskNotFound),
  • a task already in a terminal state answers -32004 (UnsupportedOperation) - it accepts no further input,
  • a message.contextId that disagrees with the task's own answers -32602 (InvalidParams). Omitting contextId is fine; the server infers it from the task.

SendMessageRequest.message is a required value-typed Message - pass it directly, not wrapped in Some(..). Every tenant field is optional (Option<String>), as are the filters: page_size / page_token on ListTaskPushNotificationConfigsRequest, and context_id, status and last_updated_after on ListTasksRequest. The resource identifiers stay plain String - id on GetTaskRequest / CancelTaskRequest / SubscribeToTaskRequest, task_id and id on the pushNotificationConfig get and delete requests, and task_id on the list request.

Sending images to an agent

A user message can carry image file parts alongside its text. The agent forwards them to the configured LLM as OpenAI-compatible image_url content parts, so a vision-capable model can read a screenshot, a diagram, or a captcha:

use base64::{Engine as _, engine::general_purpose::STANDARD};
use inference_gateway_adk::a2a_types::{FilePart, Message, Part, Role, SendMessageRequest};

let encoded = STANDARD.encode(std::fs::read("captcha.png")?);

let message = Message {
    context_id: None,
    extensions: vec![],
    message_id: uuid::Uuid::new_v4().to_string(),
    metadata: None,
    parts: vec![
        Part {
            text: Some("What characters are in this captcha?".to_string()),
            ..Default::default()
        },
        Part {
            file: Some(FilePart {
                file_with_bytes: Some(encoded.parse()?),
                file_with_uri: None,
                media_type: Some("image/png".to_string()),
                name: Some("captcha.png".to_string()),
            }),
            ..Default::default()
        },
    ],
    reference_task_ids: vec![],
    role: Role::RoleUser,
    task_id: None,
};

let response = client
    .send_message(SendMessageRequest {
        configuration: None,
        message,
        metadata: None,
        tenant: Some("example".to_string()),
    })
    .await?;
  • Optional metadata. FilePart::media_type and FilePart::name are Option<String>, so a literal takes Some(..) and reading one back means matching on the Option. Set media_type on image parts - it is what the agent uses to build the data: URL it forwards to the model.
  • Bytes vs URI. fileWithBytes is inlined as a data:<mediaType>;base64,<bytes> URL. fileWithUri is passed through unchanged, so the provider has to fetch it. Artifact-server URLs that only resolve inside your cluster are typically unreachable from a hosted provider; send bytes in that case. When both are set, the bytes win.
  • Ordering. Text and image parts reach the model in the order they appear in the A2A message. Text-only messages are unchanged - they keep plain string content.
  • What is forwarded. Only image/* parts on user-role messages. Other media types (PDF, audio) are skipped, and file parts on agent-role messages are never converted, because OpenAI-compatible assistant messages cannot carry images. A message whose only part is a skipped file is dropped from the history entirely.
  • Model. There is no config flag for this. The operator picks the model with A2A_AGENT_CLIENT_MODEL, and a model without vision support rejects the request - the task ends as failed carrying the provider error. That is the expected outcome, not a bug.
  • Agent card. An agent that accepts images should advertise it in its card JSON, e.g. "defaultInputModes": ["text/plain", "image/png"], so clients can discover the capability.

No extra wiring is needed: this works with the default task handlers and with custom handlers that run the agent's LLM client over the task history.

Health monitoring

get_health() returns the agent's HealthStatus for service discovery and load-balancer probes. The status string is one of:

  • healthy - fully operational.
  • degraded - partially operational; some functionality may be limited.
  • unhealthy - not operational or experiencing significant issues.

Push notifications

A2A servers persist per-task webhook configurations through the four push-notification-config control-plane methods on A2AClient (set, get, list, delete). Each uses the typed structs from a2a_types and is exercised by a dedicated binary under examples/a2a-methods/.

The card must advertise the capability. All four methods - CreateTaskPushNotificationConfig, GetTaskPushNotificationConfig, ListTaskPushNotificationConfigs, DeleteTaskPushNotificationConfig - are gated on capabilities.pushNotifications: true in the served agent card. Against a card that omits the flag or sets it to false, the server rejects every one of them with -32003 (PushNotificationNotSupported) instead of storing the config. Earlier releases accepted and stored configs regardless of the card.

{
  "capabilities": {
    "pushNotifications": true
  }
}
use inference_gateway_adk::a2a_types::TaskPushNotificationConfig;

client
    .set_task_push_notification_config(TaskPushNotificationConfig {
        authentication: None,
        id: Some("primary".to_string()),
        task_id: Some(task_id.clone()),
        tenant: Some("example".to_string()),
        token: Some("shared-secret".to_string()),
        url: "https://your-app.example/webhooks/a2a".to_string(),
    })
    .await?;

The same struct can ride along with the message that creates the task - see taskPushNotificationConfig on SendMessageConfiguration - which saves the extra round trip and guarantees no state change is missed between creating the task and registering the webhook.

Webhook delivery

Every task update is POSTed to each webhook registered for that task. The body is a JSON StreamResponse whose task field holds the task snapshot, which is the same payload shape the SSE stream and the Go and TypeScript ADKs use, so one receiver serves all three:

{
  "task": {
    "id": "f81d4fae-7dec-11d0-a765-00a0c91e6bf6",
    "contextId": "6c8f1b2e-0f25-4f28-9f61-2b8a2d1fdd0c",
    "status": { "state": "TASK_STATE_COMPLETED", "timestamp": "2026-10-03T12:00:00Z" },
    "history": [],
    "artifacts": []
  }
}

Two headers are set from the stored config, each only when that field is present:

Header Source Value
Authorization authentication <scheme> <credentials> from AuthenticationInfo
X-A2A-Notification-Token token The token verbatim

A webhook that needs its deliveries gated must reject unauthenticated requests itself: with neither field set, the POST goes out unauthenticated.

Delivery is at most once. Each POST gets a 10-second timeout; a failure is logged and not retried, and a slow webhook never fails the task that triggered it. Treat notifications as a hint to call GetTask, not as the system of record - a receiver that must not miss a transition should reconcile with GetTask or SubscribeToTask.

Agent card and metadata

The agent card served at /.well-known/agent-card.json is the discovery document for your agent. Its name, description, version, supportedInterfaces, and capabilities come from the card JSON you hand the builder - either inline via with_agent_card(...) or from disk via with_agent_card_from_file(path, overrides). There are no environment variables for card fields: the card is the single source of truth, and the path is the path argument, not a configured value.

Card caching

The card endpoint is served with HTTP caching headers per A2A spec section 8.6: Cache-Control: public, max-age=300, an ETag derived from the card body, and a Last-Modified of server start. The card is fixed for the life of the process, so a conditional request carrying a matching If-None-Match - or an If-Modified-Since equal to the advertised Last-Modified - is answered with 304 Not Modified and no body. The five-minute freshness window is fixed, not configurable.

$ curl -sI http://localhost:8080/.well-known/agent-card.json
HTTP/1.1 200 OK
cache-control: public, max-age=300
etag: "a3f1c0d49b2e7f58"
last-modified: Sat, 03 Oct 2026 12:00:00 GMT

Runtime overrides layer on top of whatever was loaded from disk. Pass AgentCardOverrides to with_agent_card_from_file(...); the file supplies the baseline and each explicitly-set override wins:

use inference_gateway_adk::{A2AServerBuilder, AgentCardOverrides, Config};

let config: Config = envy::prefixed("A2A_").from_env()?;

let server = A2AServerBuilder::new()
    .with_config(config)
    .with_agent_card_from_file(
        ".well-known/agent-card.json",
        Some(
            AgentCardOverrides::new()
                .with_name("Development Weather Assistant")
                .with_description("Development version with debug features")
                .with_version("dev-1.0.0")
                .with_url("http://localhost:8080/a2a"),
        ),
    )
    .build()
    .await?;

AgentCardOverrides exposes with_name, with_description, with_version, and with_url - with_url rewrites the url of the first entry in the card's supportedInterfaces list, appending a JSONRPC / 1.0 interface when the list is empty. See examples/static-agent-card/ for a runnable demo.

The advertised URL in supportedInterfaces[0].url is resolved at startup, highest precedence first:

  1. AgentCardOverrides::with_url
  2. A2A_AGENT_URL
  3. The card's own supportedInterfaces[0].url, when non-empty
  4. http://localhost:<A2A_SERVER_PORT>/a2a - https:// when A2A_SERVER_TLS_ENABLED=true

The same resolution fills in the URL when the card declares no supportedInterfaces entry at all, so an agent serving on the default port advertises http://localhost:8080/a2a without any configuration. Set A2A_AGENT_URL to the externally reachable address whenever the agent runs behind a container name, service DNS record, or ingress - clients and the gateway dial exactly what the card advertises.

Capability-gated methods

Two capabilities flags gate methods outright - a call whose flag is absent or false answers a JSON-RPC error instead of running:

Method Required flag Error when the flag is missing
CreateTaskPushNotificationConfig, GetTaskPushNotificationConfig, ListTaskPushNotificationConfigs, DeleteTaskPushNotificationConfig capabilities.pushNotifications -32003 (PushNotificationNotSupported)
GetExtendedAgentCard capabilities.extendedAgentCard -32004 (UnsupportedOperation)

See Push notifications for the config methods and Card-driven authentication flow for how the extended-card flag is set, its full error contract, and how clients discover which schemes the agent accepts.

Authentication

When A2A_AUTH_ENABLED=true, the server gates POST /a2a behind an Authorization: Bearer <token> header. GET /health and GET /.well-known/agent-card.json stay public so probes and discovery clients keep working without a credential. Tokens that fail validation get HTTP 401 with a WWW-Authenticate: Bearer realm="a2a" header.

The bundled OidcJwtVerifier:

  1. Performs OIDC discovery at <A2A_AUTH_ISSUER_URL>/.well-known/openid-configuration.
  2. Fetches and caches the JWKS advertised by the discovery document.
  3. Validates the JWT signature, iss, exp, and the aud claim - always checked against A2A_AUTH_CLIENT_ID.
Variable Default Purpose
A2A_AUTH_ENABLED false When true, POST /a2a requires a valid bearer token.
A2A_AUTH_ISSUER_URL (empty) OIDC issuer; the server performs discovery + JWKS lookup against it. Required when auth is enabled.
A2A_AUTH_CLIENT_ID (empty) Validated as the JWT audience (aud). Required when auth is enabled.
A2A_AUTH_CLIENT_SECRET (empty) Required when auth is enabled, even though token verification never uses it. Reserved for client-side OAuth2 flows.

All three of A2A_AUTH_ISSUER_URL, A2A_AUTH_CLIENT_ID, and A2A_AUTH_CLIENT_SECRET must be non-empty when A2A_AUTH_ENABLED=true. If any is missing, building the verifier fails with AUTH_ISSUER_URL, AUTH_CLIENT_ID, and AUTH_CLIENT_SECRET are required when AUTH_ENABLED=true and build() returns that error instead of starting the server.

On success the verifier produces an AuthenticatedPrincipal - subject (sub), tenant (first of tenant/tid/organization), issuer, and the full claims map - and attaches it to the request as an Axum extension so the JSON-RPC dispatcher can scope behaviour by tenant.

To plug in a custom backend (a static signing key, an internal identity service, a mock for tests), implement AuthVerifier and pass it to with_auth_verifier(...). This overrides whatever A2A_AUTH_ENABLED selects, the same way with_storage(...) overrides the task store:

use async_trait::async_trait;
use inference_gateway_adk::{AuthError, AuthVerifier, AuthenticatedPrincipal};

#[derive(Debug)]
struct StaticToken(&'static str);

#[async_trait]
impl AuthVerifier for StaticToken {
    async fn verify(&self, token: &str) -> Result<AuthenticatedPrincipal, AuthError> {
        if token == self.0 {
            Ok(AuthenticatedPrincipal {
                subject: "demo-user".to_string(),
                tenant: "demo-tenant".to_string(),
                issuer: "static".to_string(),
                claims: Default::default(),
            })
        } else {
            Err(AuthError::InvalidToken("unrecognized token".to_string()))
        }
    }
}

let server = A2AServerBuilder::new()
    .with_agent_card_from_file(".well-known/agent-card.json", None)
    .with_default_task_handlers()
    .with_auth_verifier(std::sync::Arc::new(StaticToken("demo-token-123")))
    .build()
    .await?;

When auth is disabled the middleware is not attached, and GetExtendedAgentCard returns the configured card whenever capabilities.extendedAgentCard == true. See examples/auth/ for an end-to-end demo that runs both a static-token verifier and a Keycloak-backed OidcJwtVerifier.

Card-driven authentication flow

Beyond validating bearer tokens, the ADK implements the A2A spec, section 7 card-driven auth model: the agent card advertises how to authenticate, and clients transmit credentials obtained out-of-band on every request. A2A does not run OAuth flows in-protocol.

  1. Discovery - the client fetches the public card from /.well-known/agent-card.json (always unauthenticated). The card declares securitySchemes (named schemes the agent accepts: apiKey, http, oauth2, openIdConnect, mutualTLS) and securityRequirements (a requirement list with OR-of-ANDs semantics - satisfying any one entry is sufficient).
  2. Credential acquisition is out-of-band - the client obtains a token/key however the chosen scheme dictates.
  3. Transmission - the client sends the credential (e.g. Authorization: Bearer <token>) on every request.
  4. Server enforcement - with A2A_AUTH_ENABLED=true the POST /a2a endpoint is protected; unauthenticated requests get 401 with a WWW-Authenticate: Bearer realm="a2a" challenge.
  5. Extended card - if the card sets capabilities.extendedAgentCard: true, an authenticated client MAY call GetExtendedAgentCard for a richer card and SHOULD replace its cached public card with the response.

Declaring security schemes on the card

With A2A_AUTH_ENABLED=true the served card must declare securitySchemes so clients can discover how to authenticate. The oidc_security_schemes helper derives that declaration from the auth config at startup - keyed "openId" with a discovery URL built from the OIDC issuer, plus a matching securityRequirements entry:

use inference_gateway_adk::{oidc_security_schemes, Config};

let config: Config = envy::prefixed("A2A_").from_env()?;

let (schemes, security) = oidc_security_schemes(&config.auth);
card.security_schemes = schemes;
card.security_requirements = security;

This mirrors Go's server.OIDCSecuritySchemes. It matters because ADL manifests deliberately exclude OIDC/OAuth2 from their card definitions - so an ADL-generated agent derives its securitySchemes/securityRequirements from the running auth config at startup rather than hard-coding an issuer into the manifest. The resulting card fragment looks like:

{
  "securitySchemes": {
    "openId": {
      "openIdConnectSecurityScheme": {
        "openIdConnectUrl": "https://issuer.example.com/realms/app/.well-known/openid-configuration"
      }
    }
  },
  "securityRequirements": [{ "schemes": { "openId": { "list": [] } } }]
}

If A2A_AUTH_ENABLED=true but the card declares no securitySchemes (or the inverse - a card advertises schemes while auth is disabled), build() logs a startup warning: the two declarations disagree and discovery would otherwise be broken.

The authenticated extended card

The extended card is served only to authenticated callers via GetExtendedAgentCard. Configure it with the builder; with_extended_agent_card also forces capabilities.extendedAgentCard: true on the public card:

let server = A2AServerBuilder::new()
    .with_agent_card_from_file(".well-known/agent-card.json", None)
    .with_extended_agent_card(extended_card) // extra skills, capability detail, ...
    .with_default_task_handlers()
    .build()
    .await?;

The error contract (spec 3.3.4) for GetExtendedAgentCard:

Card state Result
capabilities.extendedAgentCard absent or false -32004 (UnsupportedOperation)
capabilities.extendedAgentCard: true, no extended configured -32007 (AuthenticatedExtendedCardNotConfigured)
capabilities.extendedAgentCard: true, extended configured the extended card is returned

Card shape. From 0.15.0 the Rust ADK card type is the A2A v1.0.1 AgentCard: the extended-card flag lives on capabilities.extendedAgentCard, security requirements on securityRequirements, and the endpoint list on the required supportedInterfaces array (there is no top-level url / preferredTransport / protocolVersion, and no capabilities.stateTransitionHistory). See Deprecated card fields for the mapping from the older names.

TLS and mTLS

When A2A_SERVER_TLS_ENABLED=true, A2AServer::serve swaps its plaintext listener for axum-server backed by rustls 0.23 (with the ring crypto provider) and serves the same router over HTTPS. Rustls was chosen over native-tls because it is pure Rust - avoiding the OpenSSL toolchain on container builds - and because it gives programmatic access to the negotiated connection, which is what makes the mTLS subject extraction below tractable.

Variable Default Purpose
A2A_SERVER_TLS_ENABLED false When true, A2AServer::serve binds an HTTPS listener.
A2A_SERVER_TLS_CERT_PATH (empty) PEM file with the server certificate chain.
A2A_SERVER_TLS_KEY_PATH (empty) PEM file with the server private key (PKCS#1, PKCS#8, or SEC1).
A2A_SERVER_TLS_CLIENT_CA_PATH (unset) When set, the server requires mTLS and trusts client certificates signed by any CA in this PEM bundle - the MutualTlsSecurityScheme the A2A spec describes. Leave it out entirely for plain TLS: any value, an empty string included, makes TlsConfig::client_ca_path Some(..) and turns mTLS on, so A2A_SERVER_TLS_CLIENT_CA_PATH="" fails at startup trying to read a certificate from an empty path.

When mTLS is enabled, the TLS acceptor parses the peer's leaf certificate and exposes it to handlers as an axum::Extension<PeerCert> - the same plumbing pattern the bearer-token middleware uses for AuthenticatedPrincipal. The wrapped ClientCertPrincipal carries the subject DN, the Common Name (when present), the issuer DN, and the raw DER bytes of the leaf:

use axum::Extension;
use inference_gateway_adk::PeerCert;

async fn my_handler(Extension(peer): Extension<PeerCert>) {
    if let Some(p) = peer.0 {
        tracing::info!("authenticated client: {} (issued by {})", p.subject, p.issuer);
    }
}

PeerCert is injected on every TLS connection; for plain HTTPS (no A2A_SERVER_TLS_CLIENT_CA_PATH) its inner Option is None because the client presented no certificate. Because the artifacts server reuses the same build_server_config machinery, enabling TLS/mTLS on the agent covers the artifacts endpoint too.

See examples/tls/ for an end-to-end demo with a make-certs.sh script that mints a self-signed CA plus server and client certificates, and exercises both modes via the tls and mtls Compose profiles.

Artifacts

The ADK ships a first-class artifacts subsystem so agents can produce downloadable file artifacts (reports, images, structured-data dumps) and hand A2A clients a URI rather than inline base64 bytes embedded in JSON-RPC responses. It mirrors the Go ADK artifacts surface, so artifact-producing agents behave identically on the wire across ADKs.

The subsystem has four moving parts, each behind a trait so production deployments can swap in their own backends:

Layer Trait / type Default
Configuration ArtifactsConfig (in config.rs) disabled (ARTIFACTS_ENABLED=false)
Storage ArtifactStorage (store, retrieve, exists, delete, cleanup_*) FilesystemArtifactStorage
Service ArtifactService (create_*_artifact, add_artifact_to_task, retention) DefaultArtifactService
HTTP surface ArtifactsServer (GET /health, GET /artifacts/:context_id/:artifact_id/:filename) 0.0.0.0:8081 listener with byte-range support

When ARTIFACTS_ENABLED=true, A2AServer::serve(...) starts the artifacts HTTP server on its own socket alongside the main A2A JSON-RPC server and runs a background retention loop that prunes expired and over-cap blobs. The artifacts server reuses the same TLS machinery as the A2A endpoint, so it can sit behind TLS/mTLS too.

flowchart LR
    client["A2A client"]
    subgraph proc["Agent process"]
      a2a["A2A JSON-RPC server :8080"]
      artsrv["ArtifactsServer :8081"]
      svc["ArtifactService"]
      store[("ArtifactStorage")]
    end
    client -->|"SendStreamingMessage"| a2a
    a2a -->|"emit_file_artifact"| svc
    svc -->|"store bytes"| store
    a2a -.->|"TaskArtifactUpdateEvent (FilePart.fileWithUri)"| client
    client -->|"GET /artifacts/:context/:id/:file"| artsrv
    artsrv -->|"retrieve"| store
Loading

The streaming handler writes the artifact through the service into storage, attaches an Artifact (carrying a FilePart with fileWithUri set) to the stored task, and emits a TaskArtifactUpdateEvent over the SSE stream. The client treats the URI as opaque and downloads the bytes directly - either from the ADK's artifacts server or, with the MinIO backend, straight from the object store.

All artifact types live behind the public exports from the crate root (inference_gateway_adk): ArtifactStorage, FilesystemArtifactStorage, MinioArtifactStorage, ArtifactService, DefaultArtifactService, ArtifactsServer, plus the config types ArtifactsConfig, ArtifactsServerConfig, ArtifactsStorageConfig, ArtifactRetentionConfig, and ArtifactsStorageProvider.

Storage backends

ArtifactStorage is the pluggable backend trait. Its surface is intentionally small - store, retrieve, exists, delete (each taking the contextId as their first argument), plus cleanup_expired / cleanup_oldest for the retention loop and a url(...) helper that builds the public URI baked into file artifacts. Two backends ship in the box.

Filesystem (default)

FilesystemArtifactStorage is the zero-config default. It lays files out under <base_path>/<context_id>/<artifact_id>/<filename> and sanitizes every path segment to prevent traversal. Deleting an artifact also removes the now-empty artifact and context directories, stopping at the configured root. The artifacts HTTP server streams blobs back out of this directory with content-type inference, a Content-Disposition header, and HTTP byte-range support.

MinIO (behind the minio Cargo feature)

MinioArtifactStorage implements the same ArtifactStorage trait against an S3-compatible MinIO server. It is gated behind the minio Cargo feature, which pulls in the minio crate:

[dependencies]
inference-gateway-adk = { version = "0.4", features = ["minio"] }

Selecting ARTIFACTS_STORAGE_PROVIDER=minio without compiling the minio feature is not an error - the builder logs a warn! and falls back to the filesystem store, so the ARTIFACTS_STORAGE_* env surface stays valid either way.

On startup MinioArtifactStorage::from_config checks for the target bucket and creates it if missing. With ARTIFACTS_STORAGE_BASE_URL pointed at the MinIO endpoint, url(...) emits a path-style http://<endpoint>/<bucket>/<context_id>/<artifact_id>/<filename> so clients download directly from MinIO, bypassing the ADK's artifacts HTTP server entirely - the way you would offload bulk transfer in production. See Production notes for the private-bucket trade-off.

The artifact service

ArtifactService is the helper layer between your handler and the storage backend. The bundled DefaultArtifactService covers the full artifact lifecycle:

  • create_text_artifact, create_file_artifact, create_uri_artifact, and create_data_artifact - mint the different Part kinds. File and data artifacts are persisted to storage and returned as a FilePart with fileWithUri set (data artifacts also serialize as A2A DataParts).
  • add_artifact_to_task - attach a created Artifact to a stored task so it is included in GetTask responses.
  • retrieve / exists / cleanup - read-side and retention helpers used by the artifacts server and the background cleanup loop. create_file_artifact, retrieve, and exists take the contextId the artifact belongs to.

Most handlers never call the service directly; they go through the StreamEmitter helpers below, which wrap it.

The artifacts HTTP server

ArtifactsServer is a standalone Axum app on its own socket (default 0.0.0.0:8081), kept separate from the A2A JSON-RPC surface so bulk-download traffic does not entangle the protocol endpoint. It exposes two routes:

Route Purpose
GET /health Liveness probe for load balancers and orchestrators.
GET /artifacts/:context_id/:artifact_id/:filename Streams a stored blob with content-type inference, Content-Disposition, and byte ranges.

Because it reuses the A2A endpoint's build_server_config TLS machinery, enabling TLS/mTLS on the agent also covers the artifacts server.

Enabling artifacts on the server

A2AServerBuilder::with_config(config) auto-wires the artifacts subsystem from config.artifacts_config whenever enable is true - no extra builder calls are required. A2AServer::serve(addr) then spawns the artifacts server and the retention loop next to the A2A server:

use inference_gateway_adk::{
    A2AServerBuilder, ArtifactsConfig, ArtifactsServerConfig, ArtifactsStorageConfig, Config,
};

let config = Config {
    artifacts_config: ArtifactsConfig {
        enable: true,
        server: ArtifactsServerConfig {
            port: 8088,
            ..Default::default()
        },
        storage: ArtifactsStorageConfig {
            base_path: "./artifacts-data".to_string(),
            base_url: "http://localhost:8088".to_string(),
            ..Default::default()
        },
        retention: Default::default(),
    },
    ..Config::default()
};

let server = A2AServerBuilder::new()
    .with_config(config)
    .with_agent_card_from_file(".well-known/agent-card.json", None)
    .with_default_task_handlers()
    .build()
    .await?;

// Serves the A2A JSON-RPC API on this address AND the artifacts
// server on `config.artifacts_config.server` (here :8088).
server.serve("0.0.0.0:8087".parse()?).await?;

To supply a custom backend - your own ArtifactStorage or a fully custom ArtifactService - pass it via A2AServerBuilder::with_artifact_service(...), the same way with_storage(...) overrides the task store.

Loading config from the environment

The artifacts subsystem uses its own ARTIFACTS_ env prefix (matching the Go ADK and the bundled examples) rather than the A2A_ prefix the rest of Config uses. Load it independently and assign the result onto Config::artifacts_config:

use inference_gateway_adk::{ArtifactsConfig, Config};

let artifacts_config = envy::prefixed("ARTIFACTS_")
    .from_env::<ArtifactsConfig>()
    .unwrap_or_default();

let config = Config {
    artifacts_config,
    ..Config::default()
};

Config::artifacts_config is #[serde(skip)], so an envy::prefixed("A2A_").from_env::<Config>() load never touches it. Load the two prefixes separately.

Emitting artifacts from a streaming handler

Streaming task handlers mint artifacts mid-stream through the StreamEmitter:

  • emit_file_artifact(task_id, context_id, filename, bytes, content_type, last_chunk) - persists raw bytes and emits a file artifact whose FilePart.fileWithUri points at the artifacts server (URL prefix taken from ARTIFACTS_STORAGE_BASE_URL).
  • emit_data_artifact(...) - emits a structured-data artifact as an A2A DataPart.

Both routes write to storage, attach the Artifact to the stored task, and publish a TaskArtifactUpdateEvent to the SSE stream. The following handler emits a one-shot text report as report.txt:

use inference_gateway_adk::a2a_types::{Message as A2AMessage, Task, TaskState};
use inference_gateway_adk::{StreamEmitter, StreamableTaskHandler};

#[derive(Debug)]
struct ReportHandler;

#[async_trait::async_trait]
impl StreamableTaskHandler for ReportHandler {
    async fn handle_streaming_task(
        &self,
        task: Task,
        _message: Option<A2AMessage>,
        emitter: StreamEmitter,
    ) -> anyhow::Result<()> {
        // 1. Announce that work has started.
        emitter
            .emit_status(&task.id, &task.context_id, TaskState::TaskStateWorking, None, false)
            .await?;

        // 2. Produce and persist a file artifact. The trailing `true`
        //    marks this as the final (and only) chunk of the artifact.
        let report = format!(
            "# Generated Report\n\nTask id: {}\nContext id: {}\n",
            task.id, task.context_id,
        );
        emitter
            .emit_file_artifact(
                &task.id,
                &task.context_id,
                "report.txt",
                report.into_bytes(),
                Some("text/plain"),
                true,
            )
            .await?;

        // 3. Close out the task.
        emitter
            .emit_status(&task.id, &task.context_id, TaskState::TaskStateCompleted, None, true)
            .await
    }
}

Wire the handler into the builder with .with_streaming_task_handler(ReportHandler). The client opens SendStreamingMessage, reads the FilePart.fileWithUri off the TaskArtifactUpdateEvent, and fetches it with a plain HTTP GET - it never needs to know about artifact IDs or storage layout.

Artifacts configuration reference

The artifacts subsystem is configured entirely through the ARTIFACTS_* environment-variable surface, loaded via envy::prefixed("ARTIFACTS_").from_env::<ArtifactsConfig>(). Every value falls back to the default below when unset.

Variable Default Description
ARTIFACTS_ENABLED false Master switch. When true, A2AServer::serve(...) spawns the artifacts server and retention loop.
ARTIFACTS_SERVER_HOST 0.0.0.0 Bind address of the artifacts HTTP server.
ARTIFACTS_SERVER_PORT 8081 Port of the artifacts HTTP server.
ARTIFACTS_SERVER_READ_TIMEOUT 30s Parsed but not yet applied - ArtifactsServer sets no read timeout.
ARTIFACTS_SERVER_WRITE_TIMEOUT 30s Parsed but not yet applied - ArtifactsServer sets no write timeout.
ARTIFACTS_STORAGE_PROVIDER filesystem filesystem or minio. The minio provider requires the minio Cargo feature; without it, requests fall back to filesystem storage with a warn!.
ARTIFACTS_STORAGE_BASE_PATH ./artifacts On-disk root for the filesystem provider.
ARTIFACTS_STORAGE_BASE_URL http://localhost:8081 Public URL prefix baked into file artifact URIs. Point it at wherever the artifacts server (or MinIO endpoint) is externally reachable.
ARTIFACTS_STORAGE_ENDPOINT unset MinIO endpoint URL. A https:// scheme implies SSL.
ARTIFACTS_STORAGE_ACCESS_KEY unset MinIO access key.
ARTIFACTS_STORAGE_SECRET_KEY unset MinIO secret key.
ARTIFACTS_STORAGE_BUCKET_NAME unset MinIO bucket name. Created on startup if missing.
ARTIFACTS_STORAGE_REGION unset MinIO region.
ARTIFACTS_STORAGE_USE_SSL false Whether to use TLS when talking to the MinIO endpoint.
ARTIFACTS_RETENTION_MAX_ARTIFACTS 5 Cap on artifacts kept per contextId (matching the Go ADK); the oldest beyond the cap are pruned. 0 means unlimited.
ARTIFACTS_RETENTION_MAX_AGE 168h Maximum age before an artifact is pruned.
ARTIFACTS_RETENTION_CLEANUP_INTERVAL 24h Frequency of the retention loop.

ARTIFACTS_RETENTION_MAX_ARTIFACTS is store-wide here: each cleanup run keeps the newest N artifacts across the whole store and prunes the oldest beyond that, regardless of contextId. The Go ADK shares the variable name and the default of 5 but applies it per context, so the same value retains more artifacts there.

Duration values (*_TIMEOUT, *_MAX_AGE, *_CLEANUP_INTERVAL) accept Go-style suffixes - 30s, 15m, 2h, 7d - or a bare integer interpreted as seconds. An unknown suffix such as 5w is rejected at load time.

Upgrading to contextId-scoped artifacts

Artifact paths used to be /artifacts/{artifactId}/{filename} (filesystem and MinIO keys <artifactId>/<filename>), and ARTIFACTS_RETENTION_MAX_ARTIFACTS capped the store as a whole. Both are now scoped by contextId. This is a breaking change with no dual-read fallback: artifacts written under the old layout are not readable after upgrading.

  • Drain or wipe the artifact store before or during the upgrade if artifacts are enabled - old blobs are unreachable and the retention loop skips keys that do not match the three-segment shape.
  • A custom ArtifactStorage or ArtifactService needs its store / retrieve / exists / delete / url and create_file_artifact signatures updated to take the contextId. Streaming handlers using emit_file_artifact already pass one, so they need no change.
  • Clients that parse artifact URLs rather than treating them as opaque need updating for the extra path segment; the bundled Rust client treats them as opaque.

Example: filesystem backend

A runnable end-to-end demo lives at examples/artifacts-filesystem. The streaming handler emits a small text report and the client downloads it directly from the artifacts server. Two HTTP servers run in the same process: A2A JSON-RPC on :8087 and the artifacts server on :8088.

# examples/artifacts-filesystem/docker-compose.yaml (env excerpt)
services:
  server:
    ports:
      - '8087:8087' # A2A JSON-RPC
      - '8088:8088' # Artifacts HTTP (exposed for host-side curl/debug)
    environment:
      ARTIFACTS_ENABLED: 'true'
      ARTIFACTS_SERVER_HOST: '0.0.0.0'
      ARTIFACTS_SERVER_PORT: '8088'
      ARTIFACTS_STORAGE_PROVIDER: filesystem
      ARTIFACTS_STORAGE_BASE_PATH: /data/artifacts
      # Baked into FilePart.fileWithUri. Uses the Docker service name so the
      # client container can resolve it; host-side curl still works via 8088.
      ARTIFACTS_STORAGE_BASE_URL: http://server:8088
      ARTIFACTS_RETENTION_MAX_ARTIFACTS: '5'
      ARTIFACTS_RETENTION_MAX_AGE: '168h'
      ARTIFACTS_RETENTION_CLEANUP_INTERVAL: '24h'
    volumes:
      - ./server/artifacts-data:/data/artifacts

Run it:

cd examples/artifacts-filesystem
docker compose up --build

No .env and no provider keys are required. The filesystem provider lays files out under <base_path>/<context_id>/<artifact_id>/<filename>; the compose stack bind-mounts the container store to ./server/artifacts-data/ so produced files are inspectable after the run. The full artifact URI is printed by the client log line beginning received file artifact.

Example: MinIO backend

examples/artifacts-minio demonstrates the same flow against a MinIO container, with FilePart.fileWithUri pointing directly at MinIO. The compose stack adds a minio/minio service plus a one-shot minio/mc init container that creates the artifacts bucket with an anonymous-download policy, and builds the server with CARGO_FEATURES=minio.

# examples/artifacts-minio/docker-compose.yaml (excerpt)
services:
  minio:
    image: minio/minio:latest
    command: server /data --console-address ':9001'
    ports:
      - '9000:9000' # MinIO API
      - '9001:9001' # MinIO console
    environment:
      MINIO_ROOT_USER: minioadmin
      MINIO_ROOT_PASSWORD: minioadmin

  createbucket:
    image: minio/mc:latest
    depends_on:
      minio:
        condition: service_healthy
    entrypoint:
      - /bin/sh
      - -c
      - |
        /usr/bin/mc alias set local http://minio:9000 minioadmin minioadmin &&
        /usr/bin/mc mb --ignore-existing local/artifacts &&
        /usr/bin/mc anonymous set download local/artifacts &&
        exit 0

  server:
    build:
      args:
        CARGO_FEATURES: minio # compiles the `minio` crate in
    environment:
      ARTIFACTS_STORAGE_ENDPOINT: http://minio:9000
      ARTIFACTS_STORAGE_ACCESS_KEY: minioadmin
      ARTIFACTS_STORAGE_SECRET_KEY: minioadmin
      ARTIFACTS_STORAGE_BUCKET_NAME: artifacts
      # Points at MinIO via the Docker service name; downloads bypass the
      # artifacts HTTP server entirely. The bucket has an anonymous-download
      # policy (set by `createbucket` above).
      ARTIFACTS_STORAGE_BASE_URL: http://minio:9000

The MinIO server's handler picks ArtifactsStorageProvider::Minio and runs the same emit_file_artifact code as the filesystem example - only the storage backend and ARTIFACTS_STORAGE_* env differ:

use inference_gateway_adk::{ArtifactsConfig, ArtifactsStorageProvider, Config};

let mut artifacts_config = envy::prefixed("ARTIFACTS_")
    .from_env::<ArtifactsConfig>()?;
artifacts_config.enable = true;
artifacts_config.storage.provider = ArtifactsStorageProvider::Minio;

let config = Config { artifacts_config, ..Config::default() };

Run it:

cd examples/artifacts-minio
docker compose up --build

The MinIO-specific env vars default to a local-friendly setup:

Variable Example default Description
ARTIFACTS_STORAGE_ENDPOINT http://localhost:9000 MinIO endpoint URL. https:// implies SSL.
ARTIFACTS_STORAGE_ACCESS_KEY minioadmin Static access key.
ARTIFACTS_STORAGE_SECRET_KEY minioadmin Static secret key.
ARTIFACTS_STORAGE_BUCKET_NAME artifacts Target bucket. Created on startup if missing.
ARTIFACTS_STORAGE_BASE_URL http://localhost:9000 Public URL prefix baked into FilePart.fileWithUri.

The retention and artifacts-server bind variables from the artifacts configuration reference apply here too.

Production notes

  • Anonymous-read buckets are a deployment choice, not a default. If your MinIO bucket is private, point ARTIFACTS_STORAGE_BASE_URL at the ADK's artifacts HTTP server instead of the object store; the server then proxies the fetch through ArtifactStorage::retrieve(). The trade-off is bulk traffic flowing through your agent process rather than straight from MinIO.
  • Pre-signed URLs are not yet wired in. A future enhancement could have MinioArtifactStorage::url(...) mint a time-limited pre-signed GET so a private bucket needs no proxying.
  • Retention runs cleanup_expired / cleanup_oldest over a listing of stored objects - fine for thousands of artifacts. Past that, prefer MinIO bucket lifecycle policies for the bulk of expiry.

Telemetry

The ADK bridges its tracing instrumentation to OpenTelemetry, exporting spans to an OTLP collector over HTTP/protobuf. It is a traces-only signal - there is no metrics export and no gRPC/tonic transport - and it lands behind the optional telemetry Cargo feature so the default build stays lean. It mirrors the Go ADK: A2A_TELEMETRY_ENABLED is the sole switch, and A2A_OTEL_TRACES_EXPORTER=none opts the trace signal out while telemetry stays enabled.

Enable the feature at build time:

[dependencies]
inference-gateway-adk = { version = "0.4", features = ["telemetry"] }
cargo build --features telemetry

telemetry::init installs the tracing subscriber and, when telemetry is enabled, the OTLP span exporter. It returns a TelemetryGuard that owns the exporter, so bind it to a variable that lives for the whole process (let _guard = ...) - dropping the guard early, or writing let _ = telemetry::init(...), tears the exporter down immediately and loses buffered spans on shutdown.

use inference_gateway_adk::{A2AServerBuilder, Config, telemetry};

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let config: Config = envy::prefixed("A2A_").from_env()?;

    // Always installs the fmt layer; adds the OTLP exporter when
    // A2A_TELEMETRY_ENABLED=true AND the `telemetry` feature is compiled in.
    // Hold the guard for the process lifetime so batched spans flush on exit.
    let _guard = telemetry::init(
        &config.telemetry_config,
        env!("CARGO_PKG_NAME"),
        env!("CARGO_PKG_VERSION"),
    )?;

    let server = A2AServerBuilder::new()
        .with_config(config)
        .with_agent_card_from_file(".well-known/agent-card.json", None)
        .with_default_task_handlers()
        .build()
        .await?;

    server.serve("0.0.0.0:8080".parse()?).await?;
    Ok(())
}

init always installs the fmt layer, so logging works with or without the feature. When A2A_TELEMETRY_ENABLED=true but the telemetry feature was not compiled in, init logs a warn! and skips export rather than failing - rebuild with --features telemetry to turn spans on.

Emitted spans

With export active, the server emits three spans:

  • a2a.request - a server span around each POST /a2a request. It extracts the caller's W3C tracecontext from the request headers (so a client-initiated trace continues across the hop), records the JSON-RPC method, route, and status, and flags 5xx responses as errors.
  • task.process - a fresh root span wrapping each background task's execution.
  • tool.<name> - one span per tool call dispatched by the default task handlers, parented on task.process. No instrumentation is required in your tool implementations; the handler wraps every dispatch automatically. Mirrors the Go ADK toolbox spans.

Each tool.<name> span carries:

Attribute Source
gen_ai.tool.name The tool name, matching the <name> in the span name.
gen_ai.tool.call.id The tool-call id assigned by the model.
session.id Copied from the W3C baggage member of the same name, when the incoming request has it.

A tool that returns an error marks its span as errored, so failed calls stand out in the trace without extra wiring. All of this lands behind the telemetry Cargo feature - build without it and the spans compile down to no-ops.

Configuration

Telemetry loads from the same A2A_ Config as the rest of the server (config.telemetry_config). Endpoint precedence: A2A_TELEMETRY_ENDPOINT wins when set; otherwise the exporter falls back to the SDK's standard OTEL_EXPORTER_OTLP_ENDPOINT resolution, which defaults to http://localhost:4318. Standard OTEL_* variables are honored by the OpenTelemetry SDK as usual.

Variable Default Purpose
A2A_TELEMETRY_ENABLED false Sole telemetry switch. When true, traces export over OTLP.
A2A_OTEL_TRACES_EXPORTER otlp Trace exporter: otlp or none. none opts traces out while telemetry stays on.
A2A_TELEMETRY_ENDPOINT (unset) OTLP collector endpoint. Takes precedence over OTEL_EXPORTER_OTLP_ENDPOINT.
OTEL_EXPORTER_OTLP_ENDPOINT http://localhost:4318 Standard OTLP endpoint, used when A2A_TELEMETRY_ENDPOINT is unset.

Examples

The examples/ directory ships fourteen runnable scenarios, each its own Cargo package with a docker-compose.yaml. The catalogue is grouped by whether the scenario needs an LLM provider.

Without AI - no Inference Gateway, no provider keys:

Example What it shows
minimal Bare A2A server + client with the built-in echo reply - no agent wired up.
static-agent-card Load agent metadata from JSON and patch it at runtime with AgentCardOverrides.
streaming A custom streaming handler emits a sentence word-by-word over SSE.
input-required A handler parks a task in TaskStateInputRequired when the user message is incomplete.

With AI - Inference Gateway container plus a provider key:

Example What it shows
default-handlers An LLM agent with with_default_task_handlers() - no custom handler code.
ai-powered An LLM agent with custom function tools (weather, math, search).
ai-powered-streaming The same agent streamed over SendStreamingMessage.
usage-metadata Default handlers attach token usage and execution_stats to task.metadata on terminal states.

Storage and protocol coverage:

Example What it shows
queue-storage Queue-driven SendMessage with in-memory or Redis storage, selectable via Compose profiles.
a2a-methods One client binary per JSON-RPC method in the A2A spec, sharing a single offline server.
auth Bearer-token auth on POST /a2a with public /health and /.well-known/agent-card.json; static-token or Keycloak OIDC.
tls TLS termination via axum-server + rustls, with optional mTLS exposing the client-cert subject as the principal.
artifacts-filesystem A streaming handler emits a FilePart served by the standalone artifacts HTTP server, backed by an on-disk store.
artifacts-minio The same flow backed by a MinIO bucket instead of the local filesystem.

Environment variable reference

The library never reads the environment itself. You pick a loader - typically envy::prefixed("A2A_").from_env::<Config>() - and hand the resulting Config to A2AServerBuilder::with_config. Every variable below is optional and falls back to the default shown; switch prefixes by changing the loader, not the code.

*_SECS variables are plain integer seconds. The Go-style duration grammar (30s, 15m, 2h, 7d) applies to the MCP_ durations below and to the ARTIFACTS_ durations in the artifacts configuration reference. The artifacts subsystem loads under its own ARTIFACTS_ prefix and is not part of the A2A_ surface. The agent card's identity fields and advertised capabilities are not configurable through the environment - they come from the card JSON (see Agent card and metadata).

Log verbosity is controlled by RUST_LOG, which telemetry::init reads through EnvFilter::try_from_default_env() - for example RUST_LOG=inference_gateway_adk=debug.

Server and core - the listener and top-level toggles.

Variable Default Purpose
A2A_SERVER_HOST 0.0.0.0 Inert - ServerConfig::host is never read. The bind address comes solely from the SocketAddr passed to A2AServer::serve.
A2A_SERVER_PORT 8080 Feeds the default advertised URL http(s)://localhost:<port>/a2a; it does not bind the listener either.
A2A_AGENT_URL (derived) Public URL advertised in the card - see below.
A2A_STREAMING_STATUS_UPDATE_INTERVAL_SECS 1 Seconds between TaskStatusUpdateEvents on a streamed task.

A2A_AGENT_URL has no fixed default - the server resolves the advertised URL at startup, see Agent card and metadata for the full precedence chain.

Agent (LLM client) - these mirror the AgentBuilder setters; an explicit setter overrides the env value.

Variable Default Purpose
A2A_AGENT_CLIENT_PROVIDER (empty) LLM provider id (e.g. openai, nvidia, ollama, groq).
A2A_AGENT_CLIENT_MODEL (empty) Model name.
A2A_AGENT_CLIENT_BASE_URL (unset) Override the gateway/provider base URL.
A2A_AGENT_CLIENT_API_KEY (unset) Provider API key, sent to the gateway as a bearer token.
A2A_AGENT_CLIENT_TIMEOUT_SECS 30 Bounds each LLM request and each wait for a streamed event. 0 disables the bound.
A2A_AGENT_CLIENT_MAX_RETRIES 3 Retry budget for failed LLM calls.
A2A_AGENT_CLIENT_MAX_CHAT_COMPLETION_ITERATIONS 10 Cap on the agent tool loop - chat-completion round trips per turn.
A2A_AGENT_CLIENT_MAX_TOKENS 4096 Max tokens per completion. Non-streaming completions only - the gateway SDK omits max_tokens from streaming requests.
A2A_AGENT_CLIENT_TEMPERATURE (unset) Sampling temperature, 0.0 - 2.0. Sent with streaming and non-streaming completions; unset keeps the gateway default.
A2A_AGENT_CLIENT_SYSTEM_PROMPT (unset) System prompt prepended to conversations.
A2A_AGENT_CLIENT_ENABLE_USAGE_METADATA true Serve the usage extension on terminal tasks.

Capabilities are not environment-configurable. The served card and the builder's streaming-handler validation read the capabilities object of your agent card JSON, so set streaming, pushNotifications, and extendedAgentCard there. pushNotifications gates the push-notification-config methods and extendedAgentCard gates GetExtendedAgentCard; a method whose flag is missing answers a JSON-RPC error rather than running. The A2A v1.0.1 card has no stateTransitionHistory flag - a card still carrying it fails to deserialize into the ADK's AgentCapabilities.

MCP client (MCP_ prefix) - connect the agent to MCP servers; disabled by default. Loaded under its own MCP_ prefix (like ARTIFACTS_), separate from the A2A_ Config.

Variable Default Purpose
MCP_ENABLED false Enable the MCP client.
MCP_SERVERS (unset) Comma-separated MCP server base URLs, e.g. http://mcp:8080.
MCP_ENDPOINT /mcp Path appended to each server URL for the Streamable HTTP endpoint.
MCP_REFRESH_INTERVAL 5m How often to refresh the tool catalog from each server.
MCP_DIAL_TIMEOUT 30s Timeout for initializing / listing tools.
MCP_CALL_TIMEOUT 30s Timeout for a single tool invocation.
MCP_MAX_RETRIES 0 Max initial connection attempts per server (0 = retry forever).
MCP_RETRY_INTERVAL 2s Initial backoff between connection/refresh retries (doubles).
MCP_RETRY_MAX_INTERVAL 30s Maximum backoff between retries.

Queue and storage - selects the Storage backend the server factory wires up.

Variable Default Purpose
A2A_QUEUE_PROVIDER memory memory or redis.
A2A_QUEUE_URL (unset) Backend URL, e.g. redis://host:6379. Required when provider is redis.
A2A_QUEUE_NAMESPACE a2a Key prefix for backend keys.
A2A_QUEUE_WORKERS 1 Number of DefaultTaskManager workers draining the queue.
A2A_QUEUE_MAX_SIZE 1000 Advisory in-flight cap for the in-memory backend.
A2A_QUEUE_TIMEOUT_SECS 30 Per-operation backend timeout.

Authentication - see Authentication for the full flow.

Variable Default Purpose
A2A_AUTH_ENABLED false Enable OIDC bearer-token auth on POST /a2a.
A2A_AUTH_ISSUER_URL (empty) OIDC issuer used for discovery + JWKS. Required when auth is enabled.
A2A_AUTH_CLIENT_ID (empty) Expected aud claim on incoming tokens. Required when auth is enabled.
A2A_AUTH_CLIENT_SECRET (empty) Required when auth is enabled; build() fails without it, though verification never reads it.

TLS and mTLS - see TLS and mTLS.

Variable Default Purpose
A2A_SERVER_TLS_ENABLED false Terminate TLS on the A2A listener.
A2A_SERVER_TLS_CERT_PATH (empty) PEM file with the server certificate chain.
A2A_SERVER_TLS_KEY_PATH (empty) PEM file with the server private key.
A2A_SERVER_TLS_CLIENT_CA_PATH (omit) Trusted client-CA bundle. Any value, an empty string included, flips the server into mTLS - omit the variable for plain TLS.

Telemetry - OpenTelemetry OTLP trace export; see Telemetry for telemetry::init usage and the emitted spans. Traces-only over HTTP/protobuf, behind the optional telemetry Cargo feature. A2A_TELEMETRY_ENABLED is the sole switch, matching the Go ADK; A2A_OTEL_TRACES_EXPORTER=none opts the trace signal out while telemetry stays enabled.

Variable Default Purpose
A2A_TELEMETRY_ENABLED false Sole telemetry switch. When true, traces export over OTLP.
A2A_OTEL_TRACES_EXPORTER otlp Trace exporter: otlp or none. none opts traces out while telemetry stays on.
A2A_TELEMETRY_ENDPOINT (unset) OTLP collector endpoint. Takes precedence over the standard OTEL_EXPORTER_OTLP_ENDPOINT.
OTEL_EXPORTER_OTLP_ENDPOINT http://localhost:4318 Standard OTLP endpoint, used when A2A_TELEMETRY_ENDPOINT is unset. Part of the OTEL_* passthrough.

Related

  • Agent Definition Language (ADL) - define an agent declaratively in YAML instead of wiring the builders by hand.
  • ADL CLI - scaffold a Rust A2A agent project (built on this ADK) from an ADL file, then fill in the generated tool stubs.
  • A2A Integration - the Agent-to-Agent protocol these agents speak.
  • TypeScript ADK - the sibling ADK for Node/TypeScript agents.
  • Go ADK - the reference ADK whose semantics the Rust port mirrors.
  • Rust ADK on GitHub - source, issues, and the full example catalogue.
  • Inference Gateway - the gateway the agents call for LLM completions.