Is your feature request related to a problem? Please describe.
When Agent completes its loop, its terminal output is plain freeform text in ChatMessage.text. In pipelines and production applications where downstream components expect structured data (e.g., Pydantic models or JSON schemas), users currently have to attach separate validator components or write custom parsing logic outside the agent.
Furthermore, LLMs frequently produce slight schema violations, miss required keys, or wrap responses in markdown fences (json ... ). When parsing fails downstream, the entire pipeline fails, and the agent's multi-turn conversational context is already lost—making prompt-based self-correction impossible without writing an external retry loop around the entire agent.
Describe the solution you'd like
Add response_schema and max_schema_retries parameters directly to Agent:
response_schema: type[BaseModel] | dict[str, Any] | None = None:
Accepts either a Pydantic BaseModel subclass or a standard JSON Schema dictionary.
max_schema_retries: int = 3:
Maximum number of self-correction attempts when the model's text response fails schema validation.
- Execution Behavior:
- Gated strictly on the model's genuine text termination (
model_exit_reason == "text").
- Strips optional markdown code fences prior to JSON decoding.
- On success: Emits the parsed Pydantic instance or dict under a new
structured_output key (and corresponding component output socket).
- On validation failure: If retries and step budget remain, appends a targeted correction prompt (
ChatMessage.from_user) containing the error and schema, and continues the agent loop so the model can correct its output.
- On retry exhaustion or step budget cutoff: Gracefully logs a warning and returns
structured_output = None without raising an unhandled exception.
- When omitted (
response_schema=None): Behavior is 100% backward compatible; no "structured_output" key is registered or returned.
Example usage:
from pydantic import BaseModel
from haystack.components.agents import Agent
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
class EntityExtractor(BaseModel):
entities: list[str]
confidence: float
agent = Agent(
chat_generator=OpenAIChatGenerator(),
response_schema=EntityExtractor,
max_schema_retries=2,
)
result = agent.run(messages=[ChatMessage.from_user("Extract entities from: 'Apple released M3 chip in California.'")])
data: EntityExtractor | None = result["structured_output"]
Describe alternatives you've considered
- Provider-specific structured output APIs (e.g. OpenAI
response_format): Modifying individual generators couples the feature to specific providers and prevents use with other generators (Anthropic, Bedrock, Ollama, HuggingFace, etc.). Post-hoc validation in Agent is completely provider-agnostic.
- Downstream OutputValidator component: Moving validation to a pipeline component outside
Agent loses the agent's active execution context and cannot prompt the model to fix its response without re-invoking the entire agent from scratch.
Additional context
- Uses existing dependencies already in Haystack (
pydantic and jsonschema).
- Validates JSON schema dicts at
__init__ via Draft202012Validator.check_schema for fail-fast error reporting.
👋 Hello there! This issue will be handled internally and isn't open for external contributions. If you'd like to contribute, please take a look at issues labeled contributions welcome or good first issue. We'd really appreciate it!
Is your feature request related to a problem? Please describe.
When
Agentcompletes its loop, its terminal output is plain freeform text inChatMessage.text. In pipelines and production applications where downstream components expect structured data (e.g., Pydantic models or JSON schemas), users currently have to attach separate validator components or write custom parsing logic outside the agent.Furthermore, LLMs frequently produce slight schema violations, miss required keys, or wrap responses in markdown fences (
json ...). When parsing fails downstream, the entire pipeline fails, and the agent's multi-turn conversational context is already lost—making prompt-based self-correction impossible without writing an external retry loop around the entire agent.Describe the solution you'd like
Add
response_schemaandmax_schema_retriesparameters directly toAgent:response_schema: type[BaseModel] | dict[str, Any] | None = None:Accepts either a Pydantic
BaseModelsubclass or a standard JSON Schema dictionary.max_schema_retries: int = 3:Maximum number of self-correction attempts when the model's text response fails schema validation.
model_exit_reason == "text").structured_outputkey (and corresponding component output socket).ChatMessage.from_user) containing the error and schema, and continues the agent loop so the model can correct its output.structured_output = Nonewithout raising an unhandled exception.response_schema=None): Behavior is 100% backward compatible; no"structured_output"key is registered or returned.Example usage:
Describe alternatives you've considered
response_format): Modifying individual generators couples the feature to specific providers and prevents use with other generators (Anthropic, Bedrock, Ollama, HuggingFace, etc.). Post-hoc validation inAgentis completely provider-agnostic.Agentloses the agent's active execution context and cannot prompt the model to fix its response without re-invoking the entire agent from scratch.Additional context
pydanticandjsonschema).__init__viaDraft202012Validator.check_schemafor fail-fast error reporting.👋 Hello there! This issue will be handled internally and isn't open for external contributions. If you'd like to contribute, please take a look at issues labeled contributions welcome or good first issue. We'd really appreciate it!