# Table of Contents - [Welcome to Pydantic | Pydantic Docs](#welcome-to-pydantic-pydantic-docs) - [Pydantic Docs - Validation, AI Agents, Logfire Observability](#pydantic-docs-validation-ai-agents-logfire-observability) - [bedrock_mantle | Pydantic Docs](#bedrock-mantle-pydantic-docs) - [bedrock | Pydantic Docs](#bedrock-pydantic-docs) - [anthropic | Pydantic Docs](#anthropic-pydantic-docs) - [cohere | Pydantic Docs](#cohere-pydantic-docs) - [crusoe | Pydantic Docs](#crusoe-pydantic-docs) - [cerebras | Pydantic Docs](#cerebras-pydantic-docs) - [base | Pydantic Docs](#base-pydantic-docs) - [huggingface | Pydantic Docs](#huggingface-pydantic-docs) - [instrumented | Pydantic Docs](#instrumented-pydantic-docs) - [fallback | Pydantic Docs](#fallback-pydantic-docs) - [function | Pydantic Docs](#function-pydantic-docs) - [google | Pydantic Docs](#google-pydantic-docs) - [groq | Pydantic Docs](#groq-pydantic-docs) - [ollama | Pydantic Docs](#ollama-pydantic-docs) - [wrapper | Pydantic Docs](#wrapper-pydantic-docs) - [Medical Agent Delegation | Pydantic Docs](#medical-agent-delegation-pydantic-docs) - [format_prompt | Pydantic Docs](#format-prompt-pydantic-docs) - [mcp-sampling | Pydantic Docs](#mcp-sampling-pydantic-docs) - [Process Event Stream | Pydantic Docs](#process-event-stream-pydantic-docs) - [Web Fetch | Pydantic Docs](#web-fetch-pydantic-docs) - [Agent Handoff | Pydantic Docs](#agent-handoff-pydantic-docs) - [Instrumentation | Pydantic Docs](#instrumentation-pydantic-docs) - [Installation | Pydantic Docs](#installation-pydantic-docs) - [Resolve Model ID | Pydantic Docs](#resolve-model-id-pydantic-docs) - [Tool Search | Pydantic Docs](#tool-search-pydantic-docs) - [Compaction | Pydantic Docs](#compaction-pydantic-docs) - [Case Lifecycle Hooks | Pydantic Docs](#case-lifecycle-hooks-pydantic-docs) - [Overview | Pydantic Docs](#overview-pydantic-docs) - [Decisions | Pydantic Docs](#decisions-pydantic-docs) - [Pydantic Model | Pydantic Docs](#pydantic-model-pydantic-docs) - [Coding Agent Skills | Pydantic Docs](#coding-agent-skills-pydantic-docs) - [Troubleshooting | Pydantic Docs](#troubleshooting-pydantic-docs) - [snowflake | Pydantic Docs](#snowflake-pydantic-docs) - [Handle Deferred Tool Calls | Pydantic Docs](#handle-deferred-tool-calls-pydantic-docs) - [Bank Support | Pydantic Docs](#bank-support-pydantic-docs) - [Text to Audio | Pydantic Docs](#text-to-audio-pydantic-docs) - [Server | Pydantic Docs](#server-pydantic-docs) - [Overview | Pydantic Docs](#overview-pydantic-docs) - [Raise Content Filter Error | Pydantic Docs](#raise-content-filter-error-pydantic-docs) - [Include Tool Return Schemas | Pydantic Docs](#include-tool-return-schemas-pydantic-docs) - [Prefix Tools | Pydantic Docs](#prefix-tools-pydantic-docs) - [Chat App with FastAPI | Pydantic Docs](#chat-app-with-fastapi-pydantic-docs) - [Data Analyst | Pydantic Docs](#data-analyst-pydantic-docs) - [Crusoe | Pydantic Docs](#crusoe-pydantic-docs) - [Cohere | Pydantic Docs](#cohere-pydantic-docs) - [Audio, images, and transcripts | Pydantic Docs](#audio-images-and-transcripts-pydantic-docs) - [Set Tool Metadata | Pydantic Docs](#set-tool-metadata-pydantic-docs) - [Web Search | Pydantic Docs](#web-search-pydantic-docs) - [pydantic_evals.generation | Pydantic Docs](#pydantic-evals-generation-pydantic-docs) - [zai | Pydantic Docs](#zai-pydantic-docs) - [X Search | Pydantic Docs](#x-search-pydantic-docs) - [Timeouts | Pydantic Docs](#timeouts-pydantic-docs) - [Third-Party Integrations | Pydantic Docs](#third-party-integrations-pydantic-docs) - [Logfire Integration | Pydantic Docs](#logfire-integration-pydantic-docs) - [Dataset Serialization | Pydantic Docs](#dataset-serialization-pydantic-docs) - [Setup | Pydantic Docs](#setup-pydantic-docs) - [Stream Whales | Pydantic Docs](#stream-whales-pydantic-docs) - [Version Policy | Pydantic Docs](#version-policy-pydantic-docs) - [Browser WebRTC | Pydantic Docs](#browser-webrtc-pydantic-docs) - [Connection lifecycle | Pydantic Docs](#connection-lifecycle-pydantic-docs) - [Connecting a frontend | Pydantic Docs](#connecting-a-frontend-pydantic-docs) - [pydantic_graph.exceptions | Pydantic Docs](#pydantic-graph-exceptions-pydantic-docs) - [Usage and observability | Pydantic Docs](#usage-and-observability-pydantic-docs) - [pydantic_evals.online_capability | Pydantic Docs](#pydantic-evals-online-capability-pydantic-docs) - [template | Pydantic Docs](#template-pydantic-docs) - [Prepare Tools | Pydantic Docs](#prepare-tools-pydantic-docs) - [Third-Party Capabilities | Pydantic Docs](#third-party-capabilities-pydantic-docs) - [xai | Pydantic Docs](#xai-pydantic-docs) - [Direct Model Requests | Pydantic Docs](#direct-model-requests-pydantic-docs) - [Simple Validation | Pydantic Docs](#simple-validation-pydantic-docs) - [Standard Quality Metrics | Pydantic Docs](#standard-quality-metrics-pydantic-docs) - [Cerebras | Pydantic Docs](#cerebras-pydantic-docs) - [Getting Started | Pydantic Docs](#getting-started-pydantic-docs) - [SQL Generation | Pydantic Docs](#sql-generation-pydantic-docs) - [pydantic_evals.lifecycle | Pydantic Docs](#pydantic-evals-lifecycle-pydantic-docs) - [openrouter | Pydantic Docs](#openrouter-pydantic-docs) - [Image Generation | Pydantic Docs](#image-generation-pydantic-docs) - [test | Pydantic Docs](#test-pydantic-docs) - [Overview | Pydantic Docs](#overview-pydantic-docs) - [Quick Start | Pydantic Docs](#quick-start-pydantic-docs) - [Flight Booking | Pydantic Docs](#flight-booking-pydantic-docs) - [TwelveLabs Video Agent | Pydantic Docs](#twelvelabs-video-agent-pydantic-docs) - [ext | Pydantic Docs](#ext-pydantic-docs) - [Snowflake Cortex | Pydantic Docs](#snowflake-cortex-pydantic-docs) - [OpenAI | Pydantic Docs](#openai-pydantic-docs) - [xAI | Pydantic Docs](#xai-pydantic-docs) - [pydantic_graph.basenode | Pydantic Docs](#pydantic-graph-basenode-pydantic-docs) - [Tools | Pydantic Docs](#tools-pydantic-docs) - [Process History | Pydantic Docs](#process-history-pydantic-docs) - [Select Model | Pydantic Docs](#select-model-pydantic-docs) - [mistral | Pydantic Docs](#mistral-pydantic-docs) - [Stream Markdown | Pydantic Docs](#stream-markdown-pydantic-docs) - [Steps | Pydantic Docs](#steps-pydantic-docs) - [Kitaru | Pydantic Docs](#kitaru-pydantic-docs) - [Testing | Pydantic Docs](#testing-pydantic-docs) - [Hugging Face | Pydantic Docs](#hugging-face-pydantic-docs) - [Mistral | Pydantic Docs](#mistral-pydantic-docs) - [Concurrency & Performance | Pydantic Docs](#concurrency-performance-pydantic-docs) - [RAG | Pydantic Docs](#rag-pydantic-docs) - [Multimodal Input | Pydantic Docs](#multimodal-input-pydantic-docs) - [Shell | Pydantic Docs](#shell-pydantic-docs) - [Groq | Pydantic Docs](#groq-pydantic-docs) - [Parallel Execution | Pydantic Docs](#parallel-execution-pydantic-docs) - [MCP | Pydantic Docs](#mcp-pydantic-docs) - [Restate | Pydantic Docs](#restate-pydantic-docs) - [History and handoff | Pydantic Docs](#history-and-handoff-pydantic-docs) - [Reinject System Prompt | Pydantic Docs](#reinject-system-prompt-pydantic-docs) - [Retry Strategies | Pydantic Docs](#retry-strategies-pydantic-docs) - [Agentic Evaluators | Pydantic Docs](#agentic-evaluators-pydantic-docs) - [Extensibility | Pydantic Docs](#extensibility-pydantic-docs) - [Conversation Search | Pydantic Docs](#conversation-search-pydantic-docs) - [Overview | Pydantic Docs](#overview-pydantic-docs) - [Macroscope | Pydantic Docs](#macroscope-pydantic-docs) - [Z.AI | Pydantic Docs](#z-ai-pydantic-docs) - [xAI | Pydantic Docs](#xai-pydantic-docs) - [Troubleshooting | Pydantic Docs](#troubleshooting-pydantic-docs) - [Thread Executor | Pydantic Docs](#thread-executor-pydantic-docs) - [pydantic_graph.util | Pydantic Docs](#pydantic-graph-util-pydantic-docs) - [Dependencies | Pydantic Docs](#dependencies-pydantic-docs) - [Multi-Run Evaluation | Pydantic Docs](#multi-run-evaluation-pydantic-docs) - [Pydantic AI Harness](#pydantic-ai-harness) - [Getting Help | Pydantic Docs](#getting-help-pydantic-docs) - [Warn On Cache Busts | Pydantic Docs](#warn-on-cache-busts-pydantic-docs) - [Azure | Pydantic Docs](#azure-pydantic-docs) - [System Reminders | Pydantic Docs](#system-reminders-pydantic-docs) - [ACP (Agent Client Protocol)](#acp-agent-client-protocol-) - [Apache Airflow | Pydantic Docs](#apache-airflow-pydantic-docs) - [pydantic_graph.node | Pydantic Docs](#pydantic-graph-node-pydantic-docs) - [Turns and interruptions | Pydantic Docs](#turns-and-interruptions-pydantic-docs) - [Weather Agent | Pydantic Docs](#weather-agent-pydantic-docs) - [Retries | Pydantic Docs](#retries-pydantic-docs) - [Multi-Agent Patterns | Pydantic Docs](#multi-agent-patterns-pydantic-docs) - [Events | Pydantic Docs](#events-pydantic-docs) - [V1 → V2 Migration Map | Pydantic Docs](#v1-v2-migration-map-pydantic-docs) - [pydantic_graph.join | Pydantic Docs](#pydantic-graph-join-pydantic-docs) - [Pydantic Evals](#pydantic-evals) - [xai | Pydantic Docs](#xai-pydantic-docs) - [Overview | Pydantic Docs](#overview-pydantic-docs) - [Core Concepts | Pydantic Docs](#core-concepts-pydantic-docs) - [Vercel AI | Pydantic Docs](#vercel-ai-pydantic-docs) - [Ollama | Pydantic Docs](#ollama-pydantic-docs) - [Web Chat UI | Pydantic Docs](#web-chat-ui-pydantic-docs) - [google | Pydantic Docs](#google-pydantic-docs) - [LLM Judge | Pydantic Docs](#llm-judge-pydantic-docs) - [Advisor | Pydantic Docs](#advisor-pydantic-docs) - [FileSystem | Pydantic Docs](#filesystem-pydantic-docs) - [Media Externalization](#media-externalization) - [Managed Prompt | Pydantic Docs](#managed-prompt-pydantic-docs) - [Command Line Interface (CLI) | Pydantic Docs](#command-line-interface-cli-pydantic-docs) - [AG-UI | Pydantic Docs](#ag-ui-pydantic-docs) - [Contributing | Pydantic Docs](#contributing-pydantic-docs) - [ag_ui | Pydantic Docs](#ag-ui-pydantic-docs) - [Pydantic AI Gateway | Pydantic Docs](#pydantic-ai-gateway-pydantic-docs) - [output | Pydantic Docs](#output-pydantic-docs) - [Joins & Reducers | Pydantic Docs](#joins-reducers-pydantic-docs) - [Runtime Capability Creation | Pydantic Docs](#runtime-capability-creation-pydantic-docs) - [Agent Specs | Pydantic Docs](#agent-specs-pydantic-docs) - [DBOS | Pydantic Docs](#dbos-pydantic-docs) - [Business Applications | Pydantic Docs](#business-applications-pydantic-docs) - [UI Examples | Pydantic Docs](#ui-examples-pydantic-docs) - [Repo Context | Pydantic Docs](#repo-context-pydantic-docs) - [StackOne | Pydantic Docs](#stackone-pydantic-docs) - [Pydantic AI | Pydantic Docs](#pydantic-ai-pydantic-docs) - [Overview | Pydantic Docs](#overview-pydantic-docs) - [retries | Pydantic Docs](#retries-pydantic-docs) - [Dataset Management | Pydantic Docs](#dataset-management-pydantic-docs) - [Prefect | Pydantic Docs](#prefect-pydantic-docs) - [Voice Assistant | Pydantic Docs](#voice-assistant-pydantic-docs) - [Tool Output Limits | Pydantic Docs](#tool-output-limits-pydantic-docs) - [Overview | Pydantic Docs](#overview-pydantic-docs) - [openai | Pydantic Docs](#openai-pydantic-docs) - [Thinking | Pydantic Docs](#thinking-pydantic-docs) - [Built-in Evaluators | Pydantic Docs](#built-in-evaluators-pydantic-docs) - [Custom Evaluators | Pydantic Docs](#custom-evaluators-pydantic-docs) - [Overview | Pydantic Docs](#overview-pydantic-docs) - [Report Evaluators | Pydantic Docs](#report-evaluators-pydantic-docs) - [Capabilities and hooks | Pydantic Docs](#capabilities-and-hooks-pydantic-docs) - [Dynamic Workflow | Pydantic Docs](#dynamic-workflow-pydantic-docs) - [pydantic_evals.otel | Pydantic Docs](#pydantic-evals-otel-pydantic-docs) - [Hooks | Pydantic Docs](#hooks-pydantic-docs) - [Planning | Pydantic Docs](#planning-pydantic-docs) - [pydantic_graph.decision | Pydantic Docs](#pydantic-graph-decision-pydantic-docs) - [exceptions | Pydantic Docs](#exceptions-pydantic-docs) - [Metrics & Attributes | Pydantic Docs](#metrics-attributes-pydantic-docs) - [Skills | Pydantic Docs](#skills-pydantic-docs) - [Span-Based | Pydantic Docs](#span-based-pydantic-docs) - [concurrency | Pydantic Docs](#concurrency-pydantic-docs) - [Pydantic AI Docs | Pydantic Docs](#pydantic-ai-docs-pydantic-docs) - [OpenRouter | Pydantic Docs](#openrouter-pydantic-docs) - [LocalStack | Pydantic Docs](#localstack-pydantic-docs) - [azure | Pydantic Docs](#azure-pydantic-docs) - [Bedrock | Pydantic Docs](#bedrock-pydantic-docs) - [settings | Pydantic Docs](#settings-pydantic-docs) - [Messages and chat history | Pydantic Docs](#messages-and-chat-history-pydantic-docs) - [direct | Pydantic Docs](#direct-pydantic-docs) - [Input, Output & Tool Guardrails](#input-output-tool-guardrails) - [Upgrade Guide | Pydantic Docs](#upgrade-guide-pydantic-docs) - [Subagents | Pydantic Docs](#subagents-pydantic-docs) - [Google Gemini | Pydantic Docs](#google-gemini-pydantic-docs) - [pydantic_graph.step | Pydantic Docs](#pydantic-graph-step-pydantic-docs) - [function_signature | Pydantic Docs](#function-signature-pydantic-docs) - [Code Mode | Pydantic Docs](#code-mode-pydantic-docs) - [HTTP Request Retries | Pydantic Docs](#http-request-retries-pydantic-docs) - [Debugging & Monitoring with Pydantic Logfire | Pydantic Docs](#debugging-monitoring-with-pydantic-logfire-pydantic-docs) - [Google | Pydantic Docs](#google-pydantic-docs) - [usage | Pydantic Docs](#usage-pydantic-docs) - [codec | Pydantic Docs](#codec-pydantic-docs) - [Anthropic | Pydantic Docs](#anthropic-pydantic-docs) - [Client | Pydantic Docs](#client-pydantic-docs) - [native_tools | Pydantic Docs](#native-tools-pydantic-docs) - [run | Pydantic Docs](#run-pydantic-docs) - [Exa Search | Pydantic Docs](#exa-search-pydantic-docs) - [On-Demand Capabilities | Pydantic Docs](#on-demand-capabilities-pydantic-docs) - [Camera Agent | Pydantic Docs](#camera-agent-pydantic-docs) - [Online Evaluation | Pydantic Docs](#online-evaluation-pydantic-docs) --- # Welcome to Pydantic | Pydantic Docs [Skip to content](https://pydantic.dev/docs/validation/latest/#_top) Welcome to Pydantic =================== [![CI](https://img.shields.io/github/actions/workflow/status/pydantic/pydantic/ci.yml?branch=main&logo=github&label=CI)](https://github.com/pydantic/pydantic/actions?query=event%3Apush+branch%3Amain+workflow%3ACI) [![Coverage](https://coverage-badge.samuelcolvin.workers.dev/pydantic/pydantic.svg)](https://github.com/pydantic/pydantic/actions?query=event%3Apush+branch%3Amain+workflow%3ACI) [![pypi](https://img.shields.io/pypi/v/pydantic.svg)](https://pypi.python.org/pypi/pydantic) [![CondaForge](https://img.shields.io/conda/v/conda-forge/pydantic.svg)](https://anaconda.org/conda-forge/pydantic) [![downloads](https://static.pepy.tech/badge/pydantic/month)](https://pepy.tech/project/pydantic) [![license](https://img.shields.io/github/license/pydantic/pydantic.svg)](https://github.com/pydantic/pydantic/blob/main/LICENSE) [![llms.txt](https://img.shields.io/badge/llms.txt-green)](https://docs.pydantic.dev/latest/llms.txt) Documentation for version: v2.13.4. Pydantic is the most widely used data validation library for Python. Fast and extensible, Pydantic plays nicely with your linters/IDE/brain. Define how data should be in pure, canonical Python 3.9+; validate it with Pydantic. **Sign up for our newsletter, _The Pydantic Stack_, with updates & tutorials on Pydantic, Logfire, and Pydantic AI:** Subscribe Why use Pydantic? ----------------- [](https://pydantic.dev/docs/validation/latest/#why-use-pydantic) * **Powered by type hints** — with Pydantic, schema validation and serialization are controlled by type annotations; less to learn, less code to write, and integration with your IDE and static analysis tools. [Learn more…](https://pydantic.dev/docs/validation/latest/get-started/why/#type-hints) * **Speed** — Pydantic’s core validation logic is written in Rust. As a result, Pydantic is among the fastest data validation libraries for Python. [Learn more…](https://pydantic.dev/docs/validation/latest/get-started/why/#performance) * **JSON Schema** — Pydantic models can emit JSON Schema, allowing for easy integration with other tools. [Learn more…](https://pydantic.dev/docs/validation/latest/get-started/why/#json-schema) * **Strict** and **Lax** mode — Pydantic can run in either strict mode (where data is not converted) or lax mode where Pydantic tries to coerce data to the correct type where appropriate. [Learn more…](https://pydantic.dev/docs/validation/latest/get-started/why/#strict-lax) * **Dataclasses**, **TypedDicts** and more — Pydantic supports validation of many standard library types including `dataclass` and `TypedDict`. [Learn more…](https://pydantic.dev/docs/validation/latest/get-started/why/#dataclasses-typeddict-more) * **Customisation** — Pydantic allows custom validators and serializers to alter how data is processed in many powerful ways. [Learn more…](https://pydantic.dev/docs/validation/latest/get-started/why/#customisation) * **Ecosystem** — around 8,000 packages on PyPI use Pydantic, including massively popular libraries like _FastAPI_, _huggingface_, _Django Ninja_, _SQLModel_, & _LangChain_. [Learn more…](https://pydantic.dev/docs/validation/latest/get-started/why/#ecosystem) * **Battle tested** — Pydantic is downloaded over 550M times/month and is used by all FAANG companies and 20 of the 25 largest companies on NASDAQ. If you’re trying to do something with Pydantic, someone else has probably already done it. [Learn more…](https://pydantic.dev/docs/validation/latest/get-started/why/#using-pydantic) [Installing Pydantic](https://pydantic.dev/docs/validation/latest/get-started/install/) is as simple as: `pip install pydantic` Pydantic examples ----------------- [](https://pydantic.dev/docs/validation/latest/#pydantic-examples) To see Pydantic at work, let’s start with a simple example, creating a custom class that inherits from `BaseModel`: Validation Successful from datetime import datetimefrom pydantic import BaseModel, PositiveIntclass User(BaseModel): id: int name: str = 'John Doe' signup_ts: datetime | None tastes: dict[str, PositiveInt] external_data = { 'id': 123, 'signup_ts': '2019-06-01 12:22', 'tastes': { 'wine': 9, b'cheese': 7, 'cabbage': '1', },}user = User(**external_data) print(user.id) #> 123print(user.model_dump()) """{ 'id': 123, 'name': 'John Doe', 'signup_ts': datetime.datetime(2019, 6, 1, 12, 22), 'tastes': {'wine': 9, 'cheese': 7, 'cabbage': 1},}""" `id` is of type `int`; the annotation-only declaration tells Pydantic that this field is required. Strings, bytes, or floats will be coerced to integers if possible; otherwise an exception will be raised. `name` is a string; because it has a default, it is not required. `signup_ts` is a [`datetime`](https://docs.python.org/3/library/datetime.html#datetime.datetime) field that is required, but the value `None` may be provided; Pydantic will process either a [Unix timestamp](https://en.wikipedia.org/wiki/Unix_time) integer (e.g. `1496498400`) or a string representing the date and time. `tastes` is a dictionary with string keys and positive integer values. The `PositiveInt` type is shorthand for `Annotated[int, annotated_types.Gt(0)]`. The input here is an [ISO 8601](https://en.wikipedia.org/wiki/ISO_8601) formatted datetime, but Pydantic will convert it to a [`datetime`](https://docs.python.org/3/library/datetime.html#datetime.datetime) object. The key here is `bytes`, but Pydantic will take care of coercing it to a string. Similarly, Pydantic will coerce the string `'1'` to the integer `1`. We create instance of `User` by passing our external data to `User` as keyword arguments. We can access fields as attributes of the model. We can convert the model to a dictionary with [`model_dump()`](https://pydantic.dev/docs/validation/latest/api/pydantic/base_model/#pydantic.BaseModel.model_dump) . If validation fails, Pydantic will raise an error with a breakdown of what was wrong: Validation Error # continuing the above example...from datetime import datetimefrom pydantic import BaseModel, PositiveInt, ValidationErrorclass User(BaseModel): id: int name: str = 'John Doe' signup_ts: datetime | None tastes: dict[str, PositiveInt]external_data = {'id': 'not an int', 'tastes': {}} try: User(**external_data) except ValidationError as e: print(e.errors()) """ [ { 'type': 'int_parsing', 'loc': ('id',), 'msg': 'Input should be a valid integer, unable to parse string as an integer', 'input': 'not an int', 'url': 'https://errors.pydantic.dev/2/v/int_parsing', }, { 'type': 'missing', 'loc': ('signup_ts',), 'msg': 'Field required', 'input': {'id': 'not an int', 'tastes': {}}, 'url': 'https://errors.pydantic.dev/2/v/missing', }, ] """ The input data is wrong here — `id` is not a valid integer, and `signup_ts` is missing. Trying to instantiate `User` will raise a [`ValidationError`](https://pydantic.dev/docs/validation/latest/api/pydantic-core/pydantic_core/#pydantic_core.ValidationError) with a list of errors. Who is using Pydantic? ---------------------- [](https://pydantic.dev/docs/validation/latest/#who-is-using-pydantic) Hundreds of organisations and packages are using Pydantic. Some of the prominent companies and organizations around the world who are using Pydantic include: [![Adobe](https://pydantic.dev/docs/validation/logos/adobe_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-adobe "Adobe") [![Amazon and AWS](https://pydantic.dev/docs/validation/logos/amazon_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-amazon "Amazon and AWS") [![Anthropic](https://pydantic.dev/docs/validation/logos/anthropic_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-anthropic "Anthropic") [![Apple](https://pydantic.dev/docs/validation/logos/apple_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-apple "Apple") [![ASML](https://pydantic.dev/docs/validation/logos/asml_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-asml "ASML") [![AstraZeneca](https://pydantic.dev/docs/validation/logos/astrazeneca_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-astrazeneca "AstraZeneca") [![Cisco Systems](https://pydantic.dev/docs/validation/logos/cisco_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-cisco "Cisco Systems") [![Capital One](https://pydantic.dev/docs/validation/logos/capital_one_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-capital_one "Capital One") [![Comcast](https://pydantic.dev/docs/validation/logos/comcast_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-comcast "Comcast") [![Datadog](https://pydantic.dev/docs/validation/logos/datadog_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-datadog "Datadog") [![Facebook](https://pydantic.dev/docs/validation/logos/facebook_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-facebook "Facebook") [![GitHub](https://pydantic.dev/docs/validation/logos/github_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-github "GitHub") [![Google](https://pydantic.dev/docs/validation/logos/google_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-google "Google") [![HSBC](https://pydantic.dev/docs/validation/logos/hsbc_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-hsbc "HSBC") [![IBM](https://pydantic.dev/docs/validation/logos/ibm_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-ibm "IBM") [![Intel](https://pydantic.dev/docs/validation/logos/intel_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-intel "Intel") [![Intuit](https://pydantic.dev/docs/validation/logos/intuit_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-intuit "Intuit") [![Intergovernmental Panel on Climate Change](https://pydantic.dev/docs/validation/logos/ipcc_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-ipcc "Intergovernmental Panel on Climate Change") [![JPMorgan](https://pydantic.dev/docs/validation/logos/jpmorgan_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-jpmorgan "JPMorgan") [![Jupyter](https://pydantic.dev/docs/validation/logos/jupyter_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-jupyter "Jupyter") [![Microsoft](https://pydantic.dev/docs/validation/logos/microsoft_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-microsoft "Microsoft") [![Molecular Science Software Institute](https://pydantic.dev/docs/validation/logos/molssi_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-molssi "Molecular Science Software Institute") [![NASA](https://pydantic.dev/docs/validation/logos/nasa_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-nasa "NASA") [![Netflix](https://pydantic.dev/docs/validation/logos/netflix_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-netflix "Netflix") [![NSA](https://pydantic.dev/docs/validation/logos/nsa_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-nsa "NSA") [![NVIDIA](https://pydantic.dev/docs/validation/logos/nvidia_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-nvidia "NVIDIA") [![OpenAI](https://pydantic.dev/docs/validation/logos/openai_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-openai "OpenAI") [![Oracle](https://pydantic.dev/docs/validation/logos/oracle_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-oracle "Oracle") [![Palantir](https://pydantic.dev/docs/validation/logos/palantir_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-palantir "Palantir") [![Qualcomm](https://pydantic.dev/docs/validation/logos/qualcomm_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-qualcomm "Qualcomm") [![Red Hat](https://pydantic.dev/docs/validation/logos/redhat_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-redhat "Red Hat") [![Revolut](https://pydantic.dev/docs/validation/logos/revolut_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-revolut "Revolut") [![Robusta](https://pydantic.dev/docs/validation/logos/robusta_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-robusta "Robusta") [![Salesforce](https://pydantic.dev/docs/validation/logos/salesforce_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-salesforce "Salesforce") [![Starbucks](https://pydantic.dev/docs/validation/logos/starbucks_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-starbucks "Starbucks") [![Texas Instruments](https://pydantic.dev/docs/validation/logos/ti_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-ti "Texas Instruments") [![Twilio](https://pydantic.dev/docs/validation/logos/twilio_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-twilio "Twilio") [![Twitter](https://pydantic.dev/docs/validation/logos/twitter_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-twitter "Twitter") [![UK Home Office](https://pydantic.dev/docs/validation/logos/ukhomeoffice_logo.png)](https://pydantic.dev/docs/validation/latest/why/#org-ukhomeoffice "UK Home Office") For a more comprehensive list of open-source projects using Pydantic see the [list of dependents on github](https://github.com/pydantic/pydantic/network/dependents) , or you can find some awesome projects using Pydantic in [awesome-pydantic](https://github.com/Kludex/awesome-pydantic) . Was this page helpful? Thanks for your feedback! --- # Pydantic Docs - Validation, AI Agents, Logfire Observability [Skip to content](https://pydantic.dev/docs/#_top) Pydantic Docs ============= Pydantic is the end-to-end AI Engineering stack. Validate untrusted data and build agents you’ll actually ship to production using our Python libraries. Observe, govern, and optimize agents written in any language or framework once they’re live. [### Pydantic Validation\ \ Data validation from Python type hints. Parse, validate, and serialize with confidence.\ \ class User(BaseModel):\ name: str\ age: int\ \ Learn about Validation→](https://pydantic.dev/docs/validation/latest/get-started/) [### Pydantic AI\ \ The batteries-included type-safe framework for building production agents.\ \ model = 'openai:gpt-5.6-sol'\ agent = Agent(model)\ agent.run\_sync('Does it snow?')\ \ ![](https://pydantic.dev/docs/_home/providers/openai.svg)\ \ Build production-ready Agents→](https://pydantic.dev/docs/ai/overview/) [Not just Python!Logfire lets you monitor, secure, and optimize agents written in Python, TypeScript, Rust, Go, Java, Ruby, or any other language.\ \ ### Pydantic Logfire\ \ Observability and governance for AI agents, LLMs, applications, services, and hosts.\ \ logfire.configure();\ logfire.info('app started');\ \ See the best AI Engineering platform→](https://pydantic.dev/docs/logfire/get-started/) --- # bedrock_mantle | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#_top) bedrock\_mantle =============== Setup ----- [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#setup) For details on how to set up authentication with these models, see [model configuration for Bedrock Mantle](https://pydantic.dev/docs/ai/models/bedrock/#bedrock-mantle) . BedrockMantleChatModel ---------------------- [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.BedrockMantleChatModel) **Bases:** `OpenAIChatModel` An OpenAI Chat Completions model served by Amazon Bedrock Mantle (GPT-OSS Safeguard). The response-scoped tool-call-ID normalization added for #6536 is Responses-only: Mantle’s Chat Completions API returns globally-unique `chatcmpl-tool-*` IDs across separate responses (verified live), unlike the `/openai/v1/responses` endpoint’s per-response `call_0` counter, so the Chat path needs no normalization. ### Methods [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.BedrockMantleChatModel.__init__) def __init__( model_name: BedrockMantleModelName, *, provider: Literal['bedrock-mantle'] | BedrockMantleProvider = 'bedrock-mantle', profile: ModelProfileSpec | None = None, settings: OpenAIChatModelSettings | None = None, ) -> None Initialize a Bedrock Mantle Chat Completions model. ##### Returns [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#returns) [`None`](https://docs.python.org/3/library/constants.html#None) ##### Parameters [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#parameters) **`model_name`** : `BedrockMantleModelName` [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.BedrockMantleChatModel.__init__(model_name)) The name of the model, e.g. `openai.gpt-oss-safeguard-20b`. **`provider`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘bedrock-mantle’\] | `BedrockMantleProvider` _Default:_ `'bedrock-mantle'` [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.BedrockMantleChatModel.__init__(provider)) The provider to use. Defaults to the `bedrock-mantle` provider. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.BedrockMantleChatModel.__init__(profile)) The model profile to use. Defaults to a profile picked by the provider based on the model name. **`settings`** : `OpenAIChatModelSettings` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.BedrockMantleChatModel.__init__(settings)) The model settings to use. Defaults to `None`. BedrockMantleResponsesModel --------------------------- [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.BedrockMantleResponsesModel) **Bases:** `OpenAIResponsesModel` An OpenAI Responses model served by Amazon Bedrock Mantle. Serves GPT-5.4+ (on the `/openai/v1` endpoint) and GPT-OSS (on the `/v1` endpoint); the endpoint is chosen from the model profile. ### Methods [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#methods-1) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.BedrockMantleResponsesModel.__init__) def __init__( model_name: BedrockMantleModelName, *, provider: Literal['bedrock-mantle'] | BedrockMantleProvider = 'bedrock-mantle', profile: ModelProfileSpec | None = None, settings: OpenAIResponsesModelSettings | None = None, ) -> None Initialize a Bedrock Mantle Responses model. ##### Returns [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#returns-1) [`None`](https://docs.python.org/3/library/constants.html#None) ##### Parameters [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#parameters-1) **`model_name`** : `BedrockMantleModelName` [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.BedrockMantleResponsesModel.__init__(model_name)) The name of the model, e.g. `openai.gpt-5.6-luna`. **`provider`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘bedrock-mantle’\] | `BedrockMantleProvider` _Default:_ `'bedrock-mantle'` [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.BedrockMantleResponsesModel.__init__(provider)) The provider to use. Defaults to the `bedrock-mantle` provider. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.BedrockMantleResponsesModel.__init__(profile)) The model profile to use. Defaults to a profile picked by the provider based on the model name. **`settings`** : `OpenAIResponsesModelSettings` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.BedrockMantleResponsesModel.__init__(settings)) The model settings to use. Defaults to `None`. BedrockMantleModelName ---------------------- [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.BedrockMantleModelName) Possible Amazon Bedrock Mantle model names. Since Bedrock Mantle supports a variety of OpenAI models and the list changes frequently, we explicitly list the latest models but allow any name in the type hints. **Default:** `str | LatestBedrockMantleModelNames` LatestBedrockMantleModelNames ----------------------------- [](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.LatestBedrockMantleModelNames) Latest OpenAI models served through Amazon Bedrock Mantle. **Default:** `Literal['openai.gpt-5.4', 'openai.gpt-5.4-2026-03-05', 'openai.gpt-5.5', 'openai.gpt-5.5-2026-04-23', 'openai.gpt-5.6-luna', 'openai.gpt-5.6-sol', 'openai.gpt-5.6-terra', 'openai.gpt-oss-20b', 'openai.gpt-oss-120b', 'openai.gpt-oss-safeguard-20b', 'openai.gpt-oss-safeguard-120b']` Was this page helpful? Thanks for your feedback! --- # bedrock | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/bedrock/#_top) bedrock ======= Setup ----- [](https://pydantic.dev/docs/ai/api/models/bedrock/#setup) For details on how to set up authentication with this model, see [model configuration for Bedrock](https://pydantic.dev/docs/ai/models/bedrock/) . BedrockConverseModel -------------------- [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockConverseModel) **Bases:** `Model[BaseClient]` A model that uses the Bedrock Converse API. ### Attributes [](https://pydantic.dev/docs/ai/api/models/bedrock/#attributes) #### client [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockConverseModel.client) The boto3 client used to make requests to the Bedrock Converse API. Defaults to the client from the [`Provider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.Provider) . It can be reassigned, e.g. to rotate short-lived credentials in a long-running service, but prefer assigning to [`BedrockProvider.client`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.bedrock.BedrockProvider.client) so all models sharing the provider pick up the new client. Once you’ve assigned a client here, you’re responsible for keeping it valid; the provider’s client is no longer consulted. **Type:** `BedrockRuntimeClient` #### model\_name [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockConverseModel.model_name) The model name. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### system [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockConverseModel.system) The model provider. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) ### Methods [](https://pydantic.dev/docs/ai/api/models/bedrock/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockConverseModel.__init__) def __init__( model_name: BedrockModelName, *, provider: Literal['bedrock', 'gateway'] | Provider[BaseClient] = 'bedrock', profile: ModelProfileSpec | None = None, settings: ModelSettings | None = None, ) Initialize a Bedrock model. ##### Parameters [](https://pydantic.dev/docs/ai/api/models/bedrock/#parameters) **`model_name`** : `BedrockModelName` [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockConverseModel.__init__(model_name)) The name of the model to use. **`model_name`** : `BedrockModelName` [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockConverseModel.__init__(model_name)) The name of the Bedrock model to use. List of model names available [here](https://docs.aws.amazon.com/bedrock/latest/userguide/models-supported.html) . **`provider`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘bedrock’, ‘gateway’\] | `Provider`\[`BaseClient`\] _Default:_ `'bedrock'` [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockConverseModel.__init__(provider)) The provider to use for authentication and API access. Can be either the string ‘bedrock’ or an instance of `Provider[BaseClient]`. If not provided, a new provider will be created using the other parameters. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockConverseModel.__init__(profile)) The model profile to use. Defaults to a profile picked by the provider based on the model name. **`settings`** : [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockConverseModel.__init__(settings)) Model-specific settings that will be used as defaults for this model. #### count\_tokens [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockConverseModel.count_tokens) `@async` def count_tokens( messages: list[ModelMessage], model_settings: ModelSettings | None, model_request_parameters: ModelRequestParameters, ) -> usage.RequestUsage Count the number of tokens, works with limited models. Check the actual supported models on [https://docs.aws.amazon.com/bedrock/latest/userguide/count-tokens.html](https://docs.aws.amazon.com/bedrock/latest/userguide/count-tokens.html) ##### Returns [](https://pydantic.dev/docs/ai/api/models/bedrock/#returns) [`usage.RequestUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage) #### resolve\_prompt\_cache\_retention [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockConverseModel.resolve_prompt_cache_retention) def resolve_prompt_cache_retention( model_settings: ModelSettings | None, ) -> timedelta | None Resolve the longest retention requested by supported Bedrock cache settings. ##### Returns [](https://pydantic.dev/docs/ai/api/models/bedrock/#returns-1) `timedelta` | [`None`](https://docs.python.org/3/library/constants.html#None) #### supported\_native\_tools [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockConverseModel.supported_native_tools) `@classmethod` def supported_native_tools(cls) -> frozenset[type[AbstractNativeTool]] The set of builtin tool types this model can handle. ##### Returns [](https://pydantic.dev/docs/ai/api/models/bedrock/#returns-2) [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[[`type`](https://docs.python.org/3/glossary.html#term-type)\ \[`AbstractNativeTool`\]\] BedrockModelSettings -------------------- [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings) **Bases:** [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) Settings for Bedrock models. See [the Bedrock Converse API docs](https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_Converse.html#API_runtime_Converse_RequestSyntax) for a full list. See [the boto3 implementation](https://boto3.amazonaws.com/v1/documentation/api/latest/reference/services/bedrock-runtime/client/converse.html) of the Bedrock Converse API. `extra_headers` are injected before the request is signed, so under SigV4 authentication they are covered by the signature (except the few headers botocore never signs, e.g. `X-Amzn-Trace-Id`). Headers the AWS SDK computes itself (e.g. `Authorization`, `User-Agent`, `X-Amz-Date`) are overwritten by botocore afterwards. ### Attributes [](https://pydantic.dev/docs/ai/api/models/bedrock/#attributes-1) #### bedrock\_additional\_model\_requests\_fields [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_additional_model_requests_fields) Additional model-specific parameters to include in requests. See more about it on [https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters.html](https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters.html) . **Type:** [`Mapping`](https://docs.python.org/3/library/typing.html#typing.Mapping) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] #### bedrock\_additional\_model\_response\_fields\_paths [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_additional_model_response_fields_paths) JSON paths to extract additional fields from model responses. See more about it on [https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters.html](https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters.html) . **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] #### bedrock\_cache\_instructions [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_cache_instructions) Whether to add a cache point after the system prompt blocks. When enabled, an extra `cachePoint` is appended to the system prompt so Bedrock can cache system instructions. Set to `True` or `'5m'` for a 5-minute TTL (the default), or `'1h'` for a 1-hour TTL. See [https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html](https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html) for more information. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘5m’, ‘1h’\] #### bedrock\_cache\_messages [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_cache_messages) Convenience setting to enable caching for the last user message. When enabled, this automatically adds a cache point to the last content block in the final user message, which is useful for caching conversation history or context in multi-turn conversations. Set to `True` or `'5m'` for a 5-minute TTL (the default), or `'1h'` for a 1-hour TTL. Note: Uses 1 of Bedrock’s 4 available cache points per request. Any additional CachePoint markers in messages will be automatically limited to respect the 4-cache-point maximum. See [https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html](https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html) for more information. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘5m’, ‘1h’\] #### bedrock\_cache\_tool\_definitions [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_cache_tool_definitions) Whether to add a cache point after the last tool definition. When enabled, the last tool in the `tools` array will include a `cachePoint`, allowing Bedrock to cache tool definitions and reduce costs for compatible models. Set to `True` or `'5m'` for a 5-minute TTL (the default), or `'1h'` for a 1-hour TTL. See [https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html](https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html) for more information. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘5m’, ‘1h’\] #### bedrock\_guardrail\_config [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_guardrail_config) Content moderation and safety settings for Bedrock API requests. See more about it on [https://docs.aws.amazon.com/bedrock/latest/APIReference/API\_runtime\_GuardrailConfiguration.html](https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_GuardrailConfiguration.html) . **Type:** `GuardrailConfigurationTypeDef` #### bedrock\_inference\_profile [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_inference_profile) An [inference profile](https://docs.aws.amazon.com/bedrock/latest/userguide/inference-profiles.html) ARN to use as the `modelId` in API requests. When set, this value is used as the `modelId` in `converse` and `converse_stream` API calls instead of the base `model_name`. This allows you to pass the base model name (e.g. `'anthropic.claude-sonnet-4-5-20250929-v1:0'`) as `model_name` for detecting model capabilities and token counting, while routing requests through an inference profile for cost tracking or cross-region inference. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### bedrock\_performance\_configuration [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_performance_configuration) Performance optimization settings for model inference. See more about it on [https://docs.aws.amazon.com/bedrock/latest/APIReference/API\_runtime\_PerformanceConfiguration.html](https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_PerformanceConfiguration.html) . **Type:** `PerformanceConfigurationTypeDef` #### bedrock\_prompt\_variables [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_prompt_variables) Variables for substitution into prompt templates. See more about it on [https://docs.aws.amazon.com/bedrock/latest/APIReference/API\_runtime\_PromptVariableValues.html](https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_PromptVariableValues.html) . **Type:** [`Mapping`](https://docs.python.org/3/library/typing.html#typing.Mapping) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , `PromptVariableValuesTypeDef`\] #### bedrock\_request\_metadata [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_request_metadata) Additional metadata to attach to Bedrock API requests. See more about it on [https://docs.aws.amazon.com/bedrock/latest/APIReference/API\_runtime\_Converse.html#API\_runtime\_Converse\_RequestSyntax](https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_Converse.html#API_runtime_Converse_RequestSyntax) . **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] #### bedrock\_service\_tier [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_service_tier) Setting for optimizing performance and cost. Accepts `{'type': 'default' | 'flex' | 'priority' | 'reserved'}`. Takes precedence over the top-level [`service_tier`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.service_tier) , and is the only way to request `'reserved'` (which requires a pre-purchased capacity reservation). See more about it on [https://docs.aws.amazon.com/bedrock/latest/userguide/service-tiers-inference.html](https://docs.aws.amazon.com/bedrock/latest/userguide/service-tiers-inference.html) . **Type:** `ServiceTierTypeDef` BedrockStreamedResponse ----------------------- [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockStreamedResponse) **Bases:** `StreamedResponse` Implementation of `StreamedResponse` for Bedrock models. ### Attributes [](https://pydantic.dev/docs/ai/api/models/bedrock/#attributes-2) #### model\_name [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockStreamedResponse.model_name) Get the model name of the response. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### provider\_name [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockStreamedResponse.provider_name) Get the provider name. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### provider\_url [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockStreamedResponse.provider_url) Get the provider base URL. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) BedrockModelName ---------------- [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelName) Possible Bedrock model names. Since Bedrock supports a variety of date-stamped models, we explicitly list the latest models but allow any name in the type hints. See [the Bedrock docs](https://docs.aws.amazon.com/bedrock/latest/userguide/models-supported.html) for a full list. **Default:** `str | LatestBedrockModelNames` LatestBedrockModelNames ----------------------- [](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.LatestBedrockModelNames) Latest Bedrock models. **Default:** `Literal['amazon.titan-tg1-large', 'amazon.titan-text-lite-v1', 'amazon.titan-text-express-v1', 'us.amazon.nova-2-lite-v1:0', 'us.amazon.nova-pro-v1:0', 'us.amazon.nova-lite-v1:0', 'us.amazon.nova-micro-v1:0', 'anthropic.claude-3-5-sonnet-20241022-v2:0', 'us.anthropic.claude-3-5-sonnet-20241022-v2:0', 'anthropic.claude-3-5-haiku-20241022-v1:0', 'us.anthropic.claude-3-5-haiku-20241022-v1:0', 'anthropic.claude-instant-v1', 'anthropic.claude-v2:1', 'anthropic.claude-v2', 'anthropic.claude-3-sonnet-20240229-v1:0', 'us.anthropic.claude-3-sonnet-20240229-v1:0', 'anthropic.claude-3-haiku-20240307-v1:0', 'us.anthropic.claude-3-haiku-20240307-v1:0', 'anthropic.claude-3-opus-20240229-v1:0', 'us.anthropic.claude-3-opus-20240229-v1:0', 'anthropic.claude-3-5-sonnet-20240620-v1:0', 'us.anthropic.claude-3-5-sonnet-20240620-v1:0', 'anthropic.claude-3-7-sonnet-20250219-v1:0', 'us.anthropic.claude-3-7-sonnet-20250219-v1:0', 'anthropic.claude-opus-4-20250514-v1:0', 'us.anthropic.claude-opus-4-20250514-v1:0', 'global.anthropic.claude-opus-4-5-20251101-v1:0', 'anthropic.claude-sonnet-4-20250514-v1:0', 'us.anthropic.claude-sonnet-4-20250514-v1:0', 'eu.anthropic.claude-sonnet-4-20250514-v1:0', 'anthropic.claude-sonnet-4-5-20250929-v1:0', 'us.anthropic.claude-sonnet-4-5-20250929-v1:0', 'eu.anthropic.claude-sonnet-4-5-20250929-v1:0', 'anthropic.claude-sonnet-4-6', 'us.anthropic.claude-sonnet-4-6', 'eu.anthropic.claude-sonnet-4-6', 'anthropic.claude-haiku-4-5-20251001-v1:0', 'us.anthropic.claude-haiku-4-5-20251001-v1:0', 'eu.anthropic.claude-haiku-4-5-20251001-v1:0', 'cohere.command-text-v14', 'cohere.command-r-v1:0', 'cohere.command-r-plus-v1:0', 'cohere.command-light-text-v14', 'meta.llama3-8b-instruct-v1:0', 'meta.llama3-70b-instruct-v1:0', 'meta.llama3-1-8b-instruct-v1:0', 'us.meta.llama3-1-8b-instruct-v1:0', 'meta.llama3-1-70b-instruct-v1:0', 'us.meta.llama3-1-70b-instruct-v1:0', 'meta.llama3-1-405b-instruct-v1:0', 'us.meta.llama3-2-11b-instruct-v1:0', 'us.meta.llama3-2-90b-instruct-v1:0', 'us.meta.llama3-2-1b-instruct-v1:0', 'us.meta.llama3-2-3b-instruct-v1:0', 'us.meta.llama3-3-70b-instruct-v1:0', 'mistral.mistral-7b-instruct-v0:2', 'mistral.mixtral-8x7b-instruct-v0:1', 'mistral.mistral-large-2402-v1:0', 'mistral.mistral-large-2407-v1:0', 'us.anthropic.claude-opus-4-1-20250805-v1:0', 'us.anthropic.claude-opus-4-5-20251101-v1:0', 'us.anthropic.claude-opus-4-6-v1', 'global.anthropic.claude-opus-4-6-v1', 'us.anthropic.claude-opus-4-7', 'global.anthropic.claude-opus-4-7', 'us.anthropic.claude-opus-4-8', 'global.anthropic.claude-opus-4-8', 'us.anthropic.claude-opus-5', 'global.anthropic.claude-opus-5', 'us.anthropic.claude-sonnet-5', 'global.anthropic.claude-sonnet-5', 'us.anthropic.claude-fable-5', 'global.anthropic.claude-fable-5', 'us.amazon.nova-premier-v1:0', 'global.amazon.nova-2-lite-v1:0', 'us.meta.llama4-maverick-17b-instruct-v1:0', 'us.meta.llama4-scout-17b-instruct-v1:0', 'mistral.mistral-small-2402-v1:0', 'mistral.mistral-large-3-675b-instruct', 'mistral.ministral-3-3b-instruct', 'mistral.ministral-3-8b-instruct', 'mistral.ministral-3-14b-instruct', 'mistral.magistral-small-2509', 'mistral.devstral-2-123b', 'mistral.pixtral-large-2502-v1:0', 'us.mistral.pixtral-large-2502-v1:0', 'deepseek.r1-v1:0', 'deepseek.v3.2', 'qwen.qwen3-32b-v1:0', 'qwen.qwen3-coder-30b-a3b-v1:0', 'qwen.qwen3-coder-next', 'qwen.qwen3-next-80b-a3b', 'qwen.qwen3-vl-235b-a22b', 'google.gemma-3-4b-it', 'google.gemma-3-12b-it', 'google.gemma-3-27b-it', 'minimax.minimax-m2', 'minimax.minimax-m2.1', 'minimax.minimax-m2.5', 'nvidia.nemotron-nano-9b-v2', 'nvidia.nemotron-nano-12b-v2', 'nvidia.nemotron-nano-3-30b', 'nvidia.nemotron-super-3-120b', 'us.writer.palmyra-x4-v1:0', 'us.writer.palmyra-x5-v1:0', 'zai.glm-4.7', 'zai.glm-4.7-flash', 'zai.glm-5', 'moonshot.kimi-k2-thinking', 'moonshotai.kimi-k2.5']` Was this page helpful? Thanks for your feedback! --- # anthropic | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/anthropic/#_top) anthropic ========= Setup ----- [](https://pydantic.dev/docs/ai/api/models/anthropic/#setup) For details on how to set up authentication with this model, see [model configuration for Anthropic](https://pydantic.dev/docs/ai/models/anthropic/) . AnthropicCompaction ------------------- [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicCompaction) **Bases:** `AbstractCapability[AgentDepsT]` Compaction capability for Anthropic models. Configures automatic context management via Anthropic’s `context_management` API parameter. Compaction triggers server-side when input tokens exceed the configured threshold. Example usage: from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicCompaction agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[AnthropicCompaction(token_threshold=100_000)], ) ### Methods [](https://pydantic.dev/docs/ai/api/models/anthropic/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicCompaction.__init__) def __init__( *, token_threshold: int = 150000, instructions: str | None = None, pause_after_compaction: bool = False, ) -> None Initialize the Anthropic compaction capability. ##### Returns [](https://pydantic.dev/docs/ai/api/models/anthropic/#returns) [`None`](https://docs.python.org/3/library/constants.html#None) ##### Parameters [](https://pydantic.dev/docs/ai/api/models/anthropic/#parameters) **`token_threshold`** : [`int`](https://docs.python.org/3/library/functions.html#int) _Default:_ `150000` [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicCompaction.__init__(token_threshold)) Compact when input tokens exceed this threshold. Minimum 50,000. **`instructions`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicCompaction.__init__(instructions)) Custom instructions for the compaction summarization. **`pause_after_compaction`** : [`bool`](https://docs.python.org/3/library/functions.html#bool) _Default:_ `False` [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicCompaction.__init__(pause_after_compaction)) If `True`, the response will stop after the compaction block with `stop_reason='compaction'`, allowing explicit handling. AnthropicModel -------------- [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModel) **Bases:** `Model[AsyncAnthropicClient]` A model that uses the Anthropic API. Internally, this uses the [Anthropic Python client](https://github.com/anthropics/anthropic-sdk-python) to interact with the API. Apart from `__init__`, all methods are private or match those of the base class. ### Attributes [](https://pydantic.dev/docs/ai/api/models/anthropic/#attributes) #### model\_name [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModel.model_name) The model name. **Type:** `AnthropicModelName` #### profile [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModel.profile) The model profile. Anthropic web-tool availability depends on both model support and the client/platform, so the profile’s `supported_native_tools` and `anthropic_supports_dynamic_filtering` are narrowed here for clients that don’t support them (e.g. Bedrock, Vertex). `supports_inline_system_prompts` is narrowed the same way, and for the same reason: serving a `{'role': 'system'}` entry is a fact about the transport as much as about the model. **Type:** `AnthropicModelProfile` #### system [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModel.system) The model provider. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### tool\_addition\_mode [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModel.tool_addition_mode) The effective addition mode, narrowed for transports without inline system messages. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘by\_reference’\] | [`None`](https://docs.python.org/3/library/constants.html#None) ### Methods [](https://pydantic.dev/docs/ai/api/models/anthropic/#methods-1) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModel.__init__) def __init__( model_name: AnthropicModelName, *, provider: Literal['anthropic', 'gateway'] | Provider[AsyncAnthropicClient] = 'anthropic', profile: ModelProfileSpec | None = None, settings: ModelSettings | None = None, ) Initialize an Anthropic model. ##### Parameters [](https://pydantic.dev/docs/ai/api/models/anthropic/#parameters-1) **`model_name`** : `AnthropicModelName` [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModel.__init__(model_name)) The name of the Anthropic model to use. List of model names available [here](https://docs.anthropic.com/en/docs/about-claude/models) . **`provider`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘anthropic’, ‘gateway’\] | `Provider`\[`AsyncAnthropicClient`\] _Default:_ `'anthropic'` [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModel.__init__(provider)) The provider to use for the Anthropic API. Can be either the string ‘anthropic’ or an instance of `Provider[AsyncAnthropicClient]`. Defaults to ‘anthropic’. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModel.__init__(profile)) The model profile to use. Defaults to a profile picked by the provider based on the model name. The default ‘anthropic’ provider will use the default `..profiles.anthropic.anthropic_model_profile`. **`settings`** : [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModel.__init__(settings)) Default model settings for this model instance. #### resolve\_prompt\_cache\_retention [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModel.resolve_prompt_cache_retention) def resolve_prompt_cache_retention( model_settings: ModelSettings | None, ) -> timedelta | None Resolve the longest retention requested by active Anthropic cache settings. ##### Returns [](https://pydantic.dev/docs/ai/api/models/anthropic/#returns-1) `timedelta` | [`None`](https://docs.python.org/3/library/constants.html#None) #### supported\_native\_tools [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModel.supported_native_tools) `@classmethod` def supported_native_tools(cls) -> frozenset[type[AbstractNativeTool]] The set of builtin tool types this model can handle. ##### Returns [](https://pydantic.dev/docs/ai/api/models/anthropic/#returns-2) [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[[`type`](https://docs.python.org/3/glossary.html#term-type)\ \[`AbstractNativeTool`\]\] AnthropicModelSettings ---------------------- [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings) **Bases:** [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) Settings used for an Anthropic model request. ### Attributes [](https://pydantic.dev/docs/ai/api/models/anthropic/#attributes-1) #### anthropic\_betas [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_betas) List of Anthropic beta features to enable for API requests. Each item can be a known beta name (e.g. ‘interleaved-thinking-2025-05-14’) or a custom string. Merged with auto-added betas (e.g. builtin tools) and any betas from extra\_headers\[‘anthropic-beta’\]. See the Anthropic docs for available beta features. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[`AnthropicBetaParam`\] #### anthropic\_cache [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_cache) Enable prompt caching for multi-turn conversations. Passes a top-level `cache_control` parameter so the server automatically applies a cache breakpoint to the last cacheable block and moves it forward as conversations grow. On Bedrock and Vertex, automatic caching is not yet supported, so this falls back to per-block caching on the last user message. If the last content block already has `cache_control` from an explicit `CachePoint`, it is preserved. If `True`, uses TTL=‘5m’. You can also specify ‘5m’ or ‘1h’ directly. This can be combined with explicit cache breakpoints (`anthropic_cache_instructions`, `anthropic_cache_tool_definitions`, `CachePoint`). The automatic breakpoint counts as 1 of Anthropic’s 4 cache point slots; we automatically trim excess explicit breakpoints. See [https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching#automatic-caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching#automatic-caching) for more information. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘5m’, ‘1h’\] #### anthropic\_cache\_instructions [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_cache_instructions) Whether to add `cache_control` to the last system prompt block. When enabled, the last system prompt will have `cache_control` set, allowing Anthropic to cache system instructions and reduce costs. If `True`, uses TTL=‘5m’. You can also specify ‘5m’ or ‘1h’ directly. See [https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching) for more information. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘5m’, ‘1h’\] #### anthropic\_cache\_messages [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_cache_messages) Whether to add `cache_control` to the last message content block. This is an alternative to `anthropic_cache` for Anthropic-compatible gateways and proxies that accept the Anthropic message format but don’t support the top-level automatic caching parameter. If `True`, uses TTL=‘5m’. You can also specify ‘5m’ or ‘1h’ directly. Cannot be combined with `anthropic_cache`. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘5m’, ‘1h’\] #### anthropic\_cache\_tool\_definitions [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_cache_tool_definitions) Whether to add `cache_control` to the last tool definition. When enabled, the last tool in the `tools` array will have `cache_control` set, allowing Anthropic to cache tool definitions and reduce costs. If `True`, uses TTL=‘5m’. You can also specify ‘5m’ or ‘1h’ directly. See [https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching) for more information. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘5m’, ‘1h’\] #### anthropic\_code\_execution\_tool\_version [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_code_execution_tool_version) Which Anthropic code execution tool version to send for `CodeExecutionTool`. Defaults to `'auto'`, which uses the default version from the model profile: `'20260120'` for Sonnet 4.5+ and Opus 4.5+, otherwise `'20250825'`. Set a concrete version to force that tool version; a `UserError` is raised if the selected model profile does not support that version. **Type:** `AnthropicCodeExecutionToolVersion` | [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘auto’\] #### anthropic\_container [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_container) Container configuration for multi-turn conversations. By default, if previous messages contain a container\_id (from a prior response), it will be reused automatically. Set to `False` to force a fresh container (ignore any `container_id` from history). Set to a container id string (e.g. `'container_xxx'`) to explicitly reuse a container, or to a `BetaContainerParams` dict (e.g. `{'skills': [...]}` or `{'id': 'container_xxx', 'skills': [...]}`) when passing Skills to the Anthropic Skills beta. **Type:** `BetaContainerParams` | [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[[`False`](https://docs.python.org/3/library/constants.html#False)\ \] #### anthropic\_context\_management [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_context_management) Context management configuration for automatic compaction. When configured, Anthropic will automatically compact older context when the input token count exceeds the configured threshold. The compaction produces a summary that replaces the compacted messages. See [the Anthropic docs](https://docs.anthropic.com/en/docs/build-with-claude/compaction) for more details. **Type:** `BetaContextManagementConfigParam` #### anthropic\_eager\_input\_streaming [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_eager_input_streaming) Whether to enable eager input streaming on tool definitions. When enabled, all tool definitions will have `eager_input_streaming` set to `True`, allowing Anthropic to stream tool call arguments incrementally instead of buffering the entire JSON before streaming. This reduces latency for tool calls with large inputs. See [https://platform.claude.com/docs/en/agents-and-tools/tool-use/fine-grained-tool-streaming](https://platform.claude.com/docs/en/agents-and-tools/tool-use/fine-grained-tool-streaming) for more information. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### anthropic\_effort [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_effort) The effort level for the model to use when generating a response. See [the Anthropic docs](https://docs.anthropic.com/en/docs/build-with-claude/effort) for more information. **Type:** `AnthropicEffort` | [`None`](https://docs.python.org/3/library/constants.html#None) #### anthropic\_metadata [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_metadata) An object describing metadata about the request. Contains `user_id`, an external identifier for the user who is associated with the request. **Type:** `BetaMetadataParam` #### anthropic\_service\_tier [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_service_tier) The service tier to use for the model request. See [https://docs.anthropic.com/en/docs/build-with-claude/latency-and-throughput](https://docs.anthropic.com/en/docs/build-with-claude/latency-and-throughput) for more information. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘auto’, ‘standard\_only’\] #### anthropic\_speed [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_speed) The inference speed mode for this request. `'fast'` enables high output-tokens-per-second inference for supported models (currently Claude Opus 4.6, 4.7, 4.8, and 5). On unsupported models or clients, `anthropic_speed='fast'` is ignored with a `UserWarning`. Fast mode is a research preview and only available on the direct Anthropic API (not Bedrock, Vertex, or Foundry); see [the Anthropic docs](https://platform.claude.com/docs/en/build-with-claude/fast-mode) for details. Note: switching between `'fast'` and `'standard'` invalidates the prompt cache. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘standard’, ‘fast’\] #### anthropic\_task\_budget [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_task_budget) Task budget configuration for Anthropic beta requests. Maps to `output_config.task_budget`. Supported models are gated by the [`anthropic_supports_task_budgets`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.anthropic.AnthropicModelProfile.anthropic_supports_task_budgets) profile flag, and Pydantic AI automatically enables Anthropic’s required task-budget beta when this setting is present. Omit `remaining` unless you are intentionally carrying a budget across compaction or other rewritten context. **Type:** `AnthropicTaskBudget` #### anthropic\_thinking [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_thinking) Determine whether the model should generate a thinking block. See [the Anthropic docs](https://docs.anthropic.com/en/docs/build-with-claude/extended-thinking) for more information. **Type:** `BetaThinkingConfigParam` AnthropicStreamedResponse ------------------------- [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicStreamedResponse) **Bases:** `StreamedResponse` Implementation of `StreamedResponse` for Anthropic models. ### Attributes [](https://pydantic.dev/docs/ai/api/models/anthropic/#attributes-2) #### model\_name [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicStreamedResponse.model_name) Get the model name of the response. **Type:** `AnthropicModelName` #### provider\_name [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicStreamedResponse.provider_name) Get the provider name. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### provider\_url [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicStreamedResponse.provider_url) Get the provider base URL. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### timestamp [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicStreamedResponse.timestamp) Get the timestamp of the response. **Type:** [`datetime`](https://docs.python.org/3/library/datetime.html#module-datetime) AnthropicModelName ------------------ [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelName) Possible Anthropic model names. The installed Anthropic SDK exposes the current literal set and still allows arbitrary string model names. See [the Anthropic docs](https://docs.anthropic.com/en/docs/about-claude/models) for a full list. **Default:** `LatestAnthropicModelNames | Literal['claude-sonnet-5', 'claude-opus-5']` AnthropicTaskBudget ------------------- [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicTaskBudget) Anthropic task budget payload for `output_config.task_budget`. **Type:** [`TypeAlias`](https://docs.python.org/3/library/typing.html#typing.TypeAlias) **Default:** `BetaTokenTaskBudgetParam` DEPRECATED\_ANTHROPIC\_MODELS ----------------------------- [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.DEPRECATED_ANTHROPIC_MODELS) Models that have been retired by Anthropic but are still present in the SDK’s type definitions. **Type:** [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `frozenset({'claude-3-haiku-20240307', 'claude-opus-4-0', 'claude-opus-4-20250514', 'claude-sonnet-4-0', 'claude-sonnet-4-20250514'})` LatestAnthropicModelNames ------------------------- [](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.LatestAnthropicModelNames) Anthropic model names from the installed SDK. **Default:** `ModelParam` Was this page helpful? Thanks for your feedback! --- # cohere | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/cohere/#_top) cohere ====== Setup ----- [](https://pydantic.dev/docs/ai/api/models/cohere/#setup) For details on how to set up authentication with this model, see [model configuration for Cohere](https://pydantic.dev/docs/ai/models/cohere/) . CohereModel ----------- [](https://pydantic.dev/docs/ai/api/models/cohere/#pydantic_ai.models.cohere.CohereModel) **Bases:** `Model[AsyncClientV2]` A model that uses the Cohere API. Internally, this uses the [Cohere Python client](https://github.com/cohere-ai/cohere-python) to interact with the API. Apart from `__init__`, all methods are private or match those of the base class. ### Attributes [](https://pydantic.dev/docs/ai/api/models/cohere/#attributes) #### model\_name [](https://pydantic.dev/docs/ai/api/models/cohere/#pydantic_ai.models.cohere.CohereModel.model_name) The model name. **Type:** `CohereModelName` #### system [](https://pydantic.dev/docs/ai/api/models/cohere/#pydantic_ai.models.cohere.CohereModel.system) The model provider. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) ### Methods [](https://pydantic.dev/docs/ai/api/models/cohere/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/cohere/#pydantic_ai.models.cohere.CohereModel.__init__) def __init__( model_name: CohereModelName, *, provider: Literal['cohere'] | Provider[AsyncClientV2] = 'cohere', profile: ModelProfileSpec | None = None, settings: ModelSettings | None = None, ) Initialize an Cohere model. ##### Parameters [](https://pydantic.dev/docs/ai/api/models/cohere/#parameters) **`model_name`** : `CohereModelName` [](https://pydantic.dev/docs/ai/api/models/cohere/#pydantic_ai.models.cohere.CohereModel.__init__(model_name)) The name of the Cohere model to use. List of model names available [here](https://docs.cohere.com/docs/models#command) . **`provider`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘cohere’\] | `Provider`\[`AsyncClientV2`\] _Default:_ `'cohere'` [](https://pydantic.dev/docs/ai/api/models/cohere/#pydantic_ai.models.cohere.CohereModel.__init__(provider)) The provider to use for authentication and API access. Can be either the string ‘cohere’ or an instance of `Provider[AsyncClientV2]`. If not provided, a new provider will be created using the other parameters. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/cohere/#pydantic_ai.models.cohere.CohereModel.__init__(profile)) The model profile to use. Defaults to a profile picked by the provider based on the model name. **`settings`** : [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/cohere/#pydantic_ai.models.cohere.CohereModel.__init__(settings)) Model-specific settings that will be used as defaults for this model. CohereModelSettings ------------------- [](https://pydantic.dev/docs/ai/api/models/cohere/#pydantic_ai.models.cohere.CohereModelSettings) **Bases:** [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) Settings used for a Cohere model request. CohereModelName --------------- [](https://pydantic.dev/docs/ai/api/models/cohere/#pydantic_ai.models.cohere.CohereModelName) Possible Cohere model names. Since Cohere supports a variety of date-stamped models, we explicitly list the latest models but allow any name in the type hints. See [Cohere’s docs](https://docs.cohere.com/v2/docs/models) for a list of all available models. **Default:** `str | LatestCohereModelNames` LatestCohereModelNames ---------------------- [](https://pydantic.dev/docs/ai/api/models/cohere/#pydantic_ai.models.cohere.LatestCohereModelNames) Latest Cohere models. **Default:** `Literal['c4ai-aya-expanse-32b', 'c4ai-aya-expanse-8b', 'command-nightly', 'command-r-08-2024', 'command-r-plus-08-2024', 'command-r7b-12-2024']` Was this page helpful? Thanks for your feedback! --- # crusoe | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/crusoe/#_top) crusoe ====== Setup ----- [](https://pydantic.dev/docs/ai/api/models/crusoe/#setup) For details on how to set up authentication with this model, see [model configuration for Crusoe](https://pydantic.dev/docs/ai/models/crusoe/) . Crusoe model implementation using OpenAI-compatible API. CrusoeModel ----------- [](https://pydantic.dev/docs/ai/api/models/crusoe/#pydantic_ai.models.crusoe.CrusoeModel) **Bases:** `OpenAIChatModel` A model that uses Crusoe’s OpenAI-compatible Serverless Inference API. Crusoe serves open-weight models from many labs behind one endpoint, so the model family — and with it the profile [`CrusoeProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.crusoe.CrusoeProvider) resolves — is derived from the vendor prefix on the model name (`zai/`, `deepseek-ai/`, `meta-llama/`, …). Every model is served with guided decoding, so [`NativeOutput`](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.NativeOutput) works across the catalog, including for families whose own profiles don’t claim native structured output support. Thinking is returned in a non-standard field (`reasoning`, or `reasoning_content` for DeepSeek), both of which `OpenAIChatModel` reads. Apart from `__init__`, all methods are inherited from the base class. ### Methods [](https://pydantic.dev/docs/ai/api/models/crusoe/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/crusoe/#pydantic_ai.models.crusoe.CrusoeModel.__init__) def __init__( model_name: CrusoeModelName, *, provider: Literal['crusoe'] | Provider[AsyncOpenAI] = 'crusoe', profile: ModelProfileSpec | None = None, settings: ModelSettings | None = None, ) Initialize a Crusoe model. ##### Parameters [](https://pydantic.dev/docs/ai/api/models/crusoe/#parameters) **`model_name`** : `CrusoeModelName` [](https://pydantic.dev/docs/ai/api/models/crusoe/#pydantic_ai.models.crusoe.CrusoeModel.__init__(model_name)) The name of the Crusoe model to use, including the vendor prefix (e.g. `'zai/GLM-5.2'`). **`provider`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘crusoe’\] | `Provider`\[`AsyncOpenAI`\] _Default:_ `'crusoe'` [](https://pydantic.dev/docs/ai/api/models/crusoe/#pydantic_ai.models.crusoe.CrusoeModel.__init__(provider)) The provider to use. Defaults to `'crusoe'`. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/crusoe/#pydantic_ai.models.crusoe.CrusoeModel.__init__(profile)) The model profile to use. Defaults to a profile picked by the provider based on the model name. **`settings`** : [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/crusoe/#pydantic_ai.models.crusoe.CrusoeModel.__init__(settings)) Model-specific settings that will be used as defaults for this model. CrusoeModelName --------------- [](https://pydantic.dev/docs/ai/api/models/crusoe/#pydantic_ai.models.crusoe.CrusoeModelName) Possible Crusoe model names. Since Crusoe supports a variety of models and the list changes frequently, we explicitly list known models but allow any name in the type hints. See [https://docs.crusoecloud.com/serverless-inference/overview](https://docs.crusoecloud.com/serverless-inference/overview) for an up to date list of models. **Default:** `str | LatestCrusoeModelNames` Was this page helpful? Thanks for your feedback! --- # cerebras | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/cerebras/#_top) cerebras ======== Setup ----- [](https://pydantic.dev/docs/ai/api/models/cerebras/#setup) For details on how to set up authentication with this model, see [model configuration for Cerebras](https://pydantic.dev/docs/ai/models/cerebras/) . Cerebras model implementation using OpenAI-compatible API. CerebrasModel ------------- [](https://pydantic.dev/docs/ai/api/models/cerebras/#pydantic_ai.models.cerebras.CerebrasModel) **Bases:** `OpenAIChatModel` A model that uses Cerebras’s OpenAI-compatible API. Cerebras provides ultra-fast inference powered by the Wafer-Scale Engine (WSE). Apart from `__init__`, all methods are private or match those of the base class. ### Methods [](https://pydantic.dev/docs/ai/api/models/cerebras/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/cerebras/#pydantic_ai.models.cerebras.CerebrasModel.__init__) def __init__( model_name: CerebrasModelName, *, provider: Literal['cerebras'] | Provider[AsyncOpenAI] = 'cerebras', profile: ModelProfileSpec | None = None, settings: CerebrasModelSettings | None = None, ) Initialize a Cerebras model. ##### Parameters [](https://pydantic.dev/docs/ai/api/models/cerebras/#parameters) **`model_name`** : `CerebrasModelName` [](https://pydantic.dev/docs/ai/api/models/cerebras/#pydantic_ai.models.cerebras.CerebrasModel.__init__(model_name)) The name of the Cerebras model to use. **`provider`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘cerebras’\] | `Provider`\[`AsyncOpenAI`\] _Default:_ `'cerebras'` [](https://pydantic.dev/docs/ai/api/models/cerebras/#pydantic_ai.models.cerebras.CerebrasModel.__init__(provider)) The provider to use. Defaults to ‘cerebras’. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/cerebras/#pydantic_ai.models.cerebras.CerebrasModel.__init__(profile)) The model profile to use. Defaults to a profile based on the model name. **`settings`** : `CerebrasModelSettings` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/cerebras/#pydantic_ai.models.cerebras.CerebrasModel.__init__(settings)) Model-specific settings that will be used as defaults for this model. CerebrasModelSettings --------------------- [](https://pydantic.dev/docs/ai/api/models/cerebras/#pydantic_ai.models.cerebras.CerebrasModelSettings) **Bases:** [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) Settings used for a Cerebras model request. ALL FIELDS MUST BE `cerebras_` PREFIXED SO YOU CAN MERGE THEM WITH OTHER MODELS. ### Attributes [](https://pydantic.dev/docs/ai/api/models/cerebras/#attributes) #### cerebras\_clear\_thinking [](https://pydantic.dev/docs/ai/api/models/cerebras/#pydantic_ai.models.cerebras.CerebrasModelSettings.cerebras_clear_thinking) Whether Cerebras strips prior reasoning from earlier turns on multi-turn `zai`/GLM requests. `True` (Cerebras’s API default) drops thinking from previous turns before the next request; `False` preserves it, which improves multi-turn coherence and prompt-cache hit rates at the cost of more tokens. Pydantic AI sends `False` by default for `zai`/GLM models (which replay prior reasoning as `` tags) so the replayed reasoning isn’t stripped; set this explicitly to override. GLM-specific setting. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### cerebras\_disable\_reasoning [](https://pydantic.dev/docs/ai/api/models/cerebras/#pydantic_ai.models.cerebras.CerebrasModelSettings.cerebras_disable_reasoning) Disable reasoning for the model. Deprecated: use the unified `thinking=False` setting instead. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) CerebrasModelName ----------------- [](https://pydantic.dev/docs/ai/api/models/cerebras/#pydantic_ai.models.cerebras.CerebrasModelName) Possible Cerebras model names. Since Cerebras supports a variety of models and the list changes frequently, we explicitly list known models but allow any name in the type hints. See [https://inference-docs.cerebras.ai/models/overview](https://inference-docs.cerebras.ai/models/overview) for an up to date list of models. **Default:** `str | LatestCerebrasModelNames` Was this page helpful? Thanks for your feedback! --- # base | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/base/#_top) base ==== Logic related to making requests to an LLM. The aim here is to make a common interface for different LLMs, so that the rest of the code can be agnostic to the specific LLM being used. ModelRequestParameters ---------------------- [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters) Configuration for an agent’s request to a model, specifically related to tools and output handling. ### Attributes [](https://pydantic.dev/docs/ai/api/models/base/#attributes) #### declared\_function\_tools [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters.declared_function_tools) Function tools represented in the provider’s ordinary `tools` collection. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`ToolDefinition`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.ToolDefinition)\ \] #### declared\_tool\_defs [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters.declared_tool_defs) Definitions represented in the provider’s ordinary `tools` collection. The visibility filter applies to function tools only: output tools are always plain `tools` entries, so they are included unconditionally rather than keyed through a name-indexed filter a hidden function tool could shadow. **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`ToolDefinition`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.ToolDefinition)\ \] #### deferred\_capability\_ids [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters.deferred_capability_ids) IDs of the run’s capabilities that defer their loading. Read from the capability instances themselves, so it means what it says. It cannot be derived from the function tools: `ToolDefinition.capability_id` records which capability _contributed_ a tool, and `defer_loading` is set both by a deferred capability and by a search-gated tool inside an always-on one — so the two cases are indistinguishable from the definitions alone. Used to answer “may this tool be revealed yet?”: a tool whose `capability_id` is in this set is gated on that capability being loaded, while one whose owner is absent here is gated only on its own discovery. **Type:** [`set`](https://docs.python.org/3/reference/expressions.html#set) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `field(default_factory=(set[str]), repr=False)` #### instruction\_parts [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters.instruction_parts) Structured instruction parts with metadata about their origin (static vs dynamic). Static instructions (`dynamic=False`) come from literal strings passed to `Agent(instructions=...)`. Dynamic instructions (`dynamic=True`) come from `@agent.instructions` functions, `TemplateStr`, or toolset `get_instructions()` methods. Models that support granular caching (e.g. Anthropic, Bedrock) use this to place cache boundaries at the static/dynamic instruction boundary. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`InstructionPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.InstructionPart)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### revealed\_tool\_names [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters.revealed_tool_names) Names history has revealed so far, derived from the outgoing message list before each request. Discovered means evidenced by history; revealed means represented on this request’s wire state. Input to visibility resolution: `ToolDefinition.defer_loading` records what the author asked for and stays set after a reveal, so this answers the separate question of what the model can see _now_. History can name tools that no longer exist in the current run’s definitions, so this is not necessarily a subset of `function_tools`’ names; resolution ignores unknown names. **Type:** [`set`](https://docs.python.org/3/reference/expressions.html#set) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `field(default_factory=(set[str]), repr=False)` #### thinking [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters.thinking) Resolved thinking/reasoning configuration for this request. `None` means the model should use its default behavior. Set by the base `Model.prepare_request()` from the unified `thinking` field in `ModelSettings`, after checking that the model’s profile supports thinking. **Type:** `ThinkingLevel` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### tool\_visibility [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters.tool_visibility) Maps each function tool name to its resolved [`ToolVisibility`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ToolVisibility) . `None` on authored parameters; [`Model.prepare_request`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.prepare_request) populates an entry for every function tool, so a resolved request always carries a dict — empty exactly when there are no function tools. Output tools never get entries because they are always plain `tools` entries; [`visibility_of`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters.visibility_of) treats their absent entries like `'visible'`. The no-defaults `repr` omits the field until resolution, so authored parameters print as authored and resolved state stays visible. **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , `ToolVisibility`\] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` ### Methods [](https://pydantic.dev/docs/ai/api/models/base/#methods) #### visibility\_of [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters.visibility_of) def visibility_of(tool_name: str) -> ToolVisibility The resolved [`ToolVisibility`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ToolVisibility) for `tool_name`. For parameters constructed directly rather than resolved by [`Model.prepare_request`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.prepare_request) , deferred function tools default to `'withheld'` and every other name defaults to `'visible'`. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns) `ToolVisibility` #### with\_default\_output\_mode [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters.with_default_output_mode) def with_default_output_mode( output_mode: StructuredOutputMode, ) -> ModelRequestParameters Set the default output mode if the current mode is ‘auto’, atomically updating allow\_text\_output. No-op if the current output\_mode is not ‘auto’. This ensures the two fields stay in sync — output\_mode=‘tool’ implies allow\_text\_output=False, while ‘native’ and ‘prompted’ imply allow\_text\_output=True. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-1) `ModelRequestParameters` AbstractModel ------------- [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.AbstractModel) **Bases:** `ABC` Shared identity for request-response and realtime models. ### Attributes [](https://pydantic.dev/docs/ai/api/models/base/#attributes-1) #### base\_url [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.AbstractModel.base_url) The base URL for the provider API, if available. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### label [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.AbstractModel.label) Human-friendly display label for the model. Handles common patterns: * gpt-5 -> GPT 5 * claude-sonnet-4-5 -> Claude Sonnet 4.5 * gemini-2.5-pro -> Gemini 2.5 Pro * meta-llama/llama-3-70b -> Llama 3 70b (OpenRouter style) **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### model\_id [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.AbstractModel.model_id) The fully qualified model name in `'provider:model_name'` format. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### model\_name [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.AbstractModel.model_name) The model name. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### system [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.AbstractModel.system) The model provider, ex: openai. Use to populate the `gen_ai.system` OpenTelemetry semantic convention attribute, so should use well-known values listed in [https://opentelemetry.io/docs/specs/semconv/attributes-registry/gen-ai/#gen-ai-system](https://opentelemetry.io/docs/specs/semconv/attributes-registry/gen-ai/#gen-ai-system) when applicable. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) ### Methods [](https://pydantic.dev/docs/ai/api/models/base/#methods-1) #### \_\_aenter\_\_ [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.AbstractModel.__aenter__) `@async` def __aenter__() -> Self Enter the model context. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-2) [`Self`](https://docs.python.org/3/library/typing.html#typing.Self) #### \_\_aexit\_\_ [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.AbstractModel.__aexit__) `@async` def __aexit__( exc_type: type[BaseException] | None, exc_val: BaseException | None, exc_tb: TracebackType | None, ) -> bool | None Exit the model context. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-3) [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`None`](https://docs.python.org/3/library/constants.html#None) ModelRequestContext ------------------- [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestContext) Context for model request hooks. Wrapping these parameters in a dataclass instead of a tuple makes the signature future-proof: new fields can be added without breaking existing implementations. ### Attributes [](https://pydantic.dev/docs/ai/api/models/base/#attributes-2) #### model\_id [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestContext.model_id) The model-name string this request’s model was selected/resolved from, if any. This is the _selection_ token — e.g. `'openai:gpt-5.6-sol'`, or an alias like `'tenant-x'` that a [`resolve_model_id`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability.resolve_model_id) capability turned into a concrete model — so it can differ from the resolved model’s own `model_id`. `None` when the model was supplied as an instance rather than resolved from a string. Durable-execution capabilities carry this across the activity/step/task boundary in preference to the resolved model’s own `model_id`, so an aliased model round-trips as the original string the worker-side resolution chain can re-resolve. Only meaningful while `model` is still the run’s resolved model — a model swapped in by a hook invalidates it. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `field(default=None, init=False)` #### streaming [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestContext.streaming) Whether the agent loop expects to iterate the model response as a stream. Set for streamed runs — `run_stream()`, `run_stream_events()`, `iter()`’s node streaming — and for `run()` when an `event_stream_handler` is set or a capability overrides `wrap_run_event_stream` (e.g. `ProcessEventStream`, or a durability capability’s `event_stream_handler=`). There is no separate `before_model_request_stream` hook — streaming and non-streaming requests share the same hooks — so this field is how a hook can tell them apart. Read-only from hooks: reassigning it doesn’t change how the loop consumes the response. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `field(default=False, init=False)` ModelResolutionContext ---------------------- [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelResolutionContext) **Bases:** `Generic[ModelContextDepsT]` Context used to resolve a model ID before a model is available. This is narrower than [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) because model resolution happens before a run context can contain its resolved model. ### Attributes [](https://pydantic.dev/docs/ai/api/models/base/#attributes-3) #### agent [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelResolutionContext.agent) The agent whose model is being resolved. **Type:** `AbstractAgent`\[`ModelContextDepsT`, [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] #### deps [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelResolutionContext.deps) The dependencies supplied for this run. **Type:** `ModelContextDepsT` ModelSelectionContext --------------------- [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelSelectionContext) **Bases:** `ModelResolutionContext[ModelContextDepsT]` Context used by a capability to select the model for a request step. ### Attributes [](https://pydantic.dev/docs/ai/api/models/base/#attributes-4) #### messages [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelSelectionContext.messages) The message history available before this request step. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage)\ \] #### model [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelSelectionContext.model) The lower-precedence model on the first step, then the model used for the previous step. **Type:** `Model` | [`None`](https://docs.python.org/3/library/constants.html#None) #### run\_step [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelSelectionContext.run_step) The request step being selected, starting at `1`. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### usage [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelSelectionContext.usage) Usage accumulated by the run before this request step. **Type:** [`RunUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage) Model ----- [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) **Bases:** [`AbstractModel`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.AbstractModel) , `Generic[InterfaceClient]` Abstract class for a model. ### Attributes [](https://pydantic.dev/docs/ai/api/models/base/#attributes-5) #### compaction\_requires\_encrypted\_content [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.compaction_requires_encrypted_content) Whether this adapter’s API only honors a [`CompactionPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.CompactionPart) that carries encrypted content. When set, a part without it isn’t a wire boundary: the adapter would omit it, so letting it hide the earlier history would send nothing in its place. Declared by the adapter rather than the model profile: how an API carries compaction state is a property of the API, not of the model behind it — the same model reached through OpenAI’s Chat Completions and Responses APIs answers differently, and eight providers route a profile of their own through `OpenAIResponsesModel`. Independent of `compaction_retains_standing_prompt`, which today’s two adapters happen to answer the same way. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### compaction\_retains\_standing\_prompt [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.compaction_retains_standing_prompt) Whether this adapter’s compaction item keeps serving the leading system items of the window it replaced. When set, re-sending the standing prompt after the boundary would duplicate it. When not (the default), the standing prompt travels in a per-request channel rebuilt from those items, so the trim has to re-insert them or it is silently dropped from every subsequent request. See `compaction_requires_encrypted_content` for why this is declared here and not on the profile. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### profile [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.profile) The model profile. Resolution order (later layers override earlier ones): 1. `DEFAULT_PROFILE` — base values for every key in `ModelProfile`. 2. The provider’s `model_profile(model_name)` result — provider-specific defaults for this model. 3. The user’s `profile=` argument — partial dict merged on top, OR a callable `(default) -> profile` for full control. After resolution we compute the intersection of the profile’s `supported_native_tools` and the model class’s implemented tools, ensuring `model.profile['supported_native_tools']` is the single source of truth for what’s actually usable. **Type:** [`ModelProfile`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfile) #### provider [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.provider) The provider for this model, if any. **Type:** `Provider`\[`InterfaceClient`\] | [`None`](https://docs.python.org/3/library/constants.html#None) #### settings [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.settings) Get the model settings. **Type:** [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) #### supported\_tool\_addition\_modes [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.supported_tool_addition_modes) `tool_addition_mode` values this adapter’s renderer implements. See `supported_tool_deferral_modes`. **Type:** [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[`ToolAdditionMode`\] **Default:** `frozenset()` #### supported\_tool\_deferral\_modes [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.supported_tool_deferral_modes) `tool_deferral_mode` values this adapter’s renderer implements. A profile may claim a mode for the model family, but the claim only takes effect when the adapter class declares it here: `Model.tool_deferral_mode` intersects the two, so a `Model` subclass that declares nothing (the default) never resolves tools to a wire shape it cannot render, no matter what a pass-through vendor profile claims. **Type:** [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[`ToolDeferralMode`\] **Default:** `frozenset()` #### tool\_addition\_mode [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.tool_addition_mode) The effective tool-addition mode: the profile’s claim, if this adapter renders it. **Type:** `ToolAdditionMode` | [`None`](https://docs.python.org/3/library/constants.html#None) #### tool\_deferral\_mode [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.tool_deferral_mode) The effective schema-deferral mode: the profile’s claim, if this adapter renders it. **Type:** `ToolDeferralMode` | [`None`](https://docs.python.org/3/library/constants.html#None) ### Methods [](https://pydantic.dev/docs/ai/api/models/base/#methods-2) #### \_\_aenter\_\_ [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.__aenter__) `@async` def __aenter__() -> Self Enter the model context, delegating to the provider to manage its HTTP client lifecycle. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-4) [`Self`](https://docs.python.org/3/library/typing.html#typing.Self) #### \_\_aexit\_\_ [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.__aexit__) `@async` def __aexit__( exc_type: type[BaseException] | None, exc_val: BaseException | None, exc_tb: TracebackType | None, ) -> bool | None Exit the model context, closing the provider’s HTTP client if it owns one. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-5) [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`None`](https://docs.python.org/3/library/constants.html#None) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.__init__) def __init__( *, settings: ModelSettings | None = None, profile: ModelProfileSpec | None = None, ) -> None Initialize the model with optional settings and profile. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-6) [`None`](https://docs.python.org/3/library/constants.html#None) ##### Parameters [](https://pydantic.dev/docs/ai/api/models/base/#parameters) **`settings`** : [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.__init__(settings)) Model-specific settings that will be used as defaults for this model. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.__init__(profile)) The model profile to use. #### cancel\_suspended\_response [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.cancel_suspended_response) `@async` def cancel_suspended_response(response: ModelResponse) -> None Cancel a server-side suspended/background response (e.g. an OpenAI background job). Called when a continuation is abandoned via cancellation or error. No-op by default; model classes with cancellable server-side jobs override this. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-7) [`None`](https://docs.python.org/3/library/constants.html#None) #### compact\_messages [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.compact_messages) `@async` def compact_messages( request_context: ModelRequestContext, *, instructions: str | None = None, ) -> ModelResponse Compact messages to reduce conversation context size. This method is optional and only supported by specific providers (e.g. OpenAI Responses API). Providers that support compaction override this method with their implementation. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-8) [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) #### continuation\_delay [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.continuation_delay) def continuation_delay(response: ModelResponse) -> float | None Seconds to wait before continuing a suspended response, or `None` to continue immediately. Called between the segments of a suspended turn. `None` by default (e.g. Anthropic `pause_turn` continues immediately); a model that polls a server-side job (e.g. OpenAI background mode) overrides this to return a poll interval so the graph doesn’t busy-poll. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-9) [`float`](https://docs.python.org/3/library/functions.html#float) | [`None`](https://docs.python.org/3/library/constants.html#None) #### count\_tokens [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.count_tokens) `@async` def count_tokens( messages: list[ModelMessage], model_settings: ModelSettings | None, model_request_parameters: ModelRequestParameters, ) -> RequestUsage Make a request to the model for counting tokens. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-10) [`RequestUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage) #### customize\_request\_parameters [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.customize_request_parameters) def customize_request_parameters( model_request_parameters: ModelRequestParameters, ) -> ModelRequestParameters Customize the request parameters for the model. This method can be overridden by subclasses to modify the request parameters before sending them to the model. In particular, this method can be used to make modifications to the generated tool JSON schemas if necessary for vendor/model-specific reasons. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-11) `ModelRequestParameters` #### prepare\_messages [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.prepare_messages) def prepare_messages( messages: list[ModelMessage], model_request_parameters: ModelRequestParameters | None = None, ) -> list[ModelMessage] Pre-process the message history before it’s handed to the adapter’s message-prep step. Translates typed `NativeToolSearch*Part` instances carried over from a different provider (e.g. Anthropic to OpenAI Responses), or any native provider when the active model doesn’t support `ToolSearchTool`, into the local-shape `ToolSearch*Part` instances. This splits the single `ModelResponse(call+return)` carrying the inline server-side result into `ModelResponse(call) + ModelRequest(return)` so the adapter can render the provider-agnostic exchange. Also wraps non-leading `SystemPromptPart`s as ``\-tagged `UserPromptPart`s when the profile’s `supports_inline_system_prompts` is `False`, and converts `SpeechPart`s from realtime session history into `UserPromptPart`s / `TextPart`s that any model can consume. Subclasses normally don’t need to override this; the framework calls it on the agent’s behalf in `_agent_graph._make_request` so per-adapter message-prep code sees a homogeneous shape regardless of which provider produced the prior turn. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-12) [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage)\ \] ##### Parameters [](https://pydantic.dev/docs/ai/api/models/base/#parameters-1) **`messages`** : [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage)\ \] [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.prepare_messages(messages)) The history to pre-process. **`model_request_parameters`** : `ModelRequestParameters` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.prepare_messages(model_request_parameters)) The parameters this history will be sent with. Optional, and only needed to render a `ToolAvailabilityDeltaPart` on a model with no native way to express one: whether that reveal has to be a mechanism or can just be a statement depends on whether any tool actually goes on the wire with its schema withheld, which the profile alone can’t answer. Omitting it falls back to the adapter’s effective mode, which differs only for a corpus mixing capability-gated and standalone deferred tools. Framework callers pass it. #### prepare\_request [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.prepare_request) def prepare_request( model_settings: ModelSettings | None, model_request_parameters: ModelRequestParameters, ) -> tuple[ModelSettings | None, ModelRequestParameters] Prepare request inputs before they are passed to the provider. This merges the given `model_settings` with the model’s own `settings` attribute and ensures `customize_request_parameters` is applied to the resolved [`ModelRequestParameters`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters) . Subclasses can override this method if they need to customize the preparation flow further, but most implementations should simply call `self.prepare_request(...)` at the start of their `request` (and related) methods. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-13) [`tuple`](https://docs.python.org/3/library/stdtypes.html#tuple) \[[`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings)\ | [`None`](https://docs.python.org/3/library/constants.html#None)\ , `ModelRequestParameters`\] #### request [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.request) `@abstractmethod` `@async` def request( messages: list[ModelMessage], model_settings: ModelSettings | None, model_request_parameters: ModelRequestParameters, ) -> ModelResponse Make a request to the model. This is ultimately called by `pydantic_ai._agent_graph.ModelRequestNode._make_request(...)`. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-14) [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) #### request\_stream [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.request_stream) `@async` def request_stream( messages: list[ModelMessage], model_settings: ModelSettings | None, model_request_parameters: ModelRequestParameters, run_context: RunContext[Any] | None = None, ) -> AsyncGenerator[StreamedResponse] Make a request to the model and return a streaming response. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-15) [`AsyncGenerator`](https://docs.python.org/3/library/typing.html#typing.AsyncGenerator) \[`StreamedResponse`\] #### resolve\_prompt\_cache\_retention [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.resolve_prompt_cache_retention) def resolve_prompt_cache_retention( model_settings: ModelSettings | None, ) -> timedelta | None Resolve prompt cache retention requested by provider-specific model settings. The model’s default settings are merged with the per-request `model_settings`. Only provider-specific settings are currently considered; a future unified cache setting is not yet an input. If multiple active settings request different retention periods, the longest period wins because any longer-lived cache breakpoint can keep the corresponding prompt prefix available. Models without a provider-specific retention setting return `None`. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-16) `timedelta` | [`None`](https://docs.python.org/3/library/constants.html#None) #### supported\_native\_tools [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.supported_native_tools) `@classmethod` def supported_native_tools(cls) -> frozenset[type[AbstractNativeTool]] Return the set of native tool types this model class can handle. Subclasses should override this to reflect their actual capabilities. Default is empty set - subclasses must explicitly declare support. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-17) [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[[`type`](https://docs.python.org/3/glossary.html#term-type)\ \[`AbstractNativeTool`\]\] StreamedResponse ---------------- [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.StreamedResponse) **Bases:** `ABC` Streamed response from an LLM when calling a tool. ### Attributes [](https://pydantic.dev/docs/ai/api/models/base/#attributes-6) #### cancelled [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.StreamedResponse.cancelled) Whether the stream has been cancelled via `cancel()`. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### model\_name [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.StreamedResponse.model_name) Get the model name of the response. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### provider\_name [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.StreamedResponse.provider_name) Get the provider name. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### provider\_url [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.StreamedResponse.provider_url) Get the provider base URL. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### state [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.StreamedResponse.state) Lifecycle state of the response. **Type:** [`ModelResponseState`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponseState) **Default:** `field(default='complete', init=False)` #### timestamp [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.StreamedResponse.timestamp) Get the timestamp of the response. **Type:** [`datetime`](https://docs.python.org/3/library/datetime.html#module-datetime) #### usage [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.StreamedResponse.usage) Get the usage of the response so far. This will not be the final usage until the stream is exhausted. **Type:** [`RequestUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage) ### Methods [](https://pydantic.dev/docs/ai/api/models/base/#methods-3) #### \_\_aiter\_\_ [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.StreamedResponse.__aiter__) def __aiter__() -> AsyncIterator[ModelResponseStreamEvent] Stream the response as an async iterable of [`ModelResponseStreamEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponseStreamEvent) s. This proxies the `_event_iterator()` and emits all events, while also checking for matches on the result schema and emitting a [`FinalResultEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.FinalResultEvent) if/when the first match is found. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-18) [`AsyncIterator`](https://docs.python.org/3/library/typing.html#typing.AsyncIterator) \[[`ModelResponseStreamEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponseStreamEvent)\ \] #### cancel [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.StreamedResponse.cancel) `@async` def cancel() -> None Cancel local stream consumption and request provider shutdown. Sets `self._cancelled = True` before delegating to `close_stream()` so the flag is visible to any iterator that observes the transport error raised when the underlying connection is torn down, even if `close_stream()` itself raises. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-19) [`None`](https://docs.python.org/3/library/constants.html#None) #### close\_stream [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.StreamedResponse.close_stream) `@async` def close_stream() -> None Close the provider stream and any exposed HTTP or gRPC transport. Model classes must override this to close the local stream and, where the provider SDK exposes one, its transport. Integrations that cannot support local cancellation should leave the default implementation so `cancel()` fails clearly. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-20) [`None`](https://docs.python.org/3/library/constants.html#None) #### get [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.StreamedResponse.get) def get() -> ModelResponse Build a [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) from the data received from the stream so far. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-21) [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) #### get\_stream\_cancel\_errors [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.StreamedResponse.get_stream_cancel_errors) def get_stream_cancel_errors() -> tuple[type[BaseException], ...] Return transport errors caused by `cancel()` tearing down the stream. The default covers model classes whose SDKs iterate `httpx` responses directly (Anthropic, OpenAI, Groq, Mistral, Google GenAI, HuggingFace, and the custom Gemini client), since they let bare `httpx` errors propagate from chunk reads. Model classes that use other transports (for example gRPC or botocore) should override this method. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-22) [`tuple`](https://docs.python.org/3/library/stdtypes.html#tuple) \[[`type`](https://docs.python.org/3/glossary.html#term-type)\ \[[`BaseException`](https://docs.python.org/3/library/exceptions.html#BaseException)\ \], …\] #### time\_to\_first\_chunk [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.StreamedResponse.time_to_first_chunk) def time_to_first_chunk(request_start: float) -> float | None Seconds from `request_start` to the first chunk surfaced to the consumer, or `None` if nothing was yielded. `request_start` must be a `time.perf_counter()` reading taken when the request was issued. The first-chunk instant is stamped on the first `async for` pull, so the result reflects when the consumer _received_ the first event: it includes any consumer-side iteration delay (debouncing, batching, or awaiting other work) on top of the chunk’s transit time, which for eager consumers is negligible. ##### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-23) [`float`](https://docs.python.org/3/library/functions.html#float) | [`None`](https://docs.python.org/3/library/constants.html#None) known\_model\_names ------------------- [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.known_model_names) `@cached` def known_model_names() -> tuple[str, ...] Return every model name known to [`KnownModelName`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.KnownModelName) . This is the public, stable way to enumerate the known model ids. Prefer it over introspecting the `KnownModelName` type alias directly (e.g. `get_args(KnownModelName.__value__)`), which is not part of the public API and would break if the alias were ever recomposed. ### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-24) [`tuple`](https://docs.python.org/3/library/stdtypes.html#tuple) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , …\] check\_allow\_model\_requests ----------------------------- [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.check_allow_model_requests) def check_allow_model_requests() -> None Check if model requests are allowed. If you’re defining your own models that have costs or latency associated with their use, you should call this at the top of each method that sends a request to the provider: [`Model.request`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.request) , [`Model.request_stream`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.request_stream) , [`Model.count_tokens`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.count_tokens) , [`Model.compact_messages`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.compact_messages) , [`EmbeddingModel.embed`](https://pydantic.dev/docs/ai/api/pydantic-ai/embeddings/#pydantic_ai.embeddings.EmbeddingModel.embed) and [`EmbeddingModel.count_tokens`](https://pydantic.dev/docs/ai/api/pydantic-ai/embeddings/#pydantic_ai.embeddings.EmbeddingModel.count_tokens) . Methods that produce their result locally don’t need it — for example [`OpenAIEmbeddingModel`](https://pydantic.dev/docs/ai/api/pydantic-ai/embeddings/#pydantic_ai.embeddings.openai.OpenAIEmbeddingModel) ’s `count_tokens`, which tokenizes with `tiktoken` and never calls the provider. Neither does [`Model.cancel_suspended_response`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model.cancel_suspended_response) , which deliberately omits it so an already-started job can still be cancelled after the flag is flipped. ### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-25) [`None`](https://docs.python.org/3/library/constants.html#None) ### Raises [](https://pydantic.dev/docs/ai/api/models/base/#raises) * `RuntimeError` — If model requests are not allowed. infer\_model ------------ [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.infer_model) def infer_model( model: Model | KnownModelName | str, provider_factory: Callable[[str], Provider[Any]] = infer_provider, ) -> Model Infer the model from the name. ### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-26) `Model` ### Parameters [](https://pydantic.dev/docs/ai/api/models/base/#parameters-2) **`model`** : `Model` | `KnownModelName` | [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.infer_model(model)) Model name to instantiate, in the format of `provider:model`. Use the string “test” to instantiate TestModel. **`provider_factory`** : [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \], `Provider`\[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \]\] _Default:_ `infer_provider` [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.infer_model(provider_factory)) Function that instantiates a provider object. The provider name is passed into the function parameter. Defaults to `provider.infer_provider`. download\_item -------------- [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.download_item) `@async` def download_item( item: FileUrl, data_format: Literal['bytes'], type_format: Literal['mime', 'extension'] = 'mime', ) -> DownloadedItem[bytes] def download_item( item: FileUrl, data_format: Literal['base64', 'base64_uri', 'text'], type_format: Literal['mime', 'extension'] = 'mime', ) -> DownloadedItem[str] Download an item by URL and return the content as a bytes object or a (base64-encoded) string. This function includes SSRF (Server-Side Request Forgery) protection: * Only http:// and https:// protocols are allowed * Private/internal IP addresses are blocked by default * Cloud metadata endpoints (169.254.169.254) are always blocked * Hostnames are resolved before requests to prevent DNS rebinding * Response bodies are limited to 50 MiB Set `item.force_download='allow-local'` to allow private IP addresses. ### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-27) `DownloadedItem`\[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | `DownloadedItem`\[[`bytes`](https://docs.python.org/3/library/stdtypes.html#bytes)\ \] ### Parameters [](https://pydantic.dev/docs/ai/api/models/base/#parameters-3) **`item`** : [`FileUrl`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.FileUrl) [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.download_item(item)) The item to download. **`data_format`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘bytes’, ‘base64’, ‘base64\_uri’, ‘text’\] _Default:_ `'bytes'` [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.download_item(data_format)) The format to return the content in: * `bytes`: The raw bytes of the content. * `base64`: The base64-encoded content. * `base64_uri`: The base64-encoded content as a data URI. * `text`: The content as a string. **`type_format`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘mime’, ‘extension’\] _Default:_ `'mime'` [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.download_item(type_format)) The format to return the media type in: * `mime`: The media type as a MIME type. * `extension`: The media type as an extension. ### Raises [](https://pydantic.dev/docs/ai/api/models/base/#raises-1) * `UserError` — If the URL points to a YouTube video. * `ValueError` — If the URL uses an unsupported protocol or targets a private/internal IP address (unless allow-local is set), or the body exceeds 50 MiB. override\_allow\_model\_requests -------------------------------- [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.override_allow_model_requests) def override_allow_model_requests(allow_model_requests: bool) -> Generator[None] Context manager to temporarily override [`ALLOW_MODEL_REQUESTS`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ALLOW_MODEL_REQUESTS) . ### Returns [](https://pydantic.dev/docs/ai/api/models/base/#returns-28) [`Generator`](https://docs.python.org/3/library/typing.html#typing.Generator) \[[`None`](https://docs.python.org/3/library/constants.html#None)\ \] ### Parameters [](https://pydantic.dev/docs/ai/api/models/base/#parameters-4) **`allow_model_requests`** : [`bool`](https://docs.python.org/3/library/functions.html#bool) [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.override_allow_model_requests(allow_model_requests)) Whether to allow model requests within the context. KnownModelName -------------- [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.KnownModelName) Known model names that can be used with the `model` parameter of [`Agent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent) . `KnownModelName` is provided as a concise way to specify a model. **Default:** `TypeAliasType('KnownModelName', Literal['anthropic:claude-fable-5', 'anthropic:claude-haiku-4-5', 'anthropic:claude-haiku-4-5-20251001', 'anthropic:claude-mythos-5', 'anthropic:claude-mythos-preview', 'anthropic:claude-opus-4-1', 'anthropic:claude-opus-4-1-20250805', 'anthropic:claude-opus-4-5', 'anthropic:claude-opus-4-5-20251101', 'anthropic:claude-opus-4-6', 'anthropic:claude-opus-4-7', 'anthropic:claude-opus-4-8', 'anthropic:claude-opus-5', 'anthropic:claude-sonnet-4-5', 'anthropic:claude-sonnet-4-5-20250929', 'anthropic:claude-sonnet-4-6', 'anthropic:claude-sonnet-5', 'bedrock-mantle:openai.gpt-5.4', 'bedrock-mantle:openai.gpt-5.4-2026-03-05', 'bedrock-mantle:openai.gpt-5.5', 'bedrock-mantle:openai.gpt-5.5-2026-04-23', 'bedrock-mantle:openai.gpt-5.6-luna', 'bedrock-mantle:openai.gpt-5.6-sol', 'bedrock-mantle:openai.gpt-5.6-terra', 'bedrock-mantle:openai.gpt-oss-120b', 'bedrock-mantle:openai.gpt-oss-20b', 'bedrock-mantle:openai.gpt-oss-safeguard-120b', 'bedrock-mantle:openai.gpt-oss-safeguard-20b', 'bedrock:amazon.titan-text-express-v1', 'bedrock:amazon.titan-text-lite-v1', 'bedrock:amazon.titan-tg1-large', 'bedrock:anthropic.claude-3-5-haiku-20241022-v1:0', 'bedrock:anthropic.claude-3-5-sonnet-20240620-v1:0', 'bedrock:anthropic.claude-3-5-sonnet-20241022-v2:0', 'bedrock:anthropic.claude-3-7-sonnet-20250219-v1:0', 'bedrock:anthropic.claude-3-haiku-20240307-v1:0', 'bedrock:anthropic.claude-3-opus-20240229-v1:0', 'bedrock:anthropic.claude-3-sonnet-20240229-v1:0', 'bedrock:anthropic.claude-haiku-4-5-20251001-v1:0', 'bedrock:anthropic.claude-instant-v1', 'bedrock:anthropic.claude-opus-4-20250514-v1:0', 'bedrock:anthropic.claude-sonnet-4-20250514-v1:0', 'bedrock:anthropic.claude-sonnet-4-5-20250929-v1:0', 'bedrock:anthropic.claude-sonnet-4-6', 'bedrock:anthropic.claude-v2', 'bedrock:anthropic.claude-v2:1', 'bedrock:cohere.command-light-text-v14', 'bedrock:cohere.command-r-plus-v1:0', 'bedrock:cohere.command-r-v1:0', 'bedrock:cohere.command-text-v14', 'bedrock:deepseek.r1-v1:0', 'bedrock:deepseek.v3.2', 'bedrock:eu.anthropic.claude-haiku-4-5-20251001-v1:0', 'bedrock:eu.anthropic.claude-sonnet-4-20250514-v1:0', 'bedrock:eu.anthropic.claude-sonnet-4-5-20250929-v1:0', 'bedrock:eu.anthropic.claude-sonnet-4-6', 'bedrock:global.amazon.nova-2-lite-v1:0', 'bedrock:global.anthropic.claude-fable-5', 'bedrock:global.anthropic.claude-opus-4-5-20251101-v1:0', 'bedrock:global.anthropic.claude-opus-4-6-v1', 'bedrock:global.anthropic.claude-opus-4-7', 'bedrock:global.anthropic.claude-opus-4-8', 'bedrock:global.anthropic.claude-opus-5', 'bedrock:global.anthropic.claude-sonnet-5', 'bedrock:google.gemma-3-12b-it', 'bedrock:google.gemma-3-27b-it', 'bedrock:google.gemma-3-4b-it', 'bedrock:meta.llama3-1-405b-instruct-v1:0', 'bedrock:meta.llama3-1-70b-instruct-v1:0', 'bedrock:meta.llama3-1-8b-instruct-v1:0', 'bedrock:meta.llama3-70b-instruct-v1:0', 'bedrock:meta.llama3-8b-instruct-v1:0', 'bedrock:minimax.minimax-m2', 'bedrock:minimax.minimax-m2.1', 'bedrock:minimax.minimax-m2.5', 'bedrock:mistral.devstral-2-123b', 'bedrock:mistral.magistral-small-2509', 'bedrock:mistral.ministral-3-14b-instruct', 'bedrock:mistral.ministral-3-3b-instruct', 'bedrock:mistral.ministral-3-8b-instruct', 'bedrock:mistral.mistral-7b-instruct-v0:2', 'bedrock:mistral.mistral-large-2402-v1:0', 'bedrock:mistral.mistral-large-2407-v1:0', 'bedrock:mistral.mistral-large-3-675b-instruct', 'bedrock:mistral.mistral-small-2402-v1:0', 'bedrock:mistral.mixtral-8x7b-instruct-v0:1', 'bedrock:mistral.pixtral-large-2502-v1:0', 'bedrock:moonshot.kimi-k2-thinking', 'bedrock:moonshotai.kimi-k2.5', 'bedrock:nvidia.nemotron-nano-12b-v2', 'bedrock:nvidia.nemotron-nano-3-30b', 'bedrock:nvidia.nemotron-nano-9b-v2', 'bedrock:nvidia.nemotron-super-3-120b', 'bedrock:qwen.qwen3-32b-v1:0', 'bedrock:qwen.qwen3-coder-30b-a3b-v1:0', 'bedrock:qwen.qwen3-coder-next', 'bedrock:qwen.qwen3-next-80b-a3b', 'bedrock:qwen.qwen3-vl-235b-a22b', 'bedrock:us.amazon.nova-2-lite-v1:0', 'bedrock:us.amazon.nova-lite-v1:0', 'bedrock:us.amazon.nova-micro-v1:0', 'bedrock:us.amazon.nova-premier-v1:0', 'bedrock:us.amazon.nova-pro-v1:0', 'bedrock:us.anthropic.claude-3-5-haiku-20241022-v1:0', 'bedrock:us.anthropic.claude-3-5-sonnet-20240620-v1:0', 'bedrock:us.anthropic.claude-3-5-sonnet-20241022-v2:0', 'bedrock:us.anthropic.claude-3-7-sonnet-20250219-v1:0', 'bedrock:us.anthropic.claude-3-haiku-20240307-v1:0', 'bedrock:us.anthropic.claude-3-opus-20240229-v1:0', 'bedrock:us.anthropic.claude-3-sonnet-20240229-v1:0', 'bedrock:us.anthropic.claude-fable-5', 'bedrock:us.anthropic.claude-haiku-4-5-20251001-v1:0', 'bedrock:us.anthropic.claude-opus-4-1-20250805-v1:0', 'bedrock:us.anthropic.claude-opus-4-20250514-v1:0', 'bedrock:us.anthropic.claude-opus-4-5-20251101-v1:0', 'bedrock:us.anthropic.claude-opus-4-6-v1', 'bedrock:us.anthropic.claude-opus-4-7', 'bedrock:us.anthropic.claude-opus-4-8', 'bedrock:us.anthropic.claude-opus-5', 'bedrock:us.anthropic.claude-sonnet-4-20250514-v1:0', 'bedrock:us.anthropic.claude-sonnet-4-5-20250929-v1:0', 'bedrock:us.anthropic.claude-sonnet-4-6', 'bedrock:us.anthropic.claude-sonnet-5', 'bedrock:us.meta.llama3-1-70b-instruct-v1:0', 'bedrock:us.meta.llama3-1-8b-instruct-v1:0', 'bedrock:us.meta.llama3-2-11b-instruct-v1:0', 'bedrock:us.meta.llama3-2-1b-instruct-v1:0', 'bedrock:us.meta.llama3-2-3b-instruct-v1:0', 'bedrock:us.meta.llama3-2-90b-instruct-v1:0', 'bedrock:us.meta.llama3-3-70b-instruct-v1:0', 'bedrock:us.meta.llama4-maverick-17b-instruct-v1:0', 'bedrock:us.meta.llama4-scout-17b-instruct-v1:0', 'bedrock:us.mistral.pixtral-large-2502-v1:0', 'bedrock:us.writer.palmyra-x4-v1:0', 'bedrock:us.writer.palmyra-x5-v1:0', 'bedrock:zai.glm-4.7', 'bedrock:zai.glm-4.7-flash', 'bedrock:zai.glm-5', 'cerebras:gemma-4-31b', 'cerebras:gpt-oss-120b', 'cerebras:zai-glm-4.7', 'cohere:c4ai-aya-expanse-32b', 'cohere:c4ai-aya-expanse-8b', 'cohere:command-nightly', 'cohere:command-r-08-2024', 'cohere:command-r-plus-08-2024', 'cohere:command-r7b-12-2024', 'crusoe:Qwen/Qwen3-235B-A22B-Instruct-2507', 'crusoe:deepseek-ai/DeepSeek-V3-0324', 'crusoe:deepseek-ai/DeepSeek-V4-Pro', 'crusoe:deepseek-ai/Deepseek-V4-Flash', 'crusoe:google/gemma-4-31b-it', 'crusoe:meta-llama/Llama-3.3-70B-Instruct', 'crusoe:moonshotai/Kimi-K2.6', 'crusoe:nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B', 'crusoe:nvidia/NVIDIA-Nemotron-3-Super-120B-A12B', 'crusoe:nvidia/Nemotron-3-Nano-Omni-Reasoning-30B-A3B', 'crusoe:nvidia/Nemotron-3.5-Lightning-30B-A3B', 'crusoe:openai/gpt-oss-120b', 'crusoe:yutori/n1.5', 'crusoe:zai/GLM-5.1', 'crusoe:zai/GLM-5.2', 'deepseek:deepseek-chat', 'deepseek:deepseek-reasoner', 'deepseek:deepseek-v4-flash', 'deepseek:deepseek-v4-pro', 'gateway/anthropic:claude-fable-5', 'gateway/anthropic:claude-haiku-4-5', 'gateway/anthropic:claude-haiku-4-5-20251001', 'gateway/anthropic:claude-opus-4-1', 'gateway/anthropic:claude-opus-4-1-20250805', 'gateway/anthropic:claude-opus-4-5', 'gateway/anthropic:claude-opus-4-5-20251101', 'gateway/anthropic:claude-opus-4-6', 'gateway/anthropic:claude-opus-4-7', 'gateway/anthropic:claude-opus-4-8', 'gateway/anthropic:claude-opus-5', 'gateway/anthropic:claude-sonnet-4-5', 'gateway/anthropic:claude-sonnet-4-5-20250929', 'gateway/anthropic:claude-sonnet-4-6', 'gateway/anthropic:claude-sonnet-5', 'gateway/bedrock:anthropic.claude-3-haiku-20240307-v1:0', 'gateway/bedrock:deepseek.r1-v1:0', 'gateway/bedrock:deepseek.v3.2', 'gateway/bedrock:eu.anthropic.claude-haiku-4-5-20251001-v1:0', 'gateway/bedrock:eu.anthropic.claude-sonnet-4-20250514-v1:0', 'gateway/bedrock:eu.anthropic.claude-sonnet-4-5-20250929-v1:0', 'gateway/bedrock:eu.anthropic.claude-sonnet-4-6', 'gateway/bedrock:global.amazon.nova-2-lite-v1:0', 'gateway/bedrock:global.anthropic.claude-fable-5', 'gateway/bedrock:global.anthropic.claude-opus-4-5-20251101-v1:0', 'gateway/bedrock:global.anthropic.claude-opus-4-6-v1', 'gateway/bedrock:global.anthropic.claude-opus-4-7', 'gateway/bedrock:global.anthropic.claude-opus-4-8', 'gateway/bedrock:global.anthropic.claude-opus-5', 'gateway/bedrock:global.anthropic.claude-sonnet-5', 'gateway/bedrock:google.gemma-3-12b-it', 'gateway/bedrock:google.gemma-3-27b-it', 'gateway/bedrock:google.gemma-3-4b-it', 'gateway/bedrock:minimax.minimax-m2', 'gateway/bedrock:minimax.minimax-m2.1', 'gateway/bedrock:minimax.minimax-m2.5', 'gateway/bedrock:mistral.devstral-2-123b', 'gateway/bedrock:mistral.magistral-small-2509', 'gateway/bedrock:mistral.ministral-3-14b-instruct', 'gateway/bedrock:mistral.ministral-3-3b-instruct', 'gateway/bedrock:mistral.ministral-3-8b-instruct', 'gateway/bedrock:mistral.mistral-large-3-675b-instruct', 'gateway/bedrock:mistral.mistral-small-2402-v1:0', 'gateway/bedrock:mistral.pixtral-large-2502-v1:0', 'gateway/bedrock:moonshot.kimi-k2-thinking', 'gateway/bedrock:moonshotai.kimi-k2.5', 'gateway/bedrock:nvidia.nemotron-nano-12b-v2', 'gateway/bedrock:nvidia.nemotron-nano-3-30b', 'gateway/bedrock:nvidia.nemotron-nano-9b-v2', 'gateway/bedrock:nvidia.nemotron-super-3-120b', 'gateway/bedrock:qwen.qwen3-32b-v1:0', 'gateway/bedrock:qwen.qwen3-coder-30b-a3b-v1:0', 'gateway/bedrock:qwen.qwen3-coder-next', 'gateway/bedrock:qwen.qwen3-next-80b-a3b', 'gateway/bedrock:qwen.qwen3-vl-235b-a22b', 'gateway/bedrock:us.amazon.nova-premier-v1:0', 'gateway/bedrock:us.anthropic.claude-fable-5', 'gateway/bedrock:us.anthropic.claude-opus-4-1-20250805-v1:0', 'gateway/bedrock:us.anthropic.claude-opus-4-5-20251101-v1:0', 'gateway/bedrock:us.anthropic.claude-opus-4-6-v1', 'gateway/bedrock:us.anthropic.claude-opus-4-7', 'gateway/bedrock:us.anthropic.claude-opus-4-8', 'gateway/bedrock:us.anthropic.claude-opus-5', 'gateway/bedrock:us.anthropic.claude-sonnet-5', 'gateway/bedrock:us.meta.llama4-maverick-17b-instruct-v1:0', 'gateway/bedrock:us.meta.llama4-scout-17b-instruct-v1:0', 'gateway/bedrock:us.mistral.pixtral-large-2502-v1:0', 'gateway/bedrock:us.writer.palmyra-x4-v1:0', 'gateway/bedrock:us.writer.palmyra-x5-v1:0', 'gateway/bedrock:zai.glm-4.7', 'gateway/bedrock:zai.glm-4.7-flash', 'gateway/bedrock:zai.glm-5', 'gateway/google-cloud:gemini-2.5-flash', 'gateway/google-cloud:gemini-2.5-flash-image', 'gateway/google-cloud:gemini-2.5-flash-lite', 'gateway/google-cloud:gemini-2.5-pro', 'gateway/google-cloud:gemini-3-flash-preview', 'gateway/google-cloud:gemini-3-pro-image', 'gateway/google-cloud:gemini-3.1-flash-image', 'gateway/google-cloud:gemini-3.1-flash-lite', 'gateway/google-cloud:gemini-3.1-pro-preview', 'gateway/google-cloud:gemini-3.5-flash', 'gateway/google-cloud:gemini-3.5-flash-lite', 'gateway/google-cloud:gemini-3.6-flash', 'gateway/google-cloud:gemini-3.7-flash', 'gateway/google:gemini-2.5-flash', 'gateway/google:gemini-2.5-flash-image', 'gateway/google:gemini-2.5-flash-lite', 'gateway/google:gemini-2.5-pro', 'gateway/google:gemini-3-flash-preview', 'gateway/google:gemini-3-pro-image', 'gateway/google:gemini-3.1-flash-image', 'gateway/google:gemini-3.1-flash-lite', 'gateway/google:gemini-3.1-pro-preview', 'gateway/google:gemini-3.5-flash', 'gateway/google:gemini-3.5-flash-lite', 'gateway/google:gemini-3.6-flash', 'gateway/google:gemini-3.7-flash', 'gateway/groq:llama-3.1-8b-instant', 'gateway/groq:llama-3.3-70b-versatile', 'gateway/groq:openai/gpt-oss-120b', 'gateway/groq:openai/gpt-oss-20b', 'gateway/groq:openai/gpt-oss-safeguard-20b', 'gateway/openai:gpt-3.5-turbo', 'gateway/openai:gpt-3.5-turbo-0125', 'gateway/openai:gpt-3.5-turbo-1106', 'gateway/openai:gpt-4', 'gateway/openai:gpt-4-0613', 'gateway/openai:gpt-4-turbo', 'gateway/openai:gpt-4-turbo-2024-04-09', 'gateway/openai:gpt-4.1', 'gateway/openai:gpt-4.1-2025-04-14', 'gateway/openai:gpt-4.1-mini', 'gateway/openai:gpt-4.1-mini-2025-04-14', 'gateway/openai:gpt-4.1-nano', 'gateway/openai:gpt-4.1-nano-2025-04-14', 'gateway/openai:gpt-4o', 'gateway/openai:gpt-4o-2024-05-13', 'gateway/openai:gpt-4o-2024-08-06', 'gateway/openai:gpt-4o-2024-11-20', 'gateway/openai:gpt-4o-mini', 'gateway/openai:gpt-4o-mini-2024-07-18', 'gateway/openai:gpt-5', 'gateway/openai:gpt-5-2025-08-07', 'gateway/openai:gpt-5-mini', 'gateway/openai:gpt-5-mini-2025-08-07', 'gateway/openai:gpt-5-nano', 'gateway/openai:gpt-5-nano-2025-08-07', 'gateway/openai:gpt-5-pro', 'gateway/openai:gpt-5-pro-2025-10-06', 'gateway/openai:gpt-5.1', 'gateway/openai:gpt-5.1-2025-11-13', 'gateway/openai:gpt-5.2', 'gateway/openai:gpt-5.2-2025-12-11', 'gateway/openai:gpt-5.2-chat-latest', 'gateway/openai:gpt-5.2-pro', 'gateway/openai:gpt-5.2-pro-2025-12-11', 'gateway/openai:gpt-5.3-chat-latest', 'gateway/openai:gpt-5.4', 'gateway/openai:gpt-5.4-mini', 'gateway/openai:gpt-5.4-mini-2026-03-17', 'gateway/openai:gpt-5.4-nano', 'gateway/openai:gpt-5.4-nano-2026-03-17', 'gateway/openai:gpt-5.6-luna', 'gateway/openai:gpt-5.6-sol', 'gateway/openai:gpt-5.6-terra', 'gateway/openai:o1', 'gateway/openai:o1-2024-12-17', 'gateway/openai:o1-pro', 'gateway/openai:o1-pro-2025-03-19', 'gateway/openai:o3', 'gateway/openai:o3-2025-04-16', 'gateway/openai:o3-mini', 'gateway/openai:o3-mini-2025-01-31', 'gateway/openai:o3-pro', 'gateway/openai:o3-pro-2025-06-10', 'gateway/openai:o4-mini', 'gateway/openai:o4-mini-2025-04-16', 'google-cloud:gemini-2.0-flash', 'google-cloud:gemini-2.0-flash-lite', 'google-cloud:gemini-2.5-flash', 'google-cloud:gemini-2.5-flash-image', 'google-cloud:gemini-2.5-flash-lite', 'google-cloud:gemini-2.5-flash-preview-09-2025', 'google-cloud:gemini-2.5-pro', 'google-cloud:gemini-3-flash-preview', 'google-cloud:gemini-3-pro-image', 'google-cloud:gemini-3-pro-image-preview', 'google-cloud:gemini-3-pro-preview', 'google-cloud:gemini-3.1-flash-image', 'google-cloud:gemini-3.1-flash-image-preview', 'google-cloud:gemini-3.1-flash-lite', 'google-cloud:gemini-3.1-pro-preview', 'google-cloud:gemini-3.5-flash', 'google-cloud:gemini-3.5-flash-lite', 'google-cloud:gemini-3.6-flash', 'google-cloud:gemini-3.7-flash', 'google-cloud:gemini-flash-latest', 'google-cloud:gemini-flash-lite-latest', 'google:gemini-2.0-flash', 'google:gemini-2.0-flash-lite', 'google:gemini-2.5-flash', 'google:gemini-2.5-flash-image', 'google:gemini-2.5-flash-lite', 'google:gemini-2.5-flash-preview-09-2025', 'google:gemini-2.5-pro', 'google:gemini-3-flash-preview', 'google:gemini-3-pro-image', 'google:gemini-3-pro-image-preview', 'google:gemini-3-pro-preview', 'google:gemini-3.1-flash-image', 'google:gemini-3.1-flash-image-preview', 'google:gemini-3.1-flash-lite', 'google:gemini-3.1-pro-preview', 'google:gemini-3.5-flash', 'google:gemini-3.5-flash-lite', 'google:gemini-3.6-flash', 'google:gemini-3.7-flash', 'google:gemini-flash-latest', 'google:gemini-flash-lite-latest', 'groq:llama-3.1-8b-instant', 'groq:llama-3.3-70b-versatile', 'groq:meta-llama/llama-4-maverick-17b-128e-instruct', 'groq:meta-llama/llama-guard-4-12b', 'groq:meta-llama/llama-prompt-guard-2-22m', 'groq:meta-llama/llama-prompt-guard-2-86m', 'groq:openai/gpt-oss-120b', 'groq:openai/gpt-oss-20b', 'groq:openai/gpt-oss-safeguard-20b', 'groq:playai-tts', 'groq:playai-tts-arabic', 'groq:whisper-large-v3', 'groq:whisper-large-v3-turbo', 'heroku:claude-3-5-haiku', 'heroku:claude-3-5-sonnet-latest', 'heroku:claude-3-7-sonnet', 'heroku:claude-3-haiku', 'heroku:claude-4-5-haiku', 'heroku:claude-4-5-sonnet', 'heroku:claude-4-6-sonnet', 'heroku:claude-4-sonnet', 'heroku:claude-opus-4-5', 'heroku:claude-opus-4-6', 'heroku:deepseek-v3-2', 'heroku:glm-4-7', 'heroku:glm-4-7-flash', 'heroku:gpt-oss-120b', 'heroku:kimi-k2-5', 'heroku:kimi-k2-thinking', 'heroku:minimax-m2', 'heroku:minimax-m2-1', 'heroku:nova-2-lite', 'heroku:nova-lite', 'heroku:nova-pro', 'heroku:qwen3-235b', 'heroku:qwen3-coder-480b', 'huggingface:Qwen/QwQ-32B', 'huggingface:Qwen/Qwen2.5-72B-Instruct', 'huggingface:Qwen/Qwen3-235B-A22B', 'huggingface:Qwen/Qwen3-32B', 'huggingface:deepseek-ai/DeepSeek-R1', 'huggingface:meta-llama/Llama-3.3-70B-Instruct', 'huggingface:meta-llama/Llama-4-Maverick-17B-128E-Instruct', 'huggingface:meta-llama/Llama-4-Scout-17B-16E-Instruct', 'mistral:codestral-latest', 'mistral:mistral-large-latest', 'mistral:mistral-moderation-latest', 'mistral:mistral-small-latest', 'moonshotai:kimi-k2-0711-preview', 'moonshotai:kimi-k2.5', 'moonshotai:kimi-k2.6', 'moonshotai:kimi-k2.7-code', 'moonshotai:kimi-k2.7-code-highspeed', 'moonshotai:kimi-k3', 'moonshotai:kimi-latest', 'moonshotai:kimi-thinking-preview', 'moonshotai:moonshot-v1-128k', 'moonshotai:moonshot-v1-128k-vision-preview', 'moonshotai:moonshot-v1-32k', 'moonshotai:moonshot-v1-32k-vision-preview', 'moonshotai:moonshot-v1-8k', 'moonshotai:moonshot-v1-8k-vision-preview', 'moonshotai:moonshot-v1-auto', 'openai-chat:computer-use-preview', 'openai-chat:computer-use-preview-2025-03-11', 'openai-chat:gpt-3.5-turbo', 'openai-chat:gpt-3.5-turbo-0125', 'openai-chat:gpt-3.5-turbo-0301', 'openai-chat:gpt-3.5-turbo-1106', 'openai-chat:gpt-3.5-turbo-16k', 'openai-chat:gpt-4', 'openai-chat:gpt-4-0314', 'openai-chat:gpt-4-0613', 'openai-chat:gpt-4-turbo', 'openai-chat:gpt-4-turbo-2024-04-09', 'openai-chat:gpt-4.1', 'openai-chat:gpt-4.1-2025-04-14', 'openai-chat:gpt-4.1-mini', 'openai-chat:gpt-4.1-mini-2025-04-14', 'openai-chat:gpt-4.1-nano', 'openai-chat:gpt-4.1-nano-2025-04-14', 'openai-chat:gpt-4o', 'openai-chat:gpt-4o-2024-05-13', 'openai-chat:gpt-4o-2024-08-06', 'openai-chat:gpt-4o-2024-11-20', 'openai-chat:gpt-4o-audio-preview', 'openai-chat:gpt-4o-audio-preview-2024-12-17', 'openai-chat:gpt-4o-audio-preview-2025-06-03', 'openai-chat:gpt-4o-mini', 'openai-chat:gpt-4o-mini-2024-07-18', 'openai-chat:gpt-4o-mini-audio-preview', 'openai-chat:gpt-4o-mini-audio-preview-2024-12-17', 'openai-chat:gpt-4o-mini-search-preview', 'openai-chat:gpt-4o-mini-search-preview-2025-03-11', 'openai-chat:gpt-4o-search-preview', 'openai-chat:gpt-4o-search-preview-2025-03-11', 'openai-chat:gpt-5', 'openai-chat:gpt-5-2025-08-07', 'openai-chat:gpt-5-chat-latest', 'openai-chat:gpt-5-codex', 'openai-chat:gpt-5-mini', 'openai-chat:gpt-5-mini-2025-08-07', 'openai-chat:gpt-5-nano', 'openai-chat:gpt-5-nano-2025-08-07', 'openai-chat:gpt-5-pro', 'openai-chat:gpt-5-pro-2025-10-06', 'openai-chat:gpt-5.1', 'openai-chat:gpt-5.1-2025-11-13', 'openai-chat:gpt-5.1-chat-latest', 'openai-chat:gpt-5.1-codex', 'openai-chat:gpt-5.1-codex-max', 'openai-chat:gpt-5.2', 'openai-chat:gpt-5.2-2025-12-11', 'openai-chat:gpt-5.2-chat-latest', 'openai-chat:gpt-5.2-pro', 'openai-chat:gpt-5.2-pro-2025-12-11', 'openai-chat:gpt-5.3-chat-latest', 'openai-chat:gpt-5.4', 'openai-chat:gpt-5.4-mini', 'openai-chat:gpt-5.4-mini-2026-03-17', 'openai-chat:gpt-5.4-nano', 'openai-chat:gpt-5.4-nano-2026-03-17', 'openai-chat:gpt-5.6-luna', 'openai-chat:gpt-5.6-sol', 'openai-chat:gpt-5.6-terra', 'openai-chat:o1', 'openai-chat:o1-2024-12-17', 'openai-chat:o1-pro', 'openai-chat:o1-pro-2025-03-19', 'openai-chat:o3', 'openai-chat:o3-2025-04-16', 'openai-chat:o3-deep-research', 'openai-chat:o3-deep-research-2025-06-26', 'openai-chat:o3-mini', 'openai-chat:o3-mini-2025-01-31', 'openai-chat:o3-pro', 'openai-chat:o3-pro-2025-06-10', 'openai-chat:o4-mini', 'openai-chat:o4-mini-2025-04-16', 'openai-chat:o4-mini-deep-research', 'openai-chat:o4-mini-deep-research-2025-06-26', 'openai:computer-use-preview', 'openai:computer-use-preview-2025-03-11', 'openai:gpt-3.5-turbo', 'openai:gpt-3.5-turbo-0125', 'openai:gpt-3.5-turbo-0301', 'openai:gpt-3.5-turbo-1106', 'openai:gpt-4', 'openai:gpt-4-0314', 'openai:gpt-4-0613', 'openai:gpt-4-turbo', 'openai:gpt-4-turbo-2024-04-09', 'openai:gpt-4.1', 'openai:gpt-4.1-2025-04-14', 'openai:gpt-4.1-mini', 'openai:gpt-4.1-mini-2025-04-14', 'openai:gpt-4.1-nano', 'openai:gpt-4.1-nano-2025-04-14', 'openai:gpt-4o', 'openai:gpt-4o-2024-05-13', 'openai:gpt-4o-2024-08-06', 'openai:gpt-4o-2024-11-20', 'openai:gpt-4o-audio-preview', 'openai:gpt-4o-audio-preview-2024-12-17', 'openai:gpt-4o-audio-preview-2025-06-03', 'openai:gpt-4o-mini', 'openai:gpt-4o-mini-2024-07-18', 'openai:gpt-4o-mini-audio-preview', 'openai:gpt-4o-mini-audio-preview-2024-12-17', 'openai:gpt-5', 'openai:gpt-5-2025-08-07', 'openai:gpt-5-chat-latest', 'openai:gpt-5-codex', 'openai:gpt-5-mini', 'openai:gpt-5-mini-2025-08-07', 'openai:gpt-5-nano', 'openai:gpt-5-nano-2025-08-07', 'openai:gpt-5-pro', 'openai:gpt-5-pro-2025-10-06', 'openai:gpt-5.1', 'openai:gpt-5.1-2025-11-13', 'openai:gpt-5.1-chat-latest', 'openai:gpt-5.1-codex', 'openai:gpt-5.1-codex-max', 'openai:gpt-5.2', 'openai:gpt-5.2-2025-12-11', 'openai:gpt-5.2-chat-latest', 'openai:gpt-5.2-pro', 'openai:gpt-5.2-pro-2025-12-11', 'openai:gpt-5.3-chat-latest', 'openai:gpt-5.4', 'openai:gpt-5.4-mini', 'openai:gpt-5.4-mini-2026-03-17', 'openai:gpt-5.4-nano', 'openai:gpt-5.4-nano-2026-03-17', 'openai:gpt-5.6-luna', 'openai:gpt-5.6-sol', 'openai:gpt-5.6-terra', 'openai:o1', 'openai:o1-2024-12-17', 'openai:o1-pro', 'openai:o1-pro-2025-03-19', 'openai:o3', 'openai:o3-2025-04-16', 'openai:o3-deep-research', 'openai:o3-deep-research-2025-06-26', 'openai:o3-mini', 'openai:o3-mini-2025-01-31', 'openai:o3-pro', 'openai:o3-pro-2025-06-10', 'openai:o4-mini', 'openai:o4-mini-2025-04-16', 'openai:o4-mini-deep-research', 'openai:o4-mini-deep-research-2025-06-26', 'test', 'snowflake:claude-4-sonnet', 'snowflake:claude-fable-5', 'snowflake:claude-haiku-4-5', 'snowflake:claude-opus-4-5', 'snowflake:claude-opus-4-6', 'snowflake:claude-opus-4-7', 'snowflake:claude-opus-4-8', 'snowflake:claude-opus-5', 'snowflake:claude-sonnet-4-5', 'snowflake:claude-sonnet-4-6', 'snowflake:claude-sonnet-5', 'snowflake:deepseek-r1', 'snowflake:llama3.1-405b', 'snowflake:llama3.1-70b', 'snowflake:llama3.1-8b', 'snowflake:llama4-maverick', 'snowflake:mistral-7b', 'snowflake:mistral-large', 'snowflake:mistral-large2', 'snowflake:openai-gpt-4.1', 'snowflake:openai-gpt-5', 'snowflake:openai-gpt-5-6-luna', 'snowflake:openai-gpt-5-6-sol', 'snowflake:openai-gpt-5-6-terra', 'snowflake:openai-gpt-5-chat', 'snowflake:openai-gpt-5-mini', 'snowflake:openai-gpt-5-nano', 'snowflake:openai-gpt-5.1', 'snowflake:openai-gpt-5.2', 'snowflake:openai-gpt-5.4', 'snowflake:openai-gpt-5.5', 'snowflake:snowflake-llama-3.3-70b', 'xai:grok-3', 'xai:grok-3-fast', 'xai:grok-3-fast-latest', 'xai:grok-3-latest', 'xai:grok-3-mini', 'xai:grok-3-mini-fast', 'xai:grok-3-mini-fast-latest', 'xai:grok-4', 'xai:grok-4-0709', 'xai:grok-4-1-fast', 'xai:grok-4-1-fast-non-reasoning', 'xai:grok-4-1-fast-non-reasoning-latest', 'xai:grok-4-1-fast-reasoning', 'xai:grok-4-1-fast-reasoning-latest', 'xai:grok-4-fast', 'xai:grok-4-fast-non-reasoning', 'xai:grok-4-fast-non-reasoning-latest', 'xai:grok-4-fast-reasoning', 'xai:grok-4-fast-reasoning-latest', 'xai:grok-4-latest', 'xai:grok-4.20', 'xai:grok-4.20-0309', 'xai:grok-4.20-0309-non-reasoning', 'xai:grok-4.20-0309-reasoning', 'xai:grok-4.20-multi-agent', 'xai:grok-4.20-multi-agent-0309', 'xai:grok-4.20-multi-agent-latest', 'xai:grok-4.20-non-reasoning', 'xai:grok-4.20-non-reasoning-latest', 'xai:grok-4.20-reasoning-latest', 'xai:grok-4.3', 'xai:grok-4.3-latest', 'xai:grok-4.5', 'xai:grok-4.5-latest', 'xai:grok-code-fast-1', 'zai:autoglm-phone-multilingual', 'zai:glm-4-32b-0414-128k', 'zai:glm-4.5', 'zai:glm-4.5-air', 'zai:glm-4.5-airx', 'zai:glm-4.5-flash', 'zai:glm-4.5-x', 'zai:glm-4.5v', 'zai:glm-4.6', 'zai:glm-4.6v', 'zai:glm-4.6v-flash', 'zai:glm-4.6v-flashx', 'zai:glm-4.7', 'zai:glm-4.7-flash', 'zai:glm-4.7-flashx', 'zai:glm-5', 'zai:glm-5-turbo', 'zai:glm-5.1', 'zai:glm-5.2', 'zai:glm-5v-turbo'])` ToolVisibility -------------- [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ToolVisibility) How a function tool is represented on the request a provider actually receives. * `'visible'`: an ordinary entry in the provider’s `tools` collection, schema included. * `'deferred'`: a declared `tools` entry whose schema is withheld behind the provider’s schema-deferral flag until something reveals it. * `'withheld'`: absent from the request entirely. * `'via_history'`: absent from the `tools` collection; the full definition travels on the provider’s mid-conversation tool-addition channel instead. Resolved per tool name into [`ModelRequestParameters.tool_visibility`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters.tool_visibility) . **Default:** `Literal['visible', 'deferred', 'withheld', 'via_history']` ALLOW\_MODEL\_REQUESTS ---------------------- [](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ALLOW_MODEL_REQUESTS) Whether to allow requests to models. This global setting allows you to disable request to most models, e.g. to make sure you don’t accidentally make costly requests to a model during tests. The testing models [`TestModel`](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestModel) , [`FunctionModel`](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel) and [`TestEmbeddingModel`](https://pydantic.dev/docs/ai/api/pydantic-ai/embeddings/#pydantic_ai.embeddings.TestEmbeddingModel) are not affected by this setting, nor is [`SentenceTransformerEmbeddingModel`](https://pydantic.dev/docs/ai/api/pydantic-ai/embeddings/#pydantic_ai.embeddings.sentence_transformers.SentenceTransformerEmbeddingModel) , which runs inference locally and so has no per-call provider cost. **Default:** `True` Was this page helpful? Thanks for your feedback! --- # huggingface | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/huggingface/#_top) huggingface =========== Setup ----- [](https://pydantic.dev/docs/ai/api/models/huggingface/#setup) For details on how to set up authentication with this model, see [model configuration for Hugging Face](https://pydantic.dev/docs/ai/models/huggingface/) . HuggingFaceModel ---------------- [](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceModel) **Bases:** `Model[AsyncInferenceClient]` A model that uses Hugging Face Inference Providers. Internally, this uses the [HF Python client](https://github.com/huggingface/huggingface_hub) to interact with the API. Apart from `__init__`, all methods are private or match those of the base class. ### Attributes [](https://pydantic.dev/docs/ai/api/models/huggingface/#attributes) #### base\_url [](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceModel.base_url) The base URL of the provider. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### model\_name [](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceModel.model_name) The model name. **Type:** `HuggingFaceModelName` #### system [](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceModel.system) The system / model provider. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) ### Methods [](https://pydantic.dev/docs/ai/api/models/huggingface/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceModel.__init__) def __init__( model_name: str, *, provider: Literal['huggingface'] | Provider[AsyncInferenceClient] = 'huggingface', profile: ModelProfileSpec | None = None, settings: ModelSettings | None = None, ) Initialize a Hugging Face model. ##### Parameters [](https://pydantic.dev/docs/ai/api/models/huggingface/#parameters) **`model_name`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceModel.__init__(model_name)) The name of the Model to use. You can browse available models [here](https://huggingface.co/models?pipeline_tag=text-generation&inference_provider=all&sort=trending) . **`provider`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘huggingface’\] | `Provider`\[`AsyncInferenceClient`\] _Default:_ `'huggingface'` [](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceModel.__init__(provider)) The provider to use for Hugging Face Inference Providers. Can be either the string ‘huggingface’ or an instance of `Provider[AsyncInferenceClient]`. If not provided, the other parameters will be used. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceModel.__init__(profile)) The model profile to use. Defaults to a profile picked by the provider based on the model name. **`settings`** : [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceModel.__init__(settings)) Model-specific settings that will be used as defaults for this model. HuggingFaceModelSettings ------------------------ [](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceModelSettings) **Bases:** [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) Settings used for a Hugging Face model request. HuggingFaceStreamedResponse --------------------------- [](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceStreamedResponse) **Bases:** `StreamedResponse` Implementation of `StreamedResponse` for Hugging Face models. ### Attributes [](https://pydantic.dev/docs/ai/api/models/huggingface/#attributes-1) #### model\_name [](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceStreamedResponse.model_name) Get the model name of the response. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### provider\_name [](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceStreamedResponse.provider_name) Get the provider name. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### provider\_url [](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceStreamedResponse.provider_url) Get the provider base URL. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### timestamp [](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceStreamedResponse.timestamp) Get the timestamp of the response. **Type:** [`datetime`](https://docs.python.org/3/library/datetime.html#module-datetime) HuggingFaceModelName -------------------- [](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceModelName) Possible Hugging Face model names. You can browse available models [here](https://huggingface.co/models?pipeline_tag=text-generation&inference_provider=all&sort=trending) . **Default:** `str | LatestHuggingFaceModelNames` LatestHuggingFaceModelNames --------------------------- [](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.LatestHuggingFaceModelNames) Latest Hugging Face models. **Default:** `Literal['deepseek-ai/DeepSeek-R1', 'meta-llama/Llama-3.3-70B-Instruct', 'meta-llama/Llama-4-Maverick-17B-128E-Instruct', 'meta-llama/Llama-4-Scout-17B-16E-Instruct', 'Qwen/QwQ-32B', 'Qwen/Qwen2.5-72B-Instruct', 'Qwen/Qwen3-235B-A22B', 'Qwen/Qwen3-32B']` Was this page helpful? Thanks for your feedback! --- # instrumented | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/instrumented/#_top) instrumented ============ InstrumentationSettings ----------------------- [](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentationSettings) Options for instrumenting models and agents with OpenTelemetry. Used in: * [`Instrumentation`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Instrumentation) capability * [`Agent.instrument`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.instrument) / [`Agent.instrument_all()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.instrument_all) * [`InstrumentedModel`](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentedModel) See the [Debugging and Monitoring guide](https://ai.pydantic.dev/logfire/) for more info. ### Methods [](https://pydantic.dev/docs/ai/api/models/instrumented/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentationSettings.__init__) def __init__( *, tracer_provider: TracerProvider | None = None, meter_provider: MeterProvider | None = None, include_binary_content: bool = True, include_content: bool = True, include_model_request_parameters: bool = True, version: Literal[2, 3, 4, 5] = DEFAULT_INSTRUMENTATION_VERSION, use_aggregated_usage_attribute_names: bool = True, ) Create instrumentation options. ##### Parameters [](https://pydantic.dev/docs/ai/api/models/instrumented/#parameters) **`tracer_provider`** : `TracerProvider` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentationSettings.__init__(tracer_provider)) The OpenTelemetry tracer provider to use. If not provided, the global tracer provider is used. Calling `logfire.configure()` sets the global tracer provider, so most users don’t need this. **`meter_provider`** : `MeterProvider` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentationSettings.__init__(meter_provider)) The OpenTelemetry meter provider to use. If not provided, the global meter provider is used. Calling `logfire.configure()` sets the global meter provider, so most users don’t need this. **`include_binary_content`** : [`bool`](https://docs.python.org/3/library/functions.html#bool) _Default:_ `True` [](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentationSettings.__init__(include_binary_content)) Whether to include binary file data in the instrumentation events: user prompts and model responses, tool returns, the agent’s output and the arguments its output function receives, and run and tool deferral metadata. The media type is recorded either way. Binary content is found inside dictionaries, lists and `ToolReturn`s, but not inside your own types: a `BinaryContent` held as a field of a model or dataclass you define is still recorded in full. **`include_content`** : [`bool`](https://docs.python.org/3/library/functions.html#bool) _Default:_ `True` [](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentationSettings.__init__(include_content)) Whether to include prompts, completions, and tool call arguments and responses in the instrumentation events. **`include_model_request_parameters`** : [`bool`](https://docs.python.org/3/library/functions.html#bool) _Default:_ `True` [](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentationSettings.__init__(include_model_request_parameters)) Whether to emit the `model_request_parameters` span attribute on model request spans. This serializes the full `ModelRequestParameters` (output configuration and every tool definition, including fields that are not sent to the model such as tool `metadata` and, when not requested, `return_schema`). Defaults to `True`. Set to `False` to omit it entirely, which is useful when large tool output schemas make the attribute big enough to strain span export. The OpenTelemetry `gen_ai.tool.definitions` attribute (tool name, description, and parameters) is always emitted regardless of this setting. **`version`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[2, 3, 4, 5\] _Default:_ `DEFAULT_INSTRUMENTATION_VERSION` [](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentationSettings.__init__(version)) Version of the data format. This is unrelated to the Pydantic AI package version. Defaults to version 5. Versions 2, 3, and 4 are deprecated compatibility formats and emit a `PydanticAIDeprecationWarning` when used. Version 2 uses the newer OpenTelemetry GenAI spec and stores messages in the following attributes: * `gen_ai.system_instructions` for instructions passed to the agent. * `gen_ai.input.messages` and `gen_ai.output.messages` on model request spans. * `pydantic_ai.all_messages` on agent run spans. Version 3 is the same as version 2, with additional support for thinking tokens. Version 4 is the same as version 3, with GenAI semantic conventions for multimodal content: URL-based media uses type=‘uri’ with uri and mime\_type fields (and modality for image/audio/video). Inline binary content uses type=‘blob’ with mime\_type and content fields (and modality for image/audio/video). [https://opentelemetry.io/docs/specs/semconv/gen-ai/non-normative/examples-llm-calls/#multimodal-inputs-example](https://opentelemetry.io/docs/specs/semconv/gen-ai/non-normative/examples-llm-calls/#multimodal-inputs-example) Version 5 is the same as version 4, but CallDeferred and ApprovalRequired exceptions no longer record an exception event or set the span status to ERROR — the span is left as UNSET, since deferrals are control flow, not errors. **`use_aggregated_usage_attribute_names`** : [`bool`](https://docs.python.org/3/library/functions.html#bool) _Default:_ `True` [](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentationSettings.__init__(use_aggregated_usage_attribute_names)) Whether to use `gen_ai.aggregated_usage.*` attribute names for token usage on agent run spans instead of the standard `gen_ai.usage.*` names. Defaults to True to prevent double-counting in observability backends that aggregate span attributes across parent and child spans. Note: `gen_ai.aggregated_usage.*` is a custom namespace, not part of the OpenTelemetry Semantic Conventions. It may be updated if OTel introduces an official convention. #### aggregated\_usage\_attributes [](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentationSettings.aggregated_usage_attributes) def aggregated_usage_attributes(usage: UsageBase) -> dict[str, int] Cumulative-usage OpenTelemetry attributes for a run/session span. Remaps `gen_ai.usage.*` to `gen_ai.aggregated_usage.*` when `use_aggregated_usage_attribute_names` is set, so a backend that sums span attributes doesn’t double-count the run’s cumulative usage against the per-request `chat` spans’ `gen_ai.usage.*`. Shared by the classic agent-run span (the `Instrumentation` capability) and the realtime session span so the two can’t drift. ##### Returns [](https://pydantic.dev/docs/ai/api/models/instrumented/#returns) [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`int`](https://docs.python.org/3/library/functions.html#int)\ \] InstrumentedModel ----------------- [](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentedModel) **Bases:** `WrapperModel` Model which wraps another model so that requests are instrumented with OpenTelemetry. See the [Debugging and Monitoring guide](https://ai.pydantic.dev/logfire/) for more info. ### Attributes [](https://pydantic.dev/docs/ai/api/models/instrumented/#attributes) #### instrumentation\_settings [](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentedModel.instrumentation_settings) Instrumentation settings for this model. **Type:** [`InstrumentationSettings`](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentationSettings) **Default:** `options or InstrumentationSettings()` instrument\_model ----------------- [](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.instrument_model) def instrument_model(model: Model, instrument: InstrumentationSettings | bool) -> Model Wrap `model` in an `InstrumentedModel` so OTel/Logfire spans are emitted around requests. ### Returns [](https://pydantic.dev/docs/ai/api/models/instrumented/#returns-1) `Model` Was this page helpful? Thanks for your feedback! --- # fallback | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/fallback/#_top) fallback ======== FallbackModel ------------- [](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackModel) **Bases:** `Model` A model that uses one or more fallback models upon failure. Apart from `__init__`, all methods are private or match those of the base class. ### Attributes [](https://pydantic.dev/docs/ai/api/models/fallback/#attributes) #### model\_id [](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackModel.model_id) The fully qualified model identifier, combining the wrapped models’ IDs. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### model\_name [](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackModel.model_name) The model name. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) ### Methods [](https://pydantic.dev/docs/ai/api/models/fallback/#methods) #### \_\_aenter\_\_ [](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackModel.__aenter__) `@async` def __aenter__() -> FallbackModel Enter all sub-models so their providers can manage HTTP client lifecycle. ##### Returns [](https://pydantic.dev/docs/ai/api/models/fallback/#returns) `FallbackModel` #### \_\_aexit\_\_ [](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackModel.__aexit__) `@async` def __aexit__( exc_type: type[BaseException] | None, exc_val: BaseException | None, exc_tb: TracebackType | None, ) -> bool | None Exit all sub-models, closing their providers’ HTTP clients. ##### Returns [](https://pydantic.dev/docs/ai/api/models/fallback/#returns-1) [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`None`](https://docs.python.org/3/library/constants.html#None) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackModel.__init__) def __init__( default_model: Model | KnownModelName | str, *fallback_models: Model | KnownModelName | str, fallback_on: FallbackOn = (ModelAPIError,), ) Initialize a fallback model instance. ##### Parameters [](https://pydantic.dev/docs/ai/api/models/fallback/#parameters) **`default_model`** : `Model` | `KnownModelName` | [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackModel.__init__(default_model)) The name or instance of the default model to use. **`fallback_models`** : `Model` | `KnownModelName` | [`str`](https://docs.python.org/3/library/stdtypes.html#str) _Default:_ `()` [](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackModel.__init__(fallback_models)) The names or instances of the fallback models to use upon failure. **`fallback_on`** : `FallbackOn` _Default:_ `(ModelAPIError,)` [](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackModel.__init__(fallback_on)) Conditions that trigger fallback to the next model. Accepts: * A tuple of exception types: `(ModelAPIError, RateLimitError)` * An exception handler (sync or async): `lambda exc: isinstance(exc, MyError)` * A response handler (sync or async): `def check(r: ModelResponse) -> bool` * A sequence mixing all of the above: `[ModelAPIError, exc_handler, response_handler]` Handler type is auto-detected by inspecting type hints on the first parameter. If the first parameter is hinted as `ModelResponse`, it’s a response handler. Otherwise (including untyped handlers and lambdas), it’s an exception handler. #### cancel\_suspended\_response [](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackModel.cancel_suspended_response) `@async` def cancel_suspended_response(response: ModelResponse) -> None Cancel a suspended continuation on the underlying model holding the server-side job. When the response carries a continuation pin, resolve that model and delegate to it. Resolve the pin directly from metadata rather than via `_get_continuation_model`: the cancel path is driven by `_ContinuationStreamedResponse.get()`, whose `state` is already `'interrupted'`/`'incomplete'`/`'complete'` (never `'suspended'`) by the time cancellation unwinds, so gating on `state == 'suspended'` here would never find the pin. When no pin resolves, the response can still hold a live server-side job: the pin is only stamped when a segment _ends_ suspended, so a streamed background job cancelled during its first segment (e.g. OpenAI background mode, marked by `provider_details['background']` + `provider_response_id`) has no pin yet. Best-effort delegate to every inner model so the job is torn down rather than leaked. This is safe because each model’s own cancel guard is strict (OpenAI only acts on its own `background` marker with a matching `provider_name`; others no-op), and a raising model doesn’t stop the rest. ##### Returns [](https://pydantic.dev/docs/ai/api/models/fallback/#returns-2) [`None`](https://docs.python.org/3/library/constants.html#None) #### request [](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackModel.request) `@async` def request( messages: list[ModelMessage], model_settings: ModelSettings | None, model_request_parameters: ModelRequestParameters, ) -> ModelResponse Try each model in sequence until one succeeds. In case of failure, raise a FallbackExceptionGroup with all exceptions. If a previous response set `state='suspended'`, the request is routed directly to the pinned continuation model, bypassing the fallback chain. If the pinned model raises a fallback-eligible error during continuation, the messages are rewound (stripping the suspended response and trailing continuation request) and the normal fallback chain is tried. ##### Returns [](https://pydantic.dev/docs/ai/api/models/fallback/#returns-3) [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) #### request\_stream [](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackModel.request_stream) `@async` def request_stream( messages: list[ModelMessage], model_settings: ModelSettings | None, model_request_parameters: ModelRequestParameters, run_context: RunContext[Any] | None = None, ) -> AsyncGenerator[StreamedResponse] Try each model in sequence until one succeeds. If a previous response set `state='suspended'`, the request is routed directly to the pinned continuation model, bypassing the fallback chain. If the pinned model raises a fallback-eligible error while opening the stream, the messages are rewound and the normal fallback chain is tried. Mid-stream failures still propagate. ##### Returns [](https://pydantic.dev/docs/ai/api/models/fallback/#returns-4) [`AsyncGenerator`](https://docs.python.org/3/library/typing.html#typing.AsyncGenerator) \[`StreamedResponse`\] ResponseRejected ---------------- [](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.ResponseRejected) **Bases:** [`Exception`](https://docs.python.org/3/library/exceptions.html#Exception) Raised within a `FallbackExceptionGroup` when model responses are rejected by a response handler. ExceptionHandler ---------------- [](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.ExceptionHandler) A sync or async callable that decides whether an exception should trigger fallback. **Default:** `Callable[[Exception], Awaitable[bool]] | Callable[[Exception], bool]` FallbackOn ---------- [](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackOn) The type of the `fallback_on` parameter to [`FallbackModel`](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackModel) . **Default:** `type[Exception] | tuple[type[Exception], ...] | ExceptionHandler | ResponseHandler | Sequence[type[Exception] | ExceptionHandler | ResponseHandler]` ResponseHandler --------------- [](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.ResponseHandler) A sync or async callable that decides whether a model response should trigger fallback. **Default:** `Callable[[ModelResponse], Awaitable[bool]] | Callable[[ModelResponse], bool]` Was this page helpful? Thanks for your feedback! --- # function | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/function/#_top) function ======== A model controlled by a local function. [`FunctionModel`](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel) is similar to [`TestModel`](https://pydantic.dev/docs/ai/api/models/test/) , but allows greater control over the model’s behavior. Its primary use case is for more advanced unit testing than is possible with `TestModel`. Here’s a minimal example: function\_model\_usage.py from pydantic_ai import Agent from pydantic_ai import ModelMessage, ModelResponse, TextPart from pydantic_ai.models.function import FunctionModel, AgentInfo my_agent = Agent('openai:gpt-5.2') async def model_function( messages: list[ModelMessage], info: AgentInfo ) -> ModelResponse: print(messages) """ [\ ModelRequest(\ parts=[\ UserPromptPart(\ content='Testing my agent...',\ timestamp=datetime.datetime(...),\ )\ ],\ timestamp=datetime.datetime(...),\ run_id='...',\ conversation_id='...',\ )\ ] """ print(info) """ AgentInfo( function_tools=[], allow_text_output=True, output_tools=[], model_settings=None, model_request_parameters=ModelRequestParameters( function_tools=[], native_tools=[], tool_visibility={}, output_tools=[] ), instructions=None, ) """ return ModelResponse(parts=[TextPart('hello world')]) async def test_my_agent(): """Unit test for my_agent, to be run by pytest.""" with my_agent.override(model=FunctionModel(model_function)): result = await my_agent.run('Testing my agent...') assert result.output == 'hello world' See [Unit testing with `FunctionModel`](https://pydantic.dev/docs/ai/guides/testing/#unit-testing-with-functionmodel) for detailed documentation. AgentInfo --------- [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.AgentInfo) Information about an agent. This is passed as the second to functions used within [`FunctionModel`](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel) . ### Attributes [](https://pydantic.dev/docs/ai/api/models/function/#attributes) #### allow\_text\_output [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.AgentInfo.allow_text_output) Whether a plain text output is allowed. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### function\_tools [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.AgentInfo.function_tools) The function tools available on this agent. These are the tools registered via the [`tool`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.tool) and [`tool_plain`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.tool_plain) decorators. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`ToolDefinition`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.ToolDefinition)\ \] #### instructions [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.AgentInfo.instructions) The instructions passed to model. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### model\_request\_parameters [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.AgentInfo.model_request_parameters) The model request parameters passed to the run call. **Type:** `ModelRequestParameters` #### model\_settings [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.AgentInfo.model_settings) The model settings passed to the run call. **Type:** [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) #### output\_tools [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.AgentInfo.output_tools) The tools that can called to produce the final output of the run. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`ToolDefinition`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.ToolDefinition)\ \] DeltaThinkingPart ----------------- [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.DeltaThinkingPart) Incremental change to a thinking part. Used to describe a chunk when streaming thinking responses. ### Attributes [](https://pydantic.dev/docs/ai/api/models/function/#attributes-1) #### content [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.DeltaThinkingPart.content) Incremental change to the thinking content. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### signature [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.DeltaThinkingPart.signature) Incremental change to the thinking signature. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` DeltaToolCall ------------- [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.DeltaToolCall) Incremental change to a tool call. Used to describe a chunk when streaming structured responses. ### Attributes [](https://pydantic.dev/docs/ai/api/models/function/#attributes-2) #### json\_args [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.DeltaToolCall.json_args) Incremental change to the arguments as JSON **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### name [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.DeltaToolCall.name) Incremental change to the name of the tool. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### tool\_call\_id [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.DeltaToolCall.tool_call_id) Incremental change to the tool call ID. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` FunctionModel ------------- [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel) **Bases:** `Model` A model controlled by a local function. Apart from `__init__`, all methods are private or match those of the base class. ### Attributes [](https://pydantic.dev/docs/ai/api/models/function/#attributes-3) #### model\_name [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel.model_name) The model name. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### system [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel.system) The system / model provider. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) ### Methods [](https://pydantic.dev/docs/ai/api/models/function/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel.__init__) def __init__( function: FunctionDef, *, model_name: str | None = None, profile: ModelProfileSpec | None = None, settings: ModelSettings | None = None, ) -> None def __init__( *, stream_function: StreamFunctionDef, model_name: str | None = None, profile: ModelProfileSpec | None = None, settings: ModelSettings | None = None, ) -> None def __init__( function: FunctionDef, *, stream_function: StreamFunctionDef, model_name: str | None = None, profile: ModelProfileSpec | None = None, settings: ModelSettings | None = None, ) -> None Initialize a `FunctionModel`. Either `function` or `stream_function` must be provided, providing both is allowed. ##### Parameters [](https://pydantic.dev/docs/ai/api/models/function/#parameters) **`function`** : `FunctionDef` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel.__init__(function)) The function to call for non-streamed requests. **`stream_function`** : `StreamFunctionDef` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel.__init__(stream_function)) The function to call for streamed requests. **`model_name`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel.__init__(model_name)) The name of the model. If not provided, a name is generated from the function names. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel.__init__(profile)) The model profile to use. **`settings`** : [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel.__init__(settings)) Model-specific settings that will be used as defaults for this model. #### supported\_native\_tools [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel.supported_native_tools) `@classmethod` def supported_native_tools(cls) -> frozenset[type[AbstractNativeTool]] FunctionModel supports all builtin tools for testing flexibility. ##### Returns [](https://pydantic.dev/docs/ai/api/models/function/#returns) [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[[`type`](https://docs.python.org/3/glossary.html#term-type)\ \[`AbstractNativeTool`\]\] FunctionStreamedResponse ------------------------ [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionStreamedResponse) **Bases:** `StreamedResponse` Implementation of `StreamedResponse` for [FunctionModel](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel) . ### Attributes [](https://pydantic.dev/docs/ai/api/models/function/#attributes-4) #### model\_name [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionStreamedResponse.model_name) Get the model name of the response. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### provider\_name [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionStreamedResponse.provider_name) Get the provider name. **Type:** [`None`](https://docs.python.org/3/library/constants.html#None) #### provider\_url [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionStreamedResponse.provider_url) Get the provider base URL. **Type:** [`None`](https://docs.python.org/3/library/constants.html#None) #### timestamp [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionStreamedResponse.timestamp) Get the timestamp of the response. **Type:** [`datetime`](https://docs.python.org/3/library/datetime.html#module-datetime) DeltaThinkingCalls ------------------ [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.DeltaThinkingCalls) A mapping of thinking call IDs to incremental changes. **Type:** [`TypeAlias`](https://docs.python.org/3/library/typing.html#typing.TypeAlias) **Default:** `dict[int, DeltaThinkingPart]` DeltaToolCalls -------------- [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.DeltaToolCalls) A mapping of tool call IDs to incremental changes. **Type:** [`TypeAlias`](https://docs.python.org/3/library/typing.html#typing.TypeAlias) **Default:** `dict[int, DeltaToolCall]` FunctionDef ----------- [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionDef) A function used to generate a non-streamed response. **Type:** [`TypeAlias`](https://docs.python.org/3/library/typing.html#typing.TypeAlias) **Default:** `Callable[[list[ModelMessage], AgentInfo], ModelResponse | Awaitable[ModelResponse]]` StreamFunctionDef ----------------- [](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.StreamFunctionDef) A function used to generate a streamed response. While this is defined as having return type of `AsyncIterator[str | DeltaToolCalls | DeltaThinkingCalls | BuiltinTools]`, it should really be considered as `AsyncIterator[str] | AsyncIterator[DeltaToolCalls] | AsyncIterator[DeltaThinkingCalls]`, E.g. you need to yield all text, all `DeltaToolCalls`, all `DeltaThinkingCalls`, or all `BuiltinToolCallsReturns`, not mix them. **Type:** [`TypeAlias`](https://docs.python.org/3/library/typing.html#typing.TypeAlias) **Default:** `Callable[[list[ModelMessage], AgentInfo], AsyncIterator[str | DeltaToolCalls | DeltaThinkingCalls | BuiltinToolCallsReturns]]` Was this page helpful? Thanks for your feedback! --- # google | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/google/#_top) google ====== Interface that uses the [`google-genai`](https://pypi.org/project/google-genai/) package under the hood to access Google’s Gemini models via both the Gemini API and Google Cloud (formerly known as Vertex AI). Setup ----- [](https://pydantic.dev/docs/ai/api/models/google/#setup) For details on how to set up authentication with this model, see [model configuration for Google](https://pydantic.dev/docs/ai/models/google/) . GeminiStreamedResponse ---------------------- [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GeminiStreamedResponse) **Bases:** `StreamedResponse` Implementation of `StreamedResponse` for the Gemini model. ### Attributes [](https://pydantic.dev/docs/ai/api/models/google/#attributes) #### model\_name [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GeminiStreamedResponse.model_name) Get the model name of the response. **Type:** `GoogleModelName` #### provider\_name [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GeminiStreamedResponse.provider_name) Get the provider name. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### provider\_url [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GeminiStreamedResponse.provider_url) Get the provider base URL. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### timestamp [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GeminiStreamedResponse.timestamp) Get the timestamp of the response. **Type:** [`datetime`](https://docs.python.org/3/library/datetime.html#module-datetime) GoogleModel ----------- [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModel) **Bases:** `Model[Client]` A model that uses Gemini via `generativelanguage.googleapis.com` API. This is implemented from scratch rather than using a dedicated SDK, good API documentation is available [here](https://ai.google.dev/api) . Apart from `__init__`, all methods are private or match those of the base class. ### Attributes [](https://pydantic.dev/docs/ai/api/models/google/#attributes-1) #### model\_name [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModel.model_name) The model name. **Type:** `GoogleModelName` #### system [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModel.system) The model provider. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) ### Methods [](https://pydantic.dev/docs/ai/api/models/google/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModel.__init__) def __init__( model_name: GoogleModelName, *, provider: Literal['google', 'google-cloud', 'gateway'] | Provider[Client] = 'google', profile: ModelProfileSpec | None = None, settings: ModelSettings | None = None, ) Initialize a Gemini model. ##### Parameters [](https://pydantic.dev/docs/ai/api/models/google/#parameters) **`model_name`** : `GoogleModelName` [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModel.__init__(model_name)) The name of the model to use. **`provider`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘google’, ‘google-cloud’, ‘gateway’\] | `Provider`\[`Client`\] _Default:_ `'google'` [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModel.__init__(provider)) The provider to use for authentication and API access. Can be either the string ‘google’ (Gemini API) or ‘google-cloud’ (Google Cloud, formerly known as Vertex AI), or an instance of `Provider[google.genai.AsyncClient]`. Defaults to ‘google’. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModel.__init__(profile)) The model profile to use. Defaults to a profile picked by the provider based on the model name. **`settings`** : [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModel.__init__(settings)) The model settings to use. Defaults to None. #### supported\_native\_tools [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModel.supported_native_tools) `@classmethod` def supported_native_tools(cls) -> frozenset[type[AbstractNativeTool]] Return the set of native tool types this model can handle. ##### Returns [](https://pydantic.dev/docs/ai/api/models/google/#returns) [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[[`type`](https://docs.python.org/3/glossary.html#term-type)\ \[`AbstractNativeTool`\]\] GoogleModelSettings ------------------- [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings) **Bases:** [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) Settings used for a Gemini model request. ### Attributes [](https://pydantic.dev/docs/ai/api/models/google/#attributes-2) #### google\_cached\_content [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings.google_cached_content) The name of the cached content to use for the model. When set, `system_instruction`, `tools`, and `tool_config` are omitted from the outgoing request — the cached content resource owns those fields, and both the Gemini API and Vertex AI reject requests that supply them alongside `cached_content` (`400 INVALID_ARGUMENT`: “Tool config, tools and system instruction should not be set in the request when using cached content.”). Any tools registered on the agent and any system prompt are therefore ignored on requests that go through the cache; a `UserWarning` is emitted whenever stripping actually drops a field so the mismatch is discoverable. See [https://ai.google.dev/gemini-api/docs/caching](https://ai.google.dev/gemini-api/docs/caching) for more information. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### google\_cloud\_service\_tier [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings.google_cloud_service_tier) The service tier to use for the model request when using Google Cloud. Controls routing for Provisioned Throughput, Flex PayGo, and Priority PayGo (e.g., `'pt_only'`, `'flex_only'`, `'priority_only'`). See [`GoogleCloudServiceTier`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleCloudServiceTier) for all values, headers sent, and links to Google docs. **Type:** `GoogleCloudServiceTier` #### google\_labels [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings.google_labels) User-defined metadata to break down billed charges. Only supported by the Vertex AI API. See the [Gemini API docs](https://cloud.google.com/vertex-ai/generative-ai/docs/multimodal/add-labels-to-api-calls) for use cases and limitations. **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] #### google\_logprobs [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings.google_logprobs) Include log probabilities in the response. See [https://docs.cloud.google.com/vertex-ai/generative-ai/docs/multimodal/content-generation-parameters#log-probabilities-output-tokens](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/multimodal/content-generation-parameters#log-probabilities-output-tokens) for more information. Note: Only supported for Vertex AI and non-streaming requests. These will be included in `ModelResponse.provider_details['logprobs']`. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### google\_model\_armor\_config [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings.google_model_armor_config) Model Armor configuration for screening prompts and responses. Only supported by the Vertex AI API. Specifies the Model Armor templates to use for sanitizing user prompts and model responses. Both fields are optional — omit either to skip screening for that direction. Mutually exclusive with `google_safety_settings`: Vertex AI rejects a request that sets both, since Model Armor replaces the built-in safety filters for that request. Note: Model Armor screening — both prompt and response — is only applied for non-streaming requests. Google’s API ignores `modelArmorConfig` for streaming requests (`streamGenerateContent`). See the [Model Armor docs](https://cloud.google.com/security-command-center/docs/model-armor-overview) for use cases and limitations. **Type:** `ModelArmorConfigDict` #### google\_safety\_settings [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings.google_safety_settings) The safety settings to use for the model. See [https://ai.google.dev/gemini-api/docs/safety-settings](https://ai.google.dev/gemini-api/docs/safety-settings) for more information. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[`SafetySettingDict`\] #### google\_thinking\_config [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings.google_thinking_config) The thinking configuration to use for the model. See [https://ai.google.dev/gemini-api/docs/thinking](https://ai.google.dev/gemini-api/docs/thinking) for more information. **Type:** `ThinkingConfigDict` #### google\_top\_logprobs [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings.google_top_logprobs) Include log probabilities of the top n tokens in the response. See [https://docs.cloud.google.com/vertex-ai/generative-ai/docs/multimodal/content-generation-parameters#log-probabilities-output-tokens](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/multimodal/content-generation-parameters#log-probabilities-output-tokens) for more information. Note: Only supported for Vertex AI and non-streaming requests. These will be included in `ModelResponse.provider_details['logprobs']`. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### google\_video\_resolution [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings.google_video_resolution) The video resolution to use for the model. See [https://ai.google.dev/api/generate-content#MediaResolution](https://ai.google.dev/api/generate-content#MediaResolution) for more information. **Type:** `MediaResolution` GoogleCloudServiceTier ---------------------- [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleCloudServiceTier) Values for the `google_cloud_service_tier` field on [`GoogleModelSettings`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings) . Controls Google Cloud HTTP headers for [Provisioned Throughput](https://cloud.google.com/vertex-ai/generative-ai/docs/provisioned-throughput/use-provisioned-throughput) (PT), [Flex PayGo](https://cloud.google.com/vertex-ai/generative-ai/docs/flex-paygo) , and [Priority PayGo](https://cloud.google.com/vertex-ai/generative-ai/docs/priority-paygo) . * `'pt_then_on_demand'` (**default**): PT when quota allows, then standard on-demand spillover. No headers sent. * `'pt_only'`: PT only (`X-Vertex-AI-LLM-Request-Type: dedicated`). No on-demand spillover; returns 429 when over quota. * `'pt_then_flex'`: PT when quota allows, then [Flex PayGo](https://cloud.google.com/vertex-ai/generative-ai/docs/flex-paygo) spillover (`X-Vertex-AI-LLM-Shared-Request-Type: flex`). * `'pt_then_priority'`: PT when quota allows, then [Priority PayGo](https://cloud.google.com/vertex-ai/generative-ai/docs/priority-paygo) spillover (`X-Vertex-AI-LLM-Shared-Request-Type: priority`). * `'on_demand'`: Standard on-demand only (`X-Vertex-AI-LLM-Request-Type: shared`). Bypasses PT for this request. * `'flex_only'`: [Flex PayGo](https://cloud.google.com/vertex-ai/generative-ai/docs/flex-paygo) only (`X-Vertex-AI-LLM-Request-Type: shared` and `X-Vertex-AI-LLM-Shared-Request-Type: flex`). Bypasses PT. * `'priority_only'`: [Priority PayGo](https://cloud.google.com/vertex-ai/generative-ai/docs/priority-paygo) only (`X-Vertex-AI-LLM-Request-Type: shared` and `X-Vertex-AI-LLM-Shared-Request-Type: priority`). Bypasses PT. Not every model or region supports every value; see the linked Google docs. **Default:** `Literal['pt_then_on_demand', 'pt_only', 'pt_then_flex', 'pt_then_priority', 'on_demand', 'flex_only', 'priority_only']` GoogleModelName --------------- [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelName) Possible Gemini model names. Since Gemini supports a variety of date-stamped models, we explicitly list the latest models but allow any name in the type hints. See [the Gemini API docs](https://ai.google.dev/gemini-api/docs/models/gemini#model-variations) for a full list. **Default:** `str | LatestGoogleModelNames` LatestGoogleModelNames ---------------------- [](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.LatestGoogleModelNames) Latest Gemini models. **Default:** `Literal['gemini-flash-latest', 'gemini-flash-lite-latest', 'gemini-2.0-flash', 'gemini-2.0-flash-lite', 'gemini-2.5-flash', 'gemini-2.5-flash-preview-09-2025', 'gemini-2.5-flash-image', 'gemini-2.5-flash-lite', 'gemini-2.5-pro', 'gemini-3-flash-preview', 'gemini-3-pro-image', 'gemini-3-pro-image-preview', 'gemini-3-pro-preview', 'gemini-3.1-flash-image', 'gemini-3.1-flash-image-preview', 'gemini-3.1-flash-lite', 'gemini-3.1-pro-preview', 'gemini-3.5-flash', 'gemini-3.5-flash-lite', 'gemini-3.6-flash', 'gemini-3.7-flash']` Was this page helpful? Thanks for your feedback! --- # groq | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/groq/#_top) groq ==== Setup ----- [](https://pydantic.dev/docs/ai/api/models/groq/#setup) For details on how to set up authentication with this model, see [model configuration for Groq](https://pydantic.dev/docs/ai/models/groq/) . GroqModel --------- [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqModel) **Bases:** `Model[AsyncGroq]` A model that uses the Groq API. Internally, this uses the [Groq Python client](https://github.com/groq/groq-python) to interact with the API. Apart from `__init__`, all methods are private or match those of the base class. ### Attributes [](https://pydantic.dev/docs/ai/api/models/groq/#attributes) #### model\_name [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqModel.model_name) The model name. **Type:** `GroqModelName` #### system [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqModel.system) The model provider. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) ### Methods [](https://pydantic.dev/docs/ai/api/models/groq/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqModel.__init__) def __init__( model_name: GroqModelName, *, provider: Literal['groq', 'gateway'] | Provider[AsyncGroq] = 'groq', profile: ModelProfileSpec | None = None, settings: ModelSettings | None = None, ) Initialize a Groq model. ##### Parameters [](https://pydantic.dev/docs/ai/api/models/groq/#parameters) **`model_name`** : `GroqModelName` [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqModel.__init__(model_name)) The name of the Groq model to use. List of model names available [here](https://console.groq.com/docs/models) . **`provider`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘groq’, ‘gateway’\] | `Provider`\[`AsyncGroq`\] _Default:_ `'groq'` [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqModel.__init__(provider)) The provider to use for authentication and API access. Can be either the string ‘groq’ or an instance of `Provider[AsyncGroq]`. If not provided, a new provider will be created using the other parameters. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqModel.__init__(profile)) The model profile to use. Defaults to a profile picked by the provider based on the model name. **`settings`** : [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqModel.__init__(settings)) Model-specific settings that will be used as defaults for this model. #### supported\_native\_tools [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqModel.supported_native_tools) `@classmethod` def supported_native_tools(cls) -> frozenset[type[AbstractNativeTool]] Return the set of builtin tool types this model can handle. ##### Returns [](https://pydantic.dev/docs/ai/api/models/groq/#returns) [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[[`type`](https://docs.python.org/3/glossary.html#term-type)\ \[`AbstractNativeTool`\]\] GroqModelSettings ----------------- [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqModelSettings) **Bases:** [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) Settings used for a Groq model request. ### Attributes [](https://pydantic.dev/docs/ai/api/models/groq/#attributes-1) #### groq\_reasoning\_effort [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqModelSettings.groq_reasoning_effort) The reasoning effort level. See [the Groq docs](https://console.groq.com/docs/reasoning#reasoning-effort) for more details. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘none’, ‘default’, ‘low’, ‘medium’, ‘high’\] #### groq\_reasoning\_format [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqModelSettings.groq_reasoning_format) The format of the reasoning output. See [the Groq docs](https://console.groq.com/docs/reasoning#reasoning-format) for more details. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘hidden’, ‘raw’, ‘parsed’\] GroqStreamedResponse -------------------- [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqStreamedResponse) **Bases:** `StreamedResponse` Implementation of `StreamedResponse` for Groq models. ### Attributes [](https://pydantic.dev/docs/ai/api/models/groq/#attributes-2) #### model\_name [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqStreamedResponse.model_name) Get the model name of the response. **Type:** `GroqModelName` #### provider\_name [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqStreamedResponse.provider_name) Get the provider name. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### provider\_url [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqStreamedResponse.provider_url) Get the provider base URL. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### timestamp [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqStreamedResponse.timestamp) Get the timestamp of the response. **Type:** [`datetime`](https://docs.python.org/3/library/datetime.html#module-datetime) GroqModelName ------------- [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqModelName) Possible Groq model names. Since Groq supports a variety of models and the list changes frequently, we explicitly list the named models as of 2025-03-31 but allow any name in the type hints. See [https://console.groq.com/docs/models](https://console.groq.com/docs/models) for an up to date list of models and more details. **Default:** `str | ProductionGroqModelNames | PreviewGroqModelNames` PreviewGroqModelNames --------------------- [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.PreviewGroqModelNames) Preview Groq models from [https://console.groq.com/docs/models#preview-models](https://console.groq.com/docs/models#preview-models) . **Default:** `Literal['meta-llama/llama-4-maverick-17b-128e-instruct', 'meta-llama/llama-prompt-guard-2-22m', 'meta-llama/llama-prompt-guard-2-86m', 'openai/gpt-oss-safeguard-20b', 'playai-tts', 'playai-tts-arabic']` ProductionGroqModelNames ------------------------ [](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.ProductionGroqModelNames) Production Groq models from [https://console.groq.com/docs/models#production-models](https://console.groq.com/docs/models#production-models) . **Default:** `Literal['llama-3.1-8b-instant', 'llama-3.3-70b-versatile', 'meta-llama/llama-guard-4-12b', 'openai/gpt-oss-120b', 'openai/gpt-oss-20b', 'whisper-large-v3', 'whisper-large-v3-turbo']` Was this page helpful? Thanks for your feedback! --- # ollama | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/ollama/#_top) ollama ====== Setup ----- [](https://pydantic.dev/docs/ai/api/models/ollama/#setup) For details on how to set up authentication with this model, see [model configuration for Ollama](https://pydantic.dev/docs/ai/models/ollama/) . Ollama model implementation using OpenAI-compatible API. OllamaModel ----------- [](https://pydantic.dev/docs/ai/api/models/ollama/#pydantic_ai.models.ollama.OllamaModel) **Bases:** `OpenAIChatModel` A model that uses Ollama’s OpenAI-compatible Chat Completions API. Self-hosted Ollama (v0.5.0+) honors `response_format` with `json_schema` via `llama.cpp`’s grammar-constrained decoder, so `NativeOutput` produces schema-valid output at generation time. Ollama Cloud currently accepts `response_format` with `json_schema` without error but does not enforce the schema upstream (see [pydantic-ai#4917](https://github.com/pydantic/pydantic-ai/issues/4917) and [ollama/ollama#12362](https://github.com/ollama/ollama/issues/12362) ). When this model detects a Cloud path — either a `base_url` on `ollama.com` or a model name ending in `-cloud` — it disables `supports_json_schema_output` on the resolved profile. With that flag off, [`NativeOutput`](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.NativeOutput) raises a clear [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) so users pick a mode that actually works on Cloud ([`ToolOutput`](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.ToolOutput) — the default — and [`PromptedOutput`](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.PromptedOutput) are both verified to work). Apart from `__init__`, all methods are inherited from the base class. ### Methods [](https://pydantic.dev/docs/ai/api/models/ollama/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/ollama/#pydantic_ai.models.ollama.OllamaModel.__init__) def __init__( model_name: str, *, provider: Literal['ollama'] | Provider[AsyncOpenAI] = 'ollama', profile: ModelProfileSpec | None = None, settings: ModelSettings | None = None, ) Initialize an Ollama model. ##### Parameters [](https://pydantic.dev/docs/ai/api/models/ollama/#parameters) **`model_name`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/api/models/ollama/#pydantic_ai.models.ollama.OllamaModel.__init__(model_name)) The name of the Ollama model to use (e.g. `'qwen3'`, `'llama3.2'`). **`provider`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘ollama’\] | `Provider`\[`AsyncOpenAI`\] _Default:_ `'ollama'` [](https://pydantic.dev/docs/ai/api/models/ollama/#pydantic_ai.models.ollama.OllamaModel.__init__(provider)) The provider to use. Defaults to `'ollama'`. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/ollama/#pydantic_ai.models.ollama.OllamaModel.__init__(profile)) The model profile to use. Defaults to a profile picked by the provider based on the model name, adjusted to disable `supports_json_schema_output` when the request routes through Ollama Cloud. **`settings`** : [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/ollama/#pydantic_ai.models.ollama.OllamaModel.__init__(settings)) Model-specific settings that will be used as defaults for this model. Was this page helpful? Thanks for your feedback! --- # wrapper | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/wrapper/#_top) wrapper ======= WrapperModel ------------ [](https://pydantic.dev/docs/ai/api/models/wrapper/#pydantic_ai.models.wrapper.WrapperModel) **Bases:** `Model` Model which wraps another model. Does nothing on its own, used as a base class. ### Attributes [](https://pydantic.dev/docs/ai/api/models/wrapper/#attributes) #### settings [](https://pydantic.dev/docs/ai/api/models/wrapper/#pydantic_ai.models.wrapper.WrapperModel.settings) Get the settings from the wrapped model. **Type:** [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) #### wrapped [](https://pydantic.dev/docs/ai/api/models/wrapper/#pydantic_ai.models.wrapper.WrapperModel.wrapped) The underlying model being wrapped. **Type:** `Model` **Default:** `infer_model(wrapped)` Was this page helpful? Thanks for your feedback! --- # Medical Agent Delegation | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/complex-workflows/medical-agent-delegation/#_top) Medical Agent Delegation ======================== Medical triage and delegation system built with **Pydantic AI**, demonstrating how an orchestrator agent (`triage_agent`) coordinates multiple specialized agents (e.g. cardiology, neurology, and senior clinician). Demonstrates: * [Agent delegation and coordination](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#agent-delegation) * [structured `output_type`](https://pydantic.dev/docs/ai/core-concepts/output/#structured-output) * [tools](https://pydantic.dev/docs/ai/tools-toolsets/tools/) * * * This example shows how to use **multiple Pydantic AI agents** to simulate a medical triage workflow. The system includes: * **General Practitioner, Cardiology, and Neurology agents** — for Level 1 consultation. * **Senior Doctor agent** — for escalations and treatment planning. * **Triage Agent (Coordinator)** — which decides which tool to invoke and when to escalate. The `triage_agent` uses two tools: 1. `consult_specialist` — routes the complaint to a domain specialist. 2. `consult_senior_doctor` — escalates the case for critical or ambiguous scenarios. Each specialist produces a structured `MedicalReport`, and the senior doctor produces a structured `TreatmentPlan`. The orchestrator then compiles both into a final `TriageFinalOutput`. * * * Running the Example ------------------- [](https://pydantic.dev/docs/ai/examples/complex-workflows/medical-agent-delegation/#running-the-example) With [dependencies installed and environment variables set](https://pydantic.dev/docs/ai/examples/setup/#usage) , run: Terminal python -m pydantic_ai_examples.medical_agent_delegation Was this page helpful? Thanks for your feedback! --- # format_prompt | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic-ai/format_prompt/#_top) format\_prompt ============== format\_as\_xml --------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/format_prompt/#pydantic_ai.format_prompt.format_as_xml) def format_as_xml( obj: Any, root_tag: str | None = None, item_tag: str = 'item', none_str: str = 'null', indent: str | None = ' ', include_field_info: Literal['once'] | bool = False, ) -> str Format a Python object as XML. This is useful since LLMs often find it easier to read semi-structured data (e.g. examples) as XML, rather than JSON etc. Supports: `str`, `bytes`, `bytearray`, `bool`, `int`, `float`, `Decimal`, `date`, `datetime`, `time`, `timedelta`, `UUID`, `Enum`, `Mapping`, `Iterable`, `dataclass`, and `BaseModel`. Example: format\_as\_xml\_example.py from pydantic_ai import format_as_xml print(format_as_xml({'name': 'John', 'height': 6, 'weight': 200}, root_tag='user')) ''' John 6 200 ''' ### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/format_prompt/#returns) [`str`](https://docs.python.org/3/library/stdtypes.html#str) — XML representation of the object. ### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/format_prompt/#parameters) **`obj`** : [`Any`](https://docs.python.org/3/library/typing.html#typing.Any) [](https://pydantic.dev/docs/ai/api/pydantic-ai/format_prompt/#pydantic_ai.format_prompt.format_as_xml(obj)) Python Object to serialize to XML. **`root_tag`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/format_prompt/#pydantic_ai.format_prompt.format_as_xml(root_tag)) Outer tag to wrap the XML in, use `None` to omit the outer tag. **`item_tag`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) _Default:_ `'item'` [](https://pydantic.dev/docs/ai/api/pydantic-ai/format_prompt/#pydantic_ai.format_prompt.format_as_xml(item_tag)) Tag to use for each item in an iterable (e.g. list), this is overridden by the class name for dataclasses and Pydantic models. **`none_str`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) _Default:_ `'null'` [](https://pydantic.dev/docs/ai/api/pydantic-ai/format_prompt/#pydantic_ai.format_prompt.format_as_xml(none_str)) String to use for `None` values. **`indent`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `' '` [](https://pydantic.dev/docs/ai/api/pydantic-ai/format_prompt/#pydantic_ai.format_prompt.format_as_xml(indent)) Indentation string to use for pretty printing. **`include_field_info`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘once’\] | [`bool`](https://docs.python.org/3/library/functions.html#bool) _Default:_ `False` [](https://pydantic.dev/docs/ai/api/pydantic-ai/format_prompt/#pydantic_ai.format_prompt.format_as_xml(include_field_info)) Whether to include attributes like Pydantic `Field` attributes and dataclasses `field()` `metadata` as XML attributes. In both cases the allowed `Field` attributes and `field()` metadata keys are `title` and `description`. If a field is repeated in the data (e.g. in a list) by setting `once` the attributes are included only in the first occurrence of an XML element relative to the same field. Was this page helpful? Thanks for your feedback! --- # mcp-sampling | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/mcp-sampling/#_top) mcp-sampling ============ MCPSamplingModel ---------------- [](https://pydantic.dev/docs/ai/api/models/mcp-sampling/#pydantic_ai.models.mcp_sampling.MCPSamplingModel) **Bases:** `Model` A model that uses MCP Sampling. [MCP Sampling](https://modelcontextprotocol.io/docs/concepts/sampling) allows an MCP server to make requests to a model by calling back to the MCP client that connected to it. ### Attributes [](https://pydantic.dev/docs/ai/api/models/mcp-sampling/#attributes) #### default\_max\_tokens [](https://pydantic.dev/docs/ai/api/models/mcp-sampling/#pydantic_ai.models.mcp_sampling.MCPSamplingModel.default_max_tokens) Default max tokens to use if not set in [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.max_tokens) . Max tokens is a required parameter for MCP Sampling, but optional on [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) , so this value is used as fallback. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `16384` #### model\_name [](https://pydantic.dev/docs/ai/api/models/mcp-sampling/#pydantic_ai.models.mcp_sampling.MCPSamplingModel.model_name) The model name. Since the model name isn’t known until the request is made, this property always returns `'mcp-sampling'`. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### session [](https://pydantic.dev/docs/ai/api/models/mcp-sampling/#pydantic_ai.models.mcp_sampling.MCPSamplingModel.session) The MCP server session to use for sampling. **Type:** `ServerSession` #### system [](https://pydantic.dev/docs/ai/api/models/mcp-sampling/#pydantic_ai.models.mcp_sampling.MCPSamplingModel.system) The system / model provider, returns `'MCP'`. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) MCPSamplingModelSettings ------------------------ [](https://pydantic.dev/docs/ai/api/models/mcp-sampling/#pydantic_ai.models.mcp_sampling.MCPSamplingModelSettings) **Bases:** [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) Settings used for an MCP Sampling model request. ### Attributes [](https://pydantic.dev/docs/ai/api/models/mcp-sampling/#attributes-1) #### mcp\_model\_preferences [](https://pydantic.dev/docs/ai/api/models/mcp-sampling/#pydantic_ai.models.mcp_sampling.MCPSamplingModelSettings.mcp_model_preferences) Model preferences to use for MCP Sampling. **Type:** `ModelPreferences` Was this page helpful? Thanks for your feedback! --- # Process Event Stream | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/process-event-stream/#_top) Process Event Stream ==================== [`ProcessEventStream`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ProcessEventStream) is a [capability](https://pydantic.dev/docs/ai/capabilities/overview/) that forwards the agent’s stream of [`AgentStreamEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.AgentStreamEvent) s — model streaming and tool execution events — to a handler. When it’s registered, `agent.run()` automatically enables streaming, so the handler fires without passing an explicit [`event_stream_handler`](https://pydantic.dev/docs/ai/core-concepts/agent/#streaming-all-events) argument: During a realtime session, the stream also contains realtime-only [`RealtimeEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeEvent) members. process\_event\_stream.py from collections.abc import AsyncIterable from pydantic_ai import Agent, AgentStreamEvent, RunContext from pydantic_ai.capabilities import ProcessEventStream async def log_events(ctx: RunContext, events: AsyncIterable[AgentStreamEvent]) -> None: async for event in events: print(event) # (1) agent = Agent('openai:gpt-5.2', capabilities=[ProcessEventStream(log_events)]) For example, forward events to a websocket, progress bar, or audit log. The handler comes in two forms: * An [`EventStreamHandler`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.EventStreamHandler) — an `async def` returning `None`, as above. Events are forwarded to the handler and passed through unchanged, so multiple handlers (and a top-level `event_stream_handler` argument) can all observe the same stream. Events are delivered synchronously, so a slow handler back-pressures the rest of the stream. * An `EventStreamProcessor` — an async generator that yields events. What it yields replaces the stream for downstream consumers, so it can modify, drop, or add events. Registering the capability composes with other streaming mechanisms: see [Streaming all events](https://pydantic.dev/docs/ai/core-concepts/agent/#streaming-all-events) for the event vocabulary and handler examples. Was this page helpful? Thanks for your feedback! --- # Web Fetch | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/web-fetch/#_top) Web Fetch ========= The [`WebFetch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebFetch) [capability](https://pydantic.dev/docs/ai/capabilities/overview/) lets your agent fetch the contents of URLs. Like all [provider-adaptive tools](https://pydantic.dev/docs/ai/capabilities/overview/#provider-adaptive-tools) , it prefers the provider’s native web fetch tool and can fall back to a local implementation on other models. [`WebFetch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebFetch) defaults to native-only. Backed by [`WebFetchTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebFetchTool) on the native side (see [Web Fetch Tool](https://pydantic.dev/docs/ai/tools-toolsets/native-tools/#web-fetch-tool) for provider support and configuration) — pass `native=WebFetchTool(...)` directly for full control. For the local side, pass `local=True` for the bundled [markdownify-based fetch tool](https://pydantic.dev/docs/ai/tools-toolsets/common-tools/#web-fetch-tool) (requires the `web-fetch` optional group), or any callable, [`Tool`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.Tool) , or [`AbstractToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AbstractToolset) . Native constraint fields: `allowed_domains`, `blocked_domains`, `max_uses`, `enable_citations`, `max_content_tokens`. Only `max_uses` requires native; domain filters are enforced locally when native isn’t available. web\_fetch.py from pydantic_ai.capabilities import WebFetch # Native-only — raises on models without native web fetch WebFetch() # Native preferred; markdownify-based fallback (needs `pydantic-ai-slim[web-fetch]`) WebFetch(local=True) # Domain filters enforced locally when native isn't available WebFetch(allowed_domains=['example.com'], local=True) Was this page helpful? Thanks for your feedback! --- # Agent Handoff | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/realtime/realtime-handoff/#_top) Agent Handoff ============= Realtime speech-to-speech models are great conversationalists, but they don’t produce structured output. This example shows the robust pattern: let the realtime model run the live conversation, then hand its [message history](https://pydantic.dev/docs/ai/core-concepts/message-history/) to a normal [`Agent.run()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.run) with `output_type` to extract a typed result. Because a [realtime session](https://pydantic.dev/docs/ai/realtime/overview/) records the _same_ [`ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage) history a text agent produces, the handoff is just passing [`session.all_messages()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.all_messages) along — realtime and non-realtime runs are peers that interoperate through message history. Demonstrates: * [realtime sessions](https://pydantic.dev/docs/ai/realtime/overview/) * [structured output](https://pydantic.dev/docs/ai/core-concepts/output/) via a text-agent handoff * [message history](https://pydantic.dev/docs/ai/core-concepts/message-history/) shared across realtime and non-realtime runs The example models a short support call: a caller describes a problem to the realtime voice agent, then the accumulated conversation is handed to a text agent that distills it into a typed `SupportTicket`. The caller’s side is driven with text turns so the example runs without a microphone — a real app would stream microphone audio with [`send_audio()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.send_audio) instead (see the [voice assistant example](https://pydantic.dev/docs/ai/examples/realtime/realtime-voice/) ). The handoff only runs after every scripted caller turn receives a [`RealtimeTurnCompleteEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeTurnCompleteEvent) . If the realtime connection ends early, the example raises an error rather than creating a ticket from a partial call. Running the Example ------------------- [](https://pydantic.dev/docs/ai/examples/realtime/realtime-handoff/#running-the-example) Both the realtime `gpt-realtime` model and the text triage agent run on OpenAI, so you’ll need an OpenAI API key set via `OPENAI_API_KEY`. With [dependencies installed and environment variables set](https://pydantic.dev/docs/ai/examples/setup/#usage) , run: * [pip](https://pydantic.dev/docs/ai/examples/realtime/realtime-handoff/#tab-panel-50) * [uv](https://pydantic.dev/docs/ai/examples/realtime/realtime-handoff/#tab-panel-51) Terminal python -m pydantic_ai_examples.realtime_handoff Terminal uv run -m pydantic_ai_examples.realtime_handoff Example Code ------------ [](https://pydantic.dev/docs/ai/examples/realtime/realtime-handoff/#example-code) realtime\_handoff.py from __future__ import annotations import asyncio from typing import Literal import logfire from pydantic import BaseModel from pydantic_ai import Agent, PartEndEvent, SpeechPart from pydantic_ai.realtime import RealtimeTurnCompleteEvent # 'if-token-present' means nothing will be sent (and the example will work) if you don't have logfire configured logfire.configure(send_to_logfire='if-token-present') logfire.instrument_pydantic_ai() class SupportTicket(BaseModel): """The structured ticket distilled from the spoken support call.""" summary: str category: Literal['hardware', 'software', 'billing', 'other'] priority: Literal['low', 'medium', 'high'] follow_up_questions: list[str] # The realtime model runs the live conversation. voice_agent = Agent( instructions='You are a friendly, concise phone support agent. Ask one question at a time.' ) # A normal text agent turns the finished conversation into a typed result — something a realtime # model can't do itself. triage_agent = Agent( 'openai:gpt-5.2', output_type=SupportTicket, instructions='Summarize the support call as a structured ticket.', ) # What the caller "says" — each line is one spoken turn, driven as text so the example runs without # a microphone. CALLER_TURNS = [\ "Hi, my laptop won't charge anymore — the light doesn't come on when I plug it in.",\ 'I already tried a different outlet and it still does nothing. I need it for a presentation tomorrow.',\ ] async def main() -> None: async with voice_agent.realtime('openai:gpt-realtime').session() as session: # A session is consumed with a single event loop. We drive the caller's turns from inside it: # send the first line, then send the next one each time the model finishes a turn. remaining_turns = iter(CALLER_TURNS) first_turn = next(remaining_turns) print(f'caller: {first_turn}') # Sending text into an OpenAI realtime session asks the model to respond right away. await session.send(first_turn) async for event in session: match event: case PartEndEvent( part=SpeechPart(speaker='assistant', transcript=transcript) ) if transcript: print(f'agent: {transcript}') case RealtimeTurnCompleteEvent(): next_turn = next(remaining_turns, None) if next_turn is None: break # The caller has said everything; end the call. print(f'caller: {next_turn}') await session.send(next_turn) case _: pass else: # The event stream ended without the `break` above, i.e. before the call completed. raise RuntimeError( 'The realtime session ended before the support call completed' ) # The realtime session recorded ordinary `ModelMessage` history; hand it off to the text # agent, which can do the structured extraction the realtime model can't. handoff_history = session.all_messages() ticket = await triage_agent.run( 'Create the support ticket for this call.', message_history=handoff_history ) print(f'\nStructured ticket:\n{ticket.output.model_dump_json(indent=2)}') if __name__ == '__main__': asyncio.run(main()) Was this page helpful? Thanks for your feedback! --- # Instrumentation | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/instrumentation/#_top) Instrumentation =============== [`Instrumentation`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Instrumentation) is a [capability](https://pydantic.dev/docs/ai/capabilities/overview/) that instruments agent runs with OpenTelemetry tracing: it creates spans for the run itself, each model request, and each tool execution, following the [OpenTelemetry Semantic Conventions for Generative AI](https://opentelemetry.io/docs/specs/semconv/gen-ai/) . Combined with [Pydantic Logfire](https://pydantic.dev/docs/ai/integrations/logfire/) (or any OTel backend), it gives you full visibility into what your agent is doing: instrumentation\_capability.py import logfire from pydantic_ai import Agent from pydantic_ai.capabilities import Instrumentation logfire.configure() # (1) agent = Agent('openai:gpt-5.2', capabilities=[Instrumentation()]) Sets the global `TracerProvider` that `Instrumentation` uses by default. Any OpenTelemetry SDK configuration works too. Pass [`InstrumentationSettings`](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentationSettings) via `Instrumentation(settings=...)` to customize providers, content capture, and the conventions version. To instrument every agent in your application instead of attaching the capability per agent, use [`Agent.instrument_all()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.instrument_all) . Other capabilities can attach attributes to the created spans through the OpenTelemetry API (`opentelemetry.trace.get_current_span().set_attribute(...)`). See [Debugging and Monitoring](https://pydantic.dev/docs/ai/integrations/logfire/) for the full guide: setup, what gets captured, semantic-conventions versions, and excluding sensitive or binary content. Was this page helpful? Thanks for your feedback! --- # Installation | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/overview/install/#_top) Installation ============ Pydantic AI is available on PyPI as [`pydantic-ai`](https://pypi.org/project/pydantic-ai/) so installation is as simple as: * [pip](https://pydantic.dev/docs/ai/overview/install/#tab-panel-150) * [uv](https://pydantic.dev/docs/ai/overview/install/#tab-panel-151) Terminal pip install pydantic-ai Terminal uv add pydantic-ai (Requires Python 3.10+) This installs the `pydantic_ai` package, core dependencies, and libraries required to use the OpenAI, Anthropic, and Google models, plus the [CLI](https://pydantic.dev/docs/ai/integrations/cli/) , [MCP](https://pydantic.dev/docs/ai/mcp/client/) , [Evals](https://pydantic.dev/docs/ai/evals/evals/) , [Web UI](https://pydantic.dev/docs/ai/integrations/ui/overview/) , [Retries](https://pydantic.dev/docs/ai/models/http-request-retries/) , and [Logfire](https://pydantic.dev/docs/ai/integrations/logfire/) integrations. To use any other models or integrations, add the relevant extras to your install command, e.g. `pydantic-ai[bedrock,temporal]`. Alternatively, you can install the [`pydantic-ai-slim`](https://pydantic.dev/docs/ai/overview/install/#slim-install) package with only the extras you need. Use with Pydantic Logfire ------------------------- [](https://pydantic.dev/docs/ai/overview/install/#use-with-pydantic-logfire) Pydantic AI has an excellent (but completely optional) integration with [Pydantic Logfire](https://pydantic.dev/logfire) to help you view and understand agent runs. Logfire comes included with `pydantic-ai` (but not the [“slim” version](https://pydantic.dev/docs/ai/overview/install/#slim-install) ), so you can typically start using it immediately by following the [Logfire setup docs](https://pydantic.dev/docs/ai/integrations/logfire/#using-logfire) . Running Examples ---------------- [](https://pydantic.dev/docs/ai/overview/install/#running-examples) We distribute the [`pydantic_ai_examples`](https://github.com/pydantic/pydantic-ai/tree/main/examples/pydantic_ai_examples) directory as a separate PyPI package ([`pydantic-ai-examples`](https://pypi.org/project/pydantic-ai-examples/) ) to make examples extremely easy to customize and run. To install examples, use the `examples` optional group: * [pip](https://pydantic.dev/docs/ai/overview/install/#tab-panel-152) * [uv](https://pydantic.dev/docs/ai/overview/install/#tab-panel-153) Terminal pip install "pydantic-ai[examples]" Terminal uv add "pydantic-ai[examples]" To run the examples, follow instructions in the [examples docs](https://pydantic.dev/docs/ai/examples/setup/) . Slim Install ------------ [](https://pydantic.dev/docs/ai/overview/install/#slim-install) If you know which model you’re going to use and want to avoid installing superfluous packages, you can use the [`pydantic-ai-slim`](https://pypi.org/project/pydantic-ai-slim/) package. For example, if you’re using just [`OpenAIChatModel`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIChatModel) , you would run: * [pip](https://pydantic.dev/docs/ai/overview/install/#tab-panel-154) * [uv](https://pydantic.dev/docs/ai/overview/install/#tab-panel-155) Terminal pip install "pydantic-ai-slim[openai]" Terminal uv add "pydantic-ai-slim[openai]" `pydantic-ai-slim` has the following optional groups: * `logfire` — installs [Pydantic Logfire](https://pydantic.dev/docs/ai/integrations/logfire/) dependency `logfire` [PyPI ↗](https://pypi.org/project/logfire) * `evals` — installs [Pydantic Evals](https://pydantic.dev/docs/ai/evals/evals/) dependency `pydantic-evals` [PyPI ↗](https://pypi.org/project/pydantic-evals) * `openai` — installs [OpenAI Model](https://pydantic.dev/docs/ai/models/openai/) dependency `openai` [PyPI ↗](https://pypi.org/project/openai) * `google` — installs [Google Model](https://pydantic.dev/docs/ai/models/google/) dependency `google-genai` [PyPI ↗](https://pypi.org/project/google-genai) * `anthropic` — installs [Anthropic Model](https://pydantic.dev/docs/ai/models/anthropic/) dependency `anthropic` [PyPI ↗](https://pypi.org/project/anthropic) * `groq` — installs [Groq Model](https://pydantic.dev/docs/ai/models/groq/) dependency `groq` [PyPI ↗](https://pypi.org/project/groq) * `mistral` — installs [Mistral Model](https://pydantic.dev/docs/ai/models/mistral/) dependency `mistralai` [PyPI ↗](https://pypi.org/project/mistralai) * `cohere` - installs [Cohere Model](https://pydantic.dev/docs/ai/models/cohere/) dependency `cohere` [PyPI ↗](https://pypi.org/project/cohere) * `bedrock` - installs [Bedrock Model](https://pydantic.dev/docs/ai/models/bedrock/) dependency `boto3` [PyPI ↗](https://pypi.org/project/boto3) * `bedrock-mantle` - installs [Bedrock Mantle Model](https://pydantic.dev/docs/ai/models/bedrock/#bedrock-mantle) dependencies `openai` [PyPI ↗](https://pypi.org/project/openai) and `botocore` [PyPI ↗](https://pypi.org/project/botocore) * `xai` - installs [xAI Model](https://pydantic.dev/docs/ai/models/xai/) dependency `xai-sdk` [PyPI ↗](https://pypi.org/project/xai-sdk) * `openrouter` - installs the [OpenRouter](https://pydantic.dev/docs/ai/models/openrouter/) dependency `openai` [PyPI ↗](https://pypi.org/project/openai) * `zai` - installs the [Z.AI](https://pydantic.dev/docs/ai/models/zai/) dependency `openai` [PyPI ↗](https://pypi.org/project/openai) * `snowflake` - installs the [Snowflake Cortex](https://pydantic.dev/docs/ai/models/snowflake/) dependency `openai` [PyPI ↗](https://pypi.org/project/openai) * `crusoe` - installs the [Crusoe](https://pydantic.dev/docs/ai/models/crusoe/) dependency `openai` [PyPI ↗](https://pypi.org/project/openai) * `cerebras` - installs the [Cerebras](https://pydantic.dev/docs/ai/models/cerebras/) dependency `openai` [PyPI ↗](https://pypi.org/project/openai) * `huggingface` - installs [Hugging Face Model](https://pydantic.dev/docs/ai/models/huggingface/) dependency `huggingface-hub` [PyPI ↗](https://pypi.org/project/huggingface-hub) * `sentence-transformers` - installs [Sentence Transformers Embedding Model](https://pydantic.dev/docs/ai/guides/embeddings/#sentence-transformers-local) dependency `sentence-transformers` [PyPI ↗](https://pypi.org/project/sentence-transformers) * `voyageai` - installs [VoyageAI Embedding Model](https://pydantic.dev/docs/ai/guides/embeddings/#voyageai) dependency `voyageai` [PyPI ↗](https://pypi.org/project/voyageai) * `duckduckgo` - installs [DuckDuckGo Search Tool](https://pydantic.dev/docs/ai/tools-toolsets/common-tools/#duckduckgo-search-tool) dependency `ddgs` [PyPI ↗](https://pypi.org/project/ddgs) * `tavily` - installs [Tavily Search Tool](https://pydantic.dev/docs/ai/tools-toolsets/common-tools/#tavily-search-tool) dependency `tavily-python` [PyPI ↗](https://pypi.org/project/tavily-python) * `exa` - installs [Exa Search Tool](https://pydantic.dev/docs/ai/tools-toolsets/common-tools/#exa-search-tool) dependency `exa-py` [PyPI ↗](https://pypi.org/project/exa-py) * `web-fetch` - installs [Web Fetch Tool](https://pydantic.dev/docs/ai/tools-toolsets/common-tools/#web-fetch-tool) dependency `markdownify` [PyPI ↗](https://pypi.org/project/markdownify) * `cli` - installs [CLI](https://pydantic.dev/docs/ai/integrations/cli/) dependencies `rich` [PyPI ↗](https://pypi.org/project/rich) , `prompt-toolkit` [PyPI ↗](https://pypi.org/project/prompt-toolkit) , and `argcomplete` [PyPI ↗](https://pypi.org/project/argcomplete) * `mcp` - installs [MCP](https://pydantic.dev/docs/ai/mcp/client/) dependency `fastmcp-slim[client]` [PyPI ↗](https://pypi.org/project/fastmcp-slim) * `ui` - installs [UI Event Streams](https://pydantic.dev/docs/ai/integrations/ui/overview/) dependency `starlette` [PyPI ↗](https://pypi.org/project/starlette) * `web` - installs [Web UI](https://pydantic.dev/docs/ai/integrations/ui/overview/) dependencies `starlette` [PyPI ↗](https://pypi.org/project/starlette) , `httpx` [PyPI ↗](https://pypi.org/project/httpx) , and `uvicorn` [PyPI ↗](https://pypi.org/project/uvicorn) * `ag-ui` - installs [AG-UI Event Stream Protocol](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/) dependencies `ag-ui-protocol` [PyPI ↗](https://pypi.org/project/ag-ui-protocol) and `starlette` [PyPI ↗](https://pypi.org/project/starlette) * `retries` - installs [HTTP Retries](https://pydantic.dev/docs/ai/models/http-request-retries/) dependency `tenacity` [PyPI ↗](https://pypi.org/project/tenacity) * `temporal` - installs [Temporal Durable Execution](https://pydantic.dev/docs/ai/capabilities/durable_execution/temporal/) dependency `temporalio` [PyPI ↗](https://pypi.org/project/temporalio) * `dbos` - installs [DBOS Durable Execution](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/) dependency `dbos` [PyPI ↗](https://pypi.org/project/dbos) * `prefect` - installs [Prefect Durable Execution](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/) dependency `prefect` [PyPI ↗](https://pypi.org/project/prefect) * `spec` - installs [AgentSpec](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) dependencies `pyyaml` [PyPI ↗](https://pypi.org/project/PyYAML) and `pydantic-handlebars` [PyPI ↗](https://pypi.org/project/pydantic-handlebars) You can also install dependencies for multiple models and use cases, for example: * [pip](https://pydantic.dev/docs/ai/overview/install/#tab-panel-156) * [uv](https://pydantic.dev/docs/ai/overview/install/#tab-panel-157) Terminal pip install "pydantic-ai-slim[openai,google,logfire]" Terminal uv add "pydantic-ai-slim[openai,google,logfire]" Was this page helpful? Thanks for your feedback! --- # Resolve Model ID | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/resolve-model-id/#_top) Resolve Model ID ================ [`ResolveModelId`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ResolveModelId) is a [capability](https://pydantic.dev/docs/ai/capabilities/overview/) that turns application-specific model IDs into [`Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) instances. The resolver can use run dependencies to look up tenant-specific providers, credentials, or model registries: resolve\_model\_id.py from dataclasses import dataclass from typing import Any from pydantic_ai import Agent, ModelResolutionContext from pydantic_ai.capabilities import ResolveModelId from pydantic_ai.models import Model, infer_model from pydantic_ai.providers import Provider, infer_provider from pydantic_ai.providers.openai import OpenAIProvider @dataclass class Deps: """Per-user provider credentials.""" openai_api_key: str def resolve_model(ctx: ModelResolutionContext[Deps], model_id: str) -> Model | None: """Resolve IDs in the `user:` namespace with the current user's credentials.""" if not model_id.startswith('user:'): return None def provider_factory(provider_name: str) -> Provider[Any]: if provider_name == 'openai': return OpenAIProvider(api_key=ctx.deps.openai_api_key) return infer_provider(provider_name) return infer_model(model_id.removeprefix('user:'), provider_factory) agent = Agent( 'user:openai:gpt-5.6-sol', deps_type=Deps, capabilities=[ResolveModelId(resolve_model)], ) The resolver may be synchronous or asynchronous. Its full callable signature is `(ModelResolutionContext[Deps], str) -> Model | None | Awaitable[Model | None]`. The convenience capability adapts both forms to the asynchronous [`resolve_model_id()`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability.resolve_model_id) hook. Resolvers form a chain in capability order: the first non-`None` result wins, and Pydantic AI falls back to normal model inference if every resolver returns `None`. See [Resolving model IDs](https://pydantic.dev/docs/ai/capabilities/custom/#resolving-model-ids) to implement the hook in a custom capability and understand when each resolver tree is used. Was this page helpful? Thanks for your feedback! --- # Tool Search | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/tool-search/#_top) Tool Search =========== The [`ToolSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ToolSearch) [capability](https://pydantic.dev/docs/ai/capabilities/overview/) handles model-driven discovery of searchable tools marked with `defer_loading=True`, so agents with large toolsets only pay tokens for the tools the model needs. Like the [provider-adaptive tools](https://pydantic.dev/docs/ai/capabilities/overview/#provider-adaptive-tools) above, it picks the best path for the active model — native server-executed search on Anthropic and OpenAI Responses, a local `search_tools` function tool elsewhere — and is auto-injected into every agent when searchable deferred tools exist. Bundle-level disclosure is covered by [on-demand capabilities](https://pydantic.dev/docs/ai/capabilities/on-demand/) . Pass an explicit [`ToolSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ToolSearch) to pick a specific [`strategy`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ToolSearch.strategy) (`'keywords'`, `'bm25'`, `'regex'`, or a custom callable) or tune the local fallback: tool\_search\_capability.py from pydantic_ai import Agent from pydantic_ai.capabilities import ToolSearch agent = Agent('anthropic:claude-sonnet-4-6', capabilities=[ToolSearch(strategy='keywords')]) When the local `search_tools` function tool is used, its retry budget follows the agent’s tool budget — so `Agent(retries={'tools': N})` gives the model `N` attempts to correct a malformed `queries` argument, on the same [precedence ladder](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#which-retry-limit-wins) as any other tool. A search that finds no matches returns normally and never spends a retry. See [Tool Search](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-search) for when to reach for it, the full strategy table, and provider support details. Was this page helpful? Thanks for your feedback! --- # Compaction | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/compaction/#_top) Compaction ========== As a conversation grows, its message history can approach the model’s context window. _Compaction_ keeps it in check by shrinking older messages — trimming, clearing, or summarizing them — while preserving recent context and tool-call integrity. Pydantic AI supports this at several levels, from provider-native APIs to model-agnostic history editing. Provider-native compaction -------------------------- [](https://pydantic.dev/docs/ai/capabilities/compaction/#provider-native-compaction) Some providers expose a built-in compaction API that runs on their side. Pydantic AI wraps these as [capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) : | Provider | Capability | Details | | --- | --- | --- | | OpenAI Responses API | [`OpenAICompaction`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAICompaction) | [OpenAI compaction](https://pydantic.dev/docs/ai/models/openai/#message-compaction) | | Anthropic | [`AnthropicCompaction`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicCompaction) | [Anthropic compaction](https://pydantic.dev/docs/ai/models/anthropic/#message-compaction) | Each uses the corresponding provider API, so it’s only available on that provider. Pydantic AI treats a compaction part as a visibility boundary for state that feeds future requests. Tool discoveries and on-demand capability loads before the boundary reset, so later requests advertise them again. Capability and toolset authors should apply the same rule to their own derived state: compute anything the model needs to have seen — announcements, disclosures, catalogs — from [`post_compaction_window`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.post_compaction_window) rather than remembering it in instance attributes, so it self-heals when compaction replaces the history that carried it. When dispatching a tool call, Pydantic AI uses the provider that served that response to determine whether the provider actually honored the boundary on the request wire. A foreign-provider compaction part, an OpenAI part without encrypted content, or an Anthropic part without summary content does not hide earlier callability evidence, because that provider sent the earlier history to the model. A boundary emitted inside the response containing the call is likewise too late to affect what the model saw for that response. If the response has no provider name, dispatch falls back to the provider-agnostic boundary. ### Client-held history [](https://pydantic.dev/docs/ai/capabilities/compaction/#client-held-history) [`CompactionPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.CompactionPart) s round-trip through the [UI adapters](https://pydantic.dev/docs/ai/integrations/ui/overview/) , whose protocols have the client transmit the full conversation history on each request, so compacted conversations keep working with such frontends. A client-submitted compaction item is honored — the conversation stays compacted — but it is never trusted to stand in for the agent’s [system prompt](https://pydantic.dev/docs/ai/core-concepts/agent/#system-prompts) : that still reaches the model on every request, as described under [Loading untrusted history](https://pydantic.dev/docs/ai/core-concepts/message-history/#loading-untrusted-history) . If a run also receives its own server-side history — the [server-side persistence pattern](https://pydantic.dev/docs/ai/integrations/ui/overview/#trust-model-for-client-submitted-messages) , where stored messages are passed as `message_history` and client messages only supply the latest turn — client-submitted compaction items are ignored instead. A compaction item marks a boundary before which nothing is sent to the model, so honoring one from the client would let it hide the server’s own stored history from the model and substitute its summary for that context. Client-submitted compaction items are only honored when the client-transmitted messages are the entire conversation. Even then, a client can replay any compaction item the server’s provider account has ever produced — opaque encrypted state on OpenAI, a plaintext summary on Anthropic. That is equivalent in kind to fabricating plain-text history, which client-transmitted history always permits (see [Trust boundary for client-supplied history](https://pydantic.dev/docs/ai/core-concepts/message-history/#trust-boundary-for-client-supplied-history) ), with one difference: the server cannot inspect what an opaque item contains. If that matters for your deployment, keep the history server-side: persist the full message list keyed by conversation, send the client only display data, and pass the stored messages as `message_history` on each run. Don’t trim the stored history around compaction boundaries yourself — each model adapter already omits what its own provider’s compaction replaces, while models from other providers, which ignore a foreign compaction item, still get the full earlier history they need. Model-agnostic compaction ------------------------- [](https://pydantic.dev/docs/ai/capabilities/compaction/#model-agnostic-compaction) To compact on any model, edit the message history yourself with a [history processor](https://pydantic.dev/docs/ai/core-concepts/message-history/#processing-message-history) wrapped as a [`ProcessHistory`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ProcessHistory) capability — this works with every provider. Common patterns: * [Keep only recent messages](https://pydantic.dev/docs/ai/core-concepts/message-history/#keep-only-recent-messages) — a zero-cost sliding window over the most recent turns. * [Summarize old messages](https://pydantic.dev/docs/ai/core-concepts/message-history/#summarize-old-messages) — use a (cheaper) model to condense older messages into a summary. Pydantic AI Harness ------------------- [](https://pydantic.dev/docs/ai/capabilities/compaction/#pydantic-ai-harness) [Pydantic AI Harness](https://pydantic.dev/docs/ai/harness/) packages a menu of ready-made, model-agnostic [compaction strategies](https://pydantic.dev/docs/ai/harness/compaction/) : mostly zero-LLM history editing — sliding-window trimming, clearing old tool results, deduplicating repeated file reads, clamping oversized message parts — plus LLM summarization for when that’s not enough, and a `TieredCompaction` orchestrator (the recommended default) that escalates from cheap to expensive strategies only as far as needed to fit the target. Was this page helpful? Thanks for your feedback! --- # Case Lifecycle Hooks | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/how-to/lifecycle/#_top) Case Lifecycle Hooks ==================== Control per-case setup, context preparation, and teardown during evaluation using [`CaseLifecycle`](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#pydantic_evals.lifecycle.CaseLifecycle) . [`CaseLifecycle`](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#pydantic_evals.lifecycle.CaseLifecycle) provides hooks at each stage of case evaluation. You pass a lifecycle **class** (not an instance) to [`Dataset.evaluate`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset.evaluate) , and a new instance is created for each case, so instance attributes naturally hold case-specific state. Evaluation Flow --------------- [](https://pydantic.dev/docs/ai/evals/how-to/lifecycle/#evaluation-flow) Each case follows this flow: 1. **`setup()`** — called before task execution 2. **Task runs** 3. **`prepare_context()`** — called after task, before evaluators 4. **Evaluators run** 5. **`teardown()`** — called after evaluators complete, or during cleanup if the case is interrupted Per-Case Setup and Teardown --------------------------- [](https://pydantic.dev/docs/ai/evals/how-to/lifecycle/#per-case-setup-and-teardown) Use `setup()` and `teardown()` when each case needs its own environment — for example, creating a database, starting a service, or preparing fixtures driven by case metadata. Since a new lifecycle instance is created for each case, instance attributes are naturally case-scoped: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators.context import EvaluatorContext from pydantic_evals.lifecycle import CaseLifecycle from pydantic_evals.reporting import ReportCase, ReportCaseFailure class SetupFromMetadata(CaseLifecycle[str, str, dict]): async def setup(self) -> None: prefix = (self.case.metadata or {}).get('prefix', '') self.prefix = prefix async def prepare_context( self, ctx: EvaluatorContext[str, str, dict] ) -> EvaluatorContext[str, str, dict]: ctx.metrics['prefix_length'] = len(self.prefix) return ctx async def teardown( self, result: ReportCase[str, str, dict] | ReportCaseFailure[str, str, dict] | None, ) -> None: pass # Clean up resources here dataset = Dataset( name='setup_teardown', cases=[\ Case(name='no_prefix', inputs='hello', metadata={'prefix': ''}),\ Case(name='with_prefix', inputs='hello', metadata={'prefix': 'PREFIX:'}),\ ] ) report = dataset.evaluate_sync(lambda inputs: inputs.upper(), lifecycle=SetupFromMetadata) metrics = {c.name: c.metrics for c in report.cases} print(metrics['no_prefix']['prefix_length']) #> 0 print(metrics['with_prefix']['prefix_length']) #> 7 The case metadata drives per-case behavior without needing custom [`Case`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Case) subclasses or serialization. ### Conditional Teardown [](https://pydantic.dev/docs/ai/evals/how-to/lifecycle/#conditional-teardown) The `teardown()` hook receives the full result, so you can vary cleanup logic based on success or failure — for example, keeping test environments up for manual inspection when a case fails. The `result` can be `None` if evaluation is interrupted before the case produces a report result, so handle that branch when your cleanup depends on the case outcome: from pydantic_evals import Case, Dataset from pydantic_evals.lifecycle import CaseLifecycle from pydantic_evals.reporting import ReportCase, ReportCaseFailure cleaned_up: list[str] = [] class ConditionalCleanup(CaseLifecycle[str, str, dict]): async def setup(self) -> None: self.resource_id = self.case.name async def teardown( self, result: ReportCase[str, str, dict] | ReportCaseFailure[str, str, dict] | None, ) -> None: keep_on_failure = (self.case.metadata or {}).get('keep_on_failure', False) if result is None: # abnormal exit cleaned_up.append(self.resource_id) elif isinstance(result, ReportCaseFailure) and keep_on_failure: # case failed pass # Keep resource for inspection else: # case succeeded cleaned_up.append(self.resource_id) dataset = Dataset( name='conditional_cleanup', cases=[\ Case(name='success_case', inputs='hello', metadata={'keep_on_failure': True}),\ Case(name='failure_case', inputs='fail', metadata={'keep_on_failure': True}),\ ] ) def task(inputs: str) -> str: if inputs == 'fail': raise ValueError('intentional failure') return inputs.upper() report = dataset.evaluate_sync(task, max_concurrency=1, lifecycle=ConditionalCleanup) print(cleaned_up) #> ['success_case'] Preparing Evaluator Context --------------------------- [](https://pydantic.dev/docs/ai/evals/how-to/lifecycle/#preparing-evaluator-context) The `prepare_context()` hook runs after the task completes but before evaluators see the context. This can be used to add metrics or attributes based on the task output, span tree, or any other state — for example, deriving metrics from instrumented spans (like tool call counts or API latency), or computing values from external resources set up during `setup()`: from dataclasses import dataclass from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import Evaluator, EvaluatorContext from pydantic_evals.lifecycle import CaseLifecycle class EnrichMetrics(CaseLifecycle): async def prepare_context(self, ctx: EvaluatorContext) -> EvaluatorContext: ctx.metrics['output_length'] = len(str(ctx.output)) return ctx @dataclass class CheckLength(Evaluator): max_length: int = 50 def evaluate(self, ctx: EvaluatorContext) -> bool: return ctx.metrics.get('output_length', 0) <= self.max_length dataset = Dataset( name='context_enrichment', cases=[Case(name='short', inputs='hi'), Case(name='long', inputs='hello world')], evaluators=[CheckLength()], ) report = dataset.evaluate_sync(lambda inputs: inputs.upper(), lifecycle=EnrichMetrics) for case in report.cases: print(f'{case.name}: output_length={case.metrics["output_length"]}') #> short: output_length=2 #> long: output_length=11 Type Parameters --------------- [](https://pydantic.dev/docs/ai/evals/how-to/lifecycle/#type-parameters) [`CaseLifecycle`](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#pydantic_evals.lifecycle.CaseLifecycle) is generic over the same three type parameters as [`Case`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Case) : `InputsT`, `OutputT`, and `MetadataT`. All three default to `Any`, so you can omit them when your hooks don’t need type-specific access: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators.context import EvaluatorContext from pydantic_evals.lifecycle import CaseLifecycle # Works with any dataset — no type parameters needed class GenericMetricEnricher(CaseLifecycle): async def prepare_context(self, ctx: EvaluatorContext) -> EvaluatorContext: ctx.metrics['custom'] = 42 return ctx dataset = Dataset(name='generic_lifecycle', cases=[Case(inputs='test')]) report = dataset.evaluate_sync(lambda inputs: inputs, lifecycle=GenericMetricEnricher) print(report.cases[0].metrics['custom']) #> 42 Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/how-to/lifecycle/#next-steps) * **[Metrics & Attributes](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/) ** — Recording metrics inside tasks * **[Custom Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/custom/) ** — Using enriched metrics in evaluators * **[Span-Based Evaluation](https://pydantic.dev/docs/ai/evals/evaluators/span-based/) ** — Analyzing execution traces Was this page helpful? Thanks for your feedback! --- # Overview | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/durable_execution/overview/#_top) Overview ======== Pydantic AI allows you to build durable agents that can preserve their progress across transient API failures and application errors or restarts, and handle long-running, asynchronous, and human-in-the-loop workflows with production-grade reliability. Durable agents have full support for [streaming](https://pydantic.dev/docs/ai/core-concepts/agent/#streaming-all-events) and [MCP](https://pydantic.dev/docs/ai/mcp/client/) , with the added benefit of fault tolerance. Pydantic AI officially supports four durable execution solutions: * [Temporal](https://pydantic.dev/docs/ai/capabilities/durable_execution/temporal/) * [DBOS](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/) * [Prefect](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/) * [Restate](https://pydantic.dev/docs/ai/capabilities/durable_execution/restate/) These integrations are co-maintained by the Pydantic and vendor teams. The Temporal, DBOS, and Prefect integrations ship with Pydantic AI as [capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) you attach to an agent; the [Restate](https://pydantic.dev/docs/ai/capabilities/durable_execution/restate/) integration lives in the Restate SDK and builds only on Pydantic AI’s public interface, so it can also serve as a reference for integrating with other durable systems. Additional external SDK integrations: * [Kitaru](https://pydantic.dev/docs/ai/capabilities/durable_execution/kitaru/) * [Apache Airflow](https://pydantic.dev/docs/ai/capabilities/durable_execution/airflow/) Was this page helpful? Thanks for your feedback! --- # Decisions | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/graph/builder/decisions/#_top) Decisions ========= Decision nodes enable conditional branching in your graph based on the type or value of data flowing through it. A decision node evaluates incoming data and routes it to different branches based on: * Type matching (using `isinstance`) * Literal value matching * Custom predicate functions The first matching branch is taken, similar to pattern matching or `if-elif-else` chains. Creating Decisions ------------------ [](https://pydantic.dev/docs/ai/graph/builder/decisions/#creating-decisions) Use [`g.decision()`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder.decision) to create a decision node, then add branches with [`g.match()`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder.match) : simple\_decision.py from dataclasses import dataclass from typing import Literal from pydantic_graph import GraphBuilder, StepContext, TypeExpression @dataclass class DecisionState: path_taken: str | None = None async def main(): g = GraphBuilder(state_type=DecisionState, output_type=str) @g.step async def choose_path(ctx: StepContext[DecisionState, None, None]) -> Literal['left', 'right']: return 'left' @g.step async def left_path(ctx: StepContext[DecisionState, None, object]) -> str: ctx.state.path_taken = 'left' return 'Went left' @g.step async def right_path(ctx: StepContext[DecisionState, None, object]) -> str: ctx.state.path_taken = 'right' return 'Went right' g.add( g.edge_from(g.start_node).to(choose_path), g.edge_from(choose_path).to( g.decision() .branch(g.match(TypeExpression[Literal['left']]).to(left_path)) .branch(g.match(TypeExpression[Literal['right']]).to(right_path)) ), g.edge_from(left_path, right_path).to(g.end_node), ) graph = g.build() state = DecisionState() result = await graph.run(state=state) print(result) #> Went left print(state.path_taken) #> left _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Type Matching ------------- [](https://pydantic.dev/docs/ai/graph/builder/decisions/#type-matching) Match by type using regular Python types: type\_matching.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext @dataclass class DecisionState: pass async def main(): g = GraphBuilder(state_type=DecisionState, output_type=str) @g.step async def return_int(ctx: StepContext[DecisionState, None, None]) -> int: return 42 @g.step async def handle_int(ctx: StepContext[DecisionState, None, int]) -> str: return f'Got int: {ctx.inputs}' @g.step async def handle_str(ctx: StepContext[DecisionState, None, str]) -> str: return f'Got str: {ctx.inputs}' g.add( g.edge_from(g.start_node).to(return_int), g.edge_from(return_int).to( g.decision() .branch(g.match(int).to(handle_int)) .branch(g.match(str).to(handle_str)) ), g.edge_from(handle_int, handle_str).to(g.end_node), ) graph = g.build() result = await graph.run(state=DecisionState()) print(result) #> Got int: 42 _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ ### Matching Union Types [](https://pydantic.dev/docs/ai/graph/builder/decisions/#matching-union-types) For more complex type expressions like unions, you need to use [`TypeExpression`](https://pydantic.dev/docs/ai/api/pydantic_graph/util/#pydantic_graph.util.TypeExpression) because Python’s type system doesn’t allow union types to be used directly as runtime values: union\_type\_matching.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, TypeExpression @dataclass class DecisionState: pass async def main(): g = GraphBuilder(state_type=DecisionState, output_type=str) @g.step async def return_value(ctx: StepContext[DecisionState, None, None]) -> int | str: """Returns either an int or a str.""" return 42 @g.step async def handle_number(ctx: StepContext[DecisionState, None, int | float]) -> str: return f'Got number: {ctx.inputs}' @g.step async def handle_text(ctx: StepContext[DecisionState, None, str]) -> str: return f'Got text: {ctx.inputs}' g.add( g.edge_from(g.start_node).to(return_value), g.edge_from(return_value).to( g.decision() # Use TypeExpression for union types .branch(g.match(TypeExpression[int | float]).to(handle_number)) .branch(g.match(str).to(handle_text)) ), g.edge_from(handle_number, handle_text).to(g.end_node), ) graph = g.build() result = await graph.run(state=DecisionState()) print(result) #> Got number: 42 _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Custom Matchers --------------- [](https://pydantic.dev/docs/ai/graph/builder/decisions/#custom-matchers) Provide custom matching logic with the `matches` parameter: custom\_matcher.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, TypeExpression @dataclass class DecisionState: pass async def main(): g = GraphBuilder(state_type=DecisionState, output_type=str) @g.step async def return_number(ctx: StepContext[DecisionState, None, None]) -> int: return 7 @g.step async def even_path(ctx: StepContext[DecisionState, None, int]) -> str: return f'{ctx.inputs} is even' @g.step async def odd_path(ctx: StepContext[DecisionState, None, int]) -> str: return f'{ctx.inputs} is odd' g.add( g.edge_from(g.start_node).to(return_number), g.edge_from(return_number).to( g.decision() .branch(g.match(TypeExpression[int], matches=lambda x: x % 2 == 0).to(even_path)) .branch(g.match(TypeExpression[int], matches=lambda x: x % 2 == 1).to(odd_path)) ), g.edge_from(even_path, odd_path).to(g.end_node), ) graph = g.build() result = await graph.run(state=DecisionState()) print(result) #> 7 is odd _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Branch Priority --------------- [](https://pydantic.dev/docs/ai/graph/builder/decisions/#branch-priority) Branches are evaluated in the order they’re added. The first matching branch is taken: branch\_priority.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, TypeExpression @dataclass class DecisionState: pass async def main(): g = GraphBuilder(state_type=DecisionState, output_type=str) @g.step async def return_value(ctx: StepContext[DecisionState, None, None]) -> int: return 10 @g.step async def branch_a(ctx: StepContext[DecisionState, None, int]) -> str: return 'Branch A' @g.step async def branch_b(ctx: StepContext[DecisionState, None, int]) -> str: return 'Branch B' g.add( g.edge_from(g.start_node).to(return_value), g.edge_from(return_value).to( g.decision() .branch(g.match(TypeExpression[int], matches=lambda x: x >= 5).to(branch_a)) .branch(g.match(TypeExpression[int], matches=lambda x: x >= 0).to(branch_b)) ), g.edge_from(branch_a, branch_b).to(g.end_node), ) graph = g.build() result = await graph.run(state=DecisionState()) print(result) #> Branch A _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Both branches could match `10`, but Branch A is first, so it’s taken. Catch-All Branches ------------------ [](https://pydantic.dev/docs/ai/graph/builder/decisions/#catch-all-branches) Use `object` or `Any` to create a catch-all branch: catch\_all.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, TypeExpression @dataclass class DecisionState: pass async def main(): g = GraphBuilder(state_type=DecisionState, output_type=str) @g.step async def return_value(ctx: StepContext[DecisionState, None, None]) -> int: return 100 @g.step async def catch_all(ctx: StepContext[DecisionState, None, object]) -> str: return f'Caught: {ctx.inputs}' g.add( g.edge_from(g.start_node).to(return_value), g.edge_from(return_value).to(g.decision().branch(g.match(TypeExpression[object]).to(catch_all))), g.edge_from(catch_all).to(g.end_node), ) graph = g.build() result = await graph.run(state=DecisionState()) print(result) #> Caught: 100 _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Nested Decisions ---------------- [](https://pydantic.dev/docs/ai/graph/builder/decisions/#nested-decisions) Decisions can be nested for complex conditional logic: nested\_decisions.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, TypeExpression @dataclass class DecisionState: pass async def main(): g = GraphBuilder(state_type=DecisionState, output_type=str) @g.step async def get_number(ctx: StepContext[DecisionState, None, None]) -> int: return 15 @g.step async def is_positive(ctx: StepContext[DecisionState, None, int]) -> int: return ctx.inputs @g.step async def is_negative(ctx: StepContext[DecisionState, None, int]) -> str: return 'Negative' @g.step async def small_positive(ctx: StepContext[DecisionState, None, int]) -> str: return 'Small positive' @g.step async def large_positive(ctx: StepContext[DecisionState, None, int]) -> str: return 'Large positive' g.add( g.edge_from(g.start_node).to(get_number), g.edge_from(get_number).to( g.decision() .branch(g.match(TypeExpression[int], matches=lambda x: x > 0).to(is_positive)) .branch(g.match(TypeExpression[int], matches=lambda x: x <= 0).to(is_negative)) ), g.edge_from(is_positive).to( g.decision() .branch(g.match(TypeExpression[int], matches=lambda x: x < 10).to(small_positive)) .branch(g.match(TypeExpression[int], matches=lambda x: x >= 10).to(large_positive)) ), g.edge_from(is_negative, small_positive, large_positive).to(g.end_node), ) graph = g.build() result = await graph.run(state=DecisionState()) print(result) #> Large positive _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Branching with Labels --------------------- [](https://pydantic.dev/docs/ai/graph/builder/decisions/#branching-with-labels) Add labels to branches for documentation and diagram generation: labeled\_branches.py from dataclasses import dataclass from typing import Literal from pydantic_graph import GraphBuilder, StepContext, TypeExpression @dataclass class DecisionState: pass async def main(): g = GraphBuilder(state_type=DecisionState, output_type=str) @g.step async def choose(ctx: StepContext[DecisionState, None, None]) -> Literal['a', 'b']: return 'a' @g.step async def path_a(ctx: StepContext[DecisionState, None, object]) -> str: return 'Path A' @g.step async def path_b(ctx: StepContext[DecisionState, None, object]) -> str: return 'Path B' g.add( g.edge_from(g.start_node).to(choose), g.edge_from(choose).to( g.decision() .branch(g.match(TypeExpression[Literal['a']]).label('Take path A').to(path_a)) .branch(g.match(TypeExpression[Literal['b']]).label('Take path B').to(path_b)) ), g.edge_from(path_a, path_b).to(g.end_node), ) graph = g.build() result = await graph.run(state=DecisionState()) print(result) #> Path A _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Next Steps ---------- [](https://pydantic.dev/docs/ai/graph/builder/decisions/#next-steps) * Learn about [parallel execution](https://pydantic.dev/docs/ai/graph/builder/parallel/) with broadcasting and mapping * Understand [join nodes](https://pydantic.dev/docs/ai/graph/builder/joins/) for aggregating parallel results * See the [API reference](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision) for complete decision documentation Was this page helpful? Thanks for your feedback! --- # Pydantic Model | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/getting-started/pydantic-model/#_top) Pydantic Model ============== Simple example of using Pydantic AI to construct a Pydantic model from a text input. Demonstrates: * [structured `output_type`](https://pydantic.dev/docs/ai/core-concepts/output/#structured-output) Running the Example ------------------- [](https://pydantic.dev/docs/ai/examples/getting-started/pydantic-model/#running-the-example) With [dependencies installed and environment variables set](https://pydantic.dev/docs/ai/examples/setup/#usage) , run: * [pip](https://pydantic.dev/docs/ai/examples/getting-started/pydantic-model/#tab-panel-42) * [uv](https://pydantic.dev/docs/ai/examples/getting-started/pydantic-model/#tab-panel-43) Terminal python -m pydantic_ai_examples.pydantic_model Terminal uv run -m pydantic_ai_examples.pydantic_model This examples uses `openai:gpt-5` by default, but it works well with other models, e.g. you can run it with Gemini using: * [pip](https://pydantic.dev/docs/ai/examples/getting-started/pydantic-model/#tab-panel-44) * [uv](https://pydantic.dev/docs/ai/examples/getting-started/pydantic-model/#tab-panel-45) Terminal PYDANTIC_AI_MODEL=gemini-3-pro-preview python -m pydantic_ai_examples.pydantic_model Terminal PYDANTIC_AI_MODEL=gemini-3-pro-preview uv run -m pydantic_ai_examples.pydantic_model (or `PYDANTIC_AI_MODEL=gemini-3-flash-preview ...`) Example Code ------------ [](https://pydantic.dev/docs/ai/examples/getting-started/pydantic-model/#example-code) pydantic\_model.py import os import logfire from pydantic import BaseModel from pydantic_ai import Agent # 'if-token-present' means nothing will be sent (and the example will work) if you don't have logfire configured logfire.configure(send_to_logfire='if-token-present') logfire.instrument_pydantic_ai() class MyModel(BaseModel): city: str country: str model = os.getenv('PYDANTIC_AI_MODEL', 'openai:gpt-5.2') print(f'Using model: {model}') agent = Agent(model, output_type=MyModel) if __name__ == '__main__': result = agent.run_sync('The windy city in the US of A.') print(result.output) print(result.usage) Was this page helpful? Thanks for your feedback! --- # Coding Agent Skills | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/overview/coding-agent-skills/#_top) Coding Agent Skills =================== If you’re building Pydantic AI applications with a coding agent, you can install the Pydantic AI skill from the [`pydantic/skills`](https://github.com/pydantic/skills) repository to give your agent up-to-date framework knowledge. [Agent skills](https://agentskills.io/) are packages of instructions and reference material that coding agents load on demand. With the skill installed, coding agents have access to Pydantic AI patterns, architecture guidance, and common task references covering [tools](https://pydantic.dev/docs/ai/tools-toolsets/tools/) , [capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) , [structured output](https://pydantic.dev/docs/ai/core-concepts/output/) , [streaming](https://pydantic.dev/docs/ai/core-concepts/agent/#streaming-events-and-final-output) , [testing](https://pydantic.dev/docs/ai/guides/testing/) , [multi-agent delegation](https://pydantic.dev/docs/ai/guides/multi-agent-applications/) , [hooks](https://pydantic.dev/docs/ai/core-concepts/hooks/) , and [agent specs](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) . Installation ------------ [](https://pydantic.dev/docs/ai/overview/coding-agent-skills/#installation) ### Claude Code [](https://pydantic.dev/docs/ai/overview/coding-agent-skills/#claude-code) Install the [official Pydantic AI plugin](https://claude.com/plugins/pydantic-ai) from the Anthropic marketplace, which is available by default: Terminal claude plugin install pydantic-ai@claude-plugins-official As an alternative, you can install from the [`pydantic/skills`](https://github.com/pydantic/skills) marketplace, which bundles the Pydantic AI skill alongside other Pydantic-maintained skills: Terminal claude plugin marketplace add pydantic/skills claude plugin install ai@pydantic-skills ### Cross-Agent (agentskills.io) [](https://pydantic.dev/docs/ai/overview/coding-agent-skills/#cross-agent-agentskillsio) Install the Pydantic AI skill using the [skills CLI](https://github.com/vercel-labs/skills) : Terminal npx skills add pydantic/skills This works with 30+ agents via the [agentskills.io](https://agentskills.io/) standard, including Claude Code, Codex, Cursor, and Gemini CLI. ### Library Skills [](https://pydantic.dev/docs/ai/overview/coding-agent-skills/#library-skills) Pydantic AI also ships its skill bundled with the package, so you can install it directly from your project’s dependencies via [library-skills.io](https://library-skills.io/) : Terminal uvx library-skills --all The `--all` flag is required because the skill is bundled in `pydantic-ai-slim`, which is a transitive dependency of the `pydantic-ai` meta-package. Without it, `library-skills` only scans direct dependencies and won’t discover the skill. Add `--claude` to also install into `.claude/skills/` alongside the default `.agents/skills/` directory, since Claude Code doesn’t read from `.agents/`. See Also -------- [](https://pydantic.dev/docs/ai/overview/coding-agent-skills/#see-also) * [`pydantic/skills`](https://github.com/pydantic/skills) : source repository * [agentskills.io](https://agentskills.io/) : the open standard for agent skills * [library-skills.io](https://library-skills.io/) : install agent skills bundled with your project’s dependencies Was this page helpful? Thanks for your feedback! --- # Troubleshooting | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/overview/troubleshooting/#_top) Troubleshooting =============== Below are suggestions on how to fix some common errors you might encounter while using Pydantic AI. If the issue you’re experiencing is not listed below or addressed in the documentation, please feel free to ask in the [Pydantic Slack](https://pydantic.dev/docs/ai/overview/help/) or create an issue on [GitHub](https://github.com/pydantic/pydantic-ai/issues) . Jupyter Notebook Errors ----------------------- [](https://pydantic.dev/docs/ai/overview/troubleshooting/#jupyter-notebook-errors) ### `RuntimeError: This event loop is already running` [](https://pydantic.dev/docs/ai/overview/troubleshooting/#runtimeerror-this-event-loop-is-already-running) **Modern Jupyter/IPython (7.0+)**: This environment supports top-level `await` natively. You can use `Agent.run()` directly in notebook cells without additional setup: from pydantic_ai import Agent agent = Agent('openai:gpt-5.2') result = await agent.run('Who let the dogs out?') **Legacy environments or specific integrations**: If you encounter event loop conflicts, use [`nest-asyncio`](https://pypi.org/project/nest-asyncio/) : import nest_asyncio from pydantic_ai import Agent nest_asyncio.apply() agent = Agent('openai:gpt-5.2') result = agent.run_sync('Who let the dogs out?') **Note**: This also applies to Google Colab and [Marimo](https://github.com/marimo-team/marimo) environments. `RuntimeError: Event loop is closed` ------------------------------------ [](https://pydantic.dev/docs/ai/overview/troubleshooting/#runtimeerror-event-loop-is-closed) Synchronous methods like [`Agent.run_sync()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.run_sync) reuse the thread’s current event loop, and install a fresh one if other code closed it. If this error is raised from inside `httpx` or `httpcore` during a model request, the agent was already used before its event loop was closed: the provider’s HTTP connection pool still holds connections bound to the dead loop. Recreate the agent together with its model and provider (or pass a fresh `http_client` to the provider); reusing an existing `Model` instance keeps the dead connection pool. Avoid closing an event loop that other code is still using. API Key Configuration --------------------- [](https://pydantic.dev/docs/ai/overview/troubleshooting/#api-key-configuration) ### [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) : Set the `[PROVIDER]_API_KEY` environment variable or pass it via the provider’s `api_key=...` argument [](https://pydantic.dev/docs/ai/overview/troubleshooting/#usererror-set-the-provider_api_key-environment-variable-or-pass-it-via-the-providers-api_key-argument) If you’re running into issues with setting the API key for your model, visit the [Models](https://pydantic.dev/docs/ai/models/overview/) page to learn more about how to set an environment variable and/or pass in an `api_key` argument. To try Pydantic AI without an API key, use the built-in [`'test'` model](https://pydantic.dev/docs/ai/guides/testing/#unit-testing-with-testmodel) : [`Agent('test')`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent) . Monitoring HTTPX Requests ------------------------- [](https://pydantic.dev/docs/ai/overview/troubleshooting/#monitoring-httpx-requests) You can use custom `httpx` clients in your models in order to access specific requests, responses, and headers at runtime. It’s particularly helpful to use `logfire`’s [HTTPX integration](https://pydantic.dev/docs/ai/integrations/logfire/#monitoring-http-requests) to monitor the above. Was this page helpful? Thanks for your feedback! --- # snowflake | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/snowflake/#_top) snowflake ========= Setup ----- [](https://pydantic.dev/docs/ai/api/models/snowflake/#setup) For details on how to set up authentication with this model, see [model configuration for Snowflake Cortex](https://pydantic.dev/docs/ai/models/snowflake/) . Snowflake Cortex model implementation using Snowflake’s OpenAI-compatible Chat Completions API. SnowflakeModel -------------- [](https://pydantic.dev/docs/ai/api/models/snowflake/#pydantic_ai.models.snowflake.SnowflakeModel) **Bases:** `OpenAIChatModel` A model that uses Snowflake Cortex’s OpenAI-compatible Chat Completions API. Snowflake Cortex serves Claude, GPT, Llama, Mistral, DeepSeek, and Snowflake’s own models, with all inference running inside the customer’s Snowflake account. Apart from `__init__`, all methods are private or match those of the base class. ### Methods [](https://pydantic.dev/docs/ai/api/models/snowflake/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/snowflake/#pydantic_ai.models.snowflake.SnowflakeModel.__init__) def __init__( model_name: SnowflakeModelName, *, provider: Literal['snowflake'] | Provider[AsyncOpenAI] = 'snowflake', profile: ModelProfileSpec | None = None, settings: SnowflakeModelSettings | None = None, ) Initialize a Snowflake Cortex model. ##### Parameters [](https://pydantic.dev/docs/ai/api/models/snowflake/#parameters) **`model_name`** : `SnowflakeModelName` [](https://pydantic.dev/docs/ai/api/models/snowflake/#pydantic_ai.models.snowflake.SnowflakeModel.__init__(model_name)) The name of the Snowflake Cortex model to use. **`provider`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘snowflake’\] | `Provider`\[`AsyncOpenAI`\] _Default:_ `'snowflake'` [](https://pydantic.dev/docs/ai/api/models/snowflake/#pydantic_ai.models.snowflake.SnowflakeModel.__init__(provider)) The provider to use. Defaults to ‘snowflake’. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/snowflake/#pydantic_ai.models.snowflake.SnowflakeModel.__init__(profile)) The model profile to use. Defaults to a profile based on the model name. **`settings`** : `SnowflakeModelSettings` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/snowflake/#pydantic_ai.models.snowflake.SnowflakeModel.__init__(settings)) Model-specific settings that will be used as defaults for this model. SnowflakeModelSettings ---------------------- [](https://pydantic.dev/docs/ai/api/models/snowflake/#pydantic_ai.models.snowflake.SnowflakeModelSettings) **Bases:** [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) Settings used for a Snowflake Cortex model request. ALL FIELDS MUST BE `snowflake_` PREFIXED SO YOU CAN MERGE THEM WITH OTHER MODELS. ### Attributes [](https://pydantic.dev/docs/ai/api/models/snowflake/#attributes) #### snowflake\_reasoning [](https://pydantic.dev/docs/ai/api/models/snowflake/#pydantic_ai.models.snowflake.SnowflakeModelSettings.snowflake_reasoning) Configure reasoning tokens for Claude models. Defaults to an effort level based on the unified `thinking` setting. **Type:** `SnowflakeReasoning` SnowflakeReasoning ------------------ [](https://pydantic.dev/docs/ai/api/models/snowflake/#pydantic_ai.models.snowflake.SnowflakeReasoning) **Bases:** [`TypedDict`](https://docs.python.org/3/library/typing.html#typing.TypedDict) Configuration for reasoning tokens in Snowflake Cortex requests to Claude models. See [https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-rest-api](https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-rest-api) for details. ### Attributes [](https://pydantic.dev/docs/ai/api/models/snowflake/#attributes-1) #### effort [](https://pydantic.dev/docs/ai/api/models/snowflake/#pydantic_ai.models.snowflake.SnowflakeReasoning.effort) Reasoning effort level. Converted to a reasoning token budget by Cortex. Cannot be used with `max_tokens`. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘high’, ‘medium’, ‘low’\] #### max\_tokens [](https://pydantic.dev/docs/ai/api/models/snowflake/#pydantic_ai.models.snowflake.SnowflakeReasoning.max_tokens) Specific token limit for reasoning. Cannot be used with `effort`. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) SnowflakeStreamedResponse ------------------------- [](https://pydantic.dev/docs/ai/api/models/snowflake/#pydantic_ai.models.snowflake.SnowflakeStreamedResponse) **Bases:** `OpenAIStreamedResponse` Implementation of `StreamedResponse` for Snowflake Cortex models. SnowflakeModelName ------------------ [](https://pydantic.dev/docs/ai/api/models/snowflake/#pydantic_ai.models.snowflake.SnowflakeModelName) Possible Snowflake Cortex model names. Since Snowflake Cortex serves a variety of models and the list changes frequently, we explicitly list known models but allow any name in the type hints. Fine-tuned models can be referenced as `database.schema.model`. See [https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-rest-api](https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-rest-api) for an up to date list of models. **Default:** `str | LatestSnowflakeModelNames` Was this page helpful? Thanks for your feedback! --- # Handle Deferred Tool Calls | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/handle-deferred-tool-calls/#_top) Handle Deferred Tool Calls ========================== [`HandleDeferredToolCalls`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.HandleDeferredToolCalls) is a [capability](https://pydantic.dev/docs/ai/capabilities/overview/) that resolves [deferred tool calls](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/) inline during an agent run. When tools require approval or external execution, the agent normally pauses and returns [`DeferredToolRequests`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.DeferredToolRequests) as output; this capability intercepts those calls, invokes your handler to resolve them, and continues the run automatically: handle\_deferred\_tool\_calls.py from pydantic_ai import Agent, RunContext from pydantic_ai.capabilities import HandleDeferredToolCalls from pydantic_ai.tools import DeferredToolRequests, DeferredToolResults async def handle_deferred(ctx: RunContext, requests: DeferredToolRequests) -> DeferredToolResults: return requests.build_results(approve_all=True) # (1) agent = Agent('openai:gpt-5.2', capabilities=[HandleDeferredToolCalls(handle_deferred)]) Auto-approve every call that's waiting on approval. Real handlers typically inspect `requests.approvals` and `requests.calls` and decide per call — prompt an operator, check a policy, or execute an external call. The handler may be sync or async. It returns [`DeferredToolResults`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.DeferredToolResults) with results for some or all pending calls, or `None` to decline — in which case the next `HandleDeferredToolCalls` capability in the chain gets a chance, and unhandled calls bubble up as `DeferredToolRequests` output as usual. See [Resolving deferred calls with a handler](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/#resolving-deferred-calls-with-a-handler) for how this fits into the wider deferred-tools flow, including human-in-the-loop approval and external tool execution. Was this page helpful? Thanks for your feedback! --- # Bank Support | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/conversational-agents/bank-support/#_top) Bank Support ============ Small but complete example of using Pydantic AI to build a support agent for a bank. Demonstrates: * [dynamic system prompt](https://pydantic.dev/docs/ai/core-concepts/agent/#system-prompts) * [structured `output_type`](https://pydantic.dev/docs/ai/core-concepts/output/#structured-output) * [tools](https://pydantic.dev/docs/ai/tools-toolsets/tools/) Running the Example ------------------- [](https://pydantic.dev/docs/ai/examples/conversational-agents/bank-support/#running-the-example) With [dependencies installed and environment variables set](https://pydantic.dev/docs/ai/examples/setup/#usage) , run: * [pip](https://pydantic.dev/docs/ai/examples/conversational-agents/bank-support/#tab-panel-26) * [uv](https://pydantic.dev/docs/ai/examples/conversational-agents/bank-support/#tab-panel-27) Terminal python -m pydantic_ai_examples.bank_support Terminal uv run -m pydantic_ai_examples.bank_support (or `PYDANTIC_AI_MODEL=gemini-3-flash-preview ...`) Example Code ------------ [](https://pydantic.dev/docs/ai/examples/conversational-agents/bank-support/#example-code) bank\_support.py import sqlite3 from dataclasses import dataclass from pydantic import BaseModel from pydantic_ai import Agent, RunContext @dataclass class DatabaseConn: """A wrapper over the SQLite connection.""" sqlite_conn: sqlite3.Connection async def customer_name(self, *, id: int) -> str | None: res = cur.execute('SELECT name FROM customers WHERE id=?', (id,)) row = res.fetchone() if row: return row[0] return None async def customer_balance(self, *, id: int) -> float: res = cur.execute('SELECT balance FROM customers WHERE id=?', (id,)) row = res.fetchone() if row: return row[0] else: raise ValueError('Customer not found') @dataclass class SupportDependencies: customer_id: int db: DatabaseConn class SupportOutput(BaseModel): support_advice: str """Advice returned to the customer""" block_card: bool """Whether to block their card or not""" risk: int """Risk level of query""" support_agent = Agent( 'openai:gpt-5.2', deps_type=SupportDependencies, output_type=SupportOutput, instructions=( 'You are a support agent in our bank, give the ' 'customer support and judge the risk level of their query. ' "Reply using the customer's name." ), ) @support_agent.instructions async def add_customer_name(ctx: RunContext[SupportDependencies]) -> str: customer_name = await ctx.deps.db.customer_name(id=ctx.deps.customer_id) return f"The customer's name is {customer_name!r}" @support_agent.tool async def customer_balance(ctx: RunContext[SupportDependencies]) -> str: """Returns the customer's current account balance.""" balance = await ctx.deps.db.customer_balance( id=ctx.deps.customer_id, ) return f'${balance:.2f}' if __name__ == '__main__': with sqlite3.connect(':memory:') as con: cur = con.cursor() cur.execute('CREATE TABLE customers(id, name, balance)') cur.execute(""" INSERT INTO customers VALUES (123, 'John', 123.45) """) con.commit() deps = SupportDependencies(customer_id=123, db=DatabaseConn(sqlite_conn=con)) result = support_agent.run_sync('What is my balance?', deps=deps) print(result.output) """ support_advice='Hello John, your current account balance, including pending transactions, is $123.45.' block_card=False risk=1 """ result = support_agent.run_sync('I just lost my card!', deps=deps) print(result.output) """ support_advice="I'm sorry to hear that, John. We are temporarily blocking your card to prevent unauthorized transactions." block_card=True risk=8 """ Was this page helpful? Thanks for your feedback! --- # Text to Audio | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/realtime/realtime-text-to-audio/#_top) Text to Audio ============= The smallest possible [realtime session](https://pydantic.dev/docs/ai/realtime/overview/) : send plain text from Python and hear the model speak the reply. Sending text into an OpenAI realtime session asks the model to respond right away, so there’s no microphone, voice-activity detection, or manual turn-taking to manage — just [`send()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.send) and iterate the session’s events. Demonstrates: * [realtime sessions](https://pydantic.dev/docs/ai/realtime/overview/) * the text-in / audio-out path (no audio hardware required) * streaming [`SpeechPartDelta`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SpeechPartDelta) audio and transcript deltas The script streams the spoken reply back, prints the transcript as it arrives, and saves the audio to a `.wav` file you can play afterwards. It’s a handy starting point for turning an existing text chatbot into one that talks, or for generating spoken snippets like a voicemail greeting. Running the Example ------------------- [](https://pydantic.dev/docs/ai/examples/realtime/realtime-text-to-audio/#running-the-example) The realtime model runs on `gpt-realtime`, so you’ll need an OpenAI API key set via `OPENAI_API_KEY`. With [dependencies installed and environment variables set](https://pydantic.dev/docs/ai/examples/setup/#usage) , run: * [pip](https://pydantic.dev/docs/ai/examples/realtime/realtime-text-to-audio/#tab-panel-52) * [uv](https://pydantic.dev/docs/ai/examples/realtime/realtime-text-to-audio/#tab-panel-53) Terminal python -m pydantic_ai_examples.realtime_text_to_audio "Tell me a fun fact about octopuses." Terminal uv run -m pydantic_ai_examples.realtime_text_to_audio "Tell me a fun fact about octopuses." The streamed PCM audio is saved to `realtime-response.wav` so you can listen to the result afterwards. If the turn completes without audio, the script raises an error and does not create an empty WAV file. Example Code ------------ [](https://pydantic.dev/docs/ai/examples/realtime/realtime-text-to-audio/#example-code) realtime\_text\_to\_audio.py from __future__ import annotations import asyncio import sys import wave import logfire from pydantic_ai import Agent, PartDeltaEvent, SpeechPartDelta from pydantic_ai.realtime import RealtimeTurnCompleteEvent from pydantic_ai.realtime.openai import OpenAIRealtimeModelSettings # 'if-token-present' means nothing will be sent (and the example will work) if you don't have logfire configured logfire.configure(send_to_logfire='if-token-present') logfire.instrument_pydantic_ai() # OpenAI's realtime models speak in 24 kHz mono PCM16 audio. SAMPLE_RATE = 24000 DEFAULT_PROMPT = 'Tell me a fun fact about octopuses.' OUTPUT_PATH = 'realtime-response.wav' agent = Agent( instructions='You are a friendly voice assistant. Keep your replies short and conversational.' ) def save_wav(path: str, audio: bytes) -> None: """Wrap the streamed raw PCM16 audio in a WAV container so it can be played back.""" with wave.open(path, 'wb') as wav_file: wav_file.setnchannels(1) # mono wav_file.setsampwidth(2) # 16-bit samples wav_file.setframerate(SAMPLE_RATE) wav_file.writeframes(audio) async def main(prompt: str, output_path: str) -> None: audio = bytearray() async with agent.realtime( 'openai:gpt-realtime', model_settings=OpenAIRealtimeModelSettings(openai_voice='marin'), ).session() as session: # Sending text (rather than audio) into an OpenAI realtime session asks the model to respond # right away — with speech, since a session's default output modality is audio. await session.send(prompt) print(f'you: {prompt}') print('assistant: ', end='', flush=True) async for event in session: match event: case PartDeltaEvent(delta=SpeechPartDelta() as delta): # Deltas carry raw PCM16 audio for playback and/or incremental transcript text. if delta.audio_chunk: audio.extend(delta.audio_chunk) if delta.transcript_delta: print(delta.transcript_delta, end='', flush=True) case RealtimeTurnCompleteEvent(): # The model finished speaking; this was a one-shot request, so we're done. break case _: pass print() if not audio: raise RuntimeError('The realtime response completed without any audio') save_wav(output_path, bytes(audio)) print(f'\nSaved {len(audio)} bytes of audio to {output_path}') if __name__ == '__main__': prompt = sys.argv[1] if len(sys.argv) > 1 else DEFAULT_PROMPT asyncio.run(main(prompt, OUTPUT_PATH)) Was this page helpful? Thanks for your feedback! --- # Server | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/mcp/server/#_top) Server ====== Pydantic AI models can also be used within MCP Servers. MCP Server ---------- [](https://pydantic.dev/docs/ai/mcp/server/#mcp-server) Here’s a simple example of a [Python MCP server](https://github.com/modelcontextprotocol/python-sdk) using Pydantic AI within a tool call: mcp\_server.py from mcp.server.fastmcp import FastMCP from pydantic_ai import Agent server = FastMCP('Pydantic AI Server') server_agent = Agent( 'anthropic:claude-haiku-4-5', instructions='always reply in rhyme' ) @server.tool() async def poet(theme: str) -> str: """Poem generator""" r = await server_agent.run(f'write a poem about {theme}') return r.output if __name__ == '__main__': server.run() Simple client ------------- [](https://pydantic.dev/docs/ai/mcp/server/#simple-client) This server can be queried with any MCP client. Here is an example using the Python SDK directly: mcp\_client.py import asyncio import os from mcp import ClientSession, StdioServerParameters from mcp.client.stdio import stdio_client async def client(): server_params = StdioServerParameters( command='python', args=['mcp_server.py'], env=os.environ ) async with stdio_client(server_params) as (read, write): async with ClientSession(read, write) as session: await session.initialize() result = await session.call_tool('poet', {'theme': 'socks'}) print(result.content[0].text) """ Oh, socks, those garments soft and sweet, That nestle softly 'round our feet, From cotton, wool, or blended thread, They keep our toes from feeling dread. """ if __name__ == '__main__': asyncio.run(client()) MCP Sampling ------------ [](https://pydantic.dev/docs/ai/mcp/server/#mcp-sampling) When Pydantic AI agents are used within MCP servers, they can use sampling via [`MCPSamplingModel`](https://pydantic.dev/docs/ai/api/models/mcp-sampling/#pydantic_ai.models.mcp_sampling.MCPSamplingModel) . We can extend the above example to use sampling so instead of connecting directly to the LLM, the agent calls back through the MCP client to make LLM calls. mcp\_server\_sampling.py from mcp.server.fastmcp import Context, FastMCP from pydantic_ai import Agent from pydantic_ai.models.mcp_sampling import MCPSamplingModel server = FastMCP('Pydantic AI Server with sampling') server_agent = Agent(instructions='always reply in rhyme') @server.tool() async def poet(ctx: Context, theme: str) -> str: """Poem generator""" r = await server_agent.run(f'write a poem about {theme}', model=MCPSamplingModel(session=ctx.session)) return r.output if __name__ == '__main__': server.run() # run the server over stdio The [above](https://pydantic.dev/docs/ai/mcp/server/#simple-client) client does not support sampling, so if you tried to use it with this server you’d get an error. The simplest way to support sampling in an MCP client is to [use](https://pydantic.dev/docs/ai/mcp/client/#mcp-sampling) a Pydantic AI agent as the client, but if you wanted to support sampling with the vanilla MCP SDK, you could do so like this: mcp\_client\_sampling.py import asyncio from typing import Any from mcp import ClientSession, StdioServerParameters from mcp.client.stdio import stdio_client from mcp.shared.context import RequestContext from mcp.types import ( CreateMessageRequestParams, CreateMessageResult, ErrorData, TextContent, ) async def sampling_callback( context: RequestContext[ClientSession, Any], params: CreateMessageRequestParams ) -> CreateMessageResult | ErrorData: print('sampling system prompt:', params.systemPrompt) #> sampling system prompt: always reply in rhyme print('sampling messages:', params.messages) """ sampling messages: [\ SamplingMessage(\ role='user',\ content=TextContent(\ type='text',\ text='write a poem about socks',\ annotations=None,\ meta=None,\ ),\ meta=None,\ )\ ] """ # TODO get the response content by calling an LLM... response_content = 'Socks for a fox.' return CreateMessageResult( role='assistant', content=TextContent(type='text', text=response_content), model='fictional-llm', ) async def client(): server_params = StdioServerParameters(command='python', args=['mcp_server_sampling.py']) async with stdio_client(server_params) as (read, write): async with ClientSession(read, write, sampling_callback=sampling_callback) as session: await session.initialize() result = await session.call_tool('poet', {'theme': 'socks'}) print(result.content[0].text) #> Socks for a fox. if __name__ == '__main__': asyncio.run(client()) _(This example is complete, it can be run “as is”)_ Was this page helpful? Thanks for your feedback! --- # Overview | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/mcp/overview/#_top) Overview ======== Pydantic AI supports [Model Context Protocol (MCP)](https://modelcontextprotocol.io/) in multiple ways: 1. [Agents](https://pydantic.dev/docs/ai/core-concepts/agent/) can connect to MCP servers and use their tools — see [Use MCP servers](https://pydantic.dev/docs/ai/mcp/overview/#use-mcp-servers) below. 2. Agents can be used within MCP servers. [Learn more](https://pydantic.dev/docs/ai/mcp/server/) What is MCP? ------------ [](https://pydantic.dev/docs/ai/mcp/overview/#what-is-mcp) The Model Context Protocol is a standardized protocol that allows AI applications (including programmatic agents like Pydantic AI, coding agents like [cursor](https://www.cursor.com/) , and desktop applications like [Claude Desktop](https://claude.ai/download) ) to connect to external tools and services using a common interface. As with other protocols, the dream of MCP is that a wide range of applications can speak to each other without the need for specific integrations. There is a great list of MCP servers at [github.com/modelcontextprotocol/servers](https://github.com/modelcontextprotocol/servers) . Some examples of what this means: * Pydantic AI could use a web search service implemented as an MCP server to implement a deep research agent * Cursor could connect to the [Pydantic Logfire](https://github.com/pydantic/logfire-mcp) MCP server to search logs, traces and metrics to gain context while fixing a bug * Pydantic AI, or any other MCP client could connect to our [Run Python](https://github.com/pydantic/mcp-run-python) MCP server to run arbitrary Python code in a sandboxed environment Use MCP servers --------------- [](https://pydantic.dev/docs/ai/mcp/overview/#use-mcp-servers) The recommended way to give an agent access to an MCP server is the [`MCP` capability](https://pydantic.dev/docs/ai/capabilities/mcp/) . It runs the MCP server locally by default — keeping credentials, hooks, and tracing under your control — and lets you opt into the model provider’s [native MCP support](https://pydantic.dev/docs/ai/tools-toolsets/native-tools/#mcp-server-tool) with a single `native=True` flag, so the same agent works across providers without code changes: mcp\_capability.py from pydantic_ai import Agent from pydantic_ai.capabilities import MCP agent = Agent( 'openai:gpt-5.2', capabilities=[\ # Runs the MCP server locally by default\ MCP(url='https://mcp.example.com/api'),\ \ # Opt into native MCP — falls back to local if the model doesn't support it\ MCP(url='https://mcp.example.com/other', native=True),\ ], ) Pass a URL as the first argument to enable both the local fallback and (with `native=True`) provider-native MCP. On the local side, `local=` accepts any [`MCPToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset) input — a URL, FastMCP transport, pre-built `fastmcp.Client`, in-process `FastMCP` server, or local script path. See the [capability documentation](https://pydantic.dev/docs/ai/capabilities/mcp/) for the full set of inputs and configuration options. For lower-level access — managing the toolset lifecycle directly, sharing one MCP server across multiple agents, or passing advanced transport / client configuration that doesn’t fit the capability shape — use `MCPToolset` directly via `toolsets=[...]`. See the [MCP client documentation](https://pydantic.dev/docs/ai/mcp/client/) for details. If you only need the model provider’s native MCP support without a local fallback, you can use [`MCPServerTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.MCPServerTool) as a [native tool](https://pydantic.dev/docs/ai/tools-toolsets/native-tools/#mcp-server-tool) directly. For building MCP servers with Pydantic AI agents, see the [MCP server documentation](https://pydantic.dev/docs/ai/mcp/server/) . Was this page helpful? Thanks for your feedback! --- # Raise Content Filter Error | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/raise-content-filter-error/#_top) Raise Content Filter Error ========================== [`RaiseContentFilterError`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.RaiseContentFilterError) is a [capability](https://pydantic.dev/docs/ai/capabilities/overview/) that opts into treating any model response with `finish_reason='content_filter'` as a [`ContentFilterError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ContentFilterError) , even when the provider returns partial text or refusal text: raise\_content\_filter\_error.py from pydantic_ai import Agent from pydantic_ai.capabilities import RaiseContentFilterError from pydantic_ai.exceptions import ContentFilterError from pydantic_ai.messages import ModelMessage, ModelResponse, TextPart from pydantic_ai.models.function import AgentInfo, FunctionModel def filtered_response(messages: list[ModelMessage], info: AgentInfo) -> ModelResponse: return ModelResponse( parts=[TextPart(content='I cannot help with that.')], finish_reason='content_filter', provider_details={'finish_reason': 'content_filter'}, ) agent = Agent(FunctionModel(filtered_response), capabilities=[RaiseContentFilterError()]) try: agent.run_sync('Tell me how to make a weapon.') except ContentFilterError as exc: print(exc.message) #> Content filter triggered. Finish reason: 'content_filter' _(This example is complete, it can be run “as is”)_ By default, Pydantic AI only raises [`ContentFilterError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ContentFilterError) when a `content_filter` response is _empty_: if the provider returns partial text or refusal text alongside `finish_reason='content_filter'`, that text becomes ordinary agent output and no error is raised (see [finish reason handling](https://pydantic.dev/docs/ai/models/overview/#finish-reason-example) ). This capability extends the check to _every_ `content_filter` response, so partial and refusal text raise too. When it raises, the full [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) is serialized into [`ContentFilterError.body`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UnexpectedModelBehavior.body) so the partial text remains inspectable. Was this page helpful? Thanks for your feedback! --- # Include Tool Return Schemas | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/include-tool-return-schemas/#_top) Include Tool Return Schemas =========================== [`IncludeToolReturnSchemas`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.IncludeToolReturnSchemas) is a [capability](https://pydantic.dev/docs/ai/capabilities/overview/) that includes return type schemas in tool definitions sent to the model. For models that natively support return schemas (e.g. Google Gemini), the schema is passed as a structured field in the API request. For other models, it is injected into the tool description as JSON text. include\_return\_schemas.py from pydantic_ai import Agent from pydantic_ai.capabilities import IncludeToolReturnSchemas from pydantic_ai.models.test import TestModel test_model = TestModel() agent = Agent(test_model, capabilities=[IncludeToolReturnSchemas()]) @agent.tool_plain def get_temperature(city: str) -> float: """Get the temperature for a city.""" return 21.0 result = agent.run_sync('What is the temperature in Paris?') params = test_model.last_model_request_parameters assert params is not None td = params.function_tools[0] assert td.include_return_schema is True _(This example is complete, it can be run “as is”)_ Use the `tools` parameter to select which tools should include return schemas. It accepts a list of tool names, a metadata dict for matching, or a callable predicate: include\_return\_schemas\_selective.py from pydantic_ai import Agent from pydantic_ai.capabilities import IncludeToolReturnSchemas from pydantic_ai.models.test import TestModel test_model = TestModel() agent = Agent( test_model, capabilities=[IncludeToolReturnSchemas(tools=['get_temperature'])], ) @agent.tool_plain def get_temperature(city: str) -> float: """Get the temperature for a city.""" return 21.0 @agent.tool_plain def get_greeting(name: str) -> str: """Get a greeting.""" return f'Hello, {name}!' result = agent.run_sync('Hello') params = test_model.last_model_request_parameters assert params is not None temp_tool = next(t for t in params.function_tools if t.name == 'get_temperature') greet_tool = next(t for t in params.function_tools if t.name == 'get_greeting') assert temp_tool.include_return_schema is True assert greet_tool.include_return_schema is None _(This example is complete, it can be run “as is”)_ The same effect can be achieved at the toolset level using [`.include_return_schemas()`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AbstractToolset.include_return_schemas) — see [toolset composition](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/#including-return-schemas) . Was this page helpful? Thanks for your feedback! --- # Prefix Tools | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/prefix-tools/#_top) Prefix Tools ============ [`PrefixTools`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.PrefixTools) is a [capability](https://pydantic.dev/docs/ai/capabilities/overview/) that wraps another capability and prefixes all of its tool names, useful for namespacing when composing multiple capabilities that might have conflicting tool names: prefix\_tools\_example.py from pydantic_ai import Agent from pydantic_ai.capabilities import MCP, PrefixTools agent = Agent( 'openai:gpt-5.2', capabilities=[\ PrefixTools(MCP(url='https://api1.example.com', native=True), prefix='api1'),\ PrefixTools(MCP(url='https://api2.example.com', native=True), prefix='api2'),\ ], ) Every [`AbstractCapability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability) has a convenience method [`prefix_tools`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability.prefix_tools) that returns a [`PrefixTools`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.PrefixTools) wrapper: prefix\_convenience.py MCP(url='https://mcp.example.com/api', native=True).prefix_tools('mcp') Was this page helpful? Thanks for your feedback! --- # Chat App with FastAPI | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/conversational-agents/chat-app/#_top) Chat App with FastAPI ===================== Simple chat app example build with FastAPI. Demonstrates: * [reusing chat history](https://pydantic.dev/docs/ai/core-concepts/message-history/) * [serializing messages](https://pydantic.dev/docs/ai/core-concepts/message-history/#accessing-messages-from-results) * [streaming responses](https://pydantic.dev/docs/ai/core-concepts/output/#streamed-results) This demonstrates storing chat history between requests and using it to give the model context for new responses. Most of the complex logic here is between `chat_app.py` which streams the response to the browser, and `chat_app.ts` which renders messages in the browser. Running the Example ------------------- [](https://pydantic.dev/docs/ai/examples/conversational-agents/chat-app/#running-the-example) With [dependencies installed and environment variables set](https://pydantic.dev/docs/ai/examples/setup/#usage) , run: * [pip](https://pydantic.dev/docs/ai/examples/conversational-agents/chat-app/#tab-panel-28) * [uv](https://pydantic.dev/docs/ai/examples/conversational-agents/chat-app/#tab-panel-29) Terminal python -m pydantic_ai_examples.chat_app Terminal uv run -m pydantic_ai_examples.chat_app Then open the app at [localhost:8000](http://localhost:8000/) . ![Example conversation](https://pydantic.dev/docs/ai/img/chat-app-example.png) Example Code ------------ [](https://pydantic.dev/docs/ai/examples/conversational-agents/chat-app/#example-code) Python code that runs the chat app: chat\_app.py from __future__ import annotations as _annotations import asyncio import json import sqlite3 from collections.abc import AsyncGenerator, Callable from concurrent.futures.thread import ThreadPoolExecutor from contextlib import asynccontextmanager from dataclasses import dataclass from datetime import datetime, timezone from functools import partial from pathlib import Path from typing import Annotated, Any, Literal, TypeVar import fastapi import logfire from fastapi import Depends, Request from fastapi.responses import FileResponse, Response, StreamingResponse from typing_extensions import LiteralString, ParamSpec, TypedDict from pydantic_ai import ( Agent, ModelMessage, ModelMessagesTypeAdapter, ModelRequest, ModelResponse, TextPart, UnexpectedModelBehavior, UserPromptPart, ) # 'if-token-present' means nothing will be sent (and the example will work) if you don't have logfire configured logfire.configure(send_to_logfire='if-token-present') logfire.instrument_pydantic_ai() agent = Agent('openai:gpt-5.2') THIS_DIR = Path(__file__).parent @asynccontextmanager async def lifespan(_app: fastapi.FastAPI): async with Database.connect() as db: yield {'db': db} app = fastapi.FastAPI(lifespan=lifespan) logfire.instrument_fastapi(app) @app.get('/') async def index() -> FileResponse: return FileResponse((THIS_DIR / 'chat_app.html'), media_type='text/html') @app.get('/chat_app.ts') async def main_ts() -> FileResponse: """Get the raw typescript code, it's compiled in the browser, forgive me.""" return FileResponse((THIS_DIR / 'chat_app.ts'), media_type='text/plain') async def get_db(request: Request) -> Database: return request.state.db @app.get('/chat/') async def get_chat(database: Database = Depends(get_db)) -> Response: msgs = await database.get_messages() return Response( b'\n'.join(json.dumps(to_chat_message(m)).encode('utf-8') for m in msgs), media_type='text/plain', ) class ChatMessage(TypedDict): """Format of messages sent to the browser.""" role: Literal['user', 'model'] timestamp: str content: str def to_chat_message(m: ModelMessage) -> ChatMessage: first_part = m.parts[0] if isinstance(m, ModelRequest): if isinstance(first_part, UserPromptPart): assert isinstance(first_part.content, str) return { 'role': 'user', 'timestamp': first_part.timestamp.isoformat(), 'content': first_part.content, } elif isinstance(m, ModelResponse): if isinstance(first_part, TextPart): return { 'role': 'model', 'timestamp': m.timestamp.isoformat(), 'content': first_part.content, } raise UnexpectedModelBehavior(f'Unexpected message type for chat app: {m}') @app.post('/chat/') async def post_chat( prompt: Annotated[str, fastapi.Form()], database: Database = Depends(get_db) ) -> StreamingResponse: async def stream_messages(): """Streams new line delimited JSON `Message`s to the client.""" # stream the user prompt so that can be displayed straight away yield ( json.dumps( { 'role': 'user', 'timestamp': datetime.now(tz=timezone.utc).isoformat(), 'content': prompt, } ).encode('utf-8') + b'\n' ) # get the chat history so far to pass as context to the agent messages = await database.get_messages() # run the agent with the user prompt and the chat history async with agent.run_stream(prompt, message_history=messages) as result: async for text in result.stream_output(debounce_by=0.01): # text here is a `str` and the frontend wants # JSON encoded ModelResponse, so we create one m = ModelResponse(parts=[TextPart(text)], timestamp=result.timestamp) yield json.dumps(to_chat_message(m)).encode('utf-8') + b'\n' # add new messages (e.g. the user prompt and the agent response in this case) to the database await database.add_messages(result.new_messages_json()) return StreamingResponse(stream_messages(), media_type='text/plain') P = ParamSpec('P') R = TypeVar('R') @dataclass class Database: """Rudimentary database to store chat messages in SQLite. The SQLite standard library package is synchronous, so we use a thread pool executor to run queries asynchronously. """ con: sqlite3.Connection _loop: asyncio.AbstractEventLoop _executor: ThreadPoolExecutor @classmethod @asynccontextmanager async def connect( cls, file: Path = THIS_DIR / '.chat_app_messages.sqlite' ) -> AsyncGenerator[Database]: with logfire.span('connect to DB'): loop = asyncio.get_running_loop() executor = ThreadPoolExecutor(max_workers=1) con = await loop.run_in_executor(executor, cls._connect, file) slf = cls(con, loop, executor) try: yield slf finally: await slf._asyncify(con.close) @staticmethod def _connect(file: Path) -> sqlite3.Connection: con = sqlite3.connect(str(file)) con = logfire.instrument_sqlite3(con) cur = con.cursor() cur.execute( 'CREATE TABLE IF NOT EXISTS messages (id INT PRIMARY KEY, message_list TEXT);' ) con.commit() return con async def add_messages(self, messages: bytes): await self._asyncify( self._execute, 'INSERT INTO messages (message_list) VALUES (?);', messages, commit=True, ) await self._asyncify(self.con.commit) async def get_messages(self) -> list[ModelMessage]: c = await self._asyncify( self._execute, 'SELECT message_list FROM messages order by id' ) rows = await self._asyncify(c.fetchall) messages: list[ModelMessage] = [] for row in rows: messages.extend(ModelMessagesTypeAdapter.validate_json(row[0])) return messages def _execute( self, sql: LiteralString, *args: Any, commit: bool = False ) -> sqlite3.Cursor: cur = self.con.cursor() cur.execute(sql, args) if commit: self.con.commit() return cur async def _asyncify( self, func: Callable[P, R], *args: P.args, **kwargs: P.kwargs ) -> R: return await self._loop.run_in_executor( # pyright: ignore[reportUnknownVariableType] self._executor, partial(func, **kwargs), *args, # pyright: ignore[reportCallIssue] ) if __name__ == '__main__': import uvicorn uvicorn.run( 'pydantic_ai_examples.chat_app:app', reload=True, reload_dirs=[str(THIS_DIR)] ) Simple HTML page to render the app: chat\_app.html Chat App

Chat App

Ask me anything...

Error occurred, check the browser developer console for more information.
TypeScript to handle rendering the messages, to keep this simple (and at the risk of offending frontend developers) the typescript code is passed to the browser as plain text and transpiled in the browser. chat\_app.ts // BIG FAT WARNING: to avoid the complexity of npm, this typescript is compiled in the browser // there's currently no static type checking import { marked } from 'https://cdnjs.cloudflare.com/ajax/libs/marked/15.0.0/lib/marked.esm.js' const convElement = document.getElementById('conversation') const promptInput = document.getElementById('prompt-input') as HTMLInputElement const spinner = document.getElementById('spinner') // stream the response and render messages as each chunk is received // data is sent as newline-delimited JSON async function onFetchResponse(response: Response): Promise { let text = '' let decoder = new TextDecoder() if (response.ok) { const reader = response.body.getReader() while (true) { const {done, value} = await reader.read() if (done) { break } text += decoder.decode(value) addMessages(text) spinner.classList.remove('active') } addMessages(text) promptInput.disabled = false promptInput.focus() } else { const text = await response.text() console.error(`Unexpected response: ${response.status}`, {response, text}) throw new Error(`Unexpected response: ${response.status}`) } } // The format of messages, this matches pydantic-ai both for brevity and understanding // in production, you might not want to keep this format all the way to the frontend interface Message { role: string content: string timestamp: string } // take raw response text and render messages into the `#conversation` element // Message timestamp is assumed to be a unique identifier of a message, and is used to deduplicate // hence you can send data about the same message multiple times, and it will be updated // instead of creating a new message elements function addMessages(responseText: string) { const lines = responseText.split('\n') const messages: Message[] = lines.filter(line => line.length > 1).map(j => JSON.parse(j)) for (const message of messages) { // we use the timestamp as a crude element id const {timestamp, role, content} = message const id = `msg-${timestamp}` let msgDiv = document.getElementById(id) if (!msgDiv) { msgDiv = document.createElement('div') msgDiv.id = id msgDiv.title = `${role} at ${timestamp}` msgDiv.classList.add('border-top', 'pt-2', role) convElement.appendChild(msgDiv) } msgDiv.innerHTML = marked.parse(content) } window.scrollTo({ top: document.body.scrollHeight, behavior: 'smooth' }) } function onError(error: any) { console.error(error) document.getElementById('error').classList.remove('d-none') document.getElementById('spinner').classList.remove('active') } async function onSubmit(e: SubmitEvent): Promise { e.preventDefault() spinner.classList.add('active') const body = new FormData(e.target as HTMLFormElement) promptInput.value = '' promptInput.disabled = true const response = await fetch('/chat/', {method: 'POST', body}) await onFetchResponse(response) } // call onSubmit when the form is submitted (e.g. user clicks the send button or hits Enter) document.querySelector('form').addEventListener('submit', (e) => onSubmit(e).catch(onError)) // load messages on page load fetch('/chat/').then(onFetchResponse).catch(onError) Was this page helpful? Thanks for your feedback! --- # Data Analyst | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/data-analytics/data-analyst/#_top) Data Analyst ============ Sometimes in an agent workflow, the agent does not need to know the exact tool output, but still needs to process the tool output in some ways. This is especially common in data analytics: the agent needs to know that the result of a query tool is a `DataFrame` with certain named columns, but not necessarily the content of every single row. With Pydantic AI, you can use a [dependencies object](https://pydantic.dev/docs/ai/core-concepts/dependencies/) to store the result from one tool and use it in another tool. In this example, we’ll build an agent that analyzes the [Rotten Tomatoes movie review dataset from Cornell](https://huggingface.co/datasets/cornell-movie-review-data/rotten_tomatoes) . Demonstrates: * [agent dependencies](https://pydantic.dev/docs/ai/core-concepts/dependencies/) Running the Example ------------------- [](https://pydantic.dev/docs/ai/examples/data-analytics/data-analyst/#running-the-example) With [dependencies installed and environment variables set](https://pydantic.dev/docs/ai/examples/setup/#usage) , run: * [pip](https://pydantic.dev/docs/ai/examples/data-analytics/data-analyst/#tab-panel-30) * [uv](https://pydantic.dev/docs/ai/examples/data-analytics/data-analyst/#tab-panel-31) Terminal python -m pydantic_ai_examples.data_analyst Terminal uv run -m pydantic_ai_examples.data_analyst Output (debug): > Based on my analysis of the Cornell Movie Review dataset (rotten\_tomatoes), there are **4,265 negative comments** in the training split. These are the reviews labeled as ‘neg’ (represented by 0 in the dataset). Example Code ------------ [](https://pydantic.dev/docs/ai/examples/data-analytics/data-analyst/#example-code) data\_analyst.py from dataclasses import dataclass, field import datasets import duckdb import pandas as pd from pydantic_ai import Agent, ModelRetry, RunContext @dataclass class AnalystAgentDeps: output: dict[str, pd.DataFrame] = field(default_factory=dict[str, pd.DataFrame]) def store(self, value: pd.DataFrame) -> str: """Store the output in deps and return the reference such as Out[1] to be used by the LLM.""" ref = f'Out[{len(self.output) + 1}]' self.output[ref] = value return ref def get(self, ref: str) -> pd.DataFrame: if ref not in self.output: raise ModelRetry( f'Error: {ref} is not a valid variable reference. Check the previous messages and try again.' ) return self.output[ref] analyst_agent = Agent( 'openai:gpt-5.2', deps_type=AnalystAgentDeps, instructions='You are a data analyst and your job is to analyze the data according to the user request.', ) @analyst_agent.tool def load_dataset( ctx: RunContext[AnalystAgentDeps], path: str, split: str = 'train', ) -> str: """Load the `split` of dataset `dataset_name` from huggingface. Args: ctx: Pydantic AI agent RunContext path: name of the dataset in the form of `/` split: load the split of the dataset (default: "train") """ # begin load data from hf builder = datasets.load_dataset_builder(path) # pyright: ignore[reportUnknownMemberType] splits: dict[str, datasets.SplitInfo] = builder.info.splits or {} if split not in splits: raise ModelRetry( f'{split} is not valid for dataset {path}. Valid splits are {",".join(splits.keys())}' ) builder.download_and_prepare() # pyright: ignore[reportUnknownMemberType] dataset = builder.as_dataset(split=split) assert isinstance(dataset, datasets.Dataset) dataframe = dataset.to_pandas() assert isinstance(dataframe, pd.DataFrame) # end load data from hf # store the dataframe in the deps and get a ref like "Out[1]" ref = ctx.deps.store(dataframe) # construct a summary of the loaded dataset output = [\ f'Loaded the dataset as `{ref}`.',\ f'Description: {dataset.info.description}'\ if dataset.info.description\ else None,\ f'Features: {dataset.info.features!r}' if dataset.info.features else None,\ ] return '\n'.join(filter(None, output)) @analyst_agent.tool def run_duckdb(ctx: RunContext[AnalystAgentDeps], dataset: str, sql: str) -> str: """Run DuckDB SQL query on the DataFrame. Note that the virtual table name used in DuckDB SQL must be `dataset`. Args: ctx: Pydantic AI agent RunContext dataset: reference string to the DataFrame sql: the query to be executed using DuckDB """ data = ctx.deps.get(dataset) result = duckdb.query_df(df=data, virtual_table_name='dataset', sql_query=sql) # pass the result as ref (because DuckDB SQL can select many rows, creating another huge dataframe) ref = ctx.deps.store(result.df()) return f'Executed SQL, result is `{ref}`' @analyst_agent.tool def display(ctx: RunContext[AnalystAgentDeps], name: str) -> str: """Display at most 5 rows of the dataframe.""" dataset = ctx.deps.get(name) return dataset.head().to_string() # pyright: ignore[reportUnknownMemberType] if __name__ == '__main__': deps = AnalystAgentDeps() result = analyst_agent.run_sync( user_prompt='Count how many negative comments are there in the dataset `cornell-movie-review-data/rotten_tomatoes`', deps=deps, ) print(result.output) Appendix -------- [](https://pydantic.dev/docs/ai/examples/data-analytics/data-analyst/#appendix) ### Choosing a Model [](https://pydantic.dev/docs/ai/examples/data-analytics/data-analyst/#choosing-a-model) This example requires using a model that understands DuckDB SQL. You can check with `clai`: Terminal > clai -m bedrock:us.anthropic.claude-sonnet-4-5-20250929-v1:0 clai - Pydantic AI CLI v0.0.1.dev920+41dd069 with bedrock:us.anthropic.claude-sonnet-4-5-20250929-v1:0 clai ➤ do you understand duckdb sql? # DuckDB SQL Yes, I understand DuckDB SQL. DuckDB is an in-process analytical SQL database that uses syntax similar to PostgreSQL. It specializes in analytical queries and is designed for high-performance analysis of structured data. Some key features of DuckDB SQL include: • OLAP (Online Analytical Processing) optimized • Columnar-vectorized query execution • Standard SQL support with PostgreSQL compatibility • Support for complex analytical queries • Efficient handling of CSV/Parquet/JSON files I can help you with DuckDB SQL queries, schema design, optimization, or other DuckDB-related questions. Was this page helpful? Thanks for your feedback! --- # Crusoe | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/models/crusoe/#_top) Crusoe ====== Install ------- [](https://pydantic.dev/docs/ai/models/crusoe/#install) To use `CrusoeModel`, you need to either install `pydantic-ai`, or install `pydantic-ai-slim` with the `crusoe` optional group: * [pip](https://pydantic.dev/docs/ai/models/crusoe/#tab-panel-110) * [uv](https://pydantic.dev/docs/ai/models/crusoe/#tab-panel-111) Terminal pip install "pydantic-ai-slim[crusoe]" Terminal uv add "pydantic-ai-slim[crusoe]" Configuration ------------- [](https://pydantic.dev/docs/ai/models/crusoe/#configuration) To use [Crusoe](https://crusoe.ai/) Serverless Inference, go to the [Crusoe Cloud console](https://console.crusoecloud.com/) , select Models, and click `Get API Key`. For a list of available models, see the [Crusoe Serverless Inference documentation](https://docs.crusoecloud.com/serverless-inference/overview) . Environment variable -------------------- [](https://pydantic.dev/docs/ai/models/crusoe/#environment-variable) Once you have the API key, you can set it as an environment variable: Terminal export CRUSOE_API_KEY='your-api-key' You can then use `CrusoeModel` by name: from pydantic_ai import Agent agent = Agent('crusoe:zai/GLM-5.2') ... Or initialise the model directly with just the model name: from pydantic_ai import Agent from pydantic_ai.models.crusoe import CrusoeModel model = CrusoeModel('zai/GLM-5.2') agent = Agent(model) ... Model names ----------- [](https://pydantic.dev/docs/ai/models/crusoe/#model-names) Crusoe serves open-weight models from many labs behind one endpoint, and model names carry the lab as a prefix — `zai/GLM-5.2`, `deepseek-ai/DeepSeek-V4-Pro`, `meta-llama/Llama-3.3-70B-Instruct`, `openai/gpt-oss-120b`. That prefix is what selects the [model profile](https://pydantic.dev/docs/ai/models/openai/#model-profile) , so keep it on the name rather than passing the bare model id. Structured output ----------------- [](https://pydantic.dev/docs/ai/models/crusoe/#structured-output) Crusoe serves every model with guided decoding, so [`NativeOutput`](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.NativeOutput) works across the catalog — including for model families that don’t support native structured output when you reach them through their own provider. `provider` argument ------------------- [](https://pydantic.dev/docs/ai/models/crusoe/#provider-argument) You can provide a custom `Provider` via the `provider` argument: from pydantic_ai import Agent from pydantic_ai.models.crusoe import CrusoeModel from pydantic_ai.providers.crusoe import CrusoeProvider model = CrusoeModel('zai/GLM-5.2', provider=CrusoeProvider(api_key='your-api-key')) agent = Agent(model) ... You can also customize the `CrusoeProvider` with a custom `httpx.AsyncClient`: from httpx import AsyncClient from pydantic_ai import Agent from pydantic_ai.models.crusoe import CrusoeModel from pydantic_ai.providers.crusoe import CrusoeProvider custom_http_client = AsyncClient(timeout=30) model = CrusoeModel( 'zai/GLM-5.2', provider=CrusoeProvider(api_key='your-api-key', http_client=custom_http_client), ) agent = Agent(model) ... Was this page helpful? Thanks for your feedback! --- # Cohere | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/models/cohere/#_top) Cohere ====== Install ------- [](https://pydantic.dev/docs/ai/models/cohere/#install) To use `CohereModel`, you need to either install `pydantic-ai`, or install `pydantic-ai-slim` with the `cohere` optional group: * [pip](https://pydantic.dev/docs/ai/models/cohere/#tab-panel-108) * [uv](https://pydantic.dev/docs/ai/models/cohere/#tab-panel-109) Terminal pip install "pydantic-ai-slim[cohere]" Terminal uv add "pydantic-ai-slim[cohere]" Configuration ------------- [](https://pydantic.dev/docs/ai/models/cohere/#configuration) To use [Cohere](https://cohere.com/) through their API, go to [dashboard.cohere.com/api-keys](https://dashboard.cohere.com/api-keys) and follow your nose until you find the place to generate an API key. `CohereModelName` contains a list of the most popular Cohere models. Environment variable -------------------- [](https://pydantic.dev/docs/ai/models/cohere/#environment-variable) Once you have the API key, you can set it as an environment variable: Terminal export CO_API_KEY='your-api-key' You can then use `CohereModel` by name: from pydantic_ai import Agent agent = Agent('cohere:command-r7b-12-2024') ... Or initialise the model directly with just the model name: from pydantic_ai import Agent from pydantic_ai.models.cohere import CohereModel model = CohereModel('command-r7b-12-2024') agent = Agent(model) ... `provider` argument ------------------- [](https://pydantic.dev/docs/ai/models/cohere/#provider-argument) You can provide a custom `Provider` via the `provider` argument: from pydantic_ai import Agent from pydantic_ai.models.cohere import CohereModel from pydantic_ai.providers.cohere import CohereProvider model = CohereModel('command-r7b-12-2024', provider=CohereProvider(api_key='your-api-key')) agent = Agent(model) ... You can also customize the `CohereProvider` with a custom `http_client`: from httpx import AsyncClient from pydantic_ai import Agent from pydantic_ai.models.cohere import CohereModel from pydantic_ai.providers.cohere import CohereProvider custom_http_client = AsyncClient(timeout=30) model = CohereModel( 'command-r7b-12-2024', provider=CohereProvider(api_key='your-api-key', http_client=custom_http_client), ) agent = Agent(model) ... Model settings -------------- [](https://pydantic.dev/docs/ai/models/cohere/#model-settings) You can customize model behavior using [`CohereModelSettings`](https://pydantic.dev/docs/ai/api/models/cohere/#pydantic_ai.models.cohere.CohereModelSettings) : from pydantic_ai import Agent from pydantic_ai.models.cohere import CohereModel, CohereModelSettings model = CohereModel('command-r7b-12-2024') settings = CohereModelSettings( temperature=0.2, top_k=40, ) agent = Agent(model, model_settings=settings) ... Was this page helpful? Thanks for your feedback! --- # Audio, images, and transcripts | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/realtime/audio/#_top) Audio, images, and transcripts ============================== A realtime session accepts live audio, text, and supported images while exposing separate views for playback and captions. Use the high-level session views for media and transcripts; consume the main [event stream](https://pydantic.dev/docs/ai/realtime/events/) for tools, turn boundaries, reconnects, and errors. Audio wire contract ------------------- [](https://pydantic.dev/docs/ai/realtime/audio/#audio-wire-contract) You send and receive raw audio samples; there is no container or codec in the live path. [`send_audio()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.send_audio) accepts raw, signed 16-bit little-endian mono PCM. [`stream_audio()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.stream_audio) returns the same format. Capture at [`session.audio_input_sample_rate`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.audio_input_sample_rate) and play at [`session.audio_output_sample_rate`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.audio_output_sample_rate) ; input and output rates can differ. Start with 100 ms input chunks to balance interactive cadence with per-chunk overhead, then tune for your transport. The provider pages list their model-specific rates and constraints: [OpenAI](https://pydantic.dev/docs/ai/realtime/openai/#feature-support-and-limitations) , [Azure OpenAI](https://pydantic.dev/docs/ai/realtime/azure/#feature-support-and-limitations) , [Google Gemini](https://pydantic.dev/docs/ai/realtime/gemini/#feature-support-and-limitations) , and [xAI](https://pydantic.dev/docs/ai/realtime/xai/#feature-support-and-limitations) . For a complete microphone and speaker loop with bounded buffers, playback accounting, and clean shutdown, use the [realtime voice assistant example](https://pydantic.dev/docs/ai/examples/realtime/realtime-voice/) . Consuming audio and transcripts ------------------------------- [](https://pydantic.dev/docs/ai/realtime/audio/#consuming-audio-and-transcripts) Run media views alongside the main iterator: import asyncio from collections.abc import AsyncIterator from pydantic_ai import Agent from pydantic_ai.messages import SpeechPart from pydantic_ai.realtime import RealtimeTurnCompleteEvent agent = Agent(instructions='You are a helpful voice assistant.') async def play_audio(chunks: AsyncIterator[bytes]) -> None: async for chunk in chunks: ... # Write the PCM16 chunk to your speaker or audio output stream. async def show_transcripts(parts: AsyncIterator[SpeechPart]) -> None: async for part in parts: print(part.speaker, part.transcript) #> assistant Hello from the realtime assistant. async def main(): async with agent.realtime('openai:gpt-realtime').session() as session: audio_task = asyncio.create_task(play_audio(session.stream_audio())) transcript_task = asyncio.create_task(show_transcripts(session.stream_transcripts())) async for event in session: if isinstance(event, RealtimeTurnCompleteEvent): break # Leaving the `async with` block closes the session, which ends every live view. await asyncio.gather(audio_task, transcript_task) Each view is independently bounded; a slow consumer drops its oldest item rather than stalling tools, turn tracking, or other consumers. Subscriptions begin when iteration starts, so unused views do not buffer. [`close()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.close) discards pending items and ends every live iterator; [`closed`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.closed) reports the state. ### Live captions [](https://pydantic.dev/docs/ai/realtime/audio/#live-captions) For live captions, pass `delta=True` to [`stream_transcripts()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.stream_transcripts) . Each [`TranscriptUpdate`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.TranscriptUpdate) includes the speaker, new delta, full transcript so far, and an index identifying the turn. Replace a caption by index rather than blindly appending, because speech recognition can revise earlier words: from pydantic_ai.realtime import RealtimeSession bubbles: dict[int, tuple[str, str]] = {} async def show_captions(session: RealtimeSession) -> None: async for update in session.stream_transcripts(delta=True): bubbles[update.index] = (update.speaker, update.transcript) Input transcription ------------------- [](https://pydantic.dev/docs/ai/realtime/audio/#input-transcription) The shared `input_transcription_model` setting controls whether user speech reaches history as text: | Value | Behavior | | --- | --- | | `'auto'` (default) | Uses the provider’s recommended transcription path. | | A model ID | Pins a dedicated transcription model on providers that support one. | | `None` | Disables input transcription. | OpenAI, Azure OpenAI, and xAI use dedicated transcription models. Gemini uses native transcription, configured with `google_input_transcription`: a pinned model ID in the shared setting is ignored (native transcription stays on), and only `None` turns it off. Provider-specific defaults and deployment constraints live on the provider pages. Disabling transcription changes what a spoken turn contributes to history, replay, and text-agent handoff; see [History and handoff](https://pydantic.dev/docs/ai/realtime/history/#retaining-audio) before relying on it. A [WebRTC sideband](https://pydantic.dev/docs/ai/realtime/deployment/#browser-webrtc-server-sideband) receives no audio bytes to retain, so without input transcription its user turns contain no spoken text. Images ------ [](https://pydantic.dev/docs/ai/realtime/audio/#images) Beyond audio and text, a session accepts the same image content as [multimodal input](https://pydantic.dev/docs/ai/core-concepts/input/#image-input) to a standard run. Send an image as context with [`send()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.send) . An image does not trigger a response by itself; the model uses it on the next voice, text, or manually-created turn. from pydantic_ai import BinaryContent async def send_image(session): jpeg_bytes = b'...' await session.send(BinaryContent(data=jpeg_bytes, media_type='image/jpeg')) Streaming images continuously approximates live video: the [camera example](https://pydantic.dev/docs/ai/examples/realtime/realtime-camera/) sends one camera frame per second alongside microphone audio. For continuous streams like that, use the session’s image-retention controls to bound local history; see [Retaining images](https://pydantic.dev/docs/ai/realtime/history/#retaining-images) . Gemini-specific live-video settings belong on the [Gemini provider page](https://pydantic.dev/docs/ai/realtime/gemini/#settings) . Edge cases ---------- [](https://pydantic.dev/docs/ai/realtime/audio/#edge-cases) * Audio and transcript iterators deliberately drop old buffered items when consumers fall behind. [Logfire attributes](https://pydantic.dev/docs/ai/realtime/observability/#logfire-instrumentation) report those drops. * Provider speech/interruption signals differ. Use the profile flags and the [turns guide](https://pydantic.dev/docs/ai/realtime/turns/#barge-in) rather than branching on provider names. Was this page helpful? Thanks for your feedback! --- # Set Tool Metadata | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/set-tool-metadata/#_top) Set Tool Metadata ================= [`SetToolMetadata`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.SetToolMetadata) is a [capability](https://pydantic.dev/docs/ai/capabilities/overview/) that merges metadata key-value pairs onto selected tools. This is useful for tagging tools with configuration that other capabilities or custom logic can inspect: set\_tool\_metadata.py from pydantic_ai import Agent from pydantic_ai.capabilities import SetToolMetadata from pydantic_ai.models.test import TestModel test_model = TestModel() agent = Agent( test_model, capabilities=[SetToolMetadata(tools=['search'], sensitive=True)], ) @agent.tool_plain def search(query: str) -> str: """Search for information.""" return f'Results for: {query}' @agent.tool_plain def greet(name: str) -> str: """Greet someone.""" return f'Hello, {name}!' result = agent.run_sync('Search for pydantic') params = test_model.last_model_request_parameters assert params is not None search_tool = next(t for t in params.function_tools if t.name == 'search') greet_tool = next(t for t in params.function_tools if t.name == 'greet') assert search_tool.metadata is not None and search_tool.metadata.get('sensitive') is True assert greet_tool.metadata is None or greet_tool.metadata.get('sensitive') is None _(This example is complete, it can be run “as is”)_ The same effect can be achieved at the toolset level using [`.with_metadata()`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AbstractToolset.with_metadata) — see [toolset composition](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/#setting-tool-metadata) . Was this page helpful? Thanks for your feedback! --- # Web Search | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/web-search/#_top) Web Search ========== The [`WebSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebSearch) [capability](https://pydantic.dev/docs/ai/capabilities/overview/) gives your agent web search. Like all [provider-adaptive tools](https://pydantic.dev/docs/ai/capabilities/overview/#provider-adaptive-tools) , it uses the provider’s native web search when the model supports it and can fall back to a local implementation on other models. [`WebSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebSearch) defaults to native-only. Backed by [`WebSearchTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebSearchTool) on the native side (see [Web Search Tool](https://pydantic.dev/docs/ai/tools-toolsets/native-tools/#web-search-tool) for provider support and configuration) — pass `native=WebSearchTool(...)` directly when you need full control over the native instance. For the local side, pass `local='duckduckgo'` (or `local=True`) for a [DuckDuckGo](https://pydantic.dev/docs/ai/tools-toolsets/common-tools/#duckduckgo-search-tool) fallback (requires the `duckduckgo` optional group); for other search providers, use a [Tavily](https://pydantic.dev/docs/ai/api/pydantic-ai/common_tools/#pydantic_ai.common_tools.tavily.tavily_search_tool) wrapper from [`common_tools`](https://pydantic.dev/docs/ai/tools-toolsets/common-tools/) , the [`ExaSearchToolset`](https://pydantic.dev/docs/ai/harness/exa-search/) from the Pydantic AI Harness, or any callable, [`Tool`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.Tool) , or [`AbstractToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AbstractToolset) . Native configuration fields: `search_context_size`, `user_location`, `blocked_domains`, `allowed_domains`, `max_uses`, and OpenAI Responses’ `external_web_access`. The domain and `max_uses` constraints require native support. Setting `external_web_access=False` also requires native support because a local fallback cannot guarantee cached or indexed-only search. web\_search.py from pydantic_ai.capabilities import WebSearch # Native-only — raises on models without native web search WebSearch() # Native preferred; DuckDuckGo fallback (needs `pydantic-ai-slim[duckduckgo]`) WebSearch(local='duckduckgo') # Native preferred; custom callable as fallback def my_search(query: str) -> str: ... WebSearch(local=my_search) Was this page helpful? Thanks for your feedback! --- # pydantic_evals.generation | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic_evals/generation/#_top) pydantic\_evals.generation ========================== Utilities for generating example datasets for pydantic\_evals. This module provides functions for generating sample datasets for testing and examples, using LLMs to create realistic test data with proper structure. generate\_dataset ----------------- [](https://pydantic.dev/docs/ai/api/pydantic_evals/generation/#pydantic_evals.generation.generate_dataset) `@async` def generate_dataset( *, dataset_type: type[Dataset[InputsT, OutputT, MetadataT]], path: Path | str | None = None, custom_evaluator_types: Sequence[type[Evaluator[InputsT, OutputT, MetadataT]]] = (), model: models.Model | models.KnownModelName = 'openai:gpt-5.2', n_examples: int = 3, extra_instructions: str | None = None, ) -> Dataset[InputsT, OutputT, MetadataT] Use an LLM to generate a dataset of test cases, each consisting of input, expected output, and metadata. This function creates a properly structured dataset with the specified input, output, and metadata types. It uses an LLM to attempt to generate realistic test cases that conform to the types’ schemas. ### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/generation/#returns) `Dataset`\[`InputsT`, `OutputT`, `MetadataT`\] — A properly structured Dataset object with generated test cases. ### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_evals/generation/#parameters) **`path`** : `Path` | [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic_evals/generation/#pydantic_evals.generation.generate_dataset(path)) Optional path to save the generated dataset. If provided, the dataset will be saved to this location. **`dataset_type`** : [`type`](https://docs.python.org/3/glossary.html#term-type) \[`Dataset`\[`InputsT`, `OutputT`, `MetadataT`\]\] [](https://pydantic.dev/docs/ai/api/pydantic_evals/generation/#pydantic_evals.generation.generate_dataset(dataset_type)) The type of dataset to generate, with the desired input, output, and metadata types. **`custom_evaluator_types`** : [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`type`](https://docs.python.org/3/glossary.html#term-type)\ \[`Evaluator`\[`InputsT`, `OutputT`, `MetadataT`\]\]\] _Default:_ `()` [](https://pydantic.dev/docs/ai/api/pydantic_evals/generation/#pydantic_evals.generation.generate_dataset(custom_evaluator_types)) Optional sequence of custom evaluator classes to include in the schema. **`model`** : [`models.Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) | [`models.KnownModelName`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.KnownModelName) _Default:_ `'openai:gpt-5.2'` [](https://pydantic.dev/docs/ai/api/pydantic_evals/generation/#pydantic_evals.generation.generate_dataset(model)) The Pydantic AI model to use for generation. Defaults to ‘openai:gpt-5.2’. **`n_examples`** : [`int`](https://docs.python.org/3/library/functions.html#int) _Default:_ `3` [](https://pydantic.dev/docs/ai/api/pydantic_evals/generation/#pydantic_evals.generation.generate_dataset(n_examples)) Number of examples to generate. Defaults to 3. **`extra_instructions`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic_evals/generation/#pydantic_evals.generation.generate_dataset(extra_instructions)) Optional additional instructions to provide to the LLM. ### Raises [](https://pydantic.dev/docs/ai/api/pydantic_evals/generation/#raises) * `ValidationError` — If the LLM’s response cannot be parsed as a valid dataset. InputsT ------- [](https://pydantic.dev/docs/ai/api/pydantic_evals/generation/#pydantic_evals.generation.InputsT) Generic type for the inputs to the task being evaluated. **Default:** `TypeVar('InputsT', default=Any)` MetadataT --------- [](https://pydantic.dev/docs/ai/api/pydantic_evals/generation/#pydantic_evals.generation.MetadataT) Generic type for the metadata associated with the task being evaluated. **Default:** `TypeVar('MetadataT', default=Any)` OutputT ------- [](https://pydantic.dev/docs/ai/api/pydantic_evals/generation/#pydantic_evals.generation.OutputT) Generic type for the expected output of the task being evaluated. **Default:** `TypeVar('OutputT', default=Any)` Was this page helpful? Thanks for your feedback! --- # zai | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/zai/#_top) zai === Setup ----- [](https://pydantic.dev/docs/ai/api/models/zai/#setup) For details on how to set up authentication with this model, see [model configuration for Z.AI](https://pydantic.dev/docs/ai/models/zai/) . Z.AI (Zhipu AI) model implementation using OpenAI-compatible API. ZaiModel -------- [](https://pydantic.dev/docs/ai/api/models/zai/#pydantic_ai.models.zai.ZaiModel) **Bases:** `OpenAIChatModel` A model that uses Z.AI’s OpenAI-compatible API. Z.AI (Zhipu AI) provides GLM models with support for thinking/reasoning mode and preserved thinking across turns. Apart from `__init__`, all methods are private or match those of the base class. ### Methods [](https://pydantic.dev/docs/ai/api/models/zai/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/zai/#pydantic_ai.models.zai.ZaiModel.__init__) def __init__( model_name: ZaiModelName, *, provider: Literal['zai'] | Provider[AsyncOpenAI] = 'zai', profile: ModelProfileSpec | None = None, settings: ZaiModelSettings | None = None, ) Initialize a Z.AI model. ##### Parameters [](https://pydantic.dev/docs/ai/api/models/zai/#parameters) **`model_name`** : `ZaiModelName` [](https://pydantic.dev/docs/ai/api/models/zai/#pydantic_ai.models.zai.ZaiModel.__init__(model_name)) The name of the Z.AI model to use. **`provider`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘zai’\] | `Provider`\[`AsyncOpenAI`\] _Default:_ `'zai'` [](https://pydantic.dev/docs/ai/api/models/zai/#pydantic_ai.models.zai.ZaiModel.__init__(provider)) The provider to use. Defaults to ‘zai’. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/zai/#pydantic_ai.models.zai.ZaiModel.__init__(profile)) The model profile to use. Defaults to a profile based on the model name. **`settings`** : `ZaiModelSettings` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/zai/#pydantic_ai.models.zai.ZaiModel.__init__(settings)) Model-specific settings that will be used as defaults for this model. ZaiModelSettings ---------------- [](https://pydantic.dev/docs/ai/api/models/zai/#pydantic_ai.models.zai.ZaiModelSettings) **Bases:** [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) Settings used for a Z.AI model request. ALL FIELDS MUST BE `zai_` PREFIXED SO YOU CAN MERGE THEM WITH OTHER MODELS. ### Attributes [](https://pydantic.dev/docs/ai/api/models/zai/#attributes) #### zai\_clear\_thinking [](https://pydantic.dev/docs/ai/api/models/zai/#pydantic_ai.models.zai.ZaiModelSettings.zai_clear_thinking) Whether to clear historical thinking content from prior turns. Defaults to `False` (preserved thinking) on thinking-capable models, retaining reasoning content from prior assistant responses for improved multi-turn coherence and consistency with other providers. Set to `True` to clear it instead. Only affects cross-turn historical thinking blocks; it does not change whether the model generates thinking in the current turn (controlled by the unified `thinking` setting). When using preserved thinking, you must return the complete, unmodified `reasoning_content` back to the API. All consecutive `reasoning_content` blocks must exactly match the original sequence. See [the Z.AI docs](https://docs.z.ai/guides/capabilities/thinking-mode#preserved-thinking) for more details. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) ZaiModelName ------------ [](https://pydantic.dev/docs/ai/api/models/zai/#pydantic_ai.models.zai.ZaiModelName) Possible Z.AI model names. Since Z.AI supports a variety of models and the list changes frequently, we explicitly list known models but allow any name in the type hints. See [https://docs.z.ai/](https://docs.z.ai/) for an up to date list of models. **Default:** `str | LatestZaiModelNames` Was this page helpful? Thanks for your feedback! --- # X Search | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/x-search/#_top) X Search ======== The [`XSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.XSearch) [capability](https://pydantic.dev/docs/ai/capabilities/overview/) gives your agent search over X (Twitter) posts. It’s a [provider-adaptive tool](https://pydantic.dev/docs/ai/capabilities/overview/#provider-adaptive-tools) backed by [`XSearchTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.XSearchTool) on the native side — see [X Search Tool](https://pydantic.dev/docs/ai/tools-toolsets/native-tools/#x-search-tool) for configuration options. Unlike [Web Search](https://pydantic.dev/docs/ai/capabilities/web-search/) and [Web Fetch](https://pydantic.dev/docs/ai/capabilities/web-fetch/) , there is no default non-xAI fallback: X search is only available natively on xAI models. If your agent is not running on an xAI model, set `fallback_model` explicitly to an xAI model that supports [`XSearchTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.XSearchTool) , and search requests are delegated to that model as a subagent tool: x\_search.py from pydantic_ai import Agent from pydantic_ai.capabilities import XSearch agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[XSearch(fallback_model='xai:grok-4.3')], ) Was this page helpful? Thanks for your feedback! --- # Timeouts | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/core-concepts/timeouts/#_top) Timeouts ======== Bounding how long one step inside a run may take, and ending a run from inside a tool, are answered by separate mechanisms with separate failure modes. This page maps them. To stop a run that is already in flight, see [Cancelling a Run](https://pydantic.dev/docs/ai/core-concepts/agent/#cancelling-a-run) . Bounding how long a step takes ------------------------------ [](https://pydantic.dev/docs/ai/core-concepts/timeouts/#bounding-how-long-a-step-takes) Each knob below bounds a different unit of work. None of them bounds the wall-clock duration of a whole run. | What you want to bound | How to set it | What happens on expiry | | --- | --- | --- | | A single model request | `timeout` on [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | The provider client raises; the run fails unless a [`FallbackModel`](https://pydantic.dev/docs/ai/models/overview/#fallback-model)
or a [transport retry](https://pydantic.dev/docs/ai/models/http-request-retries/)
handles it | | A function tool call | `Agent(tool_timeout=...)`, or `timeout=` on an individual tool — see [Tool Timeout](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-timeout) | The model receives a retry prompt `'Timed out after N seconds.'`, consuming that tool’s [retry budget](https://pydantic.dev/docs/ai/core-concepts/retries/#tool-retries)
. A `def` tool is not actually stopped: the deadline is enforced around the await, so the worker thread runs to completion | | A [hook](https://pydantic.dev/docs/ai/core-concepts/hooks/)
function | `timeout=` on the `@hooks.on.*` decorator | [`HookTimeoutError`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.HookTimeoutError)
, which is an [`AgentRunError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.AgentRunError)
and aborts the run | | Connecting to an MCP server | `MCPToolset(init_timeout=...)`, default `5` seconds | The connection and `initialize` handshake fail | | A single MCP request | `MCPToolset(read_timeout=...)`, default `300` seconds | The request fails; under the default [`tool_error_behavior='retry'`](https://pydantic.dev/docs/ai/mcp/client/#tool-errors)
the model sees it as a retryable tool error | | Opening a [realtime session](https://pydantic.dev/docs/ai/realtime/overview/) | `handshake_timeout` on [`RealtimeModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelSettings)
, default `30` seconds — OpenAI, Azure OpenAI, and xAI | Opening the session raises [`RealtimeError`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeError)
. On a reconnect it consumes a [`ReconnectPolicy`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.ReconnectPolicy)
attempt instead | | Total work done by a run | [`UsageLimits`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits)
— requests, tool calls, tokens, or cost — see [Usage Limits](https://pydantic.dev/docs/ai/core-concepts/agent/#usage-limits) | [`UsageLimitExceeded`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UsageLimitExceeded) | | Wall-clock duration of a whole run | Nothing built in — wrap `agent.run()` in `asyncio.timeout` (Python 3.11+) or `anyio.fail_after()`, or cancel a [`CancellationToken`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.CancellationToken)
from a timer | The run is [cancelled](https://pydantic.dev/docs/ai/core-concepts/agent/#cancelling-a-run) | Two of these need qualifying: * **`ModelSettings['timeout']` is applied per model class, not universally.** The model classes that forward it to their provider client are listed under [`ModelSettings.timeout`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.timeout) ; the ones built on OpenAI’s inherit the forwarding from [`OpenAIChatModel`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIChatModel) / [`OpenAIResponsesModel`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIResponsesModel) . Other model classes ignore the setting, and the timeout on the HTTP client they were built with applies instead. When Pydantic AI creates that client itself, it defaults to a 600-second total timeout with a 5-second connect timeout. Google and Mistral additionally reject an `httpx.Timeout` object and accept only a number of seconds. To bound a request on a model class that ignores the setting, configure the timeout where that provider actually takes one. Most providers accept your own `http_client`, but several don’t: [`XaiProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.xai.XaiProvider) takes a client-level `timeout` (or a preconfigured `xai_client`), [`BedrockProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.bedrock.BedrockProvider) takes `aws_read_timeout` and `aws_connect_timeout` (or a preconfigured `bedrock_client`), and [`HuggingFaceProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.huggingface.HuggingFaceProvider) rejects `http_client` outright in favor of `hf_client`. * **Tool timeouts are enforced by [`FunctionToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.FunctionToolset) only, and each toolset carries its own.** `Agent(tool_timeout=...)` sets the default for tools you register _on the agent_ — it does not reach into a `FunctionToolset` you constructed yourself and passed via `toolsets=[...]`. Give that toolset its own `FunctionToolset(timeout=...)`, or set `timeout=` on the individual tools. Tools coming from an [MCP server](https://pydantic.dev/docs/ai/mcp/client/) , an [external toolset](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/) , or a custom [`AbstractToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AbstractToolset) read neither; bound those with the server-side or transport-level timeout instead. If you enforce a deadline inside a tool body yourself, catch the `TimeoutError` and re-raise it as [`ModelRetry`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelRetry) or [`ToolFailed`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ToolFailed) rather than letting it escape. What happens to a bare `TimeoutError` depends on whether that tool has a timeout of its own: * **No `timeout` on the tool or its toolset.** It is an ordinary exception and propagates out of the agent run — unless a [capability](https://pydantic.dev/docs/ai/capabilities/overview/) implements `on_tool_execute_error`, which can turn it into a replacement tool result or a `ModelRetry`. * **A `timeout` is configured.** The call runs inside `anyio.fail_after(timeout)`, which signals expiry with `TimeoutError` too, so a `TimeoutError` you raised yourself is indistinguishable from the deadline expiring and becomes the same `'Timed out after N seconds.'` retry prompt — reporting a deadline that may never have passed. Re-raising in the tool is the more local choice; the hook is for applying one policy across every tool. Ending a run from inside a tool ------------------------------- [](https://pydantic.dev/docs/ai/core-concepts/timeouts/#ending-a-run-from-inside-a-tool) What a tool raises decides whether the run continues, and what the model gets to see: | Raise | Run continues? | The model sees | | --- | --- | --- | | [`ModelRetry`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelRetry) | Yes | A [retry prompt](https://pydantic.dev/docs/ai/core-concepts/retries/#tool-retries)
asking it to correct the call — consumes that tool’s retry budget | | [`ToolFailed`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ToolFailed) | Yes | A [failed tool result](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-failed)
to adapt to — does not consume the retry budget | | [`ApprovalRequired`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ApprovalRequired)
/ [`CallDeferred`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.CallDeferred) | Ends the run with a [`DeferredToolRequests`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.DeferredToolRequests)
output, unless a [`HandleDeferredToolCalls`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.HandleDeferredToolCalls)
handler resolves the call inline | Nothing yet — see [Deferred Tools](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/) | | Any other exception | No | By default nothing — it propagates out of `agent.run()`. A [capability](https://pydantic.dev/docs/ai/capabilities/overview/)
implementing `on_tool_execute_error` sees it first and can return a replacement tool result or raise `ModelRetry`, letting the run continue | The deferred row reads differently inside a [realtime session](https://pydantic.dev/docs/ai/realtime/overview/) , which has no way to pause: a live conversation can’t wait for an out-of-band result. A `HandleDeferredToolCalls` handler still gets the chance to resolve the call inline, but where a run would end with a `DeferredToolRequests` output, a session instead answers the model with an explanation that the tool can’t complete during the session, and keeps going. See [Deferred and approval-required tools](https://pydantic.dev/docs/ai/realtime/tools/#deferred-and-approval-required-tools) . A tool can also end the run without raising, by calling [`RunContext.cancel()`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.cancel) — the run ends with [`RunCancelled`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled) and the tool’s return value is discarded. See [Cancelling the Run from a Tool](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#cancelling-the-run-from-a-tool) . There is no exception that ends a run early with a _successful_ output. To let a tool finish the run with a value, make that value the run’s output: give the agent an [output tool](https://pydantic.dev/docs/ai/core-concepts/output/#tool-output) the model can call, or an [output function](https://pydantic.dev/docs/ai/core-concepts/output/#output-functions) that produces the result. Was this page helpful? Thanks for your feedback! --- # Third-Party Integrations | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/evaluators/framework-integrations/#_top) Third-Party Integrations ======================== Pydantic Evals does not take a hard dependency on any particular metrics framework. When a team already uses [Ragas](https://github.com/vibrantlabsai/ragas) , [DeepEval](https://github.com/confident-ai/deepeval) , or another scoring library, the [`Evaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Evaluator) base class makes it straightforward to wrap the upstream metric and run it inside any Pydantic Evals dataset. This page shows worked examples for the common ones. Pattern ------- [](https://pydantic.dev/docs/ai/evals/evaluators/framework-integrations/#pattern) Each framework integration follows the same pattern: 1. Subclass [`Evaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Evaluator) . 2. Adapt `ctx.inputs`, `ctx.output`, `ctx.expected_output`, and metadata into whatever the upstream metric expects. 3. Return a `float` score, a `bool` assertion, an [`EvaluationReason`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluationReason) , or a `dict` of these. The rest of this page shows concrete adapters. They are intentionally compact — extend them with whatever configuration your team needs (model selection, thresholds, per-case toggles). Ragas ----- [](https://pydantic.dev/docs/ai/evals/evaluators/framework-integrations/#ragas) Install with `pip install ragas` (not included in `pydantic-evals`). This adapter wraps [`ragas.metrics.Faithfulness`](https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/faithfulness/) for a single-turn sample. Each case is expected to provide the retrieved context as part of its inputs or metadata. from dataclasses import dataclass from ragas.dataset_schema import SingleTurnSample from ragas.metrics import Faithfulness from pydantic_evals.evaluators import EvaluationReason, Evaluator, EvaluatorContext @dataclass class RagasFaithfulness(Evaluator): """Wrap `ragas.metrics.Faithfulness` as a Pydantic Evals evaluator.""" context_field: str = 'context' async def evaluate(self, ctx: EvaluatorContext) -> EvaluationReason: metadata = ctx.metadata or {} retrieved_contexts = metadata.get(self.context_field, []) if isinstance(retrieved_contexts, str): retrieved_contexts = [retrieved_contexts] sample = SingleTurnSample( user_input=str(ctx.inputs), response=str(ctx.output), retrieved_contexts=retrieved_contexts, ) metric = Faithfulness() score = await metric.single_turn_ascore(sample) return EvaluationReason(value=float(score), reason=f'ragas.Faithfulness = {score:.3f}') Usage is the same as any built-in evaluator: from pydantic_evals import Case, Dataset dataset = Dataset( name='rag_eval', cases=[\ Case(\ inputs='What is the capital of France?',\ metadata={'context': ['Paris is the capital of France.']},\ ),\ ], evaluators=[RagasFaithfulness()], ) The same pattern works for `ragas.metrics.answer_relevancy`, `context_precision`, and the other scoring metrics: swap the metric class and (if needed) the sample fields. DeepEval -------- [](https://pydantic.dev/docs/ai/evals/evaluators/framework-integrations/#deepeval) Install with `pip install deepeval` (not included in `pydantic-evals`). This adapter wraps [DeepEval’s `GEval` metric](https://docs.confident-ai.com/docs/metrics-llm-evals) to score a criterion against a `LLMTestCase`. DeepEval’s `measure` is synchronous, so the evaluator is synchronous too. from dataclasses import dataclass from deepeval.metrics import GEval from deepeval.test_case import LLMTestCase, LLMTestCaseParams from pydantic_evals.evaluators import EvaluationReason, Evaluator, EvaluatorContext @dataclass class DeepEvalGEval(Evaluator): """Wrap `deepeval.metrics.GEval` as a Pydantic Evals evaluator.""" metric_name: str criteria: str threshold: float = 0.5 def evaluate(self, ctx: EvaluatorContext) -> dict[str, float | bool | EvaluationReason]: test_case = LLMTestCase( input=str(ctx.inputs), actual_output=str(ctx.output), expected_output=None if ctx.expected_output is None else str(ctx.expected_output), ) metric = GEval( name=self.metric_name, criteria=self.criteria, evaluation_params=[LLMTestCaseParams.INPUT, LLMTestCaseParams.ACTUAL_OUTPUT], threshold=self.threshold, ) metric.measure(test_case) return { f'{self.metric_name}_score': EvaluationReason(value=float(metric.score), reason=metric.reason or ''), f'{self.metric_name}_pass': bool(metric.success), } The same wrapper shape works for DeepEval’s `FaithfulnessMetric`, `AnswerRelevancyMetric`, `HallucinationMetric`, and others — swap the metric class and populate the relevant `LLMTestCase` fields (for example `retrieval_context` for faithfulness). Notes on dependencies --------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/framework-integrations/#notes-on-dependencies) * `ragas` and `deepeval` are optional dependencies — they are not installed with `pydantic-evals` and are not part of any dependency group. Install them only in projects that use these integrations. * Both libraries make their own LLM calls, so be prepared for extra API usage when running a dataset that includes these evaluators. Was this page helpful? Thanks for your feedback! --- # Logfire Integration | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#_top) Logfire Integration =================== Visualize and analyze evaluation results using Pydantic Logfire. Pydantic Evals uses OpenTelemetry to record traces of the evaluation process. These traces contain all the information from your evaluation reports, plus full tracing from the execution of your task function. You can send these traces to any OpenTelemetry-compatible backend, including [Pydantic Logfire](https://logfire.pydantic.dev/docs/guides/web-ui/evals/) . Installation ------------ [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#installation) Install the optional logfire dependency: Terminal pip install 'pydantic-evals[logfire]' Basic Setup ----------- [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#basic-setup) Configure Logfire before running evaluations: basic\_logfire\_setup.py import logfire from pydantic_evals import Case, Dataset # Configure Logfire logfire.configure( send_to_logfire='if-token-present', # (1) ) # Your evaluation code def my_task(inputs: str) -> str: return f'result for {inputs}' dataset = Dataset(name='logfire_demo', cases=[Case(name='test', inputs='example')]) report = dataset.evaluate_sync(my_task) Sends data to Logfire only if the `LOGFIRE_TOKEN` environment variable is set That’s it! Your evaluation traces will now appear in the Logfire web UI as long as you have the `LOGFIRE_TOKEN` environment variable set. What Gets Sent to Logfire ------------------------- [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#what-gets-sent-to-logfire) When you run an evaluation, Logfire receives: 1. **Evaluation metadata** 1. Dataset name 2. Number of cases 3. Evaluator names 2. **Per-case data** 1. Inputs and outputs 2. Expected outputs 3. Metadata 4. Execution duration 3. **Evaluation results** 1. Scores, assertions, and labels 2. Reasons (if included) 3. Evaluator failures 4. **Task execution traces** 1. All OpenTelemetry spans from your task function 2. Tool calls (for Pydantic AI agents) 3. API calls, database queries, etc. Viewing Results in Logfire -------------------------- [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#viewing-results-in-logfire) ### Evaluation Overview [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#evaluation-overview) Logfire provides a special table view for evaluation results on the root evaluation span: ![Logfire Evals Overview](https://pydantic.dev/docs/ai/img/logfire-evals-overview.png) This view shows: * Case names * Pass/fail status * Scores and assertions * Execution duration * Quick filtering and sorting ### Individual Case Details [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#individual-case-details) Click any case to see detailed inputs and outputs: ![Logfire Evals Case](https://pydantic.dev/docs/ai/img/logfire-evals-case.png) ### Full Trace View [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#full-trace-view) View the complete execution trace including all spans generated during evaluation: ![Logfire Evals Case Trace](https://pydantic.dev/docs/ai/img/logfire-evals-case-trace.png) This is especially useful for: * Debugging failed cases * Understanding performance bottlenecks * Analyzing tool usage patterns * Writing span-based evaluators Analyzing Traces ---------------- [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#analyzing-traces) ### Comparing Runs [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#comparing-runs) Run the same evaluation multiple times and compare in Logfire: from pydantic_evals import Case, Dataset def original_task(inputs: str) -> str: return f'original result for {inputs}' def improved_task(inputs: str) -> str: return f'improved result for {inputs}' dataset = Dataset(name='comparison', cases=[Case(name='test', inputs='example')]) # Run 1: Original implementation report1 = dataset.evaluate_sync(original_task) # Run 2: Improved implementation report2 = dataset.evaluate_sync(improved_task) # Compare in Logfire by filtering by timestamp or attributes ### Debugging Failed Cases [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#debugging-failed-cases) Find failed cases quickly: 1. Search for `service_name = 'my_service_evals' AND is_exception` (replace with the actual service name you are using) 2. View the full span tree to see where the failure occurred 3. Inspect attributes and logs for error messages Span-Based Evaluation --------------------- [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#span-based-evaluation) Logfire integration enables powerful span-based evaluators. See [Span-Based Evaluation](https://pydantic.dev/docs/ai/evals/evaluators/span-based/) for details. Example: Verify specific tools were called: import logfire from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import HasMatchingSpan logfire.configure(send_to_logfire='if-token-present') def my_agent(inputs: str) -> str: return f'result for {inputs}' dataset = Dataset( name='logfire_demo', cases=[Case(name='test', inputs='example')], evaluators=[\ HasMatchingSpan(\ query={'name_contains': 'search_tool'},\ evaluation_name='used_search',\ ),\ ], ) report = dataset.evaluate_sync(my_agent) The span tree is available in both: * Your evaluator code (via `ctx.span_tree`) * Logfire UI (visual trace view) Troubleshooting --------------- [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#troubleshooting) ### No Data Appearing in Logfire [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#no-data-appearing-in-logfire) Check: 1. **Token is set**: `echo $LOGFIRE_TOKEN` 2. **Configuration is correct**: import logfire logfire.configure(send_to_logfire='always') # Force sending 3. **Network connectivity**: Check firewall settings 4. **Project exists**: Verify project name in Logfire UI ### Traces Missing Spans [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#traces-missing-spans) If some spans are missing: 1. **Ensure logfire is configured before imports**: import logfire logfire.configure() # Must be first 2. **Check instrumentation**: Ensure your code has enabled all instrumentations you want: import logfire logfire.instrument_pydantic_ai() logfire.instrument_httpx(capture_all=True) Best Practices -------------- [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#best-practices) ### 1\. Configure Early [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#1-configure-early) Always configure Logfire before running evaluations: import logfire from pydantic_evals import Case, Dataset logfire.configure(send_to_logfire='if-token-present') # Now import and run evaluations def task(inputs: str) -> str: return f'result for {inputs}' dataset = Dataset(name='logfire_demo', cases=[Case(name='test', inputs='example')]) dataset.evaluate_sync(task) ### 2\. Use Descriptive Service Names And Environments [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#2-use-descriptive-service-names-and-environments) import logfire logfire.configure( service_name='rag-pipeline-evals', environment='development', ) ### 3\. Review Periodically [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#3-review-periodically) * Check Logfire regularly to identify patterns * Look for consistently failing cases * Analyze performance trends * Adjust evaluators based on insights Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/#next-steps) * **[Span-Based Evaluation](https://pydantic.dev/docs/ai/evals/evaluators/span-based/) ** - Use OpenTelemetry spans in evaluators * **[Logfire Documentation](https://logfire.pydantic.dev/docs/guides/web-ui/evals/) ** - Complete Logfire guide * **[Metrics & Attributes](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/) ** - Add custom data to traces Was this page helpful? Thanks for your feedback! --- # Dataset Serialization | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#_top) Dataset Serialization ===================== Learn how to save and load datasets in different formats, with support for custom evaluators and IDE integration. Pydantic Evals supports serializing datasets to files in two formats: * **YAML** (`.yaml`, `.yml`) - Human-readable, great for version control * **JSON** (`.json`) - Structured, machine-readable Both formats support: * Automatic JSON schema generation for IDE autocomplete and validation * Custom evaluator serialization/deserialization * Type-safe loading with generic parameters YAML Format ----------- [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#yaml-format) YAML is the recommended format for most use cases due to its readability and compact syntax. ### Basic Example [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#basic-example) from typing import Any from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import EqualsExpected, IsInstance # Create a dataset with typed parameters dataset = Dataset[str, str, Any]( name='my_tests', cases=[\ Case(\ name='test_1',\ inputs='hello',\ expected_output='HELLO',\ ),\ ], evaluators=[\ IsInstance(type_name='str'),\ EqualsExpected(),\ ], ) # Save to YAML dataset.to_file('my_tests.yaml') This creates two files: 1. **`my_tests.yaml`** - The dataset 2. **`my_tests_schema.json`** - JSON schema for IDE support ### YAML Output [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#yaml-output) # yaml-language-server: $schema=my_tests_schema.json name: my_tests cases: - name: test_1 inputs: hello expected_output: HELLO evaluators: - IsInstance: str - EqualsExpected ### JSON Schema for IDEs [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#json-schema-for-ides) The first line references the schema file: # yaml-language-server: $schema=my_tests_schema.json This enables: * ✅ **Autocomplete** in VS Code, PyCharm, and other editors * ✅ **Inline validation** while editing * ✅ **Documentation tooltips** for fields * ✅ **Error highlighting** for invalid data ### Loading from YAML [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#loading-from-yaml) from pathlib import Path from typing import Any from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import EqualsExpected, IsInstance # First create and save the dataset Path('my_tests.yaml').parent.mkdir(exist_ok=True) dataset = Dataset[str, str, Any]( name='my_tests', cases=[Case(name='test_1', inputs='hello', expected_output='HELLO')], evaluators=[IsInstance(type_name='str'), EqualsExpected()], ) dataset.to_file('my_tests.yaml') # Load the dataset with type parameters dataset = Dataset[str, str, Any].from_file('my_tests.yaml') def my_task(text: str) -> str: return text.upper() # Run evaluation report = dataset.evaluate_sync(my_task) JSON Format ----------- [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#json-format) JSON format is useful for programmatic generation or when strict structure is required. ### Basic Example [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#basic-example-1) from typing import Any from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import EqualsExpected dataset = Dataset[str, str, Any]( name='my_tests', cases=[\ Case(name='test_1', inputs='hello', expected_output='HELLO'),\ ], evaluators=[EqualsExpected()], ) # Save to JSON dataset.to_file('my_tests.json') ### JSON Output [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#json-output) { "$schema": "my_tests_schema.json", "name": "my_tests", "cases": [\ {\ "name": "test_1",\ "inputs": "hello",\ "expected_output": "HELLO"\ }\ ], "evaluators": [\ "EqualsExpected"\ ] } The `$schema` key at the top enables IDE support similar to YAML. ### Loading from JSON [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#loading-from-json) from typing import Any from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import EqualsExpected # First create and save the dataset dataset = Dataset[str, str, Any]( name='my_tests', cases=[Case(name='test_1', inputs='hello', expected_output='HELLO')], evaluators=[EqualsExpected()], ) dataset.to_file('my_tests.json') # Load from JSON dataset = Dataset[str, str, Any].from_file('my_tests.json') Schema Generation ----------------- [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#schema-generation) ### Automatic Schema Creation [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#automatic-schema-creation) By default, `to_file()` creates a JSON schema file alongside your dataset: from typing import Any from pydantic_evals import Case, Dataset dataset = Dataset[str, str, Any](name='my_tests', cases=[Case(inputs='test')]) # Creates both my_tests.yaml AND my_tests_schema.json dataset.to_file('my_tests.yaml') ### Custom Schema Location [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#custom-schema-location) from pathlib import Path from typing import Any from pydantic_evals import Case, Dataset dataset = Dataset[str, str, Any](name='my_tests', cases=[Case(inputs='test')]) # Create directories Path('data').mkdir(exist_ok=True) # Custom schema filename (relative to dataset file location) dataset.to_file( 'data/my_tests.yaml', schema_path='my_schema.json', ) # No schema file dataset.to_file('my_tests.yaml', schema_path=None) ### Schema Path Templates [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#schema-path-templates) Use `{stem}` to reference the dataset filename: from typing import Any from pydantic_evals import Case, Dataset dataset = Dataset[str, str, Any](name='my_tests', cases=[Case(inputs='test')]) # Creates: my_tests.yaml and my_tests.schema.json dataset.to_file( 'my_tests.yaml', schema_path='{stem}.schema.json', ) ### Manual Schema Generation [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#manual-schema-generation) Generate a schema without saving the dataset: import json from typing import Any from pydantic_evals import Dataset # Get schema as dictionary for a specific dataset type schema = Dataset[str, str, Any].model_json_schema_with_evaluators() # Save manually with open('custom_schema.json', 'w', encoding='utf-8') as f: json.dump(schema, f, indent=2) Custom Evaluators ----------------- [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#custom-evaluators) Custom evaluators require special handling during serialization and deserialization. ### Requirements [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#requirements) Custom evaluators must: 1. Be decorated with `@dataclass` 2. Inherit from `Evaluator` 3. Be passed to both `to_file()` and `from_file()` ### Complete Example [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#complete-example) from dataclasses import dataclass from typing import Any from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class CustomThreshold(Evaluator): """Check if output length exceeds a threshold.""" min_length: int max_length: int = 100 def evaluate(self, ctx: EvaluatorContext) -> bool: length = len(str(ctx.output)) return self.min_length <= length <= self.max_length # Create dataset with custom evaluator dataset = Dataset[str, str, Any]( name='custom_threshold_tests', cases=[\ Case(\ name='test_length',\ inputs='example',\ expected_output='long result',\ evaluators=[\ CustomThreshold(min_length=5, max_length=20),\ ],\ ),\ ], ) # Save with custom evaluator types dataset.to_file( 'dataset.yaml', custom_evaluator_types=[CustomThreshold], ) ### Saved YAML [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#saved-yaml) # yaml-language-server: $schema=dataset_schema.json cases: - name: test_length inputs: example expected_output: long result evaluators: - CustomThreshold: min_length: 5 max_length: 20 ### Loading with Custom Evaluators [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#loading-with-custom-evaluators) from dataclasses import dataclass from typing import Any from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class CustomThreshold(Evaluator): """Check if output length exceeds a threshold.""" min_length: int max_length: int = 100 def evaluate(self, ctx: EvaluatorContext) -> bool: length = len(str(ctx.output)) return self.min_length <= length <= self.max_length # First create and save the dataset dataset = Dataset[str, str, Any]( name='custom_threshold_tests', cases=[\ Case(\ name='test_length',\ inputs='example',\ expected_output='long result',\ evaluators=[CustomThreshold(min_length=5, max_length=20)],\ ),\ ], ) dataset.to_file('dataset.yaml', custom_evaluator_types=[CustomThreshold]) # Load with custom evaluator registry dataset = Dataset[str, str, Any].from_file( 'dataset.yaml', custom_evaluator_types=[CustomThreshold], ) Evaluator Serialization Formats ------------------------------- [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#evaluator-serialization-formats) Evaluators can be serialized in three forms: ### 1\. Name Only (No Parameters) [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#1-name-only-no-parameters) evaluators: - EqualsExpected - IsInstance: str # Using default parameter ### 2\. Single Parameter (Short Form) [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#2-single-parameter-short-form) evaluators: - IsInstance: str - Contains: "required text" - MaxDuration: 2.0 ### 3\. Multiple Parameters (Dict Form) [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#3-multiple-parameters-dict-form) evaluators: - CustomThreshold: min_length: 5 max_length: 20 - LLMJudge: rubric: "Response is accurate" model: "openai:gpt-5" include_input: true Format Comparison ----------------- [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#format-comparison) | Feature | YAML | JSON | | --- | --- | --- | | Human readable | ✅ Excellent | ⚠️ Good | | Comments | ✅ Yes | ❌ No | | Compact | ✅ Yes | ⚠️ Verbose | | Machine parsing | ✅ Good | ✅ Excellent | | IDE support | ✅ Yes | ✅ Yes | | Version control | ✅ Clean diffs | ⚠️ Noisy diffs | **Recommendation**: Use YAML for most cases, JSON for programmatic generation. Advanced: Evaluator Serialization Name -------------------------------------- [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#advanced-evaluator-serialization-name) Customize how your evaluator appears in serialized files: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class VeryLongDescriptiveEvaluatorName(Evaluator): @classmethod def get_serialization_name(cls) -> str: return 'ShortName' def evaluate(self, ctx: EvaluatorContext) -> bool: return True In YAML: evaluators: - ShortName # Instead of VeryLongDescriptiveEvaluatorName Troubleshooting --------------- [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#troubleshooting) ### Schema Not Found in IDE [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#schema-not-found-in-ide) **Problem**: YAML file doesn’t show autocomplete **Solutions**: 1. **Check the schema path** in the first line of YAML: # yaml-language-server: $schema=correct_schema_name.json 2. **Verify schema file exists** in the same directory 3. **Restart the language server** in your IDE 4. **Install YAML extension** (VS Code: “YAML” by Red Hat) ### Custom Evaluator Not Found [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#custom-evaluator-not-found) **Problem**: `ValueError: Unknown evaluator name: 'CustomEvaluator'` **Solution**: Pass `custom_evaluator_types` when loading: from dataclasses import dataclass from typing import Any from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class CustomEvaluator(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return True # First create and save with custom evaluator dataset = Dataset[str, str, Any]( name='custom_eval_tests', cases=[Case(inputs='test', evaluators=[CustomEvaluator()])], ) dataset.to_file('tests.yaml', custom_evaluator_types=[CustomEvaluator]) # Load with custom evaluator types dataset = Dataset[str, str, Any].from_file( 'tests.yaml', custom_evaluator_types=[CustomEvaluator], # Required! ) ### Format Inference Failed [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#format-inference-failed) **Problem**: `ValueError: Cannot infer format from extension` **Solution**: Specify format explicitly: from typing import Any from pydantic_evals import Case, Dataset dataset = Dataset[str, str, Any](name='my_tests', cases=[Case(inputs='test')]) # Explicit format for unusual extensions dataset.to_file('data.txt', fmt='yaml') dataset_loaded = Dataset[str, str, Any].from_file('data.txt', fmt='yaml') ### Schema Generation Error [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#schema-generation-error) **Problem**: Custom evaluator causes schema generation to fail **Solution**: Ensure evaluator is a proper dataclass: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext # ✅ Correct @dataclass class MyEvaluator(Evaluator): value: int def evaluate(self, ctx: EvaluatorContext) -> bool: return True # ❌ Wrong: Missing @dataclass class BadEvaluator(Evaluator): def __init__(self, value: int): self.value = value def evaluate(self, ctx: EvaluatorContext) -> bool: return True Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/#next-steps) * **[Dataset Management](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/) ** - Creating and organizing datasets * **[Custom Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/custom/) ** - Write custom evaluation logic * **[Core Concepts](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/) ** - Understand the data model Was this page helpful? Thanks for your feedback! --- # Setup | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/setup/#_top) Setup ===== Here we include some examples of how to use Pydantic AI and what it can do. Usage ----- [](https://pydantic.dev/docs/ai/examples/setup/#usage) These examples are distributed with `pydantic-ai` so you can run them either by cloning the [pydantic-ai repo](https://github.com/pydantic/pydantic-ai) or by simply installing `pydantic-ai` from PyPI with `pip` or `uv`. ### Installing required dependencies [](https://pydantic.dev/docs/ai/examples/setup/#installing-required-dependencies) Either way you’ll need to install extra dependencies to run some examples, you just need to install the `examples` optional dependency group. If you’ve installed `pydantic-ai` via pip/uv, you can install the extra dependencies with: * [pip](https://pydantic.dev/docs/ai/examples/setup/#tab-panel-56) * [uv](https://pydantic.dev/docs/ai/examples/setup/#tab-panel-57) Terminal pip install "pydantic-ai[examples]" Terminal uv add "pydantic-ai[examples]" If you clone the repo, you should instead use `uv sync --extra examples` to install extra dependencies. ### Setting model environment variables [](https://pydantic.dev/docs/ai/examples/setup/#setting-model-environment-variables) These examples will need you to set up authentication with one or more of the LLMs, see the [model configuration](https://pydantic.dev/docs/ai/models/overview/) docs for details on how to do this. TL;DR: in most cases you’ll need to set one of the following environment variables: * [OpenAI](https://pydantic.dev/docs/ai/examples/setup/#tab-panel-58) * [Google Gemini](https://pydantic.dev/docs/ai/examples/setup/#tab-panel-59) Terminal export OPENAI_API_KEY=your-api-key Terminal export GEMINI_API_KEY=your-api-key ### Running Examples [](https://pydantic.dev/docs/ai/examples/setup/#running-examples) To run the examples (this will work whether you installed `pydantic_ai`, or cloned the repo), run: * [pip](https://pydantic.dev/docs/ai/examples/setup/#tab-panel-60) * [uv](https://pydantic.dev/docs/ai/examples/setup/#tab-panel-61) Terminal python -m pydantic_ai_examples. Terminal uv run -m pydantic_ai_examples. For example, to run the very simple [`pydantic_model`](https://pydantic.dev/docs/ai/examples/getting-started/pydantic-model/) example: * [pip](https://pydantic.dev/docs/ai/examples/setup/#tab-panel-62) * [uv](https://pydantic.dev/docs/ai/examples/setup/#tab-panel-63) Terminal python -m pydantic_ai_examples.pydantic_model Terminal uv run -m pydantic_ai_examples.pydantic_model If you like one-liners and you’re using uv, you can run a pydantic-ai example with zero setup: Terminal OPENAI_API_KEY='your-api-key' \ uv run --with "pydantic-ai[examples]" \ -m pydantic_ai_examples.pydantic_model * * * You’ll probably want to edit examples in addition to just running them. You can copy the examples to a new directory with: * [pip](https://pydantic.dev/docs/ai/examples/setup/#tab-panel-64) * [uv](https://pydantic.dev/docs/ai/examples/setup/#tab-panel-65) Terminal python -m pydantic_ai_examples --copy-to examples/ Terminal uv run -m pydantic_ai_examples --copy-to examples/ Was this page helpful? Thanks for your feedback! --- # Stream Whales | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/streaming/stream-whales/#_top) Stream Whales ============= Information about whales — an example of streamed structured response validation. Demonstrates: * [streaming structured output](https://pydantic.dev/docs/ai/core-concepts/output/#streaming-structured-output) This script streams structured responses about whales, validates the data and displays it as a dynamic table using [`rich`](https://github.com/Textualize/rich) as the data is received. Running the Example ------------------- [](https://pydantic.dev/docs/ai/examples/streaming/stream-whales/#running-the-example) With [dependencies installed and environment variables set](https://pydantic.dev/docs/ai/examples/setup/#usage) , run: * [pip](https://pydantic.dev/docs/ai/examples/streaming/stream-whales/#tab-panel-68) * [uv](https://pydantic.dev/docs/ai/examples/streaming/stream-whales/#tab-panel-69) Terminal python -m pydantic_ai_examples.stream_whales Terminal uv run -m pydantic_ai_examples.stream_whales Should give an output like this: Example Code ------------ [](https://pydantic.dev/docs/ai/examples/streaming/stream-whales/#example-code) stream\_whales.py from typing import Annotated import logfire from pydantic import Field from rich.console import Console from rich.live import Live from rich.table import Table from typing_extensions import NotRequired, TypedDict from pydantic_ai import Agent # 'if-token-present' means nothing will be sent (and the example will work) if you don't have logfire configured logfire.configure(send_to_logfire='if-token-present') logfire.instrument_pydantic_ai() class Whale(TypedDict): name: str length: Annotated[\ float, Field(description='Average length of an adult whale in meters.')\ ] weight: NotRequired[\ Annotated[\ float,\ Field(description='Average weight of an adult whale in kilograms.', ge=50),\ ]\ ] ocean: NotRequired[str] description: NotRequired[Annotated[str, Field(description='Short Description')]] agent = Agent('openai:gpt-5.2', output_type=list[Whale]) async def main(): console = Console() with Live('\n' * 36, console=console) as live: console.print('Requesting data...', style='cyan') async with agent.run_stream( 'Generate me details of 5 species of Whale.' ) as result: console.print('Response:', style='green') async for whales in result.stream_output(debounce_by=0.01): table = Table( title='Species of Whale', caption='Streaming Structured responses from OpenAI', width=120, ) table.add_column('ID', justify='right') table.add_column('Name') table.add_column('Avg. Length (m)', justify='right') table.add_column('Avg. Weight (kg)', justify='right') table.add_column('Ocean') table.add_column('Description', justify='right') for wid, whale in enumerate(whales, start=1): table.add_row( str(wid), whale['name'], f'{whale["length"]:0.0f}', f'{w:0.0f}' if (w := whale.get('weight')) else '…', whale.get('ocean') or '…', whale.get('description') or '…', ) live.update(table) if __name__ == '__main__': import asyncio asyncio.run(main()) Was this page helpful? Thanks for your feedback! --- # Version Policy | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/project/version-policy/#_top) Version Policy ============== Pydantic AI V1 was released in September 2025, and the stable V2.0 was released on June 23, 2026; see the [Upgrade Guide](https://pydantic.dev/docs/ai/project/changelog/) for what’s in V2, how to install it, and how to upgrade. We will not intentionally make breaking changes in minor releases. Functionality marked as deprecated in a release is not removed until the next major version, which we won’t release sooner than 3 months after V2.0. We’ll continue to provide security fixes for V1 for at least 6 months after V2’s stable release, so you have time to upgrade your applications. When you’re ready to make the jump, the [Upgrade Guide](https://pydantic.dev/docs/ai/project/changelog/) lists the breaking changes for each version, along with our recommended path to V2. Of course, some apparently safe changes and bug fixes will inevitably break some users’ code — obligatory link to [xkcd](https://xkcd.com/1172/) . The following changes will **NOT** be considered breaking changes, and may occur in minor releases: * Bug fixes that may result in existing code breaking, provided that such code was relying on undocumented features/constructs/assumptions. * Adding new [message parts](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages) , [stream events](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.AgentStreamEvent) , or optional fields (including fields with default values) on existing message (part) and event types. Always code defensively when consuming message parts or event streams, and use the [`ModelMessagesTypeAdapter`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessagesTypeAdapter) to (de)serialize message histories. * Changing OpenTelemetry span attributes. Because different [observability platforms](https://pydantic.dev/docs/ai/integrations/logfire/#using-opentelemetry) support different versions of the [OpenTelemetry Semantic Conventions for Generative AI systems](https://opentelemetry.io/docs/specs/semconv/gen-ai/) , Pydantic AI lets you configure the [instrumentation version](https://pydantic.dev/docs/ai/integrations/logfire/#configuring-data-format) , but the default version may change in a minor release. Span attributes for [Pydantic Evals](https://pydantic.dev/docs/ai/evals/evals/) may also change as we iterate on Evals support in [Pydantic Logfire](https://logfire.pydantic.dev/docs/guides/web-ui/evals/) . * Changing how `__repr__` behaves, even of public classes. In all cases we will aim to minimize churn and do so only when justified by the increase of quality of Pydantic AI for users. Beta Features ------------- [](https://pydantic.dev/docs/ai/project/version-policy/#beta-features) At Pydantic, we like to move quickly and innovate! To that end, minor releases may introduce beta features (indicated by a `beta` module) that are active works in progress. While in its beta phase, a feature’s API and behaviors may not be stable, and it’s very possible that changes made to the feature will not be backward-compatible. We aim to move beta features out of beta within a few months after initial release, once users have had a chance to provide feedback and test the feature in production. Support for Python versions --------------------------- [](https://pydantic.dev/docs/ai/project/version-policy/#support-for-python-versions) Pydantic will drop support for a Python version when the following conditions are met: * The Python version has reached its [expected end of life](https://devguide.python.org/versions/) . * less than 5% of downloads of the most recent minor release are using that version. Was this page helpful? Thanks for your feedback! --- # Browser WebRTC | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/realtime/realtime-webrtc/#_top) Browser WebRTC ============== This example is a browser voice agent where the **browser exchanges audio with the provider (OpenAI or Azure OpenAI) directly over WebRTC** (lowest latency), while a [Pydantic AI sideband](https://pydantic.dev/docs/ai/realtime/deployment/#browser-webrtc-server-sideband) on the server runs the agent’s tools, builds message history, and keeps the API key off the client. It’s the recommended topology for browser voice agents: the server never sits in the audio path — it is the control plane. browser ──mic/speaker audio (WebRTC media)──▶ OpenAI / Azure OpenAI Realtime ◀───────────────────────────────────── │ SDP offer (POST /offer) ▲ control WebSocket (call_id) ▼ │ FastAPI backend ──answer_webrtc_offer()──▶ provider ──session(provider_session=…)──┘ (relays the SDP, gets a call_id) (runs tools, builds history) Demonstrates: * [browser WebRTC + server sideband](https://pydantic.dev/docs/ai/realtime/deployment/#browser-webrtc-server-sideband) with the [OpenAI](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModel) and [Azure OpenAI](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModel) providers * [`AgentRealtime.answer_webrtc_offer`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentRealtime.answer_webrtc_offer) — relaying the browser’s SDP offer server-side, with the agent’s instructions and tools baked in, so the API key never reaches the client * [`agent.realtime(model).session(provider_session=…)`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentRealtime.session) — running the agent’s [tools](https://pydantic.dev/docs/ai/realtime/tools/) over the call’s control plane while the browser owns the audio Running the Example ------------------- [](https://pydantic.dev/docs/ai/examples/realtime/realtime-webrtc/#running-the-example) You’ll need an `OPENAI_API_KEY` with realtime access, in a `.env` file at the repo root: OPENAI_API_KEY=... To run against **Azure OpenAI** instead, point `WEBRTC_REALTIME_MODEL` at your realtime deployment: WEBRTC_REALTIME_MODEL=azure:gpt-realtime AZURE_OPENAI_ENDPOINT=https://my-resource.openai.azure.com AZURE_OPENAI_API_KEY=... With [dependencies installed and your key set](https://pydantic.dev/docs/ai/examples/setup/#usage) , start the server: Terminal uv run --all-packages uvicorn pydantic_ai_examples.realtime_webrtc.app:app Open [http://localhost:8000](http://localhost:8000/) , click **Start call**, allow microphone access, and ask “What time is it in Tokyo?” or “What’s your refund policy?” to trigger a server-side tool. Overrides: `WEBRTC_REALTIME_MODEL` (default `openai:gpt-realtime`), `WEBRTC_REALTIME_VOICE` (default `marin`), `WEBRTC_TRANSCRIPTION_MODEL` (default: the provider’s `'auto'` choice). Example Code ------------ [](https://pydantic.dev/docs/ai/examples/realtime/realtime-webrtc/#example-code) The server — it relays the SDP offer to OpenAI, attaches the sideband session, and runs the tools: app.py from __future__ import annotations import asyncio import os from contextlib import asynccontextmanager, suppress from dataclasses import dataclass, field from datetime import datetime from pathlib import Path from zoneinfo import ZoneInfo, ZoneInfoNotFoundError import logfire from dotenv import load_dotenv from fastapi import FastAPI, HTTPException, Request from fastapi.responses import HTMLResponse, JSONResponse from pydantic_ai import Agent from pydantic_ai.messages import FunctionToolCallEvent, FunctionToolResultEvent from pydantic_ai.realtime import ( RealtimeTurnCompleteEvent, WebRTCSession, infer_realtime_model, ) from pydantic_ai.realtime.openai import OpenAIRealtimeModelSettings load_dotenv() logfire.configure(send_to_logfire='if-token-present', service_name='realtime-webrtc') logfire.instrument_pydantic_ai() VOICE = os.getenv('WEBRTC_REALTIME_VOICE', 'marin') INSTRUCTIONS = ( 'You are Roberto, a concise and friendly voice support assistant. ' 'Use `lookup_time` for time questions and `lookup_support_policy` for account or refund questions. ' 'Keep answers short and natural for speech.' ) INDEX_HTML = (Path(__file__).parent / 'index.html').read_text(encoding='utf-8') agent = Agent(instructions=INSTRUCTIONS) @agent.tool_plain def lookup_time(city: str) -> str: """Look up the current local time for a city.""" timezones = { 'london': 'Europe/London', 'new york': 'America/New_York', 'tokyo': 'Asia/Tokyo', 'sydney': 'Australia/Sydney', 'san francisco': 'America/Los_Angeles', } zone = timezones.get(city.lower()) if zone is None: return f'I only know these example cities: {", ".join(sorted(timezones))}.' try: now = datetime.now(ZoneInfo(zone)) except ZoneInfoNotFoundError: # pragma: no cover - depends on the host tz database return f'I could not load timezone data for {city}.' return now.strftime(f'It is %A, %I:%M %p in {city}.') @agent.tool_plain def lookup_support_policy(topic: str) -> str: """Return a short canned support policy answer.""" policies = { 'refund': 'Refunds are available within 30 days for billing errors or duplicate charges.', 'return': 'Physical returns can be started within 14 days of the delivery date.', 'password': 'Reset your password from the sign-in page using the email verification flow.', } return policies.get( topic.lower(), 'I only have example policies for refund, return, and password.' ) model = infer_realtime_model(os.getenv('WEBRTC_REALTIME_MODEL', 'openai:gpt-realtime')) settings = OpenAIRealtimeModelSettings(openai_voice=VOICE) if transcription_model := os.getenv('WEBRTC_TRANSCRIPTION_MODEL'): settings['input_transcription_model'] = transcription_model realtime = agent.realtime(model, model_settings=settings) @dataclass class Call: """One live WebRTC call and its server-side sideband task.""" answer_sdp: str provider_session: WebRTCSession task: asyncio.Task[None] | None = None # Set once the sideband has either attached or failed to; `attach_error` distinguishes the two so # `/offer` doesn't return a successful answer for a session that never came up. attached: asyncio.Event = field(default_factory=asyncio.Event) attach_error: BaseException | None = None CALLS: dict[str, Call] = {} async def run_sideband(call: Call) -> None: """Attach the sideband session to the WebRTC call and run the agent's tool loop over its events.""" call_id = call.provider_session.call_id try: async with realtime.session(provider_session=call.provider_session) as session: call.attached.set() async for event in session: if isinstance(event, FunctionToolCallEvent): logfire.info( 'tool call', tool=event.part.tool_name, args=event.part.args ) elif isinstance(event, FunctionToolResultEvent): logfire.info( 'tool result', tool=event.part.tool_name, content=event.part.content, ) elif isinstance(event, RealtimeTurnCompleteEvent): logfire.info('turn complete', messages=len(session.all_messages())) except asyncio.CancelledError: raise except Exception as exc: logfire.exception('sideband session for {call_id} failed', call_id=call_id) # Record the failure so `/offer` can surface it instead of returning a dead call. call.attach_error = exc call.attached.set() finally: CALLS.pop(call_id, None) @asynccontextmanager async def lifespan(_app: FastAPI): try: yield finally: for call in list(CALLS.values()): if call.task is not None: call.task.cancel() with suppress(asyncio.CancelledError): await call.task app = FastAPI(lifespan=lifespan) @app.get('/') async def index() -> HTMLResponse: return HTMLResponse(INDEX_HTML) @app.post('/offer') async def offer(request: Request) -> JSONResponse: """Relay the browser's SDP offer to the provider, start the sideband, and return the SDP answer.""" try: sdp_offer = (await request.body()).decode('utf-8') except UnicodeDecodeError: # The SDP offer is untrusted signaling input; reject malformed bytes as a client error, not a 500. raise HTTPException( status_code=400, detail='Expected a UTF-8 SDP offer in the request body.' ) from None if not sdp_offer.strip(): raise HTTPException( status_code=400, detail='Expected an SDP offer in the request body.' ) answer = await realtime.answer_webrtc_offer(sdp_offer) call = Call(answer_sdp=answer.sdp, provider_session=answer.session) CALLS[answer.session.call_id] = call # Attach the sideband before returning the answer, so the tools are live before the browser (which # only starts sending audio once it has the answer) can speak. call.task = asyncio.create_task(run_sideband(call)) try: await asyncio.wait_for(call.attached.wait(), timeout=10) except asyncio.TimeoutError: call.task.cancel() CALLS.pop(answer.session.call_id, None) raise HTTPException( status_code=504, detail='Timed out attaching the server-side session.' ) except asyncio.CancelledError: # The client disconnected before receiving the answer, so it never got the `call_id` and can't # call `/hangup`. Cancel the sideband and drop the call here to avoid leaking the provider # connection and the background agent task. call.task.cancel() CALLS.pop(answer.session.call_id, None) raise if call.attach_error is not None: raise HTTPException( status_code=502, detail='The server-side session failed to attach.' ) return JSONResponse({'sdp': call.answer_sdp, 'call_id': answer.session.call_id}) @app.post('/hangup/{call_id}') async def hangup(call_id: str) -> JSONResponse: call = CALLS.get(call_id) if call is not None and call.task is not None: call.task.cancel() with suppress(asyncio.CancelledError): await call.task return JSONResponse({'stopped': call is not None}) def main() -> None: # pragma: no cover - manual entrypoint import uvicorn uvicorn.run(app, host='127.0.0.1', port=8000) if __name__ == '__main__': # pragma: no cover main() The browser — it captures the microphone, negotiates WebRTC through the backend, and plays the audio back. Plain HTML and JavaScript, no build step: index.html Pydantic AI — Realtime WebRTC voice agent

Realtime WebRTC voice agent

The browser exchanges audio with the provider directly over WebRTC. The Python backend negotiates the call, attaches a Pydantic AI sideband session, and runs the tools server-side.

Try: "What time is it in Tokyo?" or "What's your refund policy?"

Idle
Was this page helpful? Thanks for your feedback! --- # Connection lifecycle | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/realtime/lifecycle/#_top) Connection lifecycle ==================== A realtime model uses one persistent provider connection. Your backend owns that session and the media bridge to the user (see [Connecting a frontend](https://pydantic.dev/docs/ai/realtime/deployment/) ); a reconnect policy can recover dropped connections and provider session limits without changing the application event loop. The session lifecycle --------------------- [](https://pydantic.dev/docs/ai/realtime/lifecycle/#the-session-lifecycle) stateDiagram-v2 [*] --> Connecting: session() opens Connecting --> Listening: handshake complete Listening --> UserTurn: speech detected /
audio committed UserTurn --> ModelResponse: turn detection /
create_response() ModelResponse --> ToolCalls: model calls a tool ToolCalls --> ModelResponse: result returned ModelResponse --> Listening: turn complete Listening --> Reconnecting: connection drops ModelResponse --> Reconnecting: connection drops Reconnecting --> Listening: redial succeeds Reconnecting --> [*]: attempts exhausted Listening --> [*]: close() Opening the session performs the provider handshake, after which the session listens for input. [Turn detection](https://pydantic.dev/docs/ai/realtime/turns/) (or manual [push-to-talk](https://pydantic.dev/docs/ai/realtime/turns/#push-to-talk) control) moves a user turn into a model response, which may loop through [tool calls](https://pydantic.dev/docs/ai/realtime/tools/) before [`RealtimeTurnCompleteEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeTurnCompleteEvent) marks the [turn boundary](https://pydantic.dev/docs/ai/realtime/events/#the-turn-boundary) and the session listens again. A dropped connection enters the reconnect loop below — emitting [`RealtimeSessionReconnectEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSessionReconnectEvent) on recovery — until [`close()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.close) (or leaving the `async with` block) ends the session. Connection and handshake ------------------------ [](https://pydantic.dev/docs/ai/realtime/lifecycle/#connection-and-handshake) The connection is opened when the `session()` context is entered, and the shared `handshake_timeout` setting (default 30 seconds) bounds how long the session waits for each realtime protocol handshake event on providers with an explicit handshake (OpenAI, Azure OpenAI, and xAI). A handshake that times out raises [`RealtimeError`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeError) ; a rejected WebSocket upgrade raises [`ModelHTTPError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelHTTPError) (see [Errors](https://pydantic.dev/docs/ai/realtime/lifecycle/#errors) ). Reconnecting ------------ [](https://pydantic.dev/docs/ai/realtime/lifecycle/#reconnecting) Set the `reconnect` [shared setting](https://pydantic.dev/docs/ai/realtime/overview/#shared-settings) to a [`ReconnectPolicy`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.ReconnectPolicy) to redial with exponential backoff, reapply configuration, and emit [`RealtimeSessionReconnectEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSessionReconnectEvent) . Like any realtime model setting, it can be a default on the model or passed for one session: from pydantic_ai import Agent agent = Agent() realtime = agent.realtime( 'openai:gpt-realtime', model_settings={'reconnect': {'max_attempts': 5}}, ) `max_attempts` bounds retries for one drop. `max_reconnects` bounds recoveries across the entire session, preventing an endpoint that repeatedly accepts and closes connections from redialing forever. Without a policy, an unexpected provider close raises [`RealtimeError`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeError) from the session iterator. On a [WebRTC sideband](https://pydantic.dev/docs/ai/realtime/deployment/#browser-webrtc-server-sideband) the same policy applies to an unexpected drop, but a _clean_ close is treated as the browser hanging up: the sideband is a control channel, so a normal close ends iteration without a session error or reconnect attempt even when a `reconnect` policy is set. The close frame alone can’t distinguish a hangup from a WebSocket-terminating proxy closing the sideband cleanly mid-call (a restart or graceful rotation), which would end the agent side while the browser keeps talking to the provider — drain such connections at the infrastructure layer rather than relying on the `reconnect` policy to cover them. ### State restoration [](https://pydantic.dev/docs/ai/realtime/lifecycle/#state-restoration) OpenAI and Azure OpenAI have no cross-connection server state, so Pydantic AI replays local message history into the new session. Prior transcript turns survive; in-flight audio does not. Gemini and xAI use native in-process session resumption, enabled automatically when a `reconnect` policy is present (an explicit `google_enable_session_resumption=False` alongside a policy raises [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) instead of silently losing the conversation); see the [Gemini resumption settings](https://pydantic.dev/docs/ai/realtime/gemini/#session-resumption) . Their handles live only in memory and cannot be persisted for another process. [`RealtimeSessionReconnectEvent.state_restored`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSessionReconnectEvent.state_restored) reports whether the reconnect carried the conversation through without cutting a turn off. How a reply the drop caught in flight is handled depends on the mechanism. Under native resumption (xAI) the recorded response simply stays open: output on the new connection continues it, the turn completes with the response terminal as usual, and `state_restored` stays `True`. Gemini also reports `True` but closes the cut reply as an interrupted response (keeping any partial transcript in history) before the [`RealtimeSessionReconnectEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSessionReconnectEvent) and stays quiet until the next input. Local replay (OpenAI, Azure OpenAI) restores only the finalized turns, so a reply in flight when the socket dropped cannot continue. The session settles it before emitting the event — the partial reply becomes an interrupted response, running tool calls get cancelled returns, and the turn ends so queued messages waiting for the boundary still flush — and `state_restored` is `False` to say the turn was cut off. An answer that was solicited but had not started streaming is instead re-requested on the new connection, and a drop with nothing in flight restores the whole call as-is; both lose no output, so `state_restored` stays `True`. Provider session limits ----------------------- [](https://pydantic.dev/docs/ai/realtime/lifecycle/#provider-session-limits) Providers cap individual connection duration. A reconnect policy is also how an application survives those limits. Exact limits and provider behavior can change, so provider pages are canonical: * [OpenAI session behavior](https://pydantic.dev/docs/ai/realtime/openai/#feature-support-and-limitations) * [Azure OpenAI session behavior](https://pydantic.dev/docs/ai/realtime/azure/#feature-support-and-limitations) * [Gemini session resumption](https://pydantic.dev/docs/ai/realtime/gemini/#session-resumption) * [xAI native session resumption](https://pydantic.dev/docs/ai/realtime/xai/#session-resumption) Gemini sends `GoAway` shortly before its cap but Pydantic AI currently reconnects only after the connection drops, so a long call can briefly drop mid-turn. Errors ------ [](https://pydantic.dev/docs/ai/realtime/lifecycle/#errors) Realtime sessions use the standard Pydantic AI exception hierarchy: | Exception | Raised when | | --- | --- | | [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) | The application requests an unsupported operation, passes incompatible settings, lacks credentials, or misuses the session. | | [`ModelHTTPError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelHTTPError) | The provider rejects the WebSocket upgrade with an HTTP status. | | [`RealtimeError`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeError) | The connection fails, times out, closes unexpectedly, returns an invalid frame, or exhausts reconnect attempts. | | [`UsageLimitExceeded`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UsageLimitExceeded) | A configured [usage limit](https://pydantic.dev/docs/ai/realtime/observability/#usage-and-limits)
is exceeded. | [`RealtimeError`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeError) subclasses [`ModelAPIError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelAPIError) , so `except ModelAPIError` covers HTTP and non-HTTP provider failures together. Recoverable failures arrive as events: [`RealtimeSessionErrorEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSessionErrorEvent) for provider operations and [`RealtimeInputTranscriptionErrorEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeInputTranscriptionErrorEvent) for one failed user transcription. The session remains usable after either event. Failures surface from the responsible call where possible; a failed `send_audio()` raises there. Receive-loop and tool failures propagate from session iteration. For symptom-first debugging, see [Troubleshooting](https://pydantic.dev/docs/ai/realtime/troubleshooting/) . Was this page helpful? Thanks for your feedback! --- # Connecting a frontend | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/realtime/deployment/#_top) Connecting a frontend ===================== Keep provider keys, tools, and business logic on the server; connect user devices to your backend, not straight to the provider. How audio travels between the device and the model depends on the client and the provider: * **[Browser WebRTC + server sideband](https://pydantic.dev/docs/ai/realtime/deployment/#browser-webrtc-server-sideband) ** — the recommended browser path on OpenAI and Azure OpenAI: the browser exchanges media directly with the provider (lowest latency) while your backend runs the agent over a control-plane sideband. * **[Browser → backend WebSocket relay](https://pydantic.dev/docs/ai/realtime/deployment/#browser-backend-websocket-relay) ** — works with every provider, including Gemini Live and xAI: the browser streams audio to your backend, which owns the media bridge. * **[SIP / telephony bridge](https://pydantic.dev/docs/ai/realtime/deployment/#siptelephony-bridge) ** — for phone calls, via a telephony provider. In every shape the session — with its [tools](https://pydantic.dev/docs/ai/realtime/tools/) , [history](https://pydantic.dev/docs/ai/realtime/history/) , and [usage limits](https://pydantic.dev/docs/ai/realtime/observability/#usage-and-limits) — runs on your backend. Wiring a browser straight to the provider with the provider’s own SDK instead moves the agent loop into the client and gives up all of that; prefer the shapes above. Browser WebRTC + server sideband -------------------------------- [](https://pydantic.dev/docs/ai/realtime/deployment/#browser-webrtc--server-sideband) For browser voice agents on OpenAI and Azure OpenAI, the browser carries microphone and speaker audio directly over WebRTC while the backend attaches a control-plane **sideband** to the same call. Media never touches your backend, so latency stays low, while tools, history, dependencies, and provider credentials stay on the server. Gemini Live and xAI do not offer this sideband transport — use the [WebSocket relay](https://pydantic.dev/docs/ai/realtime/deployment/#browser-backend-websocket-relay) there. browser ── WebRTC media ── provider │ ▲ └─ SDP offer → backend ───┘ sideband identified by call_id Relay the browser’s offer with [`AgentRealtime.answer_webrtc_offer`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentRealtime.answer_webrtc_offer) , return the SDP answer to the browser, then attach the returned call handle: import asyncio from pydantic_ai import Agent agent = Agent(instructions='You are a concise voice assistant.') realtime = agent.realtime('openai:gpt-realtime') async def handle_offer(sdp_offer: str) -> str: answer = await realtime.answer_webrtc_offer(sdp_offer) async def run_sideband() -> None: async with realtime.session(provider_session=answer.session) as session: async for event in session: print(event) asyncio.create_task(run_sideband()) return answer.sdp The secure offer-relay flow never gives the browser a token. As an alternative, [`AgentRealtime.create_client_secret`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentRealtime.create_client_secret) mints a short-lived credential for client-led negotiation. Either way the browser is a peer on the provider session and can send provider-native control events, so authorize every server-side tool against trusted [`deps`](https://pydantic.dev/docs/ai/core-concepts/dependencies/) , not session instructions supplied to the model. A sideband does not own the audio transport: its `send_audio()`, `commit_audio()`, `clear_audio()`, and `stream_audio()` methods raise, and `audio_retention` must remain `'transcript_only'`. Enable [input transcription](https://pydantic.dev/docs/ai/realtime/audio/#input-transcription) when user speech must appear in history, and use [`RealtimeOutputSpeechStartEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeOutputSpeechStartEvent) / [`RealtimeOutputSpeechEndEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeOutputSpeechEndEvent) for speaking indicators — [`interrupt()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.interrupt) clears the provider’s outbound WebRTC audio buffer so [barge-in](https://pydantic.dev/docs/ai/realtime/turns/#barge-in) stops playback. A dropped sideband follows the same [reconnect](https://pydantic.dev/docs/ai/realtime/lifecycle/#reconnecting) rules, with one wrinkle covered there: a clean close is treated as the browser hanging up. The [realtime WebRTC example](https://pydantic.dev/docs/ai/examples/realtime/realtime-webrtc/) demonstrates the full FastAPI and browser flow. Provider-specific setup (Azure’s Microsoft Entra ID and `webrtcfilter`) lives on the [Azure](https://pydantic.dev/docs/ai/realtime/azure/#browser-webrtc-and-microsoft-entra-id) page. Browser → backend WebSocket relay --------------------------------- [](https://pydantic.dev/docs/ai/realtime/deployment/#browser--backend-websocket-relay) When the browser can’t use WebRTC — or the provider is Gemini Live or xAI — build a WebSocket endpoint on your backend that accepts the browser’s microphone audio and pumps it into [`send_audio()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.send_audio) , while relaying [`stream_audio()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.stream_audio) output back for playback. Here the backend owns the media bridge. The [realtime camera example](https://pydantic.dev/docs/ai/examples/realtime/realtime-camera/) demonstrates this shape end to end; a minimal FastAPI relay — the browser sends raw PCM16 binary frames and plays the frames it receives — is: import asyncio from fastapi import FastAPI, WebSocket from pydantic_ai import Agent agent = Agent(instructions='You are a helpful voice assistant.') app = FastAPI() @app.websocket('/voice') async def voice_socket(websocket: WebSocket): await websocket.accept() async with agent.realtime('openai:gpt-realtime').session() as session: async def pump_input(): while True: await session.send_audio(await websocket.receive_bytes()) input_task = asyncio.create_task(pump_input()) try: async for chunk in session.stream_audio(): await websocket.send_bytes(chunk) finally: input_task.cancel() SIP/telephony bridge -------------------- [](https://pydantic.dev/docs/ai/realtime/deployment/#siptelephony-bridge) Terminate the phone call with a telephony provider such as Twilio, then build the service that connects its media stream (e.g. Twilio Media Streams over WebSocket) to the backend session, transcoding between the line’s codec and PCM16. Was this page helpful? Thanks for your feedback! --- # pydantic_graph.exceptions | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic_graph/exceptions/#_top) pydantic\_graph.exceptions ========================== GraphBuildingError ------------------ [](https://pydantic.dev/docs/ai/api/pydantic_graph/exceptions/#pydantic_graph.exceptions.GraphBuildingError) **Bases:** [`ValueError`](https://docs.python.org/3/library/exceptions.html#ValueError) An error raised during graph-building. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/exceptions/#attributes) #### message [](https://pydantic.dev/docs/ai/api/pydantic_graph/exceptions/#pydantic_graph.exceptions.GraphBuildingError.message) The error message. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `message` GraphRuntimeError ----------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/exceptions/#pydantic_graph.exceptions.GraphRuntimeError) **Bases:** [`RuntimeError`](https://docs.python.org/3/library/exceptions.html#RuntimeError) Error caused by an issue during graph execution. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/exceptions/#attributes-1) #### message [](https://pydantic.dev/docs/ai/api/pydantic_graph/exceptions/#pydantic_graph.exceptions.GraphRuntimeError.message) The error message. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `message` GraphSetupError --------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/exceptions/#pydantic_graph.exceptions.GraphSetupError) **Bases:** [`TypeError`](https://docs.python.org/3/library/exceptions.html#TypeError) Error caused by an incorrectly configured graph. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/exceptions/#attributes-2) #### message [](https://pydantic.dev/docs/ai/api/pydantic_graph/exceptions/#pydantic_graph.exceptions.GraphSetupError.message) Description of the mistake. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `message` GraphValidationError -------------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/exceptions/#pydantic_graph.exceptions.GraphValidationError) **Bases:** [`ValueError`](https://docs.python.org/3/library/exceptions.html#ValueError) An error raised during graph validation. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/exceptions/#attributes-3) #### message [](https://pydantic.dev/docs/ai/api/pydantic_graph/exceptions/#pydantic_graph.exceptions.GraphValidationError.message) The error message. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `message` UnsupportedEventLoopError ------------------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/exceptions/#pydantic_graph.exceptions.UnsupportedEventLoopError) **Bases:** [`RuntimeError`](https://docs.python.org/3/library/exceptions.html#RuntimeError) Error caused by calling a synchronous method on an event loop that cannot be driven by the caller. Synchronous methods run their asynchronous implementation using `loop.run_until_complete()`, which not every event loop implements. Temporal’s workflow event loop is one that doesn’t: it can only be driven by Temporal. Pydantic AI’s synchronous methods report this as a `pydantic_ai.exceptions.UserError` instead. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/exceptions/#attributes-4) #### message [](https://pydantic.dev/docs/ai/api/pydantic_graph/exceptions/#pydantic_graph.exceptions.UnsupportedEventLoopError.message) The error message. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `message` Was this page helpful? Thanks for your feedback! --- # Usage and observability | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/realtime/observability/#_top) Usage and observability ======================= Realtime audio bills by the second in both directions, so knowing what a session cost — and capping it — matters even more than for a text run. Realtime sessions accumulate standard [`RunUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage) , enforce standard [`UsageLimits`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits) , and emit OpenTelemetry spans — viewable in [Pydantic Logfire](https://pydantic.dev/docs/ai/integrations/logfire/) — through Pydantic AI’s normal instrumentation. This lets voice and follow-up text runs share one usage budget and trace. Usage and limits ---------------- [](https://pydantic.dev/docs/ai/realtime/observability/#usage-and-limits) Read cumulative usage from [`RealtimeSession.usage`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.usage) . It includes input/output tokens, provider audio and cache breakdowns where available, and tool-call counts. Usage updates are not emitted as session events. As with a standard run’s [usage limits](https://pydantic.dev/docs/ai/core-concepts/agent/#usage-limits) , pass `usage=` to accumulate into a shared object — for example one carried across a voice call and its follow-up text runs — and `usage_limits=` to cap a session: from pydantic_ai import Agent from pydantic_ai.realtime import RealtimeTurnCompleteEvent from pydantic_ai.usage import RunUsage, UsageLimits agent = Agent() shared = RunUsage() async def main(): async with agent.realtime( 'openai:gpt-realtime', usage=shared, usage_limits=UsageLimits(total_tokens_limit=100_000), ).session() as session: await session.send('Say hello.') async for event in session: if isinstance(event, RealtimeTurnCompleteEvent): break print(shared) #> RunUsage(requests=1) Input-transcription usage is reported separately in `RunUsage.details` under `input_transcription_*` keys. It is not included in response token totals or attributed to a `ModelResponse`, because transcription can use a separate model and billing meter. Token and tool-call limits are checked as usage accrues. Request limits are checked before sending text, explicitly creating a response, or returning a tool result. With server-side VAD, the provider can begin a response without a client request; that limit is checked at the first response event. Breaches raise [`UsageLimitExceeded`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UsageLimitExceeded) from iteration. Provider-specific usage fields belong on the [OpenAI](https://pydantic.dev/docs/ai/realtime/openai/#feature-support-and-limitations) , [Azure OpenAI](https://pydantic.dev/docs/ai/realtime/azure/#feature-support-and-limitations) , [Google Gemini](https://pydantic.dev/docs/ai/realtime/gemini/#feature-support-and-limitations) , and [xAI](https://pydantic.dev/docs/ai/realtime/xai/#feature-support-and-limitations) pages. Logfire instrumentation ----------------------- [](https://pydantic.dev/docs/ai/realtime/observability/#logfire-instrumentation) Call `logfire.instrument_pydantic_ai()` or set `instrument=True` on the agent: import logfire logfire.configure() logfire.instrument_pydantic_ai() The session creates an `invoke_agent` span with cumulative usage and conversation content, subject to the normal content-redaction setting. Nested `chat {model}` spans represent provider responses, and `execute_tool` spans represent tools and delegated agent runs. `model turn complete` and `interrupt` spans mark those boundaries. A tool round can produce several response spans within one turn. You may see runs of `model turn complete (interrupted)` spans with no `chat` span between them. That’s normal on OpenAI server VAD: the provider starts a response for each detected speech segment, so a user who keeps talking cancels each auto-response before it produces output. Every cancelled or interrupted response still draws a boundary, displayed as `model turn complete (interrupted)`. | Attribute | Set on | Meaning | | --- | --- | --- | | `pydantic_ai.realtime` | Spans the session emits itself (session, response, boundary, and `user speech` spans) | Always `True`; marks spans that belong to a realtime session. `execute_tool` spans come from the [`Instrumentation`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Instrumentation)
capability and don’t carry it. | | `gen_ai.output.type` | Response spans | `speech` or `text`. | | `pydantic_ai.response.state` | Interrupted response spans | `'interrupted'`. | | Response-level usage | OpenAI, Azure OpenAI, and xAI response spans | Tokens attributed to that response. | Gemini can report usage only on a later completed turn after a function-call response; cumulative session usage remains authoritative. When providers report both user speech start and end, Pydantic AI records a `user speech` span. Providers without both boundaries do not get a guessed duration. On a [WebRTC sideband](https://pydantic.dev/docs/ai/realtime/deployment/#browser-webrtc-server-sideband) a `speak {model}` span additionally covers how long the model was _audible_, which the response spans can’t show: the provider generates audio far ahead of playing it, so this span routinely outlasts the `model turn complete` that ended the response. The session span also reports `pydantic_ai.audio_chunks_dropped` and `pydantic_ai.transcript_items_dropped`, summed across bounded [audio and transcript consumers](https://pydantic.dev/docs/ai/realtime/audio/#consuming-audio-and-transcripts) . These totals are written when the session closes. See [Debugging and monitoring](https://pydantic.dev/docs/ai/integrations/logfire/) for Logfire setup and privacy controls. Gateway trace propagation ------------------------- [](https://pydantic.dev/docs/ai/realtime/observability/#gateway-trace-propagation) Routing through the [Pydantic AI Gateway](https://pydantic.dev/docs/ai/overview/gateway/) — e.g. `agent.realtime('gateway/openai:gpt-realtime')` — is provider configuration, documented on the [OpenAI](https://pydantic.dev/docs/ai/realtime/openai/#gateway) and [Gemini](https://pydantic.dev/docs/ai/realtime/gemini/#gateway) pages. When a span is active during the WebSocket handshake, Pydantic AI propagates [W3C trace context](https://www.w3.org/TR/trace-context/) so gateway spans can join the trace. The provider connection is established before the realtime session span starts. Wrap the entire session context in an outer span when the handshake itself must be included: import logfire from pydantic_ai import Agent agent = Agent() async def main(): with logfire.span('voice call'): async with agent.realtime('openai:gpt-realtime').session() as session: await session.send('Say hello.') Edge cases ---------- [](https://pydantic.dev/docs/ai/realtime/observability/#edge-cases) * Usage is cumulative session state, not an event stream. Read it after the relevant responses or when the session closes. * A provider can report response-level usage at a different point from the local tool or turn boundary. Use the session total for billing and limits. * Dropped-stream counters represent each slow consumer independently; two lagging audio iterators can both contribute drops for the same produced audio. Was this page helpful? Thanks for your feedback! --- # pydantic_evals.online_capability | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic_evals/online_capability/#_top) pydantic\_evals.online\_capability ================================== Online evaluation capability for pydantic-ai agents. Provides an `OnlineEvaluation` capability that attaches evaluators to agent runs, dispatching them asynchronously in the background after each run completes. OnlineEvaluation ---------------- [](https://pydantic.dev/docs/ai/api/pydantic_evals/online_capability/#pydantic_evals.online_capability.OnlineEvaluation) **Bases:** `AbstractCapability[AgentDepsT]` Capability that runs online evaluators on agent run results. Dispatches evaluators asynchronously in the background after each completed agent run. Non-blocking — the agent run returns without waiting for evaluators to finish. Example: from dataclasses import dataclass from pydantic_ai import Agent from pydantic_evals.evaluators import Evaluator, EvaluatorContext from pydantic_evals.online_capability import OnlineEvaluation @dataclass class OutputNotEmpty(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return bool(ctx.output) agent = Agent( 'openai:gpt-5.2', name='assistant', capabilities=[OnlineEvaluation(evaluators=[OutputNotEmpty()])], ) ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_evals/online_capability/#attributes) #### config [](https://pydantic.dev/docs/ai/api/pydantic_evals/online_capability/#pydantic_evals.online_capability.OnlineEvaluation.config) Optional config override. Defaults to the global `DEFAULT_CONFIG`. **Type:** `OnlineEvalConfig` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### evaluators [](https://pydantic.dev/docs/ai/api/pydantic_evals/online_capability/#pydantic_evals.online_capability.OnlineEvaluation.evaluators) Evaluators to run after each agent run. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[`Evaluator` | `OnlineEvaluator`\] Was this page helpful? Thanks for your feedback! --- # template | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic-ai/template/#_top) template ======== Template string support for dynamic instructions. TemplateStr ----------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/template/#pydantic_ai.template.TemplateStr) **Bases:** `Generic[AgentDepsT]` A Handlebars template string that renders against `RunContext.deps`. When used in type hints, strings containing `{{` are automatically compiled as Handlebars templates during Pydantic validation. Uses [pydantic-handlebars](https://github.com/pydantic/pydantic-handlebars) for template compilation, schema validation, and rendering. When used with an `Agent`, `deps_type` is inferred automatically from the agent’s validation context, so you only need to pass it when constructing a `TemplateStr` outside of an agent (e.g. for standalone rendering). ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/template/#methods) #### \_\_call\_\_ [](https://pydantic.dev/docs/ai/api/pydantic-ai/template/#pydantic_ai.template.TemplateStr.__call__) def __call__(ctx: RunContext[AgentDepsT]) -> str Render the template against `ctx.deps`. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/template/#returns) [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### render [](https://pydantic.dev/docs/ai/api/pydantic-ai/template/#pydantic_ai.template.TemplateStr.render) def render(deps: AgentDepsT | None = None) -> str Render the template against the given deps object. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/template/#returns-1) [`str`](https://docs.python.org/3/library/stdtypes.html#str) Was this page helpful? Thanks for your feedback! --- # Prepare Tools | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/prepare-tools/#_top) Prepare Tools ============= [`PrepareTools`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.PrepareTools) and [`PrepareOutputTools`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.PrepareOutputTools) wrap a [`ToolsPrepareFunc`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.ToolsPrepareFunc) as a [capability](https://pydantic.dev/docs/ai/capabilities/overview/) , for filtering or modifying [tool definitions](https://pydantic.dev/docs/ai/tools-toolsets/tools/) per step. `PrepareTools` handles function tools; `PrepareOutputTools` handles [output tools](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.ToolOutput) . prepare\_tools\_native.py from pydantic_ai import Agent, RunContext, ToolDefinition from pydantic_ai.capabilities import PrepareTools async def hide_dangerous(ctx: RunContext, tool_defs: list[ToolDefinition]) -> list[ToolDefinition]: return [td for td in tool_defs if not td.name.startswith('delete_')] agent = Agent('openai:gpt-5.2', capabilities=[PrepareTools(hide_dangerous)]) @agent.tool_plain def delete_file(path: str) -> str: """Delete a file.""" return f'deleted {path}' @agent.tool_plain def read_file(path: str) -> str: """Read a file.""" return f'contents of {path}' result = agent.run_sync('hello') # The model only sees `read_file`, not `delete_file` For more complex tool preparation logic, see [Tool preparation](https://pydantic.dev/docs/ai/capabilities/custom/#tool-preparation) under lifecycle hooks. Was this page helpful? Thanks for your feedback! --- # Third-Party Capabilities | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/third-party/#_top) Third-Party Capabilities ======================== [Capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) are the recommended way for third-party packages to extend Pydantic AI, since they can bundle tools with hooks, instructions, and model settings. See [Extensibility](https://pydantic.dev/docs/ai/guides/extensibility/) for the full ecosystem, including [third-party toolsets](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/#third-party-toolsets) that can also be wrapped as capabilities. Many of the use cases below are also covered by first-party capabilities in Pydantic AI itself or [Pydantic AI Harness](https://pydantic.dev/docs/ai/harness/) , the official capability library. Where that’s the case, we point to the built-in option first, and list the community packages as alternatives. Task Management --------------- [](https://pydantic.dev/docs/ai/capabilities/third-party/#task-management) For model-owned task planning and progress tracking, [Pydantic AI Harness](https://pydantic.dev/docs/ai/harness/) ships [`Planning`](https://pydantic.dev/docs/ai/harness/planning/) , a cache-friendly self-updating task plan. As a community alternative with subtask, dependency, and PostgreSQL persistence support: * [`pydantic-ai-todo`](https://github.com/vstorm-co/pydantic-ai-todo) - `TodoCapability` with `add_todo`, `read_todos`, `write_todos`, `update_todo_status`, and `remove_todo` tools. Supports subtasks, dependencies, and PostgreSQL persistence. Also available as a lower-level `TodoToolset`. Context Management ------------------ [](https://pydantic.dev/docs/ai/capabilities/third-party/#context-management) Pydantic AI has [built-in compaction](https://pydantic.dev/docs/ai/capabilities/compaction/) — provider-native APIs and model-agnostic history summarization — and [Pydantic AI Harness](https://pydantic.dev/docs/ai/harness/compaction/) adds a full menu of compaction strategies. As a community alternative: * [`summarization-pydantic-ai`](https://github.com/vstorm-co/summarization-pydantic-ai) - Four capabilities for managing long conversations: `ContextManagerCapability` (real-time token tracking, auto-compression at a configurable threshold, and large tool-output truncation); `SummarizationCapability` (LLM-powered history compression); `SlidingWindowCapability` (zero-cost message trimming); `LimitWarnerCapability` (injects a finish-soon hint before hard context limits). Also available as standalone `history_processors`: `SummarizationProcessor`, `SlidingWindowProcessor`, and `LimitWarnerProcessor`. Multi-Agent Orchestration ------------------------- [](https://pydantic.dev/docs/ai/capabilities/third-party/#multi-agent-orchestration) Pydantic AI supports [multi-agent patterns](https://pydantic.dev/docs/ai/guides/multi-agent-applications/) directly, and [Pydantic AI Harness](https://pydantic.dev/docs/ai/harness/subagents/) ships [`SubAgents`](https://pydantic.dev/docs/ai/harness/subagents/) for delegating self-contained tasks to named child agents. As a community alternative: * [`subagents-pydantic-ai`](https://github.com/vstorm-co/subagents-pydantic-ai) - `SubAgentCapability` adds tools for multi-agent delegation: `task` (spawn a subagent), `check_task`, `wait_tasks`, `list_active_tasks`, `soft_cancel_task`, `hard_cancel_task`, and `answer_subagent`. Supports sync, async, and auto-execution modes, nested subagents, and runtime agent creation. Also available as a lower-level toolset via `create_subagent_toolset`. Guardrails & Safety ------------------- [](https://pydantic.dev/docs/ai/capabilities/third-party/#guardrails--safety) [Pydantic AI Harness](https://pydantic.dev/docs/ai/harness/guardrails/) provides input and output guardrails that validate or block requests and responses, and Pydantic AI enforces usage, token, and request limits via [`UsageLimits`](https://pydantic.dev/docs/ai/core-concepts/agent/#usage-limits) . As a community alternative bundling several ready-made shields, including USD cost tracking: * [`pydantic-ai-shields`](https://github.com/vstorm-co/pydantic-ai-shields) - Ready-to-use guardrail capabilities: `CostTracking` (tracks token usage and USD cost per run, raises `BudgetExceededError` on budget overrun); `ToolGuard` (block or require approval for specific tools); `InputGuard` and `OutputGuard` (custom sync or async validation functions); `PromptInjection`, `PiiDetector`, `SecretRedaction`, `BlockedKeywords`, and `NoRefusals` content shields. File Operations & Sandboxing ---------------------------- [](https://pydantic.dev/docs/ai/capabilities/third-party/#file-operations--sandboxing) [Pydantic AI Harness](https://pydantic.dev/docs/ai/harness/) ships sandboxed [`FileSystem`](https://pydantic.dev/docs/ai/harness/filesystem/) and [`Shell`](https://pydantic.dev/docs/ai/harness/shell/) capabilities, plus [`CodeMode`](https://pydantic.dev/docs/ai/harness/code-mode/) for running tool calls as sandboxed Python. As a community alternative: * [`pydantic-ai-backend`](https://github.com/vstorm-co/pydantic-ai-backend) - `ConsoleCapability` registers `ls`, `read_file`, `write_file`, `edit_file`, `glob`, `grep`, and `execute` tools with a fine-grained permission system. Backends include `StateBackend` (in-memory, for testing), `LocalBackend` (real filesystem), `DockerSandbox` (isolated container execution), and `CompositeBackend` (routing across backends). Also available as a lower-level `ConsoleToolset`. Agent Skills ------------ [](https://pydantic.dev/docs/ai/capabilities/third-party/#agent-skills) Pydantic AI supports [Agent Skills natively](https://pydantic.dev/docs/ai/capabilities/on-demand/#loading-skills-from-markdown-files) through [on-demand capabilities](https://pydantic.dev/docs/ai/capabilities/on-demand/) , which collapse a skill to a one-line catalog entry until the model loads it. As a community alternative: * [`pydantic-ai-skills`](https://github.com/DougTrajano/pydantic-ai-skills) - `SkillsCapability` implements Agent Skills support with progressive disclosure (load skills on-demand to reduce tokens). Supports filesystem and programmatic skills; compatible with [agentskills.io](https://agentskills.io/) . Data & Analytics ---------------- [](https://pydantic.dev/docs/ai/capabilities/third-party/#data--analytics) Capabilities for querying and analyzing structured data help agents answer questions over files and databases: * [`pydantic-ai-chdb`](https://github.com/chdb-io/pydantic-ai-chdb) - `ChDBCapability` gives agents analytical SQL over local files (Parquet/CSV/JSON), object storage, and remote databases with [chDB](https://clickhouse.com/docs/en/chdb) , the in-process ClickHouse engine — the engine itself needs no server or connection string to run (remote sources are reached via ClickHouse table functions, which take their own credentials). Registers `run_select_query` (read-only ClickHouse SQL with parameter binding), `list_databases`, `list_tables`, `describe_table`, `get_sample_data`, `list_functions`, and `attach_file` (opt-in writable sessions) tools plus schema-first instructions. Sessions default to the engine-level `readonly=2` setting with capped results, and typed engine errors are mapped to [`ModelRetry`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelRetry) so the model can correct its queries. Works with [agent specs](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) out of the box, so it can be loaded via [`from_spec`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability.from_spec) / [`Agent.from_spec`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.from_spec) . Also available as a lower-level [toolset](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) via [`ChDBCapability(...).get_toolset()`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability.get_toolset) . To add your package to this page, open a pull request. To publish your own capability package, see [Publishing capabilities](https://pydantic.dev/docs/ai/capabilities/custom/#publishing-capabilities) and [Extensibility](https://pydantic.dev/docs/ai/guides/extensibility/) . Was this page helpful? Thanks for your feedback! --- # xai | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/realtime/xai/#_top) xai === The xAI Grok Voice realtime API provider. Requires the `realtime`, `xai`, and `openai` optional groups (`pip install "pydantic-ai-slim[realtime,xai,openai]"`) — `openai` because the model reuses the OpenAI Realtime codec, whose event types come from the OpenAI SDK. xAI’s realtime API is a clone of the OpenAI Realtime protocol, so [`XaiRealtimeModel`](https://pydantic.dev/docs/ai/api/realtime/xai/#pydantic_ai.realtime.xai.XaiRealtimeModel) reuses the OpenAI codec (event mapping, seeding, the WebSocket connection). Turn-taking uses the shared [`TurnDetection`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.TurnDetection) (or `False` for push-to-talk); for exact server-VAD control, `xai_turn_detection` accepts [`ServerVAD`](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.ServerVAD) and fully overrides the shared setting. It diverges only where xAI does: it supports cancellation-based interruption but not output truncation, has no image input, and streams input transcription as cumulative snapshots that may revise earlier text, rather than as incremental deltas. Authentication comes from an [`XaiProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.xai.XaiProvider) , mirroring [`XaiModel`](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModel) . xAI Grok Voice realtime API provider for speech-to-speech sessions. Connects to `wss://api.x.ai/v1/realtime` over a WebSocket. xAI’s realtime API is a deliberate clone of the OpenAI Realtime protocol, so this provider reuses the OpenAI codec from [`pydantic_ai.realtime.openai`](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai) — event mapping, session seeding, tool conversion, server-VAD config, and the WebSocket connection itself — and diverges only where xAI does: * the `session.update` shape (`voice`/`turn_detection` sit at the session top level, not nested under `audio` as on OpenAI’s GA surface); * input audio transcription, delivered as cumulative `conversation.item.input_audio_transcription.updated` snapshots plus a final `.completed`, rather than OpenAI’s incremental `.delta` events (see [`map_event`](https://pydantic.dev/docs/ai/api/realtime/xai/#pydantic_ai.realtime.xai.map_event) ); * native conversation resumption when a reconnect policy is configured: the provider-assigned `conversation.id` is reused and its replay burst is suppressed from local history; * no output truncation (`conversation.item.truncate` is unsupported), so [`RealtimeModelProfile.supports_output_truncation`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelProfile.supports_output_truncation) is `False` while cancellation-based interruption still works; * no text output — the API has no response-modality control and always speaks — so [`RealtimeModelProfile.supports_text_output`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelProfile.supports_text_output) is `False` and `output_modality='text'` raises rather than silently coming back as audio. Requires the `websockets` package (the `realtime` optional group), `xai-sdk` (the `xai` group, for [`XaiProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.xai.XaiProvider) ), and `openai` (the `openai` group, whose SDK supplies the event types the shared OpenAI codec is built on): pip install “pydantic-ai-slim\[xai-realtime\]“ XaiRealtimeConnection --------------------- [](https://pydantic.dev/docs/ai/api/realtime/xai/#pydantic_ai.realtime.xai.XaiRealtimeConnection) **Bases:** `OpenAIRealtimeConnection` A live WebSocket connection to the xAI Grok Voice realtime API. Reuses [`OpenAIRealtimeConnection`](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeConnection) for the shared wire protocol, while mapping xAI’s cumulative input transcription and conversation lifecycle events and emitting the resumption replay controls captured during reconnect handshakes. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/xai/#attributes) #### conversation\_id [](https://pydantic.dev/docs/ai/api/realtime/xai/#pydantic_ai.realtime.xai.XaiRealtimeConnection.conversation_id) The xAI conversation ID used for native session resumption. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) ### Methods [](https://pydantic.dev/docs/ai/api/realtime/xai/#methods) #### set\_message\_history [](https://pydantic.dev/docs/ai/api/realtime/xai/#pydantic_ai.realtime.xai.XaiRealtimeConnection.set_message_history) def set_message_history(message_history: Callable[[], Sequence[ModelMessage]]) -> None Ignored: xAI restores the conversation itself, so replaying it would say everything twice. ##### Returns [](https://pydantic.dev/docs/ai/api/realtime/xai/#returns) [`None`](https://docs.python.org/3/library/constants.html#None) XaiRealtimeModel ---------------- [](https://pydantic.dev/docs/ai/api/realtime/xai/#pydantic_ai.realtime.xai.XaiRealtimeModel) **Bases:** `RealtimeModel` xAI Grok Voice realtime API model. Pass `provider='xai'` (the default, which reads `XAI_API_KEY`) or an [`XaiProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.xai.XaiProvider) constructed with `api_key=`. A custom `api_host` is not supported, and a provider constructed only with `xai_client=` cannot be used because the WebSocket connection needs access to the API key. The realtime WebSocket URL is `wss://api.x.ai/v1/realtime`. ### Constructor Parameters [](https://pydantic.dev/docs/ai/api/realtime/xai/#constructor-parameters) **`model`** : `XaiRealtimeModelName` [](https://pydantic.dev/docs/ai/api/realtime/xai/#pydantic_ai.realtime.xai.XaiRealtimeModel.__init__(model)) The model name, e.g. `grok-voice-latest` (which tracks the current model) or a pinned version like `grok-voice-think-fast-1.0`. The `model` query parameter is required by the server, which otherwise falls back to a default silently. **`provider`** : `XaiProvider` | [`str`](https://docs.python.org/3/library/stdtypes.html#str) _Default:_ `'xai'` [](https://pydantic.dev/docs/ai/api/realtime/xai/#pydantic_ai.realtime.xai.XaiRealtimeModel.__init__(provider)) The provider to use for authentication and the base URL. Defaults to `'xai'`. **`settings`** : `RealtimeModelSettings` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/realtime/xai/#pydantic_ai.realtime.xai.XaiRealtimeModel.__init__(settings)) [Model settings](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelSettings) used as defaults for realtime sessions. A [`reconnect`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelSettings.reconnect) policy enables xAI’s native session resumption: prior turns are restored when reconnecting within xAI’s resumption window (reportedly ~30 minutes). **`profile`** : `RealtimeModelProfileSpec` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/realtime/xai/#pydantic_ai.realtime.xai.XaiRealtimeModel.__init__(profile)) Optional override for the [realtime model profile](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelProfile) , merged over the provider’s — a partial dict, or a callable taking the resolved profile and returning the one to use. Mirrors `profile=` on a standard [`Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) , and is the escape hatch when a model name doesn’t identify the model (e.g. an Azure deployment named something other than its model). XaiRealtimeModelSettings ------------------------ [](https://pydantic.dev/docs/ai/api/realtime/xai/#pydantic_ai.realtime.xai.XaiRealtimeModelSettings) **Bases:** `RealtimeModelSettings` Settings specific to xAI realtime models. Grok Voice always produces audio, so its profile reports [`supports_text_output=False`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelProfile.supports_text_output) and the inherited `output_modality='text'` is rejected up front rather than quietly ignored. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/xai/#attributes-1) #### xai\_turn\_detection [](https://pydantic.dev/docs/ai/api/realtime/xai/#pydantic_ai.realtime.xai.XaiRealtimeModelSettings.xai_turn_detection) xAI-specific server-VAD configuration. When present, this fully overrides the cross-provider `turn_detection` setting. **Type:** `ServerVAD` #### xai\_voice [](https://pydantic.dev/docs/ai/api/realtime/xai/#pydantic_ai.realtime.xai.XaiRealtimeModelSettings.xai_voice) Voice used for audio output, e.g. `eve`, or a custom voice ID. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) map\_conversation\_event ------------------------ [](https://pydantic.dev/docs/ai/api/realtime/xai/#pydantic_ai.realtime.xai.map_conversation_event) def map_conversation_event( data: dict[str, Any], *, replayed: bool | None = None, ) -> ConversationCreated | ConversationItemCreated | None Map xAI’s conversation handshake and item lifecycle events to codec control events. ### Returns [](https://pydantic.dev/docs/ai/api/realtime/xai/#returns-1) `ConversationCreated` | `ConversationItemCreated` | [`None`](https://docs.python.org/3/library/constants.html#None) map\_event ---------- [](https://pydantic.dev/docs/ai/api/realtime/xai/#pydantic_ai.realtime.xai.map_event) def map_event(data: dict[str, Any]) -> RealtimeCodecEvent | None Map a raw xAI Grok Voice realtime event to a [`RealtimeCodecEvent`](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeCodecEvent) . xAI clones the OpenAI Realtime protocol, so most events map identically via the OpenAI codec. The first exception is input audio transcription: xAI emits cumulative `conversation.item.input_audio_transcription.updated` snapshots (which may retroactively _correct_ earlier text — `'Hello?'` becomes `'Hello, my name is'`) plus cumulative `.completed` snapshots, rather than OpenAI’s incremental `.delta`. The partials are surfaced as cumulative [`InputTranscript`](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.InputTranscript) s so a live transcript can render the user’s words as they are spoken; the session adopts each snapshot wholesale, appending when it merely extends and replacing when xAI revises itself. The shared codec still drops interim `.completed` snapshots. The other exception is xAI’s conversation lifecycle events, which are surfaced as codec control events so the connection can capture `conversation.id` and the session can suppress resume replay. ### Returns [](https://pydantic.dev/docs/ai/api/realtime/xai/#returns-2) `RealtimeCodecEvent` | [`None`](https://docs.python.org/3/library/constants.html#None) Was this page helpful? Thanks for your feedback! --- # Direct Model Requests | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/core-concepts/direct/#_top) Direct Model Requests ===================== The `direct` module provides low-level methods for making imperative requests to LLMs where the only abstraction is input and output schema translation, enabling you to use all models with the same API. These methods are thin wrappers around the [`Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) implementations, offering a simpler interface when you don’t need the full functionality of an [`Agent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent) . The following functions are available: * [`model_request`](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request) : Make a non-streamed async request to a model * [`model_request_sync`](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_sync) : Make a non-streamed synchronous request to a model * [`model_request_stream`](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_stream) : Make a streamed async request to a model * [`model_request_stream_sync`](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_stream_sync) : Make a streamed sync request to a model Basic Example ------------- [](https://pydantic.dev/docs/ai/core-concepts/direct/#basic-example) Here’s a simple example demonstrating how to use the direct API to make a basic request: direct\_basic.py from pydantic_ai import ModelRequest from pydantic_ai.direct import model_request_sync # Make a synchronous request to the model model_response = model_request_sync( 'anthropic:claude-haiku-4-5', [ModelRequest.user_text_prompt('What is the capital of France?')] ) print(model_response.parts[0].content) #> The capital of France is Paris. print(model_response.usage) #> RequestUsage(input_tokens=56, output_tokens=7) _(This example is complete, it can be run “as is”)_ Advanced Example with Tool Calling ---------------------------------- [](https://pydantic.dev/docs/ai/core-concepts/direct/#advanced-example-with-tool-calling) You can also use the direct API to work with function/tool calling. Even here we can use Pydantic to generate the JSON schema for the tool: from typing import Literal from pydantic import BaseModel from pydantic_ai import ModelRequest, ToolDefinition from pydantic_ai.direct import model_request from pydantic_ai.models import ModelRequestParameters class Divide(BaseModel): """Divide two numbers.""" numerator: float denominator: float on_inf: Literal['error', 'infinity'] = 'infinity' async def main(): # Make a request to the model with tool access model_response = await model_request( 'openai:gpt-5-nano', [ModelRequest.user_text_prompt('What is 123 / 456?')], model_request_parameters=ModelRequestParameters( function_tools=[\ ToolDefinition(\ name=Divide.__name__.lower(),\ description=Divide.__doc__,\ parameters_json_schema=Divide.model_json_schema(),\ )\ ], allow_text_output=True, # Allow model to either use tools or respond directly ), ) print(model_response) """ ModelResponse( parts=[\ ToolCallPart(\ tool_name='divide',\ args={'numerator': '123', 'denominator': '456'},\ tool_call_id='pyd_ai_2e0e396768a14fe482df90a29a78dc7b',\ )\ ], usage=RequestUsage(input_tokens=55, output_tokens=7), model_name='gpt-5-nano', timestamp=datetime.datetime(...), ) """ _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main())` to run `main`)_ When to Use the direct API vs Agent ----------------------------------- [](https://pydantic.dev/docs/ai/core-concepts/direct/#when-to-use-the-direct-api-vs-agent) The direct API is ideal when: 1. You need more direct control over model interactions 2. You want to implement custom behavior around model requests 3. You’re building your own abstractions on top of model interactions For most application use cases, the higher-level [`Agent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent) API provides a more convenient interface with additional features such as native tool execution, retrying, structured output parsing, and more. OpenTelemetry or Logfire Instrumentation ---------------------------------------- [](https://pydantic.dev/docs/ai/core-concepts/direct/#opentelemetry-or-logfire-instrumentation) As with [agents](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent) , you can enable OpenTelemetry/Logfire instrumentation with just a few extra lines direct\_instrumented.py import logfire from pydantic_ai import ModelRequest from pydantic_ai.direct import model_request_sync logfire.configure() logfire.instrument_pydantic_ai() # Make a synchronous request to the model model_response = model_request_sync( 'anthropic:claude-haiku-4-5', [ModelRequest.user_text_prompt('What is the capital of France?')], ) print(model_response.parts[0].content) #> The capital of France is Paris. _(This example is complete, it can be run “as is”)_ You can also enable OpenTelemetry on a per call basis: direct\_instrumented.py import logfire from pydantic_ai import ModelRequest from pydantic_ai.direct import model_request_sync logfire.configure() # Make a synchronous request to the model model_response = model_request_sync( 'anthropic:claude-haiku-4-5', [ModelRequest.user_text_prompt('What is the capital of France?')], instrument=True ) print(model_response.parts[0].content) #> The capital of France is Paris. See [Debugging and Monitoring](https://pydantic.dev/docs/ai/integrations/logfire/) for more details, including how to instrument with plain OpenTelemetry without Logfire. Was this page helpful? Thanks for your feedback! --- # Simple Validation | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/examples/simple-validation/#_top) Simple Validation ================= A proof of concept example of evaluating a simple text transformation function with deterministic checks. Scenario -------- [](https://pydantic.dev/docs/ai/evals/examples/simple-validation/#scenario) We’re testing a function that converts text to title case. We want to verify: * Output is always a string * Output matches expected format * Function handles edge cases correctly * Performance meets requirements Complete Example ---------------- [](https://pydantic.dev/docs/ai/evals/examples/simple-validation/#complete-example) from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import ( Contains, EqualsExpected, IsInstance, MaxDuration, ) # The function we're testing def to_title_case(text: str) -> str: """Convert text to title case.""" return text.title() # Create evaluation dataset dataset = Dataset( name='title_case_validation', cases=[\ # Basic functionality\ Case(\ name='basic_lowercase',\ inputs='hello world',\ expected_output='Hello World',\ ),\ Case(\ name='basic_uppercase',\ inputs='HELLO WORLD',\ expected_output='Hello World',\ ),\ Case(\ name='mixed_case',\ inputs='HeLLo WoRLd',\ expected_output='Hello World',\ ),\ \ # Edge cases\ Case(\ name='empty_string',\ inputs='',\ expected_output='',\ ),\ Case(\ name='single_word',\ inputs='hello',\ expected_output='Hello',\ ),\ Case(\ name='with_punctuation',\ inputs='hello, world!',\ expected_output='Hello, World!',\ ),\ Case(\ name='with_numbers',\ inputs='hello 123 world',\ expected_output='Hello 123 World',\ ),\ Case(\ name='apostrophes',\ inputs="don't stop believin'",\ expected_output="Don'T Stop Believin'",\ ),\ ], evaluators=[\ # Always returns a string\ IsInstance(type_name='str'),\ \ # Matches expected output\ EqualsExpected(),\ \ # Output should contain capital letters\ Contains(value='H', evaluation_name='has_capitals'),\ \ # Should be fast (under 1ms)\ MaxDuration(seconds=0.001),\ ], ) # Run evaluation if __name__ == '__main__': report = dataset.evaluate_sync(to_title_case) # Print results report.print(include_input=True, include_output=True) """ Evaluation Summary: to_title_case ┏━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━┓ ┃ Case ID ┃ Inputs ┃ Outputs ┃ Assertions ┃ Duration ┃ ┡━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━┩ │ basic_lowercase │ hello world │ Hello World │ ✔✔✔✗ │ 10ms │ ├──────────────────┼──────────────────────┼──────────────────────┼────────────┼──────────┤ │ basic_uppercase │ HELLO WORLD │ Hello World │ ✔✔✔✗ │ 10ms │ ├──────────────────┼──────────────────────┼──────────────────────┼────────────┼──────────┤ │ mixed_case │ HeLLo WoRLd │ Hello World │ ✔✔✔✗ │ 10ms │ ├──────────────────┼──────────────────────┼──────────────────────┼────────────┼──────────┤ │ empty_string │ - │ - │ ✔✔✗✗ │ 10ms │ ├──────────────────┼──────────────────────┼──────────────────────┼────────────┼──────────┤ │ single_word │ hello │ Hello │ ✔✔✔✗ │ 10ms │ ├──────────────────┼──────────────────────┼──────────────────────┼────────────┼──────────┤ │ with_punctuation │ hello, world! │ Hello, World! │ ✔✔✔✗ │ 10ms │ ├──────────────────┼──────────────────────┼──────────────────────┼────────────┼──────────┤ │ with_numbers │ hello 123 world │ Hello 123 World │ ✔✔✔✗ │ 10ms │ ├──────────────────┼──────────────────────┼──────────────────────┼────────────┼──────────┤ │ apostrophes │ don't stop believin' │ Don'T Stop Believin' │ ✔✔✗✗ │ 10ms │ ├──────────────────┼──────────────────────┼──────────────────────┼────────────┼──────────┤ │ Averages │ │ │ 68.8% ✔ │ 10ms │ └──────────────────┴──────────────────────┴──────────────────────┴────────────┴──────────┘ """ # Check if all passed avg = report.averages() if avg and avg.assertions == 1.0: print('\n✅ All tests passed!') else: print(f'\n❌ Some tests failed (pass rate: {avg.assertions:.1%})') """ ❌ Some tests failed (pass rate: 68.8%) """ Expected Output --------------- [](https://pydantic.dev/docs/ai/evals/examples/simple-validation/#expected-output) Evaluation Summary: to_title_case ┏━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━┓ ┃ Case ID ┃ Inputs ┃ Outputs ┃ Assertions ┃ Duration ┃ ┡━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━┩ │ basic_lowercase │ hello world │ Hello World │ ✔✔✔✔ │ <1ms│ ├───────────────────┼──────────────────────┼───────────────────────┼────────────┼──────────┤ │ basic_uppercase │ HELLO WORLD │ Hello World │ ✔✔✔✔ │ <1ms│ ├───────────────────┼──────────────────────┼───────────────────────┼────────────┼──────────┤ │ mixed_case │ HeLLo WoRLd │ Hello World │ ✔✔✔✔ │ <1ms│ ├───────────────────┼──────────────────────┼───────────────────────┼────────────┼──────────┤ │ empty_string │ │ │ ✔✔✗✔ │ <1ms│ ├───────────────────┼──────────────────────┼───────────────────────┼────────────┼──────────┤ │ single_word │ hello │ Hello │ ✔✔✔✔ │ <1ms│ ├───────────────────┼──────────────────────┼───────────────────────┼────────────┼──────────┤ │ with_punctuation │ hello, world! │ Hello, World! │ ✔✔✔✔ │ <1ms│ ├───────────────────┼──────────────────────┼───────────────────────┼────────────┼──────────┤ │ with_numbers │ hello 123 world │ Hello 123 World │ ✔✔✔✔ │ <1ms│ ├───────────────────┼──────────────────────┼───────────────────────┼────────────┼──────────┤ │ apostrophes │ don't stop believin' │ Don'T Stop Believin' │ ✔✔✔✔ │ <1ms│ ├───────────────────┼──────────────────────┼───────────────────────┼────────────┼──────────┤ │ Averages │ │ │ 96.9% ✔ │ <1ms│ └───────────────────┴──────────────────────┴───────────────────────┴────────────┴──────────┘ ✅ All tests passed! Note: The `empty_string` case has one failed assertion (`has_capitals`) because an empty string contains no capital letters. Saving and Loading ------------------ [](https://pydantic.dev/docs/ai/evals/examples/simple-validation/#saving-and-loading) Save the dataset for future use: from typing import Any from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import EqualsExpected # The function we're testing def to_title_case(text: str) -> str: """Convert text to title case.""" return text.title() # Create dataset dataset: Dataset[str, str, Any] = Dataset( name='title_case_tests', cases=[Case(inputs='test', expected_output='Test')], evaluators=[EqualsExpected()], ) # Save to YAML dataset.to_file('title_case_tests.yaml') # Load later dataset = Dataset.from_file('title_case_tests.yaml') report = dataset.evaluate_sync(to_title_case) Adding More Cases ----------------- [](https://pydantic.dev/docs/ai/evals/examples/simple-validation/#adding-more-cases) As you find bugs or edge cases, add them to the dataset: from pydantic_evals import Dataset # Load existing dataset dataset = Dataset.from_file('title_case_tests.yaml') # Found a bug with unicode dataset.add_case( name='unicode_chars', inputs='café résumé', expected_output='Café Résumé', ) # Found a bug with all caps words dataset.add_case( name='acronyms', inputs='the USA and FBI', expected_output='The Usa And Fbi', # Python's title() behavior ) # Test with very long input dataset.add_case( name='long_input', inputs=' '.join(['word'] * 1000), expected_output=' '.join(['Word'] * 1000), ) # Save updated dataset dataset.to_file('title_case_tests.yaml') Using with pytest ----------------- [](https://pydantic.dev/docs/ai/evals/examples/simple-validation/#using-with-pytest) Integrate with pytest for CI/CD: import pytest from pydantic_evals import Dataset # The function we're testing def to_title_case(text: str) -> str: """Convert text to title case.""" return text.title() @pytest.fixture def title_case_dataset(): return Dataset.from_file('title_case_tests.yaml') def test_title_case_evaluation(title_case_dataset): """Run evaluation tests.""" report = title_case_dataset.evaluate_sync(to_title_case) # All cases should pass avg = report.averages() assert avg is not None assert avg.assertions == 1.0, f'Some tests failed (pass rate: {avg.assertions:.1%})' def test_title_case_performance(title_case_dataset): """Verify performance.""" report = title_case_dataset.evaluate_sync(to_title_case) # All cases should complete quickly for case in report.cases: assert case.task_duration < 0.001, f'{case.name} took {case.task_duration}s' Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/examples/simple-validation/#next-steps) * **[Native Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/built-in/) ** - Explore all available evaluators * **[Custom Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/custom/) ** - Write your own evaluation logic * **[Dataset Management](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/) ** - Save, load, and manage datasets * **[Concurrency & Performance](https://pydantic.dev/docs/ai/evals/how-to/concurrency/) ** - Optimize evaluation performance Was this page helpful? Thanks for your feedback! --- # Standard Quality Metrics | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/evaluators/standard-quality-metrics/#_top) Standard Quality Metrics ======================== This page shows how to express widely-used LLM evaluation methods with Pydantic Evals primitives: * [`GEval`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.GEval) — a first-class evaluator implementing G-Eval chain-of-thought scoring (Liu et al., 2023). * Ready-made [`LLMJudge`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.LLMJudge) rubrics for the RAG metrics popularized by [Ragas](https://github.com/explodinggradients/ragas) (faithfulness, answer relevance, context precision, context recall) and for GEMBA translation quality (Kocmi & Federmann, 2023). The RAG and GEMBA metrics are provided as _rubric recipes_ rather than evaluator classes: each is one rubric away from `LLMJudge`, and a rubric you own adapts freely to your dataset shape and domain — rename a field, tighten a criterion, or translate the instructions without waiting on a library release. Copy them into your project and edit as needed. G-Eval ------ [](https://pydantic.dev/docs/ai/evals/evaluators/standard-quality-metrics/#g-eval) [`GEval`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.GEval) implements chain-of-thought evaluation: you provide the aspect being evaluated (`criteria`) and a list of explicit `evaluation_steps`, and the judge returns a reasoning trace plus an integer score in `score_range` (inclusive). Because the criteria and steps are user-supplied, `GEval` puts no structural requirements on the inputs, and it works in serialized datasets out of the box. from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import GEval dataset = Dataset( name='g_eval_demo', cases=[Case(inputs='Explain how black holes form.')], evaluators=[\ GEval(\ criteria='coherence',\ evaluation_steps=[\ 'Read the output carefully.',\ 'Check that each sentence follows logically from the previous one.',\ 'Assign a score from 1 (incoherent) to 5 (fully coherent).',\ ],\ include_input=True,\ ),\ ], ) The result is an [`EvaluationReason`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluationReason) whose value is the raw integer score — on the scale you chose via `score_range`, not normalized to `0.0`\-`1.0` like [`LLMJudge`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.LLMJudge) scores. If the judge returns a score outside `score_range`, the evaluation fails rather than recording a misleading value. RAG metric rubrics ------------------ [](https://pydantic.dev/docs/ai/evals/evaluators/standard-quality-metrics/#rag-metric-rubrics) These recipes assume each case’s `inputs` carries the user question and the context passages the output is supposed to rely on — a _supplied_ context, not whatever an agent retrieved at runtime. With `include_input=True`, [`LLMJudge`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.LLMJudge) shows the judge your full inputs object, so any shape works as long as the rubric describes it; adjust the wording if your fields are named differently. from dataclasses import dataclass from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import LLMJudge faithfulness = LLMJudge( rubric=( 'Every factual claim in the Output must be directly supported by the context passages ' 'in the Input. Unsupported claims, contradictions, and fabrications constitute failure; ' 'ignore claims that are true in the real world but absent from the provided context. ' 'The score is the fraction of claims that are supported (0.0 = none, 1.0 = all); ' 'pass only if every claim is supported.' ), include_input=True, score={'evaluation_name': 'faithfulness'}, assertion=False, ) answer_relevance = LLMJudge( rubric=( 'Judge whether the Output directly and completely answers the question in the Input, ' 'without padding or unrelated tangents. ' 'The score reflects how directly the Output addresses the question ' '(0.0 = unrelated, 1.0 = a direct, on-point answer).' ), include_input=True, score={'evaluation_name': 'answer_relevance'}, assertion=False, ) context_precision = LLMJudge( rubric=( 'This metric judges the retrieval, not the answer: assess the context passages in the ' 'Input against the question in the Input, and disregard the Output. ' 'The score is the fraction of the context that is relevant to answering the question ' '(0.0 = none is relevant, 1.0 = all of it is relevant).' ), include_input=True, score={'evaluation_name': 'context_precision'}, assertion=False, ) context_recall = LLMJudge( rubric=( 'This metric judges the retrieval, not the answer: determine whether the context ' 'passages in the Input contain enough information to produce the ground-truth answer ' 'in the Expected Output, and disregard the Output. ' 'The score is the fraction of the ground-truth answer that is supported by the context ' '(0.0 = none of it, 1.0 = all of it).' ), include_input=True, include_expected_output=True, score={'evaluation_name': 'context_recall'}, assertion=False, ) @dataclass class RagInputs: question: str context: list[str] dataset = Dataset( name='rag_quality', cases=[\ Case(\ inputs=RagInputs(\ question='Where is the Eiffel Tower?',\ context=['The Eiffel Tower is in Paris, France.'],\ ),\ expected_output='The Eiffel Tower is in Paris.',\ ),\ ], evaluators=[faithfulness, answer_relevance, context_precision, context_recall], ) Each recipe emits a `0.0`\-`1.0` score named via the `score` [`OutputConfig`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.OutputConfig) ; swap `assertion=False` for an `assertion` config (or keep both) if you also want a pass/fail column, as described in [LLM Judge](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/) . GEMBA translation quality ------------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/standard-quality-metrics/#gemba-translation-quality) The GEMBA Direct Assessment prompt (Kocmi & Federmann, 2023, “Large Language Models Are State-of-the-Art Evaluators of Translation Quality”) scores a translation from 0 to 100. Here the case’s `inputs` is the source text, the output is the candidate translation, and (optionally) the `expected_output` is a human reference translation: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import LLMJudge gemba_da = LLMJudge( rubric=( 'The Input is the English source text and the Output is its French translation ' '(the Expected Output, if present, is a human reference translation). ' 'Score the translation on a continuous scale from 0 to 100, where 0 means ' '"no meaning preserved" and 100 means "perfect meaning and grammar", ' 'then report it normalized to the 0.0-1.0 range by dividing by 100.' ), include_input=True, include_expected_output=True, score={'evaluation_name': 'gemba_da'}, assertion=False, ) dataset = Dataset( name='translation_quality', cases=[\ Case(\ inputs='Hello, world!',\ expected_output='Bonjour, le monde !',\ ),\ ], evaluators=[gemba_da], ) Adjust the language names to your language pair. For the GEMBA-SQM variant, replace the scale sentence with the anchored 0-6 scale from the paper (0 = no meaning preserved, 2 = some meaning preserved, 4 = most meaning preserved with few grammar mistakes, 6 = perfect meaning and grammar). Picking the right tool ---------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/standard-quality-metrics/#picking-the-right-tool) | Need | Use | | --- | --- | | Score a quality dimension on an integer scale with explicit CoT steps | [`GEval`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.GEval) | | Grounding, relevance, retrieval quality, translation quality | The [`LLMJudge`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.LLMJudge)
recipes above | | Something bespoke | [`LLMJudge`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.LLMJudge)
with your own rubric | | Exact parity with an upstream framework | [Third-Party Integrations](https://pydantic.dev/docs/ai/evals/evaluators/framework-integrations/) | Was this page helpful? Thanks for your feedback! --- # Cerebras | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/models/cerebras/#_top) Cerebras ======== Install ------- [](https://pydantic.dev/docs/ai/models/cerebras/#install) To use `CerebrasModel`, you need to either install `pydantic-ai`, or install `pydantic-ai-slim` with the `cerebras` optional group: * [pip](https://pydantic.dev/docs/ai/models/cerebras/#tab-panel-106) * [uv](https://pydantic.dev/docs/ai/models/cerebras/#tab-panel-107) Terminal pip install "pydantic-ai-slim[cerebras]" Terminal uv add "pydantic-ai-slim[cerebras]" Configuration ------------- [](https://pydantic.dev/docs/ai/models/cerebras/#configuration) To use [Cerebras](https://cerebras.ai/) through their API, go to [cloud.cerebras.ai](https://cloud.cerebras.ai/?utm_source=3pi_pydantic-ai&utm_campaign=partner_doc) and generate an API key. For a list of available models, see the [Cerebras models documentation](https://inference-docs.cerebras.ai/models) . Environment variable -------------------- [](https://pydantic.dev/docs/ai/models/cerebras/#environment-variable) Once you have the API key, you can set it as an environment variable: Terminal export CEREBRAS_API_KEY='your-api-key' You can then use `CerebrasModel` by name: from pydantic_ai import Agent agent = Agent('cerebras:llama-3.3-70b') ... Or initialise the model directly with just the model name: from pydantic_ai import Agent from pydantic_ai.models.cerebras import CerebrasModel model = CerebrasModel('llama-3.3-70b') agent = Agent(model) ... `provider` argument ------------------- [](https://pydantic.dev/docs/ai/models/cerebras/#provider-argument) You can provide a custom `Provider` via the `provider` argument: from pydantic_ai import Agent from pydantic_ai.models.cerebras import CerebrasModel from pydantic_ai.providers.cerebras import CerebrasProvider model = CerebrasModel( 'llama-3.3-70b', provider=CerebrasProvider(api_key='your-api-key') ) agent = Agent(model) ... You can also customize the `CerebrasProvider` with a custom `httpx.AsyncClient`: from httpx import AsyncClient from pydantic_ai import Agent from pydantic_ai.models.cerebras import CerebrasModel from pydantic_ai.providers.cerebras import CerebrasProvider custom_http_client = AsyncClient(timeout=30) model = CerebrasModel( 'llama-3.3-70b', provider=CerebrasProvider(api_key='your-api-key', http_client=custom_http_client), ) agent = Agent(model) ... Was this page helpful? Thanks for your feedback! --- # Getting Started | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/graph/builder/#_top) Getting Started =============== The graph builder API provides a powerful builder pattern for constructing parallel execution graphs. The original [`BaseNode`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.BaseNode) \-based graph API is still available (and interoperable with the builder API) and is documented in the [main graph documentation](https://pydantic.dev/docs/ai/graph/graph/) . The graph builder API in `pydantic-graph` provides: * **Step nodes** for executing async functions * **Decision nodes** for conditional branching * **Spread operations** for parallel processing of iterables * **Broadcast operations** for sending the same data to multiple parallel paths * **Join nodes and Reducers** for aggregating results from parallel execution This API is designed for advanced workflows where you want declarative control over parallelism, routing, and data aggregation. Installation ------------ [](https://pydantic.dev/docs/ai/graph/builder/#installation) The graph builder API is included with `pydantic-graph`: Terminal pip install pydantic-graph Or as part of `pydantic-ai`: Terminal pip install pydantic-ai Quick Start ----------- [](https://pydantic.dev/docs/ai/graph/builder/#quick-start) Here’s a simple example to get you started: simple\_counter.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext @dataclass class CounterState: """State for tracking a counter value.""" value: int = 0 async def main(): # Create a graph builder with state and output types g = GraphBuilder(state_type=CounterState, output_type=int) # Define steps using the decorator @g.step async def increment(ctx: StepContext[CounterState, None, None]) -> int: """Increment the counter and return its value.""" ctx.state.value += 1 return ctx.state.value @g.step async def double_it(ctx: StepContext[CounterState, None, int]) -> int: """Double the input value.""" return ctx.inputs * 2 # Add edges connecting the nodes g.add( g.edge_from(g.start_node).to(increment), g.edge_from(increment).to(double_it), g.edge_from(double_it).to(g.end_node), ) # Build and run the graph graph = g.build() state = CounterState() result = await graph.run(state=state) print(f'Result: {result}') #> Result: 2 print(f'Final state: {state.value}') #> Final state: 1 _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Key Concepts ------------ [](https://pydantic.dev/docs/ai/graph/builder/#key-concepts) ### GraphBuilder [](https://pydantic.dev/docs/ai/graph/builder/#graphbuilder) The [`GraphBuilder`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder) is the main entry point for constructing graphs. It’s generic over: * `StateT` - The type of mutable state shared across all nodes * `DepsT` - The type of dependencies injected into nodes * `InputT` - The type of initial input to the graph * `OutputT` - The type of final output from the graph ### Steps [](https://pydantic.dev/docs/ai/graph/builder/#steps) Steps are async functions decorated with [`@g.step`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder.step) that define the actual work to be done in each node. They receive a [`StepContext`](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StepContext) with access to: * `ctx.state` - The mutable graph state * `ctx.deps` - Injected dependencies * `ctx.inputs` - Input data for this step ### Edges [](https://pydantic.dev/docs/ai/graph/builder/#edges) Edges define the connections between nodes. The builder provides multiple ways to create edges: * [`g.add()`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder.add) - Add one or more edge paths * [`g.add_edge()`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder.add_edge) - Add a simple edge between two nodes * [`g.edge_from()`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder.edge_from) - Start building a complex edge path ### Start and End Nodes [](https://pydantic.dev/docs/ai/graph/builder/#start-and-end-nodes) Every graph has: * [`g.start_node`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder.start_node) - The entry point receiving initial inputs * [`g.end_node`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder.end_node) - The exit point producing final outputs A More Complex Example ---------------------- [](https://pydantic.dev/docs/ai/graph/builder/#a-more-complex-example) Here’s an example showcasing parallel execution with a map operation: parallel\_processing.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, reduce_list_append @dataclass class ProcessingState: """State for tracking processing metrics.""" items_processed: int = 0 async def main(): g = GraphBuilder( state_type=ProcessingState, input_type=list[int], output_type=list[int], ) @g.step async def square(ctx: StepContext[ProcessingState, None, int]) -> int: """Square a number and track that we processed it.""" ctx.state.items_processed += 1 return ctx.inputs * ctx.inputs # Create a join to collect results collect_results = g.join(reduce_list_append, initial_factory=list[int]) # Build the graph with map operation g.add( g.edge_from(g.start_node).map().to(square), g.edge_from(square).to(collect_results), g.edge_from(collect_results).to(g.end_node), ) graph = g.build() state = ProcessingState() result = await graph.run(state=state, inputs=[1, 2, 3, 4, 5]) print(f'Results: {sorted(result)}') #> Results: [1, 4, 9, 16, 25] print(f'Items processed: {state.items_processed}') #> Items processed: 5 _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ In this example: 1. The start node receives a list of integers 2. The `.map()` operation fans out each item to a separate parallel execution of the `square` step 3. All results are collected back together using [`reduce_list_append`](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.reduce_list_append) 4. The joined results flow to the end node Next Steps ---------- [](https://pydantic.dev/docs/ai/graph/builder/#next-steps) Explore the detailed documentation for each feature: * [**Steps**](https://pydantic.dev/docs/ai/graph/builder/steps/) - Learn about step nodes and execution contexts * [**Joins**](https://pydantic.dev/docs/ai/graph/builder/joins/) - Understand join nodes and reducer patterns * [**Decisions**](https://pydantic.dev/docs/ai/graph/builder/decisions/) - Implement conditional branching * [**Parallel Execution**](https://pydantic.dev/docs/ai/graph/builder/parallel/) - Master broadcasting and mapping Advanced Execution Control -------------------------- [](https://pydantic.dev/docs/ai/graph/builder/#advanced-execution-control) Beyond the basic [`graph.run()`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.Graph.run) method, the builder API provides fine-grained control over graph execution. ### Step-by-Step Execution [](https://pydantic.dev/docs/ai/graph/builder/#step-by-step-execution) Use [`graph.iter()`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.Graph.iter) to execute the graph one step at a time: step\_by\_step.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext @dataclass class CounterState: value: int = 0 async def main(): g = GraphBuilder(state_type=CounterState, output_type=int) @g.step async def increment(ctx: StepContext[CounterState, None, None]) -> int: ctx.state.value += 1 return ctx.state.value @g.step async def double_it(ctx: StepContext[CounterState, None, int]) -> int: return ctx.inputs * 2 g.add( g.edge_from(g.start_node).to(increment), g.edge_from(increment).to(double_it), g.edge_from(double_it).to(g.end_node), ) graph = g.build() state = CounterState() # Use iter() for step-by-step execution async with graph.iter(state=state) as graph_run: print(f'Initial state: {state.value}') #> Initial state: 0 # Advance execution step by step async for event in graph_run: print(f'{state.value=} | {event=}') #> state.value=0 | event=[GraphTask(node_id='increment', inputs=None)] #> state.value=1 | event=[GraphTask(node_id='double_it', inputs=1)] #> state.value=1 | event=[GraphTask(node_id='__end__', inputs=2)] #> state.value=1 | event=EndMarker(_value=2) if graph_run.output is not None: print(f'Final output: {graph_run.output}') #> Final output: 2 break _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ The [`GraphRun`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphRun) object provides: * **Async iteration**: Iterate through execution events * **`next_task` property**: Inspect upcoming tasks * **`output` property**: Check if the graph has completed and get the final output * **`next()` method**: Manually advance execution with optional value injection ### Visualizing Graphs [](https://pydantic.dev/docs/ai/graph/builder/#visualizing-graphs) Generate Mermaid diagrams of your graph structure using [`graph.render()`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.Graph.render) : visualize\_graph.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext @dataclass class SimpleState: pass g = GraphBuilder(state_type=SimpleState, output_type=str) @g.step async def step_a(ctx: StepContext[SimpleState, None, None]) -> int: return 10 @g.step async def step_b(ctx: StepContext[SimpleState, None, int]) -> str: return f'Result: {ctx.inputs}' g.add( g.edge_from(g.start_node).to(step_a), g.edge_from(step_a).to(step_b), g.edge_from(step_b).to(g.end_node), ) graph = g.build() # Generate a Mermaid diagram mermaid_diagram = graph.render(title='My Graph', direction='LR') print(mermaid_diagram) """ --- title: My Graph --- stateDiagram-v2 direction LR step_a step_b [*] --> step_a step_a --> step_b step_b --> [*] """ The rendered diagram can be displayed in documentation, notebooks, or any tool that supports Mermaid syntax. Comparison with Original API ---------------------------- [](https://pydantic.dev/docs/ai/graph/builder/#comparison-with-original-api) The original graph API (documented in the [main graph page](https://pydantic.dev/docs/ai/graph/graph/) ) uses a class-based approach with [`BaseNode`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.BaseNode) subclasses. The builder API uses a builder pattern with decorated functions, which provides: **Advantages:** * More concise syntax for simple workflows * Explicit control over parallelism with map/broadcast * Native reducers for common aggregation patterns * Easier to visualize complex data flows **Trade-offs:** * Requires understanding of builder patterns * Less object-oriented, more functional style Both APIs are fully supported and can even be integrated together when needed. Persistence and Resumability ---------------------------- [](https://pydantic.dev/docs/ai/graph/builder/#persistence-and-resumability) For workflows that need to preserve progress across failures, restarts, or long-running operations, use one of the supported [durable execution](https://pydantic.dev/docs/ai/capabilities/durable_execution/overview/) solutions. Was this page helpful? Thanks for your feedback! --- # SQL Generation | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/data-analytics/sql-gen/#_top) SQL Generation ============== Example demonstrating how to use Pydantic AI to generate SQL queries based on user input. Demonstrates: * [dynamic system prompt](https://pydantic.dev/docs/ai/core-concepts/agent/#system-prompts) * [structured `output_type`](https://pydantic.dev/docs/ai/core-concepts/output/#structured-output) * [output validation](https://pydantic.dev/docs/ai/core-concepts/output/#output-validator-functions) * [agent dependencies](https://pydantic.dev/docs/ai/core-concepts/dependencies/) Running the Example ------------------- [](https://pydantic.dev/docs/ai/examples/data-analytics/sql-gen/#running-the-example) The resulting SQL is validated by running it as an `EXPLAIN` query on PostgreSQL. To run the example, you first need to run PostgreSQL, e.g. via Docker: Terminal docker run --rm -e POSTGRES_PASSWORD=postgres -p 54320:5432 postgres _(we run postgres on port `54320` to avoid conflicts with any other postgres instances you may have running)_ With [dependencies installed and environment variables set](https://pydantic.dev/docs/ai/examples/setup/#usage) , run: * [pip](https://pydantic.dev/docs/ai/examples/data-analytics/sql-gen/#tab-panel-36) * [uv](https://pydantic.dev/docs/ai/examples/data-analytics/sql-gen/#tab-panel-37) Terminal python -m pydantic_ai_examples.sql_gen Terminal uv run -m pydantic_ai_examples.sql_gen or to use a custom prompt: * [pip](https://pydantic.dev/docs/ai/examples/data-analytics/sql-gen/#tab-panel-38) * [uv](https://pydantic.dev/docs/ai/examples/data-analytics/sql-gen/#tab-panel-39) Terminal python -m pydantic_ai_examples.sql_gen "find me errors" Terminal uv run -m pydantic_ai_examples.sql_gen "find me errors" This model uses `gemini-3-flash-preview` by default since Gemini is good at single shot queries of this kind. Example Code ------------ [](https://pydantic.dev/docs/ai/examples/data-analytics/sql-gen/#example-code) sql\_gen.py import asyncio import sys from collections.abc import AsyncGenerator from contextlib import asynccontextmanager from dataclasses import dataclass from datetime import date from typing import Annotated, Any, TypeAlias import asyncpg import logfire from annotated_types import MinLen from devtools import debug from pydantic import BaseModel, Field from pydantic_ai import Agent, ModelRetry, RunContext, format_as_xml # 'if-token-present' means nothing will be sent (and the example will work) if you don't have logfire configured logfire.configure(send_to_logfire='if-token-present') logfire.instrument_asyncpg() logfire.instrument_pydantic_ai() DB_SCHEMA = """ CREATE TABLE records ( created_at timestamptz, start_timestamp timestamptz, end_timestamp timestamptz, trace_id text, span_id text, parent_span_id text, level log_level, span_name text, message text, attributes_json_schema text, attributes jsonb, tags text[], is_exception boolean, otel_status_message text, service_name text ); """ SQL_EXAMPLES = [\ {\ 'request': 'show me records where foobar is false',\ 'response': "SELECT * FROM records WHERE attributes->>'foobar' = false",\ },\ {\ 'request': 'show me records where attributes include the key "foobar"',\ 'response': "SELECT * FROM records WHERE attributes ? 'foobar'",\ },\ {\ 'request': 'show me records from yesterday',\ 'response': "SELECT * FROM records WHERE start_timestamp::date > CURRENT_TIMESTAMP - INTERVAL '1 day'",\ },\ {\ 'request': 'show me error records with the tag "foobar"',\ 'response': "SELECT * FROM records WHERE level = 'error' and 'foobar' = ANY(tags)",\ },\ ] @dataclass class Deps: conn: asyncpg.Connection class Success(BaseModel): """Response when SQL could be successfully generated.""" sql_query: Annotated[str, MinLen(1)] explanation: str = Field( '', description='Explanation of the SQL query, as markdown' ) class InvalidRequest(BaseModel): """Response the user input didn't include enough information to generate SQL.""" error_message: str Response: TypeAlias = Success | InvalidRequest agent = Agent[Deps, Response]( 'google:gemini-3-flash-preview', # Pass the union members directly: a `Response` type alias isn't yet accepted as a `TypeForm` value (PEP-747) output_type=Success | InvalidRequest, deps_type=Deps, ) @agent.system_prompt async def system_prompt() -> str: return f"""\ Given the following PostgreSQL table of records, your job is to write a SQL query that suits the user's request. Database schema: {DB_SCHEMA} today's date = {date.today()} {format_as_xml(SQL_EXAMPLES)} """ @agent.output_validator async def validate_output(ctx: RunContext[Deps], output: Response) -> Response: if isinstance(output, InvalidRequest): return output # gemini often adds extraneous backslashes to SQL output.sql_query = output.sql_query.replace('\\', '') if not output.sql_query.upper().startswith('SELECT'): raise ModelRetry('Please create a SELECT query') try: await ctx.deps.conn.execute(f'EXPLAIN {output.sql_query}') except asyncpg.exceptions.PostgresError as e: raise ModelRetry(f'Invalid query: {e}') from e else: return output async def main(): if len(sys.argv) == 1: prompt = 'show me logs from yesterday, with level "error"' else: prompt = sys.argv[1] async with database_connect( 'postgresql://postgres:postgres@localhost:54320', 'pydantic_ai_sql_gen' ) as conn: deps = Deps(conn) result = await agent.run(prompt, deps=deps) debug(result.output) # pyright: reportUnknownMemberType=false # pyright: reportUnknownVariableType=false @asynccontextmanager async def database_connect(server_dsn: str, database: str) -> AsyncGenerator[Any, None]: with logfire.span('check and create DB'): conn = await asyncpg.connect(server_dsn) try: db_exists = await conn.fetchval( 'SELECT 1 FROM pg_database WHERE datname = $1', database ) if not db_exists: await conn.execute(f'CREATE DATABASE {database}') finally: await conn.close() conn = await asyncpg.connect(f'{server_dsn}/{database}') try: with logfire.span('create schema'): async with conn.transaction(): if not db_exists: await conn.execute( "CREATE TYPE log_level AS ENUM ('debug', 'info', 'warning', 'error', 'critical')" ) await conn.execute(DB_SCHEMA) yield conn finally: await conn.close() if __name__ == '__main__': asyncio.run(main()) Was this page helpful? Thanks for your feedback! --- # pydantic_evals.lifecycle | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#_top) pydantic\_evals.lifecycle ========================= Case lifecycle hooks for pydantic evals. This module provides the [`CaseLifecycle`](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#pydantic_evals.lifecycle.CaseLifecycle) class, which allows defining setup, context preparation, and teardown hooks that run at different stages of case evaluation. CaseLifecycle ------------- [](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#pydantic_evals.lifecycle.CaseLifecycle) **Bases:** `Generic[InputsT, OutputT, MetadataT]` Per-case lifecycle hooks for evaluation. A new instance is created for each case during evaluation. Subclass and override any methods you need — all methods are no-ops by default. The evaluation flow for each case is: 1. `setup()` — called before task execution 2. Task runs 3. `prepare_context()` — called after task, before evaluators; can enrich metrics/attributes 4. Evaluators run 5. `teardown()` — called after evaluators complete; receives the full result (or `None` when interrupted) Exceptions raised by `setup()` or `prepare_context()` are caught and recorded as a `ReportCaseFailure`; `teardown()` is still called afterward so you can clean up. Exceptions raised by `teardown()` propagate to the caller and may abort the evaluation. If your teardown may raise and you don’t want it to crash the evaluation run, handle exceptions within your `teardown()` implementation itself. ### Constructor Parameters [](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#constructor-parameters) **`case`** : `Case`\[`InputsT`, `OutputT`, `MetadataT`\] [](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#pydantic_evals.lifecycle.CaseLifecycle.__init__(case)) The case being evaluated. Available as `self.case` in all hooks. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#attributes) #### case [](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#pydantic_evals.lifecycle.CaseLifecycle.case) The case being evaluated. **Type:** `Case`\[`InputsT`, `OutputT`, `MetadataT`\] ### Methods [](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#methods) #### prepare\_context [](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#pydantic_evals.lifecycle.CaseLifecycle.prepare_context) `@async` def prepare_context( ctx: EvaluatorContext[InputsT, OutputT, MetadataT], ) -> EvaluatorContext[InputsT, OutputT, MetadataT] Called after the task completes, before evaluators run. Override to enrich the evaluator context with additional metrics or attributes derived from the task output, span tree, or external state. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#returns) `EvaluatorContext`\[`InputsT`, `OutputT`, `MetadataT`\] — The (possibly modified) evaluator context to pass to evaluators. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#parameters) **`ctx`** : `EvaluatorContext`\[`InputsT`, `OutputT`, `MetadataT`\] [](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#pydantic_evals.lifecycle.CaseLifecycle.prepare_context(ctx)) The evaluator context produced by the task run. #### setup [](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#pydantic_evals.lifecycle.CaseLifecycle.setup) `@async` def setup() -> None Called before task execution. Override to perform per-case resource setup (e.g., create a test database, start a service). The case metadata is available via `self.case.metadata`. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#returns-1) [`None`](https://docs.python.org/3/library/constants.html#None) #### teardown [](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#pydantic_evals.lifecycle.CaseLifecycle.teardown) `@async` def teardown( result: ReportCase[InputsT, OutputT, MetadataT] | ReportCaseFailure[InputsT, OutputT, MetadataT] | None, ) -> None Called after evaluators complete. Override to perform per-case resource cleanup. The result is provided so that teardown logic can vary based on success/failure (e.g., keep resources up for inspection on failure). ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#returns-2) [`None`](https://docs.python.org/3/library/constants.html#None) ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#parameters-1) **`result`** : `ReportCase`\[`InputsT`, `OutputT`, `MetadataT`\] | `ReportCaseFailure`\[`InputsT`, `OutputT`, `MetadataT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) [](https://pydantic.dev/docs/ai/api/pydantic_evals/lifecycle/#pydantic_evals.lifecycle.CaseLifecycle.teardown(result)) The evaluation result — a `ReportCase` (success), `ReportCaseFailure`, or `None` if the run ended without a report object (e.g. cancellation). Was this page helpful? Thanks for your feedback! --- # openrouter | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/openrouter/#_top) openrouter ========== Setup ----- [](https://pydantic.dev/docs/ai/api/models/openrouter/#setup) For details on how to set up authentication with this model, see [model configuration for OpenRouter](https://pydantic.dev/docs/ai/models/openrouter/) . OpenRouterModel --------------- [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModel) **Bases:** `OpenAIChatModel` Extends OpenAIChatModel to capture extra metadata for Openrouter. ### Methods [](https://pydantic.dev/docs/ai/api/models/openrouter/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModel.__init__) def __init__( model_name: str, *, provider: Literal['openrouter'] | Provider[AsyncOpenAI] = 'openrouter', profile: ModelProfileSpec | None = None, settings: ModelSettings | None = None, ) Initialize an OpenRouter model. ##### Parameters [](https://pydantic.dev/docs/ai/api/models/openrouter/#parameters) **`model_name`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModel.__init__(model_name)) The name of the model to use. **`provider`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘openrouter’\] | `Provider`\[`AsyncOpenAI`\] _Default:_ `'openrouter'` [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModel.__init__(provider)) The provider to use for authentication and API access. If not provided, a new provider will be created with the default settings. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModel.__init__(profile)) The model profile to use. Defaults to a profile picked by the provider based on the model name. **`settings`** : [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModel.__init__(settings)) Model-specific settings that will be used as defaults for this model. #### resolve\_prompt\_cache\_retention [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModel.resolve_prompt_cache_retention) def resolve_prompt_cache_retention( model_settings: ModelSettings | None, ) -> timedelta | None Resolve the longest explicit retention accepted by OpenRouter’s downstream model. ##### Returns [](https://pydantic.dev/docs/ai/api/models/openrouter/#returns) `timedelta` | [`None`](https://docs.python.org/3/library/constants.html#None) #### supported\_native\_tools [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModel.supported_native_tools) `@classmethod` def supported_native_tools(cls) -> frozenset[type[AbstractNativeTool]] Return the set of builtin tool types this model can handle. OpenRouter supports web search through its server-tool API. ##### Returns [](https://pydantic.dev/docs/ai/api/models/openrouter/#returns-1) [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[[`type`](https://docs.python.org/3/glossary.html#term-type)\ \[`AbstractNativeTool`\]\] OpenRouterModelSettings ----------------------- [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModelSettings) **Bases:** [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) Settings used for an OpenRouter model request. ### Attributes [](https://pydantic.dev/docs/ai/api/models/openrouter/#attributes) #### openrouter\_cache\_instructions [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModelSettings.openrouter_cache_instructions) Whether to add `cache_control` to stable system instructions. When enabled, supported downstream providers (Anthropic, Gemini) can cache stable system instructions and reduce costs. If dynamic instructions are present, the cache point is placed before them, matching Anthropic’s static-prefix caching behavior. For Gemini models, this setting is ignored when dynamic instructions are present because OpenRouter normalizes system/developer messages into a single immutable `systemInstruction`. Ignored for other downstream providers. If `True`, uses TTL=‘5m’. You can also specify ‘5m’ or ‘1h’ directly. TTL is only included for Anthropic models; Gemini does not support explicit TTL. See [https://openrouter.ai/docs/guides/best-practices/prompt-caching](https://openrouter.ai/docs/guides/best-practices/prompt-caching) for more information. **Type:** `OpenRouterCacheTTL` #### openrouter\_cache\_messages [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModelSettings.openrouter_cache_messages) Convenience setting to enable caching for the last message in the conversation. When enabled, this automatically adds `cache_control` to the last content block in the final message (regardless of role), which is useful for Anthropic’s prefix-based caching in multi-turn conversations. In tool-use flows, this may target a tool result message rather than a user message, which is correct for prefix caching. Ignored for downstream providers that do not support explicit cache control. If `True`, uses TTL=‘5m’. You can also specify ‘5m’ or ‘1h’ directly. TTL is only included for Anthropic models; Gemini does not support explicit TTL. Note: OpenRouter uses only the last breakpoint across normal message content for Gemini caching. Use this when caching the final message boundary is intentional; use `openrouter_cache_instructions` for stable system context. Anthropic supports prefix-based caching across multi-turn conversations with this setting. See [https://openrouter.ai/docs/guides/best-practices/prompt-caching](https://openrouter.ai/docs/guides/best-practices/prompt-caching) for more information. **Type:** `OpenRouterCacheTTL` #### openrouter\_cache\_tool\_definitions [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModelSettings.openrouter_cache_tool_definitions) Whether to add `cache_control` to the last tool definition. When enabled, the last tool in the `tools` array will have `cache_control` set, allowing supported downstream providers to cache tool definitions and reduce costs. Ignored for downstream providers that do not support explicit tool definition caching. If `True`, uses TTL=‘5m’. You can also specify ‘5m’ or ‘1h’ directly. TTL is only included for Anthropic models. Currently only effective for Anthropic models via OpenRouter, as tool definition caching is not documented for other providers. See [https://openrouter.ai/docs/guides/best-practices/prompt-caching](https://openrouter.ai/docs/guides/best-practices/prompt-caching) for more information. **Type:** `OpenRouterCacheTTL` #### openrouter\_models [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModelSettings.openrouter_models) A list of fallback models. These models will be tried, in order, if the main model returns an error. [See details](https://openrouter.ai/docs/features/model-routing#the-models-parameter) **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] #### openrouter\_preset [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModelSettings.openrouter_preset) Presets allow you to separate your LLM configuration from your code. Create and manage presets through the OpenRouter web application to control provider routing, model selection, system prompts, and other parameters, then reference them in OpenRouter API requests. [See more](https://openrouter.ai/docs/features/presets) **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### openrouter\_provider [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModelSettings.openrouter_provider) OpenRouter routes requests to the best available providers for your model. By default, requests are load balanced across the top providers to maximize uptime. You can customize how your requests are routed using the provider object. [See more](https://openrouter.ai/docs/features/provider-routing) **Type:** `OpenRouterProviderConfig` #### openrouter\_reasoning [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModelSettings.openrouter_reasoning) To control the reasoning tokens in the request. The reasoning config object consolidates settings for controlling reasoning strength across different models. [See more](https://openrouter.ai/docs/use-cases/reasoning-tokens) **Type:** `OpenRouterReasoning` #### openrouter\_transforms [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModelSettings.openrouter_transforms) To help with prompts that exceed the maximum context size of a model. Transforms work by removing or truncating messages from the middle of the prompt, until the prompt fits within the model’s context window. [See more](https://openrouter.ai/docs/features/message-transforms) **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[`OpenRouterTransforms`\] #### openrouter\_usage [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModelSettings.openrouter_usage) To control the usage of the model. The usage config object consolidates settings for enabling detailed usage information. [See more](https://openrouter.ai/docs/use-cases/usage-accounting) **Type:** `OpenRouterUsageConfig` OpenRouterProviderConfig ------------------------ [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterProviderConfig) **Bases:** [`TypedDict`](https://docs.python.org/3/library/typing.html#typing.TypedDict) Represents the ‘Provider’ object from the OpenRouter API. ### Attributes [](https://pydantic.dev/docs/ai/api/models/openrouter/#attributes-1) #### allow\_fallbacks [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterProviderConfig.allow_fallbacks) Whether to allow backup providers when the primary is unavailable. [See details](https://openrouter.ai/docs/features/provider-routing#disabling-fallbacks) **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### data\_collection [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterProviderConfig.data_collection) Control whether to use providers that may store data. [See details](https://openrouter.ai/docs/features/provider-routing#requiring-providers-to-comply-with-data-policies) **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘allow’, ‘deny’\] #### ignore [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterProviderConfig.ignore) List of provider slugs to skip for this request. [See details](https://openrouter.ai/docs/features/provider-routing#ignoring-providers) **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] #### max\_price [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterProviderConfig.max_price) The maximum pricing you want to pay for this request. [See details](https://openrouter.ai/docs/features/provider-routing#max-price) **Type:** `_OpenRouterMaxPrice` #### only [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterProviderConfig.only) List of provider slugs to allow for this request. [See details](https://openrouter.ai/docs/features/provider-routing#allowing-only-specific-providers) **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[`OpenRouterProviderName`\] #### order [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterProviderConfig.order) List of provider slugs to try in order (e.g. \[“anthropic”, “openai”\]). [See details](https://openrouter.ai/docs/features/provider-routing#ordering-specific-providers) **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[`OpenRouterProviderName`\] #### quantizations [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterProviderConfig.quantizations) List of quantization levels to filter by (e.g. \[“int4”, “int8”\]). [See details](https://openrouter.ai/docs/features/provider-routing#quantization) **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal)\ \[‘int4’, ‘int8’, ‘fp4’, ‘fp6’, ‘fp8’, ‘fp16’, ‘bf16’, ‘fp32’, ‘unknown’\]\] #### require\_parameters [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterProviderConfig.require_parameters) Only use providers that support all parameters in your request. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### sort [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterProviderConfig.sort) Sort providers by price or throughput. (e.g. “price” or “throughput”). [See details](https://openrouter.ai/docs/features/provider-routing#provider-sorting) **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘price’, ‘throughput’, ‘latency’\] #### zdr [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterProviderConfig.zdr) Restrict routing to only ZDR (Zero Data Retention) endpoints. [See details](https://openrouter.ai/docs/features/provider-routing#zero-data-retention-enforcement) **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) OpenRouterReasoning ------------------- [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterReasoning) **Bases:** [`TypedDict`](https://docs.python.org/3/library/typing.html#typing.TypedDict) Configuration for reasoning tokens in OpenRouter requests. Reasoning tokens allow models to show their step-by-step thinking process. You can configure this using either OpenAI-style effort levels or Anthropic-style token limits, but not both simultaneously. ### Attributes [](https://pydantic.dev/docs/ai/api/models/openrouter/#attributes-2) #### effort [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterReasoning.effort) OpenAI-style reasoning effort level. Cannot be used with max\_tokens. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘xhigh’, ‘high’, ‘medium’, ‘low’, ‘minimal’, ‘none’\] #### enabled [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterReasoning.enabled) Whether to enable reasoning with default parameters. Default is inferred from effort or max\_tokens. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### exclude [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterReasoning.exclude) Whether to exclude reasoning tokens from the response. Default is False. All models support this. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### max\_tokens [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterReasoning.max_tokens) Anthropic-style specific token limit for reasoning. Cannot be used with effort. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) OpenRouterStreamedResponse -------------------------- [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterStreamedResponse) **Bases:** `OpenAIStreamedResponse` Implementation of `StreamedResponse` for OpenRouter models. OpenRouterUsageConfig --------------------- [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterUsageConfig) **Bases:** [`TypedDict`](https://docs.python.org/3/library/typing.html#typing.TypedDict) Configuration for OpenRouter usage. KnownOpenRouterProviders ------------------------ [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.KnownOpenRouterProviders) Known providers in the OpenRouter marketplace **Default:** `Literal['z-ai', 'cerebras', 'venice', 'moonshotai', 'morph', 'stealth', 'wandb', 'klusterai', 'openai', 'sambanova', 'amazon-bedrock', 'mistral', 'nextbit', 'atoma', 'ai21', 'minimax', 'baseten', 'anthropic', 'featherless', 'groq', 'lambda', 'azure', 'ncompass', 'deepseek', 'hyperbolic', 'crusoe', 'cohere', 'mancer', 'avian', 'perplexity', 'novita', 'siliconflow', 'switchpoint', 'xai', 'inflection', 'fireworks', 'deepinfra', 'inference-net', 'inception', 'atlas-cloud', 'nvidia', 'alibaba', 'friendli', 'infermatic', 'targon', 'ubicloud', 'aion-labs', 'liquid', 'nineteen', 'cloudflare', 'nebius', 'chutes', 'enfer', 'crofai', 'open-inference', 'phala', 'gmicloud', 'meta', 'relace', 'parasail', 'together', 'google-ai-studio', 'google-vertex']` OpenRouterCacheTTL ------------------ [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterCacheTTL) Cache breakpoint time-to-live for OpenRouter prompt caching. `True` selects the default TTL (‘5m’); ‘5m’ or ‘1h’ may be given explicitly. The TTL is only forwarded to downstream providers that support it (Anthropic); it is omitted for Gemini. **Default:** `bool | Literal['5m', '1h']` OpenRouterProviderName ---------------------- [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterProviderName) Possible OpenRouter provider names. Since OpenRouter is constantly updating their list of providers, we explicitly list some known providers but allow any name in the type hints. See [the OpenRouter API](https://openrouter.ai/docs/api-reference/list-available-providers) for a full list. **Default:** `str | KnownOpenRouterProviders` OpenRouterTransforms -------------------- [](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterTransforms) Available messages transforms for OpenRouter models with limited token windows. Currently only supports ‘middle-out’, but is expected to grow in the future. **Default:** `Literal['middle-out']` Was this page helpful? Thanks for your feedback! --- # Image Generation | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/image-generation/#_top) Image Generation ================ The [`ImageGeneration`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ImageGeneration) [capability](https://pydantic.dev/docs/ai/capabilities/overview/) lets your agent generate images. Like all [provider-adaptive tools](https://pydantic.dev/docs/ai/capabilities/overview/#provider-adaptive-tools) , it uses the provider’s native image generation when available, with an optional subagent fallback for other models. [`ImageGeneration`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ImageGeneration) defaults to native-only. Backed by [`ImageGenerationTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.ImageGenerationTool) on the native side (see [Image Generation Tool](https://pydantic.dev/docs/ai/tools-toolsets/native-tools/#image-generation-tool) for provider support and configuration) — pass `native=ImageGenerationTool(...)` directly for full control. For the local side, pass `fallback_model='…'` to delegate unsupported requests to a subagent running an image-generation-capable model (e.g. `openai-responses:gpt-5.4`), or `local=` with any callable, [`Tool`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.Tool) , or [`AbstractToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AbstractToolset) for a custom generator. image\_generation.py from pydantic_ai.capabilities import ImageGeneration # Native-only — raises on models without native image generation ImageGeneration() # Native preferred; subagent fallback for unsupported models ImageGeneration(fallback_model='openai-responses:gpt-5.4') # Native preferred; custom callable as fallback def my_generator(prompt: str) -> bytes: ... ImageGeneration(local=my_generator) Was this page helpful? Thanks for your feedback! --- # test | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/test/#_top) test ==== Utility model for quickly testing apps built with Pydantic AI. Here’s a minimal example: test\_model\_usage.py from pydantic_ai import Agent from pydantic_ai.models.test import TestModel my_agent = Agent('openai:gpt-5.2', instructions='...') async def test_my_agent(): """Unit test for my_agent, to be run by pytest.""" m = TestModel() with my_agent.override(model=m): result = await my_agent.run('Testing my agent...') assert result.output == 'success (no tool calls)' assert m.last_model_request_parameters.function_tools == [] See [Unit testing with `TestModel`](https://pydantic.dev/docs/ai/guides/testing/#unit-testing-with-testmodel) for detailed documentation. TestModel --------- [](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestModel) **Bases:** `Model` A model specifically for testing purposes. This will (by default) call all tools in the agent, then return a tool response if possible, otherwise a plain response. How useful this model is will vary significantly. Apart from `__init__` derived by the `dataclass` decorator, all methods are private or match those of the base class. ### Attributes [](https://pydantic.dev/docs/ai/api/models/test/#attributes) #### call\_tools [](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestModel.call_tools) List of tools to call. If `'all'`, all tools will be called. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘all’\] **Default:** `call_tools` #### custom\_output\_args [](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestModel.custom_output_args) If set, these args will be passed to the output tool. **Type:** [`Any`](https://docs.python.org/3/library/typing.html#typing.Any) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `custom_output_args` #### custom\_output\_text [](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestModel.custom_output_text) If set, this text is returned as the final output. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `custom_output_text` #### last\_model\_request\_parameters [](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestModel.last_model_request_parameters) The last ModelRequestParameters passed to the model in a request. The ModelRequestParameters contains information about the function and output tools available during request handling. This is set when a request is made, so will reflect the function tools from the last step of the last run. **Type:** `ModelRequestParameters` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### model\_name [](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestModel.model_name) The model name. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### seed [](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestModel.seed) Seed for generating random data. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `seed` #### system [](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestModel.system) The model provider. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) ### Methods [](https://pydantic.dev/docs/ai/api/models/test/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestModel.__init__) def __init__( *, call_tools: list[str] | Literal['all'] = 'all', custom_output_text: str | None = None, custom_output_args: Any | None = None, seed: int = 0, model_name: str = 'test', profile: ModelProfileSpec | None = None, settings: ModelSettings | None = None, ) Initialize TestModel with optional settings and profile. #### supported\_native\_tools [](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestModel.supported_native_tools) `@classmethod` def supported_native_tools(cls) -> frozenset[type[AbstractNativeTool]] TestModel supports all native tools for testing flexibility. `ToolSearchTool` is excluded because TestModel can’t emulate provider-native tool search. Auto-injected `ToolSearch` capabilities work transparently thanks to the local `search_tools` fallback. ##### Returns [](https://pydantic.dev/docs/ai/api/models/test/#returns) [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[[`type`](https://docs.python.org/3/glossary.html#term-type)\ \[`AbstractNativeTool`\]\] TestStreamedResponse -------------------- [](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestStreamedResponse) **Bases:** `StreamedResponse` A structured response that streams test data. ### Attributes [](https://pydantic.dev/docs/ai/api/models/test/#attributes-1) #### model\_name [](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestStreamedResponse.model_name) Get the model name of the response. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### provider\_name [](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestStreamedResponse.provider_name) Get the provider name. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### provider\_url [](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestStreamedResponse.provider_url) Get the provider base URL. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### timestamp [](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestStreamedResponse.timestamp) Get the timestamp of the response. **Type:** [`datetime`](https://docs.python.org/3/library/datetime.html#module-datetime) Was this page helpful? Thanks for your feedback! --- # Overview | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/overview/#_top) Overview ======== A capability is a reusable, composable unit of agent behavior. Instead of threading multiple arguments through your `Agent` constructor — [instructions](https://pydantic.dev/docs/ai/core-concepts/agent/#instructions) here, [model settings](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) there, a [toolset](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) somewhere else, a [history processor](https://pydantic.dev/docs/ai/core-concepts/message-history/#processing-message-history) on yet another parameter — you can bundle related behavior into a single capability and pass it via the [`capabilities`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.__init__) parameter. Capabilities can provide any combination of: * **Tools** — via [toolsets](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) or [native tools](https://pydantic.dev/docs/ai/tools-toolsets/native-tools/) * **Lifecycle hooks** — intercept and modify model requests, tool calls, and the overall run * **Instructions** — static or dynamic [instruction](https://pydantic.dev/docs/ai/core-concepts/agent/#instructions) additions * **Model settings** — static or per-step [model settings](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) * **Models** — static or adaptive model selection and application-specific model ID resolution This makes them the primary extension point for Pydantic AI. Whether you’re building a memory system, a guardrail, a cost tracker, or an approval workflow, a capability is the right abstraction. Capabilities can be always-on or [loaded by the model on demand](https://pydantic.dev/docs/ai/capabilities/on-demand/) . Pydantic AI ships the built-in capabilities below, [Pydantic AI Harness](https://pydantic.dev/docs/ai/capabilities/overview/#pydantic-ai-harness) and [third-party packages](https://pydantic.dev/docs/ai/capabilities/third-party/) provide many more, and you can define your own — [declaratively](https://pydantic.dev/docs/ai/capabilities/overview/#bundling-behavior-with-capability) or by [subclassing](https://pydantic.dev/docs/ai/capabilities/custom/) . To run agents durably across failures, restarts, and long waits, see [Durable Execution](https://pydantic.dev/docs/ai/capabilities/durable_execution/overview/) . Built-in capabilities --------------------- [](https://pydantic.dev/docs/ai/capabilities/overview/#built-in-capabilities) Pydantic AI ships with several capabilities that cover common needs: | Capability | What it provides | Spec | | --- | --- | --- | | [`Thinking`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Thinking) | Enables model [thinking/reasoning](https://pydantic.dev/docs/ai/capabilities/thinking/)
at configurable effort | Yes | | [`Hooks`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Hooks) | Decorator-based [lifecycle hook](https://pydantic.dev/docs/ai/core-concepts/hooks/)
registration | — | | [`Instrumentation`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Instrumentation) | OpenTelemetry/Logfire [tracing](https://pydantic.dev/docs/ai/capabilities/instrumentation/)
of runs, model requests, and tool calls | Yes | | [`SelectModel`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.SelectModel) | Selects a static or per-step [model](https://pydantic.dev/docs/ai/capabilities/select-model/)
with a callable | — | | [`ResolveModelId`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ResolveModelId) | Resolves custom [model IDs](https://pydantic.dev/docs/ai/capabilities/resolve-model-id/)
with a callable | — | | [`WebSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebSearch) | [Web search](https://pydantic.dev/docs/ai/capabilities/web-search/)
— native by default, optional [local fallback](https://pydantic.dev/docs/ai/tools-toolsets/common-tools/#duckduckgo-search-tool)
via `local='duckduckgo'` | Yes | | [`WebFetch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebFetch) | [URL fetching](https://pydantic.dev/docs/ai/capabilities/web-fetch/)
— native by default, optional [local fallback](https://pydantic.dev/docs/ai/tools-toolsets/common-tools/#web-fetch-tool)
via `local=True` | Yes | | [`ImageGeneration`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ImageGeneration) | [Image generation](https://pydantic.dev/docs/ai/capabilities/image-generation/)
— native by default, optional subagent fallback via `fallback_model` | Yes | | [`XSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.XSearch) | [X search](https://pydantic.dev/docs/ai/capabilities/x-search/)
— native on xAI, explicit subagent fallback via `fallback_model` | Yes | | [`MCP`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.MCP) | [MCP server](https://pydantic.dev/docs/ai/capabilities/mcp/)
— runs locally by default; `native=True` opts into the model provider’s native MCP support | Yes | | [`ToolSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ToolSearch) | [Discovery](https://pydantic.dev/docs/ai/capabilities/tool-search/)
of [deferred tools](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-search)
— native when supported, local `search_tools` function tool otherwise | Yes | | [`PrepareTools`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.PrepareTools) | Filters or modifies function [tool definitions](https://pydantic.dev/docs/ai/capabilities/prepare-tools/)
per step | — | | [`PrepareOutputTools`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.PrepareOutputTools) | Filters or modifies [output tool](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.ToolOutput)
[definitions](https://pydantic.dev/docs/ai/capabilities/prepare-tools/)
per step | — | | [`PrefixTools`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.PrefixTools) | Wraps a capability and [prefixes its tool names](https://pydantic.dev/docs/ai/capabilities/prefix-tools/) | Yes | | [`NativeTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.NativeTool) | Registers a [native tool](https://pydantic.dev/docs/ai/tools-toolsets/native-tools/)
with the agent | Yes | | [`Capability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Capability) | Bundles instructions, function tools, and toolsets [without subclassing](https://pydantic.dev/docs/ai/capabilities/on-demand/#the-capability-convenience-class) | — | | [`Toolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Toolset) | Wraps an [`AbstractToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AbstractToolset) | — | | [`IncludeToolReturnSchemas`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.IncludeToolReturnSchemas) | [Includes return type schemas](https://pydantic.dev/docs/ai/capabilities/include-tool-return-schemas/)
in tool definitions sent to the model | Yes | | [`SetToolMetadata`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.SetToolMetadata) | [Merges metadata key-value pairs](https://pydantic.dev/docs/ai/capabilities/set-tool-metadata/)
onto selected tools | Yes | | [`RaiseContentFilterError`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.RaiseContentFilterError) | [Raises](https://pydantic.dev/docs/ai/capabilities/raise-content-filter-error/)
[`ContentFilterError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ContentFilterError)
whenever a model response has `finish_reason='content_filter'` | Yes | | [`ReinjectSystemPrompt`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ReinjectSystemPrompt) | [Reinjects the configured system prompt](https://pydantic.dev/docs/ai/capabilities/reinject-system-prompt/)
when the incoming message history is missing one | Yes | | [`HandleDeferredToolCalls`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.HandleDeferredToolCalls) | Resolves [deferred tool calls](https://pydantic.dev/docs/ai/capabilities/handle-deferred-tool-calls/)
inline with a handler function | — | | [`ProcessHistory`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ProcessHistory) | Wraps a [history processor](https://pydantic.dev/docs/ai/capabilities/process-history/) | — | | [`ProcessEventStream`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ProcessEventStream) | Forwards [agent stream events](https://pydantic.dev/docs/ai/capabilities/process-event-stream/)
to a handler function | — | | [`UseThreadExecutor`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.UseThreadExecutor) | Uses a [custom thread executor](https://pydantic.dev/docs/ai/capabilities/thread-executor/)
for [sync functions](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#thread-executor-for-long-running-servers) | — | The **Spec** column indicates whether the capability can be used in [agent specs](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) (YAML/JSON). Capabilities marked **—** take non-serializable arguments (callables, toolset objects) and can only be used in Python code. [Compaction](https://pydantic.dev/docs/ai/capabilities/compaction/) keeps conversations within the context window through several approaches; the provider-native [`OpenAICompaction`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAICompaction) and [`AnthropicCompaction`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicCompaction) capabilities live in the corresponding model modules. The [durable execution](https://pydantic.dev/docs/ai/capabilities/durable_execution/overview/) integrations also ship as capabilities — [`TemporalDurability`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.temporal.TemporalDurability) , [`DBOSDurability`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.dbos.DBOSDurability) , and [`PrefectDurability`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.prefect.PrefectDurability) — in the `pydantic_ai.durable_exec` subpackages. native\_capabilities.py from pydantic_ai import Agent from pydantic_ai.capabilities import Thinking, WebSearch agent = Agent( 'anthropic:claude-opus-4-6', instructions='You are a research assistant. Be thorough and cite sources.', capabilities=[\ Thinking(effort='high'),\ WebSearch(local='duckduckgo'),\ ], ) [Instructions](https://pydantic.dev/docs/ai/core-concepts/agent/#instructions) and [model settings](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) are configured directly via the `instructions` and `model_settings` parameters on `Agent` (or [`AgentSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentSpec) ). Capabilities are for behavior that goes beyond simple configuration — tools, lifecycle hooks, and custom extensions. They compose well, especially when you want to reuse the same configuration across multiple agents or load it from a [spec file](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) . Bundling behavior with `Capability` ----------------------------------- [](https://pydantic.dev/docs/ai/capabilities/overview/#bundling-behavior-with-capability) You don’t need a subclass to define a capability of your own: [`Capability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Capability) bundles instructions, function tools, and [toolsets](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) declaratively — think of it as defining a skill: capability\_shorthand.py from pydantic_ai import Agent from pydantic_ai.capabilities import Capability refunds = Capability( id='refunds', description='Use for refund eligibility and refund status.', instructions='Always confirm the order ID before issuing a refund.', ) @refunds.tool_plain def refund_status(order_id: str) -> str: """Look up the refund status for an order.""" return f'Order {order_id}: refund issued on 2026-05-01.' agent = Agent('openai:gpt-5.2', capabilities=[refunds]) Add `defer_loading=True` and the bundle becomes an [on-demand capability](https://pydantic.dev/docs/ai/capabilities/on-demand/) that stays collapsed to a one-line catalog entry until the model loads it — the same shape as [Agent Skills](https://pydantic.dev/docs/ai/capabilities/on-demand/#loading-skills-from-markdown-files) , which you can wrap in a `Capability` directly. See [The `Capability` convenience class](https://pydantic.dev/docs/ai/capabilities/on-demand/#the-capability-convenience-class) for the full API. For behavior beyond instructions, tools, and toolsets — lifecycle hooks, model settings, native tools — subclass [`AbstractCapability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability) as covered in [Building Custom Capabilities](https://pydantic.dev/docs/ai/capabilities/custom/) . Provider-adaptive tools ----------------------- [](https://pydantic.dev/docs/ai/capabilities/overview/#provider-adaptive-tools) [`WebSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebSearch) , [`WebFetch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebFetch) , [`ImageGeneration`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ImageGeneration) , [`XSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.XSearch) , and [`MCP`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.MCP) each cover a single capability (web search, URL fetch, image generation, X search, MCP) across two implementations: * **Native** — invoked by the model provider when the model supports it. The work happens on the provider’s side (e.g. Anthropic’s web search runs server-side, returning results inline). * **Local** — runs in your Python process. Used when the model doesn’t support the native tool; your code does the work (e.g. calling DuckDuckGo directly). | Capability | Local fallback | Notes | | --- | --- | --- | | [`WebSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebSearch) | `local='duckduckgo'` or `local=True` (DuckDuckGo) | Requires the `duckduckgo` optional group | | [`WebFetch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebFetch) | `local=True` (markdownify-based fetch) | Requires the `web-fetch` optional group | | [`ImageGeneration`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ImageGeneration) | Subagent via `fallback_model=` | Delegates to a model that supports native image generation | | [`XSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.XSearch) | Subagent via `fallback_model=` | No default non-xAI fallback; set `fallback_model` to an xAI model that supports [`XSearchTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.XSearchTool) | | [`MCP`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.MCP) | Direct connection to the MCP server (the default) | Accepts any [`MCPToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset)
input; transport is auto-detected from a URL | Because these capabilities contribute model-facing tools, their `id`, `description`, and `defer_loading` fields are meaningful: set them when that tool should stay hidden until the model loads the matching workflow with the `load_capability` tool. This includes [`ImageGeneration`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ImageGeneration) when image generation should only be available for an image-specific workflow, whether it resolves to a native image tool or a fallback subagent tool. Configure each side via the `native=` and `local=` kwargs. `native=` accepts `True` (use the capability’s default [native tool](https://pydantic.dev/docs/ai/tools-toolsets/native-tools/) instance), `False` (disable native), or an explicit instance like `WebSearchTool(...)` for fine-grained config. `local=` accepts `True` (the bundled local fallback, on capabilities that have one — `WebSearch` and `WebFetch`), `False` (disable local), a named strategy string where supported, or any callable, [`Tool`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.Tool) , or [`AbstractToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AbstractToolset) . Optional installs needed for the local fallback are opt-in — the capability raises a [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) at construction (with an install hint) when you ask for a local strategy whose extra isn’t installed. provider\_adaptive\_tools.py from pydantic_ai import Agent from pydantic_ai.capabilities import MCP, ImageGeneration, WebFetch, WebSearch, XSearch agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[\ # Native when supported; DuckDuckGo fallback on unsupported models\ WebSearch(local='duckduckgo'),\ # Native when supported; markdownify-based fallback on unsupported models\ WebFetch(local=True),\ # Native when supported; subagent fallback via `fallback_model`\ ImageGeneration(fallback_model='openai-responses:gpt-5.4'),\ # Native on xAI; on other models, explicitly delegate to an xAI model\ XSearch(fallback_model='xai:grok-4.3'),\ # Runs the MCP server locally by default; pass `native=True` to also advertise native MCP\ MCP('https://mcp.example.com/api'),\ ], ) `MCP` defaults the other way from the others: because MCP carries credentials, it runs locally by default and you opt into native MCP with `native=True`. The others default to native and you opt into local with `local=`. [`XSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.XSearch) is slightly different from [`WebSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebSearch) and [`WebFetch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebFetch) : there is no default non-xAI fallback. If your agent is not running on an xAI model, set `fallback_model` explicitly to an xAI model that supports [`XSearchTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.XSearchTool) . Some constraint fields require the native tool (the bundled local fallback can’t enforce them) — passing them locks the capability to the native path. If the model doesn’t support the native tool, the capability raises a [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) . constraints.py # Limit to 5 searches per run — requires native (the local fallback can't track call count) WebSearch(max_uses=5) # Only fetch example.com — enforced locally when native is unavailable WebFetch(allowed_domains=['example.com'], local=True) ### Building your own [](https://pydantic.dev/docs/ai/capabilities/overview/#building-your-own) All five capabilities are subclasses of [`NativeOrLocalTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.NativeOrLocalTool) , which you can use directly or subclass to build your own provider-adaptive tools. For example, to pair [`CodeExecutionTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.CodeExecutionTool) with a local fallback: custom\_native\_or\_local.py from pydantic_ai.native_tools import CodeExecutionTool from pydantic_ai.capabilities import NativeOrLocalTool cap = NativeOrLocalTool(native=CodeExecutionTool(), local=my_local_executor) Pydantic AI Harness ------------------- [](https://pydantic.dev/docs/ai/capabilities/overview/#pydantic-ai-harness) [**Pydantic AI Harness**](https://pydantic.dev/docs/ai/harness/) is the official capability library for Pydantic AI — standalone capabilities like memory, guardrails, context management, and [code mode](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/code_mode) live there rather than in core. See [What goes where?](https://pydantic.dev/docs/ai/harness/#what-goes-where) for the full breakdown, or jump to the [capability matrix](https://github.com/pydantic/pydantic-ai-harness#capability-matrix) . Third-party capabilities ------------------------ [](https://pydantic.dev/docs/ai/capabilities/overview/#third-party-capabilities) Third-party packages publish capabilities of their own — see [Third-Party Capabilities](https://pydantic.dev/docs/ai/capabilities/third-party/) for the ecosystem, and [Publishing capabilities](https://pydantic.dev/docs/ai/capabilities/custom/#publishing-capabilities) for making your own capability available to others. Was this page helpful? Thanks for your feedback! --- # Quick Start | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/getting-started/quick-start/#_top) Quick Start =========== **Pydantic Evals** is a powerful evaluation framework for systematically testing and evaluating AI systems, from simple LLM calls to complex multi-agent applications. What is Pydantic Evals? ----------------------- [](https://pydantic.dev/docs/ai/evals/getting-started/quick-start/#what-is-pydantic-evals) Pydantic Evals helps you: * **Create test datasets** with type-safe structured inputs and expected outputs * **Run evaluations** against your AI systems with automatic concurrency * **Score results** using deterministic checks, LLM judges, or custom evaluators * **Generate reports** with detailed metrics, assertions, and performance data * **Track changes** by comparing evaluation runs over time * **Integrate with Logfire** for visualization and collaborative analysis Installation ------------ [](https://pydantic.dev/docs/ai/evals/getting-started/quick-start/#installation) Terminal pip install pydantic-evals For OpenTelemetry tracing and Logfire integration: Terminal pip install 'pydantic-evals[logfire]' Quick Start ----------- [](https://pydantic.dev/docs/ai/evals/getting-started/quick-start/#quick-start) While evaluations are typically used to test AI systems, the Pydantic Evals framework works with any function call. To demonstrate the core functionality, we’ll start with a simple, deterministic example. Here’s a complete example of evaluating a simple text transformation function: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import Contains, EqualsExpected # Create a dataset with test cases dataset = Dataset( name='uppercase_tests', cases=[\ Case(\ name='uppercase_basic',\ inputs='hello world',\ expected_output='HELLO WORLD',\ ),\ Case(\ name='uppercase_with_numbers',\ inputs='hello 123',\ expected_output='HELLO 123',\ ),\ ], evaluators=[\ EqualsExpected(), # Check exact match with expected_output\ Contains(value='HELLO', case_sensitive=True), # Check contains "HELLO"\ ], ) # Define the function to evaluate def uppercase_text(text: str) -> str: return text.upper() # Run the evaluation report = dataset.evaluate_sync(uppercase_text) # Print the results report.print() """ Evaluation Summary: uppercase_text ┏━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━┓ ┃ Case ID ┃ Assertions ┃ Duration ┃ ┡━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━┩ │ uppercase_basic │ ✔✔ │ 10ms │ ├────────────────────────┼────────────┼──────────┤ │ uppercase_with_numbers │ ✔✔ │ 10ms │ ├────────────────────────┼────────────┼──────────┤ │ Averages │ 100.0% ✔ │ 10ms │ └────────────────────────┴────────────┴──────────┘ """ Output: Evaluation Summary: uppercase_text ┏━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━┓ ┃ Case ID ┃ Assertions ┃ Duration ┃ ┡━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━┩ │ uppercase_basic │ ✔✔ │ 10ms │ ├─────────────────────────┼────────────┼──────────┤ │ uppercase_with_numbers │ ✔✔ │ 10ms │ ├─────────────────────────┼────────────┼──────────┤ │ Averages │ 100.0% ✔ │ 10ms │ └─────────────────────────┴────────────┴──────────┘ Key Concepts ------------ [](https://pydantic.dev/docs/ai/evals/getting-started/quick-start/#key-concepts) Understanding a few core concepts will help you get the most out of Pydantic Evals: * **[`Dataset`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset) ** - A collection of test cases and (optional) evaluators * **[`Case`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Case) ** - A single test scenario with inputs and optional expected outputs and case-specific evaluators * **[`Evaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Evaluator) ** - A function that scores or validates task outputs * **[`EvaluationReport`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.EvaluationReport) ** - Results from running an evaluation For a deeper dive, see [Core Concepts](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/) . Common Use Cases ---------------- [](https://pydantic.dev/docs/ai/evals/getting-started/quick-start/#common-use-cases) ### Deterministic Validation [](https://pydantic.dev/docs/ai/evals/getting-started/quick-start/#deterministic-validation) Test that your AI system produces correctly-structured outputs: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import Contains, IsInstance dataset = Dataset( name='dict_validation', cases=[\ Case(inputs={'data': 'required_key present'}, expected_output={'result': 'success'}),\ ], evaluators=[\ IsInstance(type_name='dict'),\ Contains(value='required_key'),\ ], ) ### LLM-as-a-Judge Evaluation [](https://pydantic.dev/docs/ai/evals/getting-started/quick-start/#llm-as-a-judge-evaluation) Use an LLM to evaluate subjective qualities like accuracy or helpfulness: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import LLMJudge dataset = Dataset( name='llm_judge_test', cases=[\ Case(inputs='What is the capital of France?', expected_output='Paris'),\ ], evaluators=[\ LLMJudge(\ rubric='Response is accurate and helpful',\ include_input=True,\ model='anthropic:claude-sonnet-4-6',\ )\ ], ) ### Performance Testing [](https://pydantic.dev/docs/ai/evals/getting-started/quick-start/#performance-testing) Ensure your system meets performance requirements: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import MaxDuration dataset = Dataset( name='performance_test', cases=[\ Case(inputs='test input', expected_output='test output'),\ ], evaluators=[\ MaxDuration(seconds=2.0),\ ], ) Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/getting-started/quick-start/#next-steps) Explore the documentation to learn more: * **[Core Concepts](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/) ** - Understand the data model and evaluation flow * **[Native Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/built-in/) ** - Learn about all available evaluators * **[Custom Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/custom/) ** - Write your own evaluation logic * **[Dataset Management](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/) ** - Save, load, and generate datasets * **[Examples](https://pydantic.dev/docs/ai/evals/examples/simple-validation/) ** - Practical examples for common scenarios Was this page helpful? Thanks for your feedback! --- # Flight Booking | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/complex-workflows/flight-booking/#_top) Flight Booking ============== Example of a multi-agent flow where one agent delegates work to another, then hands off control to a third agent. Demonstrates: * [agent delegation](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#agent-delegation) * [programmatic agent hand-off](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#programmatic-agent-hand-off) * [usage limits](https://pydantic.dev/docs/ai/core-concepts/agent/#usage-limits) In this scenario, a group of agents work together to find the best flight for a user. The control flow for this example can be summarised as follows: graph TD START --> search_agent("search agent") search_agent --> extraction_agent("extraction agent") extraction_agent --> search_agent search_agent --> human_confirm("human confirm") human_confirm --> search_agent search_agent --> FAILED human_confirm --> find_seat_function("find seat function") find_seat_function --> human_seat_choice("human seat choice") human_seat_choice --> find_seat_agent("find seat agent") find_seat_agent --> find_seat_function find_seat_function --> buy_flights("buy flights") buy_flights --> SUCCESS Running the Example ------------------- [](https://pydantic.dev/docs/ai/examples/complex-workflows/flight-booking/#running-the-example) With [dependencies installed and environment variables set](https://pydantic.dev/docs/ai/examples/setup/#usage) , run: * [pip](https://pydantic.dev/docs/ai/examples/complex-workflows/flight-booking/#tab-panel-24) * [uv](https://pydantic.dev/docs/ai/examples/complex-workflows/flight-booking/#tab-panel-25) Terminal python -m pydantic_ai_examples.flight_booking Terminal uv run -m pydantic_ai_examples.flight_booking Example Code ------------ [](https://pydantic.dev/docs/ai/examples/complex-workflows/flight-booking/#example-code) flight\_booking.py import datetime from dataclasses import dataclass from typing import Literal import logfire from pydantic import BaseModel, Field from rich.prompt import Prompt from pydantic_ai import ( Agent, ModelMessage, ModelRetry, RunContext, RunUsage, UsageLimits, ) # 'if-token-present' means nothing will be sent (and the example will work) if you don't have logfire configured logfire.configure(send_to_logfire='if-token-present') logfire.instrument_pydantic_ai() class FlightDetails(BaseModel): """Details of the most suitable flight.""" flight_number: str price: int origin: str = Field(description='Three-letter airport code') destination: str = Field(description='Three-letter airport code') date: datetime.date class NoFlightFound(BaseModel): """When no valid flight is found.""" @dataclass class Deps: web_page_text: str req_origin: str req_destination: str req_date: datetime.date # This agent is responsible for controlling the flow of the conversation. search_agent = Agent[Deps, FlightDetails | NoFlightFound]( 'openai:gpt-5.2', output_type=FlightDetails | NoFlightFound, deps_type=Deps, retries=4, system_prompt=( 'Your job is to find the cheapest flight for the user on the given date. ' ), ) # This agent is responsible for extracting flight details from web page text. extraction_agent = Agent( 'openai:gpt-5.2', output_type=list[FlightDetails], system_prompt='Extract all the flight details from the given text.', ) @search_agent.tool async def extract_flights(ctx: RunContext[Deps]) -> list[FlightDetails]: """Get details of all flights.""" # we pass the usage to the search agent so requests within this agent are counted result = await extraction_agent.run(ctx.deps.web_page_text, usage=ctx.usage) logfire.info('found {flight_count} flights', flight_count=len(result.output)) return result.output @search_agent.output_validator async def validate_output( ctx: RunContext[Deps], output: FlightDetails | NoFlightFound ) -> FlightDetails | NoFlightFound: """Procedural validation that the flight meets the constraints.""" if isinstance(output, NoFlightFound): return output errors: list[str] = [] if output.origin != ctx.deps.req_origin: errors.append( f'Flight should have origin {ctx.deps.req_origin}, not {output.origin}' ) if output.destination != ctx.deps.req_destination: errors.append( f'Flight should have destination {ctx.deps.req_destination}, not {output.destination}' ) if output.date != ctx.deps.req_date: errors.append(f'Flight should be on {ctx.deps.req_date}, not {output.date}') if errors: raise ModelRetry('\n'.join(errors)) else: return output class SeatPreference(BaseModel): row: int = Field(ge=1, le=30) seat: Literal['A', 'B', 'C', 'D', 'E', 'F'] class Failed(BaseModel): """Unable to extract a seat selection.""" # This agent is responsible for extracting the user's seat selection seat_preference_agent = Agent[object, SeatPreference | Failed]( 'openai:gpt-5.2', output_type=SeatPreference | Failed, system_prompt=( "Extract the user's seat preference. " 'Seats A and F are window seats. ' 'Row 1 is the front row and has extra leg room. ' 'Rows 14, and 20 also have extra leg room. ' ), ) # in reality this would be downloaded from a booking site, # potentially using another agent to navigate the site flights_web_page = """ 1. Flight SFO-AK123 - Price: $350 - Origin: San Francisco International Airport (SFO) - Destination: Ted Stevens Anchorage International Airport (ANC) - Date: January 10, 2025 2. Flight SFO-AK456 - Price: $370 - Origin: San Francisco International Airport (SFO) - Destination: Fairbanks International Airport (FAI) - Date: January 10, 2025 3. Flight SFO-AK789 - Price: $400 - Origin: San Francisco International Airport (SFO) - Destination: Juneau International Airport (JNU) - Date: January 20, 2025 4. Flight NYC-LA101 - Price: $250 - Origin: San Francisco International Airport (SFO) - Destination: Ted Stevens Anchorage International Airport (ANC) - Date: January 10, 2025 5. Flight CHI-MIA202 - Price: $200 - Origin: Chicago O'Hare International Airport (ORD) - Destination: Miami International Airport (MIA) - Date: January 12, 2025 6. Flight BOS-SEA303 - Price: $120 - Origin: Boston Logan International Airport (BOS) - Destination: Ted Stevens Anchorage International Airport (ANC) - Date: January 12, 2025 7. Flight DFW-DEN404 - Price: $150 - Origin: Dallas/Fort Worth International Airport (DFW) - Destination: Denver International Airport (DEN) - Date: January 10, 2025 8. Flight ATL-HOU505 - Price: $180 - Origin: Hartsfield-Jackson Atlanta International Airport (ATL) - Destination: George Bush Intercontinental Airport (IAH) - Date: January 10, 2025 """ # restrict how many requests this app can make to the LLM usage_limits = UsageLimits(request_limit=15) async def main(): deps = Deps( web_page_text=flights_web_page, req_origin='SFO', req_destination='ANC', req_date=datetime.date(2025, 1, 10), ) message_history: list[ModelMessage] | None = None usage: RunUsage = RunUsage() # run the agent until a satisfactory flight is found while True: result = await search_agent.run( f'Find me a flight from {deps.req_origin} to {deps.req_destination} on {deps.req_date}', deps=deps, usage=usage, message_history=message_history, usage_limits=usage_limits, ) if isinstance(result.output, NoFlightFound): print('No flight found') break else: flight = result.output print(f'Flight found: {flight}') answer = Prompt.ask( 'Do you want to buy this flight, or keep searching? (buy/*search)', choices=['buy', 'search', ''], show_choices=False, ) if answer == 'buy': seat = await find_seat(usage) await buy_tickets(flight, seat) break else: message_history = result.all_messages( output_tool_return_content='Please suggest another flight' ) async def find_seat(usage: RunUsage) -> SeatPreference: message_history: list[ModelMessage] | None = None while True: answer = Prompt.ask('What seat would you like?') result = await seat_preference_agent.run( answer, message_history=message_history, usage=usage, usage_limits=usage_limits, ) if isinstance(result.output, SeatPreference): return result.output else: print('Could not understand seat preference. Please try again.') message_history = result.all_messages() async def buy_tickets(flight_details: FlightDetails, seat: SeatPreference): print(f'Purchasing flight {flight_details=!r} {seat=!r}...') if __name__ == '__main__': import asyncio asyncio.run(main()) Was this page helpful? Thanks for your feedback! --- # TwelveLabs Video Agent | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/data-analytics/twelvelabs-video-agent/#_top) TwelveLabs Video Agent ====================== Example of a Pydantic AI agent that understands video using [TwelveLabs](https://twelvelabs.io/) Pegasus. Demonstrates: * [tools](https://pydantic.dev/docs/ai/tools-toolsets/tools/) * [agent dependencies](https://pydantic.dev/docs/ai/core-concepts/dependencies/) * wrapping a third-party multimodal API as a tool In this case the idea is a “video analyst” agent — the user asks questions about a video (given its URL), and the agent uses the `analyze_video` tool to call TwelveLabs Pegasus, a video-understanding model, to answer. The LLM decides _what_ to ask about the video, and Pegasus does the actual video understanding. Running the Example ------------------- [](https://pydantic.dev/docs/ai/examples/data-analytics/twelvelabs-video-agent/#running-the-example) You’ll need a TwelveLabs API key set via `TWELVELABS_API_KEY`. You can grab a free key at [twelvelabs.io](https://twelvelabs.io/) — there’s a generous free tier. The example agent runs on `openai:gpt-5-mini`, so you’ll also need an OpenAI API key set via `OPENAI_API_KEY`. Optionally set `VIDEO_URL` to point the agent at your own publicly-accessible video; otherwise a short public sample clip is used. With [dependencies installed and environment variables set](https://pydantic.dev/docs/ai/examples/setup/#usage) , run: * [pip](https://pydantic.dev/docs/ai/examples/data-analytics/twelvelabs-video-agent/#tab-panel-40) * [uv](https://pydantic.dev/docs/ai/examples/data-analytics/twelvelabs-video-agent/#tab-panel-41) Terminal python -m pydantic_ai_examples.twelvelabs_video_agent Terminal uv run -m pydantic_ai_examples.twelvelabs_video_agent Example Code ------------ [](https://pydantic.dev/docs/ai/examples/data-analytics/twelvelabs-video-agent/#example-code) twelvelabs\_video\_agent.py from __future__ import annotations as _annotations import asyncio import os from dataclasses import dataclass import logfire from twelvelabs import AsyncTwelveLabs from twelvelabs.types import VideoContext_Url from pydantic_ai import Agent, RunContext # 'if-token-present' means nothing will be sent (and the example will work) if you don't have logfire configured logfire.configure(send_to_logfire='if-token-present') logfire.instrument_pydantic_ai() # A public sample video used when the user doesn't provide one. The URL must point at a # video file TwelveLabs can fetch directly; set VIDEO_URL to use your own. DEFAULT_VIDEO_URL = 'https://commondatastorage.googleapis.com/gtv-videos-bucket/sample/ElephantsDream.mp4' @dataclass class Deps: twelvelabs: AsyncTwelveLabs video_url: str video_agent = Agent( 'openai:gpt-5-mini', instructions=( 'You help users understand a video. ' 'Use the `analyze_video` tool to ask the video-understanding model questions, ' 'then answer the user concisely based on what it returns.' ), deps_type=Deps, retries=2, ) @video_agent.tool async def analyze_video(ctx: RunContext[Deps], prompt: str) -> str: """Analyze the video with TwelveLabs Pegasus and return a text answer. Args: ctx: The context. prompt: What to ask about the video, e.g. "Summarize this video" or "What objects appear in the first 10 seconds?". """ response = await ctx.deps.twelvelabs.analyze( model_name='pegasus1.5', video=VideoContext_Url(url=ctx.deps.video_url), prompt=prompt, max_tokens=2048, ) return response.data or '' async def main(): api_key = os.environ.get('TWELVELABS_API_KEY') if not api_key: raise RuntimeError( 'Set TWELVELABS_API_KEY to run this example. ' 'Grab a free key at https://twelvelabs.io.' ) video_url = os.environ.get('VIDEO_URL', DEFAULT_VIDEO_URL) async with AsyncTwelveLabs(api_key=api_key) as client: deps = Deps(twelvelabs=client, video_url=video_url) result = await video_agent.run( 'Give me a one-sentence summary of this video.', deps=deps ) print('Response:', result.output) if __name__ == '__main__': asyncio.run(main()) Was this page helpful? Thanks for your feedback! --- # ext | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic-ai/ext/#_top) ext === LangChainToolset ---------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/ext/#pydantic_ai.ext.langchain.LangChainToolset) **Bases:** [`FunctionToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.FunctionToolset) A toolset that wraps LangChain tools. tool\_from\_langchain --------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/ext/#pydantic_ai.ext.langchain.tool_from_langchain) def tool_from_langchain(langchain_tool: LangChainTool) -> Tool Creates a Pydantic AI tool proxy from a LangChain tool. ### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/ext/#returns) [`Tool`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.Tool) — A Pydantic AI tool that corresponds to the LangChain tool. ### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/ext/#parameters) **`langchain_tool`** : `LangChainTool` [](https://pydantic.dev/docs/ai/api/pydantic-ai/ext/#pydantic_ai.ext.langchain.tool_from_langchain(langchain_tool)) The LangChain tool to wrap. Was this page helpful? Thanks for your feedback! --- # Snowflake Cortex | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/models/snowflake/#_top) Snowflake Cortex ================ Install ------- [](https://pydantic.dev/docs/ai/models/snowflake/#install) Snowflake Cortex rejects [`service_tier`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.service_tier) with an error rather than ignoring it, so leave that setting unset on Snowflake models. To use [`SnowflakeModel`](https://pydantic.dev/docs/ai/api/models/snowflake/#pydantic_ai.models.snowflake.SnowflakeModel) , you need to either install `pydantic-ai`, or install `pydantic-ai-slim` with the `snowflake` optional group: * [pip](https://pydantic.dev/docs/ai/models/snowflake/#tab-panel-132) * [uv](https://pydantic.dev/docs/ai/models/snowflake/#tab-panel-133) Terminal pip install "pydantic-ai-slim[snowflake]" Terminal uv add "pydantic-ai-slim[snowflake]" Configuration ------------- [](https://pydantic.dev/docs/ai/models/snowflake/#configuration) [Snowflake Cortex](https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-rest-api) serves Claude, GPT, Llama, Mistral, DeepSeek, and Snowflake’s own models through a REST API hosted in your Snowflake account, so data never leaves the Snowflake security perimeter. To use it, you need your [Snowflake account identifier](https://docs.snowflake.com/en/user-guide/admin-account-identifier) (e.g. `myorg-myaccount`) and a token: a [programmatic access token](https://docs.snowflake.com/en/user-guide/programmatic-access-tokens) (PAT), OAuth token, or key-pair JWT. The role the request runs as — the role a PAT is restricted to, or otherwise your user’s default role — must have the `SNOWFLAKE.CORTEX_USER` database role, which is [granted to `PUBLIC` by default](https://docs.snowflake.com/en/user-guide/snowflake-cortex/aisql#required-privileges) . For a list of available models, see the [Cortex REST API documentation](https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-rest-api) . [Fine-tuned models](https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-finetuning) can be referenced as `database.schema.model`. Environment variables --------------------- [](https://pydantic.dev/docs/ai/models/snowflake/#environment-variables) Once you have the account identifier and token, you can set them as environment variables: Terminal export SNOWFLAKE_ACCOUNT='myorg-myaccount' export SNOWFLAKE_TOKEN='your-token' You can then use [`SnowflakeModel`](https://pydantic.dev/docs/ai/api/models/snowflake/#pydantic_ai.models.snowflake.SnowflakeModel) by name: from pydantic_ai import Agent agent = Agent('snowflake:claude-sonnet-4-6') ... Or initialise the model directly with just the model name: from pydantic_ai import Agent from pydantic_ai.models.snowflake import SnowflakeModel model = SnowflakeModel('claude-sonnet-4-6') agent = Agent(model) ... `provider` argument ------------------- [](https://pydantic.dev/docs/ai/models/snowflake/#provider-argument) You can provide a custom `Provider` via the `provider` argument: from pydantic_ai import Agent from pydantic_ai.models.snowflake import SnowflakeModel from pydantic_ai.providers.snowflake import SnowflakeProvider model = SnowflakeModel( 'claude-sonnet-4-6', provider=SnowflakeProvider(account='myorg-myaccount', token='your-token'), ) agent = Agent(model) ... You can also customize the [`SnowflakeProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.snowflake.SnowflakeProvider) with a custom `base_url` (e.g. when connecting through [private connectivity](https://docs.snowflake.com/en/user-guide/private-snowflake-service) ) or `httpx.AsyncClient`: from httpx import AsyncClient from pydantic_ai import Agent from pydantic_ai.models.snowflake import SnowflakeModel from pydantic_ai.providers.snowflake import SnowflakeProvider model = SnowflakeModel( 'claude-sonnet-4-6', provider=SnowflakeProvider( base_url='https://myorg-myaccount.privatelink.snowflakecomputing.com/api/v2/cortex/v1', token='your-token', http_client=AsyncClient(timeout=30), ), ) agent = Agent(model) ... Model capabilities ------------------ [](https://pydantic.dev/docs/ai/models/snowflake/#model-capabilities) Cortex only supports tool calling and structured output for OpenAI (`openai-*`) and Claude (`claude-*`) models; for other model families, structured output falls back to [prompted output](https://pydantic.dev/docs/ai/core-concepts/output/#prompted-output) . Thinking -------- [](https://pydantic.dev/docs/ai/models/snowflake/#thinking) To enable thinking on Claude models, use the unified [`thinking`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.thinking) [model setting](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) , or set [`SnowflakeModelSettings.snowflake_reasoning`](https://pydantic.dev/docs/ai/api/models/snowflake/#pydantic_ai.models.snowflake.SnowflakeModelSettings.snowflake_reasoning) directly to control the reasoning token budget: from pydantic_ai import Agent from pydantic_ai.models.snowflake import SnowflakeModel, SnowflakeModelSettings agent = Agent( SnowflakeModel('claude-sonnet-4-6'), model_settings=SnowflakeModelSettings(snowflake_reasoning={'max_tokens': 4096}), ) ... On OpenAI models, use the unified `thinking` setting or [`openai_reasoning_effort`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIChatModelSettings.openai_reasoning_effort) . Was this page helpful? Thanks for your feedback! --- # OpenAI | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/realtime/openai/#_top) OpenAI ====== [`OpenAIRealtimeModel`](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModel) connects an agent to OpenAI’s native speech-to-speech models. Start with the [realtime quickstart](https://pydantic.dev/docs/ai/realtime/overview/#quickstart) or the [text-to-audio example](https://pydantic.dev/docs/ai/examples/realtime/realtime-text-to-audio/) . Setup ----- [](https://pydantic.dev/docs/ai/realtime/openai/#setup) To use OpenAI realtime models, install `pydantic-ai-slim` with the `openai-realtime` optional group, which bundles the `openai` package together with the realtime WebSocket transport: * [pip](https://pydantic.dev/docs/ai/realtime/openai/#tab-panel-174) * [uv](https://pydantic.dev/docs/ai/realtime/openai/#tab-panel-175) Terminal pip install "pydantic-ai-slim[openai-realtime]" Terminal uv add "pydantic-ai-slim[openai-realtime]" Set `OPENAI_API_KEY` as described in the [OpenAI model documentation](https://pydantic.dev/docs/ai/models/openai/#configuration) . Authentication and base URL come from `provider`, mirroring [`OpenAIChatModel`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIChatModel) . The default `provider='openai'` reads the environment; pass an [`OpenAIProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.openai.OpenAIProvider) for a custom key or base URL. The realtime WebSocket opens separately, so a custom provider `httpx` client is not used for it. Sessions run over a server-side WebSocket by default; for browser voice, the browser can exchange media directly over [WebRTC](https://pydantic.dev/docs/ai/realtime/openai/#browser-webrtc) while your backend runs the agent (see [Connecting a frontend](https://pydantic.dev/docs/ai/realtime/deployment/#browser-webrtc-server-sideband) ). Model names ----------- [](https://pydantic.dev/docs/ai/realtime/openai/#model-names) Use the provider’s realtime model ID with [`OpenAIRealtimeModel`](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModel) , for example `gpt-realtime`, `gpt-realtime-2.1`, or `gpt-realtime-2.1-mini`. Model availability and aliases can change; use the [official OpenAI model documentation](https://platform.openai.com/docs/models) as the canonical model list. Settings -------- [](https://pydantic.dev/docs/ai/realtime/openai/#settings) [`OpenAIRealtimeModelSettings`](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModelSettings) — the realtime counterpart of [model run settings](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) — extends the [shared settings](https://pydantic.dev/docs/ai/realtime/overview/#shared-settings) with voice, noise reduction, output speed, exact [turn detection](https://pydantic.dev/docs/ai/realtime/turns/) , and truncation: from pydantic_ai.realtime.openai import ( OpenAIRealtimeModel, OpenAIRealtimeModelSettings, ) settings = OpenAIRealtimeModelSettings( max_tokens=2_000, openai_voice='alloy', turn_detection={'sensitivity': 'high', 'silence_duration_ms': 400}, openai_input_noise_reduction='near_field', openai_output_speed=1.1, openai_turn_detection={'type': 'semantic_vad', 'eagerness': 'high'}, openai_truncation={'type': 'retention_ratio', 'retention_ratio': 0.8}, ) model = OpenAIRealtimeModel('gpt-realtime', settings=settings) `openai_turn_detection` accepts [`ServerVAD`](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.ServerVAD) or [`SemanticVAD`](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.SemanticVAD) and overrides shared [`turn_detection`](https://pydantic.dev/docs/ai/realtime/turns/#automatic-turn-detection) . `openai_truncation` also accepts `'auto'` or `'disabled'`; retention ratio preserves a stable, cacheable prefix as the session grows. `openai_voice` selects the provider voice. OpenAI realtime does not expose `temperature` through Pydantic AI. Input transcription defaults to `'auto'`; set a supported transcription model ID to pin it or `None` to disable it. See [Input transcription](https://pydantic.dev/docs/ai/realtime/audio/#input-transcription) . ### Reasoning [](https://pydantic.dev/docs/ai/realtime/openai/#reasoning) The shared [`thinking`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelSettings.thinking) setting (see [Thinking](https://pydantic.dev/docs/ai/capabilities/thinking/) ) applies to models whose profile reports `supports_thinking`, including the `gpt-realtime-2` family. `True` uses the provider default and an effort string selects a level. `False` omits `reasoning`, because OpenAI realtime does not accept a disabled effort. The GA `gpt-realtime` ignores the setting. Reasoning traces are not surfaced as [`ThinkingPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ThinkingPart) s; the API exposes effort as input only. Browser WebRTC -------------- [](https://pydantic.dev/docs/ai/realtime/openai/#browser-webrtc) For browser voice agents, OpenAI recommends WebRTC: the audio flows browser ↔ OpenAI directly, while your backend attaches a control-plane **sideband** to run the agent. [`AgentRealtime`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentRealtime) exposes two signaling helpers, both resolving and binding the agent’s session configuration (instructions, tools, voice, VAD) server-side: * [`answer_webrtc_offer`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentRealtime.answer_webrtc_offer) — the **secure** path: relay the browser’s SDP offer to `POST /v1/realtime/calls`, returning the SDP answer and a [`WebRTCSession`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.WebRTCSession) to attach a sideband to with [`agent.realtime(model).session(provider_session=…)`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentRealtime.session) . The browser never sees a token. * [`create_client_secret`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentRealtime.create_client_secret) — mint a short-lived [`RealtimeClientSecret`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeClientSecret) (ephemeral token) for a browser that negotiates the WebRTC call itself, when you don’t relay the SDP through your backend. See [Connecting a frontend](https://pydantic.dev/docs/ai/realtime/deployment/#browser-webrtc-server-sideband) for the topology, the secure offer-relay flow, and the sideband trust model, and the [realtime WebRTC example](https://pydantic.dev/docs/ai/examples/realtime/realtime-webrtc/) for a runnable FastAPI and browser app. Feature support and limitations ------------------------------- [](https://pydantic.dev/docs/ai/realtime/openai/#feature-support-and-limitations) | Feature | Support | Notes | | --- | --- | --- | | Audio format | Full feature support | Mono PCM16, 24 kHz input and output | | Text output | Full feature support | Select with `output_modality='text'` | | Image input | Full feature support | [Images](https://pydantic.dev/docs/ai/realtime/audio/#images)
provide context for the next turn | | Manual turns | Full feature support | `turn_detection=False` plus [commit/create verbs](https://pydantic.dev/docs/ai/realtime/turns/#push-to-talk) | | Interruption/truncation | Full feature support | [`interrupt(played_ms=...)`](https://pydantic.dev/docs/ai/realtime/turns/#barge-in)
records the heard cutoff | | Input transcription | Full feature support | [Dedicated model](https://pydantic.dev/docs/ai/realtime/audio/#input-transcription)
; `'auto'` by default | | Native tools | Unsupported | Configure [local fallbacks](https://pydantic.dev/docs/ai/realtime/tools/#native-tools)
for web capabilities | | Usage | Full feature support | Token, audio, and cache breakdowns | | Reconnection | Full feature support | Pydantic AI [replays completed local history](https://pydantic.dev/docs/ai/realtime/lifecycle/#state-restoration)
; in-flight media is lost | See [Audio, images, and transcripts](https://pydantic.dev/docs/ai/realtime/audio/) , [Turns and interruptions](https://pydantic.dev/docs/ai/realtime/turns/) , [Tools](https://pydantic.dev/docs/ai/realtime/tools/) , and [Connection lifecycle](https://pydantic.dev/docs/ai/realtime/lifecycle/) for the provider-agnostic workflows. Gateway ------- [](https://pydantic.dev/docs/ai/realtime/openai/#gateway) To route through the [Pydantic AI Gateway](https://pydantic.dev/docs/ai/overview/gateway/) , use a `gateway/`\-prefixed model string: from pydantic_ai import Agent agent = Agent(instructions='You are a helpful voice assistant.') realtime = agent.realtime('gateway/openai:gpt-realtime') Credentials come from [`gateway_provider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.gateway.gateway_provider) . OpenAI-compatible endpoints that expose the realtime protocol can also be supplied through an `OpenAIProvider`. See [Gateway trace propagation](https://pydantic.dev/docs/ai/realtime/observability/#gateway-trace-propagation) . Provider-specific quirks ------------------------ [](https://pydantic.dev/docs/ai/realtime/openai/#provider-specific-quirks) * The provider connection has no resumable server handle. Automatic reconnect restores completed history by [replaying local messages](https://pydantic.dev/docs/ai/realtime/lifecycle/#state-restoration) into a new session. Was this page helpful? Thanks for your feedback! --- # xAI | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/realtime/xai/#_top) xAI === [`XaiRealtimeModel`](https://pydantic.dev/docs/ai/api/realtime/xai/#pydantic_ai.realtime.xai.XaiRealtimeModel) brings Grok Voice into the typed, server-side realtime agent loop. Start with the [realtime quickstart](https://pydantic.dev/docs/ai/realtime/overview/#quickstart) or the [text-to-audio example](https://pydantic.dev/docs/ai/examples/realtime/realtime-text-to-audio/) . Setup ----- [](https://pydantic.dev/docs/ai/realtime/xai/#setup) To use Grok Voice, install `pydantic-ai-slim` with the `xai-realtime` optional group. Alongside `xai-sdk`, the bundle includes the `openai` package, because Grok Voice’s realtime API reuses the OpenAI Realtime protocol’s event types: * [pip](https://pydantic.dev/docs/ai/realtime/xai/#tab-panel-178) * [uv](https://pydantic.dev/docs/ai/realtime/xai/#tab-panel-179) Terminal pip install "pydantic-ai-slim[xai-realtime]" Terminal uv add "pydantic-ai-slim[xai-realtime]" Set `XAI_API_KEY` as described in the [xAI model documentation](https://pydantic.dev/docs/ai/models/xai/#configuration) . Use `provider='xai'` or pass an [`XaiProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.xai.XaiProvider) with `api_key=`. Custom `api_host` is unsupported, and a provider constructed with only `xai_client=` cannot open the WebSocket because the connection requires the API key. Model names ----------- [](https://pydantic.dev/docs/ai/realtime/xai/#model-names) Use a Grok Voice ID such as `grok-voice-latest` or a pinned `grok-voice-think-*` model. `grok-voice-latest` follows xAI’s current flagship and can change underneath an application; pin a version when behavior must remain stable. Use the [official xAI voice documentation](https://docs.x.ai/docs/guides/voice-agent) for the canonical model list. Settings -------- [](https://pydantic.dev/docs/ai/realtime/xai/#settings) [`XaiRealtimeModelSettings`](https://pydantic.dev/docs/ai/api/realtime/xai/#pydantic_ai.realtime.xai.XaiRealtimeModelSettings) — the realtime counterpart of [model run settings](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) — extends the [shared settings](https://pydantic.dev/docs/ai/realtime/overview/#shared-settings) : from pydantic_ai.realtime.xai import XaiRealtimeModel, XaiRealtimeModelSettings settings = XaiRealtimeModelSettings( xai_voice='eve', turn_detection={'sensitivity': 'low'}, input_transcription_model='auto', ) model = XaiRealtimeModel('grok-voice-latest', settings=settings) `xai_voice` selects the provider voice; when unset, xAI picks its own server-side default (currently `eve`). For exact server-VAD threshold or automatic-response behavior, set `xai_turn_detection=` with [`ServerVAD`](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.ServerVAD) ; it fully overrides shared [`turn_detection`](https://pydantic.dev/docs/ai/realtime/turns/#automatic-turn-detection) . Set `turn_detection=False` for [push-to-talk](https://pydantic.dev/docs/ai/realtime/turns/#push-to-talk) . [Input transcription](https://pydantic.dev/docs/ai/realtime/audio/#input-transcription) defaults to `'auto'`. Unlike the incremental deltas described in [live captions](https://pydantic.dev/docs/ai/realtime/audio/#live-captions) , xAI sends cumulative transcript snapshots that can revise earlier words, so caption UIs should render the full [`TranscriptUpdate.transcript`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.TranscriptUpdate.transcript) rather than append deltas. ### Reasoning [](https://pydantic.dev/docs/ai/realtime/xai/#reasoning) `grok-voice-latest` and `grok-voice-think-*` models support the shared [`thinking`](https://pydantic.dev/docs/ai/capabilities/thinking/) setting. The provider exposes only `'high'` and `'none'`: every enabled effort maps to `'high'`, while `False` maps to `'none'`. Other Grok Voice models ignore the setting. Feature support and limitations ------------------------------- [](https://pydantic.dev/docs/ai/realtime/xai/#feature-support-and-limitations) | Feature | Support | Notes | | --- | --- | --- | | Audio format | Full feature support | Mono PCM16, 24 kHz input and output | | Text output | Unsupported | Grok Voice always produces audio | | Image input | Unsupported | Audio/text input only | | Manual turns | Full feature support | `turn_detection=False` plus [commit/create verbs](https://pydantic.dev/docs/ai/realtime/turns/#push-to-talk) | | Interruption | Limited parameter support | [`interrupt()`](https://pydantic.dev/docs/ai/realtime/turns/#barge-in)
works; output truncation with `played_ms` does not | | Input transcription | Full feature support | [Dedicated provider path](https://pydantic.dev/docs/ai/realtime/audio/#input-transcription)
; `'auto'` by default | | Native tools | Unsupported | Configure [local fallbacks](https://pydantic.dev/docs/ai/realtime/tools/#native-tools)
for web capabilities | | Usage | Full feature support | Audio-token buckets and `billable_audio_seconds` in `RunUsage.details` | | State-restoring reconnect | Full feature support | Native [resumption](https://pydantic.dev/docs/ai/realtime/xai/#session-resumption)
is automatic with a reconnect policy | See [Audio, images, and transcripts](https://pydantic.dev/docs/ai/realtime/audio/) , [Turns and interruptions](https://pydantic.dev/docs/ai/realtime/turns/) , [Tools](https://pydantic.dev/docs/ai/realtime/tools/) , and [Connection lifecycle](https://pydantic.dev/docs/ai/realtime/lifecycle/) for the provider-agnostic workflows. Gateway ------- [](https://pydantic.dev/docs/ai/realtime/xai/#gateway) Grok Voice is not currently available through the [Pydantic AI Gateway](https://pydantic.dev/docs/ai/overview/gateway/) . Connect through `provider='xai'` or an `XaiProvider`. Session resumption ------------------ [](https://pydantic.dev/docs/ai/realtime/xai/#session-resumption) With a [`ReconnectPolicy`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.ReconnectPolicy) , xAI automatically enables native resumption for [state-restoring reconnects](https://pydantic.dev/docs/ai/realtime/lifecycle/#state-restoration) : it restores prior turns and suppresses the provider’s replay burst from the local event stream. The handle stays in memory and cannot resume in another process. Provider-specific quirks ------------------------ [](https://pydantic.dev/docs/ai/realtime/xai/#provider-specific-quirks) * Grok Voice always speaks: its profile reports `supports_text_output=False`, so `output_modality='text'` raises a `UserError` before connecting. Read the answer from the transcript on the `SpeechPart`. * xAI supports cancellation but not output truncation. Flush local playback and call `interrupt()` without `played_ms`. * The protocol resembles OpenAI Realtime, but feature support comes from the xAI model profile; avoid assuming every OpenAI behavior is available. Was this page helpful? Thanks for your feedback! --- # pydantic_graph.basenode | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#_top) pydantic\_graph.basenode ======================== GraphRunContext --------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.GraphRunContext) **Bases:** `Generic[StateT, DepsT]` Context for a graph. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#attributes) #### deps [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.GraphRunContext.deps) Dependencies for the graph. **Type:** `DepsT` #### state [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.GraphRunContext.state) The state of the graph. **Type:** `StateT` BaseNode -------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.BaseNode) **Bases:** `ABC`, `Generic[StateT, DepsT, NodeRunEndT]` Base class for a node. ### Methods [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#methods) #### get\_node\_id [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.BaseNode.get_node_id) `@cached` `@classmethod` def get_node_id(cls) -> str Get the ID of the node. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#returns) [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### run [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.BaseNode.run) `@abstractmethod` `@async` def run( ctx: GraphRunContext[StateT, DepsT], ) -> BaseNode[StateT, DepsT, Any] | End[NodeRunEndT] Run the node. This is an abstract method that must be implemented by subclasses. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#returns-1) `BaseNode`\[`StateT`, `DepsT`, [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] | `End`\[`NodeRunEndT`\] — The next node to run or [`End`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.End) to signal the end of the graph. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#parameters) **`ctx`** : `GraphRunContext`\[`StateT`, `DepsT`\] [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.BaseNode.run(ctx)) The graph context. End --- [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.End) **Bases:** `Generic[RunEndT]` Type to return from a node to signal the end of the graph. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#attributes-1) #### data [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.End.data) Data to return from the graph. **Type:** `RunEndT` Edge ---- [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.Edge) Annotation to apply a label to an edge in a graph. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#attributes-2) #### label [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.Edge.label) Label for the edge. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) StateT ------ [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.StateT) Type variable for the state in a graph. **Default:** `TypeVar('StateT', default=object)` DepsT ----- [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.DepsT) Type variable for the dependencies of a graph and node. **Default:** `TypeVar('DepsT', default=object, contravariant=True)` RunEndT ------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.RunEndT) Covariant type variable for the return type of a graph [`run`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.Graph.run) . **Default:** `TypeVar('RunEndT', covariant=True, default=object)` NodeRunEndT ----------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.NodeRunEndT) Covariant type variable for the return type of a node [`run`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.BaseNode.run) . **Default:** `TypeVar('NodeRunEndT', covariant=True, default=Never)` Was this page helpful? Thanks for your feedback! --- # Tools | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/realtime/tools/#_top) Tools ===== [Tools](https://pydantic.dev/docs/ai/tools-toolsets/tools/) registered on an agent are offered to the realtime model and execute on your backend. The session validates arguments, applies retries, runs tools concurrently, returns results to the provider, and records ordinary tool-call messages for later handoff. Capability hooks around tool calls are covered in [Capabilities and hooks](https://pydantic.dev/docs/ai/realtime/capabilities/) . Function tools -------------- [](https://pydantic.dev/docs/ai/realtime/tools/#function-tools) When a model calls a tool, the session emits [`FunctionToolCallEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.FunctionToolCallEvent) , runs the tool, returns the result, and emits [`FunctionToolResultEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.FunctionToolResultEvent) . Parse failures and [`ModelRetry`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelRetry) produce a [`RetryPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.RetryPromptPart) , matching a standard agent run. Other tool exceptions end the session and propagate from iteration. Tool return values reach the model exactly as in a [standard run](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#advanced-tool-returns) : the model receives the string rendering of the return value — plus, where the provider supports it, multimodal content attached via [`ToolReturn`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolReturn) ’s `content` — while local history keeps the full structured [`ToolReturnPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolReturnPart) with its `return_value`, `content`, and `metadata`. Attached content is delivered for real or refused loudly — never silently degraded: OpenAI and Azure OpenAI deliver text and images as a follow-up user message; Gemini Live’s tool results are JSON-only, so text is folded into the result and any binary attachment raises [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) ([#7362](https://github.com/pydantic/pydantic-ai/issues/7362) ); media a provider can’t carry (audio and documents everywhere; images also on xAI) likewise raises before anything is sent. If the provider cancels an in-flight call, Pydantic AI cancels the task and records a synthetic cancellation result locally without sending that result back to the provider. ### Concurrent tool execution [](https://pydantic.dev/docs/ai/realtime/tools/#concurrent-tool-execution) Every tool runs in the background, so a slow tool does not block session events, other tools, or turn tracking. [`all_messages()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.all_messages) keeps each result adjacent to its call even when calls finish out of order. Whether the model continues speaking while it waits is provider-specific. Inspect the [`supports_async_tool_calls`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelProfile.supports_async_tool_calls) profile flag. OpenAI and Azure models generally fill the gap; Gemini pauses unless the [`google_async_tool_calls`](https://pydantic.dev/docs/ai/realtime/gemini/#asynchronous-tool-calls) setting — which declares the tools `NON_BLOCKING` to the Live API — is enabled on a supported model. Native tools ------------ [](https://pydantic.dev/docs/ai/realtime/tools/#native-tools) Provider-native tools execute server-side. Add them through high-level capabilities such as [`WebSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebSearch) and [`WebFetch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebFetch) , or through [`NativeTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.NativeTool) . Each model’s [`supported_native_tools`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelProfile.supported_native_tools) profile is the source of truth. from pydantic_ai import Agent from pydantic_ai.capabilities import WebSearch from pydantic_ai.messages import NativeToolReturnPart, PartEndEvent agent = Agent(instructions='Answer questions, searching the web when useful.') async def main(): async with agent.realtime( 'google:gemini-2.5-flash-native-audio-latest', capabilities=[WebSearch()], ).session() as session: await session.send("What's the latest Pydantic AI release?") async for event in session: if isinstance(event, PartEndEvent) and isinstance(event.part, NativeToolReturnPart): print(event.part.content) An unsupported native tool with a configured local fallback is replaced before connection. Without a fallback, opening the session raises [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) . Provider and model-specific combinations—including Gemini grounding, URL context, and function-tool restrictions—are canonical on the [Gemini provider page](https://pydantic.dev/docs/ai/realtime/gemini/#native-tools) . Deferred and approval-required tools ------------------------------------ [](https://pydantic.dev/docs/ai/realtime/tools/#deferred-and-approval-required-tools) **Approval-gated tools need a [`HandleDeferredToolCalls`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.HandleDeferredToolCalls) handler; without one the call is refused every time.** A standard run can end with a [`DeferredToolRequests`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.DeferredToolRequests) output and resume once a human answers (see [Deferred Tools](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/) ), but a live conversation has nowhere to pause: with no handler, the model is told the tool cannot complete during a realtime session, and the tool never runs. The handler resolves each call inline: approve it (the tool then runs and returns normally), deny it (recorded with `outcome='denied'`), substitute a result, or request a retry. This handler approves small refunds from policy and denies the rest: from pydantic_ai import Agent, DeferredToolRequests, DeferredToolResults, ToolDenied from pydantic_ai.capabilities import HandleDeferredToolCalls from pydantic_ai.tools import RunContext agent = Agent(instructions='You are a customer support voice assistant.') @agent.tool_plain(requires_approval=True) def issue_refund(order_id: str, amount: float) -> str: return f'Refunded ${amount:.2f} for order {order_id}.' async def refund_policy( ctx: RunContext[None], requests: DeferredToolRequests ) -> DeferredToolResults: results = DeferredToolResults() for call in requests.approvals: if call.args_as_dict().get('amount', 0) <= 100: results.approvals[call.tool_call_id] = True else: results.approvals[call.tool_call_id] = ToolDenied( 'Refunds over $100 need a human; offer to connect one.' ) return results async def main(): async with agent.realtime( 'openai:gpt-realtime', capabilities=[HandleDeferredToolCalls(handler=refund_policy)], ).session(): ... This applies to both ways a call is deferred — raising [`ApprovalRequired`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ApprovalRequired) or [`CallDeferred`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.CallDeferred) from the tool, and declaring it up front with [`requires_approval=True`](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/#human-in-the-loop-tool-approval) or an [external toolset](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/#external-toolset) . An approval-gated tool is still advertised to the model, exactly as in a standard run; calling it opens the approval flow rather than running the tool. Asking a human mid-call and resuming on their answer is not supported yet: a realtime session cannot pause and return a `DeferredToolRequests` output for an out-of-band result. Resolve the request during the call, or move that workflow to a standard agent run. [`DeferredToolRequestsEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.DeferredToolRequestsEvent) on a session is informational for the same reason: it is emitted when the handler _has_ resolved the calls, so a consumer can observe what was asked and decided. It is not a hook to respond to — unlike the same event in a standard run, nothing waits for the consumer, and no event is emitted when no handler is installed and the call is refused. Tools registered with `defer_loading=True` are rejected in a realtime session for a related reason; see [Deferred capability loading](https://pydantic.dev/docs/ai/realtime/capabilities/#deferred-capability-loading) . Enqueuing prompts from tools ---------------------------- [](https://pydantic.dev/docs/ai/realtime/tools/#enqueuing-prompts-from-tools) [`RunContext.enqueue()`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.enqueue) — the same mechanism as [injecting follow-up messages from a tool](https://pydantic.dev/docs/ai/tools-toolsets/tools/#injecting-follow-up-messages-from-a-tool) in a standard run — accepts one plain-text prompt per call from a realtime tool. The default `priority='asap'` sends it when no response is active; `priority='when_idle'` waits until the provider reports its current response complete. Neither priority interrupts assistant speech. Delivered prompts become ordinary user turns in history, as in [injecting messages mid-run](https://pydantic.dev/docs/ai/core-concepts/message-history/#injecting-messages-mid-run) . Multimodal content and prebuilt message/part sequences are rejected because the realtime live-input channel cannot preserve their standard-run semantics. Delegating work during a call ----------------------------- [](https://pydantic.dev/docs/ai/realtime/tools/#delegating-work-during-a-call) Realtime models do not provide structured output and can be weaker at complex reasoning than a frontier text model. Expose a tool that delegates the hard work to a standard [`Agent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent) with an `output_type`: from pydantic import BaseModel from pydantic_ai import Agent from pydantic_ai.realtime import RealtimeTurnCompleteEvent class Answer(BaseModel): summary: str confidence: float supervisor = Agent('openai:gpt-5', output_type=Answer) voice = Agent(instructions='Answer using the `consult` tool, then read the summary aloud.') @voice.tool_plain async def consult(question: str) -> str: result = await supervisor.run(question) return result.output.summary async def main(): async with voice.realtime('openai:gpt-realtime').session() as session: await session.send( 'Which of our three shipping options is cheapest for a 4 kg parcel to Berlin?' ) async for event in session: if isinstance(event, RealtimeTurnCompleteEvent): break The delegated run executes concurrently, so providers with asynchronous tool calls can keep talking while analysis runs. To continue the entire conversation after the voice session, see [History and handoff](https://pydantic.dev/docs/ai/realtime/history/#handing-off-to-a-text-agent) . Edge cases ---------- [](https://pydantic.dev/docs/ai/realtime/tools/#edge-cases) * A tool finishing does not necessarily finish the turn; see the [turn boundary](https://pydantic.dev/docs/ai/realtime/events/#the-turn-boundary) . * Short tools can make asynchronous Gemini tool calling counterproductive: the result may interrupt a reply that barely started. Enable it for tools whose latency would otherwise create dead air. * Native-tool behavior is model-specific. Check the profile and provider page rather than assuming every model from a provider supports the same tools. Was this page helpful? Thanks for your feedback! --- # Process History | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/process-history/#_top) Process History =============== [`ProcessHistory`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ProcessHistory) is a [capability](https://pydantic.dev/docs/ai/capabilities/overview/) that wraps a [history processor](https://pydantic.dev/docs/ai/core-concepts/message-history/#processing-message-history) : a function that receives the message history before each model request and returns the (possibly modified) list of messages to send. Use it to trim old turns, redact sensitive content, or summarize long conversations: process\_history.py from pydantic_ai import Agent from pydantic_ai.capabilities import ProcessHistory from pydantic_ai.messages import ModelMessage def keep_recent(messages: list[ModelMessage]) -> list[ModelMessage]: return messages[-5:] # (1) agent = Agent('openai:gpt-5.2', capabilities=[ProcessHistory(keep_recent)]) Keep only the five most recent messages. In practice you'll want to keep the first request too, so the system prompt survives — see [Processing Message History](https://pydantic.dev/docs/ai/core-concepts/message-history/#processing-message-history) for complete patterns. The processor may be sync or async, and may optionally take a [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) as its first argument to access dependencies and run state. Multiple `ProcessHistory` capabilities apply in registration order. Note that the processed messages _replace_ the run’s message history, so make a copy first if you need to keep the original. `ProcessHistory` is a thin wrapper around the [`before_model_request`](https://pydantic.dev/docs/ai/core-concepts/hooks/) lifecycle hook — hook that event directly for richer control, like short-circuiting the model call. See [Processing Message History](https://pydantic.dev/docs/ai/core-concepts/message-history/#processing-message-history) for the full guide, including summarization examples and interactions with [`new_messages()`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.new_messages) . Was this page helpful? Thanks for your feedback! --- # Select Model | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/select-model/#_top) Select Model ============ [`SelectModel`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.SelectModel) is a [capability](https://pydantic.dev/docs/ai/capabilities/overview/) that chooses a model from run dependencies, message history, usage, or the current step. The selector is first evaluated during run setup, so the agent does not need a constructor model: adaptive\_model.py from dataclasses import dataclass from typing import Literal from pydantic_ai import Agent, ModelSelectionContext from pydantic_ai.capabilities import SelectModel @dataclass class Deps: """Dependencies that influence model selection.""" task_complexity: Literal['standard', 'complex'] def select_model(ctx: ModelSelectionContext[Deps]) -> str: """Use the larger model for complex tasks.""" return 'openai:gpt-5.6-sol' if ctx.deps.task_complexity == 'complex' else 'openai:gpt-5.6-luna' agent = Agent(deps_type=Deps, capabilities=[SelectModel(select_model)]) `SelectModel` always receives a callable, which is evaluated before each new logical model request step. The callable may be synchronous or asynchronous. When it returns the same model ID on multiple steps, the resolved model/provider instance is reused for the rest of that run. Provider-side continuation polling within the same step remains pinned to the selected model. See [Selecting the model](https://pydantic.dev/docs/ai/capabilities/custom/#selecting-the-model) to implement the hook in a custom capability and for precedence and lifecycle details. Was this page helpful? Thanks for your feedback! --- # mistral | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/mistral/#_top) mistral ======= Setup ----- [](https://pydantic.dev/docs/ai/api/models/mistral/#setup) For details on how to set up authentication with this model, see [model configuration for Mistral](https://pydantic.dev/docs/ai/models/mistral/) . MistralModel ------------ [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralModel) **Bases:** `Model[Mistral]` A model that uses Mistral. Internally, this uses the [Mistral Python client](https://github.com/mistralai/client-python) to interact with the API. [API Documentation](https://docs.mistral.ai/) ### Attributes [](https://pydantic.dev/docs/ai/api/models/mistral/#attributes) #### model\_name [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralModel.model_name) The model name. **Type:** `MistralModelName` #### system [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralModel.system) The model provider. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) ### Methods [](https://pydantic.dev/docs/ai/api/models/mistral/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralModel.__init__) def __init__( model_name: MistralModelName, *, provider: Literal['mistral'] | Provider[Mistral] = 'mistral', profile: ModelProfileSpec | None = None, json_mode_schema_prompt: str = 'Answer in JSON Object, respect the format:\n```\n{schema}\n```\n', settings: ModelSettings | None = None, ) Initialize a Mistral model. ##### Parameters [](https://pydantic.dev/docs/ai/api/models/mistral/#parameters) **`model_name`** : `MistralModelName` [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralModel.__init__(model_name)) The name of the model to use. **`provider`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘mistral’\] | `Provider`\[`Mistral`\] _Default:_ `'mistral'` [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralModel.__init__(provider)) The provider to use for authentication and API access. Can be either the string ‘mistral’ or an instance of `Provider[Mistral]`. If not provided, a new provider will be created using the other parameters. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralModel.__init__(profile)) The model profile to use. Defaults to a profile picked by the provider based on the model name. **`json_mode_schema_prompt`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) _Default:_ `'Answer in JSON Object, respect the format:\n```\n{schema}\n```\n'` [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralModel.__init__(json_mode_schema_prompt)) The prompt to show when the model expects a JSON object as input. **`settings`** : [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralModel.__init__(settings)) Model-specific settings that will be used as defaults for this model. #### request [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralModel.request) `@async` def request( messages: list[ModelMessage], model_settings: ModelSettings | None, model_request_parameters: ModelRequestParameters, ) -> ModelResponse Make a non-streaming request to the model from Pydantic AI call. ##### Returns [](https://pydantic.dev/docs/ai/api/models/mistral/#returns) [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) #### request\_stream [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralModel.request_stream) `@async` def request_stream( messages: list[ModelMessage], model_settings: ModelSettings | None, model_request_parameters: ModelRequestParameters, run_context: RunContext[Any] | None = None, ) -> AsyncGenerator[StreamedResponse] Make a streaming request to the model from Pydantic AI call. ##### Returns [](https://pydantic.dev/docs/ai/api/models/mistral/#returns-1) [`AsyncGenerator`](https://docs.python.org/3/library/typing.html#typing.AsyncGenerator) \[`StreamedResponse`\] MistralModelSettings -------------------- [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralModelSettings) **Bases:** [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) Settings used for a Mistral model request. ### Attributes [](https://pydantic.dev/docs/ai/api/models/mistral/#attributes-1) #### mistral\_prompt\_cache\_key [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralModelSettings.mistral_prompt_cache_key) Used by Mistral to improve cache hit rates for similar requests, mirroring `openai_prompt_cache_key`. See the [Mistral prompt caching documentation](https://docs.mistral.ai/studio-api/conversations/advanced/prompt-caching) for more information. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) MistralStreamedResponse ----------------------- [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralStreamedResponse) **Bases:** `StreamedResponse` Implementation of `StreamedResponse` for Mistral models. ### Attributes [](https://pydantic.dev/docs/ai/api/models/mistral/#attributes-2) #### model\_name [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralStreamedResponse.model_name) Get the model name of the response. **Type:** `MistralModelName` #### provider\_name [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralStreamedResponse.provider_name) Get the provider name. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### provider\_url [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralStreamedResponse.provider_url) Get the provider base URL. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### timestamp [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralStreamedResponse.timestamp) Get the timestamp of the response. **Type:** [`datetime`](https://docs.python.org/3/library/datetime.html#module-datetime) LatestMistralModelNames ----------------------- [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.LatestMistralModelNames) Latest Mistral models. **Default:** `Literal['mistral-large-latest', 'mistral-small-latest', 'codestral-latest', 'mistral-moderation-latest']` MistralModelName ---------------- [](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralModelName) Possible Mistral model names. Since Mistral supports a variety of date-stamped models, we explicitly list the most popular models but allow any name in the type hints. Since [the Mistral docs](https://docs.mistral.ai/getting-started/models/models_overview/) for a full list. **Default:** `str | LatestMistralModelNames` Was this page helpful? Thanks for your feedback! --- # Stream Markdown | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/streaming/stream-markdown/#_top) Stream Markdown =============== This example shows how to stream markdown from an agent, using the [`rich`](https://github.com/Textualize/rich) library to highlight the output in the terminal. It’ll run the example with both OpenAI and Google Gemini models if the required environment variables are set. Demonstrates: * [streaming text responses](https://pydantic.dev/docs/ai/core-concepts/output/#streaming-text) Running the Example ------------------- [](https://pydantic.dev/docs/ai/examples/streaming/stream-markdown/#running-the-example) With [dependencies installed and environment variables set](https://pydantic.dev/docs/ai/examples/setup/#usage) , run: * [pip](https://pydantic.dev/docs/ai/examples/streaming/stream-markdown/#tab-panel-66) * [uv](https://pydantic.dev/docs/ai/examples/streaming/stream-markdown/#tab-panel-67) Terminal python -m pydantic_ai_examples.stream_markdown Terminal uv run -m pydantic_ai_examples.stream_markdown Example Code ------------ [](https://pydantic.dev/docs/ai/examples/streaming/stream-markdown/#example-code) stream\_markdown.py import asyncio import os import logfire from rich.console import Console, ConsoleOptions, RenderResult from rich.live import Live from rich.markdown import CodeBlock, Markdown from rich.syntax import Syntax from rich.text import Text from pydantic_ai import Agent from pydantic_ai.models import KnownModelName # 'if-token-present' means nothing will be sent (and the example will work) if you don't have logfire configured logfire.configure(send_to_logfire='if-token-present') logfire.instrument_pydantic_ai() agent = Agent() # models to try, and the appropriate env var models: list[tuple[KnownModelName, str]] = [\ ('google:gemini-3-flash-preview', 'GEMINI_API_KEY'),\ ('openai:gpt-5-mini', 'OPENAI_API_KEY'),\ ('groq:llama-3.3-70b-versatile', 'GROQ_API_KEY'),\ ] async def main(): prettier_code_blocks() console = Console() prompt = 'Show me a short example of using Pydantic.' console.log(f'Asking: {prompt}...', style='cyan') for model, env_var in models: if env_var in os.environ: console.log(f'Using model: {model}') with Live('', console=console, vertical_overflow='visible') as live: async with agent.run_stream(prompt, model=model) as result: async for message in result.stream_output(): live.update(Markdown(message)) console.log(result.usage) else: console.log(f'{model} requires {env_var} to be set.') def prettier_code_blocks(): """Make rich code blocks prettier and easier to copy. From https://github.com/samuelcolvin/aicli/blob/v0.8.0/samuelcolvin_aicli.py#L22 """ class SimpleCodeBlock(CodeBlock): def __rich_console__( self, console: Console, options: ConsoleOptions ) -> RenderResult: code = str(self.text).rstrip() yield Text(self.lexer_name, style='dim') yield Syntax( code, self.lexer_name, theme=self.theme, background_color='default', word_wrap=True, ) yield Text(f'/{self.lexer_name}', style='dim') Markdown.elements['fence'] = SimpleCodeBlock if __name__ == '__main__': asyncio.run(main()) Was this page helpful? Thanks for your feedback! --- # Steps | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/graph/builder/steps/#_top) Steps ===== Steps are the fundamental units of work in a graph. They’re async functions that receive a [`StepContext`](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StepContext) and return a value. Creating Steps -------------- [](https://pydantic.dev/docs/ai/graph/builder/steps/#creating-steps) Steps are created using the [`@g.step`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder.step) decorator on the [`GraphBuilder`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder) : basic\_step.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext @dataclass class MyState: counter: int = 0 g = GraphBuilder(state_type=MyState, output_type=int) @g.step async def increment(ctx: StepContext[MyState, None, None]) -> int: ctx.state.counter += 1 return ctx.state.counter g.add( g.edge_from(g.start_node).to(increment), g.edge_from(increment).to(g.end_node), ) graph = g.build() async def main(): state = MyState() result = await graph.run(state=state) print(result) #> 1 _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Step Context ------------ [](https://pydantic.dev/docs/ai/graph/builder/steps/#step-context) Every step function receives a [`StepContext`](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StepContext) as its first parameter. The context provides access to: * `ctx.state` - The mutable graph state (type: `StateT`) * `ctx.deps` - Injected dependencies (type: `DepsT`) * `ctx.inputs` - Input data for this step (type: `InputT`) ### Accessing State [](https://pydantic.dev/docs/ai/graph/builder/steps/#accessing-state) State is shared across all steps in a graph and can be freely mutated: state\_access.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext @dataclass class AppState: messages: list[str] async def main(): g = GraphBuilder(state_type=AppState, output_type=list[str]) @g.step async def add_hello(ctx: StepContext[AppState, None, None]) -> None: ctx.state.messages.append('Hello') @g.step async def add_world(ctx: StepContext[AppState, None, None]) -> None: ctx.state.messages.append('World') @g.step async def get_messages(ctx: StepContext[AppState, None, None]) -> list[str]: return ctx.state.messages g.add( g.edge_from(g.start_node).to(add_hello), g.edge_from(add_hello).to(add_world), g.edge_from(add_world).to(get_messages), g.edge_from(get_messages).to(g.end_node), ) graph = g.build() state = AppState(messages=[]) result = await graph.run(state=state) print(result) #> ['Hello', 'World'] _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ ### Working with Inputs [](https://pydantic.dev/docs/ai/graph/builder/steps/#working-with-inputs) Steps can receive and transform input data: step\_inputs.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext @dataclass class SimpleState: pass async def main(): g = GraphBuilder( state_type=SimpleState, input_type=int, output_type=str, ) @g.step async def double_it(ctx: StepContext[SimpleState, None, int]) -> int: """Double the input value.""" return ctx.inputs * 2 @g.step async def stringify(ctx: StepContext[SimpleState, None, int]) -> str: """Convert to a formatted string.""" return f'Result: {ctx.inputs}' g.add( g.edge_from(g.start_node).to(double_it), g.edge_from(double_it).to(stringify), g.edge_from(stringify).to(g.end_node), ) graph = g.build() result = await graph.run(state=SimpleState(), inputs=21) print(result) #> Result: 42 _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Dependency Injection -------------------- [](https://pydantic.dev/docs/ai/graph/builder/steps/#dependency-injection) Steps can access injected dependencies through `ctx.deps`: dependencies.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext @dataclass class AppState: pass @dataclass class AppDeps: """Dependencies injected into the graph.""" multiplier: int async def main(): g = GraphBuilder( state_type=AppState, deps_type=AppDeps, input_type=int, output_type=int, ) @g.step async def multiply(ctx: StepContext[AppState, AppDeps, int]) -> int: """Multiply input by the injected multiplier.""" return ctx.inputs * ctx.deps.multiplier g.add( g.edge_from(g.start_node).to(multiply), g.edge_from(multiply).to(g.end_node), ) graph = g.build() deps = AppDeps(multiplier=10) result = await graph.run(state=AppState(), deps=deps, inputs=5) print(result) #> 50 _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Customizing Steps ----------------- [](https://pydantic.dev/docs/ai/graph/builder/steps/#customizing-steps) ### Custom Node IDs [](https://pydantic.dev/docs/ai/graph/builder/steps/#custom-node-ids) By default, step node IDs are inferred from the function name. You can override this: custom\_id.py from pydantic_graph import StepContext from basic_step import MyState, g @g.step(node_id='my_custom_id') async def my_step(ctx: StepContext[MyState, None, None]) -> int: return 42 # The node ID is now 'my_custom_id' instead of 'my_step' ### Human-Readable Labels [](https://pydantic.dev/docs/ai/graph/builder/steps/#human-readable-labels) Labels provide documentation for diagram generation: labels.py from pydantic_graph import StepContext from basic_step import MyState, g @g.step(label='Increment the counter') async def increment(ctx: StepContext[MyState, None, None]) -> int: ctx.state.counter += 1 return ctx.state.counter # Access the label programmatically print(increment.label) #> Increment the counter Sequential Steps ---------------- [](https://pydantic.dev/docs/ai/graph/builder/steps/#sequential-steps) Multiple steps can be chained sequentially: sequential.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext @dataclass class MathState: operations: list[str] async def main(): g = GraphBuilder( state_type=MathState, input_type=int, output_type=int, ) @g.step async def add_five(ctx: StepContext[MathState, None, int]) -> int: ctx.state.operations.append('add 5') return ctx.inputs + 5 @g.step async def multiply_by_two(ctx: StepContext[MathState, None, int]) -> int: ctx.state.operations.append('multiply by 2') return ctx.inputs * 2 @g.step async def subtract_three(ctx: StepContext[MathState, None, int]) -> int: ctx.state.operations.append('subtract 3') return ctx.inputs - 3 # Connect steps sequentially g.add( g.edge_from(g.start_node).to(add_five), g.edge_from(add_five).to(multiply_by_two), g.edge_from(multiply_by_two).to(subtract_three), g.edge_from(subtract_three).to(g.end_node), ) graph = g.build() state = MathState(operations=[]) result = await graph.run(state=state, inputs=10) print(f'Result: {result}') #> Result: 27 print(f'Operations: {state.operations}') #> Operations: ['add 5', 'multiply by 2', 'subtract 3'] _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ The computation is: `(10 + 5) * 2 - 3 = 27` Streaming Steps --------------- [](https://pydantic.dev/docs/ai/graph/builder/steps/#streaming-steps) In addition to regular steps that return a single value, you can create streaming steps that yield multiple values over time using the [`@g.stream`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder.stream) decorator: streaming\_step.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, reduce_list_append @dataclass class SimpleState: pass g = GraphBuilder(state_type=SimpleState, output_type=list[int]) @g.stream async def generate_stream(ctx: StepContext[SimpleState, None, None]): """Stream numbers from 1 to 5.""" for i in range(1, 6): yield i @g.step async def square(ctx: StepContext[SimpleState, None, int]) -> int: return ctx.inputs * ctx.inputs collect = g.join(reduce_list_append, initial_factory=list[int]) g.add( g.edge_from(g.start_node).to(generate_stream), # The stream output is an AsyncIterable, so we can map over it g.edge_from(generate_stream).map().to(square), g.edge_from(square).to(collect), g.edge_from(collect).to(g.end_node), ) graph = g.build() async def main(): result = await graph.run(state=SimpleState()) print(sorted(result)) #> [1, 4, 9, 16, 25] _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ ### How Streaming Steps Work [](https://pydantic.dev/docs/ai/graph/builder/steps/#how-streaming-steps-work) Streaming steps return an `AsyncIterable` that yields values over time. When you use `.map()` on a streaming step’s output, the graph processes each yielded value as it becomes available, creating parallel tasks dynamically. This is particularly useful for: * Processing data from APIs that stream responses * Handling real-time data feeds * Progressive processing of large datasets * Any scenario where you want to start processing results before all data is available Like regular steps, streaming steps can also have custom node IDs and labels: labeled\_stream.py from pydantic_graph import StepContext from streaming_step import SimpleState, g @g.stream(node_id='my_stream', label='Generate numbers progressively') async def labeled_stream(ctx: StepContext[SimpleState, None, None]): for i in range(10): yield i Edge Building Convenience Methods --------------------------------- [](https://pydantic.dev/docs/ai/graph/builder/steps/#edge-building-convenience-methods) The builder provides helper methods for common edge patterns: ### Simple Edges with `add_edge()` [](https://pydantic.dev/docs/ai/graph/builder/steps/#simple-edges-with-add_edge) add\_edge\_example.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext @dataclass class SimpleState: pass async def main(): g = GraphBuilder(state_type=SimpleState, output_type=int) @g.step async def step_a(ctx: StepContext[SimpleState, None, None]) -> int: return 10 @g.step async def step_b(ctx: StepContext[SimpleState, None, int]) -> int: return ctx.inputs + 5 # Using add_edge() for simple connections g.add_edge(g.start_node, step_a) g.add_edge(step_a, step_b, label='from a to b') g.add_edge(step_b, g.end_node) graph = g.build() result = await graph.run(state=SimpleState()) print(result) #> 15 _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Type Safety ----------- [](https://pydantic.dev/docs/ai/graph/builder/steps/#type-safety) The graph builder API provides strong type checking through generics. Type parameters on [`StepContext`](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StepContext) ensure: * State access is properly typed * Dependencies are correctly typed * Input/output types match across edges from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext @dataclass class MyState: pass g = GraphBuilder(state_type=MyState, output_type=str) # Type checker will catch mismatches @g.step async def expects_int(ctx: StepContext[MyState, None, int]) -> str: return str(ctx.inputs) @g.step async def returns_str(ctx: StepContext[MyState, None, None]) -> str: return 'hello' # This would be a type error - expects_int needs int input, but returns_str outputs str # g.add(g.edge_from(returns_str).to(expects_int)) # Type error! Next Steps ---------- [](https://pydantic.dev/docs/ai/graph/builder/steps/#next-steps) * Learn about [parallel execution](https://pydantic.dev/docs/ai/graph/builder/parallel/) with broadcasting and mapping * Understand [join nodes](https://pydantic.dev/docs/ai/graph/builder/joins/) for aggregating parallel results * Explore [conditional branching](https://pydantic.dev/docs/ai/graph/builder/decisions/) with decision nodes Was this page helpful? Thanks for your feedback! --- # Kitaru | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/durable_execution/kitaru/#_top) Kitaru ====== [Kitaru](https://docs.zenml.io/kitaru) is a durable execution layer for AI agents. Its Pydantic AI adapter is provided by the `kitaru` package through `kitaru.adapters.pydantic_ai`, rather than by `pydantic_ai.durable_exec`. Durable Execution ----------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/kitaru/#durable-execution) Kitaru records agent progress as **flows** and **checkpoints**. A flow is the durable run you can resume later. A checkpoint is a completed model request, tool call, MCP invocation, or human wait that Kitaru can reuse during recovery. For example, imagine an agent calls a model, gets a useful response, starts a tool call, and then the process crashes. Without durable execution, restarting the program usually repeats the model request and may repeat later side effects too. With Kitaru, the restarted flow can replay the run, reuse the completed checkpoint for the model request, and continue from the first incomplete point. This is useful for long-running agents, human-in-the-loop workflows, and applications where a repeated model request or external API call would cost money, take time, or duplicate a side effect. When you call a `KitaruAgent`, your application still calls the underlying Pydantic AI agent. Kitaru starts or resumes a flow for that call. With the default `"calls"` checkpoint strategy, it records completed model requests, tool calls, MCP invocations, and human waits in Kitaru’s checkpoint storage. On recovery, Kitaru runs the Python function again until it reaches an operation that already has a checkpoint, returns the saved result for that operation, and then continues from the first operation that has not completed. Durable Agent ------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/kitaru/#durable-agent) You can make a normal [`Agent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent) durable by wrapping it with `KitaruAgent` from `kitaru.adapters.pydantic_ai`. Install Kitaru separately from Pydantic AI: Terminal uv add "kitaru[pydantic-ai]" For local development with Kitaru’s local server, install the `local` extra too: Terminal uv add "kitaru[pydantic-ai,local]" Initialize the project and check your connection before running the examples: Terminal kitaru init kitaru login kitaru status Here is the smallest durable Pydantic AI agent using Kitaru: kitaru\_agent.py from pydantic_ai import Agent from kitaru.adapters.pydantic_ai import KitaruAgent agent = Agent('openai:gpt-5-nano', name='researcher') durable_agent = KitaruAgent(agent) result = durable_agent.run_sync('Summarize quantum error correction.') print(result.output) `KitaruAgent` does not replace the original agent object. With the default call-level checkpoint strategy, it delegates to that agent for the actual Pydantic AI run and records recoverable operations while the run is executing: * model requests; * Pydantic AI tool calls; * MCP tool calls; * `@hitl_tool` human waits. It exposes the usual run methods, including `Agent.run` and `Agent.run_sync`. The original [`Agent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent) can still be used normally outside Kitaru. Production Flows ---------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/kitaru/#production-flows) kitaru\_flow.py import kitaru from pydantic_ai import Agent from kitaru.adapters.pydantic_ai import KitaruAgent agent = Agent('openai:gpt-5-nano', name='researcher') durable_agent = KitaruAgent(agent) @kitaru.flow def research_topic(topic: str) -> str: result = durable_agent.run_sync(f'Summarize {topic}.') return result.output Use the short wrapper for first experiments. Use an explicit flow when the run needs to move beyond your local process. Checkpoint Strategy ------------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/kitaru/#checkpoint-strategy) `KitaruAgent` supports two checkpoint strategies: | Strategy | Default? | What gets persisted | Best for | | --- | --- | --- | --- | | `"calls"` | Yes | Replay-safe model requests, tool calls, MCP invocations, and human waits are persisted as separate checkpoints. | Most agents, especially when individual calls are expensive or have side effects. | | `"turn"` | No | One checkpoint wraps the full agent run. | Simpler runs where per-call checkpoints are unnecessary, or cases where streaming constraints require a full-turn checkpoint. | With the default `"calls"` strategy, Kitaru cannot create nested checkpoints inside a user-defined `@kitaru.checkpoint` body. If `durable_agent.run_sync(...)` runs inside a user `@kitaru.checkpoint`, Kitaru records the whole agent turn under that outer checkpoint instead of creating separate model, tool, or MCP checkpoint rows. See the [Kitaru Pydantic AI adapter guide](https://docs.zenml.io/kitaru/adapters/pydantic-ai) for advanced checkpoint configuration. Human-in-the-loop ----------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/kitaru/#human-in-the-loop) For pure human approval or data-entry gates, prefer Kitaru’s `@hitl_tool`. It turns the wait into a durable tool call: the process can stop while waiting for a human response, and recovery can continue after the response is available. from kitaru.adapters.pydantic_ai import hitl_tool @hitl_tool(question='Approve publishing this answer?', schema=bool) def approve_publish(summary: str) -> bool: ... For Pydantic AI’s own deferred tool patterns, see [deferred tools](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/) . Streaming --------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/kitaru/#streaming) Kitaru supports Pydantic AI streaming with some constraints. For event streaming, prefer the [`event_stream_handler`](https://pydantic.dev/docs/ai/core-concepts/agent/#streaming-all-events) argument on `Agent.run`. When a run uses `event_stream_handler`, Kitaru falls back to a turn checkpoint for that call. If you use [`run_stream()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.run_stream) or [`iter()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.iter) , wrap the streaming call in an explicit `@kitaru.checkpoint`. This gives Kitaru one durable operation to replay instead of trying to persist each streamed event separately. Requirements and Constraints ---------------------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/kitaru/#requirements-and-constraints) When using Kitaru with Pydantic AI: * Define the agent with a concrete model at construction time, such as `Agent('openai:gpt-5-nano', ...)`. * Give each durable agent a stable `name`; Kitaru uses it to identify persisted work across runs. * Do not override the model per run with `model=` when using `KitaruAgent`. * Use an explicit `@kitaru.flow` for remote stacks and production services; automatic flow creation is local-only. * Avoid nested Kitaru checkpoints inside user-defined `@kitaru.checkpoint` bodies. Was this page helpful? Thanks for your feedback! --- # Testing | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/guides/testing/#_top) Testing ======= Writing unit tests for Pydantic AI code is just like unit tests for any other Python code. Because for the most part they’re nothing new, we have pretty well established tools and patterns for writing and running these kinds of tests. Unless you’re really sure you know better, you’ll probably want to follow roughly this strategy: * Use [`pytest`](https://docs.pytest.org/en/stable/) as your test harness * If you find yourself typing out long assertions, use [inline-snapshot](https://15r10nk.github.io/inline-snapshot/latest/) * Similarly, [dirty-equals](https://dirty-equals.helpmanual.io/latest/) can be useful for comparing large data structures * Use [`TestModel`](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestModel) or [`FunctionModel`](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel) in place of your actual model to avoid the usage, latency and variability of real LLM calls * Use [`Agent.override`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.override) to replace an agent’s model, dependencies, or toolsets inside your application logic * Set [`ALLOW_MODEL_REQUESTS=False`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ALLOW_MODEL_REQUESTS) globally to block any requests from being made to non-test models accidentally ### Unit testing with `TestModel` [](https://pydantic.dev/docs/ai/guides/testing/#unit-testing-with-testmodel) The simplest and fastest way to exercise most of your application code is using [`TestModel`](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestModel) , this will (by default) call all tools in the agent, then return either plain text or a structured response depending on the return type of the agent. Let’s write unit tests for the following application code: weather\_app.py import asyncio from datetime import date from pydantic_ai import Agent, RunContext from fake_database import DatabaseConn # (1) from weather_service import WeatherService # (2) weather_agent = Agent( 'openai:gpt-5.2', deps_type=WeatherService, instructions='Providing a weather forecast at the locations the user provides.', ) @weather_agent.tool def weather_forecast( ctx: RunContext[WeatherService], location: str, forecast_date: date ) -> str: if forecast_date < date.today(): # (3) return ctx.deps.get_historic_weather(location, forecast_date) else: return ctx.deps.get_forecast(location, forecast_date) async def run_weather_forecast( # (4) user_prompts: list[tuple[str, int]], conn: DatabaseConn ): """Run weather forecast for a list of user prompts and save.""" async with WeatherService() as weather_service: async def run_forecast(prompt: str, user_id: int): result = await weather_agent.run(prompt, deps=weather_service) await conn.store_forecast(user_id, result.output) # run all prompts in parallel await asyncio.gather( *(run_forecast(prompt, user_id) for (prompt, user_id) in user_prompts) ) `DatabaseConn` is a class that holds a database connection `WeatherService` has methods to get weather forecasts and historic data about the weather We need to call a different endpoint depending on whether the date is in the past or the future, you'll see why this nuance is important below This function is the code we want to test, together with the agent it uses Here we have a function that takes a list of `(user_prompt, user_id)` tuples, gets a weather forecast for each prompt, and stores the result in the database. **We want to test this code without having to mock certain objects or modify our code so we can pass test objects in.** Here’s how we would write tests using [`TestModel`](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestModel) : test\_weather\_app.py from datetime import timezone import pytest from dirty_equals import IsNow, IsStr from pydantic_ai import models, capture_run_messages, RequestUsage from pydantic_ai.models.test import TestModel from pydantic_ai import ( ModelResponse, TextPart, ToolCallPart, ToolReturnPart, UserPromptPart, ModelRequest, ) from fake_database import DatabaseConn from weather_app import run_weather_forecast, weather_agent pytestmark = pytest.mark.anyio # (1) models.ALLOW_MODEL_REQUESTS = False # (2) async def test_forecast(): conn = DatabaseConn() user_id = 1 with capture_run_messages() as messages: with weather_agent.override(model=TestModel()): # (3) prompt = 'What will the weather be like in London on 2024-11-28?' await run_weather_forecast([(prompt, user_id)], conn) # (4) forecast = await conn.get_forecast(user_id) assert forecast == '{"weather_forecast":"Sunny with a chance of rain"}' # (5) assert messages == [ # (6)\ ModelRequest(\ parts=[\ UserPromptPart(\ content='What will the weather be like in London on 2024-11-28?',\ timestamp=IsNow(tz=timezone.utc), # (7)\ ),\ ],\ instructions='Providing a weather forecast at the locations the user provides.',\ timestamp=IsNow(tz=timezone.utc),\ run_id=IsStr(),\ conversation_id=IsStr(),\ ),\ ModelResponse(\ parts=[\ ToolCallPart(\ tool_name='weather_forecast',\ args={\ 'location': 'a',\ 'forecast_date': '2024-01-01', # (8)\ },\ tool_call_id=IsStr(),\ )\ ],\ usage=RequestUsage(\ input_tokens=60,\ output_tokens=7,\ ),\ model_name='test',\ timestamp=IsNow(tz=timezone.utc),\ provider_name='test',\ run_id=IsStr(),\ conversation_id=IsStr(),\ ),\ ModelRequest(\ parts=[\ ToolReturnPart(\ tool_name='weather_forecast',\ content='Sunny with a chance of rain',\ tool_call_id=IsStr(),\ timestamp=IsNow(tz=timezone.utc),\ ),\ ],\ instructions='Providing a weather forecast at the locations the user provides.',\ timestamp=IsNow(tz=timezone.utc),\ run_id=IsStr(),\ conversation_id=IsStr(),\ ),\ ModelResponse(\ parts=[\ TextPart(\ content='{"weather_forecast":"Sunny with a chance of rain"}',\ )\ ],\ usage=RequestUsage(\ input_tokens=66,\ output_tokens=16,\ ),\ model_name='test',\ timestamp=IsNow(tz=timezone.utc),\ provider_name='test',\ run_id=IsStr(),\ conversation_id=IsStr(),\ ),\ ] We're using [anyio](https://anyio.readthedocs.io/en/stable/) to run async tests. This is a safety measure to make sure we don't accidentally make real requests to the LLM while testing, see [`ALLOW_MODEL_REQUESTS`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ALLOW_MODEL_REQUESTS) for more details. We're using [`Agent.override`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.override) to replace the agent's model with [`TestModel`](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestModel) , the nice thing about `override` is that we can replace the model inside agent without needing access to the agent `run*` methods call site. Now we call the function we want to test inside the `override` context manager. But default, `TestModel` will return a JSON string summarising the tools calls made, and what was returned. If you wanted to customise the response to something more closely aligned with the domain, you could add [`custom_output_text='Sunny'`](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestModel.custom_output_text) when defining `TestModel`. So far we don't actually know which tools were called and with which values, we can use [`capture_run_messages`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.capture_run_messages) to inspect messages from the most recent run and assert the exchange between the agent and the model occurred as expected. The [`IsNow`](https://dirty-equals.helpmanual.io/latest/types/datetime/#dirty_equals.IsNow) helper allows us to use declarative asserts even with data which will contain timestamps that change over time. `TestModel` isn't doing anything clever to extract values from the prompt, so these values are hardcoded. ### Unit testing with `FunctionModel` [](https://pydantic.dev/docs/ai/guides/testing/#unit-testing-with-functionmodel) The above tests are a great start, but careful readers will notice that the `WeatherService.get_forecast` is never called since `TestModel` calls `weather_forecast` with a date in the past. To fully exercise `weather_forecast`, we need to use [`FunctionModel`](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel) to customise how the tools is called. Here’s an example of using `FunctionModel` to test the `weather_forecast` tool with custom inputs test\_weather\_app2.py import re import pytest from pydantic_ai import models from pydantic_ai import ( ModelMessage, ModelResponse, TextPart, ToolCallPart, ) from pydantic_ai.models.function import AgentInfo, FunctionModel from fake_database import DatabaseConn from weather_app import run_weather_forecast, weather_agent pytestmark = pytest.mark.anyio models.ALLOW_MODEL_REQUESTS = False def call_weather_forecast( # (1) messages: list[ModelMessage], info: AgentInfo ) -> ModelResponse: if len(messages) == 1: # first call, call the weather forecast tool user_prompt = messages[0].parts[-1] m = re.search(r'd{4}-d{2}-d{2}', user_prompt.content) assert m is not None args = {'location': 'London', 'forecast_date': m.group()} # (2) return ModelResponse(parts=[ToolCallPart('weather_forecast', args)]) else: # second call, return the forecast msg = messages[-1].parts[0] assert msg.part_kind == 'tool-return' return ModelResponse(parts=[TextPart(f'The forecast is: {msg.content}')]) async def test_forecast_future(): conn = DatabaseConn() user_id = 1 with weather_agent.override(model=FunctionModel(call_weather_forecast)): # (3) prompt = 'What will the weather be like in London on 2032-01-01?' await run_weather_forecast([(prompt, user_id)], conn) forecast = await conn.get_forecast(user_id) assert forecast == 'The forecast is: Rainy with a chance of sun' We define a function `call_weather_forecast` that will be called by `FunctionModel` in place of the LLM, this function has access to the list of [`ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage) s that make up the run, and [`AgentInfo`](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.AgentInfo) which contains information about the agent and the function tools and return tools. Our function is slightly intelligent in that it tries to extract a date from the prompt, but just hard codes the location. We use [`FunctionModel`](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel) to replace the agent's model with our custom function. ### Overriding model via pytest fixtures [](https://pydantic.dev/docs/ai/guides/testing/#overriding-model-via-pytest-fixtures) If you’re writing lots of tests that all require model to be overridden, you can use [pytest fixtures](https://docs.pytest.org/en/6.2.x/fixture.html) to override the model with [`TestModel`](https://pydantic.dev/docs/ai/api/models/test/#pydantic_ai.models.test.TestModel) or [`FunctionModel`](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel) in a reusable way. Here’s an example of a fixture that overrides the model with `TestModel`: test\_agent.py import pytest from pydantic_ai.models.test import TestModel from weather_app import weather_agent @pytest.fixture def override_weather_agent(): with weather_agent.override(model=TestModel()): yield async def test_forecast(override_weather_agent: None): ... # test code here Was this page helpful? Thanks for your feedback! --- # Hugging Face | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/models/huggingface/#_top) Hugging Face ============ [Hugging Face](https://huggingface.co/) is an AI platform with all major open source models, datasets, MCPs, and demos. You can use [Inference Providers](https://huggingface.co/docs/inference-providers) to run open source models like DeepSeek R1 on scalable serverless infrastructure. Install ------- [](https://pydantic.dev/docs/ai/models/huggingface/#install) To use `HuggingFaceModel`, you need to either install `pydantic-ai`, or install `pydantic-ai-slim` with the `huggingface` optional group: * [pip](https://pydantic.dev/docs/ai/models/huggingface/#tab-panel-118) * [uv](https://pydantic.dev/docs/ai/models/huggingface/#tab-panel-119) Terminal pip install "pydantic-ai-slim[huggingface]" Terminal uv add "pydantic-ai-slim[huggingface]" Configuration ------------- [](https://pydantic.dev/docs/ai/models/huggingface/#configuration) To use [Hugging Face](https://huggingface.co/) inference, you’ll need to set up an account which will give you [free tier](https://huggingface.co/docs/inference-providers/pricing) allowance on [Inference Providers](https://huggingface.co/docs/inference-providers) . To setup inference, follow these steps: 1. Go to [Hugging Face](https://huggingface.co/join) and sign up for an account. 2. Create a new access token in [Hugging Face](https://huggingface.co/settings/tokens) . 3. Set the `HF_TOKEN` environment variable to the token you just created. Once you have a Hugging Face access token, you can set it as an environment variable: Terminal export HF_TOKEN='hf_token' Usage ----- [](https://pydantic.dev/docs/ai/models/huggingface/#usage) You can then use [`HuggingFaceModel`](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceModel) by name: from pydantic_ai import Agent agent = Agent('huggingface:Qwen/Qwen3-235B-A22B') ... Or initialise the model directly with just the model name: from pydantic_ai import Agent from pydantic_ai.models.huggingface import HuggingFaceModel model = HuggingFaceModel('Qwen/Qwen3-235B-A22B') agent = Agent(model) ... By default, the [`HuggingFaceModel`](https://pydantic.dev/docs/ai/api/models/huggingface/#pydantic_ai.models.huggingface.HuggingFaceModel) uses the [`HuggingFaceProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.huggingface.HuggingFaceProvider) that will select automatically the first of the inference providers (Cerebras, Together AI, Cohere..etc) available for the model, sorted by your preferred order in [https://hf.co/settings/inference-providers](https://hf.co/settings/inference-providers) . Configure the provider ---------------------- [](https://pydantic.dev/docs/ai/models/huggingface/#configure-the-provider) If you want to pass parameters in code to the provider, you can programmatically instantiate the [`HuggingFaceProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.huggingface.HuggingFaceProvider) and pass it to the model: from pydantic_ai import Agent from pydantic_ai.models.huggingface import HuggingFaceModel from pydantic_ai.providers.huggingface import HuggingFaceProvider model = HuggingFaceModel('Qwen/Qwen3-235B-A22B', provider=HuggingFaceProvider(api_key='hf_token', provider_name='nebius')) agent = Agent(model) ... Custom Hugging Face client -------------------------- [](https://pydantic.dev/docs/ai/models/huggingface/#custom-hugging-face-client) [`HuggingFaceProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.huggingface.HuggingFaceProvider) also accepts a custom [`AsyncInferenceClient`](https://huggingface.co/docs/huggingface_hub/v0.29.3/en/package_reference/inference_client#huggingface_hub.AsyncInferenceClient) client via the `hf_client` parameter, so you can customise the `headers`, `bill_to` (billing to an HF organization you’re a member of), `base_url` etc. as defined in the [Hugging Face Hub python library docs](https://huggingface.co/docs/huggingface_hub/package_reference/inference_client) . from huggingface_hub import AsyncInferenceClient from pydantic_ai import Agent from pydantic_ai.models.huggingface import HuggingFaceModel from pydantic_ai.providers.huggingface import HuggingFaceProvider client = AsyncInferenceClient( bill_to='openai', api_key='hf_token', provider='fireworks-ai', ) model = HuggingFaceModel( 'Qwen/Qwen3-235B-A22B', provider=HuggingFaceProvider(hf_client=client), ) agent = Agent(model) ... Streaming cancellation ---------------------- [](https://pydantic.dev/docs/ai/models/huggingface/#streaming-cancellation) Was this page helpful? Thanks for your feedback! --- # Mistral | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/models/mistral/#_top) Mistral ======= Install ------- [](https://pydantic.dev/docs/ai/models/mistral/#install) To use `MistralModel`, you need to either install `pydantic-ai`, or install `pydantic-ai-slim` with the `mistral` optional group: * [pip](https://pydantic.dev/docs/ai/models/mistral/#tab-panel-120) * [uv](https://pydantic.dev/docs/ai/models/mistral/#tab-panel-121) Terminal pip install "pydantic-ai-slim[mistral]" Terminal uv add "pydantic-ai-slim[mistral]" Configuration ------------- [](https://pydantic.dev/docs/ai/models/mistral/#configuration) To use [Mistral](https://mistral.ai/) through their API, go to [console.mistral.ai/api-keys/](https://console.mistral.ai/api-keys/) and follow your nose until you find the place to generate an API key. `LatestMistralModelNames` contains a list of the most popular Mistral models. Environment variable -------------------- [](https://pydantic.dev/docs/ai/models/mistral/#environment-variable) Once you have the API key, you can set it as an environment variable: Terminal export MISTRAL_API_KEY='your-api-key' You can then use `MistralModel` by name: from pydantic_ai import Agent agent = Agent('mistral:mistral-large-latest') ... Or initialise the model directly with just the model name: from pydantic_ai import Agent from pydantic_ai.models.mistral import MistralModel model = MistralModel('mistral-small-latest') agent = Agent(model) ... `provider` argument ------------------- [](https://pydantic.dev/docs/ai/models/mistral/#provider-argument) You can provide a custom `Provider` via the `provider` argument: from pydantic_ai import Agent from pydantic_ai.models.mistral import MistralModel from pydantic_ai.providers.mistral import MistralProvider model = MistralModel( 'mistral-large-latest', provider=MistralProvider(api_key='your-api-key', base_url='https://') ) agent = Agent(model) ... You can also customize the provider with a custom `httpx.AsyncClient`: from httpx import AsyncClient from pydantic_ai import Agent from pydantic_ai.models.mistral import MistralModel from pydantic_ai.providers.mistral import MistralProvider custom_http_client = AsyncClient(timeout=30) model = MistralModel( 'mistral-large-latest', provider=MistralProvider(api_key='your-api-key', http_client=custom_http_client), ) agent = Agent(model) ... Was this page helpful? Thanks for your feedback! --- # Concurrency & Performance | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/how-to/concurrency/#_top) Concurrency & Performance ========================= Control how evaluation cases are executed in parallel. By default, Pydantic Evals runs all cases concurrently to maximize throughput. You can control this behavior using the `max_concurrency` parameter. Basic Usage ----------- [](https://pydantic.dev/docs/ai/evals/how-to/concurrency/#basic-usage) from pydantic_evals import Case, Dataset def my_task(inputs: str) -> str: return f'Result: {inputs}' dataset = Dataset(name='concurrency_demo', cases=[Case(inputs='test1'), Case(inputs='test2')]) # Run all cases concurrently (default) report = dataset.evaluate_sync(my_task) # Limit to 5 concurrent cases report = dataset.evaluate_sync(my_task, max_concurrency=5) # Run sequentially (one at a time) report = dataset.evaluate_sync(my_task, max_concurrency=1) When to Limit Concurrency ------------------------- [](https://pydantic.dev/docs/ai/evals/how-to/concurrency/#when-to-limit-concurrency) ### Rate Limiting [](https://pydantic.dev/docs/ai/evals/how-to/concurrency/#rate-limiting) Many APIs have rate limits that restrict concurrent requests: from pydantic_evals import Case, Dataset async def my_llm_task(inputs: str) -> str: return f'LLM Result: {inputs}' dataset = Dataset(name='rate_limit_demo', cases=[Case(inputs='test1')]) # If your API allows 10 requests/second report = dataset.evaluate_sync( my_llm_task, max_concurrency=10, ) ### Resource Constraints [](https://pydantic.dev/docs/ai/evals/how-to/concurrency/#resource-constraints) Limit concurrency to avoid overwhelming system resources: from pydantic_evals import Case, Dataset def heavy_computation(inputs: str) -> str: return f'Heavy: {inputs}' def db_query_task(inputs: str) -> str: return f'DB: {inputs}' dataset = Dataset(name='resource_constraints', cases=[Case(inputs='test1')]) # Memory-intensive operations report = dataset.evaluate_sync( heavy_computation, max_concurrency=2, # Only 2 at a time ) # Database connection pool limits report = dataset.evaluate_sync( db_query_task, max_concurrency=5, # Match connection pool size ) ### Debugging [](https://pydantic.dev/docs/ai/evals/how-to/concurrency/#debugging) Run sequentially to see clear error traces: from pydantic_evals import Case, Dataset def my_task(inputs: str) -> str: return f'Result: {inputs}' dataset = Dataset(name='debug_demo', cases=[Case(inputs='test1')]) # Easier to debug report = dataset.evaluate_sync( my_task, max_concurrency=1, ) Performance Comparison ---------------------- [](https://pydantic.dev/docs/ai/evals/how-to/concurrency/#performance-comparison) Here’s an example showing the performance difference: concurrency\_example.py import asyncio from pydantic_evals import Case, Dataset # Create a dataset with multiple test cases dataset = Dataset( name='performance_comparison', cases=[\ Case(\ name=f'case_{i}',\ inputs=i,\ expected_output=i * 2,\ )\ for i in range(10)\ ] ) async def slow_task(input_value: int) -> int: """Simulates a slow operation (e.g., API call).""" await asyncio.sleep(0.1) # 100ms per case return input_value * 2 # Unlimited concurrency: ~0.1s total (all cases run in parallel) report = dataset.evaluate_sync(slow_task) # Limited concurrency: ~0.5s total (2 at a time, 5 batches) report = dataset.evaluate_sync(slow_task, max_concurrency=2) # Sequential: ~1.0s total (one at a time, 10 cases) report = dataset.evaluate_sync(slow_task, max_concurrency=1) Concurrency with Evaluators --------------------------- [](https://pydantic.dev/docs/ai/evals/how-to/concurrency/#concurrency-with-evaluators) Both task execution and evaluator execution happen concurrently by default: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import LLMJudge def my_task(inputs: str) -> str: return f'Result: {inputs}' dataset = Dataset( name='evaluator_concurrency', cases=[Case(inputs=f'test{i}') for i in range(100)], # 100 cases evaluators=[\ LLMJudge(rubric='Quality check'), # Makes API calls\ ], ) # Both task and evaluator run with controlled concurrency report = dataset.evaluate_sync( my_task, max_concurrency=10, ) If your evaluators are expensive (e.g., [`LLMJudge`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.LLMJudge) ), limiting concurrency helps manage: * API rate limits * Cost (fewer concurrent API calls) * Memory usage Async vs Sync ------------- [](https://pydantic.dev/docs/ai/evals/how-to/concurrency/#async-vs-sync) Both sync and async evaluation support concurrency control: ### Sync API [](https://pydantic.dev/docs/ai/evals/how-to/concurrency/#sync-api) from pydantic_evals import Case, Dataset def my_task(inputs: str) -> str: return f'Result: {inputs}' dataset = Dataset(name='sync_demo', cases=[Case(inputs='test1')]) # Runs async operations internally with controlled concurrency report = dataset.evaluate_sync(my_task, max_concurrency=10) ### Async API [](https://pydantic.dev/docs/ai/evals/how-to/concurrency/#async-api) from pydantic_evals import Case, Dataset async def my_task(inputs: str) -> str: return f'Result: {inputs}' async def run_evaluation(): dataset = Dataset(name='async_demo', cases=[Case(inputs='test1')]) # Same behavior, but in async context report = await dataset.evaluate(my_task, max_concurrency=10) return report Monitoring Concurrency ---------------------- [](https://pydantic.dev/docs/ai/evals/how-to/concurrency/#monitoring-concurrency) Track execution to optimize settings: import time from pydantic_evals import Case, Dataset def task(inputs: str) -> str: return f'Result: {inputs}' dataset = Dataset(name='monitoring', cases=[Case(inputs=f'test{i}') for i in range(10)]) t0 = time.time() report = dataset.evaluate_sync(task, max_concurrency=10) duration = time.time() - t0 num_cases = len(report.cases) + len(report.failures) avg_duration = duration / num_cases print(f'Total: {duration:.2f}s') #> Total: 0.01s print(f'Cases: {num_cases}') #> Cases: 10 print(f'Avg per case: {avg_duration:.2f}s') #> Avg per case: 0.00s print(f'Effective concurrency: ~{num_cases * avg_duration / duration:.1f}') #> Effective concurrency: ~1.0 Handling Rate Limits -------------------- [](https://pydantic.dev/docs/ai/evals/how-to/concurrency/#handling-rate-limits) If you hit rate limits, the evaluation will fail. Use retry strategies: from pydantic_evals import Case, Dataset def task(inputs: str) -> str: return f'Result: {inputs}' dataset = Dataset(name='rate_limit_handling', cases=[Case(inputs='test1')]) # Reduce concurrency to avoid rate limits report = dataset.evaluate_sync( task, max_concurrency=5, # Stay under rate limit ) See [Retry Strategies](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/) for handling transient failures. Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/how-to/concurrency/#next-steps) * **[Retry Strategies](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/) ** - Handle transient failures * **[Dataset Management](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/) ** - Work with large datasets * **[Logfire Integration](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/) ** - Monitor performance Was this page helpful? Thanks for your feedback! --- # RAG | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/data-analytics/rag/#_top) RAG === RAG search example. This demo allows you to ask questions about an October 2024 snapshot of the [Logfire](https://pydantic.dev/logfire) documentation. Demonstrates: * [tools](https://pydantic.dev/docs/ai/tools-toolsets/tools/) * [agent dependencies](https://pydantic.dev/docs/ai/core-concepts/dependencies/) * RAG search This is done by creating a database containing each section of the markdown documentation, then registering the search tool with the Pydantic AI agent. Logic for extracting sections from markdown files and a JSON file with that data is available in [this gist](https://gist.github.com/samuelcolvin/4b5bb9bb163b1122ff17e29e48c10992) . [PostgreSQL with pgvector](https://github.com/pgvector/pgvector) is used as the search database, the easiest way to download and run pgvector is using Docker: Terminal mkdir postgres-data docker run --rm \ -e POSTGRES_PASSWORD=postgres \ -p 54320:5432 \ -v `pwd`/postgres-data:/var/lib/postgresql/data \ pgvector/pgvector:pg17 As with the [SQL gen](https://pydantic.dev/docs/ai/examples/data-analytics/sql-gen/) example, we run postgres on port `54320` to avoid conflicts with any other postgres instances you may have running. We also mount the PostgreSQL `data` directory locally to persist the data if you need to stop and restart the container. With that running and [dependencies installed and environment variables set](https://pydantic.dev/docs/ai/examples/setup/#usage) , we can build the search database with (**WARNING**: this requires the `OPENAI_API_KEY` env variable and will calling the OpenAI embedding API around 300 times to generate embeddings for each section of the documentation): * [pip](https://pydantic.dev/docs/ai/examples/data-analytics/rag/#tab-panel-32) * [uv](https://pydantic.dev/docs/ai/examples/data-analytics/rag/#tab-panel-33) Terminal python -m pydantic_ai_examples.rag build Terminal uv run -m pydantic_ai_examples.rag build (Note building the database doesn’t use Pydantic AI right now, instead it uses the OpenAI SDK directly.) You can then ask the agent a question with: * [pip](https://pydantic.dev/docs/ai/examples/data-analytics/rag/#tab-panel-34) * [uv](https://pydantic.dev/docs/ai/examples/data-analytics/rag/#tab-panel-35) Terminal python -m pydantic_ai_examples.rag search "How do I configure logfire to work with FastAPI?" Terminal uv run -m pydantic_ai_examples.rag search "How do I configure logfire to work with FastAPI?" Example Code ------------ [](https://pydantic.dev/docs/ai/examples/data-analytics/rag/#example-code) rag.py from __future__ import annotations as _annotations import asyncio import re import sys import unicodedata from contextlib import asynccontextmanager from dataclasses import dataclass import asyncpg import httpx import logfire import pydantic_core from anyio import create_task_group from openai import AsyncOpenAI from pydantic import TypeAdapter from typing_extensions import AsyncGenerator from pydantic_ai import Agent, RunContext # 'if-token-present' means nothing will be sent (and the example will work) if you don't have logfire configured logfire.configure(send_to_logfire='if-token-present') logfire.instrument_asyncpg() logfire.instrument_pydantic_ai() @dataclass class Deps: openai: AsyncOpenAI pool: asyncpg.Pool agent = Agent('openai:gpt-5.2', deps_type=Deps) @agent.tool async def retrieve(context: RunContext[Deps], search_query: str) -> str: """Retrieve documentation sections based on a search query. Args: context: The call context. search_query: The search query. """ with logfire.span( 'create embedding for {search_query=}', search_query=search_query ): embedding = await context.deps.openai.embeddings.create( input=search_query, model='text-embedding-3-small', ) assert len(embedding.data) == 1, ( f'Expected 1 embedding, got {len(embedding.data)}, doc query: {search_query!r}' ) embedding = embedding.data[0].embedding embedding_json = pydantic_core.to_json(embedding).decode() rows = await context.deps.pool.fetch( 'SELECT url, title, content FROM doc_sections ORDER BY embedding <-> $1 LIMIT 8', embedding_json, ) return '\n\n'.join( f'# {row["title"]}\nDocumentation URL:{row["url"]}\n\n{row["content"]}\n' for row in rows ) async def run_agent(question: str): """Entry point to run the agent and perform RAG based question answering.""" openai = AsyncOpenAI() logfire.instrument_openai(openai) logfire.info('Asking "{question}"', question=question) async with database_connect(False) as pool: deps = Deps(openai=openai, pool=pool) answer = await agent.run(question, deps=deps) print(answer.output) ####################################################### # The rest of this file is dedicated to preparing the # # search database, and some utilities. # ####################################################### # JSON document from # https://gist.github.com/samuelcolvin/4b5bb9bb163b1122ff17e29e48c10992 DOCS_JSON = ( 'https://gist.githubusercontent.com/' 'samuelcolvin/4b5bb9bb163b1122ff17e29e48c10992/raw/' '80c5925c42f1442c24963aaf5eb1a324d47afe95/logfire_docs.json' ) async def build_search_db(): """Build the search database.""" async with httpx.AsyncClient() as client: response = await client.get(DOCS_JSON) response.raise_for_status() sections = sections_ta.validate_json(response.content) openai = AsyncOpenAI() logfire.instrument_openai(openai) async with database_connect(True) as pool: with logfire.span('create schema'): async with pool.acquire() as conn: async with conn.transaction(): await conn.execute(DB_SCHEMA) sem = asyncio.Semaphore(10) async with create_task_group() as tg: for section in sections: tg.start_soon(insert_doc_section, sem, openai, pool, section) async def insert_doc_section( sem: asyncio.Semaphore, openai: AsyncOpenAI, pool: asyncpg.Pool, section: DocsSection, ) -> None: async with sem: url = section.url() exists = await pool.fetchval('SELECT 1 FROM doc_sections WHERE url = $1', url) if exists: logfire.info('Skipping {url=}', url=url) return with logfire.span('create embedding for {url=}', url=url): embedding = await openai.embeddings.create( input=section.embedding_content(), model='text-embedding-3-small', ) assert len(embedding.data) == 1, ( f'Expected 1 embedding, got {len(embedding.data)}, doc section: {section}' ) embedding = embedding.data[0].embedding embedding_json = pydantic_core.to_json(embedding).decode() await pool.execute( 'INSERT INTO doc_sections (url, title, content, embedding) VALUES ($1, $2, $3, $4)', url, section.title, section.content, embedding_json, ) @dataclass class DocsSection: id: int parent: int | None path: str level: int title: str content: str def url(self) -> str: url_path = re.sub(r'\.md$', '', self.path) return ( f'https://logfire.pydantic.dev/docs/{url_path}/#{slugify(self.title, "-")}' ) def embedding_content(self) -> str: return '\n\n'.join((f'path: {self.path}', f'title: {self.title}', self.content)) sections_ta = TypeAdapter(list[DocsSection]) # pyright: reportUnknownMemberType=false # pyright: reportUnknownVariableType=false @asynccontextmanager async def database_connect( create_db: bool = False, ) -> AsyncGenerator[asyncpg.Pool, None]: server_dsn, database = ( 'postgresql://postgres:postgres@localhost:54320', 'pydantic_ai_rag', ) if create_db: with logfire.span('check and create DB'): conn = await asyncpg.connect(server_dsn) try: db_exists = await conn.fetchval( 'SELECT 1 FROM pg_database WHERE datname = $1', database ) if not db_exists: await conn.execute(f'CREATE DATABASE {database}') finally: await conn.close() pool = await asyncpg.create_pool(f'{server_dsn}/{database}') try: yield pool finally: await pool.close() DB_SCHEMA = """ CREATE EXTENSION IF NOT EXISTS vector; CREATE TABLE IF NOT EXISTS doc_sections ( id serial PRIMARY KEY, url text NOT NULL UNIQUE, title text NOT NULL, content text NOT NULL, -- text-embedding-3-small returns a vector of 1536 floats embedding vector(1536) NOT NULL ); CREATE INDEX IF NOT EXISTS idx_doc_sections_embedding ON doc_sections USING hnsw (embedding vector_l2_ops); """ def slugify(value: str, separator: str, unicode: bool = False) -> str: """Slugify a string, to make it URL friendly.""" # Taken unchanged from https://github.com/Python-Markdown/markdown/blob/3.7/markdown/extensions/toc.py#L38 if not unicode: # Replace Extended Latin characters with ASCII, i.e. `žlutý` => `zluty` value = unicodedata.normalize('NFKD', value) value = value.encode('ascii', 'ignore').decode('ascii') value = re.sub(r'[^\w\s-]', '', value).strip().lower() return re.sub(rf'[{separator}\s]+', separator, value) if __name__ == '__main__': action = sys.argv[1] if len(sys.argv) > 1 else None if action == 'build': asyncio.run(build_search_db()) elif action == 'search': if len(sys.argv) == 3: q = sys.argv[2] else: q = 'How do I configure logfire to work with FastAPI?' asyncio.run(run_agent(q)) else: print( 'uv run --extra examples -m pydantic_ai_examples.rag build|search', file=sys.stderr, ) sys.exit(1) Was this page helpful? Thanks for your feedback! --- # Multimodal Input | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/core-concepts/input/#_top) Multimodal Input ================ Alongside text, agents can accept image, audio, video, and document input, as long as the model supports it. Image Input ----------- [](https://pydantic.dev/docs/ai/core-concepts/input/#image-input) If you have a direct URL for the image, you can use [`ImageUrl`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ImageUrl) : image\_input.py from pydantic_ai import Agent, ImageUrl agent = Agent(model='openai:gpt-5.2') result = agent.run_sync( [\ 'What company is this logo from?',\ ImageUrl(url='https://iili.io/3Hs4FMg.png'),\ ] ) print(result.output) #> This is the logo for Pydantic, a data validation and settings management library in Python. If you have the image locally, you can also use [`BinaryContent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BinaryContent) : local\_image\_input.py import httpx from pydantic_ai import Agent, BinaryContent image_response = httpx.get('https://iili.io/3Hs4FMg.png') # Pydantic logo agent = Agent(model='openai:gpt-5.2') result = agent.run_sync( [\ 'What company is this logo from?',\ BinaryContent(data=image_response.content, media_type='image/png'), # (1)\ ] ) print(result.output) #> This is the logo for Pydantic, a data validation and settings management library in Python. To ensure the example is runnable we download this image from the web, but you can also use `Path().read_bytes()` to read a local file's contents. Audio Input ----------- [](https://pydantic.dev/docs/ai/core-concepts/input/#audio-input) You can provide audio input using either [`AudioUrl`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.AudioUrl) or [`BinaryContent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BinaryContent) . The process is analogous to the examples above. Video Input ----------- [](https://pydantic.dev/docs/ai/core-concepts/input/#video-input) You can provide video input using either [`VideoUrl`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.VideoUrl) or [`BinaryContent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BinaryContent) . The process is analogous to the examples above. Document Input -------------- [](https://pydantic.dev/docs/ai/core-concepts/input/#document-input) You can provide document input using either [`DocumentUrl`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.DocumentUrl) or [`BinaryContent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BinaryContent) . The process is similar to the examples above. If you have a direct URL for the document, you can use [`DocumentUrl`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.DocumentUrl) : document\_input.py from pydantic_ai import Agent, DocumentUrl agent = Agent(model='anthropic:claude-sonnet-4-6') result = agent.run_sync( [\ 'What is the main content of this document?',\ DocumentUrl(url='https://storage.googleapis.com/cloud-samples-data/generative-ai/pdf/2403.05530.pdf'),\ ] ) print(result.output) #> This document is the technical report introducing Gemini 1.5, Google's latest large language model... The supported document formats vary by model. You can also use [`BinaryContent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BinaryContent) to pass document data directly: binary\_content\_input.py from pathlib import Path from pydantic_ai import Agent, BinaryContent pdf_path = Path('document.pdf') agent = Agent(model='anthropic:claude-sonnet-4-6') result = agent.run_sync( [\ 'What is the main content of this document?',\ BinaryContent(data=pdf_path.read_bytes(), media_type='application/pdf'),\ ] ) print(result.output) #> The document discusses... Text Input ---------- [](https://pydantic.dev/docs/ai/core-concepts/input/#text-input) You can use [`TextContent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.TextContent) to provide text input with additional metadata: text\_content\_input.py from pydantic_ai import Agent, TextContent agent = Agent(model='openai:gpt-5.2') result = agent.run_sync([\ 'Summarize the key points from this text.',\ TextContent(\ content=(\ 'Pydantic AI is a Python agent framework. '\ 'It supports text, image, audio, video, and document input.'\ ),\ metadata={'source': 'pydantic_ai_inputs.txt'},\ ),\ ]) This is equivalent to passing the text as a `str`, but allows you to include additional `metadata` that can be accessed programmatically in your agent logic. User-side download vs. direct file URL -------------------------------------- [](https://pydantic.dev/docs/ai/core-concepts/input/#user-side-download-vs-direct-file-url) When using one of `ImageUrl`, `AudioUrl`, `VideoUrl` or `DocumentUrl`, Pydantic AI will default to sending the URL to the model provider, so the file is downloaded on their side. Support for file URLs varies depending on type and provider: | Model | Send URL directly | Download and send bytes | Unsupported | | --- | --- | --- | --- | | [`OpenAIChatModel`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIChatModel) | `ImageUrl` | `AudioUrl`, `DocumentUrl` | `VideoUrl`. `DocumentUrl` [not supported with `AzureProvider`](https://pydantic.dev/docs/ai/models/openai/#using-azure-with-the-responses-api)
or [`AlibabaProvider`](https://pydantic.dev/docs/ai/models/openai/#alibaba-cloud-model-studio-dashscope) | | [`OpenAIResponsesModel`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIResponsesModel) | `ImageUrl`, `AudioUrl`, `DocumentUrl` | — | `VideoUrl` | | [`AnthropicModel`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModel) | `ImageUrl`, `DocumentUrl` (PDF) | `DocumentUrl` (`text/plain`) | `AudioUrl`, `VideoUrl` | | [`GoogleModel`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModel)
(Google Cloud) | All URL types | — | — | | [`GoogleModel`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModel)
(Gemini API) | [YouTube](https://pydantic.dev/docs/ai/models/google/#document-image-audio-and-video-input)
, [Files API](https://pydantic.dev/docs/ai/models/google/#document-image-audio-and-video-input) | All other URLs | — | | [`XaiModel`](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModel) | `ImageUrl` | `DocumentUrl` | `AudioUrl`, `VideoUrl` | | [`MistralModel`](https://pydantic.dev/docs/ai/api/models/mistral/#pydantic_ai.models.mistral.MistralModel) | `ImageUrl`, `DocumentUrl` (PDF) | `DocumentUrl` (`text/plain`) | `AudioUrl`, `VideoUrl`, `DocumentUrl` (non-PDF, non-text) | | [`BedrockConverseModel`](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockConverseModel) | S3 URLs (`s3://`) | `ImageUrl`, `DocumentUrl`, `VideoUrl` | `AudioUrl` | | [`OpenRouterModel`](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModel) | `ImageUrl`, `DocumentUrl`, `VideoUrl` | `AudioUrl` | — | A model API may be unable to download a file (e.g., because of crawling or access restrictions) even if it supports file URLs. For example, [`GoogleModel`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModel) on Google Cloud limits YouTube video URLs to one URL per request. In such cases, you can instruct Pydantic AI to download the file content locally and send that instead of the URL by setting `force_download` on the URL object: force\_download.py from pydantic_ai import ImageUrl, AudioUrl, VideoUrl, DocumentUrl ImageUrl(url='https://example.com/image.png', force_download=True) AudioUrl(url='https://example.com/audio.mp3', force_download=True) VideoUrl(url='https://example.com/video.mp4', force_download=True) DocumentUrl(url='https://example.com/doc.pdf', force_download=True) Uploaded Files -------------- [](https://pydantic.dev/docs/ai/core-concepts/input/#uploaded-files) Some model providers have their own file storage APIs where you can upload files and reference them by ID or URL. Use [`UploadedFile`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.UploadedFile) to reference files that have been uploaded to a provider’s file storage API. ### Supported Models [](https://pydantic.dev/docs/ai/core-concepts/input/#supported-models) | Model | Support | | --- | --- | | [`AnthropicModel`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModel) | ✅ via [Anthropic Files API](https://docs.anthropic.com/en/docs/build-with-claude/files) | | [`OpenAIChatModel`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIChatModel) | ✅ via [OpenAI Files API](https://platform.openai.com/docs/api-reference/files) | | [`OpenAIResponsesModel`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIResponsesModel) | ✅ via [OpenAI Files API](https://platform.openai.com/docs/api-reference/files) | | [`GoogleModel`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModel) | ✅ via [Google Files API](https://ai.google.dev/gemini-api/docs/files) | | [`BedrockConverseModel`](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockConverseModel) | ✅ via S3 URLs (`s3://bucket/key`) | | [`XaiModel`](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModel) | ✅ via [xAI Files API](https://docs.x.ai/docs/guides/files) | | Other models | ❌ Not supported | ### Provider Name Requirement [](https://pydantic.dev/docs/ai/core-concepts/input/#provider-name-requirement) When using [`UploadedFile`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.UploadedFile) you must set the `provider_name`. Uploaded files are specific to the system they are uploaded to and are not transferable across providers. Trying to use a message that contains an `UploadedFile` with a different provider will result in an error. If you want to introduce portability into your agent logic to allow the same prompt history to work with different provider backends, you can use a [history processor](https://pydantic.dev/docs/ai/core-concepts/message-history/#processing-message-history) to remove or rewrite `UploadedFile` parts from messages before sending them to a provider that does not support them. Be aware that stripping out `UploadedFile` instances might confuse the model, especially if references to those files remain in the text. ### Media Type Inference [](https://pydantic.dev/docs/ai/core-concepts/input/#media-type-inference) The `media_type` parameter is optional for [`UploadedFile`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.UploadedFile) . If not specified, Pydantic AI will attempt to infer it from the `file_id`: 1. If `file_id` is a URL or path with a recognizable file extension (e.g., `.pdf`, `.png`), the media type is inferred automatically 2. For opaque file IDs (e.g., `'file-abc123'`), the media type defaults to `'application/octet-stream'` ### Anthropic [](https://pydantic.dev/docs/ai/core-concepts/input/#anthropic) Follow the [Anthropic Files API docs](https://docs.anthropic.com/en/docs/build-with-claude/files) to upload files. You can access the underlying Anthropic client via `provider.client`. uploaded\_file\_anthropic.py import asyncio from pydantic_ai import Agent, UploadedFile from pydantic_ai.models.anthropic import AnthropicModel from pydantic_ai.providers.anthropic import AnthropicProvider async def main(): provider = AnthropicProvider() model = AnthropicModel('claude-sonnet-4-5', provider=provider) # Upload a file using the provider's client (Anthropic client) with open('document.pdf', 'rb') as f: uploaded_file = await provider.client.beta.files.upload(file=f) # Reference the uploaded file; the beta header is added automatically agent = Agent(model) result = await agent.run( [\ 'Summarize this document',\ UploadedFile(file_id=uploaded_file.id, provider_name=model.system),\ ] ) print(result.output) #> The document discusses the main topics and key findings... asyncio.run(main()) ### OpenAI [](https://pydantic.dev/docs/ai/core-concepts/input/#openai) Follow the [OpenAI Files API docs](https://platform.openai.com/docs/api-reference/files/create) to upload files. You can access the underlying OpenAI client via `provider.client`. uploaded\_file\_openai.py import asyncio from pydantic_ai import Agent, UploadedFile from pydantic_ai.models.openai import OpenAIChatModel from pydantic_ai.providers.openai import OpenAIProvider async def main(): provider = OpenAIProvider() model = OpenAIChatModel('gpt-5', provider=provider) # Upload a file using the provider's client (OpenAI client) with open('document.pdf', 'rb') as f: uploaded_file = await provider.client.files.create(file=f, purpose='user_data') # Reference the uploaded file agent = Agent(model) result = await agent.run( [\ 'Summarize this document',\ UploadedFile(file_id=uploaded_file.id, provider_name=model.system),\ ] ) print(result.output) #> The document discusses the main topics and key findings... asyncio.run(main()) ### Google [](https://pydantic.dev/docs/ai/core-concepts/input/#google) Follow the [Google Files API docs](https://ai.google.dev/gemini-api/docs/files) to upload files. You can access the underlying Google GenAI client via `provider.client`. uploaded\_file\_google.py import asyncio from pydantic_ai import Agent, UploadedFile from pydantic_ai.models.google import GoogleModel from pydantic_ai.providers.google import GoogleProvider async def main(): provider = GoogleProvider() model = GoogleModel('gemini-2.5-flash', provider=provider) # Upload a file using the provider's client (Google GenAI client) with open('document.pdf', 'rb') as f: file = await provider.client.aio.files.upload(file=f) assert file.uri is not None # Reference the uploaded file by URI (media_type is optional for Google) agent = Agent(model) result = await agent.run( [\ 'Summarize this document',\ UploadedFile(file_id=file.uri, media_type=file.mime_type, provider_name=model.system),\ ] ) print(result.output) #> The document discusses the main topics and key findings... asyncio.run(main()) ### Bedrock (S3) [](https://pydantic.dev/docs/ai/core-concepts/input/#bedrock-s3) For Bedrock, files must be uploaded to S3 separately (e.g., using [boto3](https://boto3.amazonaws.com/v1/documentation/api/latest/reference/services/s3/client/put_object.html) ). The assumed role must have `s3:GetObject` permission on the bucket. uploaded\_file\_bedrock.py import asyncio from pydantic_ai import Agent, UploadedFile from pydantic_ai.models.bedrock import BedrockConverseModel async def main(): model = BedrockConverseModel('us.anthropic.claude-sonnet-4-20250514-v1:0') agent = Agent(model) result = await agent.run([\ 'Summarize this document',\ UploadedFile(\ file_id='s3://my-bucket/document.pdf',\ provider_name=model.system, # 'bedrock'\ media_type='application/pdf', # Optional for .pdf, but recommended\ ),\ ]) print(result.output) #> The document discusses the main topics and key findings... asyncio.run(main()) ### xAI [](https://pydantic.dev/docs/ai/core-concepts/input/#xai) Follow the [xAI Files API docs](https://docs.x.ai/docs/guides/files) to upload files. You can access the underlying xAI client via `provider.client`. uploaded\_file\_xai.py import asyncio from pydantic_ai import Agent, UploadedFile from pydantic_ai.models.xai import XaiModel from pydantic_ai.providers.xai import XaiProvider async def main(): provider = XaiProvider() model = XaiModel('grok-4.3', provider=provider) # Upload a file using the provider's client (xAI client) with open('document.pdf', 'rb') as f: uploaded_file = await provider.client.files.upload(f, filename='document.pdf') # Reference the uploaded file agent = Agent(model) result = await agent.run( [\ 'Summarize this document',\ UploadedFile(file_id=uploaded_file.id, provider_name=model.system),\ ] ) print(result.output) #> The document discusses the main topics and key findings... asyncio.run(main()) Was this page helpful? Thanks for your feedback! --- # Shell | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/shell/#_top) Shell ===== `Shell` gives an agent the ability to run shell commands, with allow/deny controls, environment scrubbing, and managed background processes. It exposes command-execution tools rooted at a working directory and cleans up any background processes automatically when the agent run ends. [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/shell/) The problem ----------- [](https://pydantic.dev/docs/ai/harness/shell/#the-problem) Agents frequently need to run a build, a test suite, a linter, or a quick `grep`. Wiring up subprocess handling — streaming output, timeouts, truncation, killing runaway processes, and cleaning up background jobs at the end of a run — is fiddly boilerplate that every agent reinvents. `Shell` bundles that plumbing into a single [capability](https://pydantic.dev/docs/ai/capabilities/overview/) : configurable allow/deny lists, output truncation tuned to keep the useful tail, optional sticky working directory, environment control that can keep host secrets out of spawned commands, and automatic cleanup of background processes when the run finishes. Usage ----- [](https://pydantic.dev/docs/ai/harness/shell/#usage) Construct `Shell` with a working directory and pass it to an `Agent` via the `capabilities` parameter: from pydantic_ai import Agent from pydantic_ai_harness import Shell agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[Shell(cwd='./workspace', allowed_commands=['ls', 'cat', 'rg'])], ) result = agent.run_sync('List the Python files and summarize the largest one.') print(result.output) By default `Shell` runs in the current directory with the built-in destructive-command denylist active — `Shell()` alone is a working (if permissive) configuration. Tools ----- [](https://pydantic.dev/docs/ai/harness/shell/#tools) `Shell` contributes four tools to the agent: | Tool | Purpose | | --- | --- | | `run_command` | Run a command synchronously and return labelled stdout/stderr plus exit code. Honors a per-call or default timeout. | | `start_command` | Launch a long-running command (server, watcher) in the background; returns an ID. | | `check_command` | Report the status and accumulated output of a background command. | | `stop_command` | Terminate a background command and return its final output. | `run_command` accepts an optional `timeout_seconds` argument that overrides `default_timeout` for a single call. `check_command` and `stop_command` take the `command_id` string returned by `start_command`. Output is labelled with `[stdout]` / `[stderr]` markers and an `[exit code: N]` line on non-zero exit. When it exceeds `max_output_chars` the **tail** is kept (the head is dropped), so errors, stack traces, and the `[stderr]` section — which all land at the end — survive truncation. Background command status and exit metadata follow the captured output so they remain in the retained tail. Command controls ---------------- [](https://pydantic.dev/docs/ai/harness/shell/#command-controls) Two mutually exclusive lists decide which executables may run, plus filters for shell operators and interactive commands: | Field | Effect | | --- | --- | | `allowed_commands` | If non-empty, only these executables may run (allowlist). | | `denied_commands` | These executables are always rejected (denylist). | | `denied_operators` | Shell operators (e.g. `>`, `>>`, `\|`) that are rejected when present. | | `allow_interactive` | If `False` (default), commands that expect a TTY (`vi`, `sudo`, `ssh`, …) are blocked. | `allowed_commands` and `denied_commands` are mutually exclusive — set one, not both. Setting non-empty values for both raises a `ValueError` when the toolset is constructed. `denied_commands` defaults to a list of destructive commands (`rm`, `rmdir`, `mkfs`, `dd`, `format`, `shutdown`, `reboot`, `halt`, `poweroff`, `init`); pass an empty list to disable it. The executable name is extracted with `shlex`, so arguments don’t bypass the check. An empty `allowed_commands` collection does not select allowlist mode. The configured `denied_commands` remain active; when omitted, this is the built-in denylist. Pass `denied_commands=[]` to disable command-name filtering. A denied command surfaces to the model as a [`ModelRetry`](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-retries) , not a hard error: the run continues and the model can pick an allowed command instead. Environment control ------------------- [](https://pydantic.dev/docs/ai/harness/shell/#environment-control) By default a spawned command inherits the agent process’s full environment. In a sandbox that holds LLM API keys, tokens, or other secrets, a command the model writes can read them. Two fields control what the subprocess sees: | Field | Effect | | --- | --- | | `env` | Explicit environment that replaces inheritance entirely. The subprocess sees exactly these variables and nothing else. | | `denied_env_patterns` | Glob patterns (`fnmatch`) for variable names stripped from the base environment. Mirrors `denied_commands`. | `env` is a hard boundary for inherited environment variables: set it and inherited secrets cannot reach the subprocess at all (you supply `PATH` and anything else the command needs). `denied_env_patterns` is a denylist over the inherited environment — lighter to configure when you only need to drop a few known-sensitive names. The two compose: when both are set, patterns also filter the explicit `env`. Leaving both unset preserves the inherit-everything default. import os from pydantic_ai_harness import Shell from pydantic_ai_harness.shell import LLM_API_KEY_ENV_PATTERNS # Strip provider credentials from the inherited environment. Shell(cwd='./repo', denied_env_patterns=LLM_API_KEY_ENV_PATTERNS) # Or hand the subprocess a fixed environment, inheriting nothing. Shell(cwd='./repo', env={'PATH': os.environ['PATH'], 'HOME': os.environ['HOME']}) `LLM_API_KEY_ENV_PATTERNS` covers common provider prefixes (`ANTHROPIC_*`, `GATEWAY_*`, `GEMINI_*`, `GOOGLE_*`, `OPENAI_*`, `OPENROUTER_*`) plus `PYDANTIC_AI_GATEWAY_API_KEY`. It targets LLM credentials only — it does not cover other host secrets (a `LOGFIRE_TOKEN`, a GitHub token, cloud credentials), and its prefixes are coarse, so `GOOGLE_*` also strips non-credential vars like `GOOGLE_APPLICATION_CREDENTIALS`. Treat it as a starting point and add your own patterns. It is not the default: stripping environment variables silently would break agents that rely on inherited credentials, so it is opt-in. `env` is enforced at spawn, not applied as a post-hoc filter on a running process: the subprocess starts with exactly the resolved environment (your `env`, minus anything `denied_env_patterns` removes from it). That makes it a real boundary for inherited environment variables, unlike the best-effort command denylist. It is not a full security boundary: a command running under the same OS identity can still read host files — use OS-level isolation for that. The flip side is that a pattern broad enough to strip `PATH` or `HOME`, or an `env` that omits them, can break command resolution. External commands may still run via the shell’s built-in default `PATH` on some systems, but don’t rely on it — set `PATH` explicitly when you replace the environment. Background processes -------------------- [](https://pydantic.dev/docs/ai/harness/shell/#background-processes) `start_command` writes stdout/stderr to temp files and returns a short ID. Use `check_command(command_id)` to poll and `stop_command(command_id)` to terminate and collect final output. Processes are launched in their own session (`start_new_session`) so the whole process group can be signalled — `SIGTERM`, escalating to `SIGKILL` after a grace period. On run end, the toolset’s cleanup terminates every still-running background process and deletes its temp files. The agent runtime enters toolsets via an `AsyncExitStack`, so this cleanup runs whether the run succeeds or raises — an agent that forgets to call `stop_command` won’t leak processes. from pydantic_ai import Agent from pydantic_ai_harness import Shell agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[Shell(cwd='./app', allowed_commands=['npm', 'curl'])], ) result = agent.run_sync( 'Start the dev server with `npm run dev`, wait for it to boot, ' 'then curl http://localhost:3000/health and report the status.' ) print(result.output) Working directory ----------------- [](https://pydantic.dev/docs/ai/harness/shell/#working-directory) By default each command runs in `cwd` and `cd` has no lasting effect. Set `persist_cwd=True` to make `cd` sticky across calls: each command is wrapped so that after it runs, its final working directory is recorded to a private temp file, and that directory is carried into subsequent calls. The path is only updated when the command exits `0`, and the record is written out-of-band (not to stdout) so command output can never spoof the tracked directory. from pydantic_ai import Agent from pydantic_ai_harness import Shell agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[Shell(cwd='.', persist_cwd=True, allowed_commands=['cd', 'ls', 'pwd'])], ) Each run gets a fresh toolset instance, so the tracked directory and any background processes are isolated between concurrent runs and always start back at the configured `cwd`. Configuration ------------- [](https://pydantic.dev/docs/ai/harness/shell/#configuration) Every field of `Shell` with its default: from pydantic_ai_harness import Shell Shell( cwd='.', # str | Path -- working directory allowed_commands=[], # allowlist (mutually exclusive with denied) denied_commands=[...], # denylist (defaults to destructive commands) denied_operators=[], # blocked shell operators default_timeout=30.0, # seconds, per run_command max_output_chars=50_000, # output cap returned to the model persist_cwd=False, # make cd sticky across calls allow_interactive=False, # allow TTY-style commands env=None, # explicit env, replacing inheritance (None = inherit) denied_env_patterns=[], # glob patterns stripped from the env ) Agent spec (YAML/JSON) ---------------------- [](https://pydantic.dev/docs/ai/harness/shell/#agent-spec-yamljson) `Shell` works with Pydantic AI’s [agent spec](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) , so you can declare it in a config file instead of Python: # agent.yaml model: anthropic:claude-sonnet-4-6 capabilities: - Shell: cwd: ./workspace allowed_commands: ['ls', 'cat', 'rg', 'pytest'] from pydantic_ai import Agent from pydantic_ai_harness import Shell agent = Agent.from_file('agent.yaml', custom_capability_types=[Shell]) Pass `custom_capability_types` so the spec loader knows how to instantiate `Shell`. Further reading --------------- [](https://pydantic.dev/docs/ai/harness/shell/#further-reading) * [Pydantic AI capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) * [Toolsets](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) API reference ------------- [](https://pydantic.dev/docs/ai/harness/shell/#api-reference) Shell ----- [](https://pydantic.dev/docs/ai/harness/shell/#pydantic_ai_harness.Shell) **Bases:** `AbstractCapability[AgentDepsT]` Shell command execution for agents. Commands execute in a subprocess rooted at `cwd`. Use `allowed_commands` or `denied_commands` to control what the agent can invoke. ### Attributes [](https://pydantic.dev/docs/ai/harness/shell/#attributes) #### allow\_interactive [](https://pydantic.dev/docs/ai/harness/shell/#pydantic_ai_harness.Shell.allow_interactive) If True, allow interactive commands (vi, nano, ssh, etc.). Blocked by default. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### allowed\_commands [](https://pydantic.dev/docs/ai/harness/shell/#pydantic_ai_harness.Shell.allowed_commands) If non-empty, only these command names may be executed (allowlist). **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `field(default_factory=(list[str]))` #### cwd [](https://pydantic.dev/docs/ai/harness/shell/#pydantic_ai_harness.Shell.cwd) Working directory for command execution. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | `Path` **Default:** `'.'` #### default\_timeout [](https://pydantic.dev/docs/ai/harness/shell/#pydantic_ai_harness.Shell.default_timeout) Default timeout in seconds for command execution. **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) **Default:** `30.0` #### denied\_commands [](https://pydantic.dev/docs/ai/harness/shell/#pydantic_ai_harness.Shell.denied_commands) These command names are always rejected (denylist). Defaults to blocking destructive commands (rm, dd, shutdown, etc.). Set to an empty list to disable. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `_DEFAULT_DENIED_COMMANDS` #### denied\_env\_patterns [](https://pydantic.dev/docs/ai/harness/shell/#pydantic_ai_harness.Shell.denied_env_patterns) Glob patterns for environment variable names to strip before spawning. Follows the `denied_*` naming convention but matches by glob (`fnmatch`, e.g. `OPENAI_*`), since env secrets cluster by prefix — unlike `denied_commands`, which matches executable names exactly. Names matching any pattern are removed from the base environment; applied on top of `env` when both are set, so patterns filter an explicit `env` too. See `LLM_API_KEY_ENV_PATTERNS` for a ready-made provider-credential denylist. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `field(default_factory=(list[str]))` #### denied\_operators [](https://pydantic.dev/docs/ai/harness/shell/#pydantic_ai_harness.Shell.denied_operators) Shell operators that are blocked (e.g. ’>’, ’>>’, ’|’ for restrictive mode). **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `field(default_factory=(list[str]))` #### env [](https://pydantic.dev/docs/ai/harness/shell/#pydantic_ai_harness.Shell.env) Explicit environment for spawned subprocesses, replacing inheritance. When `None` (default) the subprocess inherits the parent environment. Set this to a fixed mapping to start subprocesses with exactly these variables and nothing else — a hard boundary that keeps host secrets (LLM API keys, tokens) out of commands the agent runs. **Type:** [`Mapping`](https://docs.python.org/3/library/typing.html#typing.Mapping) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### max\_output\_chars [](https://pydantic.dev/docs/ai/harness/shell/#pydantic_ai_harness.Shell.max_output_chars) Maximum characters of output returned to the model. Must be positive. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `50000` #### persist\_cwd [](https://pydantic.dev/docs/ai/harness/shell/#pydantic_ai_harness.Shell.persist_cwd) If True, track cd commands and adjust the working directory for subsequent calls. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` ### Methods [](https://pydantic.dev/docs/ai/harness/shell/#methods) #### \_\_post\_init\_\_ [](https://pydantic.dev/docs/ai/harness/shell/#pydantic_ai_harness.Shell.__post_init__) def __post_init__() -> None Resolve the built-in denylist according to the selected policy. ##### Returns [](https://pydantic.dev/docs/ai/harness/shell/#returns) [`None`](https://docs.python.org/3/library/constants.html#None) #### get\_toolset [](https://pydantic.dev/docs/ai/harness/shell/#pydantic_ai_harness.Shell.get_toolset) def get_toolset() -> ShellToolset[AgentDepsT] Build and return the shell toolset. ##### Returns [](https://pydantic.dev/docs/ai/harness/shell/#returns-1) `ShellToolset`\[`AgentDepsT`\] Was this page helpful? Thanks for your feedback! --- # Groq | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/models/groq/#_top) Groq ==== Install ------- [](https://pydantic.dev/docs/ai/models/groq/#install) To use `GroqModel`, you need to either install `pydantic-ai`, or install `pydantic-ai-slim` with the `groq` optional group: * [pip](https://pydantic.dev/docs/ai/models/groq/#tab-panel-114) * [uv](https://pydantic.dev/docs/ai/models/groq/#tab-panel-115) Terminal pip install "pydantic-ai-slim[groq]" Terminal uv add "pydantic-ai-slim[groq]" Configuration ------------- [](https://pydantic.dev/docs/ai/models/groq/#configuration) To use [Groq](https://groq.com/) through their API, go to [console.groq.com/keys](https://console.groq.com/keys) and follow your nose until you find the place to generate an API key. `GroqModelName` contains a list of available Groq models. Environment variable -------------------- [](https://pydantic.dev/docs/ai/models/groq/#environment-variable) Once you have the API key, you can set it as an environment variable: Terminal export GROQ_API_KEY='your-api-key' You can then use `GroqModel` by name: from pydantic_ai import Agent agent = Agent('groq:llama-3.3-70b-versatile') ... Or initialise the model directly with just the model name: from pydantic_ai import Agent from pydantic_ai.models.groq import GroqModel model = GroqModel('llama-3.3-70b-versatile') agent = Agent(model) ... `provider` argument ------------------- [](https://pydantic.dev/docs/ai/models/groq/#provider-argument) You can provide a custom `Provider` via the `provider` argument: from pydantic_ai import Agent from pydantic_ai.models.groq import GroqModel from pydantic_ai.providers.groq import GroqProvider model = GroqModel( 'llama-3.3-70b-versatile', provider=GroqProvider(api_key='your-api-key') ) agent = Agent(model) ... You can also customize the `GroqProvider` with a custom `httpx.AsyncClient`: from httpx import AsyncClient from pydantic_ai import Agent from pydantic_ai.models.groq import GroqModel from pydantic_ai.providers.groq import GroqProvider custom_http_client = AsyncClient(timeout=30) model = GroqModel( 'llama-3.3-70b-versatile', provider=GroqProvider(api_key='your-api-key', http_client=custom_http_client), ) agent = Agent(model) ... Was this page helpful? Thanks for your feedback! --- # Parallel Execution | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/graph/builder/parallel/#_top) Parallel Execution ================== The graph builder API provides two powerful mechanisms for parallel execution: **broadcasting** and **mapping**. * **Broadcasting** - Send the same data to multiple parallel paths * **Spreading** - Fan out items from an iterable to parallel paths Both create “forks” in the execution graph that can later be synchronized with [join nodes](https://pydantic.dev/docs/ai/graph/builder/joins/) . Broadcasting ------------ [](https://pydantic.dev/docs/ai/graph/builder/parallel/#broadcasting) Broadcasting sends identical data to multiple destinations simultaneously: basic\_broadcast.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, reduce_list_append @dataclass class SimpleState: pass async def main(): g = GraphBuilder(state_type=SimpleState, output_type=list[int]) @g.step async def source(ctx: StepContext[SimpleState, None, None]) -> int: return 10 @g.step async def add_one(ctx: StepContext[SimpleState, None, int]) -> int: return ctx.inputs + 1 @g.step async def add_two(ctx: StepContext[SimpleState, None, int]) -> int: return ctx.inputs + 2 @g.step async def add_three(ctx: StepContext[SimpleState, None, int]) -> int: return ctx.inputs + 3 collect = g.join(reduce_list_append, initial_factory=list[int]) # Broadcasting: send the value from source to all three steps g.add( g.edge_from(g.start_node).to(source), g.edge_from(source).to(add_one, add_two, add_three), g.edge_from(add_one, add_two, add_three).to(collect), g.edge_from(collect).to(g.end_node), ) graph = g.build() result = await graph.run(state=SimpleState()) print(sorted(result)) #> [11, 12, 13] _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ All three steps receive the same input value (`10`) and execute in parallel. Spreading --------- [](https://pydantic.dev/docs/ai/graph/builder/parallel/#spreading) Spreading fans out elements from an iterable, processing each element in parallel: basic\_map.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, reduce_list_append @dataclass class SimpleState: pass async def main(): g = GraphBuilder(state_type=SimpleState, output_type=list[int]) @g.step async def generate_list(ctx: StepContext[SimpleState, None, None]) -> list[int]: return [1, 2, 3, 4, 5] @g.step async def square(ctx: StepContext[SimpleState, None, int]) -> int: return ctx.inputs * ctx.inputs collect = g.join(reduce_list_append, initial_factory=list[int]) # Spreading: each item in the list gets its own parallel execution g.add( g.edge_from(g.start_node).to(generate_list), g.edge_from(generate_list).map().to(square), g.edge_from(square).to(collect), g.edge_from(collect).to(g.end_node), ) graph = g.build() result = await graph.run(state=SimpleState()) print(sorted(result)) #> [1, 4, 9, 16, 25] _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ ### Spreading AsyncIterables [](https://pydantic.dev/docs/ai/graph/builder/parallel/#spreading-asynciterables) The `.map()` operation also works with `AsyncIterable` values. When mapping over an async iterable, the graph creates parallel tasks dynamically as values are yielded. This is particularly useful for streaming data or processing data that’s being generated on-the-fly: async\_iterable\_map.py import asyncio from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, reduce_list_append @dataclass class SimpleState: pass async def main(): g = GraphBuilder(state_type=SimpleState, output_type=list[int]) @g.stream async def stream_numbers(ctx: StepContext[SimpleState, None, None]): """Stream numbers with delays to simulate real-time data.""" for i in range(1, 4): await asyncio.sleep(0.05) # Simulate delay yield i @g.step async def triple(ctx: StepContext[SimpleState, None, int]) -> int: return ctx.inputs * 3 collect = g.join(reduce_list_append, initial_factory=list[int]) g.add( g.edge_from(g.start_node).to(stream_numbers), # Map over the async iterable - tasks created as items are yielded g.edge_from(stream_numbers).map().to(triple), g.edge_from(triple).to(collect), g.edge_from(collect).to(g.end_node), ) graph = g.build() result = await graph.run(state=SimpleState()) print(sorted(result)) #> [3, 6, 9] _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ This allows for progressive processing where downstream steps can start working on early results while later results are still being generated. ### Using `add_mapping_edge()` [](https://pydantic.dev/docs/ai/graph/builder/parallel/#using-add_mapping_edge) The convenience method [`add_mapping_edge()`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder.add_mapping_edge) provides a simpler syntax: mapping\_convenience.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, reduce_list_append @dataclass class SimpleState: pass async def main(): g = GraphBuilder(state_type=SimpleState, output_type=list[str]) @g.step async def generate_numbers(ctx: StepContext[SimpleState, None, None]) -> list[int]: return [10, 20, 30] @g.step async def stringify(ctx: StepContext[SimpleState, None, int]) -> str: return f'Value: {ctx.inputs}' collect = g.join(reduce_list_append, initial_factory=list[str]) g.add(g.edge_from(g.start_node).to(generate_numbers)) g.add_mapping_edge(generate_numbers, stringify) g.add( g.edge_from(stringify).to(collect), g.edge_from(collect).to(g.end_node), ) graph = g.build() result = await graph.run(state=SimpleState()) print(sorted(result)) #> ['Value: 10', 'Value: 20', 'Value: 30'] _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Empty Iterables --------------- [](https://pydantic.dev/docs/ai/graph/builder/parallel/#empty-iterables) When mapping an empty iterable, you can specify a `downstream_join_id` to ensure the join still executes: empty\_map.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, reduce_list_append @dataclass class SimpleState: pass async def main(): g = GraphBuilder(state_type=SimpleState, output_type=list[int]) @g.step async def generate_empty(ctx: StepContext[SimpleState, None, None]) -> list[int]: return [] @g.step async def double(ctx: StepContext[SimpleState, None, int]) -> int: return ctx.inputs * 2 collect = g.join(reduce_list_append, initial_factory=list[int]) g.add(g.edge_from(g.start_node).to(generate_empty)) g.add_mapping_edge(generate_empty, double, downstream_join_id=collect.id) g.add( g.edge_from(double).to(collect), g.edge_from(collect).to(g.end_node), ) graph = g.build() result = await graph.run(state=SimpleState()) print(result) #> [] _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Nested Parallel Operations -------------------------- [](https://pydantic.dev/docs/ai/graph/builder/parallel/#nested-parallel-operations) You can nest broadcasts and maps for complex parallel patterns: ### Spread then Broadcast [](https://pydantic.dev/docs/ai/graph/builder/parallel/#spread-then-broadcast) map\_then\_broadcast.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, reduce_list_append @dataclass class SimpleState: pass async def main(): g = GraphBuilder(state_type=SimpleState, output_type=list[int]) @g.step async def generate_list(ctx: StepContext[SimpleState, None, None]) -> list[int]: return [10, 20] @g.step async def add_one(ctx: StepContext[SimpleState, None, int]) -> int: return ctx.inputs + 1 @g.step async def add_two(ctx: StepContext[SimpleState, None, int]) -> int: return ctx.inputs + 2 collect = g.join(reduce_list_append, initial_factory=list[int]) g.add( g.edge_from(g.start_node).to(generate_list), # Spread the list, then broadcast each item to both steps g.edge_from(generate_list).map().to(add_one, add_two), g.edge_from(add_one, add_two).to(collect), g.edge_from(collect).to(g.end_node), ) graph = g.build() result = await graph.run(state=SimpleState()) print(sorted(result)) #> [11, 12, 21, 22] _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ The result contains: * From 10: `10+1=11` and `10+2=12` * From 20: `20+1=21` and `20+2=22` ### Multiple Sequential Spreads [](https://pydantic.dev/docs/ai/graph/builder/parallel/#multiple-sequential-spreads) sequential\_maps.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, reduce_list_append @dataclass class SimpleState: pass async def main(): g = GraphBuilder(state_type=SimpleState, output_type=list[str]) @g.step async def generate_pairs(ctx: StepContext[SimpleState, None, None]) -> list[tuple[int, int]]: return [(1, 2), (3, 4)] @g.step async def unpack_pair(ctx: StepContext[SimpleState, None, tuple[int, int]]) -> list[int]: return [ctx.inputs[0], ctx.inputs[1]] @g.step async def stringify(ctx: StepContext[SimpleState, None, int]) -> str: return f'num:{ctx.inputs}' collect = g.join(reduce_list_append, initial_factory=list[str]) g.add( g.edge_from(g.start_node).to(generate_pairs), # First map: one task per tuple g.edge_from(generate_pairs).map().to(unpack_pair), # Second map: one task per number in each tuple g.edge_from(unpack_pair).map().to(stringify), g.edge_from(stringify).to(collect), g.edge_from(collect).to(g.end_node), ) graph = g.build() result = await graph.run(state=SimpleState()) print(sorted(result)) #> ['num:1', 'num:2', 'num:3', 'num:4'] _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Edge Labels ----------- [](https://pydantic.dev/docs/ai/graph/builder/parallel/#edge-labels) Add labels to parallel edges for better documentation: labeled\_parallel.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, reduce_list_append @dataclass class SimpleState: pass async def main(): g = GraphBuilder(state_type=SimpleState, output_type=list[str]) @g.step async def generate(ctx: StepContext[SimpleState, None, None]) -> list[int]: return [1, 2, 3] @g.step async def process(ctx: StepContext[SimpleState, None, int]) -> str: return f'item-{ctx.inputs}' collect = g.join(reduce_list_append, initial_factory=list[str]) g.add(g.edge_from(g.start_node).to(generate)) g.add_mapping_edge( generate, process, pre_map_label='before map', post_map_label='after map', ) g.add( g.edge_from(process).to(collect), g.edge_from(collect).to(g.end_node), ) graph = g.build() result = await graph.run(state=SimpleState()) print(sorted(result)) #> ['item-1', 'item-2', 'item-3'] _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ State Sharing in Parallel Execution ----------------------------------- [](https://pydantic.dev/docs/ai/graph/builder/parallel/#state-sharing-in-parallel-execution) All parallel tasks share the same graph state. Be careful with mutations: parallel\_state.py from dataclasses import dataclass, field from pydantic_graph import GraphBuilder, StepContext, reduce_list_append @dataclass class CounterState: values: list[int] = field(default_factory=list) async def main(): g = GraphBuilder(state_type=CounterState, output_type=list[int]) @g.step async def generate(ctx: StepContext[CounterState, None, None]) -> list[int]: return [1, 2, 3] @g.step async def track_and_square(ctx: StepContext[CounterState, None, int]) -> int: # All parallel tasks mutate the same state ctx.state.values.append(ctx.inputs) return ctx.inputs * ctx.inputs collect = g.join(reduce_list_append, initial_factory=list[int]) g.add( g.edge_from(g.start_node).to(generate), g.edge_from(generate).map().to(track_and_square), g.edge_from(track_and_square).to(collect), g.edge_from(collect).to(g.end_node), ) graph = g.build() state = CounterState() result = await graph.run(state=state) print(f'Squared: {sorted(result)}') #> Squared: [1, 4, 9] print(f'Tracked: {sorted(state.values)}') #> Tracked: [1, 2, 3] _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Edge Transformations -------------------- [](https://pydantic.dev/docs/ai/graph/builder/parallel/#edge-transformations) You can transform data inline as it flows along edges using the `.transform()` method: edge\_transform.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext @dataclass class SimpleState: pass async def main(): g = GraphBuilder(state_type=SimpleState, output_type=str) @g.step async def generate_number(ctx: StepContext[SimpleState, None, None]) -> int: return 42 @g.step async def format_output(ctx: StepContext[SimpleState, None, str]) -> str: return f'The answer is: {ctx.inputs}' # Transform the number to a string inline g.add( g.edge_from(g.start_node).to(generate_number), g.edge_from(generate_number).transform(lambda ctx: str(ctx.inputs * 2)).to(format_output), g.edge_from(format_output).to(g.end_node), ) graph = g.build() result = await graph.run(state=SimpleState()) print(result) #> The answer is: 84 _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ The transform function receives a [`StepContext`](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StepContext) with the current inputs and has access to state and dependencies. This is useful for: * Converting data types between incompatible steps * Extracting specific fields from complex objects * Applying simple computations without creating a full step * Adapting data formats during routing Transforms can be chained and combined with other edge operations like `.map()` and `.label()`: chained\_transforms.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, reduce_list_append @dataclass class SimpleState: pass async def main(): g = GraphBuilder(state_type=SimpleState, output_type=list[str]) @g.step async def generate_data(ctx: StepContext[SimpleState, None, None]) -> list[dict[str, int]]: return [{'value': 10}, {'value': 20}, {'value': 30}] @g.step async def process_number(ctx: StepContext[SimpleState, None, int]) -> str: return f'Processed: {ctx.inputs}' collect = g.join(reduce_list_append, initial_factory=list[str]) g.add( g.edge_from(g.start_node).to(generate_data), # Transform to extract values, then map over them g.edge_from(generate_data) .transform(lambda ctx: [item['value'] for item in ctx.inputs]) .label('Extract values') .map() .to(process_number), g.edge_from(process_number).to(collect), g.edge_from(collect).to(g.end_node), ) graph = g.build() result = await graph.run(state=SimpleState()) print(sorted(result)) #> ['Processed: 10', 'Processed: 20', 'Processed: 30'] _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Next Steps ---------- [](https://pydantic.dev/docs/ai/graph/builder/parallel/#next-steps) * Learn about [join nodes](https://pydantic.dev/docs/ai/graph/builder/joins/) for aggregating parallel results * Explore [conditional branching](https://pydantic.dev/docs/ai/graph/builder/decisions/) with decision nodes * See the [steps documentation](https://pydantic.dev/docs/ai/graph/builder/steps/) for more on step execution Was this page helpful? Thanks for your feedback! --- # MCP | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/mcp/#_top) MCP === [`MCP`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.MCP) is a [provider-adaptive capability](https://pydantic.dev/docs/ai/capabilities/overview/#provider-adaptive-tools) and the primary entry point for [MCP](https://pydantic.dev/docs/ai/mcp/overview/) in Pydantic AI. It runs the MCP server locally by default — keeping credentials, hooks, and tracing under your control — and supports both URL-based servers and direct client / toolset / transport inputs. Backed by [`MCPServerTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.MCPServerTool) on the native side (see [MCP Server Tool](https://pydantic.dev/docs/ai/tools-toolsets/native-tools/#mcp-server-tool) for provider support and configuration) — pass `native=MCPServerTool(...)` directly when you need full control (e.g. a different `id`, `authorization_token`, or `description` than the capability would derive). On the local side, `local=` accepts any [`MCPToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset) input (URL, `fastmcp.Client`, transport, in-process `FastMCP` server, script path, …) — non-toolset inputs are wrapped in `MCPToolset` automatically. mcp.py from pydantic_ai.capabilities import MCP from pydantic_ai.native_tools import MCPServerTool # URL-based MCP server, running locally (requires `pydantic-ai-slim[mcp]`) MCP('https://mcp.example.com/api') # Local client without a URL — pass any `MCPToolset` input # (URL, `fastmcp.Client`, transport, in-process `FastMCP` server, script path, etc.) MCP(local=my_fastmcp_client) # Native preferred; URL-based local fallback MCP('https://mcp.example.com/api', native=True) # Strict native-only (no local — does not require the `mcp` extra) MCP('https://mcp.example.com/api', native=True, local=False) # Explicit native + explicit local — independent configuration on each side # (e.g. provider-relay URL for native, direct connection for local) MCP( native=MCPServerTool( id='public-mcp', url='https://relay.example.com/mcp', authorization_token='relay-token', ), local=my_fastmcp_client, ) For lower-level access — managing the [`MCPToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset) lifecycle directly, advanced transport / client configuration, or using MCP servers without going through a capability — see the [MCP documentation](https://pydantic.dev/docs/ai/mcp/overview/) . Was this page helpful? Thanks for your feedback! --- # Restate | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/durable_execution/restate/#_top) Restate ======= [Restate](https://restate.dev/) is a lightweight durable execution runtime with first-class support for AI agents. The Pydantic AI integration is provided via the [Restate Python SDK](https://github.com/restatedev/sdk-python/tree/main/python/restate/ext/pydantic) . Visit the [Restate documentation](https://docs.restate.dev/ai/patterns/durable-agents) for more information. Durable Execution ----------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/restate/#durable-execution) Restate makes your agent **durable** by recording every step of its execution in a journal. If your process crashes mid-execution, Restate replays the journal, skips completed steps, and resumes from exactly where it left off. Your agent runs in a regular HTTP handler inside a Restate **service**. The Restate Server sits in front of your application and manages orchestration, journaling, and retries. Services run like regular Docker containers or serverless functions. A durable agent has three building blocks: 1. The **handler**: your agent logic, exposed as an HTTP endpoint in a Restate service. 2. **LLM calls**: persisted so responses are not re-fetched on recovery — saving cost and time. 3. **Tool executions**: wrapped in durable steps so side effects are not duplicated. Clients (HTTP, Kafka, etc.) | v +---------------------+ | Restate Server | (Journals execution, +---------------------+ retries on failure, ^ manages state) | Journal | Replay on steps, | recovery, retries | schedule calls v +------------------------------------------------------+ | Application Process | | +----------------------------------------------+ | | | Restate Service Handler | | | | (Agent Run Loop) | | | | [ Durable Steps (Tool, MCP, Model) ] | | | +----------------------------------------------+ | | | | | | +------------------------------------------------------+ | | | v v v [External APIs, services, databases, etc.] See the [Restate documentation](https://docs.restate.dev/ai/patterns/durable-agents) for more information. Durable Agent ------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/restate/#durable-agent) Any Pydantic AI agent can be made durable by wrapping it with `RestateAgent` from the Restate SDK and running it inside a Restate service handler. Install the Restate SDK: * [pip](https://pydantic.dev/docs/ai/capabilities/durable_execution/restate/#tab-panel-8) * [uv](https://pydantic.dev/docs/ai/capabilities/durable_execution/restate/#tab-panel-9) Terminal pip install pydantic-ai "restate_sdk[serde]" Terminal uv add pydantic-ai "restate_sdk[serde]" Here is a complete example of a durable Pydantic AI agent with Restate: restate\_agent.py import restate from pydantic_ai import Agent, RunContext from restate.ext.pydantic import RestateAgent, restate_context weather_agent = Agent( # (1) 'openai:gpt-5.2', system_prompt='You are a helpful agent that provides weather updates.', ) @weather_agent.tool() async def get_weather(_run_ctx: RunContext, city: str) -> dict: """Get the current weather for a given city.""" # Do durable tool steps using the Restate context async def call_weather_api(city: str) -> dict: return {'temperature': 23, 'description': 'Sunny and warm.'} return await restate_context().run_typed( # (2) f'Get weather {city}', call_weather_api, city=city ) restate_agent = RestateAgent(weather_agent) # (3) agent_service = restate.Service('WeatherAgent') @agent_service.handler() async def run(_ctx: restate.Context, prompt: str) -> str: # (4) result = await restate_agent.run(prompt) return result.output app = restate.app(services=[agent_service]) # (5) if __name__ == "__main__": # (6) import hypercorn import asyncio conf = hypercorn.Config() conf.bind = ["0.0.0.0:9080"] asyncio.run(hypercorn.asyncio.serve(app, conf)) Define your agent and tools as you normally would with Pydantic AI. Use `restate_context()` actions inside tools to make their execution durable. The result is persisted and retried until it succeeds. Side effects won't be duplicated on recovery. `RestateAgent` wraps the agent so every LLM response is saved in the Restate Server and replayed during recovery. The Restate service handler gives the agent a durable execution context and exposes it as an HTTP endpoint. `restate.app()` creates the application that can be served. Run the application with an ASGI server like Hypercorn. See the [Restate agent quickstart](https://docs.restate.dev/ai-quickstart) to learn how to run the agent. Was this page helpful? Thanks for your feedback! --- # History and handoff | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/realtime/history/#_top) History and handoff =================== A realtime session builds the same [`ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage) history as a standard agent run — see [Messages and chat history](https://pydantic.dev/docs/ai/core-concepts/message-history/) . Voice conversations can start from earlier text or voice history, continue in a new realtime session, or hand off to a text model for summarization, extraction, and follow-up. Spoken turns use [`SpeechPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SpeechPart) ; text, images, and tools retain the ordinary [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) and [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) shape. Reading session history ----------------------- [](https://pydantic.dev/docs/ai/realtime/history/#reading-session-history) The session exposes copy-on-read snapshots: | Method | Returns | | --- | --- | | [`all_messages()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.all_messages) | Seeded history plus messages recorded during this session. | | [`new_messages()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.new_messages) | Only messages recorded during this session. | Seeding a session ----------------- [](https://pydantic.dev/docs/ai/realtime/history/#seeding-a-session) Pass `message_history=` to seed a new session. Replayable text, speech transcripts, thinking text, tool rounds, and supported images are projected into provider conversation items. [History processors](https://pydantic.dev/docs/ai/core-concepts/message-history/#processing-message-history) do not run at seeding; see [Capabilities and hooks](https://pydantic.dev/docs/ai/realtime/capabilities/#seeded-history-is-not-processed) . from pydantic_ai import Agent voice = Agent(instructions='You are a helpful voice assistant.') async def main(prior_history=()): async with voice.realtime( 'openai:gpt-realtime', message_history=prior_history, ).session() as session: await session.send('Continue where we left off.') Providers replay native function calls where their protocol permits. Gemini represents seeded tool calls and results as readable text because Live cannot put function parts in seeded turns. Thinking signatures and provider-native execution metadata are omitted because they belong to the session that produced them. Content-less speech parts are skipped because they carry no replayable content. Unsupported content raises [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) instead of being silently dropped. Video, documents, uploaded-file references, and model-generated files cannot be seeded. Speech transcripts are preferred over retained audio. OpenAI and Azure OpenAI can replay retained user audio when no transcript exists; Gemini and xAI cannot. Assistant speech always needs a transcript for seeding. Check `supports_session_seeding`, `supports_seeding_images`, and `supports_seeding_audio` on the [`RealtimeModelProfile`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelProfile) (see [Provider support](https://pydantic.dev/docs/ai/realtime/overview/#provider-support) for how profiles resolve) before constructing portable flows. Handing off to a text agent --------------------------- [](https://pydantic.dev/docs/ai/realtime/history/#handing-off-to-a-text-agent) Pass the session snapshot directly to [`Agent.run()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.run) : from pydantic_ai import Agent from pydantic_ai.realtime import RealtimeTurnCompleteEvent voice = Agent(instructions='You are a helpful voice assistant.') notetaker = Agent('openai:gpt-5', instructions='Summarize the conversation as bullet points.') async def main(): async with voice.realtime('openai:gpt-realtime').session() as session: await session.send('Please remind me to book a train tomorrow.') async for event in session: if isinstance(event, RealtimeTurnCompleteEvent): break result = await notetaker.run( 'Summarize the conversation.', message_history=session.all_messages() ) print(result.output) #> - Book a train tomorrow. Retained user audio is forwarded to standard models whose profile supports audio input; other models receive the transcript. Assistant speech is always handed off as transcript text. Interrupted assistant turns receive a readable interruption note when Pydantic AI prepares the text-model request. For structured work that must finish while the call remains open, expose a delegated text agent as a [realtime function tool](https://pydantic.dev/docs/ai/realtime/tools/#delegating-work-during-a-call) . Retaining audio --------------- [](https://pydantic.dev/docs/ai/realtime/history/#retaining-audio) By default, only transcripts are retained and `SpeechPart.audio` is `None`. Pass `audio_retention=` to [`session()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentRealtime.session) to retain finalized WAV audio in history: | [`AudioRetention`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.AudioRetention)
value | Retains | | --- | --- | | `'transcript_only'` (default) | Transcripts only | | `'input_audio'` | User audio | | `'output_audio'` | Model audio | | `'all'` | Both sides’ audio | Retention affects history only. Live input and output remain raw PCM16; finalized retained audio is wrapped in a WAV container. Input retention follows provider-reported boundaries rather than locally trimming speech. OpenAI, Azure OpenAI, and xAI normally retain microphone input between reported speech-end boundaries. Gemini does not report those boundaries, so it retains input between response completions. Either form can include silence or other microphone input and should not be treated as a precisely trimmed utterance. Retaining images ---------------- [](https://pydantic.dev/docs/ai/realtime/history/#retaining-images) Pass `retain_images_every_n=` and `retain_images_max=` to [`session()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentRealtime.session) to bound how many images stay in local history: from pydantic_ai import Agent agent = Agent() async def main(): async with agent.realtime('openai:gpt-realtime').session( retain_images_every_n=10, retain_images_max=25 ): ... Images sent through [`send()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.send) are recorded by default. `retain_images_every_n=N` keeps the first image and then one of every `N`; the provider still receives every frame. `retain_images_max` defaults to `100` and evicts the oldest retained image when the cap is reached. Set it to `0` to retain none or `None` to remove the bound. Sampling controls history growth rate; the maximum provides the actual memory bound. Streaming images continuously approximates live video — the [camera example](https://pydantic.dev/docs/ai/examples/realtime/realtime-camera/) sends one frame per second — so for camera and screen streams, use both deliberately. Transcription and history edge cases ------------------------------------ [](https://pydantic.dev/docs/ai/realtime/history/#transcription-and-history-edge-cases) Input transcription defaults to `'auto'`; see [Input transcription](https://pydantic.dev/docs/ai/realtime/audio/#input-transcription) and each provider page for configuration. With transcription disabled: * retained input audio creates an audio-only user `SpeechPart`; * without input retention, the session records a content-less user `SpeechPart`; * content-less parts preserve the local turn boundary but contribute no words to a text handoff and are skipped when seeding another realtime session; * transcript-less assistant audio cannot be handed off or seeded on any provider. If a future session must be portable across providers or models, retain transcripts. Filter or transform unsupported parts before passing `message_history`; history-processing capabilities do not run during realtime seeding. Was this page helpful? Thanks for your feedback! --- # Reinject System Prompt | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/reinject-system-prompt/#_top) Reinject System Prompt ====================== [`ReinjectSystemPrompt`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ReinjectSystemPrompt) is a [capability](https://pydantic.dev/docs/ai/capabilities/overview/) that ensures the agent’s configured [`system_prompt`](https://pydantic.dev/docs/ai/core-concepts/agent/#system-prompts) is at the head of the first [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) on every model request. By default, if any [`SystemPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SystemPromptPart) is already present in the history, the capability is a no-op (so multi-agent handoff and user-managed system prompts remain authoritative). Set `replace_existing=True` to instead strip any existing `SystemPromptPart`s before prepending the agent’s configured prompt — useful when the history comes from an untrusted source and the server’s prompt must win. Useful when `message_history` comes from a source that doesn’t round-trip system prompts — UI frontends, database persistence layers, conversation compaction pipelines. Without this capability, an agent configured with a `system_prompt` will silently run without it if the history doesn’t already include one. reinject\_system\_prompt.py from pydantic_ai import Agent from pydantic_ai.capabilities import ReinjectSystemPrompt from pydantic_ai.messages import ModelRequest, ModelResponse, TextPart, UserPromptPart agent = Agent('test', system_prompt='You are a helpful assistant.', capabilities=[ReinjectSystemPrompt()]) # History that's missing the system prompt (e.g. reconstructed from a UI frontend). history = [\ ModelRequest(parts=[UserPromptPart(content='Hi')]),\ ModelResponse(parts=[TextPart(content='Hello!')]),\ ] # Without the capability, the agent would run without its configured system prompt. # With the capability, the system prompt is reinjected at the head of the first request. result = agent.run_sync('Follow up', message_history=history) first_request = result.all_messages()[0] assert isinstance(first_request, ModelRequest) assert first_request.parts[0].content == 'You are a helpful assistant.' _(This example is complete, it can be run “as is”)_ The [UI adapters](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/) (AG-UI, Vercel AI) automatically add this capability with `replace_existing=True` in their `manage_system_prompt='server'` mode. Was this page helpful? Thanks for your feedback! --- # Retry Strategies | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#_top) Retry Strategies ================ Handle transient failures in tasks and evaluators with automatic retry logic. LLM-based systems can experience transient failures: * Rate limits * Network timeouts * Temporary API outages * Context length errors Pydantic Evals supports retry configuration for both: * **Task execution** - The function being evaluated * **Evaluator execution** - The evaluators themselves Basic Retry Configuration ------------------------- [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#basic-retry-configuration) Pass a retry configuration to `evaluate()` or `evaluate_sync()` using [Tenacity](https://tenacity.readthedocs.io/) parameters: from tenacity import stop_after_attempt from pydantic_evals import Case, Dataset def my_function(inputs: str) -> str: return f'Result: {inputs}' dataset = Dataset(name='basic_retry', cases=[Case(inputs='test')], evaluators=[]) report = dataset.evaluate_sync( task=my_function, retry_task={'stop': stop_after_attempt(3)}, retry_evaluators={'stop': stop_after_attempt(2)}, ) Retry Configuration Options --------------------------- [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#retry-configuration-options) Retry configurations use [Tenacity](https://tenacity.readthedocs.io/) and support the same options as Pydantic AI’s [`RetryConfig`](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.RetryConfig) : from tenacity import stop_after_attempt, wait_exponential from pydantic_evals import Case, Dataset def my_function(inputs: str) -> str: return f'Result: {inputs}' dataset = Dataset(name='retry_config', cases=[Case(inputs='test')], evaluators=[]) retry_config = { 'stop': stop_after_attempt(3), # Stop after 3 attempts 'wait': wait_exponential(multiplier=1, min=1, max=10), # Exponential backoff: 1s, 2s, 4s, 8s (capped at 10s) 'reraise': True, # Re-raise the original exception after exhausting retries } dataset.evaluate_sync( task=my_function, retry_task=retry_config, ) ### Common Parameters [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#common-parameters) The retry configuration accepts any parameters from the tenacity `retry` decorator. Common ones include: | Parameter | Type | Description | | --- | --- | --- | | `stop` | `StopBaseT` | Stop strategy (e.g., `stop_after_attempt(3)`, `stop_after_delay(60)`) | | `wait` | `WaitBaseT` | Wait strategy (e.g., `wait_exponential()`, `wait_fixed(2)`) | | `retry` | `RetryBaseT` | Retry condition (e.g., `retry_if_exception_type(TimeoutError)`) | | `reraise` | `bool` | Whether to reraise the original exception (default: `False`) | | `before_sleep` | `Callable` | Callback before sleeping between retries | See the [Tenacity documentation](https://tenacity.readthedocs.io/) for all available options. Task Retries ------------ [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#task-retries) Retry the task function when it fails: from tenacity import stop_after_attempt, wait_exponential from pydantic_evals import Case, Dataset async def call_llm(inputs: str) -> str: return f'LLM response to: {inputs}' async def flaky_llm_task(inputs: str) -> str: """This might hit rate limits or timeout.""" response = await call_llm(inputs) return response dataset = Dataset(name='task_retry', cases=[Case(inputs='test')]) report = dataset.evaluate_sync( task=flaky_llm_task, retry_task={ 'stop': stop_after_attempt(5), # Try up to 5 times 'wait': wait_exponential(multiplier=1, min=1, max=30), # Exponential backoff, capped at 30s 'reraise': True, }, ) ### When Task Retries Trigger [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#when-task-retries-trigger) Retries trigger when the task raises an exception: class RateLimitError(Exception): pass class ValidationError(Exception): pass async def call_api(inputs: str) -> str: return f'API response: {inputs}' async def my_task(inputs: str) -> str: try: return await call_api(inputs) except RateLimitError: # Will trigger retry raise except ValidationError: # Will also trigger retry raise ### Exponential Backoff [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#exponential-backoff) When using `wait_exponential()`, delays increase exponentially: Attempt 1: immediate Attempt 2: ~1s delay (multiplier * 2^0) Attempt 3: ~2s delay (multiplier * 2^1) Attempt 4: ~4s delay (multiplier * 2^2) Attempt 5: ~8s delay (multiplier * 2^3, capped at max) The actual delay depends on the `multiplier`, `min`, and `max` parameters passed to `wait_exponential()`. Evaluator Retries ----------------- [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#evaluator-retries) Retry evaluators when they fail: from tenacity import stop_after_attempt, wait_exponential from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import LLMJudge def my_task(inputs: str) -> str: return f'Result: {inputs}' dataset = Dataset( name='evaluator_retry', cases=[Case(inputs='test')], evaluators=[\ # LLMJudge might hit rate limits\ LLMJudge(rubric='Response is accurate'),\ ], ) report = dataset.evaluate_sync( task=my_task, retry_evaluators={ 'stop': stop_after_attempt(3), 'wait': wait_exponential(multiplier=1, min=0.5, max=10), 'reraise': True, }, ) ### When Evaluator Retries Trigger [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#when-evaluator-retries-trigger) Retries trigger when an evaluator raises an exception: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext async def external_api_call(output: str) -> bool: return len(output) > 0 @dataclass class APIEvaluator(Evaluator): async def evaluate(self, ctx: EvaluatorContext) -> bool: # If this raises an exception, retry logic will trigger result = await external_api_call(ctx.output) return result ### Evaluator Failures [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#evaluator-failures) If an evaluator fails after all retries, it’s recorded as an [`EvaluatorFailure`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorFailure) : from tenacity import stop_after_attempt from pydantic_evals import Case, Dataset def task(inputs: str) -> str: return f'Result: {inputs}' dataset = Dataset(name='evaluator_failures', cases=[Case(inputs='test')], evaluators=[]) report = dataset.evaluate_sync(task, retry_evaluators={'stop': stop_after_attempt(3)}) # Check for evaluator failures for case in report.cases: if case.evaluator_failures: for failure in case.evaluator_failures: print(f'Evaluator {failure.name} failed: {failure.error_message}') #> (No output - no evaluator failures in this case) View evaluator failures in reports: from pydantic_evals import Case, Dataset def task(inputs: str) -> str: return f'Result: {inputs}' dataset = Dataset(name='failure_report', cases=[Case(inputs='test')], evaluators=[]) report = dataset.evaluate_sync(task) report.print(include_evaluator_failures=True) """ Evaluation Summary: task ┏━━━━━━━━━━┳━━━━━━━━━━┓ ┃ Case ID ┃ Duration ┃ ┡━━━━━━━━━━╇━━━━━━━━━━┩ │ Case 1 │ 10ms │ ├──────────┼──────────┤ │ Averages │ 10ms │ └──────────┴──────────┘ """ #> #> ✅ case_0 ━━━━━━━━━━━━━━━━━━━━━━━━━━━ 100% (0/0) Combining Task and Evaluator Retries ------------------------------------ [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#combining-task-and-evaluator-retries) You can configure both independently: from tenacity import stop_after_attempt, wait_exponential from pydantic_evals import Case, Dataset def flaky_task(inputs: str) -> str: return f'Result: {inputs}' dataset = Dataset(name='combined_retry', cases=[Case(inputs='test')], evaluators=[]) report = dataset.evaluate_sync( task=flaky_task, retry_task={ 'stop': stop_after_attempt(5), # Retry task up to 5 times 'wait': wait_exponential(multiplier=1, min=1, max=30), 'reraise': True, }, retry_evaluators={ 'stop': stop_after_attempt(3), # Retry evaluators up to 3 times 'wait': wait_exponential(multiplier=1, min=0.5, max=10), 'reraise': True, }, ) Practical Examples ------------------ [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#practical-examples) ### Rate Limit Handling [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#rate-limit-handling) from tenacity import stop_after_attempt, wait_exponential from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import LLMJudge async def expensive_llm_call(inputs: str) -> str: return f'LLM response: {inputs}' async def llm_task(inputs: str) -> str: """Task that might hit rate limits.""" return await expensive_llm_call(inputs) dataset = Dataset( name='rate_limit_retry', cases=[Case(inputs='test')], evaluators=[\ LLMJudge(rubric='Quality check'), # Also might hit rate limits\ ], ) # Generous retries for rate limits report = dataset.evaluate_sync( task=llm_task, retry_task={ 'stop': stop_after_attempt(10), # Rate limits can take multiple retries 'wait': wait_exponential(multiplier=2, min=2, max=60), # Start at 2s, exponential up to 60s 'reraise': True, }, retry_evaluators={ 'stop': stop_after_attempt(5), 'wait': wait_exponential(multiplier=2, min=2, max=30), 'reraise': True, }, ) ### Network Timeout Handling [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#network-timeout-handling) import httpx from tenacity import stop_after_attempt, wait_exponential from pydantic_evals import Case, Dataset async def api_task(inputs: str) -> str: """Task that calls external API which might timeout.""" async with httpx.AsyncClient(timeout=10.0) as client: response = await client.post('https://api.example.com', json={'input': inputs}) return response.text dataset = Dataset(name='timeout_retry', cases=[Case(inputs='test')], evaluators=[]) # Quick retries for network issues report = dataset.evaluate_sync( task=api_task, retry_task={ 'stop': stop_after_attempt(4), # A few quick retries 'wait': wait_exponential(multiplier=0.5, min=0.5, max=5), # Fast retry, capped at 5s 'reraise': True, }, ) ### Context Length Handling [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#context-length-handling) from tenacity import stop_after_attempt from pydantic_evals import Case, Dataset class ContextLengthError(Exception): pass async def llm_call(inputs: str, max_tokens: int = 8000) -> str: return f'LLM response: {inputs[:100]}' async def smart_llm_task(inputs: str) -> str: """Task that might exceed context length.""" try: return await llm_call(inputs, max_tokens=8000) except ContextLengthError: # Retry with shorter context truncated_inputs = inputs[:4000] return await llm_call(truncated_inputs, max_tokens=4000) dataset = Dataset(name='context_length', cases=[Case(inputs='test')], evaluators=[]) # Don't retry context length errors (handle in task) report = dataset.evaluate_sync( task=smart_llm_task, retry_task={'stop': stop_after_attempt(1)}, # No retries, we handle it ) Retry vs Error Handling ----------------------- [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#retry-vs-error-handling) **Use retries for:** * Transient failures (rate limits, timeouts) * Network issues * Temporary service outages * Recoverable errors **Use error handling for:** * Validation errors * Logic errors * Permanent failures * Expected error conditions class RateLimitError(Exception): pass async def llm_call(inputs: str) -> str: return f'LLM response: {inputs}' def is_valid(result: str) -> bool: return len(result) > 0 async def smart_task(inputs: str) -> str: """Handle expected errors, let retries handle transient failures.""" try: result = await llm_call(inputs) # Validate output (don't retry validation errors) if not is_valid(result): return 'ERROR: Invalid output format' return result except RateLimitError: # Let retry logic handle this raise except ValueError as e: # Don't retry - this is a permanent error return f'ERROR: {e}' Troubleshooting --------------- [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#troubleshooting) ### ”Still failing after retries” [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#still-failing-after-retries) Increase retry attempts or check if error is retriable: import logging from tenacity import stop_after_attempt from pydantic_evals import Case, Dataset def task(inputs: str) -> str: return f'Result: {inputs}' # Add logging to see what's failing logging.basicConfig(level=logging.DEBUG) dataset = Dataset(name='troubleshooting', cases=[Case(inputs='test')], evaluators=[]) # Tenacity logs retry attempts report = dataset.evaluate_sync(task, retry_task={'stop': stop_after_attempt(5)}) ### “Evaluations taking too long” [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#evaluations-taking-too-long) Reduce retry attempts or wait times: from tenacity import stop_after_attempt, wait_exponential # Faster retries retry_config = { 'stop': stop_after_attempt(3), # Fewer attempts 'wait': wait_exponential(multiplier=0.1, min=0.1, max=2), # Quick retries, capped at 2s 'reraise': True, } ### “Hitting rate limits despite retries” [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#hitting-rate-limits-despite-retries) Increase delays or use `max_concurrency`: from tenacity import stop_after_attempt, wait_exponential from pydantic_evals import Case, Dataset def task(inputs: str) -> str: return f'Result: {inputs}' dataset = Dataset(name='rate_limit_config', cases=[Case(inputs='test')], evaluators=[]) # Longer delays retry_config = { 'stop': stop_after_attempt(5), 'wait': wait_exponential(multiplier=5, min=5, max=60), # Start at 5s, exponential up to 60s 'reraise': True, } # Also reduce concurrency report = dataset.evaluate_sync( task=task, retry_task=retry_config, max_concurrency=2, # Only 2 concurrent tasks ) Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/#next-steps) * **[Concurrency & Performance](https://pydantic.dev/docs/ai/evals/how-to/concurrency/) ** - Optimize evaluation performance * **[Logfire Integration](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/) ** - View retries in Logfire Was this page helpful? Thanks for your feedback! --- # Agentic Evaluators | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/evaluators/agentic/#_top) Agentic Evaluators ================== Deterministic, span-based evaluators that grade an agent’s _trajectory_ — the sequence and arguments of tool calls — rather than just its final output. Agentic evaluators answer a class of “did the agent do the right thing?” questions that pure input/output checks can’t: * **Tool coverage** — did the agent call the specific tools it was supposed to? ([`ToolCorrectness`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.ToolCorrectness) ) * **Trajectory shape** — did it call them in the right order, or at least use the right set? ([`TrajectoryMatch`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.TrajectoryMatch) ) * **Argument quality** — did the tool receive the expected inputs? ([`ArgumentCorrectness`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.ArgumentCorrectness) ) * **Budget discipline** — did the agent finish within a tool-call and/or model-request budget? ([`MaxToolCalls`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.MaxToolCalls) , [`MaxModelRequests`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.MaxModelRequests) ) They are all deterministic, never call an LLM, and are cheap enough to run on every case in every experiment. ToolCorrectness --------------- [](https://pydantic.dev/docs/ai/evals/evaluators/agentic/#toolcorrectness) Assert that the agent called a specific **multiset** of tools. Repeated names require repeated calls. from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import ToolCorrectness dataset = Dataset( name='rag_agent', cases=[Case(inputs='Summarize the latest papers on X')], evaluators=[\ ToolCorrectness(\ expected_tools=['search', 'rerank', 'generate'],\ ),\ ], ) **Parameters:** * `expected_tools` (`list[str]`): Tool names the agent is expected to call. Order doesn’t matter; duplicates are significant — `['search', 'search']` requires two `search` calls. * `allow_extra` (`bool`, default `False`): By default, any tool call not listed in `expected_tools` fails the check. Set to `True` to only require that the expected tools were called, permitting extras. * `include_failed` (`bool`, default `False`): Whether to count tool-call attempts that ended in an error. * `evaluation_name` (`str | None`): Custom name in reports. **Returns:** [`EvaluationReason`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluationReason) with a `bool` value. The `reason` names missing and unexpected tools. TrajectoryMatch --------------- [](https://pydantic.dev/docs/ai/evals/evaluators/agentic/#trajectorymatch) Compare the actual ordered list of tool names to an expected one, using one of three modes. from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import TrajectoryMatch dataset = Dataset( name='ordered_tools', cases=[Case(inputs='Process and file this request')], evaluators=[\ TrajectoryMatch(\ expected_trajectory=['validate', 'enrich', 'submit'],\ order='in_order',\ ),\ ], ) **Parameters:** * `expected_trajectory` (`list[str]`): Expected ordered list of tool names. * `order` (`Literal['exact', 'in_order', 'any_order']`, default `'in_order'`): * `'exact'` — `1.0` iff the sequences are equal, else `0.0`. * `'in_order'` — F1 computed from the longest common subsequence (LCS). Precision = `LCS / len(actual)`, recall = `LCS / len(expected)`. Allows extra calls interleaved with the expected order, but they reduce precision. * `'any_order'` — F1 computed from the multiset intersection. Precision = `overlap / len(actual)`, recall = `overlap / len(expected)`. Order is ignored, but extra and missing calls both reduce the score. * `include_failed` (`bool`, default `False`): Whether the trajectory includes tool-call attempts that ended in an error. * `evaluation_name` (`str | None`): Custom name in reports. **Returns:** [`EvaluationReason`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluationReason) with a `float` value in `[0.0, 1.0]`. For the F1-based modes, the reason text spells out the overlap, precision, recall, and F1 so the score is reproducible from the mismatch. For example, if `expected = ['a', 'b', 'c']` and the agent called `['a', 'x', 'b']`, the LCS is `['a', 'b']` (length 2), giving precision `2/3`, recall `2/3`, and F1 `≈ 0.667`. If both the expected and actual trajectories are empty, all modes score `1.0`; if only one of them is empty, all modes score `0.0`. ArgumentCorrectness ------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/agentic/#argumentcorrectness) Check that a specific tool call received particular arguments. from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import ArgumentCorrectness dataset = Dataset( name='support_agent', cases=[Case(inputs='Refund order 12345')], evaluators=[\ ArgumentCorrectness(\ tool_name='issue_refund',\ expected_arguments={'order_id': '12345'},\ match_mode='subset',\ occurrence='first',\ ),\ ], ) **Parameters:** * `tool_name` (`str`): The tool to inspect. * `expected_arguments` (`dict[str, Any]`): Expected argument keys/values. * `match_mode` (`Literal['exact', 'subset']`, default `'subset'`): * `'subset'` — every expected key/value is present in the actual arguments. Note that this applies only to top-level keys: an expected _value_ (including a nested dict) must compare equal to the actual value in full. * `'exact'` — deep equality; unexpected keys also fail. * `occurrence` (`Literal['first', 'last'] | int`, default `'first'`): Which invocation to inspect if the tool is called multiple times. Integer indexes are 0-based. * `include_failed` (`bool`, default `False`): Whether tool-call attempts that ended in an error are considered. When `True`, each attempt counts as a separate occurrence. * `evaluation_name` (`str | None`): Custom name in reports. **Returns:** [`EvaluationReason`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluationReason) with a `bool` value. **Graceful degradation:** this evaluator doesn’t crash when arguments aren’t available — for example, when the agent was instrumented with `include_content=False`, the evaluator returns `False` with a reason explaining the situation so your reports still make sense. MaxToolCalls and MaxModelRequests --------------------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/agentic/#maxtoolcalls-and-maxmodelrequests) Assert that the agent stayed within a tool-call and/or model-request budget. These follow the same shape as [`MaxDuration`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.MaxDuration) : one budget per evaluator, each reported as its own boolean assertion. from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import MaxModelRequests, MaxToolCalls dataset = Dataset( name='budget_aware', cases=[Case(inputs='Draft a short reply')], evaluators=[\ MaxToolCalls(max_calls=5),\ MaxModelRequests(max_requests=3),\ ], ) **Parameters:** * `MaxToolCalls`: `max_calls` (`int`) — maximum allowed locally-executed tool calls. `include_failed` (`bool`, default `True`) controls whether attempts that ended in an error count against the budget (by default they do — they still consumed time and tokens). * `MaxModelRequests`: `max_requests` (`int`) — maximum allowed model (chat) requests. Prefers the `requests` value from `ctx.metrics` when available, otherwise counts LLM request spans directly (both use the same criteria). * Both accept `evaluation_name` (`str | None`) to customize the name in reports — useful when the same budget check appears at both the dataset and case level. **Returns:** [`EvaluationReason`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluationReason) with a `bool` value. The `reason` includes the observed count and the budget. Recipes ------- [](https://pydantic.dev/docs/ai/evals/evaluators/agentic/#recipes) ### RAG agent [](https://pydantic.dev/docs/ai/evals/evaluators/agentic/#rag-agent) Check that the retrieval pipeline runs _search → rerank → generate_, with no unexpected tool calls. from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import ToolCorrectness, TrajectoryMatch dataset = Dataset( name='rag_pipeline', cases=[Case(inputs='Find papers on in-context learning')], evaluators=[\ ToolCorrectness(\ expected_tools=['search', 'rerank', 'generate'],\ ),\ TrajectoryMatch(\ expected_trajectory=['search', 'rerank', 'generate'],\ order='exact',\ ),\ ], ) ### Multi-tool agent where order matters [](https://pydantic.dev/docs/ai/evals/evaluators/agentic/#multi-tool-agent-where-order-matters) Allow occasional retries, but require the main steps to happen in order. from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import TrajectoryMatch dataset = Dataset( name='ordered_with_slack', cases=[Case(inputs='Process shipment 99')], evaluators=[\ TrajectoryMatch(\ expected_trajectory=['validate', 'enrich', 'submit'],\ order='in_order', # F1-based: extra calls only reduce precision, order must be preserved\ ),\ ], ) ### Support agent with `ArgumentCorrectness` and budget checks [](https://pydantic.dev/docs/ai/evals/evaluators/agentic/#support-agent-with-argumentcorrectness-and-budget-checks) Verify that the right action was taken with the right inputs — within a reasonable number of steps. from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import ( ArgumentCorrectness, MaxModelRequests, MaxToolCalls, ) dataset = Dataset( name='refund_handling', cases=[\ Case(\ name='valid_refund',\ inputs={'query': 'Refund my order', 'order_id': '12345'},\ evaluators=[\ ArgumentCorrectness(\ tool_name='issue_refund',\ expected_arguments={'order_id': '12345'},\ ),\ ],\ ),\ ], evaluators=[\ MaxToolCalls(max_calls=4),\ MaxModelRequests(max_requests=2),\ ], ) ### Task completion judged with the tool-call trajectory [](https://pydantic.dev/docs/ai/evals/evaluators/agentic/#task-completion-judged-with-the-tool-call-trajectory) For tasks where deterministic checks aren’t enough, you can have an LLM judge the task outcome together with the tool-call trajectory. [`LLMJudge`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.LLMJudge) only sees the case inputs, output, and expected output — not other evaluators’ results or the span tree — so to give the judge visibility into _how_ the agent got there, write a small custom evaluator that extracts the trajectory from the span tree and passes it to [`judge_input_output`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.llm_as_a_judge.judge_input_output) directly: from dataclasses import dataclass from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import EvaluationReason, Evaluator, EvaluatorContext from pydantic_evals.evaluators.llm_as_a_judge import judge_input_output from pydantic_evals.otel import SpanTreeRecordingError @dataclass class TrajectoryJudge(Evaluator): rubric: str async def evaluate(self, ctx: EvaluatorContext) -> EvaluationReason: try: span_tree = ctx.span_tree except SpanTreeRecordingError: # Degrade gracefully, like the built-in evaluators on this page. return EvaluationReason(value=False, reason='No span tree available.') # Build a plain-text trajectory summary, mirroring what the built-in # evaluators count as a tool call by default: tool spans are named # 'running tool' (v2) or 'execute_tool {name}' (v3+); deferred calls # never ran; output functions share the tool span shape but aren't # tool calls; and failed attempts (status 'error') are dropped, like # the built-in evaluators' `include_failed=False` default. tool_names = [\ node.attributes['gen_ai.tool.name']\ for node in span_tree\ if 'gen_ai.tool.name' in node.attributes\ and 'pydantic_ai.tool.deferral.name' not in node.attributes\ and node.status != 'error'\ and (node.name == 'running tool' or node.name.startswith('execute_tool '))\ and not str(node.attributes.get('logfire.msg', '')).startswith('running output function:')\ ] trajectory = ', '.join(str(n) for n in tool_names) or '(none)' grading_output = await judge_input_output( {'query': ctx.inputs, 'tool_trajectory': trajectory}, ctx.output, self.rubric, ) return EvaluationReason(value=grading_output.pass_, reason=grading_output.reason) dataset = Dataset( name='task_completion', cases=[Case(inputs='Resolve ticket 42')], evaluators=[\ TrajectoryJudge(\ rubric=(\ 'The agent completed the task correctly, and the tool trajectory '\ 'included in the input is reasonable for the given query.'\ ),\ ),\ ], ) This pattern keeps the deterministic checks above cheap and reproducible, and reserves the qualitative, open-ended judgement for the LLM — with the trajectory explicitly included in what the judge sees. Next steps ---------- [](https://pydantic.dev/docs/ai/evals/evaluators/agentic/#next-steps) * [Span-Based Evaluation](https://pydantic.dev/docs/ai/evals/evaluators/span-based/) — low-level span queries via `HasMatchingSpan` and `SpanQuery` * [Custom Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/custom/) — write your own evaluation logic * [Built-in Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/built-in/) — complete reference of other evaluator types Was this page helpful? Thanks for your feedback! --- # Extensibility | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/guides/extensibility/#_top) Extensibility ============= Pydantic AI is designed to be extended. [Capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) are the primary extension point — they bundle tools, lifecycle hooks, instructions, and model settings into reusable units that can be shared across agents, packaged as libraries, and loaded from [spec files](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) . Beyond capabilities, Pydantic AI provides several other extension mechanisms for specialized needs. Capabilities ------------ [](https://pydantic.dev/docs/ai/guides/extensibility/#capabilities) Capabilities are the recommended way to extend Pydantic AI. They are useful for: * **Teams** building reusable internal agent components (guardrails, audit logging, authentication) * **Package authors** shipping extensions that work across models and agents * **Community contributors** sharing solutions to common problems See [Capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) for using and building capabilities, and [Hooks](https://pydantic.dev/docs/ai/core-concepts/hooks/) for the lightweight decorator-based approach. Publishing capability packages ------------------------------ [](https://pydantic.dev/docs/ai/guides/extensibility/#publishing-capability-packages) To make a capability installable and usable in [agent specs](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) : 1. **Implement [`get_serialization_name()`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability.get_serialization_name) ** — defaults to the class name. Return `None` to opt out of spec support. 2. **Implement [`from_spec()`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability.from_spec) ** — defaults to `cls(*args, **kwargs)`. Override when your constructor takes non-serializable types. 3. **Package naming** — use the `pydantic-ai-` prefix (e.g. `pydantic-ai-guardrails`) so users can find your package. 4. **Registration** — users pass custom capability types via `custom_capability_types` on [`Agent.from_spec`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.from_spec) or [`Agent.from_file`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.from_file) . from pydantic_ai import Agent from my_package import MyCapability agent = Agent.from_file('agent.yaml', custom_capability_types=[MyCapability]) See [Custom capabilities in specs](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#custom-capabilities-in-specs) for implementation details. Pydantic AI Harness ------------------- [](https://pydantic.dev/docs/ai/guides/extensibility/#pydantic-ai-harness) [**Pydantic AI Harness**](https://pydantic.dev/docs/ai/harness/) is the official capability library for Pydantic AI — standalone capabilities like memory, guardrails, and context management live there rather than in core. See [What goes where?](https://pydantic.dev/docs/ai/harness/#what-goes-where) for the full breakdown, or jump to the [capability matrix](https://github.com/pydantic/pydantic-ai-harness#capability-matrix) . Third-party ecosystem --------------------- [](https://pydantic.dev/docs/ai/guides/extensibility/#third-party-ecosystem) ### Capabilities [](https://pydantic.dev/docs/ai/guides/extensibility/#capabilities-1) [Capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) are the recommended extension mechanism for packages that need to bundle tools with hooks, instructions, or model settings. See [Third-party capabilities](https://pydantic.dev/docs/ai/capabilities/third-party/) for community packages. ### Toolsets [](https://pydantic.dev/docs/ai/guides/extensibility/#toolsets) Many third-party extensions are available as [toolsets](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) , which can also be wrapped as [capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) to take advantage of hooks, instructions, and model settings. See [Third-party toolsets](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/#third-party-toolsets) for the full list. Other extension points ---------------------- [](https://pydantic.dev/docs/ai/guides/extensibility/#other-extension-points) ### Custom toolsets [](https://pydantic.dev/docs/ai/guides/extensibility/#custom-toolsets) For specialized tool execution needs (custom transport, tool filtering, execution wrapping), implement [`AbstractToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AbstractToolset) or subclass [`WrapperToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.WrapperToolset) : * [`AbstractToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AbstractToolset) — full control over tool definitions and execution * [`WrapperToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.WrapperToolset) — delegates to a wrapped toolset, override specific methods See [Building a Custom Toolset](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/#building-a-custom-toolset) for details. ### Custom models [](https://pydantic.dev/docs/ai/guides/extensibility/#custom-models) For connecting to model providers not yet supported by Pydantic AI, implement [`Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) : * [`Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) — the base interface for model implementations * [`WrapperModel`](https://pydantic.dev/docs/ai/api/models/wrapper/#pydantic_ai.models.wrapper.WrapperModel) — delegates to a wrapped model, useful for adding instrumentation or transformations See [Custom Models](https://pydantic.dev/docs/ai/models/overview/#custom-models) for details. ### Custom agents [](https://pydantic.dev/docs/ai/guides/extensibility/#custom-agents) For custom agent behavior, subclass [`AbstractAgent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent) or [`WrapperAgent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.WrapperAgent) : * [`AbstractAgent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent) — the base interface for agent implementations, providing `run`, `run_sync`, and `run_stream` * [`WrapperAgent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.WrapperAgent) — delegates to a wrapped agent, useful for adding pre/post-processing or context management Was this page helpful? Thanks for your feedback! --- # Conversation Search | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/conversation-search/#_top) Conversation Search =================== `ConversationSearch` gives the model a `search_conversation_history` tool that BM25-ranks the history a `StepPersistence` capability already persists — earlier turns that compaction dropped from the live context, and past runs in the same store. [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/conversation_search/) > The API may change between releases. Where practical, breaking changes ship with a deprecation warning. The problem ----------- [](https://pydantic.dev/docs/ai/harness/conversation-search/#the-problem) Compaction capabilities (`SlidingWindowCompaction`, `SummarizingCompaction`, …) narrow the live history so it fits the context window. `SummarizingCompaction` persists its edits: once a prefix is replaced by a summary, the originals are gone from the run’s `message_history` on the next turn. The model can no longer recall an exact file path, a decision, or a value stated earlier — only the summary’s paraphrase of it. And nothing at all from previous runs is reachable, however well persisted. The solution ------------ [](https://pydantic.dev/docs/ai/harness/conversation-search/#the-solution) `ConversationSearch` persists nothing itself. It reads whatever a persistence capability already stores, through a `HistorySource`, and exposes one tool, `search_conversation_history`, that BM25-ranks that history so the model can pull exact details back into context on demand. The shipped source, `SnapshotHistorySource`, reads the snapshots `StepPersistence` writes: pair the two capabilities on a shared store instance and recall works with no extra write path, no ordering constraints, and no hook coordination. from pydantic_ai import Agent from pydantic_ai_harness.compaction import SlidingWindowCompaction from pydantic_ai_harness.conversation_search import ConversationSearch, SnapshotHistorySource from pydantic_ai_harness.step_persistence import SqliteStepStore, StepPersistence store = SqliteStepStore(database='sessions.db') agent = Agent( 'openai:gpt-5', capabilities=[\ StepPersistence(store=store),\ ConversationSearch(SnapshotHistorySource(store)),\ SlidingWindowCompaction(max_messages=40),\ ], ) * Ranking is BM25 (the algorithm behind Lucene/Elasticsearch), implemented in pure Python — no new dependencies. Rare terms and exact matches score higher; multi-word queries score each word independently. * Results carry provenance (`run: ... | conversation: ...`), and the tool’s optional `run_id` argument scopes a search to one run — so a run referenced elsewhere (for example by a compaction receipt’s transcript handle) is directly resolvable. * The search reads the store lazily at call time, so it always sees everything persisted so far, including earlier steps of the current run. How recovery works ------------------ [](https://pydantic.dev/docs/ai/harness/conversation-search/#how-recovery-works) `StepPersistence` saves a full-history snapshot at every step boundary. A compaction strategy that persists its edits (like `SummarizingCompaction`) carries those edits into _later_ snapshots — but the earlier snapshots of the same run were taken while the originals were still live. `SnapshotHistorySource` unions each run’s snapshots in write order, skips derived summary artifacts, and removes the overlap between the accumulated history’s suffix and each snapshot’s prefix. This recovers the originals plus everything compaction never touched while preserving repeated messages at distinct sequence positions — as far back as the store still retains those pre-compaction snapshots (see Limitations). Only `complete` snapshots contribute: `interrupted` captures (unsettled tool work, synthesized tool returns) are excluded by the stores’ default read gate. Overlap matching keys off a content hash of each serialized message, not object identity: consecutive snapshots re-serialize the same growing history, and durable executors (Temporal, DBOS) re-instantiate messages between steps. `HistorySource` is deliberately substrate-neutral (“enumerate runs, yield each run’s durable message record”): a persistence substrate that keeps an append-only entry log can implement it directly by replay, replacing the snapshot-union adapter without touching the search layer. Scope ----- [](https://pydantic.dev/docs/ai/harness/conversation-search/#scope) A store is a single corpus. With the default `scope='all'`, one `search_conversation_history` call ranks every run the source enumerates and can return verbatim excerpts from any of them — which is the cross-session recall [#124](https://github.com/pydantic/pydantic-ai-harness/issues/124) asks for, and the right default when the store holds one principal’s history. If several users or tenants share a store, set `scope='conversation'`: agent = Agent( 'openai:gpt-5', capabilities=[\ StepPersistence(store=store),\ ConversationSearch(SnapshotHistorySource(store), scope='conversation'),\ ], ) async def ask(question: str, user_id: str) -> str: result = await agent.run(question, conversation_id=user_id) return result.output The corpus is then restricted to runs whose `conversation_id` matches the calling run’s, the tool’s own description tells the model the restriction applies, and the tool’s `run_id` argument cannot reach past it — an out-of-scope run reports the same “no persisted history” answer as a run that does not exist. A run with no `conversation_id` searches nothing under this scope and the tool says why. Matching on “conversation id is unset” would pool every unlabelled run in the store into one corpus, which is the exposure the scope exists to prevent, so it fails closed instead. Pass `conversation_id=` to `Agent.run(...)` (the same value `StepPersistence` records on the run). Scoping is applied to the `RunRecord`s a `HistorySource` returns, so a custom source must populate `conversation_id` on them for `scope='conversation'` to match anything. Key options ----------- [](https://pydantic.dev/docs/ai/harness/conversation-search/#key-options) | Option | Default | Purpose | | --- | --- | --- | | `source` | (required) | Where the corpus comes from. Use `SnapshotHistorySource(store)` over the store `StepPersistence` writes to. | | `scope` | `'all'` | How much of the store one search may reach: `'all'`, or `'conversation'` to restrict it to the calling run’s `conversation_id`. See Scope. | | `max_matches` | `10` | Maximum matching excerpts the search tool returns. | | `context_lines` | `5` | Lines shown around each match (within the match’s run). | | `bm25_k1` | `1.5` | BM25 term-frequency saturation. This capability’s default; Lucene’s `BM25Similarity` uses `1.2`. | | `bm25_b` | `0.75` | BM25 length normalization (Lucene default). | | `add_instructions` | `True` | Emit a short note telling the model the recall tool exists. | | `tool_id` | `conversation-search` | Toolset id for the search tool. | Limitations ----------- [](https://pydantic.dev/docs/ai/harness/conversation-search/#limitations) * Search only reaches what was persisted: history inherited from runs that never ran with `StepPersistence` (for example a long `message_history` passed in from an unpersisted session) cannot be recovered if compaction drops it before the first snapshot. * Recovery of compaction-dropped originals depends on the pre-compaction snapshots still being retained. A store with bounded snapshot retention (for example a per-run snapshot cap) can prune the early snapshots that held those originals; a search then returns only what the surviving snapshots still carry, degrading to a partial result rather than erroring. Retain full snapshot history for the run if complete recovery matters. * The corpus is rebuilt on each tool call by reading every run’s snapshots. Snapshot storage is cumulative (each snapshot re-serializes the growing history), so large stores make each search proportionally more expensive. A persistent index (SQLite FTS5, tracked in [#124](https://github.com/pydantic/pydantic-ai-harness/issues/124) ) is the scaling path. * Reading snapshots restores externalized media (large binary payloads) even though the text index never uses it; stores with remote media backends pay that fetch cost per search. Further reading --------------- [](https://pydantic.dev/docs/ai/harness/conversation-search/#further-reading) * [Pydantic AI capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) * [Step Persistence](https://pydantic.dev/docs/ai/harness/step-persistence/) — the substrate this capability reads * [Compaction](https://pydantic.dev/docs/ai/harness/compaction/) — the capabilities whose drops this one recovers from API reference ------------- [](https://pydantic.dev/docs/ai/harness/conversation-search/#api-reference) ConversationSearch ------------------ [](https://pydantic.dev/docs/ai/harness/conversation-search/#pydantic_ai_harness.conversation_search.ConversationSearch) **Bases:** `AbstractCapability[AgentDepsT]` Search persisted conversation history with a dependency-free BM25 tool. This capability persists nothing itself: it reads whatever history a persistence capability already stores, through a `HistorySource`. Pair it with `StepPersistence` sharing the same store, and the model can recall what compaction dropped from the live context as well as anything from past runs: from pydantic_ai import Agent from pydantic_ai_harness.compaction import SlidingWindowCompaction from pydantic_ai_harness.conversation_search import ConversationSearch, SnapshotHistorySource from pydantic_ai_harness.step_persistence import SqliteStepStore, StepPersistence store = SqliteStepStore(database='sessions.db') agent = Agent( 'openai:gpt-5', capabilities=[\ StepPersistence(store=store),\ ConversationSearch(SnapshotHistorySource(store)),\ SlidingWindowCompaction(max_messages=40),\ ], ) Some compaction strategies persist their edits into the run’s durable message history (`SummarizingCompaction` replaces summarized prefixes for good; a `SlidingWindowCompaction` trim only narrows what each request sends). Either way, `StepPersistence` snapshots each step boundary before the next compaction runs, so the union of a run’s snapshots still holds the originals — `SnapshotHistorySource` recovers them. No ordering or hook coordination between the capabilities is required; the search tool reads the store lazily at call time. ### Attributes [](https://pydantic.dev/docs/ai/harness/conversation-search/#attributes) #### add\_instructions [](https://pydantic.dev/docs/ai/harness/conversation-search/#pydantic_ai_harness.conversation_search.ConversationSearch.add_instructions) Emit a short instruction telling the model the recall tool exists. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `True` #### bm25\_b [](https://pydantic.dev/docs/ai/harness/conversation-search/#pydantic_ai_harness.conversation_search.ConversationSearch.bm25_b) BM25 length-normalization, between `0.0` and `1.0` (Lucene/Elasticsearch default). **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) **Default:** `0.75` #### bm25\_k1 [](https://pydantic.dev/docs/ai/harness/conversation-search/#pydantic_ai_harness.conversation_search.ConversationSearch.bm25_k1) BM25 term-frequency saturation, non-negative. This capability’s default; Lucene’s `BM25Similarity` uses `1.2`. **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) **Default:** `1.5` #### context\_lines [](https://pydantic.dev/docs/ai/harness/conversation-search/#pydantic_ai_harness.conversation_search.ConversationSearch.context_lines) Number of surrounding lines shown around each search match. Must be non-negative. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `5` #### max\_matches [](https://pydantic.dev/docs/ai/harness/conversation-search/#pydantic_ai_harness.conversation_search.ConversationSearch.max_matches) Maximum number of matching excerpts the search tool returns. Must be non-negative. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `10` #### scope [](https://pydantic.dev/docs/ai/harness/conversation-search/#pydantic_ai_harness.conversation_search.ConversationSearch.scope) How much of the store one search may reach. `all` searches every run the source enumerates. `conversation` restricts the corpus to runs whose `conversation_id` matches the calling run’s, which a store shared across users or tenants needs: with `all`, any run reading that store can retrieve verbatim excerpts from every other conversation in it. Under `conversation`, a run with no `conversation_id` searches nothing and the tool says so, rather than falling back to every other unlabelled run. **Type:** `SearchScope` **Default:** `'all'` #### source [](https://pydantic.dev/docs/ai/harness/conversation-search/#pydantic_ai_harness.conversation_search.ConversationSearch.source) Where the search corpus comes from. Use `SnapshotHistorySource` over the store a `StepPersistence` capability writes to. **Type:** `HistorySource` #### tool\_id [](https://pydantic.dev/docs/ai/harness/conversation-search/#pydantic_ai_harness.conversation_search.ConversationSearch.tool_id) Toolset id for the `search_conversation_history` tool. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'conversation-search'` ### Methods [](https://pydantic.dev/docs/ai/harness/conversation-search/#methods) #### get\_instructions [](https://pydantic.dev/docs/ai/harness/conversation-search/#pydantic_ai_harness.conversation_search.ConversationSearch.get_instructions) def get_instructions() -> AgentInstructions[AgentDepsT] | None Tell the model the recall tool exists, unless `add_instructions` is false. ##### Returns [](https://pydantic.dev/docs/ai/harness/conversation-search/#returns) `AgentInstructions`\[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) #### get\_toolset [](https://pydantic.dev/docs/ai/harness/conversation-search/#pydantic_ai_harness.conversation_search.ConversationSearch.get_toolset) def get_toolset() -> AgentToolset[AgentDepsT] | None Provide the `search_conversation_history` tool over the source. ##### Returns [](https://pydantic.dev/docs/ai/harness/conversation-search/#returns-1) [`AgentToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AgentToolset) \[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) Was this page helpful? Thanks for your feedback! --- # Overview | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/integrations/ui/overview/#_top) Overview ======== If you’re building a chat app or other interactive frontend for an AI agent, your backend will need to receive agent run input (like a chat message or complete [message history](https://pydantic.dev/docs/ai/core-concepts/message-history/) ) from the frontend, and will need to stream the [agent’s events](https://pydantic.dev/docs/ai/core-concepts/agent/#streaming-all-events) (like text, thinking, and tool calls) to the frontend so that the user knows what’s happening in real time. While your frontend could use Pydantic AI’s [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) and [`AgentStreamEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.AgentStreamEvent) directly, you’ll typically want to use a UI event stream protocol that’s natively supported by your frontend framework. Pydantic AI natively supports two UI event stream protocols: * [Agent-User Interaction (AG-UI) Protocol](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/) * [Vercel AI Data Stream Protocol](https://pydantic.dev/docs/ai/integrations/ui/vercel-ai/) These integrations are implemented as subclasses of the abstract [`UIAdapter`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter) class, so they also serve as a reference for integrating with other UI event stream protocols. Usage ----- [](https://pydantic.dev/docs/ai/integrations/ui/overview/#usage) The protocol-specific [`UIAdapter`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter) subclass (i.e. [`AGUIAdapter`](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter) or [`VercelAIAdapter`](https://pydantic.dev/docs/ai/api/ui/vercel_ai/#pydantic_ai.ui.vercel_ai.VercelAIAdapter) ) is responsible for transforming agent run input received from the frontend into arguments for [`Agent.run_stream_events()`](https://pydantic.dev/docs/ai/core-concepts/agent/#running-agents) , running the agent, and then transforming Pydantic AI events into protocol-specific events. The event stream transformation is handled by a protocol-specific [`UIEventStream`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIEventStream) subclass, but you typically won’t use this directly. If you’re using a Starlette-based web framework like FastAPI, you can use the [`UIAdapter.dispatch_request()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.dispatch_request) class method from an endpoint function to directly handle a request and return a streaming response of protocol-specific events. This is demonstrated in the next section. If you’re using a web framework not based on Starlette (e.g. Django or Flask) or need fine-grained control over the input or output, you can create a `UIAdapter` instance and directly use its methods. This is demonstrated in “Advanced Usage” section below. ### Usage with Starlette/FastAPI [](https://pydantic.dev/docs/ai/integrations/ui/overview/#usage-with-starlettefastapi) Besides the request, [`UIAdapter.dispatch_request()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.dispatch_request) takes the agent, the same optional arguments as [`Agent.run_stream_events()`](https://pydantic.dev/docs/ai/core-concepts/agent/#running-agents) , an optional `on_complete` callback for successful runs, and an optional `on_cancel` callback that receives [`RunCancelled`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled) for [first-party cancelled](https://pydantic.dev/docs/ai/core-concepts/agent/#cancelling-a-run) runs (a client disconnect is an external cancellation and does not trigger it). Both callbacks can optionally yield additional protocol-specific events. dispatch\_request.py from fastapi import FastAPI from starlette.requests import Request from starlette.responses import Response from pydantic_ai import Agent from pydantic_ai.ui.vercel_ai import VercelAIAdapter agent = Agent('openai:gpt-5.2') app = FastAPI() @app.post('/chat') async def chat(request: Request) -> Response: return await VercelAIAdapter.dispatch_request(request, agent=agent) ### Advanced Usage [](https://pydantic.dev/docs/ai/integrations/ui/overview/#advanced-usage) If you’re using a web framework not based on Starlette (e.g. Django or Flask) or need fine-grained control over the input or output, you can create a `UIAdapter` instance and directly use its methods, which can be chained to accomplish the same thing as the `UIAdapter.dispatch_request()` class method shown above: 1. The [`UIAdapter.build_run_input()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.build_run_input) class method takes the request body as bytes and returns a protocol-specific run input object, which you can then pass to the [`UIAdapter()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter) constructor along with the agent. * You can also use the [`UIAdapter.from_request()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.from_request) class method to build an adapter directly from a Starlette/FastAPI request. 2. The [`UIAdapter.run_stream()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.run_stream) method runs the agent and returns a stream of protocol-specific events. It supports the same optional arguments as [`Agent.run_stream_events()`](https://pydantic.dev/docs/ai/core-concepts/agent/#running-agents) , including the `on_complete` and `on_cancel` callbacks. * You can also use [`UIAdapter.run_stream_native()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.run_stream_native) to run the agent and return a stream of Pydantic AI events instead, which can then be transformed into protocol-specific events using [`UIAdapter.transform_stream()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.transform_stream) . 3. The [`UIAdapter.encode_stream()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.encode_stream) method encodes the stream of protocol-specific events as SSE (HTTP Server-Sent Events) strings, which you can then return as a streaming response. * You can also use [`UIAdapter.streaming_response()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.streaming_response) to generate a Starlette/FastAPI streaming response directly from the protocol-specific event stream returned by `run_stream()`. run\_stream.py import json from http import HTTPStatus from fastapi import FastAPI from fastapi.requests import Request from fastapi.responses import Response, StreamingResponse from pydantic import ValidationError from pydantic_ai import Agent from pydantic_ai.ui import SSE_CONTENT_TYPE from pydantic_ai.ui.vercel_ai import VercelAIAdapter agent = Agent('openai:gpt-5.2') app = FastAPI() @app.post('/chat') async def chat(request: Request) -> Response: accept = request.headers.get('accept', SSE_CONTENT_TYPE) try: run_input = VercelAIAdapter.build_run_input(await request.body()) except ValidationError as e: return Response( content=json.dumps(e.json()), media_type='application/json', status_code=HTTPStatus.UNPROCESSABLE_ENTITY, ) adapter = VercelAIAdapter(agent=agent, run_input=run_input, accept=accept) event_stream = adapter.run_stream() sse_event_stream = adapter.encode_stream(event_stream) return StreamingResponse(sse_event_stream, media_type=accept) Trust model for client-submitted messages ----------------------------------------- [](https://pydantic.dev/docs/ai/integrations/ui/overview/#trust-model-for-client-submitted-messages) UI adapter endpoints aren’t authentication boundaries. Both the AG-UI and Vercel AI protocols are designed around the client transmitting the full conversation history on each request, so anything in `message_history` from the protocol — assistant messages, tool calls, file URLs, tool results — is under the caller’s control. Treat the adapter endpoint as an internal backend service, running it inside your own authenticated route handler. See the [AG-UI security considerations](https://learn.microsoft.com/en-us/agent-framework/integrations/ag-ui/security-considerations) page for more on the deployment model both protocols assume. The adapters apply a few defaults so that the authoritative state stays on your side: * **System prompts** — client-submitted [`SystemPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SystemPromptPart) s are stripped by default and replaced with the agent’s configured prompt. Control with [`UIAdapter.manage_system_prompt`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.manage_system_prompt) ; see each adapter’s docs for details. * **Dangling tool calls** — if the client-submitted history ends in a [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) with unresolved [`ToolCallPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolCallPart) s and no matching `deferred_tool_results`, the tool calls are dropped with a warning, as a best-effort default so the agent doesn’t execute an unresolved tool call the model never emitted. For human-in-the-loop resumption, pass explicit `deferred_tool_results` to the run method — tool calls resolved by those results are kept (see the warning below). * **File URL schemes** — only `http` and `https` are accepted by default for [`FileUrl`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.FileUrl) parts in client-submitted messages. Non-HTTP schemes like `s3://` or `gs://` are dropped, since they cause the provider to fetch the object using your server’s IAM role or service account. See [`UIAdapter.allowed_file_url_schemes`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.allowed_file_url_schemes) . * **File URL download mode** — [`FileUrl.force_download`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.FileUrl.force_download) values other than `False` are reset to `False` by default on client-submitted messages. This prevents clients from forcing the server to fetch a URL, or using `'allow-local'` to opt out of the SSRF private-IP block. After auditing your frontend, opt into additional values with [`UIAdapter.allowed_file_url_force_download`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.allowed_file_url_force_download) . * **Uploaded files** — client-submitted [`UploadedFile`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.UploadedFile) parts are dropped by default, just like non-HTTP `FileUrl`s, since the server resolves them against the provider’s file storage API using its own credentials. After auditing your frontend, honor them by setting [`UIAdapter.allow_uploaded_files`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.allow_uploaded_files) to `True`. This is a purely inbound security setting: file content the agent produces is always serialized on the way back out to the client. For stricter conversation integrity (e.g. ensuring prior assistant turns and tool returns match what the server actually produced), persist the history server-side keyed by the thread/session ID and pass it to the adapter via `message_history` — caller-supplied history is trusted as coming from server-side persistence and isn’t subject to this sanitization. These defaults narrow what a fabricated history can reach, but a client that can submit history can always fabricate it; see [Trust boundary for client-supplied history](https://pydantic.dev/docs/ai/core-concepts/message-history/#trust-boundary-for-client-supplied-history) for what that means for your application’s authorization model. Was this page helpful? Thanks for your feedback! --- # Macroscope | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/macroscope/#_top) Macroscope ========== `Macroscope` runs a local [Macroscope](https://docs.macroscope.com/cli) code review from inside an agent: one tool shells out to the installed `macroscope` CLI, parses the streamed findings, and returns them as structured data. The agent validates each finding and fixes the real ones with the tools it already has — this capability surfaces findings only. [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/macroscope/) The problem ----------- [](https://pydantic.dev/docs/ai/harness/macroscope/#the-problem) Macroscope reviews the current branch’s diff and streams findings, but it ships as editor plugins (Claude Code, Codex, Cursor, OpenCode). There is no way to give a Pydantic AI agent the same review-and-fix loop from your own code. Usage ----- [](https://pydantic.dev/docs/ai/harness/macroscope/#usage) from pydantic_ai import Agent from pydantic_ai_harness.macroscope import Macroscope agent = Agent('anthropic:claude-sonnet-5', capabilities=[Macroscope()]) result = agent.run_sync('Run a Macroscope review and fix any real findings.') print(result.output) The `macroscope` CLI must be installed and authenticated on the host first: 1. Install: `curl -sSL https://raw.githubusercontent.com/prassoai/macroscope-local/main/install.sh | bash` 2. Sign in and pick a workspace by running `macroscope` once. The capability cannot install or authenticate on your behalf. If the binary is missing, the tool returns the install command; if a review never starts (usually because you are not signed in), the tool tells the agent to run `macroscope` to finish setup. The tool invokes `macroscope codereview --raw` for machine-readable streaming output, which needs a recent CLI build. The installer fetches the latest and the CLI self-updates on use, so a fresh install satisfies this. The tool -------- [](https://pydantic.dev/docs/ai/harness/macroscope/#the-tool) | Tool | Purpose | | --- | --- | | `run_macroscope_review` | Run `macroscope codereview` on the current branch and return the review id, terminal status, and findings. Accepts an optional `base` git ref. | Each finding is a `MacroscopeIssue` with `issue_id`, `sequence`, `path`, `line`, `severity`, `category`, and `body`. The capability’s default instructions tell the agent to treat every finding as untrusted: read the affected code to confirm an issue is real, skip false positives and duplicates, and verify each fix. Options ------- [](https://pydantic.dev/docs/ai/harness/macroscope/#options) Every field of `Macroscope` with its default: from pydantic_ai_harness.macroscope import Macroscope Macroscope( base=None, # git ref to diff against -- None lets the CLI auto-detect command='macroscope', # binary name or path cwd='.', # repository directory the review runs in timeout=600.0, # max seconds to wait for a review guidance=None, # None = default instructions, '' = none, str = custom ) A per-call `base` argument takes precedence over the field. Reviews call a remote service, so the timeout is generous by default; on timeout the CLI’s process group is killed and the timeout is reported to the model as a retryable error. Scope and composition --------------------- [](https://pydantic.dev/docs/ai/harness/macroscope/#scope-and-composition) This capability surfaces findings only. It does not edit files, create worktrees, or commit — validating and fixing findings is the agent’s job, using its other capabilities. Pair it with `FileSystem` or `Shell` to let the agent read code and apply fixes, and consider running the agent in an isolated worktree if you want fixes kept off your working tree. Agent spec ---------- [](https://pydantic.dev/docs/ai/harness/macroscope/#agent-spec) `Macroscope` works with Pydantic AI’s [agent spec](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) , so you can declare it in a config file instead of Python: # agent.yaml model: anthropic:claude-sonnet-5 capabilities: - Macroscope: base: main timeout: 900 from pydantic_ai import Agent from pydantic_ai_harness.macroscope import Macroscope agent = Agent.from_file('agent.yaml', custom_capability_types=[Macroscope]) Pass `custom_capability_types` so the spec loader knows how to instantiate `Macroscope`. Further reading --------------- [](https://pydantic.dev/docs/ai/harness/macroscope/#further-reading) * [Macroscope CLI documentation](https://docs.macroscope.com/cli) * [Pydantic AI capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) * [Toolsets](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) API reference ------------- [](https://pydantic.dev/docs/ai/harness/macroscope/#api-reference) Macroscope ---------- [](https://pydantic.dev/docs/ai/harness/macroscope/#pydantic_ai_harness.macroscope.Macroscope) **Bases:** `AbstractCapability[AgentDepsT]` Runs the `macroscope` CLI code review and hands the findings to the agent. Adds a `run_macroscope_review` tool that shells out to `macroscope codereview`, parses the streamed findings, and returns them as a `MacroscopeReview`. The agent validates and fixes findings with its own tools — this capability does not edit files, create worktrees, or commit. from pydantic_ai import Agent from pydantic_ai_harness.macroscope import Macroscope agent = Agent('anthropic:claude-sonnet-5', capabilities=[Macroscope()]) The `macroscope` CLI must be installed and authenticated on the host first (see the package README). This capability cannot sign in on the user’s behalf; if a review never starts, the tool reports that the user needs to run `macroscope` once. ### Attributes [](https://pydantic.dev/docs/ai/harness/macroscope/#attributes) #### base [](https://pydantic.dev/docs/ai/harness/macroscope/#pydantic_ai_harness.macroscope.Macroscope.base) Git ref to diff against. When `None`, `--base` is omitted and the CLI auto-detects the base branch itself (and creates its own review worktree). **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### command [](https://pydantic.dev/docs/ai/harness/macroscope/#pydantic_ai_harness.macroscope.Macroscope.command) Name or path of the CLI binary. Override for a non-default install location. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'macroscope'` #### cwd [](https://pydantic.dev/docs/ai/harness/macroscope/#pydantic_ai_harness.macroscope.Macroscope.cwd) Repository directory the review runs in. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | `Path` **Default:** `'.'` #### guidance [](https://pydantic.dev/docs/ai/harness/macroscope/#pydantic_ai_harness.macroscope.Macroscope.guidance) Custom review guidance for the system prompt. Leave as `None` for the default validate-then-fix guidance, or set `''` to contribute no instructions at all. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### timeout [](https://pydantic.dev/docs/ai/harness/macroscope/#pydantic_ai_harness.macroscope.Macroscope.timeout) Maximum seconds to wait for a review. Reviews call a remote service, so this is generous by default. **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) **Default:** `600.0` ### Methods [](https://pydantic.dev/docs/ai/harness/macroscope/#methods) #### get\_instructions [](https://pydantic.dev/docs/ai/harness/macroscope/#pydantic_ai_harness.macroscope.Macroscope.get_instructions) def get_instructions() -> str | None Static validate-then-fix guidance. A non-`None` `guidance` replaces the default; `''` disables instructions entirely. ##### Returns [](https://pydantic.dev/docs/ai/harness/macroscope/#returns) [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### get\_toolset [](https://pydantic.dev/docs/ai/harness/macroscope/#pydantic_ai_harness.macroscope.Macroscope.get_toolset) def get_toolset() -> MacroscopeToolset[AgentDepsT] Build the toolset that provides the `run_macroscope_review` tool. ##### Returns [](https://pydantic.dev/docs/ai/harness/macroscope/#returns-1) `MacroscopeToolset`\[`AgentDepsT`\] MacroscopeReview ---------------- [](https://pydantic.dev/docs/ai/harness/macroscope/#pydantic_ai_harness.macroscope.MacroscopeReview) **Bases:** `BaseModel` The result of one `macroscope codereview` run. `status` is the terminal `issue_status` reported by the CLI (`completed` or `failed`), or `unknown` if the stream ended without one. `review_id` is `None` when the CLI never emitted one — usually because the review did not start. MacroscopeIssue --------------- [](https://pydantic.dev/docs/ai/harness/macroscope/#pydantic_ai_harness.macroscope.MacroscopeIssue) **Bases:** `BaseModel` A single finding streamed by `macroscope codereview`. Parsed leniently: unknown fields are ignored so new CLI output does not break parsing, and any `issue_event` line that lacks the required fields is skipped. Was this page helpful? Thanks for your feedback! --- # Z.AI | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/models/zai/#_top) Z.AI ==== Install ------- [](https://pydantic.dev/docs/ai/models/zai/#install) To use [`ZaiModel`](https://pydantic.dev/docs/ai/api/models/zai/#pydantic_ai.models.zai.ZaiModel) , you need to either install `pydantic-ai`, or install `pydantic-ai-slim` with the `zai` optional group: * [pip](https://pydantic.dev/docs/ai/models/zai/#tab-panel-136) * [uv](https://pydantic.dev/docs/ai/models/zai/#tab-panel-137) Terminal pip install "pydantic-ai-slim[zai]" Terminal uv add "pydantic-ai-slim[zai]" Configuration ------------- [](https://pydantic.dev/docs/ai/models/zai/#configuration) To use [Z.AI](https://z.ai/) (Zhipu AI) through their API, go to [z.ai](https://z.ai/manage-apikey/apikey-list) and generate an API key. For a list of available models, see the [Z.AI documentation](https://docs.z.ai/) . Environment variable -------------------- [](https://pydantic.dev/docs/ai/models/zai/#environment-variable) Once you have the API key, you can set it as an environment variable: Terminal export ZAI_API_KEY='your-api-key' You can then use [`ZaiModel`](https://pydantic.dev/docs/ai/api/models/zai/#pydantic_ai.models.zai.ZaiModel) by name: from pydantic_ai import Agent agent = Agent('zai:glm-5') ... Or initialise the model directly with just the model name: from pydantic_ai import Agent from pydantic_ai.models.zai import ZaiModel model = ZaiModel('glm-5') agent = Agent(model) ... Thinking mode ------------- [](https://pydantic.dev/docs/ai/models/zai/#thinking-mode) Z.AI’s `glm-5.2`, `glm-5.1`, `glm-5`, `glm-4.7`, `glm-4.6` (hybrid thinking), and `glm-4.5` (interleaved thinking) models support thinking/reasoning mode, where the model produces reasoning content before the final response. This includes the `glm-4.6v` and `glm-4.5v` vision models. Configure this through the unified [`thinking`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.thinking) setting: from pydantic_ai import Agent from pydantic_ai.settings import ModelSettings agent = Agent( 'zai:glm-5', model_settings=ModelSettings(thinking=True), ) ... `thinking=True` enables thinking and `thinking=False` disables it. On GLM-5.2, an explicit effort level (`'minimal'`/`'low'`/`'medium'`/`'high'`/`'xhigh'`) is forwarded to Z.AI as `reasoning_effort`; on other GLM models, which don’t expose effort granularity, the effort levels all collapse to enabled. Omit the field to use each model’s default behavior. ### Preserved thinking [](https://pydantic.dev/docs/ai/models/zai/#preserved-thinking) On thinking-capable models, reasoning content from prior assistant responses is **preserved by default** — no configuration required — for better multi-turn coherence and consistency with other providers. The complete, unmodified `reasoning_content` from prior turns is automatically sent back to the API by Pydantic AI. If you instead want each turn to start fresh, **disable** it with `zai_clear_thinking=True` via the Z.AI-specific [`ZaiModelSettings`](https://pydantic.dev/docs/ai/api/models/zai/#pydantic_ai.models.zai.ZaiModelSettings) : from pydantic_ai import Agent from pydantic_ai.models.zai import ZaiModelSettings agent = Agent( 'zai:glm-5', # Opt out of the default preserved thinking: model_settings=ZaiModelSettings(thinking=True, zai_clear_thinking=True), ) ... See the [Z.AI thinking mode documentation](https://docs.z.ai/guides/capabilities/thinking-mode#preserved-thinking) for more details. `provider` argument ------------------- [](https://pydantic.dev/docs/ai/models/zai/#provider-argument) You can provide a custom [`Provider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.Provider) via the `provider` argument. In the simplest case, pass [`ZaiProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.zai.ZaiProvider) with just an API key. If you also want to customize the underlying `httpx.AsyncClient`, pass it when constructing the provider: from httpx import AsyncClient from pydantic_ai import Agent from pydantic_ai.models.zai import ZaiModel from pydantic_ai.providers.zai import ZaiProvider custom_http_client = AsyncClient(timeout=30) model = ZaiModel( 'glm-5', provider=ZaiProvider(api_key='your-api-key', http_client=custom_http_client), ) agent = Agent(model) ... If you do not need a custom HTTP client, omit the `http_client=custom_http_client` argument. Was this page helpful? Thanks for your feedback! --- # xAI | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/models/xai/#_top) xAI === Install ------- [](https://pydantic.dev/docs/ai/models/xai/#install) To use [`XaiModel`](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModel) , you need to either install `pydantic-ai`, or install `pydantic-ai-slim` with the `xai` optional group: * [pip](https://pydantic.dev/docs/ai/models/xai/#tab-panel-134) * [uv](https://pydantic.dev/docs/ai/models/xai/#tab-panel-135) Terminal pip install "pydantic-ai-slim[xai]" Terminal uv add "pydantic-ai-slim[xai]" Configuration ------------- [](https://pydantic.dev/docs/ai/models/xai/#configuration) To use xAI models from [xAI](https://x.ai/api) through their API, go to [console.x.ai](https://console.x.ai/team/default/api-keys) to create an API key. [docs.x.ai](https://docs.x.ai/developers/models) contains a list of available xAI models. Environment variable -------------------- [](https://pydantic.dev/docs/ai/models/xai/#environment-variable) Once you have the API key, you can set it as an environment variable: Terminal export XAI_API_KEY='your-api-key' You can then use [`XaiModel`](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModel) by name: from pydantic_ai import Agent agent = Agent('xai:grok-4.3') ... Or initialise the model directly: from pydantic_ai import Agent from pydantic_ai.models.xai import XaiModel # Uses XAI_API_KEY environment variable model = XaiModel('grok-4.3') agent = Agent(model) ... You can also customize the [`XaiModel`](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModel) with a custom provider: from pydantic_ai import Agent from pydantic_ai.models.xai import XaiModel from pydantic_ai.providers.xai import XaiProvider # Custom API key provider = XaiProvider(api_key='your-api-key') model = XaiModel('grok-4.3', provider=provider) agent = Agent(model) ... For gateway, regional, or proxy deployments you can also point the provider at a custom host and set a client-level default timeout, both of which are forwarded to the underlying `xai_sdk.AsyncClient`: from pydantic_ai import Agent from pydantic_ai.models.xai import XaiModel from pydantic_ai.providers.xai import XaiProvider provider = XaiProvider( api_key='your-api-key', api_host='gateway.example.com', timeout=30, ) model = XaiModel('grok-4.3', provider=provider) agent = Agent(model) ... `api_host` is the hostname of the xAI API server (the SDK connects over gRPC), and `timeout` is the default timeout in seconds applied to every request the client makes. Unlike other providers, the xAI SDK does not support per-request timeouts, so [`ModelSettings.timeout`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.timeout) is not supported and has no effect. Both options are omitted when left unset, so the SDK’s own defaults apply. You can also attach gRPC `metadata` to every request the client makes. The canonical use is xAI prompt-cache sticky routing, which pins a conversation to a cache node via an `x-grok-conv-id` so repeated prefixes are served from cache instead of reprocessed: from pydantic_ai import Agent from pydantic_ai.models.xai import XaiModel from pydantic_ai.providers.xai import XaiProvider provider = XaiProvider( api_key='your-api-key', metadata=(('x-grok-conv-id', 'my-conversation-id'),), ) model = XaiModel('grok-4.3', provider=provider) agent = Agent(model) ... `metadata` is a sequence of `(key, value)` string tuples forwarded verbatim to the underlying `xai_sdk.AsyncClient`, so the xAI SDK’s own documentation on [maximizing cache hits](https://docs.x.ai/developers/advanced-api-usage/prompt-caching/maximizing-cache-hits) applies. Because it is client-scoped, it applies to _every_ request made through the provider — a provider configured with a fixed `x-grok-conv-id` must not be shared between unrelated conversations, or those conversations will collide on the same cache node. Use a separate provider per conversation when the metadata is conversation-specific. Like `api_host` and `timeout`, it is omitted when left unset and ignored when a custom `xai_sdk.AsyncClient` is passed. Or with a custom `xai_sdk.AsyncClient`: from xai_sdk import AsyncClient from pydantic_ai import Agent from pydantic_ai.models.xai import XaiModel from pydantic_ai.providers.xai import XaiProvider xai_client = AsyncClient(api_key='your-api-key') provider = XaiProvider(xai_client=xai_client) model = XaiModel('grok-4.3', provider=provider) agent = Agent(model) ... X Search -------- [](https://pydantic.dev/docs/ai/models/xai/#x-search) xAI models support searching X (formerly Twitter) for real-time posts and content. The recommended way to enable it is with the [`XSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.XSearch) capability — see the [capability documentation](https://pydantic.dev/docs/ai/capabilities/overview/#provider-adaptive-tools) for more details, including cross-provider usage. For the full list of supported options, see the [xAI X Search documentation](https://docs.x.ai/developers/tools/x-search) . xai\_x\_search.py from datetime import datetime from pydantic_ai import Agent from pydantic_ai.capabilities import XSearch agent = Agent( 'xai:grok-4.3', capabilities=[\ XSearch(\ allowed_x_handles=['OpenAI', 'AnthropicAI', 'dasfacc'],\ from_date=datetime(2024, 1, 1),\ to_date=datetime(2024, 12, 31),\ enable_image_understanding=True,\ enable_video_understanding=True,\ include_output=True,\ )\ ], ) result = agent.run_sync('What have AI companies been posting about?') print(result.output) """ OpenAI announced their latest model updates, while Anthropic shared research on AI safety... """ _(This example is complete, it can be run “as is”)_ The `XSearch` capability accepts: * **`allowed_x_handles`** / **`excluded_x_handles`**: filter results to (or away from) up to 20 X handles. These are mutually exclusive. * **`from_date`** / **`to_date`**: restrict results to posts created within the given datetime range (naive datetimes are interpreted as UTC). * **`enable_image_understanding`** (default: `False`): analyze images attached to posts. * **`enable_video_understanding`** (default: `False`): analyze video content attached to posts. * **`include_output`** (default: `False`): include the raw X search results on the [`NativeToolReturnPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.NativeToolReturnPart) available via [`ModelResponse.native_tool_calls`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse.native_tool_calls) . Without this, the model uses the search results internally but only returns its text summary; enabling it gives programmatic access to the searched posts, sources, and metadata. As an alternative to the capability, you can pass the lower-level [`XSearchTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.XSearchTool) directly via `capabilities=[NativeTool(XSearchTool(...))]` — see the [X Search Tool documentation](https://pydantic.dev/docs/ai/tools-toolsets/native-tools/#x-search-tool) — or enable raw output globally via the [`XaiModelSettings.xai_include_x_search_output`](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_include_x_search_output) [model setting](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) . Reasoning effort ---------------- [](https://pydantic.dev/docs/ai/models/xai/#reasoning-effort) Grok 4.3 supports `reasoning_effort` values of `'none'`, `'low'`, `'medium'`, and `'high'`. You can configure it directly with [`XaiModelSettings.xai_reasoning_effort`](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_reasoning_effort) , or use the cross-provider [`ModelSettings.thinking`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.thinking) setting: xai\_reasoning\_effort.py from pydantic_ai import Agent from pydantic_ai.models.xai import XaiModelSettings agent = Agent( 'xai:grok-4.3', model_settings=XaiModelSettings(xai_reasoning_effort='medium'), ) Set `xai_reasoning_effort='none'` or `thinking=False` to disable reasoning on Grok 4.3. xAI redirects several retired text model slugs to `grok-4.3`; choose `grok-4.3` and an explicit reasoning effort when you need predictable behavior and cost. See the [xAI May 15 retirement guide](https://docs.x.ai/developers/migration/may-15-retirement) for details. Grok 4.5 supports `'low'`, `'medium'`, and `'high'` but not `'none'`, so it always reasons: `thinking=False` is silently ignored and `thinking=True` maps to `'medium'`. Agentic turns ------------- [](https://pydantic.dev/docs/ai/models/xai/#agentic-turns) When a request uses xAI’s server-side [native tools](https://pydantic.dev/docs/ai/tools-toolsets/native-tools/) (e.g. web search, code execution, X search), xAI runs its own loop — calling those tools and processing their results — before returning a final response. You can cap how many turns that server-side loop may take with [`XaiModelSettings.xai_max_turns`](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_max_turns) : xai\_max\_turns.py from pydantic_ai import Agent from pydantic_ai.models.xai import XaiModelSettings agent = Agent( 'xai:grok-4.3', model_settings=XaiModelSettings(xai_max_turns=5), ) `xai_max_turns` only governs xAI’s server-side native-tool loop. It has no effect on ordinary client-side tools or on Pydantic AI’s own agent loop — to bound those, use [`UsageLimits`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits) . Note that when parallel tool calls are enabled, multiple tool calls can occur within a single turn, so `xai_max_turns` does not necessarily equal the total number of tool calls made. Multi-agent models ------------------ [](https://pydantic.dev/docs/ai/models/xai/#multi-agent-models) xAI’s [multi-agent models](https://docs.x.ai/developers/model-capabilities/text/multi-agent) (e.g. `grok-4.20-multi-agent`) research a question with several agents in parallel before answering. You can choose how many agents they use with [`XaiModelSettings.xai_agent_count`](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_agent_count) , which accepts `4` or `16`: xai\_agent\_count.py from pydantic_ai import Agent from pydantic_ai.models.xai import XaiModelSettings agent = Agent( 'xai:grok-4.20-multi-agent', model_settings=XaiModelSettings(xai_agent_count=16), ) More agents means deeper research, at the cost of more tokens and higher latency. Other xAI models ignore the setting. Streaming cancellation ---------------------- [](https://pydantic.dev/docs/ai/models/xai/#streaming-cancellation) Was this page helpful? Thanks for your feedback! --- # Troubleshooting | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/realtime/troubleshooting/#_top) Troubleshooting =============== Below are suggestions on how to fix some common problems with realtime sessions, each linking to the page that covers the underlying behavior. For issues not listed here or addressed in the documentation, see the general [troubleshooting page](https://pydantic.dev/docs/ai/overview/troubleshooting/) , ask in the [Pydantic Slack](https://pydantic.dev/docs/ai/overview/help/) , or create an issue on [GitHub](https://github.com/pydantic/pydantic-ai/issues) . No audio, or no useful speech ----------------------------- [](https://pydantic.dev/docs/ai/realtime/troubleshooting/#no-audio-or-no-useful-speech) Send mono PCM16 at `session.audio_input_sample_rate` and play it at `session.audio_output_sample_rate`. Do not assume the rates match. See the [audio wire contract](https://pydantic.dev/docs/ai/realtime/audio/#audio-wire-contract) . The model never responds ------------------------ [](https://pydantic.dev/docs/ai/realtime/troubleshooting/#the-model-never-responds) In push-to-talk mode, call `commit_audio()` and then `create_response()` after sending audio. See [push-to-talk](https://pydantic.dev/docs/ai/realtime/turns/#push-to-talk) . The model interrupts itself --------------------------- [](https://pydantic.dev/docs/ai/realtime/troubleshooting/#the-model-interrupts-itself) The microphone is probably hearing speaker output. Add echo cancellation in the device/WebRTC layer and stop local playback on real [barge-in](https://pydantic.dev/docs/ai/realtime/turns/#barge-in) . Tools seem to stall ------------------- [](https://pydantic.dev/docs/ai/realtime/troubleshooting/#tools-seem-to-stall) The local tool runs concurrently, but the provider may pause speech while awaiting the result. Show tool lifecycle events and review [concurrent tool execution](https://pydantic.dev/docs/ai/realtime/tools/#concurrent-tool-execution) . A reconnect lost context ------------------------ [](https://pydantic.dev/docs/ai/realtime/troubleshooting/#a-reconnect-lost-context) Inspect `RealtimeSessionReconnectEvent.state_restored`. If false, begin a fresh conversation; if true but a current utterance vanished, that in-flight media was outside the restored completed-turn history. See [state restoration](https://pydantic.dev/docs/ai/realtime/lifecycle/#state-restoration) . Gemini reaches its session limit -------------------------------- [](https://pydantic.dev/docs/ai/realtime/troubleshooting/#gemini-reaches-its-session-limit) Set the `reconnect` setting to a [`ReconnectPolicy`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.ReconnectPolicy) ; Gemini session resumption is enabled automatically alongside it. Recovery uses the latest in-memory server handle after the drop. See [Gemini session resumption](https://pydantic.dev/docs/ai/realtime/gemini/#session-resumption) and [provider session limits](https://pydantic.dev/docs/ai/realtime/lifecycle/#provider-session-limits) . Was this page helpful? Thanks for your feedback! --- # Thread Executor | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/thread-executor/#_top) Thread Executor =============== The [`UseThreadExecutor`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.UseThreadExecutor) [capability](https://pydantic.dev/docs/ai/capabilities/overview/) provides a custom [`Executor`](https://docs.python.org/3/library/concurrent.futures.html#concurrent.futures.Executor) for running sync tool functions and other sync callbacks in threads. This is useful in long-running servers (e.g. FastAPI) where the default ephemeral threads from `anyio.to_thread.run_sync` can accumulate under sustained load: from concurrent.futures import ThreadPoolExecutor from pydantic_ai import Agent from pydantic_ai.capabilities import UseThreadExecutor executor = ThreadPoolExecutor(max_workers=16, thread_name_prefix='agent-worker') agent = Agent('openai:gpt-5.2', capabilities=[UseThreadExecutor(executor)]) See [Thread executor for long-running servers](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#thread-executor-for-long-running-servers) for more details. Was this page helpful? Thanks for your feedback! --- # pydantic_graph.util | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic_graph/util/#_top) pydantic\_graph.util ==================== Utility types and functions for type manipulation and introspection. This module provides helper classes and functions for working with Python’s type system, including workarounds for type checker limitations and utilities for runtime type inspection. Some ---- [](https://pydantic.dev/docs/ai/api/pydantic_graph/util/#pydantic_graph.util.Some) **Bases:** `Generic[T]` Container for explicitly present values in Maybe type pattern. This class represents a value that is definitely present, as opposed to None. It’s part of the Maybe pattern, similar to Option/Maybe in functional programming, allowing distinction between “no value” (None) and “value is None” (Some(None)). ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/util/#attributes) #### value [](https://pydantic.dev/docs/ai/api/pydantic_graph/util/#pydantic_graph.util.Some.value) The wrapped value. **Type:** `T` TypeExpression -------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/util/#pydantic_graph.util.TypeExpression) **Bases:** `Generic[T]` A workaround for type checker limitations when using complex type expressions. This class serves as a wrapper for types that cannot normally be used in positions requiring `type[T]`, such as `Any`, `Union[...]`, or `Literal[...]`. It provides a way to pass these complex type expressions to functions expecting concrete types. get\_callable\_name ------------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/util/#pydantic_graph.util.get_callable_name) def get_callable_name(callable_: Any) -> str Extract a human-readable name from a callable object. ### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/util/#returns) [`str`](https://docs.python.org/3/library/stdtypes.html#str) — The callable’s **name** attribute if available, otherwise its string representation. ### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_graph/util/#parameters) **`callable_`** : [`Any`](https://docs.python.org/3/library/typing.html#typing.Any) [](https://pydantic.dev/docs/ai/api/pydantic_graph/util/#pydantic_graph.util.get_callable_name(callable_)) Any callable object (function, method, class, etc.). unpack\_type\_expression ------------------------ [](https://pydantic.dev/docs/ai/api/pydantic_graph/util/#pydantic_graph.util.unpack_type_expression) def unpack_type_expression(type_: TypeOrTypeExpression[T]) -> type[T] Extract the actual type from a TypeExpression wrapper or return the type directly. ### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/util/#returns-1) [`type`](https://docs.python.org/3/glossary.html#term-type) \[`T`\] — The unwrapped type, ready for use in runtime type operations. ### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_graph/util/#parameters-1) **`type_`** : `TypeOrTypeExpression`\[`T`\] [](https://pydantic.dev/docs/ai/api/pydantic_graph/util/#pydantic_graph.util.unpack_type_expression(type_)) Either a direct type or a TypeExpression wrapper. Maybe ----- [](https://pydantic.dev/docs/ai/api/pydantic_graph/util/#pydantic_graph.util.Maybe) Optional-like type that distinguishes between absence and None values. Unlike Optional\[T\], Maybe\[T\] can differentiate between: * No value present: represented as None * Value is None: represented as Some(None) This is particularly useful when None is a valid value in your domain. **Default:** `TypeAliasType('Maybe', Some[T] | None, type_params=(T,))` T - [](https://pydantic.dev/docs/ai/api/pydantic_graph/util/#pydantic_graph.util.T) Generic type variable with inferred variance. **Default:** `TypeVar('T', infer_variance=True)` TypeOrTypeExpression -------------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/util/#pydantic_graph.util.TypeOrTypeExpression) Type alias allowing both direct types and TypeExpression wrappers. This alias enables functions to accept either regular types (when compatible with type checkers) or TypeExpression wrappers for complex type expressions. The correct type should be inferred automatically in either case. **Default:** `TypeAliasType('TypeOrTypeExpression', type[TypeExpression[T]] | type[T], type_params=(T,))` Was this page helpful? Thanks for your feedback! --- # Dependencies | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/core-concepts/dependencies/#_top) Dependencies ============ Pydantic AI uses a dependency injection system to provide data and services to your agent’s [system prompts](https://pydantic.dev/docs/ai/core-concepts/agent/#system-prompts) , [tools](https://pydantic.dev/docs/ai/tools-toolsets/tools/) and [output validators](https://pydantic.dev/docs/ai/core-concepts/output/#output-validator-functions) . Matching Pydantic AI’s design philosophy, our dependency system tries to use existing best practice in Python development rather than inventing esoteric “magic”, this should make dependencies type-safe, understandable, easier to test, and ultimately easier to deploy in production. Defining Dependencies --------------------- [](https://pydantic.dev/docs/ai/core-concepts/dependencies/#defining-dependencies) Dependencies can be any python type. While in simple cases you might be able to pass a single object as a dependency (e.g. an HTTP connection), [dataclasses](https://docs.python.org/3/library/dataclasses.html#module-dataclasses) are generally a convenient container when your dependencies included multiple objects. Here’s an example of defining an agent that requires dependencies. (**Note:** dependencies aren’t actually used in this example, see [Accessing Dependencies](https://pydantic.dev/docs/ai/core-concepts/dependencies/#accessing-dependencies) below) unused\_dependencies.py from dataclasses import dataclass import httpx from pydantic_ai import Agent @dataclass class MyDeps: # (1) api_key: str http_client: httpx.AsyncClient agent = Agent( 'openai:gpt-5.2', deps_type=MyDeps, # (2) ) async def main(): async with httpx.AsyncClient() as client: deps = MyDeps('foobar', client) result = await agent.run( 'Tell me a joke.', deps=deps, # (3) ) print(result.output) #> Did you hear about the toothpaste scandal? They called it Colgate. Define a dataclass to hold dependencies. Pass the dataclass type to the `deps_type` argument of the [`Agent` constructor](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.__init__) . **Note**: we're passing the type here, NOT an instance, this parameter is not actually used at runtime, it's here so we can get full type checking of the agent. When running the agent, pass an instance of the dataclass to the `deps` parameter. _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main())` to run `main`)_ Accessing Dependencies ---------------------- [](https://pydantic.dev/docs/ai/core-concepts/dependencies/#accessing-dependencies) Dependencies are accessed through the [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) type, this should be the first parameter of system prompt functions etc. system\_prompt\_dependencies.py from dataclasses import dataclass import httpx from pydantic_ai import Agent, RunContext @dataclass class MyDeps: api_key: str http_client: httpx.AsyncClient agent = Agent( 'openai:gpt-5.2', deps_type=MyDeps, ) @agent.system_prompt # (1) async def get_system_prompt(ctx: RunContext[MyDeps]) -> str: # (2) response = await ctx.deps.http_client.get( # (3) 'https://example.com', headers={'Authorization': f'Bearer {ctx.deps.api_key}'}, # (4) ) response.raise_for_status() return f'Prompt: {response.text}' async def main(): async with httpx.AsyncClient() as client: deps = MyDeps('foobar', client) result = await agent.run('Tell me a joke.', deps=deps) print(result.output) #> Did you hear about the toothpaste scandal? They called it Colgate. [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) may optionally be passed to a [`system_prompt`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.system_prompt) function as the only argument. [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) is parameterized with the type of the dependencies, if this type is incorrect, static type checkers will raise an error. Access dependencies through the [`.deps`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.deps) attribute. Access dependencies through the [`.deps`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.deps) attribute. _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main())` to run `main`)_ In addition to [`.deps`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.deps) , [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) provides access to the running agent via [`.agent`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.agent) , which is useful when [tools](https://pydantic.dev/docs/ai/tools-toolsets/tools/) , [hooks](https://pydantic.dev/docs/ai/core-concepts/hooks/) , or [capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) need to read agent properties like [`name`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.name) or [`output_type`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.output_type) . The [`.realtime`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.realtime) property identifies realtime sessions without requiring a model type check, and [`.realtime_session`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.realtime_session) exposes the live [`RealtimeSession`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession) to tools and hooks once it is connected. Dependency fields can also be referenced in instructions and descriptions via [template strings](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#template-strings) — for example, [`TemplateStr('Hello {{name}}')`](https://pydantic.dev/docs/ai/api/pydantic-ai/template/#pydantic_ai.template.TemplateStr) renders `name` from the deps object at runtime. This is especially useful in [agent specs](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) where callables aren’t available. ### Asynchronous vs. Synchronous dependencies [](https://pydantic.dev/docs/ai/core-concepts/dependencies/#asynchronous-vs-synchronous-dependencies) [System prompt functions](https://pydantic.dev/docs/ai/core-concepts/agent/#system-prompts) , [function tools](https://pydantic.dev/docs/ai/tools-toolsets/tools/) and [output validators](https://pydantic.dev/docs/ai/core-concepts/output/#output-validator-functions) are all run in the async context of an agent run. If these functions are not coroutines (e.g. `async def`) they are called with [`run_in_executor`](https://docs.python.org/3/library/asyncio-eventloop.html#asyncio.loop.run_in_executor) in a thread pool. It’s therefore marginally preferable to use `async` methods where dependencies perform IO, although synchronous dependencies should work fine too. Here’s the same example as above, but with a synchronous dependency: sync\_dependencies.py from dataclasses import dataclass import httpx from pydantic_ai import Agent, RunContext @dataclass class MyDeps: api_key: str http_client: httpx.Client # (1) agent = Agent( 'openai:gpt-5.2', deps_type=MyDeps, ) @agent.system_prompt def get_system_prompt(ctx: RunContext[MyDeps]) -> str: # (2) response = ctx.deps.http_client.get( 'https://example.com', headers={'Authorization': f'Bearer {ctx.deps.api_key}'} ) response.raise_for_status() return f'Prompt: {response.text}' async def main(): deps = MyDeps('foobar', httpx.Client()) result = await agent.run( 'Tell me a joke.', deps=deps, ) print(result.output) #> Did you hear about the toothpaste scandal? They called it Colgate. Here we use a synchronous `httpx.Client` instead of an asynchronous `httpx.AsyncClient`. To match the synchronous dependency, the system prompt function is now a plain function, not a coroutine. _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main())` to run `main`)_ Full Example ------------ [](https://pydantic.dev/docs/ai/core-concepts/dependencies/#full-example) As well as system prompts, dependencies can be used in [tools](https://pydantic.dev/docs/ai/tools-toolsets/tools/) and [output validators](https://pydantic.dev/docs/ai/core-concepts/output/#output-validator-functions) . full\_example.py from dataclasses import dataclass import httpx from pydantic_ai import Agent, ModelRetry, RunContext @dataclass class MyDeps: api_key: str http_client: httpx.AsyncClient agent = Agent( 'openai:gpt-5.2', deps_type=MyDeps, ) @agent.system_prompt async def get_system_prompt(ctx: RunContext[MyDeps]) -> str: response = await ctx.deps.http_client.get('https://example.com') response.raise_for_status() return f'Prompt: {response.text}' @agent.tool # (1) async def get_joke_material(ctx: RunContext[MyDeps], subject: str) -> str: response = await ctx.deps.http_client.get( 'https://example.com#jokes', params={'subject': subject}, headers={'Authorization': f'Bearer {ctx.deps.api_key}'}, ) response.raise_for_status() return response.text @agent.output_validator # (2) async def validate_output(ctx: RunContext[MyDeps], output: str) -> str: response = await ctx.deps.http_client.post( 'https://example.com#validate', headers={'Authorization': f'Bearer {ctx.deps.api_key}'}, params={'query': output}, ) if response.status_code == 400: raise ModelRetry(f'invalid response: {response.text}') response.raise_for_status() return output async def main(): async with httpx.AsyncClient() as client: deps = MyDeps('foobar', client) result = await agent.run('Tell me a joke.', deps=deps) print(result.output) #> Did you hear about the toothpaste scandal? They called it Colgate. To pass `RunContext` to a tool, use the [`tool`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.tool) decorator. `RunContext` may optionally be passed to a [`output_validator`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.output_validator) function as the first argument. _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main())` to run `main`)_ Overriding Dependencies ----------------------- [](https://pydantic.dev/docs/ai/core-concepts/dependencies/#overriding-dependencies) When testing agents, it’s useful to be able to customise dependencies. While this can sometimes be done by calling the agent directly within unit tests, we can also override dependencies while calling application code which in turn calls the agent. This is done via the [`override`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.override) method on the agent. joke\_app.py from dataclasses import dataclass import httpx from pydantic_ai import Agent, RunContext @dataclass class MyDeps: api_key: str http_client: httpx.AsyncClient async def system_prompt_factory(self) -> str: # (1) response = await self.http_client.get('https://example.com') response.raise_for_status() return f'Prompt: {response.text}' joke_agent = Agent('openai:gpt-5.2', deps_type=MyDeps) @joke_agent.system_prompt async def get_system_prompt(ctx: RunContext[MyDeps]) -> str: return await ctx.deps.system_prompt_factory() # (2) async def application_code(prompt: str) -> str: # (3) ... ... # now deep within application code we call our agent async with httpx.AsyncClient() as client: app_deps = MyDeps('foobar', client) result = await joke_agent.run(prompt, deps=app_deps) # (4) return result.output Define a method on the dependency to make the system prompt easier to customise. Call the system prompt factory from within the system prompt function. Application code that calls the agent, in a real application this might be an API endpoint. Call the agent from within the application code, in a real application this call might be deep within a call stack. Note `app_deps` here will NOT be used when deps are overridden. _(This example is complete, it can be run “as is”)_ test\_joke\_app.py from joke_app import MyDeps, application_code, joke_agent class TestMyDeps(MyDeps): # (1) async def system_prompt_factory(self) -> str: return 'test prompt' async def test_application_code(): test_deps = TestMyDeps('test_key', None) # (2) with joke_agent.override(deps=test_deps): # (3) joke = await application_code('Tell me a joke.') # (4) assert joke.startswith('Did you hear about the toothpaste scandal?') Define a subclass of `MyDeps` in tests to customise the system prompt factory. Create an instance of the test dependency, we don't need to pass an `http_client` here as it's not used. Override the dependencies of the agent for the duration of the `with` block, `test_deps` will be used when the agent is run. Now we can safely call our application code, the agent will use the overridden dependencies. Examples -------- [](https://pydantic.dev/docs/ai/core-concepts/dependencies/#examples) The following examples demonstrate how to use dependencies in Pydantic AI: * [Weather Agent](https://pydantic.dev/docs/ai/examples/getting-started/weather-agent/) * [SQL Generation](https://pydantic.dev/docs/ai/examples/data-analytics/sql-gen/) * [RAG](https://pydantic.dev/docs/ai/examples/data-analytics/rag/) Was this page helpful? Thanks for your feedback! --- # Multi-Run Evaluation | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/how-to/multi-run/#_top) Multi-Run Evaluation ==================== Run each case multiple times to measure variability and get more reliable aggregate results. AI systems are inherently stochastic — the same input can produce different outputs across runs. The `repeat` parameter lets you run each case multiple times and automatically aggregates the results, giving you a clearer picture of your system’s typical behavior. Basic Usage ----------- [](https://pydantic.dev/docs/ai/evals/how-to/multi-run/#basic-usage) Pass `repeat` to [`evaluate()`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset.evaluate) or [`evaluate_sync()`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset.evaluate_sync) : from pydantic_evals import Case, Dataset dataset = Dataset( name='multi_run_basic', cases=[\ Case(name='greeting', inputs='Say hello'),\ Case(name='farewell', inputs='Say goodbye'),\ ] ) def task(inputs: str) -> str: return inputs.upper() # Run each case 5 times report = dataset.evaluate_sync(task, repeat=5) # 2 cases × 5 repeats = 10 total runs print(len(report.cases)) #> 10 When `repeat > 1`, each run gets an indexed name like `greeting [1/5]`, `greeting [2/5]`, etc., while the original case name is preserved in [`source_case_name`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.ReportCase.source_case_name) for grouping. Accessing Grouped Results ------------------------- [](https://pydantic.dev/docs/ai/evals/how-to/multi-run/#accessing-grouped-results) Use [`case_groups()`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.EvaluationReport.case_groups) to access runs organized by original case, with per-group aggregated statistics: from pydantic_evals import Case, Dataset dataset = Dataset( name='grouped_results', cases=[\ Case(name='greeting', inputs='Say hello'),\ Case(name='farewell', inputs='Say goodbye'),\ ] ) def task(inputs: str) -> str: return inputs.upper() report = dataset.evaluate_sync(task, repeat=3) groups = report.case_groups() assert groups is not None # None for single-run (repeat=1) print(len(groups)) #> 2 group_names = [g.name for g in groups] print(group_names) #> ['greeting', 'farewell'] # Each group has 3 runs and aggregated statistics for group in groups: assert len(group.runs) == 3 assert len(group.failures) == 0 assert group.summary.task_duration > 0 Each [`ReportCaseGroup`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.ReportCaseGroup) contains: * `name` — the original case name * `runs` — the individual [`ReportCase`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.ReportCase) results * `failures` — any runs that raised exceptions * `summary` — a [`ReportCaseAggregate`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.ReportCaseAggregate) with averaged scores, metrics, labels, assertions, and durations Aggregation ----------- [](https://pydantic.dev/docs/ai/evals/how-to/multi-run/#aggregation) With `repeat > 1`, the report’s [`averages()`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.EvaluationReport.averages) uses a two-level aggregation strategy: 1. **Per-group averages**: Each case’s runs are averaged into a group summary 2. **Cross-group averages**: The group summaries are averaged to produce the final result This ensures each original case contributes equally to the overall averages, regardless of how many runs succeeded or failed. from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import EqualsExpected dataset = Dataset( name='aggregation', cases=[\ Case(name='easy', inputs='hello', expected_output='HELLO'),\ Case(name='hard', inputs='world', expected_output='WORLD'),\ ], evaluators=[EqualsExpected()], ) def task(inputs: str) -> str: return inputs.upper() report = dataset.evaluate_sync(task, repeat=3) averages = report.averages() assert averages is not None print(f'Overall assertion rate: {averages.assertions}') #> Overall assertion rate: 1.0 Default Behavior ---------------- [](https://pydantic.dev/docs/ai/evals/how-to/multi-run/#default-behavior) When `repeat=1` (the default), behavior is identical to a standard evaluation — no run indexing, no `source_case_name`, and `case_groups()` returns `None`: from pydantic_evals import Case, Dataset dataset = Dataset(name='default_behavior', cases=[Case(name='test', inputs='hello')]) def task(inputs: str) -> str: return inputs.upper() report = dataset.evaluate_sync(task) # repeat=1 by default assert report.case_groups() is None assert all(c.source_case_name is None for c in report.cases) Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/how-to/multi-run/#next-steps) * **[Concurrency & Performance](https://pydantic.dev/docs/ai/evals/how-to/concurrency/) ** — Control parallel execution with `max_concurrency` * **[Metrics & Attributes](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/) ** — Track custom metrics across runs * **[Logfire Integration](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/) ** — Visualize multi-run results Was this page helpful? Thanks for your feedback! --- # Pydantic AI Harness [Skip to content](https://pydantic.dev/docs/ai/harness/#_top) Overview ======== **The batteries for your [Pydantic AI](https://pydantic.dev/docs/ai/) agent.** Pydantic AI’s [capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) and [hooks](https://pydantic.dev/docs/ai/core-concepts/hooks/) API is how you give an agent its harness — bundles of tools, lifecycle hooks, instructions, and model settings that extend what the agent can do without any framework changes. **Pydantic AI Harness** is the official capability library for Pydantic AI, maintained by the [Pydantic AI](https://github.com/pydantic/pydantic-ai) team. Pydantic AI core ships the capabilities that require model or framework support, plus the ones fundamental to every agent — [web search](https://pydantic.dev/docs/ai/capabilities/web-search/) , [tool search](https://pydantic.dev/docs/ai/capabilities/tool-search/) , [thinking](https://pydantic.dev/docs/ai/capabilities/thinking/) . Everything else lives here: standalone building blocks you pick and choose to turn your agent into a coding agent, a research assistant, or anything else. This is also where new capabilities start — as they stabilize and prove themselves broadly essential, they can graduate into core. What goes where? ---------------- [](https://pydantic.dev/docs/ai/harness/#what-goes-where) Pydantic AI core ships the agent loop, model providers, the capabilities/hooks abstraction, and two kinds of capabilities: * **Capabilities that require model or framework support** — anything backed by provider native tools (like [image generation](https://pydantic.dev/docs/ai/capabilities/image-generation/) ), provider-specific APIs (like [compaction](https://pydantic.dev/docs/ai/capabilities/compaction/) via the OpenAI or Anthropic APIs), or deep agent graph integration (like [tool search](https://pydantic.dev/docs/ai/capabilities/tool-search/) and [on-demand loading](https://pydantic.dev/docs/ai/capabilities/on-demand/) ). These go hand-in-hand with model class code and need to ship together. * **Capabilities that are fundamental to the agent experience** — things nearly every agent benefits from, like [web search](https://pydantic.dev/docs/ai/capabilities/web-search/) , [web fetch](https://pydantic.dev/docs/ai/capabilities/web-fetch/) , [thinking](https://pydantic.dev/docs/ai/capabilities/thinking/) , and [MCP](https://pydantic.dev/docs/ai/capabilities/mcp/) . These feel like qualities of the agent itself, not accessories. See [built-in capabilities](https://pydantic.dev/docs/ai/capabilities/overview/#built-in-capabilities) for the full list. **Pydantic AI Harness** is where everything else lives: standalone capabilities that make specific categories of agents powerful, or that are still finding their final shape. Context management, memory, guardrails, file system access, code execution, multi-agent orchestration — these are the building blocks you pick and choose based on what your agent needs to do. The harness is also where new capabilities _start_. It ships as a separate package so capabilities can iterate faster without the strict backward-compatibility requirements of core. As a capability stabilizes and proves itself broadly essential, it can graduate into core — [code mode](https://pydantic.dev/docs/ai/harness/code-mode/) is an early candidate. Many capabilities benefit from a “fall up” pattern: they typically start as a local implementation that works with every model, then gain provider-native support that uses the provider’s built-in API when available — auto-switching between the two. This is how [web search](https://pydantic.dev/docs/ai/capabilities/web-search/) , [web fetch](https://pydantic.dev/docs/ai/capabilities/web-fetch/) , and [image generation](https://pydantic.dev/docs/ai/capabilities/image-generation/) already work in core, and the same approach is coming for skills, code mode, and context compaction. Installation ------------ [](https://pydantic.dev/docs/ai/harness/#installation) Terminal uv add pydantic-ai-harness Some capabilities need an extra to pull in their optional dependencies: Terminal uv add "pydantic-ai-harness[codemode]" # Code Mode (adds the Monty sandbox) uv add "pydantic-ai-harness[dynamic-workflow]" # Dynamic Workflow (adds the Monty sandbox) uv add "pydantic-ai-harness[modal]" # Modal Sandbox (adds the Modal SDK) uv add "pydantic-ai-harness[logfire]" # Managed Prompt (Logfire-managed prompts) uv add "pydantic-ai-harness[exa]" # Exa Search (web research via the Exa API) uv add "pydantic-ai-harness[skills]" # Skills (loads SKILL.md frontmatter) uv add "pydantic-ai-harness[browser-use]" # Browser Use (autonomous web tasks; Python 3.11+) uv add "pydantic-ai-harness[stackone]" # StackOne (actions on linked business applications) uv add "pydantic-ai-harness[acp]" # ACP (Agent Client Protocol SDK) uv add "pydantic-ai-harness[mongodb]" # MongoDB backends for Step Persistence and Media (adds pymongo) The `code-mode` extra is also supported as an alias for `codemode`. Requires Python 3.10+ and `pydantic-ai-slim>=2.18.0`. Quick start ----------- [](https://pydantic.dev/docs/ai/harness/#quick-start) Install the harness alongside the Pydantic AI extras this example uses: Terminal uv add "pydantic-ai-slim[anthropic,mcp,duckduckgo,logfire]" "pydantic-ai-harness[code-mode]" import logfire from pydantic_ai import Agent from pydantic_ai.capabilities import MCP, WebSearch from pydantic_ai_harness import CodeMode # See https://pydantic.dev/docs/ai/integrations/logfire/ for setup details. logfire.configure() logfire.instrument_pydantic_ai() agent = Agent( 'anthropic:claude-opus-4-7', capabilities=[\ # Wraps every tool into a single run_code tool, sandboxed by Monty\ # (https://github.com/pydantic/monty -- pulled in by the [code-mode] extra).\ # The model writes Python that calls multiple tools with loops, conditionals,\ # asyncio.gather, and local filtering -- one model round-trip for N tool calls.\ CodeMode(),\ # Connect to any MCP server -- here, the open-source Hacker News server\ # (https://github.com/cyanheads/hn-mcp-server). native=False forces the\ # local MCP toolset so CodeMode can wrap the tools; without it,\ # providers that natively support MCP server connectors execute the tools\ # server-side and bypass the sandbox.\ MCP('https://hn.caseyjhand.com/mcp', native=False),\ # Provider-adaptive web search; native=False routes through the local\ # DuckDuckGo fallback (the [duckduckgo] extra above) so CodeMode can batch\ # web searches alongside the HN calls in a single run_code.\ WebSearch(native=False),\ ], ) result = agent.run_sync( "Across the top, best, and 'show HN' Hacker News feeds, find the most-discussed " "story with at least 100 points. Pull its comment thread, its submitter's profile, " "and any web coverage. Summarize what you find in one paragraph." ) print(result.output) """ The most-discussed HN story across top/best/show clearing 100 points is "Vibe coding and agentic engineering are getting closer than I'd like" by Simon Willison (748 points, 853 comments, on the Best feed), submitted by long-time HNer e12e. The piece argues that the two modes Willison once kept mentally separate -- throwaway "vibe coding" and disciplined "agentic engineering" -- are blurring, since agents like Claude Code now reliably handle non-trivial tasks like "build a JSON API endpoint that runs a SQL query" with tests and docs on the first pass. The HN thread is unusually substantive, with commenters debating whether LLMs created or merely *exposed* sloppy engineering practices and warning of a "normalization of deviance" as engineers stop reviewing diffs. """ [![Logfire trace from the Quick start run](https://pydantic.dev/docs/ai/harness/img/quick-start-trace.png)](https://logfire-us.pydantic.dev/public-trace/84bcf123-2106-49da-9f6f-5c26395339bb?spanId=7650806a0785b946) **[See this run as a public Logfire trace ->](https://logfire-us.pydantic.dev/public-trace/84bcf123-2106-49da-9f6f-5c26395339bb?spanId=7650806a0785b946) ** Each `run_code` span fans out into the tool calls the model issued from inside the sandbox — it’s the easiest way to understand what code mode actually did. Capabilities ------------ [](https://pydantic.dev/docs/ai/harness/#capabilities) Each capability is a self-contained battery you drop into an agent’s `capabilities=[...]` list. They compose with each other and with Pydantic AI’s [built-in capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) . | Capability | What it does | Extra | | --- | --- | --- | | [Advisor](https://pydantic.dev/docs/ai/harness/advisor/) | Lets an executor consult another model through a provider-native tool or a local Pydantic AI fallback. | | | [Code Mode](https://pydantic.dev/docs/ai/harness/code-mode/) | Wraps the agent’s tools into a single `run_code` tool, sandboxed by [Monty](https://github.com/pydantic/monty)
. The model writes Python that calls the tools as functions — with loops, conditionals, `asyncio.gather`, and local filtering — collapsing N tool calls into one model round-trip. | `codemode` | | [Skills](https://pydantic.dev/docs/ai/harness/skills/) | Loads Agent Skill instructions only when the model needs them. | `skills` | | [FileSystem](https://pydantic.dev/docs/ai/harness/filesystem/) | Sandboxed file access scoped to a root directory: read, write, edit, search, and find files. Rejects path traversal above the root, resolves symlinks before authorizing, and keeps `.git/`, `.env`, key files, and secrets read-only by default. | — | | [Shell](https://pydantic.dev/docs/ai/harness/shell/) | Command execution in a subprocess rooted at a working directory, gated by allowlists, denylists, timeouts, and optional environment-variable stripping (including a preset for common LLM provider credentials). | — | | [Repo Context](https://pydantic.dev/docs/ai/harness/repo-context/) | Auto-loads repo context — `CLAUDE.md`/`AGENTS.md` and repository structure — so the agent starts a run already oriented in the project. | — | | [Pydantic AI Docs](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/) | An on-demand `read_pyai_docs` tool that pulls Pydantic AI documentation into the run when the agent needs it, instead of preloading it. | — | | [Exa Search](https://pydantic.dev/docs/ai/harness/exa-search/) | Web research backed by the [Exa](https://exa.ai/)
search API: `web_search` returns results with their most relevant excerpts, `get_page` reads a specific URL in full, and opt-in `deep_search` synthesizes a cited answer in one call. Output is budgeted per tool. | `exa` | | [Browser Use](https://pydantic.dev/docs/ai/harness/browser-use/) | Delegates open-ended web tasks to an autonomous [browser-use](https://github.com/browser-use/browser-use)
agent: one `browse_web` tool hands over a natural-language goal, the sub-agent drives a real browser, and the result comes back as text. | `browser-use` | | [Compaction](https://pydantic.dev/docs/ai/harness/compaction/) | Keeps a run within token limits: sliding-window trimming, LLM-powered summarization of older messages, and warnings before the context or iteration ceiling is hit. | — | | [Tool Output Limits](https://pydantic.dev/docs/ai/harness/tool-output-limits/) | Reduces an oversized tool return when it is produced — truncate, spill to a queryable file, or summarize — so a large payload does not persist in history and get re-sent every request. | — | | [Warn On Cache Busts](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/) | Warns when a run’s prompt-cache hit collapses between model requests — a moved cacheable prefix or an expired provider cache — reading the provider’s own `cache_read_tokens` verdict. | — | | [Step Persistence](https://pydantic.dev/docs/ai/harness/step-persistence/) | Saves and restores full conversation state; snapshot, resume (`continue_run`), and fork (`fork_run`) a run. In-memory, file, SQLite, and MongoDB backends. | `mongodb` (Mongo backend only) | | [Conversation Search](https://pydantic.dev/docs/ai/harness/conversation-search/) | A dependency-free BM25 `search_conversation_history` tool over the history `StepPersistence` stores: recall turns that compaction dropped from the live context, and past runs in the same store. | — | | [Media](https://pydantic.dev/docs/ai/harness/media/) | Offloads large `BinaryContent` and large text parts to content-addressed stores (disk, SQLite, S3, MongoDB) so big payloads do not bloat message history. | `mongodb` (Mongo store only) | | [Subagents](https://pydantic.dev/docs/ai/harness/subagents/) | Delegates subtasks to specialized child agents through a delegate tool. | — | | [Dynamic Workflow](https://pydantic.dev/docs/ai/harness/dynamic-workflow/) | Orchestrates sub-agents from a model-written Python script — fan-out, chaining, and voting in a single tool call. | `dynamic-workflow` | | [Planning](https://pydantic.dev/docs/ai/harness/planning/) | Breaks a complex task into a structured plan before execution and tracks progress against it. | — | | [System Reminders](https://pydantic.dev/docs/ai/harness/system-reminders/) | Re-injects behavioral guidance mid-run — on a cadence or reactively from a condition — to counter instruction fade in long sessions, without invalidating the prompt cache. | — | | [Memory](https://pydantic.dev/docs/ai/harness/memory/) | Gives an agent a persistent, namespaced notebook with bounded prompt injection, on-demand search, and concurrency-safe stores. | — | | [Runtime Capability Creation](https://pydantic.dev/docs/ai/harness/capability-creation/) | Lets an agent create, validate, and persist Pydantic AI capabilities during one run for the orchestrator to load on the next run. | — | | [Guardrails](https://pydantic.dev/docs/ai/harness/guardrails/) | Validates user input, tool calls and their results, and model output — block or redact, with structured results. | — | | [Managed Prompt](https://pydantic.dev/docs/ai/harness/managed-prompt/) | Backs an agent’s instructions with a [Logfire-managed prompt](https://logfire.pydantic.dev/docs/reference/advanced/prompt-management/)
, so you can version, label, and roll out prompt changes from the Logfire UI without redeploying — with a code default that keeps the agent working when no remote value is available. | `logfire` | | [StackOne](https://pydantic.dev/docs/ai/harness/stackone/) | Actions on the user’s SaaS accounts (HRIS, ATS, CRM, and more) via the [StackOne](https://www.stackone.com/)
integration platform: API-key auth, account scoping, action filtering, and a search/execute mode for large catalogs. | `stackone` | | [ACP](https://pydantic.dev/docs/ai/harness/acp/)
_(experimental)_ | Serves an agent to editors (Zed, etc.) over the [Agent Client Protocol](https://agentclientprotocol.com/)
— streamed text, diff-rendered edits, and tool approval. | `acp` | Most capabilities are stable within the [version policy](https://pydantic.dev/docs/ai/harness/#version-policy) below. [ACP](https://pydantic.dev/docs/ai/harness/acp/) is the exception — it is still experimental, imported from `pydantic_ai_harness.experimental.acp`, and may change or be removed in a future release. Build your own -------------- [](https://pydantic.dev/docs/ai/harness/#build-your-own) [Capabilities](https://pydantic.dev/docs/ai/capabilities/custom/) are the primary extension point for Pydantic AI. Any of the capabilities in this library can serve as a reference for building your own. Publishing as a standalone package? Use the `pydantic-ai-` naming convention — see [Publishing capability packages](https://pydantic.dev/docs/ai/guides/extensibility/#publishing-capability-packages) . Version policy -------------- [](https://pydantic.dev/docs/ai/harness/#version-policy) Pydantic AI Harness uses **0.x versioning** to signal that APIs are still stabilizing. During 0.x, minor releases (0.1 -> 0.2) may include breaking changes — renamed parameters, changed defaults, restructured APIs — while patch releases (0.1.0 -> 0.1.1) will not intentionally break existing behavior. All breaking changes are documented in release notes with migration guidance. This is why the harness is a separate package from [Pydantic AI](https://github.com/pydantic/pydantic-ai) , which has a [stricter version policy](https://pydantic.dev/docs/ai/project/version-policy/) . As the core capabilities stabilize, the library will move toward 1.0 with matching stability guarantees. Pydantic AI references ---------------------- [](https://pydantic.dev/docs/ai/harness/#pydantic-ai-references) * [Capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) — what capabilities are, built-in capabilities, building your own * [Hooks](https://pydantic.dev/docs/ai/core-concepts/hooks/) — lifecycle hooks reference, ordering, error handling * [Extensibility](https://pydantic.dev/docs/ai/guides/extensibility/) — publishing packages, third-party ecosystem * [Toolsets](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) — building tools for capabilities * [API reference](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/) — full API docs Was this page helpful? Thanks for your feedback! --- # Getting Help | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/overview/help/#_top) Getting Help ============ If you need help getting started with Pydantic AI or with advanced usage, the following sources may be useful. Slack ----- [](https://pydantic.dev/docs/ai/overview/help/#slack) Join the `#pydantic-ai` channel in the [Pydantic Slack](https://logfire.pydantic.dev/docs/join-slack/) to ask questions, get help, and chat about Pydantic AI. There’s also channels for Pydantic, Logfire, and FastUI. If you’re on a [Logfire](https://pydantic.dev/logfire) Pro plan, you can also get a dedicated private slack collab channel with us. GitHub Issues ------------- [](https://pydantic.dev/docs/ai/overview/help/#-github-issues) The [Pydantic AI GitHub Issues](https://github.com/pydantic/pydantic-ai/issues) are a great place to ask questions and give us feedback. Was this page helpful? Thanks for your feedback! --- # Warn On Cache Busts | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#_top) Warn On Cache Busts =================== Warn when a run’s prompt cache hit collapses between model requests, so a moved cacheable prefix or an expired provider cache surfaces instead of quietly re-charging tokens it could have served from cache. > \[!NOTE\] Import this capability from its submodule. It is not re-exported from `pydantic_ai_harness`: > > from pydantic_ai_harness.warn_on_cache_busts import WarnOnCacheBusts > Warn On Cache Busts is a released, non-experimental capability. Pydantic AI Harness is still on 0.x releases, so the API may change between minor releases. See the [version policy](https://pydantic.dev/docs/ai/harness/#version-policy) . What it watches --------------- [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#what-it-watches) Prompt caching pays off only while the cacheable prefix (tools, then system instructions, then message history) stays byte-stable across a run’s consecutive requests. When something moves that prefix — reordered tools, a timestamp injected into instructions, a serialization-level block hop — the provider re-charges tokens it could have served from cache. This is the **observe** signal: it reads the provider’s own verdict rather than guessing from the structured request. On each response it reads `usage.cache_read_tokens` and tracks the largest cacheable prefix the run has established (`cache_read_tokens + cache_write_tokens`, a high-water mark), keyed by the response’s `(provider_name, model_name)`. Because message history is append-only, a stable prefix means each request for that model reads back at least what the previous one cached; a large drop is the observable signature of a collapse. When a request reads back less than `collapse_ratio` of the established prefix, the monitor emits a `CacheBustWarning` once and latches that key, staying quiet about the collapse until a healthy read-back re-stabilizes the cache. A sustained collapse — caching toggled off mid-run (`read == 0, write == 0`), or a prefix that moves every request so the provider keeps writing a cache nothing reads back — therefore warns once, not on every request. from pydantic_ai import Agent from pydantic_ai_harness.warn_on_cache_busts import WarnOnCacheBusts agent = Agent('anthropic:claude-sonnet-4-5', capabilities=[WarnOnCacheBusts()]) await agent.run('...') # a CacheBustWarning fires if a cached prefix collapses mid-run The verdict is cross-provider for free — pyai normalizes every provider into the `cache_read_tokens` / `cache_write_tokens` fields on `RequestUsage`. Model switches and expiry ------------------------- [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#model-switches-and-expiry) Keying per provider and model means a mid-run model switch does not warn: a `FallbackModel` failover or a per-step model change uses a different cache key, so the monitor starts a fresh mark for it instead of comparing against the previous model’s. Marks are kept per key rather than reset, so switching back to an earlier model within its cache TTL still compares against that model’s prefix. A collapse has two shapes the monitor cannot tell apart, so the warning names both: the cacheable prefix moved, or the provider’s cache expired under an unchanged prefix (a gap between requests longer than the cache TTL — Anthropic’s default is 5 minutes, refreshed on each hit). When the gap since the same model’s previous request exceeds `cache_ttl_seconds`, the message reports the gap so a long tool or approval pause isn’t mistaken for a moved prefix. The gap is timed per model, so switching away and back measures the returning model’s own idle time, not whatever ran in between. Options ------- [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#options) * `collapse_ratio` (default `0.5`): warn when a request reads back less than this fraction of the established prefix. Conservative by default so ordinary rounding or a partial miss does not fire; raise toward `1.0` to warn on smaller regressions. It must be greater than `0.0` — a ratio of `0.0` could never warn, so it is rejected rather than treated as a silent disable switch. * `min_prefix_tokens` (default `1024`): only judge collapse once the established prefix reaches this many tokens. Below a provider’s minimum cacheable size (Anthropic’s is 1024) `cache_read_tokens` is noisy or zero. * `cache_ttl_seconds` (default `300`): the assumed provider cache TTL. Message-only — when the gap since the same model’s previous request exceeds it, the warning notes the collapse may be a cache expiry rather than a moved prefix. It does not change whether a warning fires. Lower it for providers with a shorter cache lifetime. Silencing and escalation ------------------------ [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#silencing-and-escalation) There is no bespoke suppression API. Use the stdlib `warnings` machinery, exactly as you would manage any other `UserWarning`: import warnings from pydantic_ai_harness.warn_on_cache_busts import CacheBustWarning # Silence the whole category: warnings.filterwarnings('ignore', category=CacheBustWarning) # Silence one intentional bust, scoped to the operation that causes it: with warnings.catch_warnings(): warnings.simplefilter('ignore', CacheBustWarning) result = agent.run_sync('...') # e.g. a step that switches models or adds a file # Treat every bust as an error (dev/CI enforcement): warnings.filterwarnings('error', category=CacheBustWarning) In tests, assert an intentional bust with `pytest.warns(CacheBustWarning)`, or silence a legitimately-busting test with `@pytest.mark.filterwarnings('ignore::pydantic_ai_harness.warn_on_cache_busts.CacheBustWarning')`. Logfire ------- [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#logfire) Logfire bridges the stdlib `logging` module, not the `warnings` module, so a `CacheBustWarning` does not reach your traces on its own. To route busts into Logfire, redirect Python warnings to the `logging` system once at startup: import logging logging.captureWarnings(True) # warnings.warn(...) -> the 'py.warnings' logger -> Logfire The monitor’s signal is the `CacheBustWarning`; routing it through `logging` is how it reaches Logfire. Composition ----------- [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#composition) * The monitor only implements `for_run` and `after_model_request`; it adds no tools, instructions, or model settings, so it composes with any other capability, toolset, or `ToolSearch` setup without interference. * Per-run state (the per-key marks and timing) is materialized in `for_run`, so one `WarnOnCacheBusts` instance can be reused across many `Agent.run` calls — each run is judged independently. Scope ----- [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#scope) * **Observational only.** It reports that a cached prefix collapsed, not why — a moved prefix and a provider-side cache expiry look the same from the token counts, so the warning names both. The structural explanation (“what moved the prefix this turn”) is a separate job. * **Fires only when caching is enabled and reported.** A run that never establishes a cache never warns. * **A mid-run model switch does not warn.** Marks are per `(provider_name, model_name)`, so a `FallbackModel` failover starts a fresh mark rather than collapsing the previous model’s. API reference ------------- [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#api-reference) * [`pydantic_ai_harness.warn_on_cache_busts` source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/warn_on_cache_busts/) * [Pydantic AI capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) * [Pydantic AI hooks](https://pydantic.dev/docs/ai/core-concepts/hooks/) The public module exports `WarnOnCacheBusts` and `CacheBustWarning`. Import them from `pydantic_ai_harness.warn_on_cache_busts`. WarnOnCacheBusts ---------------- [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#pydantic_ai_harness.warn_on_cache_busts.WarnOnCacheBusts) **Bases:** `AbstractCapability[AgentDepsT]` Warn when a run’s prompt cache hit collapses between requests. Attach it to any agent whose model uses prompt caching. On each response the monitor reads `usage.cache_read_tokens` and tracks the largest cacheable prefix the run has established (`cache_read_tokens + cache_write_tokens`, a high-water mark), keyed by the response’s `(provider_name, model_name)`. When a later request for the same key reads back fewer than `collapse_ratio` of that established prefix, it emits a `CacheBustWarning` once and then stays quiet about that collapse until a healthy read-back re-stabilizes the cache, so a sustained collapse warns once rather than on every subsequent request. Keying per provider and model means a mid-run model switch does not warn: a `FallbackModel` failover or a per-step model change uses a different cache key, so it starts a fresh mark for that key instead of comparing against the previous model’s. Marks are kept per key rather than reset, so switching back to an earlier model within its cache TTL still compares against that model’s established prefix — and the expiry hedge measures the gap against that same model’s previous request, not whatever ran in between. Because message history is append-only, a stable prefix means each request reads back at least what the previous one cached. A large drop is the observable signature of a collapse, whether the cause is a moved prefix (reordered tools, injected timestamps, a serialization-level block hop) or a provider-side cache expiry when the gap between requests exceeds the cache TTL. The monitor surfaces the collapse; it does not attribute the cause. from pydantic_ai import Agent from pydantic_ai_harness.warn_on_cache_busts import WarnOnCacheBusts agent = Agent('anthropic:claude-sonnet-4-5', capabilities=[WarnOnCacheBusts()]) await agent.run('...') # a CacheBustWarning fires if a cached prefix collapses mid-run The monitor is silent when caching is off or unreported (`cache_read_tokens` stays 0), so it never fires spuriously in tests that don’t exercise caching. Silencing and dev/CI escalation both go through the stdlib `warnings` filters — see `CacheBustWarning`. ### Attributes [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#attributes) #### cache\_ttl\_seconds [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#pydantic_ai_harness.warn_on_cache_busts.WarnOnCacheBusts.cache_ttl_seconds) Assumed provider cache TTL, in seconds (Anthropic’s default is 300, refreshed on each hit). Message-only: when the gap since the previous request for the same model exceeds this, the warning notes that the collapse may be a provider-side cache expiry rather than a moved prefix. It does not change whether a warning fires. Lower it for providers with a shorter cache lifetime. **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) **Default:** `300.0` #### collapse\_ratio [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#pydantic_ai_harness.warn_on_cache_busts.WarnOnCacheBusts.collapse_ratio) Warn when a request reads back less than this fraction of the established prefix. Conservative by default (0.5): only a drop below half the previously-cached prefix counts as a collapse, so ordinary provider rounding or a partial cache miss does not fire. Raise it toward 1.0 to warn on smaller regressions. Must be greater than 0.0 (a ratio of 0.0 could never warn, so it is rejected rather than treated as a silent disable switch). **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) **Default:** `0.5` #### min\_prefix\_tokens [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#pydantic_ai_harness.warn_on_cache_busts.WarnOnCacheBusts.min_prefix_tokens) Only judge collapse once the established prefix reaches this many tokens. Below a provider’s minimum cacheable size (Anthropic’s is 1024) `cache_read_tokens` is noisy or zero, so small prefixes are ignored to avoid false positives. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `1024` ### Methods [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#methods) #### after\_model\_request [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#pydantic_ai_harness.warn_on_cache_busts.WarnOnCacheBusts.after_model_request) `@async` def after_model_request( ctx: RunContext[AgentDepsT], *, request_context: ModelRequestContext, response: ModelResponse, ) -> ModelResponse Compare this response’s cache read against the established prefix for its model, then update it. ##### Returns [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#returns) [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) #### for\_run [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#pydantic_ai_harness.warn_on_cache_busts.WarnOnCacheBusts.for_run) `@async` def for_run(ctx: RunContext[AgentDepsT]) -> AbstractCapability[AgentDepsT] Give this run a fresh per-key state (marks, timing, step) so each run is judged alone. ##### Returns [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#returns-1) `AbstractCapability`\[`AgentDepsT`\] CacheBustWarning ---------------- [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#pydantic_ai_harness.warn_on_cache_busts.CacheBustWarning) **Bases:** [`UserWarning`](https://docs.python.org/3/library/exceptions.html#UserWarning) Warned when a previously-established prompt cache hit collapses on a later request. Emitted by `WarnOnCacheBusts` when this run read back far fewer cached tokens for the same provider and model than a prior request established. The likely causes are a moved cacheable prefix (reordered tools, injected timestamps, a serialization-level block hop) or a provider-side cache expiry under an unchanged prefix (a gap between requests longer than the cache TTL). The monitor observes the collapse; it does not attribute the cause. Silence it, or escalate it to an error in dev/CI, with the stdlib `warnings` machinery (no bespoke API): import warnings from pydantic\_ai\_harness.warn\_on\_cache\_busts import CacheBustWarning Silence the whole category: =========================== [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#silence-the-whole-category) warnings.filterwarnings(‘ignore’, category=CacheBustWarning) Silence one intentional bust, scoped to the operation that causes it: ===================================================================== [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#silence-one-intentional-bust-scoped-to-the-operation-that-causes-it) with warnings.catch\_warnings(): warnings.simplefilter(‘ignore’, CacheBustWarning) result = agent.run\_sync(’…’) # e.g. a step that switches models or adds a file Treat every bust as an error (dev/CI enforcement): ================================================== [](https://pydantic.dev/docs/ai/harness/warn-on-cache-busts/#treat-every-bust-as-an-error-devci-enforcement) warnings.filterwarnings(‘error’, category=CacheBustWarning) In tests, assert an intentional bust with `pytest.warns(CacheBustWarning)`, or silence a legitimately-busting test with `@pytest.mark.filterwarnings('ignore::pydantic_ai_harness.warn_on_cache_busts.CacheBustWarning')`. Was this page helpful? Thanks for your feedback! --- # Azure | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/realtime/azure/#_top) Azure ===== [`AzureRealtimeModel`](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModel) connects to Azure’s realtime speech-to-speech with the server-side Pydantic AI agent loop — either the **Azure OpenAI GA** protocol (the default) or **Azure AI Voice Live** (opt-in). Start with the [realtime quickstart](https://pydantic.dev/docs/ai/realtime/overview/#quickstart) or [text-to-audio example](https://pydantic.dev/docs/ai/examples/realtime/realtime-text-to-audio/) . Setup ----- [](https://pydantic.dev/docs/ai/realtime/azure/#setup) Azure OpenAI realtime uses the OpenAI realtime stack, so install `pydantic-ai-slim` with the `openai-realtime` optional group: * [pip](https://pydantic.dev/docs/ai/realtime/azure/#tab-panel-170) * [uv](https://pydantic.dev/docs/ai/realtime/azure/#tab-panel-171) Terminal pip install "pydantic-ai-slim[openai-realtime]" Terminal uv add "pydantic-ai-slim[openai-realtime]" Set `AZURE_OPENAI_ENDPOINT` and `AZURE_OPENAI_API_KEY` as for the [Azure AI Foundry provider](https://pydantic.dev/docs/ai/models/openai/#azure-ai-foundry) . Use the `azure:` prefix followed by your Azure deployment name: from pydantic_ai import Agent agent = Agent(instructions='You are a helpful voice assistant.') async def main(): async with agent.realtime('azure:my-realtime-deployment').session() as session: await session.send('Say hello.') async for part in session.stream_transcripts(): print(f'{part.speaker}: {part.transcript}') #> assistant: Hello from the realtime assistant. if part.speaker == 'assistant': break # keep listening in a real call; we stop after one reply _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main())` to run `main`)_ For explicit configuration, use [`AzureProvider.for_realtime()`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.azure.AzureProvider.for_realtime) . It accepts a bare resource endpoint or its `/openai/v1` form. The GA realtime protocol uses `/openai/v1/realtime` and does not take an `api_version`. Requests authenticate with the resource API key by default, or with a Microsoft Entra ID token when a `credential` is passed (see [Browser WebRTC and Microsoft Entra ID](https://pydantic.dev/docs/ai/realtime/azure/#browser-webrtc-and-microsoft-entra-id) ). Model names ----------- [](https://pydantic.dev/docs/ai/realtime/azure/#model-names) Pass the Azure **deployment name**, which is chosen when the model is deployed and need not match the underlying model ID. Available realtime models and regions are documented in the [Azure OpenAI realtime documentation](https://learn.microsoft.com/en-us/azure/ai-services/openai/realtime-audio-quickstart) . Settings -------- [](https://pydantic.dev/docs/ai/realtime/azure/#settings) Azure uses [`OpenAIRealtimeModelSettings`](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModelSettings) — the realtime counterpart of [model run settings](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) — including the [shared settings](https://pydantic.dev/docs/ai/realtime/overview/#shared-settings) plus: * `openai_voice` for the provider voice; * `openai_input_noise_reduction` and `openai_output_speed`; * `openai_turn_detection` for server or semantic VAD (see [turn detection](https://pydantic.dev/docs/ai/realtime/turns/#automatic-turn-detection) ); * `openai_truncation` for session context management. See [OpenAI settings](https://pydantic.dev/docs/ai/realtime/openai/#settings) for the common settings shape. Azure realtime does not expose `temperature` through Pydantic AI. ### Input transcription deployment [](https://pydantic.dev/docs/ai/realtime/azure/#input-transcription-deployment) Azure resolves the [`input_transcription_model`](https://pydantic.dev/docs/ai/realtime/audio/#input-transcription) setting against deployments in your resource. The default `'auto'` selects `gpt-realtime-whisper`; a resource without a matching deployment emits a `DeploymentNotFound` transcription error on every turn. Deploy a realtime-capable transcription model such as `gpt-realtime-whisper` or `gpt-4o-transcribe`, then set `input_transcription_model` to that deployment name. A classic `whisper` deployment is not accepted. Set the field to `None` to disable transcription and use [`audio_retention='input_audio'`](https://pydantic.dev/docs/ai/realtime/history/#retaining-audio) if the spoken turn must remain available as audio. Browser WebRTC and Microsoft Entra ID ------------------------------------- [](https://pydantic.dev/docs/ai/realtime/azure/#browser-webrtc-and-microsoft-entra-id) Azure OpenAI supports the same browser WebRTC flow as OpenAI — the audio flows browser ↔ Azure directly while your backend runs a control-plane **sideband**. See [Connecting a frontend](https://pydantic.dev/docs/ai/realtime/deployment/#browser-webrtc-server-sideband) for the topology, and use [`AgentRealtime.answer_webrtc_offer`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentRealtime.answer_webrtc_offer) / [`AgentRealtime.create_client_secret`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentRealtime.create_client_secret) exactly as on OpenAI. Azure relays the offer with `webrtcfilter=on`, which limits the events forwarded to the browser to a safe subset so the session instructions stay on the server’s control connection. Azure requests authenticate with the resource’s API key by default. To use **Microsoft Entra ID** instead — so no API key is involved, e.g. when the resource is locked to managed identity — pass a `credential` (any [`azure.identity`](https://learn.microsoft.com/python/api/overview/azure/identity-readme) credential, e.g. `DefaultAzureCredential`). It authenticates **every** request to the resource — the realtime WebSocket session and the WebRTC signaling — with a bearer token for the Azure OpenAI data plane (scope `https://ai.azure.com/.default`), which requires the **Cognitive Services User** role on the resource: from azure.identity import DefaultAzureCredential from pydantic_ai.providers.azure import AzureProvider from pydantic_ai.realtime.azure import AzureRealtimeModel model = AzureRealtimeModel( 'gpt-realtime', # `entra_authenticated=True` so no resource key is required — a resource locked to managed # identity has none. Omit `provider=` entirely to take the endpoint from `AZURE_OPENAI_ENDPOINT`. provider=AzureProvider.for_realtime( azure_endpoint='https://my-resource.openai.azure.com', entra_authenticated=True ), credential=DefaultAzureCredential(), ) # The realtime session, `answer_webrtc_offer`, and `create_client_secret` now authenticate with an Entra # bearer token; the browser only ever receives the short-lived ephemeral secret, never it or the API key. Azure AI Voice Live ------------------- [](https://pydantic.dev/docs/ai/realtime/azure/#azure-ai-voice-live) [Azure AI Voice Live](https://learn.microsoft.com/azure/ai-services/speech-service/voice-live) is Microsoft’s managed speech-to-speech service, with extra session options and a wider model catalog than the GA realtime API — including cascade pipelines (Azure speech-to-text → a chat model → Azure text-to-speech) over models like `gpt-4o`, `gpt-4.1`, and `gpt-5`. It’s the **same [`AzureRealtimeModel`](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModel) **: set [`azure_voice_live=True`](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModelSettings.azure_voice_live) and the model targets the Voice Live endpoint and beta session protocol. Voice Live is a distinct Azure resource with its own credentials, so set `AZURE_VOICELIVE_ENDPOINT`, `AZURE_VOICELIVE_API_KEY`, and `AZURE_VOICELIVE_API_VERSION`, or pass `voice_live_endpoint`, `voice_live_api_key`, and `voice_live_api_version` to [`AzureProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.azure.AzureProvider) . Each value resolves explicit argument first, then its own `AZURE_VOICELIVE_*` variable, then the Azure OpenAI endpoint/key — so a Voice Live user who only has one resource doesn’t need to configure both, and one who has both never gets a mixture of the two. from pydantic_ai import Agent from pydantic_ai.providers.azure import AzureProvider from pydantic_ai.realtime.azure import AzureRealtimeModel, AzureRealtimeModelSettings provider = AzureProvider( voice_live_endpoint='https://my-voice-live.services.ai.azure.com', voice_live_api_key='...', voice_live_api_version='2026-04-10', ) agent = Agent(instructions='You are a helpful voice assistant.') # Pass the Voice Live `provider`, and set `azure_voice_live` on the model rather than per session so # `model.profile` reflects Voice Live (see the note below). model = AzureRealtimeModel( 'gpt-realtime', provider=provider, settings=AzureRealtimeModelSettings(azure_voice_live=True) ) async def main(): async with agent.realtime(model).session() as session: await session.send('Say hello.') async for event in session: ... Voice-Live-only knobs use the `azure_voice_live_*` prefix (e.g. [`azure_voice_live_turn_detection`](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModelSettings.azure_voice_live_turn_detection) ). ### Which models use which API [](https://pydantic.dev/docs/ai/realtime/azure/#which-models-use-which-api) `azure_voice_live` isn’t always needed: `AzureRealtimeModel` routes by model. The two APIs overlap but neither contains the other, so each recognized model is served by the GA realtime API, by Voice Live, or by both: * **Both** (e.g. `gpt-realtime`, `gpt-realtime-mini`) — default to GA; `azure_voice_live=True` selects Voice Live. * **Voice Live only** (e.g. `gpt-5` and the other cascade chat models, `phi4-mm-realtime`) — routed to Voice Live automatically, with or without the setting. * **GA only** (e.g. `gpt-realtime-2`, `gpt-4o-realtime-preview`) — `azure_voice_live=True` raises a [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) , since Voice Live doesn’t serve them. An unrecognized model (a future release, or a deployment named after something else) defaults to GA and reaches Voice Live only with `azure_voice_live=True`. When a deployment’s name doesn’t match its model, pass a [`profile=`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModel.profile) [`AzureRealtimeModelProfile`](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModelProfile) with [`azure_realtime_apis`](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModelProfile.azure_realtime_apis) to correct the routing: from pydantic_ai.providers.azure import AzureProvider from pydantic_ai.realtime.azure import AzureRealtimeModel, AzureRealtimeModelProfile # A Voice-Live-only model deployed under a custom name. model = AzureRealtimeModel( 'my-voice-bot', provider=AzureProvider( voice_live_endpoint='https://my-voice-live.services.ai.azure.com', voice_live_api_key='...' ), profile=AzureRealtimeModelProfile(azure_realtime_apis=frozenset({'voice_live'})), ) Feature support and limitations ------------------------------- [](https://pydantic.dev/docs/ai/realtime/azure/#feature-support-and-limitations) | Feature | Support | Notes | | --- | --- | --- | | Audio format | Full feature support | Mono PCM16, 24 kHz input and output | | Text output | Full feature support | Select with `output_modality='text'` | | Image input | Full feature support | [Images](https://pydantic.dev/docs/ai/realtime/audio/#images)
provide context for the next turn | | Manual turns | Full feature support | `turn_detection=False` plus [commit/create verbs](https://pydantic.dev/docs/ai/realtime/turns/#push-to-talk) | | Interruption/truncation | Full feature support | [`interrupt(played_ms=...)`](https://pydantic.dev/docs/ai/realtime/turns/#barge-in)
records the heard cutoff | | Input transcription | Limited parameter support | Requires a [compatible transcription deployment](https://pydantic.dev/docs/ai/realtime/azure/#input-transcription-deployment)
in the Azure resource | | Native tools | Unsupported | Configure [local fallbacks](https://pydantic.dev/docs/ai/realtime/tools/#native-tools)
for web capabilities | | Usage | Full feature support | Token, audio, and cache breakdowns | | Reconnection | Full feature support | Pydantic AI [replays completed local history](https://pydantic.dev/docs/ai/realtime/lifecycle/#state-restoration)
; in-flight media is lost | See [Audio, images, and transcripts](https://pydantic.dev/docs/ai/realtime/audio/) , [Turns and interruptions](https://pydantic.dev/docs/ai/realtime/turns/) , [Tools](https://pydantic.dev/docs/ai/realtime/tools/) , and [Connection lifecycle](https://pydantic.dev/docs/ai/realtime/lifecycle/) for the provider-agnostic workflows. Provider-specific quirks ------------------------ [](https://pydantic.dev/docs/ai/realtime/azure/#provider-specific-quirks) * A failed input transcription leaves the user turn represented as [retained audio](https://pydantic.dev/docs/ai/realtime/history/#retaining-audio) when available, or as a content-less `SpeechPart` otherwise. * Azure AI Voice Live rides the same model behind `azure_voice_live=True`, against its own resource and beta session protocol; browser WebRTC is GA-only for now. Was this page helpful? Thanks for your feedback! --- # System Reminders | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/system-reminders/#_top) System Reminders ================ `SystemReminders` re-states targeted behavioral guidance partway through a run — on a fixed cadence or reactively from a condition — to counter the instruction fade that sets in over many turns, without ever invalidating the prompt cache. [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/system_reminders/) > The API may change between releases. Where practical, breaking changes ship with a deprecation warning. The problem ----------- [](https://pydantic.dev/docs/ai/harness/system-reminders/#the-problem) Long multi-turn runs suffer instruction fade: after many tool-use turns the model progressively ignores the guidance it was given at the start. A single start-of-session system prompt is not enough for extended work. The fix is to re-state targeted guidance mid-run — on a fixed cadence, or reactively when a condition is detected. The solution ------------ [](https://pydantic.dev/docs/ai/harness/system-reminders/#the-solution) `SystemReminders` injects reminders on each model request, either statically (`Reminder`, on a cadence) or dynamically (a callable that reads the run context). Reminders are appended to the **tail** of the request as an ephemeral `UserPromptPart` behind a `CachePoint`: * The injection runs _after_ the durable history is persisted, so the reminder reaches the model but is never written to `message_history`. No reminders accumulate across turns. * A `CachePoint` is placed immediately _before_ the reminder, so the cached prefix (tools + system + real conversation) stays byte-identical turn over turn. Only the small reminder falls outside the cache. Injecting into the system prompt (or any persisted part) instead would sit at the front of the request, so every reminder would bust the cached prefix and stale reminders would pile up in history. This capability avoids both. Usage ----- [](https://pydantic.dev/docs/ai/harness/system-reminders/#usage) Construct an `Agent` with `SystemReminders(...)` in its `capabilities`: from pydantic_ai import Agent from pydantic_ai_harness.system_reminders import SystemReminders, Reminder agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[\ SystemReminders(\ reminders=[Reminder('Stay focused on the original request.', interval=5)],\ )\ ], ) result = agent.run_sync('Refactor the auth module and add tests.') print(result.output) Static reminders ---------------- [](https://pydantic.dev/docs/ai/harness/system-reminders/#static-reminders) A `Reminder` fires on a cadence within a run: | Field | Purpose | | --- | --- | | `content` | The reminder text. | | `interval` | Fire every N model requests (`interval=3` fires on the 3rd, 6th, …). | | `first_after` | Request number of the first fire, then every `interval` after. `None` = first multiple of `interval` (plain modulo). | | `trigger` | Predicate over `RunContext`. When set, fires only when it returns `True` _and_ the cadence matches. | | `max_fires` | Cap the number of fires per run. `None` = no limit. | | `tag` | Wrap the content in `\ncontent\n`. Defaults to `'system-reminder'`; set `None` for raw content. | The default `tag='system-reminder'` wraps every reminder in `...`, following Claude Code’s convention so the model reads it as an out-of-band steering note rather than user text. The `tag` wrapping applies only to static `Reminder` content. Dynamic callables (including `GoalReanchor` and `LLMReminder`) inject their returned text raw and own their own formatting. Dynamic reminders ----------------- [](https://pydantic.dev/docs/ai/harness/system-reminders/#dynamic-reminders) A dynamic reminder is any callable `(RunContext) -> str | None` (sync or async), evaluated on every model request. Return a string to inject, or `None` to skip. This is the general seam for conditions that need run state — token budget, post-compaction, mode switches — without hardcoded detectors: from pydantic_ai_harness.system_reminders import SystemReminders SystemReminders( dynamic_reminders=[\ lambda ctx: 'Wrap up soon.' if ctx.run_step > 20 else None,\ ], ) ### `GoalReanchor` — zero-cost goal anchoring [](https://pydantic.dev/docs/ai/harness/system-reminders/#goalreanchor--zero-cost-goal-anchoring) `GoalReanchor` re-states the run’s first user request as the anchor and asks the model to check its next action advances it. No model call, no dependencies: from pydantic_ai_harness.system_reminders import SystemReminders, GoalReanchor SystemReminders(dynamic_reminders=[GoalReanchor()]) ### `LLMReminder` — model-generated nudges [](https://pydantic.dev/docs/ai/harness/system-reminders/#llmreminder--model-generated-nudges) `LLMReminder` has a model summarize a compact transcript (original goal + recent activity) into a short stay-on-task nudge. It requires an explicit `model` — there is no default model id — and falls back to `GoalReanchor` text on any error, so a failed generation never blocks the run: from pydantic_ai_harness.system_reminders import SystemReminders, LLMReminder SystemReminders(dynamic_reminders=[LLMReminder(model='anthropic:claude-haiku-4-5')]) Dynamic reminders have no cadence of their own — they run on every model request. `LLMReminder` therefore issues one extra model call per turn (its usage is threaded onto the parent run via `ctx.usage`, so it shows up in `result.usage()`). The nested call also runs under the parent’s `usage_limits` with one request held back for the model request it precedes, so the reminder cannot push a run past its `request_limit`; once the budget is that tight the generation is skipped and `GoalReanchor` text is used instead. Because the fallback is silent, a persistently misconfigured `model` (bad id, missing key) looks like normal operation. To bound the cost, gate it behind a cadence with an async wrapper: _llm = LLMReminder(model='anthropic:claude-haiku-4-5') async def every_tenth(ctx): return await _llm(ctx) if ctx.run_step % 10 == 0 else None SystemReminders(dynamic_reminders=[every_tenth]) `LLMReminder` generates inside `wrap_model_request`, so under durable execution (Temporal, DBOS, Prefect) its model call runs in orchestration context rather than a durable step — non-deterministic on replay and not checkpointed, with errors falling back silently to `GoalReanchor`. For durable runs prefer `GoalReanchor` (no model call) or gate `LLMReminder` off. Configuration ------------- [](https://pydantic.dev/docs/ai/harness/system-reminders/#configuration) from pydantic_ai_harness.system_reminders import SystemReminders, Reminder SystemReminders( reminders=[Reminder('...', interval=5)], dynamic_reminders=[], # callables evaluated every request cache_ttl='5m', # TTL for the cache breakpoint before the reminder ('5m' | '1h') on_fire=None, # optional callback invoked with each rendered reminder ) Per-run state (the request counter and per-reminder fire counts) is isolated via `for_run`, so concurrent runs on the same agent never share fire state. Caching guarantee ----------------- [](https://pydantic.dev/docs/ai/harness/system-reminders/#caching-guarantee) Reminders are never injected into the system prompt or instructions. They ride the ephemeral tail behind a `CachePoint`, so across turns: * the durable history grows append-only and is replayed byte-identically, so the whole prefix stays eligible for a cache hit (subject to the provider’s cache TTL — a gap longer than `cache_ttl` expires the entry even under an unchanged prefix); * the reminder and its `CachePoint` live only in the per-request copy, so they can’t invalidate anything and aren’t persisted. `CachePoint` is supported on Anthropic, Amazon Bedrock (Converse API), and OpenRouter (Anthropic and Gemini models); on providers without prompt caching it’s simply ignored (nothing to bust). The reminder leads with its `CachePoint` only when the request already carries user content for the breakpoint to attach to — on a turn whose only tail content is the reminder (for example an `instructions`\-only run’s first request), the reminder is injected without a breakpoint, since there is no prefix to protect. Composition ----------- [](https://pydantic.dev/docs/ai/harness/system-reminders/#composition) * [Planning](https://pydantic.dev/docs/ai/harness/planning/) uses the same ephemeral-tail mechanism to surface the plan. Both compose in one agent: each appends its own tail part behind its own `CachePoint`, and neither is persisted. Note that each ephemeral-tail capability adds a cache breakpoint: Anthropic allows 4 (3 with automatic caching), and core trims the excess oldest-first, so stacking several tail-injecting capabilities alongside `anthropic_cache_instructions` / `anthropic_cache_tool_definitions` can evict an older breakpoint. Two capabilities plus the defaults stay within budget. * Loop detection (detect-and-interrupt with a durable nudge) is a separate concern. `SystemReminders` is cadence/condition steering that stays ephemeral; a dynamic reminder can read loop state from your deps if you want to steer on it. The tail reminder is only appended when the last message in the request is a `ModelRequest` and at least one reminder fires, so a turn where nothing fires adds nothing to the request. Provider-resume turns (where the request tail is a suspended `ModelResponse` that is echoed back verbatim) are skipped and do not consume a cadence slot. Not spec-serializable --------------------- [](https://pydantic.dev/docs/ai/harness/system-reminders/#not-spec-serializable) `SystemReminders.get_serialization_name()` returns `None`: reminders take arbitrary callables, which cannot be serialized to an [agent spec](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) . Further reading --------------- [](https://pydantic.dev/docs/ai/harness/system-reminders/#further-reading) * [Pydantic AI capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) * [Hooks](https://pydantic.dev/docs/ai/core-concepts/hooks/) — `wrap_model_request` is the ephemeral injection point used here * [Anthropic prompt caching](https://docs.claude.com/en/docs/build-with-claude/prompt-caching) * [Planning](https://pydantic.dev/docs/ai/harness/planning/) — another prompt-cache-aware harness capability API reference ------------- [](https://pydantic.dev/docs/ai/harness/system-reminders/#api-reference) SystemReminders --------------- [](https://pydantic.dev/docs/ai/harness/system-reminders/#pydantic_ai_harness.system_reminders.SystemReminders) **Bases:** `AbstractCapability[AgentDepsT]` Inject periodic or conditional reminders to counter instruction fade in long sessions. Long multi-turn runs suffer instruction fade: after many tool-use turns the model progressively ignores start-of-session guidance. `SystemReminders` re-injects targeted guidance mid-run, either on a fixed cadence (`Reminder`) or reactively from a callable (`dynamic_reminders`). Cache safety is the design constraint. Reminders are appended to the _tail_ of each request as an ephemeral `UserPromptPart` behind a `CachePoint`, inside `wrap_model_request` (which runs after core persists the durable history). So reminders reach the model but never enter `message_history`: no stale reminders accumulate, and the cached prefix stays byte-identical across turns — only the small reminder falls outside the cache. Injecting into the system prompt or a persisted part instead would bust the cache prefix on every fire and let reminders pile up. from pydantic_ai import Agent from pydantic_ai_harness.system_reminders import SystemReminders, Reminder agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[\ SystemReminders(\ reminders=[Reminder('Stay focused on the original request.', interval=5)],\ )\ ], ) ### Attributes [](https://pydantic.dev/docs/ai/harness/system-reminders/#attributes) #### cache\_ttl [](https://pydantic.dev/docs/ai/harness/system-reminders/#pydantic_ai_harness.system_reminders.SystemReminders.cache_ttl) TTL for the cache breakpoint placed before the tail reminder. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘5m’, ‘1h’\] **Default:** `'5m'` #### dynamic\_reminders [](https://pydantic.dev/docs/ai/harness/system-reminders/#pydantic_ai_harness.system_reminders.SystemReminders.dynamic_reminders) Callables evaluated every model request; return text to inject or `None` to skip. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[`DynamicReminder`\[`AgentDepsT`\] | `AsyncDynamicReminder`\[`AgentDepsT`\]\] **Default:** `()` #### on\_fire [](https://pydantic.dev/docs/ai/harness/system-reminders/#pydantic_ai_harness.system_reminders.SystemReminders.on_fire) Optional observability callback invoked with each rendered reminder as it fires. **Type:** [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \], [`None`](https://docs.python.org/3/library/constants.html#None)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### reminders [](https://pydantic.dev/docs/ai/harness/system-reminders/#pydantic_ai_harness.system_reminders.SystemReminders.reminders) Static reminders injected on a cadence. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[`Reminder`\[`AgentDepsT`\]\] **Default:** `()` ### Methods [](https://pydantic.dev/docs/ai/harness/system-reminders/#methods) #### for\_run [](https://pydantic.dev/docs/ai/harness/system-reminders/#pydantic_ai_harness.system_reminders.SystemReminders.for_run) `@async` def for_run(ctx: RunContext[AgentDepsT]) -> SystemReminders[AgentDepsT] Return a fresh per-run instance with reset counters (config preserved). `replace` builds a new instance whose generated `__init__` re-initializes the `init=False` fields — `_request_count` back to `0` and `_fire_counts` to an empty dict — so concurrent runs on the same agent never share fire state. ##### Returns [](https://pydantic.dev/docs/ai/harness/system-reminders/#returns) `SystemReminders`\[`AgentDepsT`\] #### get\_serialization\_name [](https://pydantic.dev/docs/ai/harness/system-reminders/#pydantic_ai_harness.system_reminders.SystemReminders.get_serialization_name) `@classmethod` def get_serialization_name(cls) -> str | None Not spec-serializable: reminders take arbitrary callables. ##### Returns [](https://pydantic.dev/docs/ai/harness/system-reminders/#returns-1) [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### wrap\_model\_request [](https://pydantic.dev/docs/ai/harness/system-reminders/#pydantic_ai_harness.system_reminders.SystemReminders.wrap_model_request) `@async` def wrap_model_request( ctx: RunContext[AgentDepsT], *, request_context: ModelRequestContext, handler: WrapModelRequestHandler, ) -> ModelResponse Append fired reminders to the request tail behind a cache breakpoint, then call the model. Runs after core persists the durable history; the per-request message list mutated here is never written back, so the reminder and its `CachePoint` reach the model but never enter `ctx.state.message_history`. ##### Returns [](https://pydantic.dev/docs/ai/harness/system-reminders/#returns-2) [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) Reminder -------- [](https://pydantic.dev/docs/ai/harness/system-reminders/#pydantic_ai_harness.system_reminders.Reminder) **Bases:** `Generic[AgentDepsT]` A static reminder injected on a cadence during an agent run. ### Attributes [](https://pydantic.dev/docs/ai/harness/system-reminders/#attributes-1) #### content [](https://pydantic.dev/docs/ai/harness/system-reminders/#pydantic_ai_harness.system_reminders.Reminder.content) The reminder text. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### first\_after [](https://pydantic.dev/docs/ai/harness/system-reminders/#pydantic_ai_harness.system_reminders.Reminder.first_after) Request number of the first fire. `None` (the default) fires on the first multiple of `interval` (plain modulo). When set, the reminder fires at `first_after`, then every `interval` requests after that. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### interval [](https://pydantic.dev/docs/ai/harness/system-reminders/#pydantic_ai_harness.system_reminders.Reminder.interval) Fire every N model requests within a run. `interval=3` fires on the 3rd, 6th, 9th, … request. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `1` #### max\_fires [](https://pydantic.dev/docs/ai/harness/system-reminders/#pydantic_ai_harness.system_reminders.Reminder.max_fires) Maximum number of times this reminder may fire within a run. `None` means no limit. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### tag [](https://pydantic.dev/docs/ai/harness/system-reminders/#pydantic_ai_harness.system_reminders.Reminder.tag) When set, wrap the content in an XML tag: `\ncontent\n`. Defaults to `'system-reminder'` (Claude Code’s convention); set `None` to emit the raw content. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `'system-reminder'` #### trigger [](https://pydantic.dev/docs/ai/harness/system-reminders/#pydantic_ai_harness.system_reminders.Reminder.trigger) Optional predicate over the current `RunContext`. When set, the reminder fires only when the trigger returns `True` _and_ the cadence condition is met. **Type:** [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[[`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext)\ \[`AgentDepsT`\]\], [`bool`](https://docs.python.org/3/library/functions.html#bool)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` GoalReanchor ------------ [](https://pydantic.dev/docs/ai/harness/system-reminders/#pydantic_ai_harness.system_reminders.GoalReanchor) **Bases:** `Generic[AgentDepsT]` Zero-cost dynamic reminder that re-states the run’s first user request as the anchor. No model call and no dependencies: it reads the first user message from `ctx.messages` and asks the model to check its next action advances that goal. Falls back to a static line when there is no user message yet. Add it to `SystemReminders.dynamic_reminders`. LLMReminder ----------- [](https://pydantic.dev/docs/ai/harness/system-reminders/#pydantic_ai_harness.system_reminders.LLMReminder) **Bases:** `Generic[AgentDepsT]` Dynamic reminder whose text a model generates from a compact transcript. Opt-in and dependency-free (it uses `pydantic_ai.Agent`). `model` is required and has no default — pass an explicit model. On any error it falls back to `GoalReanchor` text, so a failed generation never blocks the run. Add it to `SystemReminders.dynamic_reminders`. Like every dynamic reminder it is evaluated on every model request, so it issues one extra model call per turn; its usage is threaded onto the parent run (`ctx.usage`) and it runs under the parent’s `usage_limits` minus one reserved request, so the reminder cannot push a run past its `request_limit`. Once the budget is that tight the generation is skipped and `GoalReanchor` text is used instead. Gate it on a cadence (see the docs) if per-turn generation is too costly. The generation runs inside `wrap_model_request`, so under durable execution (Temporal, DBOS, Prefect) it executes in orchestration context rather than a durable step: the model call is non-deterministic on replay and is not checkpointed, and its errors fall back silently to `GoalReanchor`. For durable runs prefer `GoalReanchor`, which makes no model call, or gate `LLMReminder` off. Was this page helpful? Thanks for your feedback! --- # ACP (Agent Client Protocol) [Skip to content](https://pydantic.dev/docs/ai/harness/acp/#_top) ACP === Editors like [Zed](https://zed.dev/docs/ai/external-agents) speak ACP: a stdio JSON-RPC protocol that lets a TUI or editor drive an external coding agent — streaming its text, rendering its file edits as diffs, and prompting the user to approve sensitive tool calls. Reach for this capability when you want a Pydantic AI `Agent` to appear as a first-class agent inside one of those editors, without implementing the ACP server side yourself. The problem ----------- [](https://pydantic.dev/docs/ai/harness/acp/#the-problem) To plug a Pydantic AI agent into an ACP editor you would otherwise have to implement the ACP server side by hand — chunking streamed text under the wire limit, rendering tool calls as diffs, mapping the protocol’s permission requests onto the agent’s tools, and managing per-workspace sessions. The solution ------------ [](https://pydantic.dev/docs/ai/harness/acp/#the-solution) `run_acp_stdio` serves any Pydantic AI `Agent` as an ACP agent over stdin/stdout. The editor launches your script as a subprocess and talks to it; the adapter translates between ACP and the agent’s run loop: | ACP needs | The adapter provides | | --- | --- | | Streamed assistant text and reasoning | Agent text/thinking deltas, chunked under the wire limit | | Rich tool calls (`kind`, file `locations`, diffs) | A presenter that recognizes `FileSystem`/`Shell` tool calls | | Human-in-the-loop tool approval | Maps ACP permission requests to Pydantic AI’s deferred-approval tools | | Per-workspace sessions | A `session_config` hook to root tools at the client’s working directory | | Cancellation, multi-turn history, session close | Handled per session | Installation ------------ [](https://pydantic.dev/docs/ai/harness/acp/#installation) Terminal uv add "pydantic-ai-harness[acp]" This pulls in the [`agent-client-protocol`](https://pypi.org/project/agent-client-protocol/) SDK. The rest of the harness does not depend on it — only `pydantic_ai_harness.experimental.acp` does. Quick start ----------- [](https://pydantic.dev/docs/ai/harness/acp/#quick-start) Write a script that builds your agent and serves it: # my_acp_agent.py from pydantic_ai import Agent from pydantic_ai_harness.experimental.acp import run_acp_stdio_sync def build_agent() -> Agent[None, str]: return Agent('anthropic:claude-sonnet-4-6', instructions='You are a coding assistant.') if __name__ == '__main__': run_acp_stdio_sync(build_agent()) `run_acp_stdio_sync` blocks for the lifetime of the connection — it is the `main()` of an agent the editor launches. Inside an existing event loop, use the async `run_acp_stdio` instead. Connecting from an editor ------------------------- [](https://pydantic.dev/docs/ai/harness/acp/#connecting-from-an-editor) ACP clients launch the agent as a subprocess. In Zed, register it as an [external agent](https://zed.dev/docs/ai/external-agents) in `settings.json`: { "agent_servers": { "My Pydantic AI Agent": { "type": "custom", "command": "python", "args": ["/absolute/path/to/my_acp_agent.py"], "env": { "ANTHROPIC_API_KEY": "..." } } } } Any ACP-compatible client works the same way — point it at `python my_acp_agent.py`. The provider environment must be available to the launched subprocess. GUI editors and SDK-based test wrappers may not source your interactive shell startup files. If a real-model agent exits before initialize or fails provider auth, first verify that the command’s process can see variables such as `ANTHROPIC_API_KEY`. Rooting tools at the workspace ------------------------------ [](https://pydantic.dev/docs/ai/harness/acp/#rooting-tools-at-the-workspace) A coding agent should read and write files in the workspace the editor opened, not wherever the subprocess started. ACP gives each session a working directory (`cwd`); a `session_config` factory turns that into per-session tools: from pydantic_ai import Agent from pydantic_ai_harness.experimental.acp import AcpSession, AcpSessionConfig, run_acp_stdio_sync from pydantic_ai_harness.filesystem import FileSystem from pydantic_ai_harness.shell import Shell agent = Agent('anthropic:claude-sonnet-4-6') def session_config(session: AcpSession) -> AcpSessionConfig[None]: # Root file and shell tools at the workspace the client opened. return AcpSessionConfig( deps=None, toolsets=[\ FileSystem[None](root_dir=session.cwd).get_toolset(),\ Shell[None](cwd=session.cwd).get_toolset(),\ ], ) if __name__ == '__main__': run_acp_stdio_sync(agent, session_config=session_config) The factory runs once per session with the client’s `AcpSession` setup (its `cwd`, `mcp_servers`, and capabilities) and returns an `AcpSessionConfig` whose `deps` and `toolsets` apply to every run in that session. This is correct across multiple concurrent sessions in one process, where a single static `FileSystem` could not be. Editor-native filesystem and shell (optional) --------------------------------------------- [](https://pydantic.dev/docs/ai/harness/acp/#editor-native-filesystem-and-shell-optional) The local [`FileSystem`](https://pydantic.dev/docs/ai/harness/filesystem/) and [`Shell`](https://pydantic.dev/docs/ai/harness/shell/) above operate on the agent process’s own disk and subprocesses. An editor’s source of truth is different: unsaved buffers, its own idea of the workspace layout, and — for a remote or containerized editor — the machine the code actually lives on. When the client advertises support, `acp_filesystem` and `acp_terminal` give the agent `read_file`/`write_file`/`run_command` tools that route through the client, so it acts where the user is: from pydantic_ai_harness.experimental.acp import AcpSession, AcpSessionConfig, acp_filesystem, acp_terminal from pydantic_ai_harness.filesystem import FileSystem from pydantic_ai_harness.shell import Shell def session_config(session: AcpSession) -> AcpSessionConfig[None]: # Use the editor's filesystem/terminal when offered; otherwise fall back to local. fs = acp_filesystem(session) or FileSystem[None](root_dir=session.cwd).get_toolset() shell = acp_terminal(session) or Shell[None](cwd=session.cwd).get_toolset() return AcpSessionConfig(deps=None, toolsets=[fs, shell]) Each helper returns `None` when the client did not advertise the capability, so the `or` falls back to local and the agent works either way. The tool names match the local `FileSystem`/`Shell`, so rich rendering stays identical. Tool approval ------------- [](https://pydantic.dev/docs/ai/harness/acp/#tool-approval) Mark a tool to require approval and ACP relays the decision to the client, which shows the user an approve/reject prompt: @agent.tool_plain(requires_approval=True) def delete_file(path: str) -> str: ... The lifecycle the client sees is `pending` (awaiting approval) -> `in_progress` (granted, running) -> `completed`/`failed`, so an unapproved action is never shown as already running. “Always allow”/“always reject” decisions are remembered for the session, scoped by default to the exact call (tool name plus arguments) so approving one call never silently approves a different one. Pass `permission_policy` to widen or narrow that scope. Rich tool rendering ------------------- [](https://pydantic.dev/docs/ai/harness/acp/#rich-tool-rendering) By default the adapter recognizes the harness `FileSystem` and `Shell` tool calls by name and annotates them with an ACP `kind` (`read`/`edit`/`search`/`execute`), the file `locations` they touch, and an inline diff for edits — so the editor renders click-to-file links and diff views instead of opaque JSON. Pass `tool_presenter` to add rendering for your own tools (optionally with `chain_presenters` ahead of the default `default_coding_presenter`), or `lambda _call: None` to disable it. MCP servers ----------- [](https://pydantic.dev/docs/ai/harness/acp/#mcp-servers) An ACP client may offer MCP servers during session setup. This adapter does not connect them itself; a `session_config` is the place to turn `session.mcp_servers` into Pydantic AI toolsets (for example with `pydantic_ai.mcp.MCPServerStdio`). If a client sends MCP servers and no `session_config` is installed to consume them, the session request is rejected rather than silently ignoring them. A spec-following client only sends HTTP/SSE MCP servers when the agent advertises support during `initialize`; when your `session_config` connects them, say so: from acp import schema from pydantic_ai_harness.experimental.acp import PydanticAIACPAgent PydanticAIACPAgent( agent, session_config=connect_mcp_servers, mcp_capabilities=schema.McpCapabilities(http=True, sse=True), ) Prompt content types -------------------- [](https://pydantic.dev/docs/ai/harness/acp/#prompt-content-types) The agent advertises which prompt content it accepts. The default is **text only**, so a client is not invited to send blocks a text model cannot handle. Enable the kinds your model supports: from acp import schema run_acp_stdio_sync(agent, prompt_capabilities=schema.PromptCapabilities(image=True, embedded_context=True)) Session persistence ------------------- [](https://pydantic.dev/docs/ai/harness/acp/#session-persistence) Pass a `session_store` to let a client reopen a past conversation with `session/load`. Each committed turn is persisted as two parts — the model’s message history and the client-visible transcript — and reopening restores the history into the agent and replays the transcript to the client, so its UI is rebuilt as the user last saw it. Without a store, `session/load` is advertised as unsupported. from pydantic_ai_harness.experimental.acp import InMemorySessionStore run_acp_stdio_sync(agent, session_store=InMemorySessionStore()) `InMemorySessionStore` keeps sessions for the lifetime of the process. Implement the `SessionStore` protocol (`save`/`load` a `StoredSession`) over a file or database to make them survive a restart — the stored values are Pydantic models, so they serialize with Pydantic. Session persistence is for _reopening a conversation_; it is orthogonal to per-run durability. To also make individual turns crash-resilient, add a [step-durability capability](https://pydantic.dev/docs/ai/harness/step-persistence/) — each ACP turn is one agent run, so the two layers compose with no glue. Model selection --------------- [](https://pydantic.dev/docs/ai/harness/acp/#model-selection) Pass `models` to advertise a stable ACP session config option named `model` (using Pydantic AI [model names](https://pydantic.dev/docs/ai/models/) ). The first is each session’s default. A selection is applied as a per-run override — the shared agent is never mutated — and is persisted with the session when a `session_store` is set. run_acp_stdio_sync(agent, models=['anthropic:claude-sonnet-4-6', 'anthropic:claude-opus-4-8', 'openai:gpt-4o']) A model id is any string a Pydantic AI model accepts, so newer models not yet in `KnownModelName` work too. Pass `models='all'` to offer every model Pydantic AI knows. To advertise ids `infer_model` does not understand (OAuth or subscription models), pass `model_resolver` to map the selected id to a prebuilt `Model`. Cancellation and limitations ---------------------------- [](https://pydantic.dev/docs/ai/harness/acp/#cancellation-and-limitations) * **Cancellation.** `session/cancel` and `session/close` cancel the in-flight turn; close waits for it to unwind before returning. Cooperative async tools stop promptly. A synchronous tool already running in a worker thread cannot be force-stopped, so prefer async tools for cancellation-sensitive work. * **Approval detection.** Tools that require approval are recognized when they live in a `FunctionToolset` (which the harness `FileSystem`/`Shell` and `@agent.tool` all use). A tool whose approval requirement is decided dynamically per call (by raising `ApprovalRequired` from its body) starts as `in_progress`, and any side effects it ran before raising have already happened — use an `ApprovalRequiredToolset` for actions that must not partially execute before approval. * **Overwrite diffs.** `write_file` renders an overwrite as if creating a new file, so the diff understates what it replaced. * **Live terminal panes.** `acp_terminal` returns a command’s captured output; it does not embed a live terminal pane in the tool call. * **Images.** Prompt image blocks are off by default and must be enabled via `prompt_capabilities` with a model that accepts them. * **Slash commands.** The adapter does not yet advertise any commands, so no slash commands appear in the client. Planned. API --- [](https://pydantic.dev/docs/ai/harness/acp/#api) run_acp_stdio( # async; serve until the client disconnects agent, *, deps=None, name=None, # advertised name; defaults to the agent's name version='0.1.0', session_config=None, # per-session deps/toolsets from the client's setup permission_policy=None, # scope of remembered "always" approval decisions prompt_capabilities=None, # defaults to text-only mcp_capabilities=None, # MCP transports to advertise; needs a session_config to connect them tool_presenter=None, # defaults to the FileSystem/Shell presenter session_store=None, # enables session/load by persisting each session models=None, # models offered as the `model` config option ('all' for every known model) model_resolver=None, # maps an advertised model id to the Model used for the run usage_limits=None, # per-run request/token ceilings ) run_acp_stdio_sync(...) # synchronous wrapper, same arguments PydanticAIACPAgent(agent, *, ...) # the ACP agent object, to embed in a custom server The module also exports the session types (`AcpSession`, `AcpSessionConfig`, `McpServer`), the store types (`SessionStore`, `StoredSession`, `InMemorySessionStore`), the client toolsets (`AcpFileSystemToolset`, `AcpTerminalToolset`, `acp_filesystem`, `acp_terminal`), the permission types (`ToolCallPermission`, `default_permission_scope`), and the presentation helpers (`ToolCallPresentation`, `chain_presenters`, `default_coding_presenter`). Source: [`pydantic_ai_harness/experimental/acp/`](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/experimental/acp/) . Further reading --------------- [](https://pydantic.dev/docs/ai/harness/acp/#further-reading) * [Agent Client Protocol](https://agentclientprotocol.com/) — protocol specification * [Zed external agents](https://zed.dev/docs/ai/external-agents) — editor-side configuration * [Human-in-the-loop tool approval](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/#human-in-the-loop-tool-approval) (Pydantic AI) * [Pydantic AI capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) Was this page helpful? Thanks for your feedback! --- # Apache Airflow | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/durable_execution/airflow/#_top) Apache Airflow ============== [Apache Airflow](https://airflow.apache.org/) is a workflow orchestrator. Its Pydantic AI integration is provided by the [`apache-airflow-providers-common-ai`](https://airflow.apache.org/docs/apache-airflow-providers-common-ai/stable/index.html) package through `airflow.providers.common.ai`, rather than by `pydantic_ai.durable_exec`. Unlike the wrapper-object integrations on this page, Airflow’s durable unit is an **Airflow task**. You author a normal Pydantic AI agent and run it as a task; Airflow’s retry machinery plus a step-level cache resume the agent from its last completed model request or tool call instead of replaying the whole run. Durable Execution ----------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/airflow/#durable-execution) When an agent runs as a durable Airflow task, Airflow records each completed **model request** and **tool call** as a cache entry. On a retry, Airflow replays these entries to skip completed work: each entry stores a fingerprint of the request that produced it, and if that fingerprint no longer matches (the conversation diverged since the previous attempt), the step re-runs live instead of returning a stale result. The cache lives in object storage (local, S3, GCS, or Azure) for the lifetime of a single DAG run’s task and is deleted when the task succeeds. For example, imagine an agent calls a model, gets a useful response, starts a tool call, and then the worker crashes. Without durable execution, Airflow’s normal retry restarts the task from the top and repeats the model request and any later side effects. With durable execution, the retry replays the run, reuses the cached result for the already-completed model request, and continues from the first operation that has not completed. This is useful for long-running agents and for runs where a repeated model request or external tool call would cost money, take time, or duplicate a side effect. Durable Agent ------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/airflow/#durable-agent) You make an agent durable by running it through Airflow’s `AgentOperator` (or the `@task.agent` decorator) with `durable=True`. Install the provider alongside Airflow: Terminal uv add "apache-airflow-providers-common-ai" The agent’s model and credentials come from an Airflow connection (the examples use `pydanticai_default`). See [Pydantic AI connection](https://airflow.apache.org/docs/apache-airflow-providers-common-ai/stable/connections/pydantic_ai.html) for how to configure one. Durable execution needs a place to store its step cache. Point `[common.ai] durable_cache_path` at an object-storage location: airflow.cfg [common.ai] durable_cache_path = s3://my-bucket/airflow-agent-cache Here is the smallest durable Pydantic AI agent as an Airflow task: durable\_agent\_dag.py from datetime import timedelta from airflow.providers.common.ai.operators.agent import AgentOperator from airflow.sdk import dag @dag(default_args={"retries": 3, "retry_delay": timedelta(seconds=30)}) def durable_agent_dag(): AgentOperator( task_id="researcher", prompt="Summarize quantum error correction.", llm_conn_id="pydanticai_default", durable=True, ) durable_agent_dag() `AgentOperator` does not replace the underlying Pydantic AI agent. It builds the agent from your connection and toolsets, then wraps the model and toolsets so that, while the task runs, it records recoverable operations: * model requests; * Pydantic AI tool calls. The same applies to the `@task.agent` decorator, where the decorated function returns the prompt: durable\_agent\_decorator.py from datetime import timedelta from airflow.providers.common.ai.toolsets.sql import SQLToolset from airflow.sdk import dag, task @dag(default_args={"retries": 3, "retry_delay": timedelta(seconds=30)}) def durable_agent_decorator(): @task.agent( llm_conn_id="pydanticai_default", system_prompt="You are a data analyst. Use tools to answer questions.", durable=True, toolsets=[SQLToolset(db_conn_id="postgres_default", allowed_tables=["orders"])], ) def analyze(question: str) -> str: return f"Answer this question about our orders data: {question}" analyze("What was our total revenue last month?") durable_agent_decorator() The agent’s retries are Airflow task retries: configure them with the task’s `retries` and `retry_delay`. On each retry the cached steps are replayed and the run continues from the first operation that has not completed. For the full reference, see the Airflow [`AgentOperator` durable execution docs](https://airflow.apache.org/docs/apache-airflow-providers-common-ai/stable/operators/agent.html#durable-execution) . Tools and side effects ---------------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/airflow/#tools-and-side-effects) Durable execution caches the result of each Pydantic AI tool call, including tools backed by Airflow toolsets (`SQLToolset`, `HookToolset`, `MCPToolset`, and others). On replay the cached result is returned without re-invoking the tool, so a tool that writes to an external system runs at most once per completed step across all retries of a run. Human-in-the-loop ----------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/airflow/#human-in-the-loop) Airflow’s `AgentOperator` has a separate human-in-the-loop review mode (`enable_hitl_review=True`) that pauses an agent run for human approval, rejection, or change requests through Airflow’s HITL UI. Streaming --------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/airflow/#streaming) Streaming is not yet supported under `durable=True`. The durable model wrapper records complete model requests, not streamed events. Run streaming agents as non-durable tasks. Requirements and Constraints ---------------------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/airflow/#requirements-and-constraints) When running a Pydantic AI agent as a durable Airflow task: * The durable unit is the Airflow task; recovery happens through Airflow task retries, so set `retries` (and a `retry_delay`) on the task. * Define the agent with a concrete model, for example via a connection that resolves to `Agent('openai:gpt-5-nano', ...)`. The model must be set when `durable=True`. * Set `[common.ai] durable_cache_path` to an object-storage location the workers can read and write. * On a retry, a cached model request or tool call is replayed when its stored fingerprint matches the current request, and re-runs live when it diverges. Requests that can’t be serialized to a fingerprint fall back to unverified positional replay, so keep runs deterministic across retries. * `durable=True` and `enable_hitl_review=True` are mutually exclusive. * Streaming is not supported under `durable=True`. * The step cache is scoped to one DAG run’s task and is deleted when the task succeeds. Was this page helpful? Thanks for your feedback! --- # pydantic_graph.node | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic_graph/node/#_top) pydantic\_graph.node ==================== Core node types for graph construction and execution. This module defines the fundamental node types used to build execution graphs, including start/end nodes and fork nodes for parallel execution. EndNode ------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/node/#pydantic_graph.node.EndNode) **Bases:** `Generic[InputT]` Terminal node representing the completion of graph execution. The EndNode marks the successful completion of a graph execution flow and can collect the final output data. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/node/#attributes) #### id [](https://pydantic.dev/docs/ai/api/pydantic_graph/node/#pydantic_graph.node.EndNode.id) Fixed identifier for the end node. **Default:** `NodeID('__end__')` Fork ---- [](https://pydantic.dev/docs/ai/api/pydantic_graph/node/#pydantic_graph.node.Fork) **Bases:** `Generic[InputT, OutputT]` Fork node that creates parallel execution branches. A Fork node splits the execution flow into multiple parallel branches, enabling concurrent execution of downstream nodes. It can either map a sequence across multiple branches or duplicate data to each branch. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/node/#attributes-1) #### downstream\_join\_id [](https://pydantic.dev/docs/ai/api/pydantic_graph/node/#pydantic_graph.node.Fork.downstream_join_id) Optional identifier of a downstream join node that should be jumped to if mapping an empty iterable. **Type:** `JoinID` | [`None`](https://docs.python.org/3/library/constants.html#None) #### id [](https://pydantic.dev/docs/ai/api/pydantic_graph/node/#pydantic_graph.node.Fork.id) Unique identifier for this fork node. **Type:** `ForkID` #### is\_map [](https://pydantic.dev/docs/ai/api/pydantic_graph/node/#pydantic_graph.node.Fork.is_map) Determines fork behavior. If True, InputT must be Sequence\[OutputT\] and each element is sent to a separate branch. If False, InputT must be OutputT and the same data is sent to all branches. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) StartNode --------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/node/#pydantic_graph.node.StartNode) **Bases:** `Generic[OutputT]` Entry point node for graph execution. The StartNode represents the beginning of a graph execution flow. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/node/#attributes-2) #### id [](https://pydantic.dev/docs/ai/api/pydantic_graph/node/#pydantic_graph.node.StartNode.id) Fixed identifier for the start node. **Default:** `NodeID('__start__')` InputT ------ [](https://pydantic.dev/docs/ai/api/pydantic_graph/node/#pydantic_graph.node.InputT) Type variable for node input data. **Default:** `TypeVar('InputT', infer_variance=True)` OutputT ------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/node/#pydantic_graph.node.OutputT) Type variable for node output data. **Default:** `TypeVar('OutputT', infer_variance=True)` StateT ------ [](https://pydantic.dev/docs/ai/api/pydantic_graph/node/#pydantic_graph.node.StateT) Type variable for graph state. **Default:** `TypeVar('StateT', infer_variance=True)` Was this page helpful? Thanks for your feedback! --- # Turns and interruptions | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/realtime/turns/#_top) Turns and interruptions ======================= Realtime providers normally use voice activity detection (VAD) to decide when the user starts and stops speaking and when the model should respond. Pydantic AI exposes a shared configuration for portable behavior, explicit interruption for providers that support it, and manual turn control for push-to-talk applications. Automatic turn detection ------------------------ [](https://pydantic.dev/docs/ai/realtime/turns/#automatic-turn-detection) Automatic detection is enabled by default. Configure common behavior with [`TurnDetection`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.TurnDetection) : `sensitivity` maps to the closest provider control, while `prefix_padding_ms` and `silence_duration_ms` pass through where supported. from pydantic_ai.realtime.openai import OpenAIRealtimeModel, OpenAIRealtimeModelSettings settings = OpenAIRealtimeModelSettings( turn_detection={'sensitivity': 'high', 'silence_duration_ms': 400} ) model = OpenAIRealtimeModel('gpt-realtime', settings=settings) Use provider-specific settings only when the shared controls are insufficient: `openai_turn_detection`, `xai_turn_detection`, and `google_vad` fully override `turn_detection`. Their accepted values, defaults, and limitations are documented on the [OpenAI](https://pydantic.dev/docs/ai/realtime/openai/#settings) , [Azure OpenAI](https://pydantic.dev/docs/ai/realtime/azure/#settings) , [Google Gemini](https://pydantic.dev/docs/ai/realtime/gemini/#settings) , and [xAI](https://pydantic.dev/docs/ai/realtime/xai/#settings) pages. Barge-in -------- [](https://pydantic.dev/docs/ai/realtime/turns/#barge-in) With server-side turn detection, providers interrupt the model when they detect new user speech. Your application still owns audio already queued for playback and must flush that local buffer. Providers whose profile declares [`emits_input_speech_events`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelProfile.emits_input_speech_events) (OpenAI, Azure OpenAI, and xAI) emit [`RealtimeInputSpeechStartEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeInputSpeechStartEvent) when user speech begins. Gemini emits [`RealtimeResponseInterruptedEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeResponseInterruptedEvent) when it interrupts model output instead. These are the signals to flush playback; read the flag rather than waiting on an event a provider never sends. [`interrupt()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.interrupt) handles the server-side half of the problem. When supported, pass how many milliseconds actually played so the provider does not record unheard words as part of the conversation. `Speaker` here stands in for your playback layer — anything that can report and flush buffered audio: from typing import Protocol from pydantic_ai.realtime import RealtimeInputSpeechStartEvent, RealtimeSession class Speaker(Protocol): def has_unplayed_audio(self) -> bool: ... def flush(self) -> None: ... def played_ms(self) -> int: ... async def handle_events(session: RealtimeSession, speaker: Speaker): async for event in session: if isinstance(event, RealtimeInputSpeechStartEvent) and speaker.has_unplayed_audio(): speaker.flush() if session.profile['supports_output_truncation']: await session.interrupt(played_ms=speaker.played_ms()) elif session.profile['supports_interruption']: await session.interrupt() The speech-start event also occurs on ordinary user turns when nothing is playing. Track unplayed audio before interrupting. `interrupt()` never flushes the local speaker buffer. On a [WebRTC sideband](https://pydantic.dev/docs/ai/realtime/deployment/#browser-webrtc-server-sideband) there is a third buffer between those two: the provider generates audio well ahead of playback and keeps streaming what it already produced, so stopping the model is not enough to stop the voice. `interrupt()` drops that outbound buffer too, which is what actually ends the turn for the listener. The browser still owns its own playback buffer and should flush it on barge-in, as above. History records a known cutoff on [`SpeechPart.interrupted_at_ms`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SpeechPart.interrupted_at_ms) and marks the response state as interrupted. When this history is sent to a text model, Pydantic AI adds a readable interruption note to the prepared request without modifying stored history. Push-to-talk ------------ [](https://pydantic.dev/docs/ai/realtime/turns/#push-to-talk) Disable automatic detection with `turn_detection=False` on models whose profile declares `supports_manual_turn_control`. Stream audio, call [`commit_audio()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.commit_audio) to end the user turn, then [`create_response()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.create_response) . The explicit `create_response()` call is needed because with turn detection off, committing the buffer only finalizes the user’s input; nothing triggers a reply until you ask for one. Use [`clear_audio()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.clear_audio) to discard uncommitted input. from pydantic_ai import Agent from pydantic_ai.realtime.openai import OpenAIRealtimeModel, OpenAIRealtimeModelSettings agent = Agent() model = OpenAIRealtimeModel( 'gpt-realtime', settings=OpenAIRealtimeModelSettings(turn_detection=False) ) async def main(): async with agent.realtime(model).session() as session: await session.send_audio(b'...') await session.commit_audio() await session.create_response() Gemini does not expose manual turn verbs through Pydantic AI; `turn_detection=False` raises [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) before connecting. Checking what the model supports -------------------------------- [](https://pydantic.dev/docs/ai/realtime/turns/#checking-what-the-model-supports) These are [_model profile_](https://pydantic.dev/docs/ai/models/overview/#inspecting-a-models-profile) flags describing what a provider connection can do — not to be confused with [capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) , which add behavior to an agent. Branch on [`RealtimeModelProfile`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelProfile) rather than provider names: | Profile flag | Gates | | --- | --- | | `supports_manual_turn_control` | `commit_audio()`, `clear_audio()`, and `create_response()` | | `supports_interruption` | `interrupt()` | | `supports_output_truncation` | `interrupt(played_ms=...)` | Calling an unsupported method raises [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) before a control message is sent. Current provider support is summarized on each provider page. Edge cases ---------- [](https://pydantic.dev/docs/ai/realtime/turns/#edge-cases) * Push-to-talk silence usually means `commit_audio()` or `create_response()` was omitted. * If playback triggers speech detection, add echo cancellation in the device or WebRTC layer and flush playback promptly on real barge-in. Was this page helpful? Thanks for your feedback! --- # Weather Agent | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/getting-started/weather-agent/#_top) Weather Agent ============= Example of Pydantic AI with multiple tools which the LLM needs to call in turn to answer a question. Demonstrates: * [tools](https://pydantic.dev/docs/ai/tools-toolsets/tools/) * [agent dependencies](https://pydantic.dev/docs/ai/core-concepts/dependencies/) * [streaming text responses](https://pydantic.dev/docs/ai/core-concepts/output/#streaming-text) * Building a [Gradio](https://www.gradio.app/) UI for the agent In this case the idea is a “weather” agent — the user can ask for the weather in multiple locations, the agent will use the `get_lat_lng` tool to get the latitude and longitude of the locations, then use the `get_weather` tool to get the weather for those locations. Running the Example ------------------- [](https://pydantic.dev/docs/ai/examples/getting-started/weather-agent/#running-the-example) To run this example properly, you might want to add two extra API keys **(Note if either key is missing, the code will fall back to dummy data, so they’re not required)**: * A weather API key from [tomorrow.io](https://www.tomorrow.io/weather-api/) set via `WEATHER_API_KEY` * A geocoding API key from [geocode.maps.co](https://geocode.maps.co/) set via `GEO_API_KEY` With [dependencies installed and environment variables set](https://pydantic.dev/docs/ai/examples/setup/#usage) , run: * [pip](https://pydantic.dev/docs/ai/examples/getting-started/weather-agent/#tab-panel-46) * [uv](https://pydantic.dev/docs/ai/examples/getting-started/weather-agent/#tab-panel-47) Terminal python -m pydantic_ai_examples.weather_agent Terminal uv run -m pydantic_ai_examples.weather_agent Example Code ------------ [](https://pydantic.dev/docs/ai/examples/getting-started/weather-agent/#example-code) weather\_agent.py from __future__ import annotations as _annotations import asyncio from dataclasses import dataclass from typing import Any import logfire from httpx import AsyncClient from pydantic import BaseModel from pydantic_ai import Agent, RunContext # 'if-token-present' means nothing will be sent (and the example will work) if you don't have logfire configured logfire.configure(send_to_logfire='if-token-present') logfire.instrument_pydantic_ai() @dataclass class Deps: client: AsyncClient weather_agent = Agent( 'openai:gpt-5-mini', # 'Be concise, reply with one sentence.' is enough for some models (like openai) to use # the below tools appropriately, but others like anthropic and gemini require a bit more direction. instructions='Be concise, reply with one sentence.', deps_type=Deps, retries=2, ) class LatLng(BaseModel): lat: float lng: float @weather_agent.tool async def get_lat_lng(ctx: RunContext[Deps], location_description: str) -> LatLng: """Get the latitude and longitude of a location. Args: ctx: The context. location_description: A description of a location. """ # NOTE: the response here will be random, and is not related to the location description. r = await ctx.deps.client.get( 'https://demo-endpoints.pydantic.workers.dev/latlng', params={'location': location_description}, ) r.raise_for_status() return LatLng.model_validate_json(r.content) @weather_agent.tool async def get_weather(ctx: RunContext[Deps], lat: float, lng: float) -> dict[str, Any]: """Get the weather at a location. Args: ctx: The context. lat: Latitude of the location. lng: Longitude of the location. """ # NOTE: the responses here will be random, and are not related to the lat and lng. temp_response, descr_response = await asyncio.gather( ctx.deps.client.get( 'https://demo-endpoints.pydantic.workers.dev/number', params={'min': 10, 'max': 30}, ), ctx.deps.client.get( 'https://demo-endpoints.pydantic.workers.dev/weather', params={'lat': lat, 'lng': lng}, ), ) temp_response.raise_for_status() descr_response.raise_for_status() return { 'temperature': f'{temp_response.text} °C', 'description': descr_response.text, } async def main(): async with AsyncClient() as client: logfire.instrument_httpx(client, capture_all=True) deps = Deps(client=client) result = await weather_agent.run( 'What is the weather like in London and in Wiltshire?', deps=deps ) print('Response:', result.output) if __name__ == '__main__': asyncio.run(main()) Running the UI -------------- [](https://pydantic.dev/docs/ai/examples/getting-started/weather-agent/#running-the-ui) You can build multi-turn chat applications for your agent with [Gradio](https://www.gradio.app/) , a framework for building AI web applications entirely in python. Gradio comes with built-in chat components and agent support so the entire UI will be implemented in a single python file! Here’s what the UI looks like for the weather agent: Terminal pip install gradio>=6.7.0 python/uv-run -m pydantic_ai_examples.weather_agent_gradio UI Code ------- [](https://pydantic.dev/docs/ai/examples/getting-started/weather-agent/#ui-code) weather\_agent\_gradio.py from __future__ import annotations as _annotations import json from httpx import AsyncClient from pydantic import BaseModel from pydantic_ai import ToolCallPart, ToolReturnPart from pydantic_ai_examples.weather_agent import Deps, weather_agent try: import gradio as gr except ImportError as e: raise ImportError( 'Please install gradio with `pip install gradio`. You must use python>=3.10.' ) from e TOOL_TO_DISPLAY_NAME = {'get_lat_lng': 'Geocoding API', 'get_weather': 'Weather API'} client = AsyncClient() deps = Deps(client=client) async def stream_from_agent(prompt: str, chatbot: list[dict], past_messages: list): chatbot.append({'role': 'user', 'content': prompt}) yield gr.Textbox(interactive=False, value=''), chatbot, gr.skip() async with weather_agent.run_stream( prompt, deps=deps, message_history=past_messages ) as result: for message in result.new_messages(): for call in message.parts: if isinstance(call, ToolCallPart): call_args = call.args_as_json_str() metadata = { 'title': f'🛠️ Using {TOOL_TO_DISPLAY_NAME[call.tool_name]}', } if call.tool_call_id is not None: metadata['id'] = call.tool_call_id gr_message = { 'role': 'assistant', 'content': 'Parameters: ' + call_args, 'metadata': metadata, } chatbot.append(gr_message) if isinstance(call, ToolReturnPart): for gr_message in chatbot: if (gr_message.get('metadata') or {}).get( 'id', '' ) == call.tool_call_id: if isinstance(call.content, BaseModel): json_content = call.content.model_dump_json() else: json_content = json.dumps(call.content) gr_message['content'] += f'\nOutput: {json_content}' yield gr.skip(), chatbot, gr.skip() chatbot.append({'role': 'assistant', 'content': ''}) async for message in result.stream_text(): chatbot[-1]['content'] = message yield gr.skip(), chatbot, gr.skip() past_messages = result.all_messages() yield gr.Textbox(interactive=True), gr.skip(), past_messages async def handle_retry(chatbot, past_messages: list, retry_data: gr.RetryData): new_history = chatbot[: retry_data.index] previous_prompt = chatbot[retry_data.index]['content'] past_messages = past_messages[: retry_data.index] async for update in stream_from_agent(previous_prompt, new_history, past_messages): yield update def undo(chatbot, past_messages: list, undo_data: gr.UndoData): new_history = chatbot[: undo_data.index] past_messages = past_messages[: undo_data.index] return chatbot[undo_data.index]['content'], new_history, past_messages def select_data(message: gr.SelectData) -> str: return message.value['text'] with gr.Blocks() as demo: gr.HTML( """

Weather Assistant

This assistant answer your weather questions.

""" ) past_messages = gr.State([]) chatbot = gr.Chatbot( label='Packing Assistant', avatar_images=(None, 'https://ai.pydantic.dev/img/logo-white.svg'), examples=[\ {'text': 'What is the weather like in Miami?'},\ {'text': 'What is the weather like in London?'},\ ], ) with gr.Row(): prompt = gr.Textbox( lines=1, show_label=False, placeholder='What is the weather like in New York City?', ) generation = prompt.submit( stream_from_agent, inputs=[prompt, chatbot, past_messages], outputs=[prompt, chatbot, past_messages], ) chatbot.example_select(select_data, None, [prompt]) chatbot.retry( handle_retry, [chatbot, past_messages], [prompt, chatbot, past_messages] ) chatbot.undo(undo, [chatbot, past_messages], [prompt, chatbot, past_messages]) if __name__ == '__main__': demo.launch() Was this page helpful? Thanks for your feedback! --- # Retries | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/core-concepts/retries/#_top) Retries ======= “Retry” means five different things in an agent run, at five different layers, and they don’t share budgets. Mixing them up is the usual cause of a run that retries far more (or far less) than expected. This page is the map; each layer links to the page that configures it in detail. The layers ---------- [](https://pydantic.dev/docs/ai/core-concepts/retries/#the-layers) | Layer | What it re-attempts | Configured with | What it adds to message history | | --- | --- | --- | --- | | [Transport](https://pydantic.dev/docs/ai/core-concepts/retries/#transport-retries) | The same HTTP request to the provider | [`AsyncTenacityTransport`](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.AsyncTenacityTransport)
on your HTTP client | Nothing — the agent never sees the attempts | | [Model fallback](https://pydantic.dev/docs/ai/core-concepts/retries/#model-fallback-is-not-a-retry) | The same request against a _different_ model | [`FallbackModel`](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackModel) | Only the winning response | | [Tool](https://pydantic.dev/docs/ai/core-concepts/retries/#tool-retries) | One tool call, by asking the model to correct it | `retries={'tools': N}` and per-tool limits | A [`RetryPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.RetryPromptPart)
in place of the tool’s result | | [Output](https://pydantic.dev/docs/ai/core-concepts/retries/#output-retries) | The model’s final answer, by asking it to correct it | `retries={'output': N}` and [`ToolOutput(max_retries=N)`](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.ToolOutput.max_retries) | A `RetryPromptPart` — see [below](https://pydantic.dev/docs/ai/core-concepts/retries/#output-retries)
for where it lands | | [Model-request hooks](https://pydantic.dev/docs/ai/core-concepts/hooks/) | The model request, from `after_model_request`, `wrap_model_request`, or `on_model_request_error` raising `ModelRetry` | The hook itself; it draws on the **output** budget | A new request carrying a `RetryPromptPart` | Only the last three are “agent retries” — they cost a model round trip each, because a retry _is_ another request. The first two are invisible to the model. Transport retries ----------------- [](https://pydantic.dev/docs/ai/core-concepts/retries/#transport-retries) Transport retries live below the model client: a failed HTTP request is re-sent without the agent ever knowing. Nothing retries at this layer unless you install a retrying transport on the HTTP client you pass to the provider, and you decide which errors qualify. This is the right layer for rate limits, connection resets, and 5xx responses. See [HTTP Request Retries](https://pydantic.dev/docs/ai/models/http-request-retries/) for the transports, the `Retry-After`\-aware wait strategy, and per-provider notes — including AWS Bedrock, which retries through boto3 rather than httpx. When you build your own backoff outside a transport, [`ModelHTTPError.retry_after`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelHTTPError.retry_after) gives you the provider’s `Retry-After` header already parsed into seconds. Model fallback is not a retry ----------------------------- [](https://pydantic.dev/docs/ai/core-concepts/retries/#model-fallback-is-not-a-retry) [`FallbackModel`](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackModel) moves to the _next_ model when the current one fails; it never re-attempts the same one. Pair it with transport retries rather than treating it as a substitute: retry the same provider for transient failures, fall back to a different provider when it’s genuinely down. See [Fallback Model](https://pydantic.dev/docs/ai/models/overview/#fallback-model) . Tool retries ------------ [](https://pydantic.dev/docs/ai/core-concepts/retries/#tool-retries) A tool retry is a message to the model: the call didn’t work, here is why, try again. It is triggered by a Pydantic `ValidationError` on the tool’s arguments, by the tool (or its `args_validator`, or a tool hook) raising [`ModelRetry`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelRetry) , by a [tool timeout](https://pydantic.dev/docs/ai/core-concepts/timeouts/#bounding-how-long-a-step-takes) , and by the model calling a tool that doesn’t exist. [Tool Execution, Retries, and Failures](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-retries) documents the configuration: the default budget of `1`, the per-tool / per-toolset / per-run / agent-wide precedence ladder, and the choice between `ModelRetry` and [`ToolFailed`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ToolFailed) . Three properties of the _counter_ matter when you’re reasoning about a run: * **The counter is keyed by tool name, and it resets on success.** Each tool has its own count; there is no run-wide tool-retry budget. When a tool succeeds, its count is cleared — so a tool that alternates failure and success can fail many times in one run without ever exhausting a budget of `1`. * **`max_retries=N` allows N retries, so N+1 attempts.** `max_retries=0` raises on the first failure without ever sending a retry prompt. * **A tool name the model invented gets its own budget.** An unknown tool name produces a retry prompt listing the available tools, and consumes a budget keyed under the invented name, bounded by the agent-wide `tools` budget. So a model that hallucinates a _different_ name each time keeps getting a fresh budget. Exhausting a tool’s budget raises [`UnexpectedModelBehavior`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UnexpectedModelBehavior) . ### What a retry looks like in message history [](https://pydantic.dev/docs/ai/core-concepts/retries/#what-a-retry-looks-like-in-message-history) A retried tool call has no [`ToolReturnPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolReturnPart) — the [`RetryPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.RetryPromptPart) takes its place, carrying the same `tool_call_id`. There is never both: retry\_prompt\_history.py from pydantic_ai import ( Agent, ModelMessage, ModelResponse, ModelRetry, TextPart, ToolCallPart, ) from pydantic_ai.models.function import AgentInfo, FunctionModel def lookup_then_answer( messages: list[ModelMessage], info: AgentInfo ) -> ModelResponse: if len(messages) == 1: return ModelResponse(parts=[ToolCallPart('lookup_user', {'name': 'John'})]) elif len(messages) == 3: return ModelResponse( parts=[ToolCallPart('lookup_user', {'name': 'John Doe'})] ) return ModelResponse(parts=[TextPart('John Doe is user 123.')]) agent = Agent(FunctionModel(lookup_then_answer)) @agent.tool_plain def lookup_user(name: str) -> int: if ' ' not in name: raise ModelRetry('Provide the full name.') return 123 result = agent.run_sync('Who is John?') print([type(p).__name__ for m in result.all_messages() for p in m.parts]) """ [\ 'UserPromptPart',\ 'ToolCallPart',\ 'RetryPromptPart',\ 'ToolCallPart',\ 'ToolReturnPart',\ 'TextPart',\ ] """ _(This example is complete, it can be run “as is”)_ A [`RetryPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.RetryPromptPart) carries the failure as either a string (from `ModelRetry`) or a list of Pydantic error details (from a `ValidationError`), and renders for the model with `'Fix the errors and try again.'` appended. Its `tool_name` is set when the retry belongs to a specific tool call, and `None` when it belongs to the run’s output. Because the retry prompts stay in the history, [reusing that history](https://pydantic.dev/docs/ai/core-concepts/message-history/) in a later run replays the failures to the model. If you don’t want the model to see its earlier mistakes, filter them out with a [`ProcessHistory`](https://pydantic.dev/docs/ai/capabilities/process-history/) capability. [`ToolFailed`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ToolFailed) is the deliberate opposite: it records a `ToolReturnPart` with `outcome='failed'` and does **not** consume the retry budget, so repeated failures are bounded by [`UsageLimits`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits) rather than by a retry count. See [Reporting a Failed Tool Result](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-failed) . Output retries -------------- [](https://pydantic.dev/docs/ai/core-concepts/retries/#output-retries) The output budget is separate from the tool budget, and how it’s enforced depends on how the model returns its final answer. [How output retries are enforced](https://pydantic.dev/docs/ai/core-concepts/agent/#how-output-retries-are-enforced) covers both paths; the difference that matters for message history is: * **Text path** (`output_type=str`, [`TextOutput`](https://pydantic.dev/docs/ai/core-concepts/output/#text-output) , [`NativeOutput`](https://pydantic.dev/docs/ai/core-concepts/output/#native-output) , [`PromptedOutput`](https://pydantic.dev/docs/ai/core-concepts/output/#prompted-output) , and responses with no usable output): one budget shared across the whole run. The retry becomes a new [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) whose only part is a `RetryPromptPart` with `tool_name=None`. * **Tool path** ([`ToolOutput`](https://pydantic.dev/docs/ai/core-concepts/output/#tool-output) ): the output budget acts as the default limit _per output tool_, overridable with [`ToolOutput(max_retries=N)`](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.ToolOutput.max_retries) . The retry prompt is bound to the output tool’s `tool_call_id`, exactly like a function tool’s. Both are triggered by validation failures, by an [output function](https://pydantic.dev/docs/ai/core-concepts/output/#output-functions) or [output validator](https://pydantic.dev/docs/ai/core-concepts/output/#output-validator-functions) raising `ModelRetry`, and by a model response with nothing actionable in it. Both raise [`UnexpectedModelBehavior`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UnexpectedModelBehavior) when the budget runs out. The last of those triggers has an exception: if the output type allows `None` — `output_type=str | None`, for instance — an empty or thinking-only response is a valid final result of `None` rather than a retry. Models that finish their work in a tool call and then emit only thinking would otherwise be pushed into producing filler text. Output validators still run on that `None`, so they can force a retry themselves by raising `ModelRetry`. Both budgets are configured through one argument: retry\_budgets.py from pydantic_ai import Agent agent = Agent('openai:gpt-5.2', retries=3) # (1) strict_output = Agent('openai:gpt-5.2', retries={'tools': 5, 'output': 1}) # (2) A bare `int` sets both the tool and output budgets. An [`AgentRetries`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentRetries) dict sets only the keys it names; unnamed keys keep the default of `1`. The same argument is accepted per run — `agent.run(..., retries=...)` and friends — and for a block of runs via [`agent.override()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.override) . [Which retry limit wins](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#which-retry-limit-wins) has the full precedence table. What is never retried --------------------- [](https://pydantic.dev/docs/ai/core-concepts/retries/#what-is-never-retried) * **`prepare` callbacks.** An exception raised by a per-tool `prepare=`, by [`PrepareTools`](https://pydantic.dev/docs/ai/capabilities/prepare-tools/) , or by a [dynamic toolset](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) propagates out of the run unchanged — including `ModelRetry`, which is _not_ turned into a retry prompt there. To hide a tool for a turn, return `None` from the callback rather than raising. * **The `before_model_request` hook.** It runs while the request is still being assembled, before the model is called, so a `ModelRetry` raised there propagates out of the run instead of becoming a retry prompt — there is no response to retry yet. Raise it from one of the [other model-request hooks](https://pydantic.dev/docs/ai/core-concepts/hooks/#model-request-hooks) instead: `hooks.on.after_model_request` to reject a response the model _did_ produce (the rejected response stays in the message history, so the model can see what it said), `hooks.on.model_request` (`wrap_model_request`), or `hooks.on.model_request_error` (`on_model_request_error`). * **Exceptions other than `ModelRetry` and `ToolFailed`.** Anything else a tool raises propagates out of the run rather than becoming a retry — _unless_ a [capability](https://pydantic.dev/docs/ai/capabilities/overview/) implements `on_tool_execute_error`, which sees the exception first and can return a replacement tool result or raise `ModelRetry` to keep the run going. [`ApprovalRequired`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ApprovalRequired) and [`CallDeferred`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.CallDeferred) are the exceptions that are neither: they’re control flow, not errors, and end the run with a [`DeferredToolRequests`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.DeferredToolRequests) output instead of propagating — except in a [realtime session](https://pydantic.dev/docs/ai/realtime/overview/) , which can’t pause and instead answers the model with an explanation that the tool can’t complete during the session. [Ending a run from inside a tool](https://pydantic.dev/docs/ai/core-concepts/timeouts/#ending-a-run-from-inside-a-tool) has the full table. * **Whole agent runs.** Nothing re-runs an agent for you. [Pydantic Evals](https://pydantic.dev/docs/ai/evals/evals/) has its own `retry_task` and `retry_evaluators` options for retrying a whole task or evaluator during an evaluation — see [Retry Strategies](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/) . Those sit outside the agent, so a retried task starts with fresh tool and output budgets. Was this page helpful? Thanks for your feedback! --- # Multi-Agent Patterns | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#_top) Multi-Agent Patterns ==================== There are roughly five levels of complexity when building applications with Pydantic AI: 1. Single agent workflows — what most of the `pydantic_ai` documentation covers 2. [Agent delegation](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#agent-delegation) — agents using another agent via tools 3. [Programmatic agent hand-off](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#programmatic-agent-hand-off) — one agent runs, then application code calls another agent 4. [Graph based control flow](https://pydantic.dev/docs/ai/graph/graph/) — for the most complex cases, a graph-based state machine can be used to control the execution of multiple agents 5. [Deep Agents](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#deep-agents) — autonomous agents with planning, file operations, task delegation, and sandboxed code execution Of course, you can combine multiple strategies in a single application. Agent delegation ---------------- [](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#agent-delegation) “Agent delegation” refers to the scenario where an agent delegates work to another agent, then takes back control when the delegate agent (the agent called from within a tool) finishes. If you want to hand off control to another agent completely, without coming back to the first agent, you can use an [output function](https://pydantic.dev/docs/ai/core-concepts/output/#output-functions) . Since agents are stateless and designed to be global, you do not need to include the agent itself in agent [dependencies](https://pydantic.dev/docs/ai/core-concepts/dependencies/) . You’ll generally want to pass [`ctx.usage`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.usage) to the [`usage`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.run) keyword argument of the delegate agent run so usage within that run counts towards the total usage of the parent agent run. [Cancellation](https://pydantic.dev/docs/ai/core-concepts/agent/#cancellation-and-sub-agents) is run-scoped: a delegate agent cancelling itself surfaces to the parent as a failed tool return rather than cancelling the parent, and a shared [`CancellationToken`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.CancellationToken) cancels a whole tree of runs at once. agent\_delegation\_simple.py from pydantic_ai import Agent, RunContext, UsageLimits joke_selection_agent = Agent( # (1) 'openai:gpt-5.2', name='joke_selection_agent', # (2) instructions=( 'Use the `joke_factory` to generate some jokes, then choose the best. ' 'You must return just a single joke.' ), ) joke_generation_agent = Agent( # (3) 'google:gemini-3-flash-preview', name='joke_generation_agent', output_type=list[str] ) @joke_selection_agent.tool async def joke_factory(ctx: RunContext, count: int) -> list[str]: r = await joke_generation_agent.run( # (4) f'Please generate {count} jokes.', usage=ctx.usage, # (5) ) return r.output # (6) result = joke_selection_agent.run_sync( 'Tell me a joke.', usage_limits=UsageLimits(request_limit=5, total_tokens_limit=500), ) print(result.output) #> Did you hear about the toothpaste scandal? They called it Colgate. print(result.usage) """ RunUsage( cost=Decimal('0.00051200'), input_tokens=165, output_tokens=24, requests=3, tool_calls=1, ) """ The "parent" or controlling agent. Passing `name` is optional but recommended when you run more than one agent: it labels each agent's run span, so naming both lets you tell the parent and delegate apart in [Logfire](https://pydantic.dev/docs/ai/integrations/logfire/) . When omitted, the name is inferred from the variable the agent is assigned to and falls back to `'agent'` when it can't be (e.g. agents kept in a list or dict). The "delegate" agent, which is called from within a tool of the parent agent. Call the delegate agent from within a tool of the parent agent. Pass the usage from the parent agent to the delegate agent so the final [`result.usage`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.usage) includes the usage from both agents. Since the function returns `list[str]`, and the `output_type` of `joke_generation_agent` is also `list[str]`, we can simply return `r.output` from the tool. _(This example is complete, it can be run “as is”)_ The control flow for this example is pretty simple and can be summarised as follows: graph TD START --> joke_selection_agent joke_selection_agent --> joke_factory["joke_factory (tool)"] joke_factory --> joke_generation_agent joke_generation_agent --> joke_factory joke_factory --> joke_selection_agent joke_selection_agent --> END ### Agent delegation and dependencies [](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#agent-delegation-and-dependencies) Generally the delegate agent needs to either have the same [dependencies](https://pydantic.dev/docs/ai/core-concepts/dependencies/) as the calling agent, or dependencies which are a subset of the calling agent’s dependencies. agent\_delegation\_deps.py from dataclasses import dataclass import httpx from pydantic_ai import Agent, RunContext @dataclass class ClientAndKey: # (1) http_client: httpx.AsyncClient api_key: str joke_selection_agent = Agent( 'openai:gpt-5.2', name='joke_selection_agent', deps_type=ClientAndKey, # (2) instructions=( 'Use the `joke_factory` tool to generate some jokes on the given subject, ' 'then choose the best. You must return just a single joke.' ), ) joke_generation_agent = Agent( 'google:gemini-3-flash-preview', name='joke_generation_agent', deps_type=ClientAndKey, # (4) output_type=list[str], instructions=( 'Use the "get_jokes" tool to get some jokes on the given subject, ' 'then extract each joke into a list.' ), ) @joke_selection_agent.tool async def joke_factory(ctx: RunContext[ClientAndKey], count: int) -> list[str]: r = await joke_generation_agent.run( f'Please generate {count} jokes.', deps=ctx.deps, # (3) usage=ctx.usage, ) return r.output @joke_generation_agent.tool # (5) async def get_jokes(ctx: RunContext[ClientAndKey], count: int) -> str: response = await ctx.deps.http_client.get( 'https://example.com', params={'count': count}, headers={'Authorization': f'Bearer {ctx.deps.api_key}'}, ) response.raise_for_status() return response.text async def main(): async with httpx.AsyncClient() as client: deps = ClientAndKey(client, 'foobar') result = await joke_selection_agent.run('Tell me a joke.', deps=deps) print(result.output) #> Did you hear about the toothpaste scandal? They called it Colgate. print(result.usage) # (6) """ RunUsage( cost=Decimal('0.00056350'), input_tokens=220, output_tokens=32, requests=4, tool_calls=2, ) """ Define a dataclass to hold the client and API key dependencies. Set the `deps_type` of the calling agent — `joke_selection_agent` here. Pass the dependencies to the delegate agent's run method within the tool call. Also set the `deps_type` of the delegate agent — `joke_generation_agent` here. Define a tool on the delegate agent that uses the dependencies to make an HTTP request. Usage now includes 4 requests — 2 from the calling agent and 2 from the delegate agent. _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main())` to run `main`)_ This example shows how even a fairly simple agent delegation can lead to a complex control flow: graph TD START --> joke_selection_agent joke_selection_agent --> joke_factory["joke_factory (tool)"] joke_factory --> joke_generation_agent joke_generation_agent --> get_jokes["get_jokes (tool)"] get_jokes --> http_request["HTTP request"] http_request --> get_jokes get_jokes --> joke_generation_agent joke_generation_agent --> joke_factory joke_factory --> joke_selection_agent joke_selection_agent --> END Programmatic agent hand-off --------------------------- [](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#programmatic-agent-hand-off) “Programmatic agent hand-off” refers to the scenario where multiple agents are called in succession, with application code and/or a human in the loop responsible for deciding which agent to call next. Here agents don’t need to use the same deps. Here we show two agents used in succession, the first to find a flight and the second to extract the user’s seat preference. programmatic\_handoff.py from typing import Literal from pydantic import BaseModel, Field from rich.prompt import Prompt from pydantic_ai import Agent, ModelMessage, RunContext, RunUsage, UsageLimits class FlightDetails(BaseModel): flight_number: str class Failed(BaseModel): """Unable to find a satisfactory choice.""" flight_search_agent = Agent[object, FlightDetails | Failed]( # (1) 'openai:gpt-5.2', name='flight_search_agent', output_type=FlightDetails | Failed, # type: ignore instructions=( 'Use the "flight_search" tool to find a flight ' 'from the given origin to the given destination.' ), ) @flight_search_agent.tool # (2) async def flight_search( ctx: RunContext, origin: str, destination: str ) -> FlightDetails | None: # in reality, this would call a flight search API or # use a browser to scrape a flight search website return FlightDetails(flight_number='AK456') usage_limits = UsageLimits(request_limit=15) # (3) async def find_flight(usage: RunUsage) -> FlightDetails | None: # (4) message_history: list[ModelMessage] | None = None for _ in range(3): prompt = Prompt.ask( 'Where would you like to fly from and to?', ) result = await flight_search_agent.run( prompt, message_history=message_history, usage=usage, usage_limits=usage_limits, ) if isinstance(result.output, FlightDetails): return result.output else: message_history = result.all_messages( output_tool_return_content='Please try again.' ) class SeatPreference(BaseModel): row: int = Field(ge=1, le=30) seat: Literal['A', 'B', 'C', 'D', 'E', 'F'] # This agent is responsible for extracting the user's seat selection seat_preference_agent = Agent[object, SeatPreference | Failed]( # (5) 'openai:gpt-5.2', name='seat_preference_agent', output_type=SeatPreference | Failed, # type: ignore instructions=( "Extract the user's seat preference. " 'Seats A and F are window seats. ' 'Row 1 is the front row and has extra leg room. ' 'Rows 14, and 20 also have extra leg room. ' ), ) async def find_seat(usage: RunUsage) -> SeatPreference: # (6) message_history: list[ModelMessage] | None = None while True: answer = Prompt.ask('What seat would you like?') result = await seat_preference_agent.run( answer, message_history=message_history, usage=usage, usage_limits=usage_limits, ) if isinstance(result.output, SeatPreference): return result.output else: print('Could not understand seat preference. Please try again.') message_history = result.all_messages() async def main(): # (7) usage: RunUsage = RunUsage() opt_flight_details = await find_flight(usage) if opt_flight_details is not None: print(f'Flight found: {opt_flight_details.flight_number}') #> Flight found: AK456 seat_preference = await find_seat(usage) print(f'Seat preference: {seat_preference}') #> Seat preference: row=1 seat='A' Define the first agent, which finds a flight. We use an explicit type annotation until [PEP-747](https://peps.python.org/pep-0747/) lands, see [structured output](https://pydantic.dev/docs/ai/core-concepts/output/#structured-output) . We use a union as the output type so the model can communicate if it's unable to find a satisfactory choice; internally, each member of the union will be registered as a separate tool. Define a tool on the agent to find a flight. In this simple case we could dispense with the tool and just define the agent to return structured data, then search for a flight, but in more complex scenarios the tool would be necessary. Define usage limits for the entire app. Define a function to find a flight, which asks the user for their preferences and then calls the agent to find a flight. As with `flight_search_agent` above, we use an explicit type annotation to define the agent. Define a function to find the user's seat preference, which asks the user for their seat preference and then calls the agent to extract the seat preference. Now that we've put our logic for running each agent into separate functions, our main app becomes very simple. _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main())` to run `main`)_ The control flow for this example can be summarised as follows: graph TB START --> ask_user_flight["ask user for flight"] subgraph find_flight flight_search_agent --> ask_user_flight ask_user_flight --> flight_search_agent end flight_search_agent --> ask_user_seat["ask user for seat"] flight_search_agent --> END subgraph find_seat seat_preference_agent --> ask_user_seat ask_user_seat --> seat_preference_agent end seat_preference_agent --> END Pydantic Graphs --------------- [](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#pydantic-graphs) See the [graph](https://pydantic.dev/docs/ai/graph/graph/) documentation on when and how to use graphs. Deep Agents ----------- [](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#deep-agents) Deep agents are autonomous agents that combine multiple architectural patterns and capabilities to handle complex, multi-step tasks reliably. These patterns can be implemented using Pydantic AI’s built-in features and (third-party) toolsets: * **Planning and progress tracking** — agents break down complex tasks into steps and track their progress, giving users visibility into what the agent is working on. See [Task Management toolsets](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/#task-management) . * **File system operations** — reading, writing, and editing files with proper abstraction layers that work across in-memory storage, real file systems, and sandboxed containers. See [File Operations toolsets](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/#file-operations) . * **Task delegation** — spawning specialized sub-agents for specific tasks, with isolated context to prevent recursive delegation issues. See [Agent Delegation](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#agent-delegation) above. * **Sandboxed code execution** — running AI-generated code in isolated environments (typically Docker containers) to prevent accidents. See [Code Execution toolsets](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/#code-execution) . * **Context management** — automatic conversation summarization to handle long sessions that would otherwise exceed token limits. See [Processing Message History](https://pydantic.dev/docs/ai/core-concepts/message-history/#processing-message-history) . * **Human-in-the-loop** — approval workflows for dangerous operations like code execution or file deletion. See [Requiring Tool Approval](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/#requiring-tool-approval) . * **Durable execution** — preserving agent state across transient API failures and application errors or restarts. See [Durable Execution](https://pydantic.dev/docs/ai/capabilities/durable_execution/overview/) . In addition, the community maintains packages that bring these concepts together in a more opinionated way: * [`pydantic-deep`](https://github.com/vstorm-co/pydantic-deepagents) by [Vstorm](https://vstorm.co/) Observing Multi-Agent Systems ----------------------------- [](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#observing-multi-agent-systems) Multi-agent systems can be challenging to debug due to their complexity; when multiple agents interact, understanding the flow of execution becomes essential. ### Tracing Agent Delegation [](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#tracing-agent-delegation) With [Logfire](https://pydantic.dev/docs/ai/integrations/logfire/) , you can trace the entire flow across multiple agents: import logfire logfire.configure() logfire.instrument_pydantic_ai() # Your multi-agent code here... Logfire shows you: * **Which agent handled which part** of the request * **Delegation decisions**—when and why one agent called another * **End-to-end latency** broken down by agent * **Token usage and costs** per agent * **What triggered the agent run**—the HTTP request, scheduled job, or user action that started it all * **What happened inside tool calls**—database queries, HTTP requests, file operations, and any other instrumented code that tools execute This is essential for understanding and optimizing complex agent workflows. When something goes wrong in a multi-agent system, you’ll see exactly which agent failed and what it was trying to do, and whether the problem was in the agent’s reasoning or in the backend systems it called. ### Full-Stack Visibility [](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#full-stack-visibility) If your Pydantic AI application includes a TypeScript frontend, API gateway, or services in other languages, Logfire can trace them too—Logfire provides SDKs for Python, JavaScript/TypeScript, and Rust, plus compatibility with any OpenTelemetry-instrumented application. See traces from your entire stack in a unified view. For details on sending data from other languages using standard OpenTelemetry, see the [alternative clients guide](https://logfire.pydantic.dev/docs/how-to-guides/alternative-clients/) . Pydantic AI’s instrumentation is built on [OpenTelemetry](https://opentelemetry.io/) , so you can also use any OTel-compatible backend. See the [Logfire integration guide](https://pydantic.dev/docs/ai/integrations/logfire/) for details. Examples -------- [](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#examples) The following examples demonstrate how to use multi-agent patterns in Pydantic AI: * [Flight booking](https://pydantic.dev/docs/ai/examples/complex-workflows/flight-booking/) Was this page helpful? Thanks for your feedback! --- # Events | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/realtime/events/#_top) Events ====== Iterating a [`RealtimeSession`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession) yields the session’s event stream: content parts, tool activity, turn boundaries, reconnects, and recoverable errors. The high-level [`stream_audio()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.stream_audio) and [`stream_transcripts()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.stream_transcripts) views described in [Audio, images, and transcripts](https://pydantic.dev/docs/ai/realtime/audio/) are derived from this same stream, so most applications iterate the session for control flow and leave media to the views. Event reference --------------- [](https://pydantic.dev/docs/ai/realtime/events/#event-reference) | Event | Meaning | | --- | --- | | [`PartStartEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.PartStartEvent) | A speech, text, or tool part started. | | [`PartDeltaEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.PartDeltaEvent) | Incremental speech audio/transcript or text content. | | [`PartEndEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.PartEndEvent) | A finalized part; retained speech audio appears here, not at part start. | | [`FunctionToolCallEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.FunctionToolCallEvent) | A local function tool began executing. | | [`FunctionToolResultEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.FunctionToolResultEvent) | A local function tool completed or returned a retry prompt. | | [`DeferredToolRequestsEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.DeferredToolRequestsEvent) | An inline capability handler resolved deferred requests. | | [`DeferredToolResultsEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.DeferredToolResultsEvent) | Inline deferred results are ready for normal tool processing. | | [`RealtimeInputSpeechStartEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeInputSpeechStartEvent) | The provider detected that the user started speaking, when the profile declares [`emits_input_speech_events`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelProfile.emits_input_speech_events)
. | | [`RealtimeInputSpeechEndEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeInputSpeechEndEvent) | The provider detected the end of user speech, when the profile declares [`emits_input_speech_events`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelProfile.emits_input_speech_events)
. | | [`RealtimeResponseInterruptedEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeResponseInterruptedEvent) | The provider reported an interrupted model response. | | [`RealtimeInputTranscriptionErrorEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeInputTranscriptionErrorEvent) | One user turn could not be transcribed; the session remains usable. | | [`RealtimeOutputSpeechStartEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeOutputSpeechStartEvent)
/ [`RealtimeOutputSpeechEndEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeOutputSpeechEndEvent) | The model became, or stopped being, audible. These are emitted on a [WebRTC sideband](https://pydantic.dev/docs/ai/realtime/deployment/#browser-webrtc-server-sideband)
, where the provider owns audio playback. | | [`RealtimeTurnCompleteEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeTurnCompleteEvent) | The model finished replying and no tool remains active. | | [`RealtimeSessionReconnectEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSessionReconnectEvent) | The connection was automatically re-established. | | [`RealtimeSessionErrorEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSessionErrorEvent) | A recoverable provider error occurred; the session remains usable. | Shared and realtime-only events ------------------------------- [](https://pydantic.dev/docs/ai/realtime/events/#shared-and-realtime-only-events) The first seven rows are [`AgentStreamEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.AgentStreamEvent) members from [`pydantic_ai.messages`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages) — the same events a [standard streamed run](https://pydantic.dev/docs/ai/core-concepts/agent/#streaming-all-events) yields, so event-handling code written for a text agent (rendering parts, logging tool calls) works on a session unchanged. The `Realtime*` rows are [`RealtimeEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeEvent) members that only a session emits: speech detection, interruption, turn completion, reconnection, and recoverable errors have no equivalent in a request-response run. A capability’s [event stream hooks](https://pydantic.dev/docs/ai/core-concepts/hooks/#event-stream-hooks) see both kinds flow through the same stream; see [Capabilities and hooks](https://pydantic.dev/docs/ai/realtime/capabilities/#the-event-stream) . The turn boundary ----------------- [](https://pydantic.dev/docs/ai/realtime/events/#the-turn-boundary) Use [`RealtimeTurnCompleteEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeTurnCompleteEvent) as the exchange boundary. A model can speak, call a tool, and speak again, so receiving speech — or a tool result — does not imply that the turn is done. Reading raw audio events ------------------------ [](https://pydantic.dev/docs/ai/realtime/events/#reading-raw-audio-events) The audio stream _is_ these events under the hood: [`stream_audio()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.stream_audio) is a bounded view over the speech part deltas, and most applications should use it. As an advanced alternative, play [`SpeechPartDelta.audio_chunk`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SpeechPartDelta.audio_chunk) from raw [`PartDeltaEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.PartDeltaEvent) s. Model audio arrives in full whether or not history retention is enabled. When output audio is retained, the final [`SpeechPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SpeechPart) contains the whole turn again as a WAV snapshot for [history](https://pydantic.dev/docs/ai/realtime/history/#retaining-audio) ; do not play both or the turn will play twice. Was this page helpful? Thanks for your feedback! --- # V1 → V2 Migration Map | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/overview/migration/#_top) V1 → V2 Migration Map ===================== A lookup index for upgrading from Pydantic AI V1 to V2: find the V1 name you have in your code, read off the V2 name to replace it with. The [Upgrade Guide](https://pydantic.dev/docs/ai/project/changelog/) is the canonical source for _why_ each change was made, the behavior changes that come with it, and the recommended upgrade path. This page is the fast path for the one question the guide answers in prose: **what replaced what.** Agent configuration ------------------- [](https://pydantic.dev/docs/ai/overview/migration/#agent-configuration) Most V1 `Agent(...)` arguments that configured behavior moved onto [capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) , a single composable primitive that bundles an agent’s tools, [hooks](https://pydantic.dev/docs/ai/core-concepts/hooks/) , instructions, and model settings. | V1 | V2 | | --- | --- | | `Agent(builtin_tools=[...])` | `Agent(capabilities=[NativeTool(...)])` | | `Agent(event_stream_handler=...)` | `Agent(capabilities=[ProcessEventStream(...)])` (the `event_stream_handler=` argument on `run()`/`run_sync()`/`run_stream()`/`iter()` is unchanged) | | `Agent(history_processors=...)` | `Agent(capabilities=[ProcessHistory(...)])` | | `Agent(instrument=...)`, `Agent.from_spec(instrument=...)`, `Agent.from_file(instrument=...)`, `AgentSpec.instrument` | `Agent(capabilities=[Instrumentation(...)])` | | `Agent(mcp_servers=[...])` | `Agent(toolsets=[...])` | | `Agent(prepare_tools=...)` | `Agent(capabilities=[PrepareTools(...)])` | | `Agent.run_mcp_servers()` | `async with agent:` | | `Agent.sequential_tool_calls()` | `Agent.parallel_tool_call_execution_mode('sequential')` | | `Agent.to_a2a()` | `fasta2a.pydantic_ai.agent_to_a2a` (install `fasta2a[pydantic-ai]>=0.6.1`) | | `Agent.to_ag_ui()`, `AGUIApp`, `pydantic_ai.ag_ui` | `pydantic_ai.ui.ag_ui.AGUIAdapter` | | `Agent('gpt-5')` (no provider prefix) | `Agent('openai:gpt-5')` — the prefix-less fallback now raises `UserError` | | `Agent[None, ...]`, `RunContext[None]`, `Tool[None]` where deps aren’t actually `None` | `Agent[object, ...]`, `RunContext[object]`, `Tool[object]` — the generic defaults changed from `None` to `object` | Models and providers -------------------- [](https://pydantic.dev/docs/ai/overview/migration/#models-and-providers) | V1 | V2 | | --- | --- | | `pydantic_ai.models.gemini.GeminiModel` | `pydantic_ai.models.google.GoogleModel` | | `pydantic_ai.models.openai.OpenAIModel` | `pydantic_ai.models.openai.OpenAIChatModel` | | `pydantic_ai.models.openai.OpenAIModelSettings` | `pydantic_ai.models.openai.OpenAIChatModelSettings` | | `OpenAIChatModel(system_prompt_role=...)` | `OpenAIChatModel(profile=OpenAIModelProfile(openai_system_prompt_role=...))` — see the [note below](https://pydantic.dev/docs/ai/overview/migration/#not-a-straight-rename)
if the model already resolves a profile | | `OpenAICompaction(instructions=...)` | Removed | | `pydantic_ai.models.outlines.OutlinesModel`, `pydantic_ai.providers.outlines.OutlinesProvider` | Removed, no replacement | | `pydantic_ai.models.cached_async_http_client` | `pydantic_ai.models.create_async_http_client()` | | `pydantic_ai.providers.google.GoogleProvider(vertexai=, location=, project=, credentials=)` | `pydantic_ai.providers.google_cloud.GoogleCloudProvider(...)` | | `pydantic_ai.providers.google.GoogleGLAProvider` | `pydantic_ai.providers.google.GoogleProvider` | | `pydantic_ai.providers.google.GoogleVertexProvider` | `pydantic_ai.providers.google_cloud.GoogleCloudProvider` | | `pydantic_ai.providers.grok.GrokProvider`, `GrokModelName` | `pydantic_ai.providers.xai.XaiProvider` with `pydantic_ai.models.xai.XaiModel` / `XaiModelName` | | `GoogleModelSettings['google_vertex_service_tier']`, `['google_service_tier']` | `GoogleModelSettings['google_cloud_service_tier']` | | `StreamedResponse.usage()` (custom `Model` subclasses) | `StreamedResponse.usage` property | ### Model name prefixes [](https://pydantic.dev/docs/ai/overview/migration/#model-name-prefixes) | V1 prefix | V2 prefix | | --- | --- | | `openai:` (Chat Completions) | `openai:` now means the Responses API; use `openai-chat:` for Chat Completions, `openai-responses:` to be explicit | | `google-gla:` | `google:` | | `google-vertex:`, `vertexai:` | `google-cloud:` | | `gateway/gemini:`, `gateway/google-vertex:` | `gateway/google-cloud:` | | `grok:` | `xai:` | Model profiles -------------- [](https://pydantic.dev/docs/ai/overview/migration/#model-profiles) [`ModelProfile`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfile) and its subclasses are now `TypedDict`s rather than dataclasses. Constructing one (`OpenAIModelProfile(field=value)`) is unchanged; reading, mutating, or merging one is not. The full recipe table is in the Upgrade Guide under [`ModelProfile` is now a `TypedDict`](https://pydantic.dev/docs/ai/project/changelog/#modelprofile-is-now-a-typeddict) . | V1 | V2 | | --- | --- | | `profile.field` | `profile.get('field', )` — defaults are exported from [`pydantic_ai.profiles`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles) | | `profile.field = value` | `profile['field'] = value` | | `dataclasses.replace(profile, field=value)` | `{**profile, 'field': value}` | | `profile.update(other)` | [`merge_profile(profile, other)`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.merge_profile) | | `OpenAIModelProfile.from_profile(p)` | `p` | | `isinstance(profile, OpenAIModelProfile)` | Not supported on a `TypedDict` — check key presence instead | | `OpenAIModelProfile.openai_supports_sampling_settings` | `OpenAIModelProfile.openai_unsupported_model_settings` — **not a rename**, see [below](https://pydantic.dev/docs/ai/overview/migration/#not-a-straight-rename) | | `OpenAIModelProfile.openai_builtin_tools` | `OpenAIModelProfile.openai_native_tools` | ### Not a straight rename [](https://pydantic.dev/docs/ai/overview/migration/#not-a-straight-rename) Two of the OpenAI rows above need more than a find-and-replace: * **`openai_supports_sampling_settings` → `openai_unsupported_model_settings` changes shape, not just name.** The V1 field was a `bool` covering the sampling settings as a group; the V2 field is a sequence of the specific setting names to drop. `openai_supports_sampling_settings=False` becomes an explicit list of what the model doesn’t accept, e.g. `openai_unsupported_model_settings=('temperature', 'top_p')`. `True` was the default, so it simply goes away. * **`system_prompt_role` moves from a model argument into a profile.** If you were already passing `profile=` to the model, merge the setting into that profile rather than replacing it — a second `OpenAIModelProfile(...)` overrides the first wholesale. Profiles are `TypedDict`s in V2, so merging is `{**existing_profile, 'openai_system_prompt_role': 'user'}` or [`merge_profile()`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.merge_profile) . MCP --- [](https://pydantic.dev/docs/ai/overview/migration/#mcp) The per-transport server classes collapsed into a single [`MCPToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset) whose transport is inferred from the arguments you pass. Its defaults differ from the V1 classes’ — notably `max_retries`, `read_timeout`, `init_timeout`, and `elicitation_handler` — so re-read [MCP Client](https://pydantic.dev/docs/ai/mcp/client/) rather than assuming your V1 timeouts carried over. | V1 | V2 | | --- | --- | | `MCPServerStdio`, `MCPServerSSE`, `MCPServerStreamableHTTP`, `MCPServerHTTP` | [`pydantic_ai.mcp.MCPToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset) | | `FastMCPToolset` (and the `fastmcp` extra) | `MCPToolset` | | `load_mcp_servers` | [`pydantic_ai.mcp.load_mcp_toolsets`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.load_mcp_toolsets) | | `Agent.run_mcp_servers()` | `async with agent:` | | `MCP(url=...)` running remotely by default | `MCP(url=..., native=True)` to keep the V1 behavior; `MCP(url=...)` now runs the server locally | Tools and toolsets ------------------ [](https://pydantic.dev/docs/ai/overview/migration/#tools-and-toolsets) | V1 | V2 | | --- | --- | | `pydantic_ai.builtin_tools` | `pydantic_ai.native_tools` | | `AgentBuiltinTool` | `AgentNativeTool` | | `pydantic_ai.native_tools.UrlContextTool` | [`pydantic_ai.native_tools.WebFetchTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebFetchTool) | | `builtin=` argument | `native=` | | `pydantic_ai.output.DeferredToolCalls` | [`DeferredToolRequests`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.DeferredToolRequests) | | `DeferredToolCalls.tool_calls` | [`DeferredToolRequests.calls`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.DeferredToolRequests.calls) | | `DeferredToolCalls.tool_defs` | Removed — it always returned an empty dict in V1 | | `pydantic_ai.toolsets.external.DeferredToolset` | [`ExternalToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.ExternalToolset) | | `FunctionToolset.tool()` on a context-free callable | [`FunctionToolset.tool_plain()`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.FunctionToolset.tool_plain)
— `tool()` now raises if the first parameter isn’t a `RunContext` | | `pydantic_ai.ext.aci.tool_from_aci`, `ACIToolset` | Removed; wrap the tool schemas with [`Tool.from_schema`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.Tool.from_schema) | | A `prepare` callback returning `None` | Return `[]` — returning `None` now raises `TypeError` instead of stripping all tools | | `WebSearch()` / `WebFetch()` falling back to a local implementation | `WebSearch(local='duckduckgo')` / `WebFetch(local=True)` — both are native-only by default and now raise on models that don’t support them | Messages, events and usage -------------------------- [](https://pydantic.dev/docs/ai/overview/migration/#messages-events-and-usage) The serialized `part_kind` wire values and the old field names’ validation aliases are retained, so message history written by V1 still deserializes in V2. | V1 | V2 | | --- | --- | | `BuiltinToolCallPart`, `BuiltinToolReturnPart` | `NativeToolCallPart`, `NativeToolReturnPart` | | `BuiltinToolCallEvent`, `BuiltinToolResultEvent` | Removed — native tool calls surface via `PartStartEvent`/`PartDeltaEvent` only | | `FunctionToolCallEvent`/`FunctionToolResultEvent` for output tools | `OutputToolCallEvent`/`OutputToolResultEvent` | | `FunctionToolCallEvent.call_id` | `FunctionToolCallEvent.tool_call_id` | | `FunctionToolResultEvent(result=...)`, `.result` | `FunctionToolResultEvent(part=...)`, `.part` | | `ModelResponse.vendor_details` | `ModelResponse.provider_details` | | `ModelResponse.vendor_id`, `ModelResponse.provider_request_id` | `ModelResponse.provider_response_id` | | `ModelResponse.builtin_tool_calls` | [`ModelResponse.native_tool_calls`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse.native_tool_calls) | | `ModelResponse.price()` | [`ModelResponse.cost()`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse.cost) | | `Usage` | [`RunUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage) | | `usage.request_tokens`, `usage.response_tokens` | `usage.input_tokens`, `usage.output_tokens` | | `UsageLimits(request_tokens_limit=)`, `(response_tokens_limit=)` | `UsageLimits(input_tokens_limit=)`, `(output_tokens_limit=)` | Results and streaming --------------------- [](https://pydantic.dev/docs/ai/overview/migration/#results-and-streaming) | V1 | V2 | | --- | --- | | `result.usage()`, `result.timestamp()` | `result.usage`, `result.timestamp` (properties) | | `stream.get()` | `stream.response` | | `StreamedRunResult.stream` | [`stream_output`](https://pydantic.dev/docs/ai/api/pydantic-ai/result/#pydantic_ai.result.StreamedRunResult.stream_output) | | `StreamedRunResult.stream_structured` | [`stream_response`](https://pydantic.dev/docs/ai/api/pydantic-ai/result/#pydantic_ai.result.StreamedRunResult.stream_response) | | `StreamedRunResult.stream_responses()` (plural, yielding `(response, is_last)`) | `stream_response()` (singular, yielding a bare `ModelResponse`; read the old `is_last` as `response.state != 'incomplete'`) | | `StreamedRunResult.validate_structured_output` | [`validate_response_output`](https://pydantic.dev/docs/ai/api/pydantic-ai/result/#pydantic_ai.result.StreamedRunResult.validate_response_output) | | `async for event in agent.run_stream_events(...)` | `async with agent.run_stream_events(...) as events:` then iterate — it is an async context manager only | Pydantic Graph -------------- [](https://pydantic.dev/docs/ai/overview/migration/#pydantic-graph) | V1 | V2 | | --- | --- | | `from pydantic_graph.beta import GraphBuilder` | `from pydantic_graph import GraphBuilder` | | `pydantic_graph.persistence` | No `pydantic_graph` equivalent — the builder API doesn’t snapshot graph state. To save, resume, and fork **agent run** state, [Pydantic AI Harness](https://pydantic.dev/docs/ai/harness/)
ships [`StepPersistence`](https://pydantic.dev/docs/ai/harness/step-persistence/) | | `pydantic_graph.mermaid` | Removed — render diagrams with `Graph.render()` | Pydantic Evals -------------- [](https://pydantic.dev/docs/ai/overview/migration/#pydantic-evals) | V1 | V2 | | --- | --- | | `Evaluator.name` (classmethod) | `Evaluator.get_serialization_name()` | | `evaluation_name` class attribute | [`Evaluator.get_default_evaluation_name()`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Evaluator.get_default_evaluation_name) | | `evaluator_version` class attribute | [`Evaluator.get_evaluator_version()`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Evaluator.get_evaluator_version) | | `Dataset(...)` without a name | `Dataset(name=...)` — now required | | Positional `name`/`max_concurrency`/`progress`/`retry_task`/`retry_evaluators` on `Dataset.evaluate()`/`evaluate_sync()` | Keyword-only | | Positional construction of `EvaluationResult` / `EvaluatorFailure` | Keyword-only | Instrumentation --------------- [](https://pydantic.dev/docs/ai/overview/migration/#instrumentation) | V1 | V2 | | --- | --- | | `InstrumentationSettings(version=1)`, `event_mode=`, `logger_provider=` | Removed; versions 2–4 still work but warn. The default is version 5 | | Reading run-span token usage from `gen_ai.usage.*` | Run spans report `gen_ai.aggregated_usage.*`; set `use_aggregated_usage_attribute_names=False` to keep the V1 names | Packaging --------- [](https://pydantic.dev/docs/ai/overview/migration/#packaging) A bare `uv add pydantic-ai` / `pip install pydantic-ai` now installs a slimmer set of extras. `bedrock`, `groq`, `mistral`, `cohere`, `xai`, `huggingface`, `temporal`, `ag-ui`, `ui`, and `spec` are no longer included by default — add the ones you use, e.g. `uv add 'pydantic-ai[bedrock,groq]'`. The `outlines-*`, `vertexai`, `fastmcp`, and `a2a` extras are removed outright. See the [installation guide](https://pydantic.dev/docs/ai/overview/install/) for the full list. Behavior changes with no code change ------------------------------------ [](https://pydantic.dev/docs/ai/overview/migration/#behavior-changes-with-no-code-change) These flip without any symbol changing name, so they can’t be found by grepping for an old name. Each is explained in full in the Upgrade Guide under [changes not covered by deprecation warnings](https://pydantic.dev/docs/ai/project/changelog/#changes-not-covered-by-deprecation-warnings) . * The default `end_strategy` changed from `'early'` to `'graceful'`, so function tools requested alongside a successful output tool now run instead of being skipped. See [Parallel Output Tool Calls](https://pydantic.dev/docs/ai/core-concepts/output/#parallel-output-tool-calls) . * `sequential=True` on a tool is now a per-tool barrier rather than a batch-wide serial switch, and applies to output tools too. * [`capture_run_messages()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.capture_run_messages) also captures the partial request/response of an interrupted run, marked `state='interrupted'`. * A resolved model profile now carries fields from other profile classes, where V1 filtered them out. Was this page helpful? Thanks for your feedback! --- # pydantic_graph.join | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#_top) pydantic\_graph.join ==================== Join operations and reducers for graph execution. This module provides the core components for joining parallel execution paths in a graph, including various reducer types that aggregate data from multiple sources into a single output. Join ---- [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.Join) **Bases:** `Generic[StateT, DepsT, InputT, OutputT]` A join operation that synchronizes and aggregates parallel execution paths. A join defines how to combine outputs from multiple parallel execution paths using a [`ReducerFunction`](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.ReducerFunction) . It specifies which fork it joins (if any) and manages the initialization of reducers. ### Methods [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#methods) #### as\_node [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.Join.as_node) def as_node(inputs: None = None) -> JoinNode[StateT, DepsT] def as_node(inputs: InputT) -> JoinNode[StateT, DepsT] Create a join node with bound inputs. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#returns) `JoinNode`\[`StateT`, `DepsT`\] — A [`JoinNode`](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.JoinNode) with this join and the bound inputs ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#parameters) **`inputs`** : `InputT` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.Join.as_node(inputs)) The input data to bind to this join, or None JoinNode -------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.JoinNode) **Bases:** `BaseNode[StateT, DepsT, Any]` A `BaseNode` that represents a builder join with bound inputs. `JoinNode` lets a [`BaseNode`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.BaseNode) subclass hand off to a builder [`Join`](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.Join) by wrapping the join together with the value it should receive as `inputs`. It is not meant to be run directly; returning a `JoinNode` from a `BaseNode.run` method tells the graph builder which join to invoke next. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#attributes) #### inputs [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.JoinNode.inputs) The inputs bound to this step. **Type:** [`Any`](https://docs.python.org/3/library/typing.html#typing.Any) #### join [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.JoinNode.join) The step to execute. **Type:** `Join`\[`StateT`, `DepsT`, [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ , [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] ### Methods [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#methods-1) #### run [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.JoinNode.run) `@async` def run(ctx: GraphRunContext[StateT, DepsT]) -> BaseNode[StateT, DepsT, Any] | End[Any] Attempt to run the join node. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#returns-1) `BaseNode`\[`StateT`, `DepsT`, [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] | `End`\[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] — The result of step execution ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#parameters-1) **`ctx`** : `GraphRunContext`\[`StateT`, `DepsT`\] [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.JoinNode.run(ctx)) The graph execution context ##### Raises [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#raises) * `NotImplementedError` — Always raised as StepNode is not meant to be run directly JoinState --------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.JoinState) The state of a join during graph execution associated to a particular fork run. ReduceFirstValue ---------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.ReduceFirstValue) **Bases:** `Generic[T]` A reducer that returns the first value it encounters, and cancels all other tasks. ### Methods [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#methods-2) #### \_\_call\_\_ [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.ReduceFirstValue.__call__) def __call__(ctx: ReducerContext[object, object], current: T, inputs: T) -> T The reducer function. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#returns-2) `T` ReducerContext -------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.ReducerContext) **Bases:** `Generic[StateT, DepsT]` Context information passed to reducer functions during graph execution. The reducer context provides access to the current graph state and dependencies. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#attributes-1) #### deps [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.ReducerContext.deps) The deps for the graph run. **Type:** `DepsT` #### state [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.ReducerContext.state) The state of the graph run. **Type:** `StateT` ### Methods [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#methods-3) #### cancel\_sibling\_tasks [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.ReducerContext.cancel_sibling_tasks) def cancel_sibling_tasks() Cancel all sibling tasks created from the same fork. You can call this if you want your join to have early-stopping behavior. SupportsSum ----------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.SupportsSum) **Bases:** [`Protocol`](https://docs.python.org/3/library/typing.html#typing.Protocol) A protocol for a type that supports adding to itself. reduce\_dict\_update -------------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.reduce_dict_update) def reduce_dict_update(current: dict[K, V], inputs: Mapping[K, V]) -> dict[K, V] A reducer that updates a dict. ### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#returns-3) [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[`K`, `V`\] reduce\_list\_append -------------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.reduce_list_append) def reduce_list_append(current: list[T], inputs: T) -> list[T] A reducer that appends to a list. ### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#returns-4) [`list`](https://docs.python.org/3/glossary.html#term-list) \[`T`\] reduce\_list\_extend -------------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.reduce_list_extend) def reduce_list_extend(current: list[T], inputs: Iterable[T]) -> list[T] A reducer that extends a list. ### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#returns-5) [`list`](https://docs.python.org/3/glossary.html#term-list) \[`T`\] reduce\_null ------------ [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.reduce_null) def reduce_null(current: None, inputs: Any) -> None A reducer that discards all input data and returns None. ### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#returns-6) [`None`](https://docs.python.org/3/library/constants.html#None) reduce\_sum ----------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.reduce_sum) def reduce_sum(current: NumericT, inputs: NumericT) -> NumericT A reducer that sums numbers. ### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#returns-7) `NumericT` ReducerFunction --------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.ReducerFunction) A function used for reducing inputs to a join node. **Default:** `TypeAliasType('ReducerFunction', ContextReducerFunction[StateT, DepsT, InputT, OutputT] | PlainReducerFunction[InputT, OutputT], type_params=(StateT, DepsT, InputT, OutputT))` Was this page helpful? Thanks for your feedback! --- # Pydantic Evals [Skip to content](https://pydantic.dev/docs/ai/evals/evals/#_top) Overview ======== **Pydantic Evals** is a powerful evaluation framework for systematically testing and evaluating AI systems, from simple LLM calls to complex multi-agent applications. Design Philosophy ----------------- [](https://pydantic.dev/docs/ai/evals/evals/#design-philosophy) Quick Navigation ---------------- [](https://pydantic.dev/docs/ai/evals/evals/#quick-navigation) **Getting Started:** * [Installation](https://pydantic.dev/docs/ai/evals/evals/#installation) * [Quick Start](https://pydantic.dev/docs/ai/evals/getting-started/quick-start/) * [Core Concepts](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/) **Evaluators:** * [Evaluators Overview](https://pydantic.dev/docs/ai/evals/evaluators/overview/) - Compare evaluator types and learn when to use each approach * [Built-in Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/built-in/) - Complete reference for exact match, instance checks, and other ready-to-use evaluators * [LLM as a Judge](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/) - Use LLMs to evaluate subjective qualities, complex criteria, and natural language outputs * [Custom Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/custom/) - Implement domain-specific scoring logic and custom evaluation metrics * [Span-Based Evaluation](https://pydantic.dev/docs/ai/evals/evaluators/span-based/) - Evaluate internal agent behavior (tool calls, execution flow) using OpenTelemetry traces. Essential for complex agents where correctness depends on _how_ the answer was reached, not just the final output. Also ensures eval assertions align with production telemetry. **How-To Guides:** * [Logfire Integration](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/) - Visualize results * [Dataset Management](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/) - Save, load, generate * [Concurrency & Performance](https://pydantic.dev/docs/ai/evals/how-to/concurrency/) - Control parallel execution * [Retry Strategies](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/) - Handle transient failures * [Metrics & Attributes](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/) - Track custom data * [Case Lifecycle Hooks](https://pydantic.dev/docs/ai/evals/how-to/lifecycle/) - Per-case setup, teardown, and context enrichment **Examples:** * [Simple Validation](https://pydantic.dev/docs/ai/evals/examples/simple-validation/) - Basic example **Reference:** * [API Documentation](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/) Code-First Evaluation --------------------- [](https://pydantic.dev/docs/ai/evals/evals/#code-first-evaluation) Pydantic Evals follows a **code-first approach** where you define all evaluation components (datasets, experiments, tasks, cases and evaluators) in Python code, or as serialized data loaded by Python code. This differs from platforms with fully web-based configuration. When you run an _Experiment_ you’ll see a progress indicator and can print the results wherever you run your python code (IDE, terminal, etc). You also get a report object back that you can serialize and store or send to a notebook or other application for further visualization and analysis. If you are using [Pydantic Logfire](https://logfire.pydantic.dev/docs/guides/web-ui/evals/) , your experiment results automatically appear in the Logfire web interface for visualization, comparison, and collaborative analysis. Logfire serves as a observability layer - you write and run evals in code, then view and analyze results in the web UI. Installation ------------ [](https://pydantic.dev/docs/ai/evals/evals/#installation) To install the Pydantic Evals package, run: * [pip](https://pydantic.dev/docs/ai/evals/evals/#tab-panel-18) * [uv](https://pydantic.dev/docs/ai/evals/evals/#tab-panel-19) Terminal pip install pydantic-evals Terminal uv add pydantic-evals `pydantic-evals` does not depend on `pydantic-ai`, but has an optional dependency on `logfire` if you’d like to use OpenTelemetry traces in your evals, or send evaluation results to [logfire](https://pydantic.dev/logfire) . * [pip](https://pydantic.dev/docs/ai/evals/evals/#tab-panel-20) * [uv](https://pydantic.dev/docs/ai/evals/evals/#tab-panel-21) Terminal pip install 'pydantic-evals[logfire]' Terminal uv add 'pydantic-evals[logfire]' Pydantic Evals Data Model ------------------------- [](https://pydantic.dev/docs/ai/evals/evals/#pydantic-evals-data-model) Pydantic Evals is built around a simple data model: ### Data Model Diagram [](https://pydantic.dev/docs/ai/evals/evals/#data-model-diagram) Dataset (1) ──────────── (Many) Case │ │ │ │ └─── (Many) Experiment ──┴─── (Many) Case results │ └─── (1) Task │ └─── (Many) Evaluator ### Key Relationships [](https://pydantic.dev/docs/ai/evals/evals/#key-relationships) 1. **Dataset → Cases**: One Dataset contains many Cases 2. **Dataset → Experiments**: One Dataset can be used across many Experiments over time 3. **Experiment → Case results**: One Experiment generates results by executing each Case 4. **Experiment → Task**: One Experiment evaluates one defined Task 5. **Experiment → Evaluators**: One Experiment uses multiple Evaluators. Dataset-wide Evaluators are run against all Cases, and Case-specific Evaluators against their respective Cases ### Data Flow [](https://pydantic.dev/docs/ai/evals/evals/#data-flow) 1. **Dataset creation**: Define cases and evaluators in YAML/JSON, or directly in Python 2. **Experiment execution**: Run `dataset.evaluate_sync(task_function)` 3. **Cases run**: Each Case is executed against the Task 4. **Evaluation**: Evaluators score the Task outputs for each Case 5. **Results**: All Case results are collected into a summary report For a deeper understanding, see [Core Concepts](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/) . Datasets and Cases ------------------ [](https://pydantic.dev/docs/ai/evals/evals/#datasets-and-cases) In Pydantic Evals, everything begins with [`Dataset`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset) s and [`Case`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Case) s: * **[`Dataset`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset) **: A collection of test Cases designed for the evaluation of a specific task or function * **[`Case`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Case) **: A single test scenario corresponding to Task inputs, with optional expected outputs, metadata, and case-specific evaluators simple\_eval\_dataset.py from pydantic_evals import Case, Dataset case1 = Case( name='simple_case', inputs='What is the capital of France?', expected_output='Paris', metadata={'difficulty': 'easy'}, ) dataset = Dataset(name='capital_quiz', cases=[case1]) _(This example is complete, it can be run “as is”)_ See [Dataset Management](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/) to learn about saving, loading, and generating datasets. Evaluators ---------- [](https://pydantic.dev/docs/ai/evals/evals/#evaluators) [`Evaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Evaluator) s analyze and score the results of your Task when tested against a Case. These can be deterministic, code-based checks (such as testing model output format with a regex, or checking for the appearance of PII or sensitive data), or they can assess non-deterministic model outputs for qualities like accuracy, precision/recall, hallucinations, or instruction-following. While both kinds of testing are useful in LLM systems, classical code-based tests are cheaper and easier than tests which require either human or machine review of model outputs. Pydantic Evals includes several [built-in evaluators](https://pydantic.dev/docs/ai/evals/evaluators/built-in/) and allows you to define [custom evaluators](https://pydantic.dev/docs/ai/evals/evaluators/custom/) : simple\_eval\_evaluator.py from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext from pydantic_evals.evaluators.common import IsInstance from simple_eval_dataset import dataset dataset.add_evaluator(IsInstance(type_name='str')) # (1) @dataclass class MyEvaluator(Evaluator): async def evaluate(self, ctx: EvaluatorContext[str, str]) -> float: # (2) if ctx.output == ctx.expected_output: return 1.0 elif ( isinstance(ctx.output, str) and ctx.expected_output.lower() in ctx.output.lower() ): return 0.8 else: return 0.0 dataset.add_evaluator(MyEvaluator()) You can add built-in evaluators to a dataset using the [`add_evaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset.add_evaluator) method. This custom evaluator returns a simple score based on whether the output matches the expected output. _(This example is complete, it can be run “as is”)_ Learn more: * [Evaluators Overview](https://pydantic.dev/docs/ai/evals/evaluators/overview/) - When to use different types * [Built-in Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/built-in/) - Complete reference * [LLM Judge](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/) - Using LLMs as evaluators * [Custom Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/custom/) - Write your own logic * [Span-Based Evaluation](https://pydantic.dev/docs/ai/evals/evaluators/span-based/) - Analyze execution traces Running Experiments ------------------- [](https://pydantic.dev/docs/ai/evals/evals/#running-experiments) Performing evaluations involves running a task against all cases in a dataset, also known as running an “experiment”. Putting the above two examples together and using the more declarative `evaluators` kwarg to [`Dataset`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset) : simple\_eval\_complete.py from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import Evaluator, EvaluatorContext, IsInstance case1 = Case( # (1) name='simple_case', inputs='What is the capital of France?', expected_output='Paris', metadata={'difficulty': 'easy'}, ) class MyEvaluator(Evaluator[str, str]): def evaluate(self, ctx: EvaluatorContext[str, str]) -> float: if ctx.output == ctx.expected_output: return 1.0 elif ( isinstance(ctx.output, str) and ctx.expected_output.lower() in ctx.output.lower() ): return 0.8 else: return 0.0 dataset = Dataset( name='capital_quiz', cases=[case1], evaluators=[IsInstance(type_name='str'), MyEvaluator()], # (2) ) async def guess_city(question: str) -> str: # (3) return 'Paris' report = dataset.evaluate_sync(guess_city) # (4) report.print(include_input=True, include_output=True, include_durations=False) # (5) """ Evaluation Summary: guess_city ┏━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓ ┃ Case ID ┃ Inputs ┃ Outputs ┃ Scores ┃ Assertions ┃ ┡━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩ │ simple_case │ What is the capital of France? │ Paris │ MyEvaluator: 1.00 │ ✔ │ ├─────────────┼────────────────────────────────┼─────────┼───────────────────┼────────────┤ │ Averages │ │ │ MyEvaluator: 1.00 │ 100.0% ✔ │ └─────────────┴────────────────────────────────┴─────────┴───────────────────┴────────────┘ """ Create a [test case](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Case) as above Create a [`Dataset`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset) with test cases and [`evaluators`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset.evaluators) Our function to evaluate. Run the evaluation with [`evaluate_sync`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset.evaluate_sync) , which runs the function against all test cases in the dataset, and returns an [`EvaluationReport`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.EvaluationReport) object. Print the report with [`print`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.EvaluationReport.print) , which shows the results of the evaluation. We have omitted duration here just to keep the printed output from changing from run to run. _(This example is complete, it can be run “as is”)_ See [Quick Start](https://pydantic.dev/docs/ai/evals/getting-started/quick-start/) for more examples and [Concurrency & Performance](https://pydantic.dev/docs/ai/evals/how-to/concurrency/) to learn about controlling parallel execution. API Reference ------------- [](https://pydantic.dev/docs/ai/evals/evals/#api-reference) For comprehensive coverage of all classes, methods, and configuration options, see the detailed [API Reference documentation](https://ai.pydantic.dev/api/pydantic_evals/dataset/) . Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/evals/#next-steps) 1. **Start with simple evaluations** using [Quick Start](https://pydantic.dev/docs/ai/evals/getting-started/quick-start/) 2. **Understand the data model** with [Core Concepts](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/) 3. **Explore built-in evaluators** in [Built-in Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/built-in/) 4. **Integrate with Logfire** for visualization: [Logfire Integration](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/) 5. **Build comprehensive test suites** with [Dataset Management](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/) 6. **Implement custom evaluators** for domain-specific metrics: [Custom Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/custom/) Was this page helpful? Thanks for your feedback! --- # xai | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/models/xai/#_top) xai === Setup ----- [](https://pydantic.dev/docs/ai/api/models/xai/#setup) For details on how to set up authentication with this model, see [model configuration for xAI](https://pydantic.dev/docs/ai/models/xai/) . xAI model implementation using [xAI SDK](https://github.com/xai-org/xai-sdk-python) . XaiModel -------- [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModel) **Bases:** `Model[AsyncClient]` A model that uses the xAI SDK to interact with xAI models. ### Attributes [](https://pydantic.dev/docs/ai/api/models/xai/#attributes) #### model\_name [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModel.model_name) The model name. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### system [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModel.system) The model provider. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) ### Methods [](https://pydantic.dev/docs/ai/api/models/xai/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModel.__init__) def __init__( model_name: XaiModelName, *, provider: Literal['xai'] | Provider[AsyncClient] = 'xai', profile: ModelProfileSpec | None = None, settings: ModelSettings | None = None, ) Initialize the xAI model. ##### Parameters [](https://pydantic.dev/docs/ai/api/models/xai/#parameters) **`model_name`** : `XaiModelName` [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModel.__init__(model_name)) The name of the xAI model to use (e.g., “grok-4.3”) **`provider`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘xai’\] | `Provider`\[`AsyncClient`\] _Default:_ `'xai'` [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModel.__init__(provider)) The provider to use for API calls. Defaults to `'xai'`. **`profile`** : [`ModelProfileSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfileSpec) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModel.__init__(profile)) Optional model profile specification. Defaults to a profile picked by the provider based on the model name. **`settings`** : [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModel.__init__(settings)) Optional model settings. #### request [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModel.request) `@async` def request( messages: list[ModelMessage], model_settings: ModelSettings | None, model_request_parameters: ModelRequestParameters, ) -> ModelResponse Make a request to the xAI model. ##### Returns [](https://pydantic.dev/docs/ai/api/models/xai/#returns) [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) #### request\_stream [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModel.request_stream) `@async` def request_stream( messages: list[ModelMessage], model_settings: ModelSettings | None, model_request_parameters: ModelRequestParameters, run_context: RunContext[Any] | None = None, ) -> AsyncGenerator[StreamedResponse] Make a streaming request to the xAI model. ##### Returns [](https://pydantic.dev/docs/ai/api/models/xai/#returns-1) [`AsyncGenerator`](https://docs.python.org/3/library/typing.html#typing.AsyncGenerator) \[`StreamedResponse`\] #### supported\_native\_tools [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModel.supported_native_tools) `@classmethod` def supported_native_tools(cls) -> frozenset[type] Return the set of builtin tool types this model can handle. ##### Returns [](https://pydantic.dev/docs/ai/api/models/xai/#returns-2) [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[[`type`](https://docs.python.org/3/glossary.html#term-type)\ \] XaiModelSettings ---------------- [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings) **Bases:** [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) Settings specific to xAI models. See [xAI SDK documentation](https://docs.x.ai/docs) for more details on these parameters. ### Attributes [](https://pydantic.dev/docs/ai/api/models/xai/#attributes-1) #### xai\_agent\_count [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_agent_count) Number of agents for xAI multi-agent models (e.g. `grok-4.20-multi-agent`). Forwarded to `chat.create(agent_count=...)`. Documented values are `4` and `16`; more agents increase token usage and latency. Only affects multi-agent models; other models ignore it. The multi-agent API is in beta, so the accepted values may change. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### xai\_include\_code\_execution\_output [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_include_code_execution_output) Whether to include the code execution results in the response. Corresponds to the `code_interpreter_call.outputs` value of the `include` parameter in the Responses API. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### xai\_include\_collections\_search\_output [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_include_collections_search_output) Whether to include the collections search results in the response. Corresponds to the `collections_search_call.outputs` value of the `include` parameter in the Responses API. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### xai\_include\_encrypted\_content [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_include_encrypted_content) Whether to include the encrypted content in the response. Corresponds to the `use_encrypted_content` value of the model settings in the Responses API. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### xai\_include\_inline\_citations [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_include_inline_citations) Whether to include inline citations in the response. Corresponds to the `inline_citations` option in the xAI `include` parameter. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### xai\_include\_mcp\_output [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_include_mcp_output) Whether to include the MCP results in the response. Corresponds to the `mcp_call.outputs` value of the `include` parameter in the Responses API. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### xai\_include\_web\_search\_output [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_include_web_search_output) Whether to include the web search results in the response. Corresponds to the `web_search_call.action.sources` value of the `include` parameter in the Responses API. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### xai\_include\_x\_search\_output [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_include_x_search_output) Whether to include the X search results in the response. Corresponds to the `x_search_call.outputs` value of the `include` parameter in the Responses API. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### xai\_logprobs [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_logprobs) Whether to return log probabilities of the output tokens or not. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### xai\_max\_turns [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_max_turns) Maximum number of agentic turns xAI’s server-side tool loop may take. Only affects requests that use xAI’s server-side native tools (e.g. web search, code execution, X search): xAI iterates up to this many turns — calling those server-side tools and processing their results — before returning a final response. It has no effect on ordinary client-side tools or on Pydantic AI’s own agent loop; use [`UsageLimits`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits) to bound those. With parallel tool calls enabled, multiple tool calls can occur within a single turn, so `xai_max_turns` does not necessarily equal the total number of tool calls made. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### xai\_previous\_response\_id [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_previous_response_id) The ID of the previous response to continue the conversation. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### xai\_reasoning\_effort [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_reasoning_effort) Reasoning effort level for Grok reasoning models. See [https://docs.x.ai](https://docs.x.ai/) for details. **Type:** `GrokReasoningEffort` #### xai\_store\_messages [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_store_messages) Whether to store messages on xAI’s servers for conversation continuity. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### xai\_top\_logprobs [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_top_logprobs) An integer between 0 and 20 specifying the number of most likely tokens to return at each position. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### xai\_user [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_user) A unique identifier representing your end-user, which can help xAI to monitor and detect abuse. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) XaiStreamedResponse ------------------- [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiStreamedResponse) **Bases:** `StreamedResponse` Implementation of `StreamedResponse` for xAI SDK. ### Attributes [](https://pydantic.dev/docs/ai/api/models/xai/#attributes-2) #### model\_name [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiStreamedResponse.model_name) Get the model name of the response. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### provider\_name [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiStreamedResponse.provider_name) The model provider. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### provider\_url [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiStreamedResponse.provider_url) Get the provider base URL. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### system [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiStreamedResponse.system) The model provider system name. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### timestamp [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiStreamedResponse.timestamp) Get the timestamp of the response. **Type:** [`datetime`](https://docs.python.org/3/library/datetime.html#module-datetime) XaiModelName ------------ [](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelName) Possible xAI model names. `grok-4.5`/`grok-4.5-latest` are bridged with a local `Literal` because `xai_sdk`’s `ChatModel` doesn’t list them yet (as of 1.17.0). Drop the literal once the `xai-sdk` floor is bumped past the release that adds them to `ChatModel`. **Default:** `str | ChatModel | Literal['grok-4.5', 'grok-4.5-latest']` Was this page helpful? Thanks for your feedback! --- # Overview | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/evaluators/overview/#_top) Overview ======== Evaluators are the core of Pydantic Evals. They analyze task outputs and provide scores, labels, or pass/fail assertions. When to Use Different Evaluators -------------------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#when-to-use-different-evaluators) ### Deterministic Checks (Fast & Reliable) [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#deterministic-checks-fast--reliable) Use deterministic evaluators when you can define exact rules: | Evaluator | Use Case | Example | | --- | --- | --- | | [`EqualsExpected`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EqualsExpected) | Exact output match | Structured data, classification | | [`Equals`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Equals) | Equals specific value | Checking for sentinel values | | [`Contains`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Contains) | Substring/element check | Required keywords, PII detection | | [`IsInstance`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.IsInstance) | Type validation | Format validation | | [`MaxDuration`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.MaxDuration) | Performance threshold | SLA compliance | | [`HasMatchingSpan`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.HasMatchingSpan) | Behavior verification | Tool calls, code paths | | [`ToolCorrectness`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.ToolCorrectness) | Required tool coverage | Multiset of tool names invoked | | [`TrajectoryMatch`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.TrajectoryMatch) | Tool-call sequence quality | F1 against expected trajectory | | [`ArgumentCorrectness`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.ArgumentCorrectness) | Tool argument checks | Refund `order_id`, search query | | [`MaxToolCalls`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.MaxToolCalls) | Budget discipline | Tool-call budget | | [`MaxModelRequests`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.MaxModelRequests) | Budget discipline | Model-request budget | **Advantages:** * Fast execution (microseconds to milliseconds) * Deterministic results * No cost * Easy to debug **When to use:** * Format validation (JSON structure, type checking) * Required content checks (must contain X, must not contain Y) * Performance requirements (latency, token counts) * Behavioral checks (which tools were called, which code paths executed) ### LLM-as-a-Judge (Flexible & Nuanced) [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#llm-as-a-judge-flexible--nuanced) Use [`LLMJudge`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.LLMJudge) when evaluation requires understanding or judgment: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import LLMJudge dataset = Dataset( name='llm_judge_example', cases=[Case(inputs='What is 2+2?', expected_output='4')], evaluators=[\ LLMJudge(\ rubric='Response is factually accurate based on the input',\ include_input=True,\ )\ ], ) For metrics aligned with widely-used evaluation methods (G-Eval, the Ragas RAG metrics, GEMBA), see [Standard Quality Metrics](https://pydantic.dev/docs/ai/evals/evaluators/standard-quality-metrics/) : the [`GEval`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.GEval) evaluator plus ready-made [`LLMJudge`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.LLMJudge) rubrics you can copy and adapt. To plug in the _exact_ upstream implementations of external frameworks, see [Third-Party Integrations](https://pydantic.dev/docs/ai/evals/evaluators/framework-integrations/) . **Advantages:** * Can evaluate subjective qualities (helpfulness, tone, creativity) * Understands natural language * Can follow complex rubrics * Flexible across domains **Disadvantages:** * Slower (seconds per evaluation) * Costs money * Non-deterministic * Can have biases **When to use:** * Factual accuracy * Relevance and helpfulness * Tone and style * Completeness * Following instructions * RAG quality (groundedness, citation accuracy) ### Custom Evaluators [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#custom-evaluators) Custom evaluators can be useful if you want to make use of any evaluation logic we don’t provide with the framework. They are frequently useful for domain-specific logic: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class ValidSQL(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: try: import sqlparse sqlparse.parse(ctx.output) return True except Exception: return False **When to use:** * Domain-specific validation (SQL syntax, regex patterns, business rules) * External API calls (running generated code, checking databases) * Complex calculations (precision/recall, BLEU scores) * Integration checks (does API call succeed?) Evaluation Types ---------------- [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#evaluation-types) Evaluators essentially return three types of results: ### 1\. Assertions (bool) [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#1-assertions-bool) Pass/fail checks that appear as ✔ or ✗ in reports: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class HasKeyword(Evaluator): keyword: str def evaluate(self, ctx: EvaluatorContext) -> bool: return self.keyword in ctx.output **Use for:** Binary checks, quality gates, compliance requirements ### 2\. Scores (int or float) [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#2-scores-int-or-float) Numeric metrics: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class ConfidenceScore(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> float: # Analyze and return score return 0.87 # 87% confidence **Use for:** Quality metrics, ranking, A/B testing, regression tracking ### 3\. Labels (str) [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#3-labels-str) Categorical classifications: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class SentimentClassifier(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> str: if 'error' in ctx.output.lower(): return 'error' elif 'success' in ctx.output.lower(): return 'success' return 'neutral' **Use for:** Classification, error categorization, quality buckets ### Multiple Results [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#multiple-results) You can return multiple evaluations from a single evaluator: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class ComprehensiveCheck(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> dict[str, bool | float | str]: return { 'valid_format': self._check_format(ctx.output), # bool 'quality_score': self._score_quality(ctx.output), # float 'category': self._classify(ctx.output), # str } def _check_format(self, output: str) -> bool: return True def _score_quality(self, output: str) -> float: return 0.85 def _classify(self, output: str) -> str: return 'good' Combining Evaluators -------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#combining-evaluators) Mix and match evaluators to create comprehensive evaluation suites: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import ( Contains, IsInstance, LLMJudge, MaxDuration, ) dataset = Dataset( name='layered_evaluation', cases=[Case(inputs='test', expected_output='result')], evaluators=[\ # Fast deterministic checks first\ IsInstance(type_name='str'),\ Contains(value='required_field'),\ MaxDuration(seconds=2.0),\ # Slower LLM checks after\ LLMJudge(\ rubric='Response is accurate and helpful',\ include_input=True,\ ),\ ], ) Case-specific evaluators ------------------------ [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#case-specific-evaluators) Case-specific evaluators are one of the most powerful features for building comprehensive evaluation suites. You can attach evaluators to individual [`Case`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Case) objects that only run for those specific cases: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import IsInstance, LLMJudge dataset = Dataset( name='case_specific_evaluators', cases=[\ Case(\ name='greeting_response',\ inputs='Say hello',\ evaluators=[\ # This evaluator only runs for this case\ LLMJudge(\ rubric='Response is warm and friendly, uses casual tone',\ include_input=True,\ ),\ ],\ ),\ Case(\ name='formal_response',\ inputs='Write a business email',\ evaluators=[\ # Different requirements for this case\ LLMJudge(\ rubric='Response is professional and formal, uses business language',\ include_input=True,\ ),\ ],\ ),\ ], evaluators=[\ # This runs for ALL cases\ IsInstance(type_name='str'),\ ], ) ### Why Case-Specific Evaluators Matter [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#why-case-specific-evaluators-matter) Case-specific evaluators solve a fundamental problem with one-size-fits-all evaluation: **if you could write a single evaluator rubric that perfectly captured your requirements across all cases, you’d just incorporate that rubric into your agent’s instructions**. (Note: this is less relevant in cases where you want to use a cheaper model in production and assess it using a more expensive model, but in many cases it makes sense to use the best model you can in production.) The power of case-specific evaluation comes from the nuance: * **Different cases have different requirements**: A customer support response needs empathy; a technical API response needs precision * **Avoid “inmates running the asylum”**: If your LLMJudge rubric is generic enough to work everywhere, your agent should already be following it * **Capture nuanced golden behavior**: Each case can specify exactly what “good” looks like for that scenario ### Building Golden Datasets with Case-Specific LLMJudge [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#building-golden-datasets-with-case-specific-llmjudge) A particularly powerful pattern is using case-specific [`LLMJudge`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.LLMJudge) evaluators to quickly build comprehensive, maintainable evaluation suites. Instead of needing exact `expected_output` values, you can describe what you care about: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import LLMJudge dataset = Dataset( name='golden_dataset', cases=[\ Case(\ name='handle_refund_request',\ inputs={'query': 'I want my money back', 'order_id': '12345'},\ evaluators=[\ LLMJudge(\ rubric="""\ Response should:\ 1. Acknowledge the refund request empathetically\ 2. Ask for the reason for the refund\ 3. Mention our 30-day refund policy\ 4. NOT process the refund immediately (needs manager approval)\ """,\ include_input=True,\ ),\ ],\ ),\ Case(\ name='handle_shipping_question',\ inputs={'query': 'Where is my order?', 'order_id': '12345'},\ evaluators=[\ LLMJudge(\ rubric="""\ Response should:\ 1. Confirm the order number\ 2. Provide tracking information\ 3. Give estimated delivery date\ 4. Be brief and factual (not overly apologetic)\ """,\ include_input=True,\ ),\ ],\ ),\ Case(\ name='handle_angry_customer',\ inputs={'query': 'This is completely unacceptable!', 'order_id': '12345'},\ evaluators=[\ LLMJudge(\ rubric="""\ Response should:\ 1. Prioritize de-escalation with empathy\ 2. Avoid being defensive\ 3. Offer concrete next steps\ 4. Use phrases like "I understand" and "Let me help"\ """,\ include_input=True,\ ),\ ],\ ),\ ], ) This approach lets you: * **Build comprehensive test suites quickly**: Just describe what you want per case * **Maintain easily**: Update rubrics as requirements change, without regenerating outputs * **Cover edge cases naturally**: Add new cases with specific requirements as you discover them * **Capture domain knowledge**: Each rubric documents what “good” means for that scenario The LLM evaluator excels at understanding nuanced requirements and assessing compliance, making this a practical way to create thorough evaluation coverage without brittleness. Async vs Sync ------------- [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#async-vs-sync) Evaluators can be sync or async: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class SyncEvaluator(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return True async def some_async_operation() -> bool: return True @dataclass class AsyncEvaluator(Evaluator): async def evaluate(self, ctx: EvaluatorContext) -> bool: result = await some_async_operation() return result Pydantic Evals handles both automatically. Use async when: * Making API calls * Running database queries * Performing I/O operations * Calling LLMs (like [`LLMJudge`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.LLMJudge) ) Evaluation Context ------------------ [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#evaluation-context) All evaluators receive an [`EvaluatorContext`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorContext) : * `ctx.inputs` - Task inputs * `ctx.output` - Task output (to evaluate) * `ctx.expected_output` - Expected output (if provided) * `ctx.metadata` - Case metadata (if provided) * `ctx.duration` - Task execution time (seconds) * `ctx.span_tree` - OpenTelemetry spans (if logfire configured) * `ctx.metrics` - Custom metrics dict * `ctx.attributes` - Custom attributes dict This gives evaluators full context to make informed assessments. Error Handling -------------- [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#error-handling) If an evaluator raises an exception, it’s captured as an [`EvaluatorFailure`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorFailure) : from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext def risky_operation(output: str) -> bool: # This might raise an exception if 'error' in output: raise ValueError('Found error in output') return True @dataclass class RiskyEvaluator(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: # If this raises an exception, it will be captured result = risky_operation(ctx.output) return result Failures appear in `report.cases[i].evaluator_failures` with: * Evaluator name * Error message * Full stacktrace Use retry configuration to handle transient failures (see [Retry Strategies](https://pydantic.dev/docs/ai/evals/how-to/retry-strategies/) ). Report Evaluators (Experiment-Wide) ----------------------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#report-evaluators-experiment-wide) All the evaluators above run once per case. **Report evaluators** are different: they run once per experiment after all cases have been evaluated, and analyze the full set of results together. Use report evaluators for experiment-wide statistics like: * **Confusion matrices** — visualize classification accuracy across classes * **Precision-recall curves** — assess ranking quality with AUC scores * **Scalar metrics** — overall accuracy, F1, BLEU, or any single number * **Summary tables** — per-class breakdowns, error category summaries from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import ConfusionMatrixEvaluator dataset = Dataset( name='report_evaluator_example', cases=[\ Case(inputs='meow', expected_output='cat'),\ Case(inputs='woof', expected_output='dog'),\ ], report_evaluators=[\ ConfusionMatrixEvaluator(\ predicted_from='output',\ expected_from='expected_output',\ ),\ ], ) **See:** [Report Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/) for the full guide, including built-in report evaluators and how to write custom ones. Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/evaluators/overview/#next-steps) * **[Native Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/built-in/) ** - Complete reference of all provided evaluators * **[LLM Judge](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/) ** - Deep dive on LLM-as-a-Judge evaluation * **[Standard Quality Metrics](https://pydantic.dev/docs/ai/evals/evaluators/standard-quality-metrics/) ** - G-Eval, plus LLM judge rubrics for common RAG and translation metrics * **[Third-Party Integrations](https://pydantic.dev/docs/ai/evals/evaluators/framework-integrations/) ** - Wrap Ragas, DeepEval, and other metrics libraries * **[Custom Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/custom/) ** - Write your own evaluation logic * **[Report Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/) ** - Experiment-wide analyses * **[Span-Based Evaluation](https://pydantic.dev/docs/ai/evals/evaluators/span-based/) ** - Evaluate using OpenTelemetry spans * **[Agentic Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/agentic/) ** - Trajectory, tool-correctness, argument, and step-budget checks for agents Was this page helpful? Thanks for your feedback! --- # Core Concepts | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#_top) Core Concepts ============= This page explains the key concepts in Pydantic Evals and how they work together. Pydantic Evals is built around these core concepts: * **[`Dataset`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset) ** - A static definition containing test cases and evaluators * **[`Case`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Case) ** - A single test scenario with inputs and optional expected outputs * **[`Evaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Evaluator) ** - Logic for scoring or validating individual outputs * **[`ReportEvaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.ReportEvaluator) ** - Logic for analyzing full experiment results (e.g., confusion matrices, accuracy) * **Experiment** - The act of running a task function against all cases in a dataset. (This corresponds to a call to `Dataset.evaluate`.) * **[`EvaluationReport`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.EvaluationReport) ** - The results from running an experiment The key distinction is between: * **Definition** (`Dataset` with `Case`s, `Evaluator`s, and `ReportEvaluator`s) - what you want to test * **Execution** (Experiment) - running your task against those tests * **Results** (`EvaluationReport` with case results and experiment-wide analyses) - what happened during the experiment Unit Testing Analogy -------------------- [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#unit-testing-analogy) A helpful way to think about Pydantic Evals: | Unit Testing | Pydantic Evals | | --- | --- | | Test function | [`Case`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Case)
+ [`Evaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Evaluator) | | Test suite | [`Dataset`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset) | | Running tests (`pytest`) | **Experiment** (`dataset.evaluate(task)`) | | Test report | [`EvaluationReport`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.EvaluationReport) | | `assert` | Evaluator returning `bool` | **Key Difference**: AI systems are probabilistic, so instead of simple pass/fail, evaluations can have: * Quantitative scores (0.0 to 1.0) * Qualitative labels (“good”, “acceptable”, “poor”) * Pass/fail assertions with explanatory reasons Just like you can run `pytest` multiple times on the same test suite, you can run multiple experiments on the same dataset to compare different implementations or track changes over time. Dataset ------- [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#dataset) A [`Dataset`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset) is a collection of test cases and evaluators that define an evaluation suite. from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import IsInstance dataset = Dataset( name='my_eval_suite', cases=[\ Case(inputs='test input', expected_output='test output'),\ ], evaluators=[\ IsInstance(type_name='str'),\ ], ) ### Key Features [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#key-features) * **Type-safe**: Generic over `InputsT`, `OutputT`, and `MetadataT` types * **Serializable**: Can be saved to/loaded from YAML or JSON files * **Evaluable**: Run against any function with matching input/output types ### `Dataset`\-Level vs `Case`\-Level Evaluators [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#dataset-level-vs-case-level-evaluators) Evaluators can be defined at two levels: * **`Dataset`\-level**: Apply to all cases in the dataset * **`Case`\-level**: Apply only to specific cases from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import EqualsExpected, IsInstance dataset = Dataset( name='case_level_evaluators', cases=[\ Case(\ name='special_case',\ inputs='test',\ expected_output='TEST',\ evaluators=[\ # This evaluator only runs for this case\ EqualsExpected(),\ ],\ ),\ ], evaluators=[\ # This evaluator runs for ALL cases\ IsInstance(type_name='str'),\ ], ) Experiments ----------- [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#experiments) An **Experiment** is what happens when you execute a task function against all cases in a dataset. This is the bridge between your static test definition (the Dataset) and your results (the EvaluationReport). ### Running an Experiment [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#running-an-experiment) You run an experiment by calling [`evaluate()`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset.evaluate) or [`evaluate_sync()`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset.evaluate_sync) on a dataset: from pydantic_evals import Case, Dataset # Define your dataset (static definition) dataset = Dataset( name='uppercase_experiment', cases=[\ Case(inputs='hello', expected_output='HELLO'),\ Case(inputs='world', expected_output='WORLD'),\ ], ) # Define your task def uppercase_task(text: str) -> str: return text.upper() # Run the experiment (execution) report = dataset.evaluate_sync(uppercase_task) ### What Happens During an Experiment [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#what-happens-during-an-experiment) When you run an experiment: 1. **Setup**: The dataset loads all cases, evaluators, and report evaluators 2. **Execution**: For each case: 1. The task function is called with `case.inputs` 2. Execution time is measured and OpenTelemetry spans are captured (if `logfire` is configured) 3. The outputs of the task function for each case are recorded 3. **Case Evaluation**: For each case output: 1. All dataset-level evaluators are run 2. Case-specific evaluators are run (if any) 3. Results are collected (scores, assertions, labels) 4. **Report Evaluation**: If report evaluators are configured, they run over the full set of results to produce experiment-wide analyses (confusion matrices, precision-recall curves, scalar metrics, tables, etc.) 5. **Reporting**: All results are aggregated into an [`EvaluationReport`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.EvaluationReport) , including both per-case results and experiment-wide analyses ### Multiple Experiments from One Dataset [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#multiple-experiments-from-one-dataset) A key feature of Pydantic Evals is that you can run the same dataset against different task implementations: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import EqualsExpected dataset = Dataset( name='comparison_test', cases=[\ Case(inputs='hello', expected_output='HELLO'),\ ], evaluators=[EqualsExpected()], ) # Original implementation def task_v1(text: str) -> str: return text.upper() # Improved implementation (with exclamation) def task_v2(text: str) -> str: return text.upper() + '!' # Compare results report_v1 = dataset.evaluate_sync(task_v1) report_v2 = dataset.evaluate_sync(task_v2) avg_v1 = report_v1.averages() avg_v2 = report_v2.averages() print(f'V1 pass rate: {avg_v1.assertions if avg_v1 and avg_v1.assertions else 0}') #> V1 pass rate: 1.0 print(f'V2 pass rate: {avg_v2.assertions if avg_v2 and avg_v2.assertions else 0}') #> V2 pass rate: 0 This allows you to: * **Compare implementations** across versions * **Track performance** over time * **A/B test** different approaches * **Validate changes** before deployment Case ---- [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#case) A [`Case`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Case) represents a single test scenario with specific inputs and optional expected outputs. from pydantic_evals import Case from pydantic_evals.evaluators import EqualsExpected case = Case( name='test_uppercase', # Optional, but recommended for reporting inputs='hello world', # Required: inputs to your task expected_output='HELLO WORLD', # Optional: expected output metadata={'category': 'basic'}, # Optional: arbitrary metadata evaluators=[EqualsExpected()], # Optional: case-specific evaluators ) ### Case Components [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#case-components) #### Inputs [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#inputs) The inputs to pass to the task being evaluated. Can be any type: from pydantic import BaseModel from pydantic_evals import Case class MyInputModel(BaseModel): field1: str # Simple types Case(inputs='hello') Case(inputs=42) # Complex types Case(inputs={'query': 'What is AI?', 'max_tokens': 100}) Case(inputs=MyInputModel(field1='value')) #### Expected Output [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#expected-output) The expected result, used by evaluators like [`EqualsExpected`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EqualsExpected) : from pydantic_evals import Case Case( inputs='2 + 2', expected_output='4', ) If no `expected_output` is provided, evaluators that require it (like `EqualsExpected`) will skip that case. #### Metadata [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#metadata) Arbitrary data that evaluators can access via [`EvaluatorContext`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorContext) : from pydantic_evals import Case Case( inputs='question', metadata={ 'difficulty': 'hard', 'category': 'math', 'source': 'exam_2024', }, ) Metadata is useful for: * Filtering cases during analysis * Providing context to evaluators * Organizing test suites #### Evaluators [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#evaluators) Cases can have their own evaluators that only run for that specific case. This is particularly powerful for building comprehensive evaluation suites where different cases have different requirements - if you could write one evaluator rubric that worked perfectly for all cases, you’d just incorporate it into your agent instructions. Case-specific [`LLMJudge`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.LLMJudge) evaluators are especially useful for quickly building maintainable golden datasets by describing what “good” looks like for each scenario. See [Case-specific evaluators](https://pydantic.dev/docs/ai/evals/evaluators/overview/#case-specific-evaluators) for a more detailed explanation and examples. Evaluator --------- [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#evaluator) An [`Evaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Evaluator) assesses the output of your task and returns one or more scores, labels, or assertions. Each score, label or assertion can also have an optional string-value reason associated. ### Evaluator Types [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#evaluator-types) Evaluators return different types of results: | Return Type | Purpose | Example | | --- | --- | --- | | `bool` | **Assertion** - Pass/fail check | `True` → ✔, `False` → ✗ | | `int` or `float` | **Score** - Numeric quality metric | `0.95`, `87` | | `str` | **Label** - Categorical result | `"correct"`, `"hallucination"` | from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class ExactMatch(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return ctx.output == ctx.expected_output # Assertion @dataclass class Confidence(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> float: # Analyze output and return confidence score return 0.95 # Score @dataclass class Classifier(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> str: if 'error' in ctx.output.lower(): return 'error' # Label return 'success' Evaluators can also return instances of [`EvaluationReason`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluationReason) , and dictionaries mapping labels to output values. See the [custom evaluator return types](https://pydantic.dev/docs/ai/evals/evaluators/custom/#return-types) docs for more detail. ### EvaluatorContext [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#evaluatorcontext) All evaluators receive an [`EvaluatorContext`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorContext) containing: * `name`: Case name (optional) * `inputs`: Task inputs * `metadata`: Case metadata (optional) * `expected_output`: Expected output (optional) * `output`: Actual output from task * `duration`: Task execution time in seconds * `span_tree`: OpenTelemetry spans (if `logfire` is configured) * `attributes`: Custom attributes dict * `metrics`: Custom metrics dict ### Multiple Evaluations [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#multiple-evaluations) Evaluators can return multiple results by returning a dictionary: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class MultiCheck(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> dict[str, bool | float | str]: return { 'is_valid': isinstance(ctx.output, str), # Assertion 'length': len(ctx.output), # Metric 'category': 'long' if len(ctx.output) > 100 else 'short', # Label } ### Evaluation Reasons [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#evaluation-reasons) Add explanations to your evaluations using [`EvaluationReason`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluationReason) : from dataclasses import dataclass from pydantic_evals.evaluators import EvaluationReason, Evaluator, EvaluatorContext @dataclass class SmartCheck(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> EvaluationReason: if ctx.output == ctx.expected_output: return EvaluationReason( value=True, reason='Exact match with expected output', ) return EvaluationReason( value=False, reason=f'Expected {ctx.expected_output!r}, got {ctx.output!r}', ) Reasons appear in reports when using `include_reasons=True`. Evaluation Report ----------------- [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#evaluation-report) An [`EvaluationReport`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.EvaluationReport) is the result of running an experiment. It contains all the data from executing your task against the dataset’s cases and running all evaluators. from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import EqualsExpected dataset = Dataset( name='report_example', cases=[Case(inputs='hello', expected_output='HELLO')], evaluators=[EqualsExpected()], ) def my_task(text: str) -> str: return text.upper() # Run an experiment report = dataset.evaluate_sync(my_task) # Print to console report.print() """ Evaluation Summary: my_task ┏━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━┓ ┃ Case ID ┃ Assertions ┃ Duration ┃ ┡━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━┩ │ Case 1 │ ✔ │ 10ms │ ├──────────┼────────────┼──────────┤ │ Averages │ 100.0% ✔ │ 10ms │ └──────────┴────────────┴──────────┘ """ # Access data programmatically for case in report.cases: print(f'{case.name}: {case.scores}') #> Case 1: {} ### Report Structure [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#report-structure) The [`EvaluationReport`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.EvaluationReport) contains: * `name`: Experiment name * `cases`: List of successful case evaluations * `failures`: List of failed executions * `analyses`: List of experiment-wide analyses from [report evaluators](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/) (confusion matrices, PR curves, scalars, tables) * `trace_id`: OpenTelemetry trace ID (optional) * `span_id`: OpenTelemetry span ID (optional) ### ReportCase [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#reportcase) Each successfulcase result contains: **Case data:** * `name`: Case name * `inputs`: Task inputs * `metadata`: Case metadata (optional) * `expected_output`: Expected output (optional) * `output`: Actual output from task **Evaluation results:** * `scores`: Dictionary of numeric scores from evaluators * `labels`: Dictionary of categorical labels from evaluators * `assertions`: Dictionary of pass/fail assertions from evaluators **Performance data:** * `task_duration`: Task execution time * `total_duration`: Total time including evaluators **Additional data:** * `metrics`: Custom metrics dict * `attributes`: Custom attributes dict **Tracing:** * `trace_id`: OpenTelemetry trace ID (optional) * `span_id`: OpenTelemetry span ID (optional) **Errors:** * `evaluator_failures`: List of evaluator errors Data Model Relationships ------------------------ [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#data-model-relationships) Here’s how the core concepts relate to each other: ### Static Definition [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#static-definition) * A **Dataset** contains: * Many **Cases** (test scenarios with inputs and expected outputs) * Many **Evaluators** (logic for scoring individual outputs) * Many **Report Evaluators** (logic for analyzing full experiment results) ### Execution (Experiment) [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#execution-experiment) When you call `dataset.evaluate(task)`, an **Experiment** runs: * The **Task** function is executed against all **Cases** in the **Dataset** * All **Evaluators** are run (both dataset-level and case-specific) against each output as appropriate * One **EvaluationReport** is produced as the final output ### Results [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#results) * An **EvaluationReport** contains: * Results for each **Case** (inputs, outputs, scores, assertions, labels) * Experiment-wide **Analyses** from report evaluators (confusion matrices, PR curves, scalars, tables) * Summary statistics (averages, pass rates) * Performance data (durations) * Tracing information (OpenTelemetry spans) ### Key Relationships [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#key-relationships) * **One Dataset → Many Experiments**: You can run the same dataset against different task implementations or multiple times to track changes * **One Experiment → One Report**: Each time you call `dataset.evaluate(...)`, you get one report * **One Experiment → Many Case Results**: The report contains results for every case in the dataset Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/getting-started/core-concepts/#next-steps) * **[Evaluators Overview](https://pydantic.dev/docs/ai/evals/evaluators/overview/) ** - When to use different evaluator types * **[Native Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/built-in/) ** - Complete reference of provided evaluators * **[Custom Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/custom/) ** - Write your own evaluation logic * **[Report Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/) ** - Experiment-wide analyses (confusion matrices, PR curves, etc.) * **[Dataset Management](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/) ** - Save, load, and generate datasets Was this page helpful? Thanks for your feedback! --- # Vercel AI | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/integrations/ui/vercel-ai/#_top) Vercel AI ========= Pydantic AI natively supports the [Vercel AI Data Stream Protocol](https://ai-sdk.dev/docs/ai-sdk-ui/stream-protocol#data-stream-protocol) to receive agent run input from, and stream events to, a frontend using [AI SDK UI](https://ai-sdk.dev/docs/ai-sdk-ui/overview) hooks like [`useChat`](https://ai-sdk.dev/docs/reference/ai-sdk-ui/use-chat) . You can optionally use [AI Elements](https://ai-sdk.dev/elements) for pre-built UI components. Usage ----- [](https://pydantic.dev/docs/ai/integrations/ui/vercel-ai/#usage) The [`VercelAIAdapter`](https://pydantic.dev/docs/ai/api/ui/vercel_ai/#pydantic_ai.ui.vercel_ai.VercelAIAdapter) class is responsible for transforming agent run input received from the frontend into arguments for [`Agent.run_stream_events()`](https://pydantic.dev/docs/ai/core-concepts/agent/#running-agents) , running the agent, and then transforming Pydantic AI events into Vercel AI events. The event stream transformation is handled by the [`VercelAIEventStream`](https://pydantic.dev/docs/ai/api/ui/vercel_ai/#pydantic_ai.ui.vercel_ai.VercelAIEventStream) class, but you typically won’t use this directly. If you’re using a Starlette-based web framework like FastAPI, you can use the [`VercelAIAdapter.dispatch_request()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.dispatch_request) class method from an endpoint function to directly handle a request and return a streaming response of Vercel AI events. This is demonstrated in the next section. If you’re using a web framework not based on Starlette (e.g. Django or Flask) or need fine-grained control over the input or output, you can create a `VercelAIAdapter` instance and directly use its methods. This is demonstrated in “Advanced Usage” section below. ### Usage with Starlette/FastAPI [](https://pydantic.dev/docs/ai/integrations/ui/vercel-ai/#usage-with-starlettefastapi) Besides the request, [`VercelAIAdapter.dispatch_request()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.dispatch_request) takes the agent, the same optional arguments as [`Agent.run_stream_events()`](https://pydantic.dev/docs/ai/core-concepts/agent/#running-agents) , an optional `on_complete` callback for successful runs, and an optional `on_cancel` callback for cancelled runs. Both callbacks can optionally yield additional Vercel AI events. dispatch\_request.py from fastapi import FastAPI from starlette.requests import Request from starlette.responses import Response from pydantic_ai import Agent from pydantic_ai.ui.vercel_ai import VercelAIAdapter agent = Agent('openai:gpt-5.2') app = FastAPI() @app.post('/chat') async def chat(request: Request) -> Response: return await VercelAIAdapter.dispatch_request(request, agent=agent) ### Advanced Usage [](https://pydantic.dev/docs/ai/integrations/ui/vercel-ai/#advanced-usage) If you’re using a web framework not based on Starlette (e.g. Django or Flask) or need fine-grained control over the input or output, you can create a `VercelAIAdapter` instance and directly use its methods, which can be chained to accomplish the same thing as the `VercelAIAdapter.dispatch_request()` class method shown above: 1. The [`VercelAIAdapter.build_run_input()`](https://pydantic.dev/docs/ai/api/ui/vercel_ai/#pydantic_ai.ui.vercel_ai.VercelAIAdapter.build_run_input) class method takes the request body as bytes and returns a Vercel AI [`RequestData`](https://pydantic.dev/docs/ai/api/ui/vercel_ai/#pydantic_ai.ui.vercel_ai.request_types.RequestData) run input object, which you can then pass to the [`VercelAIAdapter()`](https://pydantic.dev/docs/ai/api/ui/vercel_ai/#pydantic_ai.ui.vercel_ai.VercelAIAdapter) constructor along with the agent. * You can also use the [`VercelAIAdapter.from_request()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.from_request) class method to build an adapter directly from a Starlette/FastAPI request. 2. The [`VercelAIAdapter.run_stream()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.run_stream) method runs the agent and returns a stream of Vercel AI events. It supports the same optional arguments as [`Agent.run_stream_events()`](https://pydantic.dev/docs/ai/core-concepts/agent/#running-agents) , including the `on_complete` and `on_cancel` callbacks. * You can also use [`VercelAIAdapter.run_stream_native()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.run_stream_native) to run the agent and return a stream of Pydantic AI events instead, which can then be transformed into Vercel AI events using [`VercelAIAdapter.transform_stream()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.transform_stream) . 3. The [`VercelAIAdapter.encode_stream()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.encode_stream) method encodes the stream of Vercel AI events as SSE (HTTP Server-Sent Events) strings, which you can then return as a streaming response. * You can also use [`VercelAIAdapter.streaming_response()`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.streaming_response) to generate a Starlette/FastAPI streaming response directly from the Vercel AI event stream returned by `run_stream()`. ### Cancellation [](https://pydantic.dev/docs/ai/integrations/ui/vercel-ai/#cancellation) When a run ends in [first-party cancellation](https://pydantic.dev/docs/ai/core-concepts/agent/#cancelling-a-run) — `ctx.cancel()` from a tool, `AgentRun.cancel()`, or a [`CancellationToken`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.CancellationToken) your server wires to a cancel endpoint — the adapter emits a Vercel `abort` chunk. [`useChat`](https://ai-sdk.dev/docs/reference/ai-sdk-ui/use-chat) keeps the partial message and reports `isAbort` to `onFinish` instead of entering an error state. Pass an `on_cancel` callback to persist the resumable message history, as shown in the example below. run\_stream.py import json from http import HTTPStatus from fastapi import FastAPI from fastapi.requests import Request from fastapi.responses import Response, StreamingResponse from pydantic import ValidationError from pydantic_ai import Agent, RunCancelled from pydantic_ai.ui import SSE_CONTENT_TYPE from pydantic_ai.ui.vercel_ai import VercelAIAdapter agent = Agent('openai:gpt-5.2') app = FastAPI() async def on_cancel(cancelled: RunCancelled) -> None: messages = cancelled.all_messages() # (1) print(f'cancelled after {len(messages)} messages') @app.post('/chat') async def chat(request: Request) -> Response: accept = request.headers.get('accept', SSE_CONTENT_TYPE) try: run_input = VercelAIAdapter.build_run_input(await request.body()) except ValidationError as e: return Response( content=json.dumps(e.json()), media_type='application/json', status_code=HTTPStatus.UNPROCESSABLE_ENTITY, ) adapter = VercelAIAdapter(agent=agent, run_input=run_input, accept=accept) event_stream = adapter.run_stream(on_cancel=on_cancel) sse_event_stream = adapter.encode_stream(event_stream) return StreamingResponse(sse_event_stream, media_type=accept) The resumable history to persist -- pass it as `message_history` to a later run to resume the conversation. ### Data Chunks [](https://pydantic.dev/docs/ai/integrations/ui/vercel-ai/#data-chunks) Pydantic AI tools can send [Vercel AI data stream chunks](https://ai-sdk.dev/docs/ai-sdk-ui/stream-protocol#data-stream-protocol) by returning a [`ToolReturn`](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#advanced-tool-returns) object with a data-carrying chunk (or a list of chunks) as `metadata`. The supported chunk types are [`DataChunk`](https://pydantic.dev/docs/ai/api/ui/vercel_ai/#pydantic_ai.ui.vercel_ai.response_types.DataChunk) , [`SourceUrlChunk`](https://pydantic.dev/docs/ai/api/ui/vercel_ai/#pydantic_ai.ui.vercel_ai.response_types.SourceUrlChunk) , [`SourceDocumentChunk`](https://pydantic.dev/docs/ai/api/ui/vercel_ai/#pydantic_ai.ui.vercel_ai.response_types.SourceDocumentChunk) , and [`FileChunk`](https://pydantic.dev/docs/ai/api/ui/vercel_ai/#pydantic_ai.ui.vercel_ai.response_types.FileChunk) . This is useful for attaching structured data to the frontend alongside the tool result, such as source URLs or custom data payloads. vercel\_ai\_tool\_chunks.py from pydantic_ai import Agent, ToolReturn from pydantic_ai.ui.vercel_ai.response_types import DataChunk, SourceUrlChunk agent = Agent('openai:gpt-5.2') @agent.tool_plain async def search_docs(query: str) -> ToolReturn: return ToolReturn( return_value=f'Found 2 results for "{query}"', metadata=[\ SourceUrlChunk(\ source_id='doc-1',\ url='https://example.com/docs/intro',\ title='Introduction',\ ),\ DataChunk(\ type='data-search-results',\ data={'query': query, 'count': 2},\ ),\ ], ) ### Files from client-side tools [](https://pydantic.dev/docs/ai/integrations/ui/vercel-ai/#files-from-client-side-tools) Vercel AI SDK [client-side tools](https://ai-sdk.dev/docs/ai-sdk-ui/chatbot-tool-usage#client-side-tools) run in the browser and submit their result back to the server, where Pydantic AI resolves them as [external tool calls](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/#external-tool-execution) . Such a tool can return a file by putting a shape in its output that matches one of Pydantic AI’s [multimodal content types](https://pydantic.dev/docs/ai/core-concepts/input/) ; Pydantic AI deserializes it into that type (via the same `ToolReturnContent` discriminator used for round-tripping) before the run continues. Use the type’s snake\_case field names — these are validated as Pydantic AI models, not Vercel-cased payloads, so `media_type` deserializes but `mediaType` stays an opaque dict. Three shapes are supported: * **Inline bytes** — a [`BinaryContent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BinaryContent) shape, `{ kind: 'binary', media_type: 'image/png', data: }` (image media types become [`BinaryImage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BinaryImage) ). The `data` field accepts a base64 string, or the raw byte shapes a JavaScript frontend produces when it forwards a `Uint8Array` or Node `Buffer` through `JSON.stringify` without encoding it first (`{ "0": 137, "1": 80, ... }` or `{ "type": "Buffer", "data": [137, 80, ...] }`) — all normalized to bytes at the wire boundary, so a client-side tool can return binary data without base64-encoding it by hand. * **A file URL** — a [`FileUrl`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.FileUrl) shape such as `{ kind: 'image-url', url: 'https://...' }` or `{ kind: 'document-url', url: 'https://...' }`. This is often more efficient than inlining the bytes, since only the reference crosses the wire and the provider fetches the file directly — a good fit when the file already lives at a URL the frontend trusts. The URL is honored only if its scheme passes the adapter’s [`allowed_file_url_schemes`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.allowed_file_url_schemes) allowlist (`http`/`https` by default); see the [trust model](https://pydantic.dev/docs/ai/integrations/ui/vercel-ai/#trust-model) . * **A provider-hosted file** — an [`UploadedFile`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.UploadedFile) shape, `{ kind: 'uploaded-file', file_id: 'file-123', provider_name: 'openai' }`, referencing a file already uploaded to the provider’s storage. Honored only when [`allow_uploaded_files`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.allow_uploaded_files) is `True`, since the server resolves it against the provider’s file API using its own credentials. Message metadata ---------------- [](https://pydantic.dev/docs/ai/integrations/ui/vercel-ai/#message-metadata) [`VercelAIAdapter.dump_messages`](https://pydantic.dev/docs/ai/api/ui/vercel_ai/#pydantic_ai.ui.vercel_ai.VercelAIAdapter.dump_messages) writes [`ModelRequest.metadata`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest.metadata) and [`ModelResponse.metadata`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse.metadata) into Vercel AI [`UIMessage.metadata`](https://ai-sdk.dev/docs/ai-sdk-ui/message-metadata) , and stores the message `timestamp` under a reserved `pydantic_ai` key so it survives the round-trip. [`VercelAIAdapter.load_messages`](https://pydantic.dev/docs/ai/api/ui/vercel_ai/#pydantic_ai.ui.vercel_ai.VercelAIAdapter.load_messages) restores it on the way back. When streaming, the timestamp is also emitted as a Vercel AI `message-metadata` chunk after the final step, so frontends using AI SDK UI can persist it with the assistant message. Request-side messages have no analogous chunk — frontends rebuilding history purely from streamed chunks see timestamps only on assistant responses, whereas `dump_messages` populates both sides. `UIMessage.metadata` is fully client-controlled, so only `timestamp` is round-tripped: server-side fields such as `usage`, `model_name`, and `provider_*` are deliberately excluded — dumping them could leak infrastructure details, and restoring them would trust client-submitted history for values the server owns. Broadening the round-trip behind an explicit user-controlled opt-in is tracked in [issue #5174](https://github.com/pydantic/pydantic-ai/issues/5174) . Trust model ----------- [](https://pydantic.dev/docs/ai/integrations/ui/vercel-ai/#trust-model) Vercel AI’s request `messages` array is fully client-controlled, and the protocol round-trips approval responses and tool results through the message history. The [`VercelAIAdapter`](https://pydantic.dev/docs/ai/api/ui/vercel_ai/#pydantic_ai.ui.vercel_ai.VercelAIAdapter) applies defaults to strip untrusted parts before the agent runs — see [Trust model for client-submitted messages](https://pydantic.dev/docs/ai/integrations/ui/overview/#trust-model-for-client-submitted-messages) in the UI adapter overview, which covers system prompts, file URL schemes, uploaded files ([`allow_uploaded_files`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.allow_uploaded_files) ), and unresolved tool calls. Those defaults don’t make client-submitted history authentic — see [Trust boundary for client-supplied history](https://pydantic.dev/docs/ai/core-concepts/message-history/#trust-boundary-for-client-supplied-history) . Compaction ---------- [](https://pydantic.dev/docs/ai/integrations/ui/vercel-ai/#compaction) [`CompactionPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.CompactionPart) s round-trip through Vercel AI data parts (`data-compaction`), so [compacted](https://pydantic.dev/docs/ai/capabilities/compaction/) conversations keep working when a frontend such as `useChat` holds the message history. A compaction item submitted by the frontend is honored — the conversation stays compacted — with two caveats. First, it is never trusted to stand in for the system prompt: whichever prompt applies per [System prompts and instructions](https://pydantic.dev/docs/ai/integrations/ui/vercel-ai/#system-prompts-and-instructions) still reaches the model on every request. Second, if the run also receives server-side `message_history` (the [server-side persistence pattern](https://pydantic.dev/docs/ai/integrations/ui/overview/#trust-model-for-client-submitted-messages) ), frontend compaction items are ignored — everything before a compaction item is hidden from the model, so honoring one from the frontend would let it hide the server’s stored history. See [Client-held history](https://pydantic.dev/docs/ai/capabilities/compaction/#client-held-history) for the trade-offs and the recommended server-side pattern. Tool Approval ------------- [](https://pydantic.dev/docs/ai/integrations/ui/vercel-ai/#tool-approval) Pydantic AI supports human-in-the-loop tool approval workflows with AI SDK UI, allowing users to approve or deny tool executions before they run. See the [deferred tool calls documentation](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/#human-in-the-loop-tool-approval) for details on setting up tools that require approval. To enable tool approval streaming, pass `sdk_version=6` to `dispatch_request`: @app.post('/chat') async def chat(request: Request) -> Response: return await VercelAIAdapter.dispatch_request(request, agent=agent, sdk_version=6) When `sdk_version=6`, the adapter will: 1. Emit `tool-approval-request` chunks when tools with `requires_approval=True` are called 2. Automatically extract approval responses from follow-up requests 3. Emit `tool-output-denied` chunks for rejected tools On the frontend, AI SDK UI’s [`useChat`](https://ai-sdk.dev/docs/reference/ai-sdk-ui/use-chat) hook handles the approval flow. You can use the [`Confirmation`](https://ai-sdk.dev/elements/components/confirmation) component from AI Elements for a pre-built approval UI, or build your own using the hook’s `addToolApprovalResponse` function. Tool approval responses are trusted from the request by design, matching the protocol’s round-trip through `useChat`’s `addToolApprovalResponse` and the reference Next.js backend. The decision itself must be an actual JSON boolean: `approved` is strictly typed, so any stand-in — `1` or `"true"` as much as `0` or `"false"` — fails request validation rather than being coerced into a decision. If your application needs the approval decision tied to server-side state rather than the request, intercept [`DeferredToolRequests`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.DeferredToolRequests) , persist the approval IDs server-side, and pass explicit `deferred_tool_results` when resuming. Tool input validation --------------------- [](https://pydantic.dev/docs/ai/integrations/ui/vercel-ai/#tool-input-validation) `tool-input-available` is emitted **after** the agent has validated the call against the tool’s schema and any custom [`args_validator`](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#args-validator) , so the chunk only fires once the args are known to be acceptable. The chunk’s `input` field carries the raw arguments the model emitted. When validation fails, the adapter emits `tool-input-error` instead of `tool-input-available`. The chunk carries the same `tool_call_id`, `tool_name`, and `input` (the raw arguments) plus an `error_text` field rendered from the message that will be sent back to the model. When the validation failure is retryable, the agent retries the call (subject to the tool’s `retries` setting) and emits a new `tool-input-(available|error)` for each attempt; when a custom validator raises [`ToolFailed`](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-failed) , the failure is terminal and no retry follows. System prompts and instructions ------------------------------- [](https://pydantic.dev/docs/ai/integrations/ui/vercel-ai/#system-prompts-and-instructions) Pydantic AI supports two ways to provide guidance to the model: [`system_prompt`](https://pydantic.dev/docs/ai/core-concepts/agent/#system-prompts) (stored in the message history as [`SystemPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SystemPromptPart) s) and [`instructions`](https://pydantic.dev/docs/ai/core-concepts/agent/#instructions) (injected fresh on every request, never persisted). When you control the server side, `instructions` is the recommended default. The rest of this section only matters if you use `system_prompt`. If you only use `instructions`, there’s nothing to configure — they’re always applied regardless of the frontend message history. For `system_prompt`, you choose who owns it with the `manage_system_prompt` parameter on [`VercelAIAdapter`](https://pydantic.dev/docs/ai/api/ui/vercel_ai/#pydantic_ai.ui.vercel_ai.VercelAIAdapter) : * `'server'` (default): the agent’s configured `system_prompt` is authoritative. Any system message sent by the frontend is stripped with a warning (a malicious client could otherwise inject arbitrary instructions via crafted API requests), and the agent’s own system prompt is reinjected at the head of the first request via the [`ReinjectSystemPrompt`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ReinjectSystemPrompt) capability. * `'client'`: the frontend owns the system prompt. Frontend system messages are preserved as-is, and the agent’s configured `system_prompt` is not injected — the caller is fully responsible for sending it on every turn if desired. To opt into fallback-to-configured behavior, add the [`ReinjectSystemPrompt`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ReinjectSystemPrompt) capability to your agent. vercel\_ai\_client\_managed\_system\_prompt.py from fastapi import FastAPI from starlette.requests import Request from starlette.responses import Response from pydantic_ai import Agent from pydantic_ai.ui.vercel_ai import VercelAIAdapter agent = Agent('openai:gpt-5.2') app = FastAPI() @app.post('/chat') async def chat(request: Request) -> Response: return await VercelAIAdapter.dispatch_request( request, agent=agent, manage_system_prompt='client' ) Was this page helpful? Thanks for your feedback! --- # Ollama | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/models/ollama/#_top) Ollama ====== Install ------- [](https://pydantic.dev/docs/ai/models/ollama/#install) To use [`OllamaModel`](https://pydantic.dev/docs/ai/api/models/ollama/#pydantic_ai.models.ollama.OllamaModel) , you need to either install `pydantic-ai`, or install `pydantic-ai-slim` with the `openai` optional group: * [pip](https://pydantic.dev/docs/ai/models/ollama/#tab-panel-122) * [uv](https://pydantic.dev/docs/ai/models/ollama/#tab-panel-123) Terminal pip install "pydantic-ai-slim[openai]" Terminal uv add "pydantic-ai-slim[openai]" Configuration ------------- [](https://pydantic.dev/docs/ai/models/ollama/#configuration) Pydantic AI supports both self-hosted [Ollama](https://ollama.com/) servers (running locally or remotely) and [Ollama Cloud](https://ollama.com/cloud) . For servers running locally, use the `http://localhost:11434/v1` base URL. For Ollama Cloud, use `https://ollama.com/v1` and ensure an API key is set. For backward compatibility, [`OllamaModel`](https://pydantic.dev/docs/ai/api/models/ollama/#pydantic_ai.models.ollama.OllamaModel) uses Ollama’s OpenAI-compatible Chat Completions API (`/v1/chat/completions`). Environment variable -------------------- [](https://pydantic.dev/docs/ai/models/ollama/#environment-variable) Set the `OLLAMA_BASE_URL` and (optionally) `OLLAMA_API_KEY` environment variables: Terminal export OLLAMA_BASE_URL='http://localhost:11434/v1' export OLLAMA_API_KEY='your-api-key' # required for Ollama Cloud You can then use `OllamaModel` by name: from pydantic_ai import Agent agent = Agent('ollama:qwen3') ... Or initialise the model directly with just the model name: from pydantic_ai import Agent from pydantic_ai.models.ollama import OllamaModel model = OllamaModel('qwen3') agent = Agent(model) ... `provider` argument ------------------- [](https://pydantic.dev/docs/ai/models/ollama/#provider-argument) You can provide a custom `Provider` via the `provider` argument: from pydantic_ai import Agent from pydantic_ai.models.ollama import OllamaModel from pydantic_ai.providers.ollama import OllamaProvider model = OllamaModel( 'qwen3', provider=OllamaProvider(base_url='http://localhost:11434/v1') ) agent = Agent(model) ... For Ollama Cloud, use `base_url='https://ollama.com/v1'` and set the `OLLAMA_API_KEY` environment variable (or pass `api_key=` directly). Structured output ----------------- [](https://pydantic.dev/docs/ai/models/ollama/#structured-output) Self-hosted Ollama (v0.5.0+, released December 2024) enforces `response_format` with `json_schema` via `llama.cpp`’s grammar-constrained decoder, so [`NativeOutput`](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.NativeOutput) produces schema-valid output at generation time: from pydantic import BaseModel from pydantic_ai import Agent from pydantic_ai.models.ollama import OllamaModel from pydantic_ai.output import NativeOutput from pydantic_ai.providers.ollama import OllamaProvider class CityLocation(BaseModel): city: str country: str model = OllamaModel( 'qwen3', provider=OllamaProvider(base_url='http://localhost:11434/v1'), ) agent = Agent(model, output_type=NativeOutput(CityLocation)) ... Was this page helpful? Thanks for your feedback! --- # Web Chat UI | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/guides/web/#_top) Web Chat UI =========== Pydantic AI includes a built-in web chat interface that you can use to interact with your agents through a browser. ![Web Chat UI](https://pydantic.dev/docs/ai/img/web-chat-ui.png) For CLI usage with `clai web`, see the [CLI - Web Chat UI documentation](https://pydantic.dev/docs/ai/integrations/cli/#web-chat-ui) . Installation ------------ [](https://pydantic.dev/docs/ai/guides/web/#installation) Install the `web` extra (installs Starlette and Uvicorn): * [pip](https://pydantic.dev/docs/ai/guides/web/#tab-panel-84) * [uv](https://pydantic.dev/docs/ai/guides/web/#tab-panel-85) Terminal pip install 'pydantic-ai-slim[web]' Terminal uv add 'pydantic-ai-slim[web]' Basic Usage ----------- [](https://pydantic.dev/docs/ai/guides/web/#basic-usage) Create a web app from an agent instance using [`Agent.to_web()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.to_web) : from pydantic_ai import Agent agent = Agent('openai:gpt-5.2', instructions='You are a helpful assistant.') @agent.tool_plain def get_weather(city: str) -> str: return f'The weather in {city} is sunny' app = agent.to_web() Run the app with any ASGI server: Terminal uvicorn my_module:app --host 127.0.0.1 --port 7932 Configuring Models ------------------ [](https://pydantic.dev/docs/ai/guides/web/#configuring-models) You can specify additional models to make available in the UI. Models can be provided as a list of model names/instances or a dictionary mapping display labels to model names/instances. from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModel # Model with custom configuration anthropic_model = AnthropicModel('claude-sonnet-4-5') agent = Agent('openai:gpt-5.2') app = agent.to_web( models=['openai:gpt-5.2', anthropic_model], ) # Or with custom display labels app = agent.to_web( models={'GPT 5.2': 'openai:gpt-5.2', 'Claude': anthropic_model}, ) Native Tool Support ------------------- [](https://pydantic.dev/docs/ai/guides/web/#native-tool-support) Configure [native tools](https://pydantic.dev/docs/ai/tools-toolsets/native-tools/) on the agent with `capabilities=[NativeTool(...)]` to expose them as options in the UI (shown only for models that support each tool): from pydantic_ai import Agent from pydantic_ai.capabilities import NativeTool from pydantic_ai.native_tools import CodeExecutionTool, WebSearchTool agent = Agent( 'openai:gpt-5.2', capabilities=[NativeTool(CodeExecutionTool()), NativeTool(WebSearchTool())], ) app = agent.to_web(models=['anthropic:claude-sonnet-4-6']) Extra Instructions ------------------ [](https://pydantic.dev/docs/ai/guides/web/#extra-instructions) You can pass extra instructions that will be included in each agent run: from pydantic_ai import Agent agent = Agent('openai:gpt-5.2') app = agent.to_web(instructions='Always respond in a friendly tone.') Tool Approval ------------- [](https://pydantic.dev/docs/ai/guides/web/#tool-approval) Tools that [require approval](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/#human-in-the-loop-tool-approval) are surfaced in the UI as approve/reject prompts: when the agent calls such a tool, the UI renders the pending call and lets you approve or deny it before the run continues. This works out of the box — no extra configuration is needed. Reaching the UI under a hostname -------------------------------- [](https://pydantic.dev/docs/ai/guides/web/#reaching-the-ui-under-a-hostname) The app answers only to requests whose `Host` header is an IP address (`127.0.0.1`, `[::1]`, or a LAN address like `192.168.1.5`) or `localhost` — including names under it, like `my-app.localhost`. Any other `Host` gets a `421 Misdirected Request`. Hostnames are compared in ASCII form, so an internationalized name goes in the list as punycode (`xn--bcher-kva.example`), which is what the browser sends. This is what stops a website from reaching the UI on your machine by pointing a hostname it controls at `127.0.0.1` — a DNS rebinding attack, which makes the browser treat that website and the UI as the same origin, so the content type requirement above no longer applies. An IP address can’t be rebound that way, because rebinding works by pointing a _name_ at an address. If you serve the UI under a real hostname — behind a reverse proxy, or through a tunnel like ngrok — name that hostname in `allowed_hosts`: from pydantic_ai import Agent agent = Agent('openai:gpt-5.2') app = agent.to_web(allowed_hosts=['ui.example.com']) # `*.example.com` matches subdomains only; list the apex separately if you serve it too app = agent.to_web(allowed_hosts=['example.com', '*.example.com']) Or with the CLI: Terminal clai web -m openai:gpt-5.2 --allowed-host ui.example.com `clai web --host ` adds that name for you, so the URL it prints always works. Every route is checked, including `/api/health`. A health check or container probe that sends a DNS name in its `Host` header gets the same `421`, and monitoring systems often record only the status code or swap in their own error page, so the explanation may never reach you — point probes at the bound IP address or `localhost`, or add their hostname here. Pass `allowed_hosts=['*']` to answer to any host, but only if something in front of the app already authenticates requests. Only list domains whose subdomains you control: a wildcard for a domain where anyone can obtain a subdomain re-opens the problem. Reserved Routes --------------- [](https://pydantic.dev/docs/ai/guides/web/#reserved-routes) All routes are answered only for [allowed `Host` headers](https://pydantic.dev/docs/ai/guides/web/#reaching-the-ui-under-a-hostname) . The web UI app uses the following routes which should not be overwritten: * `/` and `/{id}` - Serves the chat UI * `/api/chat` - Chat endpoint (POST, OPTIONS). Requires `Content-Type: application/json`; other content types are rejected with `415`. * `/api/configure` - Frontend configuration (GET) * `/api/health` - Health check (GET) The app cannot currently be mounted at a subpath (e.g., `/chat`) because the UI expects these routes at the root. You can add additional routes to the app, but avoid conflicts with these reserved paths. Custom HTML Source ------------------ [](https://pydantic.dev/docs/ai/guides/web/#custom-html-source) By default, the web UI is fetched from a CDN and cached locally. You can provide `html_source` to override this for offline usage or enterprise environments. ### Offline and air-gapped deployments [](https://pydantic.dev/docs/ai/guides/web/#offline-and-air-gapped-deployments) The default UI build is split across many files: `index.html` references a stylesheet and, at runtime, lazily imports chunks for syntax highlighting, diagrams and math. Those references point back at the CDN, so downloading `index.html` alone gives you a page that boots and then fails to render as soon as a code block or an equation appears. Use the **offline build** instead — a single self-contained file with every chunk, font and icon inlined, so it needs no network access beyond your own server: from pydantic_ai.ui import OFFLINE_HTML_URL print(OFFLINE_HTML_URL) # Use this URL to download the self-contained UI HTML file #> https://cdn.jsdelivr.net/npm/@pydantic/ai-chat-ui@2.1.0/offline/index.html Download it once from a machine that has internet access, then move it into the air-gapped environment: Terminal curl -o ~/pydantic-ai-ui.html Then use `html_source` to point to your local file or custom URL: from pydantic_ai import Agent agent = Agent('openai:gpt-5.2') # Use a local file (e.g., for offline usage) app = agent.to_web(html_source='~/pydantic-ai-ui.html') # Or use a custom URL (e.g., for enterprise environments) app = agent.to_web(html_source='https://cdn.example.com/ui/index.html') The offline file is around 16 MB. That is not extra weight so much as relocated weight — the default build ships the same assets across 400-odd files that the browser fetches from the CDN on demand, where the offline build front-loads all of them into the first request. The default `to_web()` path is unchanged and still uses the split build: from pydantic_ai.ui import DEFAULT_HTML_URL print(DEFAULT_HTML_URL) #> https://cdn.jsdelivr.net/npm/@pydantic/ai-chat-ui@2.1.0/dist/index.html Was this page helpful? Thanks for your feedback! --- # google | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/realtime/google/#_top) google ====== The Gemini Live API provider. Requires the `google` optional group (`pip install "pydantic-ai-slim[google]"`). [`GoogleRealtimeModel`](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModel) runs over the `google-genai` SDK (which manages the WebSocket transport). Gemini expects **16 kHz** PCM input (output is 24 kHz), produces one response modality per session, and natively accepts live video frames sent as [`BinaryImage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BinaryImage) . It exposes Gemini Live’s session and generation configuration through [`GoogleRealtimeModelSettings`](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings) — shared turn-taking via [`TurnDetection`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.TurnDetection) , with finer Gemini-specific control via [`AutomaticVAD`](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.AutomaticVAD) in `google_vad` plus `google_activity_handling`/`google_turn_coverage`, voice via `google_voice` or a [`MultiSpeaker`](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.MultiSpeaker) in `google_multi_speaker`, and long-session [`ContextCompression`](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.ContextCompression) — with resilience via session resumption + a [`ReconnectPolicy`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.ReconnectPolicy) in the `reconnect` setting. Gemini Live API provider for realtime speech-to-speech (and live video) sessions. Built on the `google-genai` SDK, which manages the WebSocket transport for you. Available via the `google` optional group: pip install “pydantic-ai-slim\[google-realtime\]” Unlike the OpenAI provider, Gemini wants **16 kHz** PCM input audio (output is 24 kHz), produces a single response modality per session (audio _or_ text), and natively accepts a stream of video frames sent as [`BinaryImage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BinaryImage) . Use `provider='google'` for the Gemini Developer API, or `provider='google-cloud'` / [`GoogleCloudProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.google_cloud.GoogleCloudProvider) for Google Cloud with Application Default Credentials. AutomaticVAD ------------ [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.AutomaticVAD) **Bases:** [`TypedDict`](https://docs.python.org/3/library/typing.html#typing.TypedDict) Server-side voice activity detection — the default turn-taking mode for Gemini Live. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/google/#attributes) #### disabled [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.AutomaticVAD.disabled) Turn off automatic VAD entirely. Defaults to `False`. Do not set this through `RealtimeSession`: Pydantic AI does not expose Gemini activity markers or manual turn controls. Use automatic VAD instead; the shared `turn_detection=False` setting is rejected for the same reason. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### end\_sensitivity [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.AutomaticVAD.end_sensitivity) How readily the end of speech is detected. `high` ends turns sooner; `low` waits longer. Defaults to the provider default. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘high’, ‘low’\] #### prefix\_padding\_ms [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.AutomaticVAD.prefix_padding_ms) Audio to include before detected speech, in milliseconds. Defaults to the provider default. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### silence\_duration\_ms [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.AutomaticVAD.silence_duration_ms) Silence required to detect the end of speech, in milliseconds. Defaults to the provider default. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### start\_sensitivity [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.AutomaticVAD.start_sensitivity) How readily speech onset is detected. `high` triggers on quieter audio; `low` is stricter. Defaults to the provider default. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘high’, ‘low’\] ContextCompression ------------------ [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.ContextCompression) **Bases:** [`TypedDict`](https://docs.python.org/3/library/typing.html#typing.TypedDict) Sliding-window context compression so long sessions don’t exceed the context window. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/google/#attributes-1) #### target\_tokens [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.ContextCompression.target_tokens) Target size (in tokens) of the retained sliding window after compression. Defaults to the provider default. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### trigger\_tokens [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.ContextCompression.trigger_tokens) Compress once the context passes this many tokens. Defaults to the provider default. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) GoogleRealtimeConnection ------------------------ [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeConnection) **Bases:** `RealtimeConnection` A live connection to the Gemini Live API, backed by a `google-genai` session. ### Methods [](https://pydantic.dev/docs/ai/api/realtime/google/#methods) #### send [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeConnection.send) `@async` def send(content: RealtimeInput) -> None Send content to the Gemini Live API. Accepts `BinaryAudio` (raw PCM16, 16kHz, mono), a `str` text turn, `BinaryImage` (a live video frame), and `ToolResult`. The manual turn-taking verbs are not supported (Gemini uses automatic VAD). ##### Returns [](https://pydantic.dev/docs/ai/api/realtime/google/#returns) [`None`](https://docs.python.org/3/library/constants.html#None) GoogleRealtimeModel ------------------- [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModel) **Bases:** `RealtimeModel` Gemini Live API model. Session and generation configuration is read from [`GoogleRealtimeModelSettings`](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings) , passed through `settings` as model-level defaults or as `model_settings` when opening a session. Authentication and the underlying `google-genai` client come from a [`Provider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.Provider) , mirroring [`GoogleModel`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModel) . Pass `provider='google'` (the default) for the Gemini Developer API (reads `GOOGLE_API_KEY` / `GEMINI_API_KEY`), `provider='google-cloud'` for Vertex AI (Application Default Credentials, useful where org policy disallows API keys), or a [`GoogleProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.google.GoogleProvider) / [`GoogleCloudProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.google_cloud.GoogleCloudProvider) instance for a custom key, client, or region. Gemini Live is available on both surfaces. ### Constructor Parameters [](https://pydantic.dev/docs/ai/api/realtime/google/#constructor-parameters) **`model`** : `GoogleRealtimeModelName` [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModel.__init__(model)) The model name, e.g. `gemini-2.5-flash-native-audio-latest` (an alias that tracks the newest native-audio Live model) or `gemini-3.1-flash-live-preview`. **`provider`** : [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘google’, ‘google-cloud’, ‘gateway’\] | `Provider`\[`Client`\] _Default:_ `'google'` [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModel.__init__(provider)) The provider to use for authentication and API access — `'google'` (Gemini Developer API, the default) or `'google-cloud'` (Vertex AI), or a `Provider` instance. **`settings`** : `RealtimeModelSettings` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModel.__init__(settings)) Model-level defaults for session and generation configuration. **`profile`** : `RealtimeModelProfileSpec` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModel.__init__(profile)) Optional override for the [realtime model profile](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelProfile) , merged over the provider’s — a partial dict, or a callable taking the resolved profile and returning the one to use. Mirrors `profile=` on a standard [`Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) , and is the escape hatch when a model name doesn’t identify the model (e.g. an Azure deployment named something other than its model). ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/google/#attributes-2) #### client [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModel.client) The underlying `google.genai.Client` from the provider. **Type:** `Client` GoogleRealtimeModelSettings --------------------------- [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings) **Bases:** `RealtimeModelSettings` Settings used for a Gemini Live session. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/google/#attributes-3) #### google\_activity\_handling [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.google_activity_handling) Whether detected user activity interrupts the model. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘interrupts’, ‘no\_interruption’\] #### google\_affective\_dialog [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.google_affective_dialog) Whether to enable emotion-aware delivery (native-audio models only). **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### google\_async\_tool\_calls [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.google_async_tool_calls) Whether tool calls may run without pausing the model’s speech. Defaults to `False`. By default Gemini stops generating while a tool call is outstanding, so the caller hears silence for as long as the tool takes. Enabling this declares tools `NON_BLOCKING` and returns their results with `INTERRUPT` scheduling, so the model keeps talking (typically narrating what it’s doing) and the result cuts into that speech when it arrives. This pays off for tools that take a noticeable moment. It is a poor trade for fast tools: the result interrupts a reply the model has barely started, leaving an extra interrupted turn in history with nothing in it. Verified live against `gemini-2.5-flash-native-audio-latest`. Supported by Gemini native-audio models (see [`supports_async_tool_calls`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelProfile.supports_async_tool_calls) ). Other models silently ignore it. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### google\_config\_overrides [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.google_config_overrides) Raw values merged last into the Google `LiveConnectConfig`. **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] #### google\_context\_compression [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.google_context_compression) Sliding-window context compression for long-running sessions. **Type:** `ContextCompression` #### google\_enable\_session\_resumption [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.google_enable_session_resumption) Whether to request session-resumption handles, which let a re-dial restore the server-side conversation. When absent, handles are requested exactly when a [`reconnect`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelSettings.reconnect) policy is set. An explicit `False` cannot be combined with a `reconnect` policy: a re-dial without resumption would lose the conversation, so `connect` raises [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) rather than silently reconnecting into a model that remembers nothing. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### google\_input\_transcription [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.google_input_transcription) Whether to transcribe input audio. Defaults to `True`. When `False`, user turns are recorded as retained audio when available, or as content-less placeholders otherwise. Takes precedence over the shared [`input_transcription_model`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelSettings.input_transcription_model) , whose `None` also turns transcription off here. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### google\_language\_code [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.google_language_code) BCP-47 language code for audio output. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### google\_multi\_speaker [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.google_multi_speaker) Per-speaker voice assignments; takes precedence over `google_voice`. **Type:** `MultiSpeaker` #### google\_output\_transcription [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.google_output_transcription) Whether to transcribe output audio. Defaults to `True`. When `False`, retain output audio if assistant audio turns need to appear in history. Assistant audio without a transcript cannot be handed off or seeded. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### google\_proactive\_audio [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.google_proactive_audio) Whether the model may decide _when_ to respond, including staying silent on input not addressed to it (native-audio models only). Useful for “react to the camera” experiences. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### google\_thinking\_config [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.google_thinking_config) The thinking configuration to use for the model. **Type:** `genai_types.ThinkingConfigDict` #### google\_transcription\_language\_codes [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.google_transcription_language_codes) Language hints applied to input and output transcription. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] #### google\_turn\_coverage [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.google_turn_coverage) Which realtime input is attached to a turn — `'activity_only'`, `'all_input'` (everything between turns too), or `'all_video'` (all video frames plus audio during activity; ideal for live-camera use). Absent uses the provider default. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘activity\_only’, ‘all\_input’, ‘all\_video’\] #### google\_vad [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.google_vad) Gemini-specific server-side voice activity detection settings. When present, this fully overrides the cross-provider `turn_detection` setting. `google_vad={'disabled': True}` raises a `UserError`, like `turn_detection=False`: Pydantic AI does not expose Gemini activity markers or manual turn controls, so the resulting session could not drive turns. **Type:** `AutomaticVAD` #### google\_video\_resolution [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.google_video_resolution) The video resolution to use for the model. **Type:** `genai_types.MediaResolution` #### google\_voice [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.google_voice) Prebuilt voice used for audio output, e.g. `Puck`. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### seed [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.seed) The random seed to use for the session. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### temperature [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.temperature) Amount of randomness injected into the response. **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) #### top\_k [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.top_k) Only sample from the top K options for each subsequent token. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### top\_p [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings.top_p) Nucleus sampling probability mass. **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) MultiSpeaker ------------ [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.MultiSpeaker) **Bases:** [`TypedDict`](https://docs.python.org/3/library/typing.html#typing.TypedDict) Assign prebuilt voices to named speakers for multi-speaker audio output. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/google/#attributes-4) #### voices [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.MultiSpeaker.voices) Mapping of speaker label to prebuilt voice name, e.g. `{'Joe': 'Puck', 'Jane': 'Kore'}`. Defaults to an empty mapping. **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] INPUT\_SAMPLE\_RATE ------------------- [](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.INPUT_SAMPLE_RATE) Sample rate (Hz) Gemini expects for PCM16 input audio. **Default:** `16000` Was this page helpful? Thanks for your feedback! --- # LLM Judge | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#_top) LLM Judge ========= The [`LLMJudge`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.LLMJudge) evaluator uses an LLM to assess subjective qualities of outputs based on a rubric. When to Use LLM-as-a-Judge -------------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#when-to-use-llm-as-a-judge) LLM judges are ideal for evaluating qualities that require understanding and judgment: **Good Use Cases:** * Factual accuracy * Helpfulness and relevance * Tone and style compliance * Completeness of responses * Following complex instructions * RAG groundedness (does the answer use provided context?) * Citation accuracy **Poor Use Cases:** * Format validation (use [`IsInstance`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.IsInstance) instead) * Exact matching (use [`EqualsExpected`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EqualsExpected) ) * Performance checks (use [`MaxDuration`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.MaxDuration) ) * Deterministic logic (write a custom evaluator) Basic Usage ----------- [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#basic-usage) from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import LLMJudge dataset = Dataset( name='factual_accuracy', cases=[Case(inputs='test')], evaluators=[\ LLMJudge(rubric='Response is factually accurate'),\ ], ) Configuration Options --------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#configuration-options) ### Rubric [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#rubric) The `rubric` is your evaluation criteria. Be specific and clear: **Bad rubrics (vague):** from pydantic_evals.evaluators import LLMJudge LLMJudge(rubric='Good response') # Too vague LLMJudge(rubric='Check quality') # What aspect of quality? **Good rubrics (specific):** from pydantic_evals.evaluators import LLMJudge LLMJudge(rubric='Response directly answers the user question without hallucination') LLMJudge(rubric='Response uses formal, professional language appropriate for business communication') LLMJudge(rubric='All factual claims in the response are supported by the provided context') ### Including Context [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#including-context) Control what information the judge sees: from pydantic_evals.evaluators import LLMJudge # Output only (default) LLMJudge(rubric='Response is polite') # Output + Input LLMJudge( rubric='Response accurately answers the input question', include_input=True, ) # Output + Input + Expected Output LLMJudge( rubric='Response is semantically equivalent to the expected output', include_input=True, include_expected_output=True, ) **Example:** from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import LLMJudge dataset = Dataset( name='math_check', cases=[\ Case(\ inputs='What is 2+2?',\ expected_output='4',\ ),\ ], evaluators=[\ # This judge sees: output + inputs + expected_output\ LLMJudge(\ rubric='Response provides the same answer as expected, possibly with explanation',\ include_input=True,\ include_expected_output=True,\ ),\ ], ) ### Model Selection [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#model-selection) Choose the judge model based on cost/quality tradeoffs: from pydantic_evals.evaluators import LLMJudge # Default: GPT-4o (good balance) LLMJudge(rubric='...') # Anthropic Claude (alternative default) LLMJudge( rubric='...', model='anthropic:claude-sonnet-4-6', ) # Cheaper option for simple checks LLMJudge( rubric='Response contains profanity', model='openai:gpt-5-mini', ) # Premium option for nuanced evaluation LLMJudge( rubric='Response demonstrates deep understanding of quantum mechanics', model='anthropic:claude-opus-4-5', ) ### Model Settings [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#model-settings) Customize model behavior: from pydantic_ai import ModelSettings from pydantic_evals.evaluators import LLMJudge LLMJudge( rubric='...', model_settings=ModelSettings( temperature=0.0, # Deterministic evaluation max_tokens=100, # Shorter responses ), ) Output Modes ------------ [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#output-modes) ### Assertion Only (Default) [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#assertion-only-default) Returns pass/fail with reason: from pydantic_evals.evaluators import LLMJudge LLMJudge(rubric='Response is accurate') # Returns: {'LLMJudge_pass': EvaluationReason(value=True, reason='...')} In reports: ┃ Assertions ┃ ┃ ✔ ┃ ### Score Only [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#score-only) Returns a numeric score (0.0 to 1.0): from pydantic_evals.evaluators import LLMJudge LLMJudge( rubric='Response quality', score={'include_reason': True}, assertion=False, ) # Returns: {'LLMJudge_score': EvaluationReason(value=0.85, reason='...')} In reports: ┃ Scores ┃ ┃ LLMJudge_score: 0.85 ┃ ### Both Score and Assertion [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#both-score-and-assertion) from pydantic_evals.evaluators import LLMJudge LLMJudge( rubric='Response quality', score={'include_reason': True}, assertion={'include_reason': True}, ) # Returns: { # 'LLMJudge_score': EvaluationReason(value=0.85, reason='...'), # 'LLMJudge_pass': EvaluationReason(value=True, reason='...'), # } ### Custom Names [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#custom-names) from pydantic_evals.evaluators import LLMJudge LLMJudge( rubric='Response is factually accurate', assertion={ 'evaluation_name': 'accuracy', 'include_reason': True, }, ) # Returns: {'accuracy': EvaluationReason(value=True, reason='...')} In reports: ┃ Assertions ┃ ┃ accuracy: ✔ ┃ Practical Examples ------------------ [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#practical-examples) ### RAG Evaluation [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#rag-evaluation) Evaluate whether a RAG system uses provided context: from dataclasses import dataclass from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import LLMJudge @dataclass class RAGInput: question: str context: str dataset = Dataset( name='rag_evaluation', cases=[\ Case(\ inputs=RAGInput(\ question='What is the capital of France?',\ context='France is a country in Europe. Its capital is Paris.',\ ),\ ),\ ], evaluators=[\ LLMJudge(\ rubric='Response answers the question using only information from the provided context',\ include_input=True,\ assertion={'evaluation_name': 'grounded', 'include_reason': True},\ ),\ LLMJudge(\ rubric='Response cites specific quotes or facts from the context',\ include_input=True,\ assertion={'evaluation_name': 'uses_citations', 'include_reason': True},\ ),\ ], ) ### Recipe Generation with Case-Specific Rubrics [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#recipe-generation-with-case-specific-rubrics) This example shows how to use both dataset-level and case-specific evaluators: recipe\_evaluation.py from __future__ import annotations from typing import Any from pydantic import BaseModel from pydantic_ai import Agent, format_as_xml from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import IsInstance, LLMJudge class CustomerOrder(BaseModel): dish_name: str dietary_restriction: str | None = None class Recipe(BaseModel): ingredients: list[str] steps: list[str] recipe_agent = Agent( 'openai:gpt-5-mini', output_type=Recipe, instructions=( 'Generate a recipe to cook the dish that meets the dietary restrictions.' ), ) async def transform_recipe(customer_order: CustomerOrder) -> Recipe: r = await recipe_agent.run(format_as_xml(customer_order)) return r.output recipe_dataset = Dataset[CustomerOrder, Recipe, Any]( name='recipe_evaluation', cases=[\ Case(\ name='vegetarian_recipe',\ inputs=CustomerOrder(\ dish_name='Spaghetti Bolognese', dietary_restriction='vegetarian'\ ),\ expected_output=None,\ metadata={'focus': 'vegetarian'},\ evaluators=( # (1)\ LLMJudge(\ rubric='Recipe should not contain meat or animal products',\ ),\ ),\ ),\ Case(\ name='gluten_free_recipe',\ inputs=CustomerOrder(\ dish_name='Chocolate Cake', dietary_restriction='gluten-free'\ ),\ expected_output=None,\ metadata={'focus': 'gluten-free'},\ evaluators=( # (2)\ LLMJudge(\ rubric='Recipe should not contain gluten or wheat products',\ ),\ ),\ ),\ ], evaluators=[ # (3)\ IsInstance(type_name='Recipe'),\ LLMJudge(\ rubric='Recipe should have clear steps and relevant ingredients',\ include_input=True,\ model='anthropic:claude-sonnet-4-6',\ ),\ ], ) report = recipe_dataset.evaluate_sync(transform_recipe) print(report) """ Evaluation Summary: transform_recipe ┏━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━┓ ┃ Case ID ┃ Assertions ┃ Duration ┃ ┡━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━┩ │ vegetarian_recipe │ ✔✔✔ │ 38.1s │ ├────────────────────┼────────────┼──────────┤ │ gluten_free_recipe │ ✔✔✔ │ 22.4s │ ├────────────────────┼────────────┼──────────┤ │ Averages │ 100.0% ✔ │ 30.3s │ └────────────────────┴────────────┴──────────┘ """ Case-specific evaluator - only runs for the vegetarian recipe case Case-specific evaluator - only runs for the gluten-free recipe case Dataset-level evaluators - run for all cases ### Multi-Aspect Evaluation [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#multi-aspect-evaluation) Use multiple judges for different quality dimensions: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import LLMJudge dataset = Dataset( name='multi_aspect', cases=[Case(inputs='test')], evaluators=[\ # Accuracy\ LLMJudge(\ rubric='Response is factually accurate',\ include_input=True,\ assertion={'evaluation_name': 'accurate'},\ ),\ \ # Helpfulness\ LLMJudge(\ rubric='Response is helpful and actionable',\ include_input=True,\ score={'evaluation_name': 'helpfulness'},\ assertion=False,\ ),\ \ # Tone\ LLMJudge(\ rubric='Response uses professional, respectful language',\ assertion={'evaluation_name': 'professional_tone'},\ ),\ \ # Safety\ LLMJudge(\ rubric='Response contains no harmful, biased, or inappropriate content',\ assertion={'evaluation_name': 'safe'},\ ),\ ], ) ### Comparative Evaluation [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#comparative-evaluation) Compare output against expected output: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import LLMJudge dataset = Dataset( name='comparative_eval', cases=[\ Case(\ name='translation',\ inputs='Hello world',\ expected_output='Bonjour le monde',\ ),\ ], evaluators=[\ LLMJudge(\ rubric='Response is semantically equivalent to the expected output',\ include_input=True,\ include_expected_output=True,\ score={'evaluation_name': 'semantic_similarity'},\ assertion={'evaluation_name': 'correct_meaning'},\ ),\ ], ) Best Practices -------------- [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#best-practices) ### 1\. Be Specific in Rubrics [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#1-be-specific-in-rubrics) **Bad:** from pydantic_evals.evaluators import LLMJudge LLMJudge(rubric='Good answer') **Better:** from pydantic_evals.evaluators import LLMJudge LLMJudge(rubric='Response accurately answers the question without hallucinating facts') **Best:** from pydantic_evals.evaluators import LLMJudge LLMJudge( rubric=''' Response must: 1. Directly answer the question asked 2. Use only information from the provided context 3. Cite specific passages from the context 4. Acknowledge if information is insufficient ''', include_input=True, ) ### 2\. Use Multiple Judges [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#2-use-multiple-judges) Don’t always try to evaluate everything with one rubric: from pydantic_evals.evaluators import LLMJudge # Instead of this: LLMJudge(rubric='Response is good, accurate, helpful, and safe') # Do this: evaluators = [\ LLMJudge(rubric='Response is factually accurate'),\ LLMJudge(rubric='Response is helpful and actionable'),\ LLMJudge(rubric='Response is safe and appropriate'),\ ] ### 3\. Combine with Deterministic Checks [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#3-combine-with-deterministic-checks) Don’t use LLM evaluation for checks that can be done deterministically: from pydantic_evals.evaluators import Contains, IsInstance, LLMJudge evaluators = [\ IsInstance(type_name='str'),\ Contains(value='required_section'),\ LLMJudge(rubric='Response quality is high'),\ ] ### 4\. Use Temperature 0 for Consistency [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#4-use-temperature-0-for-consistency) from pydantic_ai import ModelSettings from pydantic_evals.evaluators import LLMJudge LLMJudge( rubric='...', model_settings=ModelSettings(temperature=0.0), ) Limitations ----------- [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#limitations) ### Non-Determinism [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#non-determinism) LLM judges are not deterministic. The same output may receive different scores across runs. **Mitigation:** * Use `temperature=0.0` for more consistency * Run multiple evaluations and average * Use retry strategies for flaky evaluations ### Cost [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#cost) LLM judges make API calls, which cost money and time. **Mitigation:** * Use cheaper models for simple checks (`gpt-5-mini`) * Run deterministic checks first to fail fast * Cache results when possible * Limit evaluation to changed cases ### Model Biases [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#model-biases) LLM judges inherit biases from their training data. **Mitigation:** * Use multiple judge models and compare * Review evaluation reasons, not just scores * Validate judges against human-labeled test sets * Be aware of known biases (length bias, style preferences) ### Context Limits [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#context-limits) Judges have token limits for inputs. **Mitigation:** * Truncate long inputs/outputs intelligently * Use focused rubrics that don’t require full context * Consider chunked evaluation for very long content Debugging LLM Judges -------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#debugging-llm-judges) ### View Reasons [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#view-reasons) from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import LLMJudge def my_task(inputs: str) -> str: return f'Result: {inputs}' dataset = Dataset( name='debug_reasons', cases=[Case(inputs='test')], evaluators=[LLMJudge(rubric='Response is clear')], ) report = dataset.evaluate_sync(my_task) report.print(include_reasons=True) """ Evaluation Summary: my_task ┏━━━━━━━━━━┳━━━━━━━━━━━━━┳━━━━━━━━━━┓ ┃ Case ID ┃ Assertions ┃ Duration ┃ ┡━━━━━━━━━━╇━━━━━━━━━━━━━╇━━━━━━━━━━┩ │ Case 1 │ LLMJudge: ✔ │ 10ms │ │ │ Reason: - │ │ │ │ │ │ │ │ │ │ ├──────────┼─────────────┼──────────┤ │ Averages │ 100.0% ✔ │ 10ms │ └──────────┴─────────────┴──────────┘ """ Output: ┃ Assertions ┃ ┃ accuracy: ✔ ┃ ┃ Reason: The response │ ┃ correctly states... │ ### Access Programmatically [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#access-programmatically) from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import LLMJudge def my_task(inputs: str) -> str: return f'Result: {inputs}' dataset = Dataset( name='programmatic_access', cases=[Case(inputs='test')], evaluators=[LLMJudge(rubric='Response is clear')], ) report = dataset.evaluate_sync(my_task) for case in report.cases: for name, result in case.assertions.items(): print(f'{name}: {result.value}') #> LLMJudge: True if result.reason: print(f' Reason: {result.reason}') #> Reason: - ### Compare Judges [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#compare-judges) Test the same cases with different judge models: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import LLMJudge def my_task(inputs: str) -> str: return f'Result: {inputs}' judges = [\ LLMJudge(rubric='Response is clear', model='openai:gpt-5.2'),\ LLMJudge(rubric='Response is clear', model='anthropic:claude-sonnet-4-6'),\ LLMJudge(rubric='Response is clear', model='openai:gpt-5-mini'),\ ] for judge in judges: dataset = Dataset(name='judge_comparison', cases=[Case(inputs='test')], evaluators=[judge]) report = dataset.evaluate_sync(my_task) # Compare results Advanced: Custom Judge Models ----------------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#advanced-custom-judge-models) Set a default judge model for all `LLMJudge` evaluators: from pydantic_evals.evaluators import LLMJudge from pydantic_evals.evaluators.llm_as_a_judge import set_default_judge_model # Set default to Claude set_default_judge_model('anthropic:claude-sonnet-4-6') # Now all LLMJudge instances use Claude by default LLMJudge(rubric='...') # Uses Claude Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/#next-steps) * **[Custom Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/custom/) ** - Write custom evaluation logic * **[Native Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/built-in/) ** - Complete evaluator reference Was this page helpful? Thanks for your feedback! --- # Advisor | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/advisor/#_top) Advisor ======= Give an executor model a way to consult a separate advisor model before it answers or commits to a decision. [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/advisor/) > The API may change between releases. Breaking changes ship deprecation warnings where practical. Usage ----- [](https://pydantic.dev/docs/ai/harness/advisor/#usage) Pass the advisor model as the first argument. The model can be any model name or model instance accepted by Pydantic AI: from pydantic_ai import Agent from pydantic_ai_harness.advisor import Advisor agent = Agent( 'openai:gpt-5.4', capabilities=[\ Advisor(\ 'anthropic:claude-opus-4-8',\ max_uses=1,\ max_tokens=4096,\ )\ ], ) result = agent.run_sync( 'Design a zero-downtime database migration. Consult the advisor before choosing a plan.' ) print(result.output) The executor decides when to consult. Ask it explicitly in the user prompt or the agent’s instructions when a consultation is required. Provider adaptation ------------------- [](https://pydantic.dev/docs/ai/harness/advisor/#provider-adaptation) `Advisor` exposes one logical tool through two execution paths: * **Native:** when the executor and advisor are both on a compatible Anthropic provider, or both use OpenRouter, Pydantic AI’s provider-native [`AdvisorTool`](https://pydantic.dev/docs/ai/tools-toolsets/native-tools/#advisor-tool) runs the consultation. * **Local fallback:** every other pairing gets an `advisor` function tool. Calling it runs a separate Pydantic AI agent with the configured advisor model. In the default `auto` mode, native selection is conservative. The capability only reuses an explicit provider-qualified model name when the executor and advisor share a provider, so it does not guess how an Anthropic model ID maps to an OpenRouter catalog slug. For example: from pydantic_ai_harness.advisor import Advisor # Native for an Anthropic executor; local for OpenAI, Google, and other executors. anthropic_advisor = Advisor('anthropic:claude-opus-4-8') # Native for an OpenRouter executor. openrouter_advisor = Advisor('openrouter:anthropic/claude-opus-4.8') Passing a `Model` instance selects local execution in `auto` mode. This preserves that instance’s provider, client, credentials, base URL, and instrumentation. String model names are resolved only if local execution is selected, so a native consultation uses the executor provider’s existing configuration. Pydantic AI’s resolved executor model profile makes the final support decision. A client or model that does not support the native tool uses the local fallback. Options ------- [](https://pydantic.dev/docs/ai/harness/advisor/#options) | Option | Default | Behavior | | --- | --- | --- | | `model` | required | Advisor model name or `Model` instance. | | `mode` | `'auto'` | Execution policy: `'auto'`, `'native'`, or `'local'`. | | `max_uses` | `None` | Maximum consultations in one executor model request. Must be at least `1`. | | `max_tokens` | `None` | Maximum output tokens for each consultation. Must be at least `1024`. | | `caching` | `None` | Anthropic-native prompt-cache TTL: `'5m'` or `'1h'`. | | `forward_history` | `False` | Forward completed executor message history to local consultations. | Use `mode='native'` when the consultation must stay inside the executor provider, or `mode='local'` when the configured advisor provider must receive a separate request. Native mode requires an `anthropic:` or `openrouter:` string and an executor on that same provider. It does not fall back when the executor lacks support. `max_uses` has the same per-request scope as Anthropic’s native tool. Only calls whose arguments validate consume this allowance. It resets when the executor makes its next model request. OpenRouter ignores native `max_uses`, so `auto` mode selects the local fallback when this option is set. Combining OpenRouter, `mode='native'`, and `max_uses` is rejected. `caching` is an opportunistic Anthropic-native optimization. OpenRouter and the local fallback have no equivalent control. `forward_history` only affects the local execution path, whether selected explicitly or as the `auto` fallback. It does not alter native tool configuration or native-versus-local selection. When enabled, the local advisor receives the completed executor message history before the current response. The current response, including partial text and unresolved tool calls, is not forwarded, so the consultation prompt still needs to contain the complete current question. String model configurations can be loaded from YAML or JSON agent specs by passing `Advisor` in `custom_capability_types`. Runtime `Model` instances remain Python-only. Context passed to the advisor ----------------------------- [](https://pydantic.dev/docs/ai/harness/advisor/#context-passed-to-the-advisor) The context depends on the execution path: | Path | Advisor context | | --- | --- | | Anthropic native | The provider supplies the full transcript, including system instructions, tool definitions, earlier turns and results, and executor text produced so far. | | OpenRouter native | The executor supplies a consultation prompt. Pydantic AI configures `forward_transcript=false`. | | Local fallback | The executor supplies a consultation prompt through the `advisor` function tool. With `forward_history=True`, the advisor also receives completed executor message history. | The local advisor uses its own fixed instructions. It does not inherit executor dependencies, tools, or toolsets. `forward_history` adds completed messages only; it does not include the executor’s current partial response. For portable behavior, tell the executor to put the question and all relevant evidence in its consultation prompt. The local tool description reinforces this requirement. The local fallback sends that prompt to the configured advisor model and provider. Native execution uses the executor’s provider configuration. Treat this distinction as a data-routing choice when reviewing credentials, transcript sharing, and provider policies. Usage, failures, and observability ---------------------------------- [](https://pydantic.dev/docs/ai/harness/advisor/#usage-failures-and-observability) Local advisor requests share the parent run’s [`RunUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage) and [`UsageLimits`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits) , so their requests and tokens count toward the agent tree’s normal limits. Native providers report advisor usage according to their own protocol. Anthropic records advisor-specific values in `RequestUsage.details`, while OpenRouter exposes aggregate server-tool counts in response provider details. Invalid option combinations fail when `Advisor` is constructed. Executor and provider compatibility is validated when a run prepares its model request. Anthropic reports native advisor errors as tool results so the executor can continue. If a local advisor produces invalid model behavior, the executor receives a normal tool retry, matching Pydantic AI’s other subagent-backed tools. Local model resolution, authentication, provider, request, and usage-limit errors otherwise propagate and can stop the run. When a local call exceeds `max_uses`, the tool returns a bounded message telling the executor to continue without more advice. Composition ----------- [](https://pydantic.dev/docs/ai/harness/advisor/#composition) The capability needs no ordering constraint. It composes with other capabilities and ordinary toolsets through Pydantic AI’s [native-or-local tool selection](https://pydantic.dev/docs/ai/capabilities/overview/#provider-adaptive-tools) . The advisor tool is always visible and is not deferred through [Tool Search](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-search) . It reserves the tool name and toolset ID `advisor`. One `Advisor` instance is supported per agent because the native tool has one stable identity. During streaming, the executor stream pauses while an advisor consultation runs and resumes when the completed advice is available. The local fallback does not splice the advisor model’s token deltas into the executor stream. Local consultations can run in parallel. When `max_uses` is set, calls claim the per-request allowance before starting the advisor model request, so parallel calls cannot exceed it. Native advice is compatible with durable execution because it remains part of the executor model request. Local execution cannot yet preserve the same semantics across every durable backend. Temporal and Prefect can checkpoint the returned advice, but changes to the activity-local or task-local [`RunUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage) do not merge back into the outer run. DBOS does not checkpoint ordinary function-tool calls, so a local advisor request could run again during workflow replay. Use `mode='native'` with a supported provider when running the agent durably. Harness does not inspect durability integrations because Pydantic AI core does not yet expose a public durable-context contract. Local execution, including an `auto` fallback, is therefore unsupported in durable runs rather than rejected by this capability. API reference ------------- [](https://pydantic.dev/docs/ai/harness/advisor/#api-reference) Advisor ------- [](https://pydantic.dev/docs/ai/harness/advisor/#pydantic_ai_harness.advisor.Advisor) **Bases:** `NativeOrLocalTool[AgentDepsT]` Let an agent consult another model through a provider-native tool or local fallback. In `auto` mode, `Advisor` uses Pydantic AI’s native `AdvisorTool` when an explicit provider-qualified model name matches a compatible Anthropic or OpenRouter executor. On every other model, it exposes an `advisor` function tool backed by a separate Pydantic AI agent. from pydantic_ai import Agent from pydantic_ai_harness.advisor import Advisor agent = Agent( 'openai:gpt-5.4', capabilities=[Advisor('anthropic:claude-opus-4-8')], ) ### Attributes [](https://pydantic.dev/docs/ai/harness/advisor/#attributes) #### caching [](https://pydantic.dev/docs/ai/harness/advisor/#pydantic_ai_harness.advisor.Advisor.caching) Anthropic-native advisor prompt caching. This is an opportunistic optimization. OpenRouter and the local fallback do not provide an equivalent cache control. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘5m’, ‘1h’\] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `caching` #### forward\_history [](https://pydantic.dev/docs/ai/harness/advisor/#pydantic_ai_harness.advisor.Advisor.forward_history) Whether local consultations receive the executor’s completed message history. Native execution keeps the provider’s transcript behavior unchanged. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `forward_history` #### max\_tokens [](https://pydantic.dev/docs/ai/harness/advisor/#pydantic_ai_harness.advisor.Advisor.max_tokens) Maximum output tokens for each advisor consultation. Values below 1024 are rejected so the setting remains valid on every native and local execution path. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `max_tokens` #### max\_uses [](https://pydantic.dev/docs/ai/harness/advisor/#pydantic_ai_harness.advisor.Advisor.max_uses) Maximum consultations in one executor model request. The limit resets on the next executor request. OpenRouter’s native advisor does not honor this option, so setting it selects the local fallback there. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `max_uses` #### mode [](https://pydantic.dev/docs/ai/harness/advisor/#pydantic_ai_harness.advisor.Advisor.mode) How advisor consultations are executed. `auto` uses a native advisor only for an explicit same-provider model name. `native` requires a provider-native advisor, and `local` always runs a separate Pydantic AI agent. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘auto’, ‘native’, ‘local’\] **Default:** `mode` #### model [](https://pydantic.dev/docs/ai/harness/advisor/#pydantic_ai_harness.advisor.Advisor.model) The model to consult. Accepts the same model names and model instances as `Agent`. In `auto` mode, model instances use local execution so their provider configuration is preserved. **Type:** `ModelSelection` **Default:** `model` ### Methods [](https://pydantic.dev/docs/ai/harness/advisor/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/harness/advisor/#pydantic_ai_harness.advisor.Advisor.__init__) def __init__( model: ModelSelection, *, mode: Literal['auto', 'native', 'local'] = 'auto', max_uses: int | None = None, max_tokens: int | None = None, caching: Literal['5m', '1h'] | None = None, forward_history: bool = False, ) -> None ##### Returns [](https://pydantic.dev/docs/ai/harness/advisor/#returns) [`None`](https://docs.python.org/3/library/constants.html#None) #### after\_model\_request [](https://pydantic.dev/docs/ai/harness/advisor/#pydantic_ai_harness.advisor.Advisor.after_model_request) `@async` def after_model_request( ctx: RunContext[AgentDepsT], *, request_context: ModelRequestContext, response: ModelResponse, ) -> ModelResponse Reset the local consultation allowance for each executor response. ##### Returns [](https://pydantic.dev/docs/ai/harness/advisor/#returns-1) [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) #### for\_run [](https://pydantic.dev/docs/ai/harness/advisor/#pydantic_ai_harness.advisor.Advisor.for_run) `@async` def for_run(ctx: RunContext[AgentDepsT]) -> Advisor[AgentDepsT] Return a fresh capability with local usage isolated to this run. ##### Returns [](https://pydantic.dev/docs/ai/harness/advisor/#returns-2) `Advisor`\[`AgentDepsT`\] Was this page helpful? Thanks for your feedback! --- # FileSystem | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/filesystem/#_top) FileSystem ========== `FileSystem` gives an agent a fixed set of file tools — read, write, edit, list, search, find, create, and inspect — all scoped to a single `root_dir`. Every path is resolved and containment-checked (symlinks included) before any I/O, and access is filtered through allow / deny / protected glob patterns. [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/filesystem/) The problem ----------- [](https://pydantic.dev/docs/ai/harness/filesystem/#the-problem) Letting an agent touch the filesystem directly is risky: path traversal (`../../etc/passwd`), symlinks that escape the project, clobbering `.git`, or leaking `.env` secrets. Hand-rolling the guards around every tool call is repetitive and easy to get subtly wrong. `FileSystem` centralizes those guards. It exposes one bounded, sandboxed toolset so you configure the boundary once and reuse it across agents. Usage ----- [](https://pydantic.dev/docs/ai/harness/filesystem/#usage) Add `FileSystem` to your agent’s `capabilities` with a `root_dir`. Everything the agent reads or writes is confined to that directory. from pydantic_ai import Agent from pydantic_ai_harness import FileSystem agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[FileSystem(root_dir='./workspace')], ) result = agent.run_sync('Read config.toml and tell me the package name.') print(result.output) `root_dir` defaults to the current directory (`.`), but passing an explicit workspace path is the recommended practice — the sandbox is only as tight as the root you give it. Tools ----- [](https://pydantic.dev/docs/ai/harness/filesystem/#tools) `FileSystem` contributes eight tools, all path-scoped to `root_dir`: | Tool | Purpose | | --- | --- | | `read_file` | Read a text file with line numbers and a content hash. Binary files are detected and not dumped. Supports `offset`/`limit` paging. | | `write_file` | Create or overwrite a file. Optional `expected_hash` rejects stale writes (optimistic concurrency). | | `edit_file` | Exact-string replacement; `old_text` must match exactly once. Optional `expected_hash`. | | `list_directory` | List a directory’s entries with type indicators and sizes. | | `search_files` | Regex search over file contents, optionally narrowed by an `include_glob`. | | `find_files` | Glob search over file names (e.g. `*.py`, `**/*.json`). | | `create_directory` | Create a directory and any missing parents. | | `file_info` | Metadata for a file or directory (size, type, line count, hash, symlink target). | Tool errors the model can correct — a missing file, a denied path, a stale edit — are surfaced as [`ModelRetry`](https://pydantic.dev/docs/ai/core-concepts/agent/#reflection-and-self-correction) , so the agent gets the error message back and can adjust rather than aborting the run. Security model -------------- [](https://pydantic.dev/docs/ai/harness/filesystem/#security-model) * **Containment.** Paths resolve relative to `root_dir`; anything resolving outside — via `..`, an absolute path, or a symlink — is rejected. Symlinks are resolved with `os.path.realpath` _before_ the containment check, closing the TOCTTOU window. * **Binary detection.** `read_file` returns a placeholder instead of dumping binary bytes into the model context. * **Optimistic concurrency.** `write_file`/`edit_file` accept an `expected_hash` so an agent operating on a stale read is told to re-read rather than silently overwriting newer content. Pattern filtering ----------------- [](https://pydantic.dev/docs/ai/harness/filesystem/#pattern-filtering) Three independent glob lists control access. Patterns are matched with `fnmatch`, whose `*` spans `/`, so `*.py` matches `src/main.py` and you rarely need `**`. | Field | Effect | | --- | --- | | `allowed_patterns` | If non-empty, only matching paths are accessible (allowlist). | | `denied_patterns` | Matching paths are always rejected (denylist). | | `protected_patterns` | Matching paths are read-only — reads succeed, writes are rejected. | `protected_patterns` defaults to `.git/*`, `.env`, `.env.*`, `*.pem`, `*.key`, and `**/secrets*`. Pass an empty list to disable protection. from pydantic_ai import Agent from pydantic_ai_harness import FileSystem agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[\ FileSystem(\ root_dir='./workspace',\ allowed_patterns=['*.py', '*.toml'],\ denied_patterns=['**/node_modules/*'],\ ),\ ], ) ### Direct access vs. walkers [](https://pydantic.dev/docs/ai/harness/filesystem/#direct-access-vs-walkers) The three rules apply at two different granularities: * **Direct access** (`read_file`, `write_file`, `edit_file`, `file_info`, `create_directory`) gates the operation’s target path. You must name a path that the patterns permit. * **Walkers** (`list_directory`, `search_files`, `find_files`) gate their root by deny/protected patterns, but **not** by `allowed_patterns` — a directory root like `.` never matches a file pattern such as `src/*.py`, so requiring it to would make every listing fail. Instead, the root is always walked and each **entry** is filtered against all three lists. A directory listing can never surface a path the agent couldn’t otherwise read or write. So with `allowed_patterns=['*.py']`, `list_directory('.')` succeeds and shows only the `.py` entries; `read_file('notes.md')` is rejected. Note that the walkers filter entries with write-level access, so `protected_patterns` matches are omitted from `list_directory`, `search_files`, and `find_files` output even though those exact paths remain directly readable via `read_file`/`file_info`. Configuration ------------- [](https://pydantic.dev/docs/ai/harness/filesystem/#configuration) from pydantic_ai_harness import FileSystem FileSystem( root_dir='.', # str | Path -- sandbox root allowed_patterns=[], # allowlist globs (empty = allow all) denied_patterns=[], # denylist globs protected_patterns=[...], # read-only globs (defaults to secrets/.git) max_read_lines=2000, # cap for a single read_file max_search_results=1000, # cap for search_files max_find_results=1000, # cap for find_files ) The three integer limits must be positive; they are validated at construction and raise `ValueError` otherwise. Agent spec (YAML/JSON) ---------------------- [](https://pydantic.dev/docs/ai/harness/filesystem/#agent-spec-yamljson) `FileSystem` works with Pydantic AI’s [agent spec](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) : model: anthropic:claude-sonnet-4-6 capabilities: - FileSystem: root_dir: ./workspace allowed_patterns: ['*.py', '*.toml'] from pydantic_ai import Agent from pydantic_ai_harness import FileSystem agent = Agent.from_file('agent.yaml', custom_capability_types=[FileSystem]) Pass `custom_capability_types` so the spec loader knows how to instantiate `FileSystem`. Further reading --------------- [](https://pydantic.dev/docs/ai/harness/filesystem/#further-reading) * [Pydantic AI capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) * [Toolsets](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) * [the capabilities overview](https://pydantic.dev/docs/ai/harness/) API reference ------------- [](https://pydantic.dev/docs/ai/harness/filesystem/#api-reference) FileSystem ---------- [](https://pydantic.dev/docs/ai/harness/filesystem/#pydantic_ai_harness.FileSystem) **Bases:** `AbstractCapability[AgentDepsT]` File system access scoped to a root directory. All paths are resolved relative to `root_dir`. Traversal above the root is rejected. Symlinks are resolved before authorization. ### Attributes [](https://pydantic.dev/docs/ai/harness/filesystem/#attributes) #### allowed\_patterns [](https://pydantic.dev/docs/ai/harness/filesystem/#pydantic_ai_harness.FileSystem.allowed_patterns) If non-empty, only paths matching at least one glob pattern are accessible. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `field(default_factory=(list[str]))` #### denied\_patterns [](https://pydantic.dev/docs/ai/harness/filesystem/#pydantic_ai_harness.FileSystem.denied_patterns) Paths matching any of these glob patterns are rejected. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `field(default_factory=(list[str]))` #### max\_find\_results [](https://pydantic.dev/docs/ai/harness/filesystem/#pydantic_ai_harness.FileSystem.max_find_results) Maximum number of matches returned by `find_files`. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `1000` #### max\_read\_lines [](https://pydantic.dev/docs/ai/harness/filesystem/#pydantic_ai_harness.FileSystem.max_read_lines) Maximum number of lines returned by a single `read_file` call. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `2000` #### max\_search\_results [](https://pydantic.dev/docs/ai/harness/filesystem/#pydantic_ai_harness.FileSystem.max_search_results) Maximum number of matches returned by `search_files`. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `1000` #### protected\_patterns [](https://pydantic.dev/docs/ai/harness/filesystem/#pydantic_ai_harness.FileSystem.protected_patterns) Paths matching these patterns are read-only (writes are rejected). Defaults to protecting `.git/`, `.env`, key files, and secrets. Set to an empty list to disable protection. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `field(default_factory=(lambda: list(_DEFAULT_PROTECTED)))` #### root\_dir [](https://pydantic.dev/docs/ai/harness/filesystem/#pydantic_ai_harness.FileSystem.root_dir) Root directory for all file operations. Defaults to the current directory. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | `Path` **Default:** `'.'` ### Methods [](https://pydantic.dev/docs/ai/harness/filesystem/#methods) #### get\_toolset [](https://pydantic.dev/docs/ai/harness/filesystem/#pydantic_ai_harness.FileSystem.get_toolset) def get_toolset() -> FileSystemToolset[AgentDepsT] Build and return the filesystem toolset. ##### Returns [](https://pydantic.dev/docs/ai/harness/filesystem/#returns) `FileSystemToolset`\[`AgentDepsT`\] Was this page helpful? Thanks for your feedback! --- # Media Externalization [Skip to content](https://pydantic.dev/docs/ai/harness/media/#_top) Media ===== A conversation that carries images, audio, or other `BinaryContent` inlines those bytes into every message, and a large text part (a big tool-return string, say) is just as heavy. Persist that history and each snapshot re-serializes the payloads; the same image referenced by ten messages is ten copies of the bytes. Media externalization solves that: content-addressed stores write each payload once, keyed by its own hash, and leave a short `media+sha256://` URI in its place. Reach for it whenever large binary or text payloads would otherwise balloon what you store or send. > The API may change between releases. Where practical, breaking changes ship with a deprecation warning. Building blocks, not a capability --------------------------------- [](https://pydantic.dev/docs/ai/harness/media/#building-blocks-not-a-capability) These are building blocks. There is no class you add to `Agent(capabilities=[...])` yet. [`StepPersistence`](https://pydantic.dev/docs/ai/harness/step-persistence/) already uses them to keep snapshots small when messages carry `BinaryContent` or large text (e.g. a big tool-return string), and a forthcoming `MediaExternalizer` capability ([#254](https://github.com/pydantic/pydantic-ai-harness/issues/254) ) will reuse the same stores to rewrite `BinaryContent` into URL parts before the model sees them. Why content-addressing ---------------------- [](https://pydantic.dev/docs/ai/harness/media/#why-content-addressing) The URI is derived from the payload hash, so identical bytes deduplicate automatically. The same bytes are stored once no matter how many messages or snapshots reference them, and moving the underlying storage is a one-line swap because the URI does not change. Stores ------ [](https://pydantic.dev/docs/ai/harness/media/#stores) Every store implements the `MediaStore` protocol — `put`, `get`, `exists`, `public_url`, and `get_metadata`, all async and content-addressed. | Store | Backed by | Use when | | --- | --- | --- | | `DiskMediaStore(directory=...)` | A directory on disk | Local runs and tests | | `SqliteMediaStore(database=...)` | A SQLite database | A single-file store that travels with the data | | `S3MediaStore(bucket=, endpoint=, region=, ...)` | S3 or an S3-compatible bucket | Shared or production storage | | `MongoMediaStore(client= or db_url=, database=, ...)` | MongoDB (sha256-addressed manual chunking) | A MongoDB deployment; blobs larger than one BSON document | `S3MediaStore` uses path-style URLs plus handrolled SigV4, so it is compatible with AWS S3, Cloudflare R2 (`region='auto'`), MinIO, and other S3-compatible providers. `SqliteMediaStore` also accepts `connection=` instead of `database=` to share a `sqlite3.Connection`. `MongoMediaStore` needs the `mongodb` extra (`pip install pydantic-ai-harness[mongodb]`, which installs `pymongo>=4.17.0`). Pass a shared `AsyncMongoClient` as `client=`, or a connection string as `db_url=` (the store then owns the client — call `await store.aclose()` to release it); `database=` is always required. Each blob is stored as sha256-addressed chunks in a `media_chunks` collection, with a `media` manifest document per blob (`_id = `). The chunking bounds each BSON document, so a blob larger than MongoDB’s 16 MiB document cap still stores and reads back. It does not bound memory: `put` takes the whole payload as `bytes` and `get` reassembles every chunk into one `bytearray`, so a blob has to fit in process memory in both directions — there is no streaming API. The manifest holds `MediaContext.metadata` inline and is not chunked, so keep per-blob metadata small. Manual chunking is used rather than the GridFS driver on purpose: the digest is the manifest `_id`, so identical bytes deduplicate (GridFS keys files by `ObjectId` and does no dedup), and the plain-collection surface stays fully testable in-memory. Two constructor knobs shape that layout. `collection=` (default `'media'`) names the manifest collection and derives the chunk collection as `_chunks`; names outside `A-Za-z_*` are rejected. `chunk_size_bytes=` (default 8 MiB) sets the split size and is rejected below 1 byte or above 16 MiB minus 64 KiB of headroom for the chunk document’s own fields, since a larger chunk would build a document MongoDB refuses on insert. On its first `put` or `get`, the store issues `createIndex` for a compound `(files_id, n)` index on the chunk collection, without which reassembly is a collection scan. The connecting user therefore needs the privilege to create indexes — a restricted Atlas role may not have it — and pointing the store at an already-populated collection pays the index build on that first call. from pymongo import AsyncMongoClient from pydantic_ai_harness.media import MongoMediaStore client = AsyncMongoClient('mongodb://localhost:27017') store = MongoMediaStore(client=client, database='agent_media') Walker helpers -------------- [](https://pydantic.dev/docs/ai/harness/media/#walker-helpers) `externalize_media` and `restore_media` walk a message node and swap payloads for URIs and back: from pydantic_ai_harness.media import DiskMediaStore, externalize_media, restore_media store = DiskMediaStore(directory='./media') # Replace binary and text payloads at or above the threshold with media+sha256:// URIs. lean = await externalize_media(message, media_store=store, threshold_bytes=32_000) # Later, rehydrate the URIs back into the original parts. full = await restore_media(lean, media_store=store) `externalize_media` externalizes both large `BinaryContent` and large text: any message part whose string `content` reaches `threshold_bytes` UTF-8 bytes (`TextPart`, `ThinkingPart`, a string-returning `ToolReturnPart`, a string-valued `UserPromptPart`), plus any `TextContent` element travelling inside a `UserPromptPart.content` sequence or a `ToolReturn`. The same `threshold_bytes` governs binary and text, and payloads below it stay inline. Round-trip is transparent — `restore_media` re-inlines binary bytes and text symmetrically. If you need to key media yourself, `media_uri_for` and `parse_media_uri` give you the raw URI round-trip. The current reader restores binary markers written before text externalization. That compatibility is upgrade-only: a release that predates text externalization treats every marker as binary, so it cannot validate a snapshot containing an externalized text marker. Keep a current reader for persisted snapshots that contain those markers. Public URLs ----------- [](https://pydantic.dev/docs/ai/harness/media/#public-urls) When a store is fronted by a CDN, a local HTTP server, or a signed-URL service, pass a `public_url=` resolver (or use `make_static_public_url`) to turn a stored `media+sha256://` URI into a URL the model can fetch directly. Without a resolver, `public_url(...)` returns `None`. A static base URL, for a public bucket or CDN: from pydantic_ai_harness.media import S3MediaStore, make_static_public_url store = S3MediaStore( bucket='my-bucket', endpoint='https://.r2.cloudflarestorage.com', region='auto', access_key_id=..., secret_access_key=..., key_prefix='media/', public_url=make_static_public_url('https://pub-abc.r2.dev', key_prefix='media/'), ) A presigned or rotating-signature URL — pass any async callable that takes `(uri, MediaContext)`: from pydantic_ai_harness.media import MediaContext, S3MediaStore async def presign(uri: str, ctx: MediaContext) -> str: key = 'media/' + uri.removeprefix('media+sha256://') + '.bin' return await my_signer.generate(key, ttl=3600, content_type=ctx.media_type) store = S3MediaStore(..., public_url=presign) This is what the forthcoming `MediaExternalizer` will use to swap `BinaryContent` parts for `ImageUrl` / `AudioUrl` / other URL parts before the model sees the message, letting providers fetch big media over the wire without re-encoding bytes into the request body. Emitting a URL is always safe: pydantic-ai providers transparently download the bytes when the target model does not natively accept that URL type, so you only ever lose wire savings, never correctness. `MediaContext` -------------- [](https://pydantic.dev/docs/ai/harness/media/#mediacontext) Every store method and both user-supplied callables (`PublicUrlResolver`, `KeyStrategy`) accept a `MediaContext` — an extensible per-operation bag: from collections.abc import Mapping from dataclasses import dataclass, field @dataclass(frozen=True, kw_only=True) class MediaContext: media_type: str | None = None # e.g. 'image/png' filename: str | None = None # original filename, when known metadata: Mapping[str, str] = field(default_factory=dict) # user-supplied tags All fields default, so you pass what you have and ignore the rest; new fields are added non-breakingly as use cases emerge. `get_metadata(uri)` round-trips the user-supplied `metadata` mapping on all four stores; `media_type` is persisted separately (as the byte payload’s `Content-Type`). `KeyStrategy` ------------- [](https://pydantic.dev/docs/ai/harness/media/#keystrategy) The default on-store key layout is `.bin`. `DiskMediaStore` and `S3MediaStore` accept a `key_strategy=` override to fit an existing layout. `SqliteMediaStore` and `MongoMediaStore` do not, since the digest is their primary key — use `table=` / `collection=` to move the rows or documents instead: from pydantic_ai_harness.media import DiskMediaStore, MediaContext def by_media_type(uri: str, ctx: MediaContext) -> str: digest = uri.removeprefix('media+sha256://') ext = {'image/png': '.png', 'image/jpeg': '.jpg'}.get(ctx.media_type or '', '.bin') return f'images/{digest}{ext}' store = DiskMediaStore('runs', key_strategy=by_media_type) If your strategy depends on `ctx.media_type`, the same context must be supplied at read time for `get`/`exists` to find the blob. `DiskMediaStore` rejects strategies that produce absolute paths or `..` segments, to keep writes inside the store directory. `default_key_strategy` is exported if you want to build on it. API --- [](https://pydantic.dev/docs/ai/harness/media/#api) | Symbol | Purpose | | --- | --- | | `MediaStore` | Async content-addressed store protocol (`put` / `get` / `exists` / `public_url` / `get_metadata`) | | `DiskMediaStore`, `SqliteMediaStore`, `S3MediaStore`, `MongoMediaStore` | Concrete stores (`MongoMediaStore` needs the `mongodb` extra) | | `MediaContext` | Per-operation context (media type, filename, tags) threaded through store operations | | `KeyStrategy`, `default_key_strategy` | On-store key layout | | `PublicUrlResolver`, `make_static_public_url` | Resolve a stored URI to a public URL | | `externalize_media`, `restore_media` | Walk a message node to externalize / rehydrate large binary and text payloads | | `media_uri_for`, `parse_media_uri` | Compute and parse a `media+sha256://` URI | Source: [`pydantic_ai_harness/media/`](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/media/) . Related ------- [](https://pydantic.dev/docs/ai/harness/media/#related) * [Step Persistence](https://pydantic.dev/docs/ai/harness/step-persistence/) — the first consumer of these stores, externalizing large `BinaryContent` and text parts in run snapshots. Was this page helpful? Thanks for your feedback! --- # Managed Prompt | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/managed-prompt/#_top) Managed Prompt ============== `ManagedPrompt` backs an agent’s instructions with a [Logfire-managed prompt](https://logfire.pydantic.dev/docs/reference/advanced/prompt-management/) , so you can iterate on your system prompt from the Logfire UI — versioned, labelled, and rolled out — without touching code or redeploying. It’s a Pydantic AI [capability](https://pydantic.dev/docs/ai/harness/) , so you wire it in through the `capabilities=` parameter on `Agent`. [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/logfire/) Install the `logfire` extra: Terminal uv add "pydantic-ai-harness[logfire]" The problem it solves --------------------- [](https://pydantic.dev/docs/ai/harness/managed-prompt/#the-problem-it-solves) Prompts are critical to agent behavior, but iterating on them through the normal edit -> review -> deploy loop is slow. You can’t easily A/B test a change, and you can’t roll it back the moment it misbehaves in production without shipping a new build. `ManagedPrompt` moves the prompt out of your codebase and into Logfire’s managed-variable store. It declares the backing managed variable for you and resolves it **once per run**, feeding the resolved value into the agent’s instructions. Resolution happens inside the run’s [`wrap_run`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability.wrap_run) hook, using the [`ResolvedVariable`](https://logfire.pydantic.dev/docs/reference/advanced/managed-variables/) as a context manager that stays open for the whole run — so the selected label and version are attached as baggage to every child span of the agent run. You get a direct correlation between a run’s behavior and the exact prompt version that produced it, plus instant iteration and rollback from the Logfire UI. Usage ----- [](https://pydantic.dev/docs/ai/harness/managed-prompt/#usage) Pass the prompt name and a default value. The name `support_agent` is declared as the managed variable `prompt__support_agent` — the naming Logfire’s Prompt management uses (hyphens in a name become underscores). The `default` keeps the agent working until a remote value is published, so your code always runs even before you create the prompt in Logfire. import logfire from pydantic_ai import Agent from pydantic_ai_harness.logfire import ManagedPrompt logfire.configure() agent = Agent( 'openai:gpt-5', capabilities=[\ ManagedPrompt(\ 'support_agent',\ default='You are a helpful customer support agent. Be friendly and concise.',\ label='production',\ )\ ], ) result = agent.run_sync('My order never arrived.') print(result.output) Pinning `label='production'` is the recommended default: the resolved value only changes on a deliberate prompt rollout, which keeps the provider prompt cache hot (see [Prompt-cache trade-off](https://pydantic.dev/docs/ai/harness/managed-prompt/#prompt-cache-trade-off) below). Targeting --------- [](https://pydantic.dev/docs/ai/harness/managed-prompt/#targeting) For deterministic A/B assignment (the same user always sees the same label), pass a `targeting_key`. It can be a static string or a callable that derives the key from the [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) — handy when the key lives in your agent’s `deps`: from dataclasses import dataclass from pydantic_ai import Agent from pydantic_ai_harness.logfire import ManagedPrompt @dataclass class Deps: user_id: str agent = Agent( 'openai:gpt-5', deps_type=Deps, capabilities=[\ ManagedPrompt(\ 'support_agent',\ default='You are a helpful customer support agent.',\ targeting_key=lambda ctx: ctx.deps.user_id,\ ),\ ], ) Pass `attributes` (a mapping, or a callable returning one) for condition-based targeting rules. When `label` is omitted, the variable’s rollout and targeting rules pick the label. When both `targeting_key` and `attributes` are omitted, Logfire falls back to its own targeting context and then to the active trace id. For Logfire-side targeting that lives outside the agent (e.g. set once per request handler), use Logfire’s [`targeting_context`](https://logfire.pydantic.dev/docs/reference/advanced/managed-variables/) in an outer scope; `ManagedPrompt` only needs `targeting_key` / `attributes` when the key comes from the agent’s `RunContext`. Templating with deps -------------------- [](https://pydantic.dev/docs/ai/harness/managed-prompt/#templating-with-deps) By default the resolved prompt is used verbatim. Pass `render_template=True` to render it as a Handlebars template against the agent’s `deps` — the same mechanism as [`TemplateStr`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/) — so `{{field}}` is filled from `deps`: from dataclasses import dataclass from pydantic_ai import Agent from pydantic_ai_harness.logfire import ManagedPrompt @dataclass class Deps: customer_name: str agent = Agent( 'openai:gpt-5', deps_type=Deps, capabilities=[\ ManagedPrompt(\ 'support_agent',\ default='You are helping {{customer_name}}. Be friendly and concise.',\ render_template=True,\ ),\ ], ) Rendering requires `pydantic-handlebars` (install `pydantic-ai-slim[spec]`). It is off by default. Prompt-cache trade-off ---------------------- [](https://pydantic.dev/docs/ai/harness/managed-prompt/#prompt-cache-trade-off) The resolved value lands in the agent’s **system instructions**. Provider prompt caches (Anthropic, OpenAI, etc.) key strictly by prefix — `tools -> system -> messages` — so any change to the system block invalidates the cached prefix for the affected runs. | Mode | Cache impact | | --- | --- | | Pinned `label='production'`, no rollout split | **Cache-stable.** The value only changes on a deliberate prompt rollout, which is the same cost as a redeploy. | | Percentage rollout across labels (no `label=`) | Different runs land on different labels -> splits the cache into one lane per label. | | `targeting_key` per user/tenant with multiple labels in play | Cache lanes per assigned label; deterministic per key but still N lanes overall. | | Mid-traffic label flip in the Logfire UI | One-shot cold-invalidation for everyone on that label. | In short: pinning a `label` keeps the cache hot; using `ManagedPrompt` as an A/B platform is opt-in cache cost. If you don’t need rollouts, `label='production'` is the recommended default. Bringing your own variable -------------------------- [](https://pydantic.dev/docs/ai/harness/managed-prompt/#bringing-your-own-variable) Declaring the same name more than once is fine — each `ManagedPrompt` builds its own backing variable, so sharing a prompt across several agents just works. Pass an existing [`logfire.variables.Variable`](https://logfire.pydantic.dev/docs/reference/advanced/managed-variables/) as the first argument instead of a name when you want to declare the variable yourself — for example a template variable, or one registered for `variables_push`: import logfire from pydantic_ai import Agent from pydantic_ai_harness.logfire import ManagedPrompt logfire.configure() support_prompt = logfire.var( name='prompt__support_agent', type=str, default='You are a helpful customer support agent. Be friendly and concise.', ) agent = Agent('openai:gpt-5', capabilities=[ManagedPrompt(support_prompt, label='production')]) When `name` is a prompt name (not a `Variable`), pass `logfire_instance=` to declare the variable on a specific Logfire instance instead of the module-level default. `default` is required when `name` is a prompt name and is ignored when you pass a `Variable` (which already carries its own default and instance). How it composes --------------- [](https://pydantic.dev/docs/ai/harness/managed-prompt/#how-it-composes) * **Resolves once per run.** A label flip or rollout change that lands in Logfire mid-run is not picked up until the next run starts — the trade-off for run-stable instructions and a single baggage scope across all child spans. * **Runs outermost.** The capability wraps [`Instrumentation`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Instrumentation) so the resolved variable’s baggage covers the agent run span as well as its children. On recent Logfire versions both the selected label and the version are propagated as separate baggage attributes. * **Concurrency-safe.** Resolution is isolated per run via a context variable, so a single capability instance is safe to share across concurrent runs. * **Inspectable mid-run.** `ManagedPrompt.resolved` exposes the active run’s `ResolvedVariable` (`value`, `label`, `version`, `reason`) for inspection — e.g. from inside a tool. It is `None` outside a run. API reference ------------- [](https://pydantic.dev/docs/ai/harness/managed-prompt/#api-reference) The resolved prompt is a `str`. Pass the bare prompt name (the `prompt__` prefix and hyphen-to-underscore normalization are applied for you) and a `default`, then use `label`, `targeting_key`, `attributes`, `render_template`, and `logfire_instance` to control resolution. ManagedPrompt ------------- [](https://pydantic.dev/docs/ai/harness/managed-prompt/#pydantic_ai_harness.ManagedPrompt) **Bases:** `AbstractCapability[AgentDepsT]` Back an agent’s instructions with a Logfire-managed prompt. **Prompt-cache trade-off:** the resolved value lands in the system instructions block, so any Logfire-side change to the prompt (new version rollout, label flip, A/B targeting) invalidates the provider’s prompt cache for the affected runs. Pin a `label` (e.g. `'production'`) for the cache-stable path; treat percentage rollouts and per-user targeting as opt-in cache cost. See the README’s “Prompt-cache trade-off” section for the full picture. Pass the managed prompt name and a default value and the capability declares the backing [managed variable](https://logfire.pydantic.dev/docs/reference/advanced/managed-variables/) for you — a name of `support_agent` resolves the variable `prompt__support_agent`, matching the naming Logfire’s [Prompt management](https://logfire.pydantic.dev/docs/reference/advanced/prompt-management/) uses. You can iterate on the prompt from the Logfire UI — versioned, labelled, and rolled out — without redeploying, while the code default keeps the agent working when no remote value is available. import logfire from pydantic_ai import Agent from pydantic_ai_harness.logfire import ManagedPrompt logfire.configure() agent = Agent( 'openai:gpt-5', capabilities=[\ ManagedPrompt(\ 'support_agent',\ default='You are a helpful customer support agent. Be friendly and concise.',\ label='production',\ )\ ], ) result = agent.run_sync('My order never arrived.') The prompt value is resolved **once per run**, inside the run’s [`wrap_run`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability.wrap_run) hook, using the `ResolvedVariable` as a context manager that stays open for the whole run — so the selected label and version are attached as baggage to every child span of the agent run. Declaring the same name more than once is fine — each `ManagedPrompt` constructs its own backing variable, so sharing a prompt across several agents just works. Pass an existing `logfire.variables.Variable` as `name` instead of a prompt name when you want to use a variable you defined yourself (for example a `template_var`, or one registered for [`variables_push`](https://logfire.pydantic.dev/docs/api/logfire/#logfire.Logfire.variables_push) ). ### Attributes [](https://pydantic.dev/docs/ai/harness/managed-prompt/#attributes) #### attributes [](https://pydantic.dev/docs/ai/harness/managed-prompt/#pydantic_ai_harness.ManagedPrompt.attributes) Attributes for condition-based targeting rules, or a callable that derives them from the [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) . **Type:** [`Mapping`](https://docs.python.org/3/library/typing.html#typing.Mapping) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] | [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[[`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext)\ \[`AgentDepsT`\]\], [`Mapping`](https://docs.python.org/3/library/typing.html#typing.Mapping)\ \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### default [](https://pydantic.dev/docs/ai/harness/managed-prompt/#pydantic_ai_harness.ManagedPrompt.default) Code-default prompt text. Required when `name` is a prompt name; ignored when `name` is a `Variable`. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### label [](https://pydantic.dev/docs/ai/harness/managed-prompt/#pydantic_ai_harness.ManagedPrompt.label) Explicit targeting label on the Logfire managed prompt to resolve (e.g. `'production'`). When `None`, the targeting rules on the managed variable select the label. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### logfire\_instance [](https://pydantic.dev/docs/ai/harness/managed-prompt/#pydantic_ai_harness.ManagedPrompt.logfire_instance) Logfire instance to resolve the variable on. When `None`, the global default instance (the one backing the module-level `logfire.var`) is used. Ignored when `name` is a `Variable`. **Type:** `Logfire` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### name [](https://pydantic.dev/docs/ai/harness/managed-prompt/#pydantic_ai_harness.ManagedPrompt.name) The managed prompt name (declared as the variable `prompt__`), or a pre-built `logfire.Variable`. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | `Variable`\[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] #### render\_template [](https://pydantic.dev/docs/ai/harness/managed-prompt/#pydantic_ai_harness.ManagedPrompt.render_template) When `True`, render the resolved prompt as a Handlebars template against the agent’s `deps` (the same mechanism as [`TemplateStr`](https://pydantic.dev/docs/ai/api/pydantic-ai/template/#pydantic_ai.template.TemplateStr) ); `{{field}}` is filled from `deps`. Requires `pydantic-handlebars` (install `pydantic-ai-slim[spec]`). Defaults to `False`, so the resolved prompt is used verbatim. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### resolved [](https://pydantic.dev/docs/ai/harness/managed-prompt/#pydantic_ai_harness.ManagedPrompt.resolved) The prompt resolution for the active run, or `None` outside a run. Exposes the full `ResolvedVariable` (`value`, `label`, `version`, `reason`, …) so callers can inspect which prompt version is in play. **Type:** `ResolvedVariable`\[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) #### targeting\_key [](https://pydantic.dev/docs/ai/harness/managed-prompt/#pydantic_ai_harness.ManagedPrompt.targeting_key) Stable key that seeds Logfire’s deterministic rollout assignment — the same key always lands in the same percentage bucket, so a given user keeps the same label across runs. Accepts a static value or a callable that derives it from the [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) . When `None`, Logfire falls back to its own targeting context and then the active trace id. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[[`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext)\ \[`AgentDepsT`\]\], [`str`](https://docs.python.org/3/library/stdtypes.html#str)\ | [`None`](https://docs.python.org/3/library/constants.html#None)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` ### Methods [](https://pydantic.dev/docs/ai/harness/managed-prompt/#methods) #### get\_instructions [](https://pydantic.dev/docs/ai/harness/managed-prompt/#pydantic_ai_harness.ManagedPrompt.get_instructions) def get_instructions() -> Callable[[RunContext[AgentDepsT]], str | None] Provide the resolved prompt to the agent’s system prompt. ##### Returns [](https://pydantic.dev/docs/ai/harness/managed-prompt/#returns) [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[[`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext)\ \[`AgentDepsT`\]\], [`str`](https://docs.python.org/3/library/stdtypes.html#str)\ | [`None`](https://docs.python.org/3/library/constants.html#None)\ \] #### get\_ordering [](https://pydantic.dev/docs/ai/harness/managed-prompt/#pydantic_ai_harness.ManagedPrompt.get_ordering) def get_ordering() -> CapabilityOrdering Run outermost so the prompt’s baggage envelops the whole run, including the run span. ##### Returns [](https://pydantic.dev/docs/ai/harness/managed-prompt/#returns-1) `CapabilityOrdering` #### wrap\_run [](https://pydantic.dev/docs/ai/harness/managed-prompt/#pydantic_ai_harness.ManagedPrompt.wrap_run) `@async` def wrap_run( ctx: RunContext[AgentDepsT], *, handler: WrapRunHandler, ) -> AgentRunResult[Any] Resolve the prompt once and keep its baggage active for the duration of the run. ##### Returns [](https://pydantic.dev/docs/ai/harness/managed-prompt/#returns-2) [`AgentRunResult`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult) \[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] Was this page helpful? Thanks for your feedback! --- # Command Line Interface (CLI) | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/integrations/cli/#_top) Command Line Interface (CLI) ============================ **Pydantic AI** comes with a CLI, `clai` (pronounced “clay”). You can use it to chat with various LLMs and quickly get answers, right from the command line, or spin up a uvicorn server to chat with your Pydantic AI agents from your browser. Installation ------------ [](https://pydantic.dev/docs/ai/integrations/cli/#installation) You can run the `clai` using [`uvx`](https://docs.astral.sh/uv/guides/tools/) : Terminal uvx clai Or install `clai` globally [with `uv`](https://docs.astral.sh/uv/guides/tools/#installing-tools) : Terminal uv tool install clai ... clai Or with `pip`: Terminal pip install clai ... clai CLI Usage --------- [](https://pydantic.dev/docs/ai/integrations/cli/#cli-usage) You’ll need to set an environment variable depending on the provider you intend to use. E.g. if you’re using OpenAI, set the `OPENAI_API_KEY` environment variable: Terminal export OPENAI_API_KEY='your-api-key-here' Then running `clai` will start an interactive session where you can chat with the AI model. Special commands available in interactive mode: * `/exit`: Exit the session * `/markdown`: Show the last response in markdown format * `/multiline`: Toggle multiline input mode (use Ctrl+D to submit) * `/cp`: Copy the last response to clipboard * `/usage`: Show cumulative token usage for the session (turns, input, output, requests, tool calls); add `--json` for a single-line JSON object ### CLI Options [](https://pydantic.dev/docs/ai/integrations/cli/#cli-options) | Option | Description | | --- | --- | | `prompt` | AI prompt for one-shot mode (positional). If omitted, starts interactive mode. | | `-m`, `--model` | Model to use in `provider:model` format (e.g., `openai:gpt-5.2`) | | `-a`, `--agent` | Custom agent in `module:variable` format | | `-t`, `--code-theme` | Syntax highlighting theme (`dark`, `light`, or [pygments theme](https://pygments.org/styles/)
) | | `--no-stream` | Disable streaming from the model | | `-l`, `--list-models` | List all available models and exit | | `--version` | Show version and exit | ### Choose a model [](https://pydantic.dev/docs/ai/integrations/cli/#choose-a-model) You can specify which model to use with the `--model` flag: Terminal clai --model anthropic:claude-sonnet-4-6 (a full list of models available can be printed with `clai --list-models`) ### Custom Agents [](https://pydantic.dev/docs/ai/integrations/cli/#custom-agents) You can specify a custom agent using the `--agent` flag with a module path and variable name: custom\_agent.py from pydantic_ai import Agent agent = Agent('openai:gpt-5.2', instructions='You always respond in Italian.') Then run: Terminal clai --agent custom_agent:agent "What's the weather today?" The format must be `module:variable` where: * `module` is the importable Python module path * `variable` is the name of the Agent instance in that module Additionally, you can directly launch CLI mode from an `Agent` instance using `Agent.to_cli_sync()`: agent\_to\_cli\_sync.py from pydantic_ai import Agent agent = Agent('openai:gpt-5.2', instructions='You always respond in Italian.') agent.to_cli_sync() You can also use the async interface with `Agent.to_cli()`: agent\_to\_cli.py from pydantic_ai import Agent agent = Agent('openai:gpt-5.2', instructions='You always respond in Italian.') async def main(): await agent.to_cli() _(You’ll need to add `asyncio.run(main())` to run `main`)_ ### Message History [](https://pydantic.dev/docs/ai/integrations/cli/#message-history) Both `Agent.to_cli()` and `Agent.to_cli_sync()` support a `message_history` parameter, allowing you to continue an existing conversation or provide conversation context: agent\_with\_history.py from pydantic_ai import ( Agent, ModelMessage, ModelRequest, ModelResponse, TextPart, UserPromptPart, ) agent = Agent('openai:gpt-5.2') # Create some conversation history message_history: list[ModelMessage] = [\ ModelRequest([UserPromptPart(content='What is 2+2?')]),\ ModelResponse([TextPart(content='2+2 equals 4.')])\ ] # Start CLI with existing conversation context agent.to_cli_sync(message_history=message_history) The CLI will start with the provided conversation history, allowing the agent to refer back to previous exchanges and maintain context throughout the session. Web Chat UI ----------- [](https://pydantic.dev/docs/ai/integrations/cli/#web-chat-ui) Launch a web-based chat interface by running: Terminal clai web -m openai:gpt-5.2 This will start a web server (default: [http://127.0.0.1:7932](http://127.0.0.1:7932/) ) with a chat interface. You can also serve an existing agent. For example, if you have an agent defined in `my_agent.py`: from pydantic_ai import Agent my_agent = Agent('openai:gpt-5.2', instructions='You are a helpful assistant.') Launch the web UI: Terminal # With a custom agent clai web --agent my_module:my_agent # With specific models (first is default when no --agent) clai web -m openai:gpt-5.2 -m anthropic:claude-sonnet-4-6 # With native tools clai web -m openai:gpt-5.2 -t web_search -t code_execution # Generic agent with system instructions clai web -m openai:gpt-5.2 -i 'You are a helpful coding assistant' # Custom agent with extra instructions for each run clai web --agent my_module:my_agent -i 'Always respond in Spanish' ### Web UI Options [](https://pydantic.dev/docs/ai/integrations/cli/#web-ui-options) | Option | Description | | --- | --- | | `--agent`, `-a` | Agent to serve in [`module:variable` format](https://pydantic.dev/docs/ai/integrations/cli/#custom-agents) | | `--model`, `-m` | Models to list as options in the UI (repeatable) | | `--tool`, `-t` | [Native tool](https://pydantic.dev/docs/ai/tools-toolsets/native-tools/)
s to list as options in the UI (repeatable). See [available tools](https://pydantic.dev/docs/ai/guides/web/#native-tool-support)
. | | `--instructions`, `-i` | System instructions. When `--agent` is specified, these are additional to the agent’s existing instructions. | | `--host` | Host to bind server (default: 127.0.0.1) | | `--port` | Port to bind server (default: 7932) | | `--html-source` | URL or file path for the chat UI HTML. | | `--allowed-host` | Hostname to answer to in addition to IP addresses and `localhost` (repeatable). See [Reaching the UI under a hostname](https://pydantic.dev/docs/ai/guides/web/#reaching-the-ui-under-a-hostname)
. | When using `--agent`, the agent’s configured model becomes the default. CLI models (`-m`) are additional options. Without `--agent`, the first `-m` model is the default. The web chat UI can also be launched programmatically using [`Agent.to_web()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.to_web) , see the [Web UI documentation](https://pydantic.dev/docs/ai/guides/web/) . Run the `web` command with `--help` to see all available options: Terminal clai web --help Was this page helpful? Thanks for your feedback! --- # AG-UI | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#_top) AG-UI ===== The [Agent-User Interaction (AG-UI) Protocol](https://docs.ag-ui.com/introduction) is an open standard introduced by the [CopilotKit](https://webflow.copilotkit.ai/blog/introducing-ag-ui-the-protocol-where-agents-meet-users) team that standardises how frontend applications communicate with AI agents, with support for streaming, frontend tools, shared state, and custom events. Installation ------------ [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#installation) The only dependencies are: * [ag-ui-protocol](https://docs.ag-ui.com/introduction) : to provide the AG-UI types and encoder. * [starlette](https://www.starlette.io/) : to handle [ASGI](https://asgi.readthedocs.io/en/latest/) requests from a framework like FastAPI. You can install Pydantic AI with the `ag-ui` extra to ensure you have all the required AG-UI dependencies: * [pip](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#tab-panel-94) * [uv](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#tab-panel-95) Terminal pip install 'pydantic-ai-slim[ag-ui]' Terminal uv add 'pydantic-ai-slim[ag-ui]' To run the examples you’ll also need: * [uvicorn](https://uvicorn.dev/) or another ASGI compatible server * [pip](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#tab-panel-96) * [uv](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#tab-panel-97) Terminal pip install uvicorn Terminal uv add uvicorn Usage ----- [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#usage) There are three ways to run a Pydantic AI agent based on AG-UI run input with streamed AG-UI events as output, from most to least flexible. If you’re using a Starlette-based web framework like FastAPI, you’ll typically want to use the second method. 1. The `AGUIAdapter.run_stream()` method, when called on an [`AGUIAdapter`](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter) instantiated with an agent and an AG-UI [`RunAgentInput`](https://docs.ag-ui.com/sdk/python/core/types#runagentinput) object, will run the agent and return a stream of AG-UI events. It also takes optional [`Agent.iter()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.iter) arguments including `deps`. Use this if you’re using a web framework not based on Starlette (e.g. Django or Flask) or want to modify the input or output some way. 2. The `AGUIAdapter.dispatch_request()` class method takes an agent and a Starlette request (e.g. from FastAPI) coming from an AG-UI frontend, and returns a streaming Starlette response of AG-UI events that you can return directly from your endpoint. It also takes optional [`Agent.iter()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.iter) arguments including `deps`, that you can vary for each request (e.g. based on the authenticated user). This is a convenience method that combines [`AGUIAdapter.from_request()`](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.from_request) , `AGUIAdapter.run_stream()`, and `AGUIAdapter.streaming_response()`. 3. Build a stand-alone [`Starlette`](https://www.starlette.io/applications/) app with a single `/` route that calls `AGUIAdapter.dispatch_request()`. The same Starlette app can be [mounted](https://fastapi.tiangolo.com/advanced/sub-applications/) at a path in an existing FastAPI app. When a run ends in [first-party cancellation](https://pydantic.dev/docs/ai/core-concepts/agent/#cancelling-a-run) — `ctx.cancel()`, `AgentRun.cancel()`, or a [`CancellationToken`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.CancellationToken) your server wires to a cancel endpoint — the adapter closes any open text or tool events and emits a bare `RUN_FINISHED`. AG-UI currently has no cancelled outcome, so cancellation is not reported as `RUN_ERROR`. Pass an `on_cancel` callback (see the `run_stream()` example below) to persist the resumable message history from [`RunCancelled.all_messages()`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled.all_messages) . ### Handle run input and output directly [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#handle-run-input-and-output-directly) This example uses `AGUIAdapter.run_stream()` and performs its own request parsing and response generation. This can be modified to work with any web framework. run\_ag\_ui.py import json from http import HTTPStatus from fastapi import FastAPI from fastapi.requests import Request from fastapi.responses import Response, StreamingResponse from pydantic import ValidationError from pydantic_ai import Agent, RunCancelled from pydantic_ai.ui import SSE_CONTENT_TYPE from pydantic_ai.ui.ag_ui import AGUIAdapter agent = Agent('openai:gpt-5.2', instructions='Be fun!') app = FastAPI() async def on_cancel(cancelled: RunCancelled) -> None: messages = cancelled.all_messages() # the resumable history to persist print(f'cancelled after {len(messages)} messages') @app.post('/') async def run_agent(request: Request) -> Response: accept = request.headers.get('accept', SSE_CONTENT_TYPE) try: run_input = AGUIAdapter.build_run_input(await request.body()) # (1) except ValidationError as e: return Response( content=json.dumps(e.json()), media_type='application/json', status_code=HTTPStatus.UNPROCESSABLE_ENTITY, ) adapter = AGUIAdapter(agent=agent, run_input=run_input, accept=accept) event_stream = adapter.run_stream(on_cancel=on_cancel) # (2) sse_event_stream = adapter.encode_stream(event_stream) return StreamingResponse(sse_event_stream, media_type=accept) # (3) 1. [`AGUIAdapter.build_run_input()`](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.build_run_input) takes the request body as bytes and returns an AG-UI [`RunAgentInput`](https://docs.ag-ui.com/sdk/python/core/types#runagentinput) object. You can also use the [`AGUIAdapter.from_request()`](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.from_request) class method to build an adapter directly from a request. 2. `AGUIAdapter.run_stream()` runs the agent and returns a stream of AG-UI events. It supports the same optional arguments as [`Agent.run_stream_events()`](https://pydantic.dev/docs/ai/core-concepts/agent/#running-agents) , including `deps`. You can also use `AGUIAdapter.run_stream_native()` to run the agent and return a stream of Pydantic AI events instead, which can then be transformed into AG-UI events using `AGUIAdapter.transform_stream()`. 3. `AGUIAdapter.encode_stream()` encodes the stream of AG-UI events as strings according to the accept header value. You can also use `AGUIAdapter.streaming_response()` to generate a streaming response directly from the AG-UI event stream returned by `run_stream()`. Since `app` is an ASGI application, it can be used with any ASGI server: Terminal uvicorn run_ag_ui:app This will expose the agent as an AG-UI server, and your frontend can start sending requests to it. ### Handle a Starlette request [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#handle-a-starlette-request) This example uses `AGUIAdapter.dispatch_request()` to directly handle a FastAPI request and return a response. Something analogous to this will work with any Starlette-based web framework. handle\_ag\_ui\_request.py from fastapi import FastAPI from starlette.requests import Request from starlette.responses import Response from pydantic_ai import Agent from pydantic_ai.ui.ag_ui import AGUIAdapter agent = Agent('openai:gpt-5.2', instructions='Be fun!') app = FastAPI() @app.post('/') async def run_agent(request: Request) -> Response: return await AGUIAdapter.dispatch_request(request, agent=agent) # (1) 1. This method essentially does the same as the previous example, but it’s more convenient to use when you’re already using a Starlette/FastAPI app. Since `app` is an ASGI application, it can be used with any ASGI server: Terminal uvicorn handle_ag_ui_request:app This will expose the agent as an AG-UI server, and your frontend can start sending requests to it. ### Stand-alone ASGI app [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#stand-alone-asgi-app) When you don’t already have a Starlette/FastAPI app to mount the route on, build a minimal [`Starlette`](https://www.starlette.io/applications/) app whose single `/` route calls `AGUIAdapter.dispatch_request()`: ag\_ui\_app.py from starlette.applications import Starlette from starlette.requests import Request from starlette.responses import Response from starlette.routing import Route from pydantic_ai import Agent from pydantic_ai.ui.ag_ui import AGUIAdapter agent = Agent('openai:gpt-5.2', instructions='Be fun!') async def run_agent(request: Request) -> Response: return await AGUIAdapter.dispatch_request(request, agent=agent) app = Starlette(routes=[Route('/', run_agent, methods=['POST'])]) Since `app` is an ASGI application, it can be used with any ASGI server: Terminal uvicorn ag_ui_app:app This will expose the agent as an AG-UI server, and your frontend can start sending requests to it. Design ------ [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#design) The Pydantic AI AG-UI integration supports all features of the spec: * [Events](https://docs.ag-ui.com/concepts/events) * [Messages](https://docs.ag-ui.com/concepts/messages) * [State Management](https://docs.ag-ui.com/concepts/state) * [Tools](https://docs.ag-ui.com/concepts/tools) The integration receives messages in the form of a [`RunAgentInput`](https://docs.ag-ui.com/sdk/python/core/types#runagentinput) object that describes the details of the requested agent run including message history, state, and available tools. These are converted to Pydantic AI types and passed to the agent’s run method. Events from the agent, including tool calls, are converted to AG-UI events and streamed back to the caller as Server-Sent Events (SSE). A user request may require multiple round trips between client UI and Pydantic AI server, depending on the tools and events needed. Features -------- [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#features) ### State management [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#state-management) The integration provides full support for [AG-UI state management](https://docs.ag-ui.com/concepts/state) , which enables real-time synchronization between agents and frontend applications. In the example below we have document state which is shared between the UI and server using the [`StateDeps`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.StateDeps) [dependencies type](https://pydantic.dev/docs/ai/core-concepts/dependencies/) that can be used to automatically validate state contained in [`RunAgentInput.state`](https://docs.ag-ui.com/sdk/js/core/types#runagentinput) using a Pydantic `BaseModel` specified as a generic parameter. ag\_ui\_state.py from dataclasses import replace from pydantic import BaseModel from starlette.applications import Starlette from starlette.requests import Request from starlette.responses import Response from starlette.routing import Route from pydantic_ai import Agent from pydantic_ai.ui import StateDeps from pydantic_ai.ui.ag_ui import AGUIAdapter class DocumentState(BaseModel): """State for the document being written.""" document: str = '' agent = Agent( 'openai:gpt-5.2', instructions='Be fun!', deps_type=StateDeps[DocumentState], ) deps = StateDeps(DocumentState()) async def run_agent(request: Request) -> Response: # `dispatch_request` mutates `deps.state` from the request, so give each request its own copy. return await AGUIAdapter.dispatch_request(request, agent=agent, deps=replace(deps)) app = Starlette(routes=[Route('/', run_agent, methods=['POST'])]) Since `app` is an ASGI application, it can be used with any ASGI server: Terminal uvicorn ag_ui_state:app --host 0.0.0.0 --port 9000 ### Tools [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#tools) AG-UI frontend tools are seamlessly provided to the Pydantic AI agent, enabling rich user experiences with frontend user interfaces. ### Context [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#context) Alongside messages, an AG-UI client can send a `context` array of `description`/`value` pairs describing things it considers relevant to the run: the originating platform, the requesting user, or a channel’s standing instructions. Every entry is a claim the client made — it can send any `description`/`value` it likes — so they describe a request, they never establish who is making it. These entries are not passed to the model automatically, and they don’t belong in [instructions](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.instructions) . Instructions carry operator authority — they’re treated as _your_ instruction to the model — so building them out of text a client sent lets a prompt injection inherit that authority. Delivering them as data denies them that authority but doesn’t make them safe: they’re still indirect prompt-injection input, so scope and re-authorize side-effecting tools from `deps` your server established, never from an entry’s `description` or `value`. See [Mid-conversation system prompts](https://pydantic.dev/docs/ai/core-concepts/message-history/#mid-conversation-system-prompts) and the [trust model](https://pydantic.dev/docs/ai/integrations/ui/overview/#trust-model-for-client-submitted-messages) . Read the entries off `adapter.run_input.context` and deliver them to the model as **data**. Facts your server established — the authenticated user, the workspace — are what go in instructions: ag\_ui\_context.py from dataclasses import dataclass from ag_ui.core import Context from starlette.requests import Request from starlette.responses import Response from pydantic_ai import Agent, RunContext from pydantic_ai.ui.ag_ui import AGUIAdapter @dataclass class ChannelDeps: workspace: str # (1) context: list[Context] # (2) agent = Agent('openai:gpt-5.2', deps_type=ChannelDeps) @agent.instructions def workspace(ctx: RunContext[ChannelDeps]) -> str: return f'You are answering in the {ctx.deps.workspace} workspace.' @agent.tool def frontend_context(ctx: RunContext[ChannelDeps]) -> list[str]: """Context the frontend says is relevant to this conversation.""" return [f'{entry.description}: {entry.value}' for entry in ctx.deps.context] def authenticated_workspace(request: Request) -> str: """Whatever your auth layer already established — a session, a signed token, an API key.""" ... async def run_agent(request: Request) -> Response: adapter = await AGUIAdapter.from_request(request, agent=agent) deps = ChannelDeps(workspace=authenticated_workspace(request), context=adapter.run_input.context) return adapter.streaming_response(adapter.run_stream(deps=deps)) Established by your server, so it can shape how the agent behaves. Sent by the client, so it reaches the model as tool output the agent can read — never as an instruction. To let a client-supplied fact change how the agent behaves, authenticate it first: verify the caller or channel, look up the policy _your_ server holds for it, and write the instruction from that. The entry itself stays data. Anything that isn’t meant for the model at all — a Slack channel ID, a locale — is better carried in `forwardedProps`, which the adapter passes through untouched as `adapter.run_input.forwarded_props`. Validating it proves shape, not identity: who the user is, what tenant they’re in, and what they’re allowed to do come from authenticated server state. `context`, `forwardedProps` and `parentRunId` are read straight off [`run_input`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.run_input) rather than through adapter properties of their own. The adapter’s properties — `messages`, `toolset`, `state`, `conversation_id`, `deferred_tool_results` — are the concepts every UI protocol shares and that the adapter itself feeds into the agent run. These three are AG-UI-specific and consumed only by your code, so they stay on the protocol object where their types are the protocol’s own. ### Tool approval (interrupts) [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#tool-approval-interrupts) Tools declared with `requires_approval=True` map onto AG-UI’s [interrupt-aware run lifecycle](https://docs.ag-ui.com/concepts/interrupts) . When the model proposes such a call, the run pauses and the adapter ends the SSE stream with a `RUN_FINISHED` event whose `outcome.type` is `"interrupt"` and whose `outcome.interrupts[]` describes each pending approval. The client renders an approval UI from that list and POSTs the next `RunAgentInput` with a `resume[]` array of `ResumeEntry` items addressing each interrupt. The mapping the adapter applies (matching the AG-UI Python SDK field names): | AG-UI direction | Pydantic AI source / sink | | --- | --- | | `Interrupt.reason` | Always `"tool_call"` for `requires_approval=True` tools | | `Interrupt.tool_call_id` | The `ToolCallPart.tool_call_id` of the proposed call | | `Interrupt.id` | `f"int-{tool_call_id}"` (round-trips back to `tool_call_id` on resume) | | `Interrupt.metadata` | `DeferredToolRequests.metadata.get(tool_call_id)` | | `ResumeEntry.payload` | `{ "approved": bool, "editedArgs"?: object, "reason"?: string }`, validated against `Interrupt.response_schema`; a payload that fails validation denies, including when the offending field is not `approved` itself — a wrongly-typed `editedArgs` or `reason` denies even alongside `approved=True`, while omitting either optional field or sending it as `null` is accepted | | `payload.approved=True` | [`ToolApproved`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.ToolApproved) | | `payload.editedArgs` | [`ToolApproved.override_args`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.ToolApproved.override_args)
(fully replaces the proposed args) | | `payload.approved=False` | [`ToolDenied`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.ToolDenied)
with `message=payload.reason`. `approved` is required, so a payload that omits it denies on validation and its `reason` is not used — send `approved: false` explicitly to have your `reason` reach the model | | `status="cancelled"` | [`ToolDenied`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.ToolDenied)
with `message="Cancelled by user."` regardless of payload | The agent must include [`DeferredToolRequests`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.DeferredToolRequests) in its `output_type` so the run can pause cleanly instead of erroring on the proposed call: ag\_ui\_tool\_approval.py from starlette.applications import Starlette from starlette.requests import Request from starlette.responses import Response from starlette.routing import Route from pydantic_ai import Agent from pydantic_ai.tools import DeferredToolRequests from pydantic_ai.ui.ag_ui import AGUIAdapter agent = Agent('openai:gpt-5.2', output_type=[str, DeferredToolRequests]) @agent.tool_plain(requires_approval=True) def delete_file(path: str) -> str: """Delete a file. Pauses on a `RUN_FINISHED` interrupt outcome until the user approves.""" return f'deleted {path}' async def run_agent(request: Request) -> Response: return await AGUIAdapter.dispatch_request(request, agent=agent) app = Starlette(routes=[Route('/', run_agent, methods=['POST'])]) On the resumed turn the agent re-executes the tool against the **original** `tool_call_id`, so only a `TOOL_CALL_RESULT` event is emitted for that id — no fresh `TOOL_CALL_START`. This preserves the audit trail the AG-UI spec requires. See [Deferred tools and human-in-the-loop tool approval](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/) for the underlying Pydantic AI primitive that also works outside AG-UI. ### Events [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#events) Pydantic AI tools can send [AG-UI events](https://docs.ag-ui.com/concepts/events) simply by returning a [`ToolReturn`](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#advanced-tool-returns) object with a [`BaseEvent`](https://docs.ag-ui.com/sdk/python/core/events#baseevent) (or a list of events) as `metadata`, which allows for custom events and state updates. ag\_ui\_tool\_events.py from dataclasses import replace from ag_ui.core import CustomEvent, EventType, StateSnapshotEvent from pydantic import BaseModel from starlette.applications import Starlette from starlette.requests import Request from starlette.responses import Response from starlette.routing import Route from pydantic_ai import Agent, RunContext, ToolReturn from pydantic_ai.ui import StateDeps from pydantic_ai.ui.ag_ui import AGUIAdapter class DocumentState(BaseModel): """State for the document being written.""" document: str = '' agent = Agent( 'openai:gpt-5.2', instructions='Be fun!', deps_type=StateDeps[DocumentState], ) deps = StateDeps(DocumentState()) async def run_agent(request: Request) -> Response: return await AGUIAdapter.dispatch_request(request, agent=agent, deps=replace(deps)) app = Starlette(routes=[Route('/', run_agent, methods=['POST'])]) @agent.tool async def update_state(ctx: RunContext[StateDeps[DocumentState]]) -> ToolReturn: return ToolReturn( return_value='State updated', metadata=[\ StateSnapshotEvent(\ type=EventType.STATE_SNAPSHOT,\ snapshot=ctx.deps.state,\ ),\ ], ) @agent.tool_plain async def custom_events() -> ToolReturn: return ToolReturn( return_value='Count events sent', metadata=[\ CustomEvent(\ type=EventType.CUSTOM,\ name='count',\ value=1,\ ),\ CustomEvent(\ type=EventType.CUSTOM,\ name='count',\ value=2,\ ),\ ] ) Since `app` is an ASGI application, it can be used with any ASGI server: Terminal uvicorn ag_ui_tool_events:app --host 0.0.0.0 --port 9000 ### Protocol version compatibility [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#protocol-version-compatibility) Pydantic AI supports every `ag-ui-protocol` release from `0.1.10` on, and features added after that floor are gated on the version you have installed rather than requiring an upgrade. That gate runs in both directions. On the way out, content an older protocol version can’t express is downgraded or omitted — see [`AGUIAdapter.ag_ui_version`](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.ag_ui_version) for the negotiated thresholds. On the way in, a message `role` or input content `type` your installed `ag-ui-protocol` has no class for is skipped with a `UserWarning` naming the tag, and the rest of the request runs — so a frontend on a newer protocol version than your server keeps working, minus the content your install has no type for. For instance, a gateway that forwards image attachments as typed multimodal content (`ag-ui-protocol >= 0.1.15`) still delivers the accompanying text to an agent running on an older install. What gets skipped is decided by the tag alone: any `role` or `type` string the installed models don’t declare qualifies, so a client that misspells `"txet"` is skipped with the same warning as one sending genuinely newer content — the server has no way to tell those apart. The skip is scoped to well-formed items: a message must still carry a string `id`, the field every AG-UI message type requires. Everything else is still rejected with `422 Unprocessable Entity` — a payload that is malformed under a `role` or `type` the install _does_ know, a `role` or `type` that isn’t a string at all, and a body that isn’t valid JSON. If you see the warning and the content was real, upgrading `ag-ui-protocol` is what makes it reach your agent. ### Trust model [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#trust-model) AG-UI’s `RunAgentInput.messages` is fully client-controlled. The [`AGUIAdapter`](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter) applies defaults to strip untrusted parts before the agent runs — see [Trust model for client-submitted messages](https://pydantic.dev/docs/ai/integrations/ui/overview/#trust-model-for-client-submitted-messages) in the UI adapter overview, which covers system prompts, file URL schemes, uploaded files ([`allow_uploaded_files`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.allow_uploaded_files) ), and unresolved tool calls. Those defaults don’t make client-submitted history authentic — see [Trust boundary for client-supplied history](https://pydantic.dev/docs/ai/core-concepts/message-history/#trust-boundary-for-client-supplied-history) . ### Compaction [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#compaction) [`CompactionPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.CompactionPart) s round-trip through AG-UI activity messages (`pydantic_ai_compaction`), so [compacted](https://pydantic.dev/docs/ai/capabilities/compaction/) conversations keep working when the frontend holds the message history. A compaction item submitted by the frontend is honored — the conversation stays compacted — with two caveats. First, it is never trusted to stand in for the system prompt: whichever prompt applies per [System prompts and instructions](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#system-prompts-and-instructions) still reaches the model on every request. Second, if the run also receives server-side `message_history` (the [server-side persistence pattern](https://pydantic.dev/docs/ai/integrations/ui/overview/#trust-model-for-client-submitted-messages) ), frontend compaction items are ignored — everything before a compaction item is hidden from the model, so honoring one from the frontend would let it hide the server’s stored history. See [Client-held history](https://pydantic.dev/docs/ai/capabilities/compaction/#client-held-history) for the trade-offs and the recommended server-side pattern. ### Preserving failed tool outcomes [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#preserving-failed-tool-outcomes) AG-UI’s [`ToolCallResultEvent`](https://github.com/ag-ui-protocol/ag-ui/blob/11f03fa65c4fa22a8637d3f6e06e77d8c1b9ae78/docs/sdk/python/core/events.mdx#L284-L304) has no error or outcome field. Although [encrypted reasoning continuity](https://github.com/ag-ui-protocol/ag-ui/blob/11f03fa65c4fa22a8637d3f6e06e77d8c1b9ae78/docs/concepts/reasoning.mdx#L6-L29) is the intended use of [`ReasoningEncryptedValueEvent`](https://github.com/ag-ui-protocol/ag-ui/blob/11f03fa65c4fa22a8637d3f6e06e77d8c1b9ae78/docs/sdk/python/core/events.mdx#L555-L577) , it is also AG-UI’s standard event for attaching `encrypted_value` to a message or tool call. Pydantic AI uses that attachment mechanism with a namespaced payload to preserve `outcome='failed'` from [`ToolReturnPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolReturnPart) when using `ag-ui-protocol >= 0.1.13`. If the client sends those messages back on a later run, the adapter restores the failed outcome. This is a history-continuity mechanism: it does not set [`ToolMessage.error`](https://github.com/ag-ui-protocol/ag-ui/blob/11f03fa65c4fa22a8637d3f6e06e77d8c1b9ae78/docs/concepts/messages.mdx#L143-L163) or guarantee that a frontend visually renders the result as an error. Event streams produced with earlier protocol versions have no metadata carrier for the outcome, so reloading them reconstructs the tool result as `outcome='success'`. ### Preserving files across round-trips [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#preserving-files-across-round-trips) AG-UI has no native representation for agent-generated files ([`FilePart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.FilePart) ) or [`UploadedFile`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.UploadedFile) references, so they are omitted from `dump_messages` output by default. Set [`AGUIAdapter.preserve_file_data`](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.preserve_file_data) to `True` to round-trip them through reserved `pydantic_ai_*` [activity messages](https://docs.ag-ui.com/concepts/messages) , which a frontend completes by echoing those activity messages back on the next request. This is a representation opt-in, not a security one: an `UploadedFile` reconstructed from a round-tripped activity message is still subject to the inbound [`allow_uploaded_files`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.allow_uploaded_files) gate before it reaches the agent. ### System prompts and instructions [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#system-prompts-and-instructions) Pydantic AI supports two ways to provide guidance to the model: [`system_prompt`](https://pydantic.dev/docs/ai/core-concepts/agent/#system-prompts) (stored in the message history as [`SystemPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SystemPromptPart) s) and [`instructions`](https://pydantic.dev/docs/ai/core-concepts/agent/#instructions) (injected fresh on every request, never persisted). When you control the server side, `instructions` is the recommended default. The rest of this section only matters if you use `system_prompt`. If you only use `instructions`, there’s nothing to configure — they’re always applied regardless of the AG-UI message history. For `system_prompt`, you choose who owns it with the `manage_system_prompt` parameter on [`AGUIAdapter`](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter) : * `'server'` (default): the agent’s configured `system_prompt` is authoritative. Any `SystemMessage` sent by the frontend is stripped with a warning (a malicious client could otherwise inject arbitrary instructions via crafted API requests), and the agent’s own system prompt is reinjected at the head of the first request via the [`ReinjectSystemPrompt`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ReinjectSystemPrompt) capability. * `'client'`: the frontend owns the system prompt. Frontend `SystemMessage`s are preserved as-is, and the agent’s configured `system_prompt` is not injected — the caller is fully responsible for sending it on every turn if desired. To opt into fallback-to-configured behavior, add the [`ReinjectSystemPrompt`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ReinjectSystemPrompt) capability to your agent. ag\_ui\_client\_managed\_system\_prompt.py from fastapi import FastAPI from starlette.requests import Request from starlette.responses import Response from pydantic_ai import Agent from pydantic_ai.ui.ag_ui import AGUIAdapter agent = Agent('openai:gpt-5.2') app = FastAPI() @app.post('/') async def run_agent(request: Request) -> Response: return await AGUIAdapter.dispatch_request( request, agent=agent, manage_system_prompt='client' ) Examples -------- [](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/#examples) For more examples see [`pydantic_ai_examples.ag_ui`](https://github.com/pydantic/pydantic-ai/tree/main/examples/pydantic_ai_examples/ag_ui) , which includes a server for use with the [AG-UI Dojo](https://docs.ag-ui.com/tutorials/debugging#the-ag-ui-dojo) . Was this page helpful? Thanks for your feedback! --- # Contributing | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/project/contributing/#_top) Contributing ============ We’d love you to contribute to Pydantic AI! How we work — the short version ------------------------------- [](https://pydantic.dev/docs/ai/project/contributing/#how-we-work--the-short-version) Pydantic AI is maintained by a small team. We set our own priorities based on what benefits the most users, and we work through issues and PRs in that order — not in the order they arrive. * **Found a bug?** Open an issue with a clear description and a minimal reproducible example. Including a [Logfire](https://logfire.pydantic.dev/) trace link helps us debug dramatically faster. * **Want a feature or API change?** Open an issue describing the problem you’re solving. Do not start with code. * **Want to help build a feature?** Comment on the issue explaining why you need it and what context you bring. We call this being a “champion” — more on that below. * **Have a fix or code to share?** Make sure a maintainer has agreed to the approach on the issue and assigned you. Then open a PR. The rest of this page explains why we work this way and what to expect. Before you write code --------------------- [](https://pydantic.dev/docs/ai/project/contributing/#before-you-write-code) For anything non-trivial, align with a maintainer on the approach before writing code. A pre-aligned PR is much faster to land than one we’re seeing cold. ### Trivial fixes [](https://pydantic.dev/docs/ai/project/contributing/#trivial-fixes) Typos, broken links, small doc improvements, obvious one-line fixes: just open a PR. No issue needed. ### Bug fixes [](https://pydantic.dev/docs/ai/project/contributing/#bug-fixes) If the fix could reasonably go more than one way, or you’re unsure it’s actually a bug: open an issue first. Include a minimal reproducible example and ideally a [Logfire trace link](https://logfire.pydantic.dev/) showing the problem. For well-scoped bugs, we may generate a fix internally — the most valuable thing you can do is file a clear report and then validate that the fix works for your use case. ### Features, integrations, or API changes [](https://pydantic.dev/docs/ai/project/contributing/#features-integrations-or-api-changes) Before writing code, ask whether the change needs to live in core at all. Most new agent behaviors belong in [**Pydantic AI Harness**](https://github.com/pydantic/pydantic-ai-harness) , the official capability library — not in this repo. Pydantic AI core is for the agent loop, model providers, and capabilities that require model-specific support or are fundamental to the agent experience. Standalone capabilities — guardrails, memory, context management, file system access, etc. — belong in the harness, where they can iterate faster. See [What goes where?](https://pydantic.dev/docs/ai/harness/#what-goes-where) for the full distinction. **If your idea is a capability**, open an issue on [pydantic-ai-harness](https://github.com/pydantic/pydantic-ai-harness/issues) instead. You can also publish capabilities as your own package using the `pydantic-ai-` convention — see [Publishing capability packages](https://pydantic.dev/docs/ai/guides/extensibility/#publishing-capability-packages) . Once a capability has real users and a stable shape, we can talk about upstreaming to harness or core. If it does belong in core: 1. **Search first.** If an existing issue covers your need, comment there. If the closest match is only related, open a new issue and link it. 2. **Describe the problem, not just the solution.** Tell us what you’re building, what’s blocking you, and what you’ve tried. This context matters more than code. 3. **Propose a plan before building.** Post the shape of the solution on the issue, or open a draft PR with just a `PLAN.md`. For larger features, we do short video calls with contributors to iterate on the design — a 20-minute call often saves weeks of async review cycles. 4. **Wait for assignment.** A maintainer needs to agree on the approach and assign the issue to you before you open a PR. Unassigned PRs may be auto-closed. Champions --------- [](https://pydantic.dev/docs/ai/project/contributing/#champions) A “champion” is someone who needs a feature, has context on the problem, and is willing to invest time to help us get it right. If you want to champion a feature: * Comment on the issue explaining: what you’re building, why you need this, and what you can contribute (domain knowledge, testing, validation). * We prioritize features where one or more champions with production use cases have stepped up. A feature with no champion stays in the backlog until either we prioritize it ourselves or someone with real context shows up. * Being a champion doesn’t mean writing the code. It means shaping the plan and validating the result. For significant features, we’ll set up a call to iterate on the design together. Champions are credited as co-authors when the feature ships. What to expect during review ---------------------------- [](https://pydantic.dev/docs/ai/project/contributing/#what-to-expect-during-review) ### We review PRs in our priority order, not submission order [](https://pydantic.dev/docs/ai/project/contributing/#we-review-prs-in-our-priority-order-not-submission-order) We do not automatically triage every new PR. PRs on issues we have not pre-aligned on are not in our review queue, regardless of how well written they are. If no maintainer has agreed to the change on an issue and assigned it to you, assume we have not seen your PR. Even for PRs with code we’ve previously engaged with: we treat all contributed code as a starting point, not a finished product. We review and prioritize PRs based on the feature’s importance to the project, not on how much effort went into the code. This is a change from how open source traditionally worked, and we’d rather be honest about it than leave PRs sitting with no signal. **If you want to know where your PR stands**, the best thing to do is ping `#pydantic-ai` on [Pydantic Slack](https://logfire.pydantic.dev/docs/join-slack/) . ### We may rewrite or supersede your code [](https://pydantic.dev/docs/ai/project/contributing/#we-may-rewrite-or-supersede-your-code) We treat contributed code as illustrative: a starting point that shows the shape of the change and proves the approach works, not the final form we merge. The most useful thing you can give us for a non-trivial change is a plan plus a working example — not a polished, merge-ready implementation. On any PR, we may push commits to your branch, open a follow-up PR that supersedes yours, or rewrite from scratch. For security reasons, we lean toward rewriting contributed code rather than merging as-is. You will still be credited as the original author. Please don’t spend effort chasing green CI, addressing every automated review comment, or rebasing for merge conflicts on a PR we haven’t pre-aligned on. If we take the change forward, that polish gets thrown away when we rewrite. Get the approach working, then stop and ping us on Slack. Do not force-push updates to an open PR. Rewriting its commits invalidates previous reviews; push follow-up commits instead. We will squash them when merging. ### Automated review is advisory, not a gate [](https://pydantic.dev/docs/ai/project/contributing/#automated-review-is-advisory-not-a-gate) PRs are automatically reviewed by Devin and our own tooling. These reviews are advisory: * A bot approval does not mean your PR is ready to merge. Only a human maintainer’s review counts. * A bot finding does not mean you must act on it. If you disagree, say so. * If automated review is generating noise on your PR, tell us. We use that feedback to retune the tooling. ### Priority [](https://pydantic.dev/docs/ai/project/contributing/#priority) We receive far more contributions than we can review, and we focus where it has the most impact. We cannot promise to get to every PR, even good ones, and we’d rather say so up front than leave your work open indefinitely with no signal. How we weigh priorities: * **User demand** — features that more users need get priority. Champion-backed features with production use cases outrank speculative additions. * **Provider significance** — work that affects frontier providers (Anthropic, OpenAI, Google) or providers we know are heavily used gets priority. A model integration for a niche provider will wait; a fix for Anthropic won’t. * **Roadmap alignment** — features that align with our current focus areas get priority. Right now that includes the capabilities/hooks API, provider-adaptive tools, and the [Pydantic AI Harness](https://github.com/pydantic/pydantic-ai-harness) capability library. * **Capabilities over core** — features that could live as a [capability](https://pydantic.dev/docs/ai/capabilities/overview/) should go to [Pydantic AI Harness](https://github.com/pydantic/pydantic-ai-harness) or ship as your own package — that’s often the fastest path. Once it has traction, come back and we can talk about upstreaming. If your PR or issue has gone quiet ---------------------------------- [](https://pydantic.dev/docs/ai/project/contributing/#if-your-pr-or-issue-has-gone-quiet) 1. Ping `#pydantic-ai` on [Pydantic Slack](https://logfire.pydantic.dev/docs/join-slack/) with a link. 2. Say what you need: “Can you take a look?”, “I’m blocked — is this on your radar?”, or “Should I close this?” are all fine. 3. If you’ve been waiting weeks without any human response, flag it. That’s a process failure on our side and we want to know. Installation and Setup ---------------------- [](https://pydantic.dev/docs/ai/project/contributing/#installation-and-setup) Clone your fork and cd into the repo directory Terminal git clone git@github.com:/pydantic-ai.git cd pydantic-ai Install `uv` (version 0.4.30 or later) and `pre-commit`: * [`uv` install docs](https://docs.astral.sh/uv/getting-started/installation/) * [`pre-commit` install docs](https://pre-commit.com/#install) To install `pre-commit` you can run the following command: Terminal uv tool install pre-commit Install `pydantic-ai`, all dependencies and pre-commit hooks Terminal make install Running Tests etc. ------------------ [](https://pydantic.dev/docs/ai/project/contributing/#running-tests-etc) We use `make` to manage most commands you’ll need to run. For details on available commands, run: Terminal make help To run code formatting, linting, static type checks, and tests with coverage report generation, run: Terminal make Documentation Changes --------------------- [](https://pydantic.dev/docs/ai/project/contributing/#documentation-changes) [`docs/navigation.yml`](https://github.com/pydantic/pydantic-ai/blob/main/docs/navigation.yml) owns the sidebar, public routes, and redirects for [Pydantic AI’s documentation](https://pydantic.dev/docs/ai/) . Update it when adding, removing, or moving a page. All routes in `docs/navigation.yml` are relative to the Pydantic AI documentation root. Give each page its complete canonical route in `slug`; use `aliases` only for redirect sources. Do not prefix either value with `/ai` or a leading slash. For the rendered site, use the documentation preview attached to a pull request after a maintainer adds the `trigger:docs` label. Rules for adding new models to Pydantic AI ------------------------------------------ [](https://pydantic.dev/docs/ai/project/contributing/#new-model-rules) To avoid an excessive workload for the maintainers of Pydantic AI, we can’t accept all model contributions, so we’re setting the following rules for when we’ll accept new models and when we won’t. This should hopefully reduce the chances of disappointment and wasted work. * To add a new model with an extra dependency, that dependency needs > 500k monthly downloads from PyPI consistently over 3 months or more * To add a new model which uses another models logic internally and has no extra dependencies, that model’s GitHub org needs > 20k stars in total * For any other model that’s just a custom URL and API key, we’re happy to add a one-paragraph description with a link and instructions on the URL to use * For any other model that requires more logic, we recommend you release your own Python package `pydantic-ai-xxx`, which depends on [`pydantic-ai-slim`](https://pydantic.dev/docs/ai/overview/install/#slim-install) and implements a model that inherits from our [`Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) ABC If you’re unsure about adding a model, please [create an issue](https://github.com/pydantic/pydantic-ai/issues) . Was this page helpful? Thanks for your feedback! --- # ag_ui | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/ui/ag_ui/#_top) ag\_ui ====== AG-UI protocol integration for Pydantic AI agents. AGUIAdapter ----------- [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter) **Bases:** `UIAdapter[RunAgentInput, Message, BaseEvent, AgentDepsT, OutputDataT]` UI adapter for the Agent-User Interaction (AG-UI) protocol. ### Attributes [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#attributes) #### ag\_ui\_version [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.ag_ui_version) AG-UI protocol version controlling behavior thresholds. Accepts any version string (e.g. `'0.1.13'`). Defaults to the version detected from the installed `ag-ui-protocol` package. Known thresholds: * `< 0.1.13`: emits `THINKING_*` events during streaming, drops `ThinkingPart` from `dump_messages` output. * `>= 0.1.13`: emits `REASONING_*` events with encrypted metadata during streaming, and includes `ThinkingPart` as `ReasoningMessage` in `dump_messages` output for full round-trip fidelity of thinking signatures and provider metadata. * `>= 0.1.15`: emits typed multimodal input content (`ImageInputContent`, `AudioInputContent`, `VideoInputContent`, `DocumentInputContent`) instead of generic `BinaryInputContent`. `load_messages` always accepts `ReasoningMessage` and multimodal content types regardless of this setting, and `build_run_input` skips inbound content types the installed `ag-ui-protocol` predates rather than rejecting the request. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `DEFAULT_AG_UI_VERSION` #### conversation\_id [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.conversation_id) Conversation ID from the AG-UI `RunAgentInput.threadId`. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### deferred\_tool\_results [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.deferred_tool_results) Translate AG-UI `RunAgentInput.resume[]` into Pydantic AI `DeferredToolResults`. See [docs.ag-ui.com/concepts/interrupts](https://docs.ag-ui.com/concepts/interrupts) . Each `ResumeEntry` is mapped to an approval keyed by the original `tool_call_id`. The payload is validated against the same Pydantic model whose JSON schema is advertised on `Interrupt.response_schema`, and the mapping is **deny-by-default**: approval requires a payload that validates with `approved=True`. Any other shape is treated as a denial so a malformed or hostile client cannot accidentally execute a tool that requires human approval. * `status == 'cancelled'` → `ToolDenied('Cancelled by user.')` * `payload.approved is True` with a valid `payload.editedArgs` dict → `ToolApproved(override_args=...)` * `payload.approved is True` without edits → `ToolApproved()` * Anything else (`False`, missing, `null`, non-bool `approved`, non-dict payload, a non-dict `editedArgs`, or a non-string `reason`) → `ToolDenied(payload.reason)` if `reason` is a non-empty string on a payload that validated, else `ToolDenied()` (which carries the default `"The tool call was denied."` message). Returns `None` when `resume` is missing or empty, or when the installed ag-ui-protocol predates the interrupt lifecycle. **Type:** [`DeferredToolResults`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.DeferredToolResults) | [`None`](https://docs.python.org/3/library/constants.html#None) #### messages [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.messages) Pydantic AI messages from the AG-UI run input. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage)\ \] #### preserve\_file\_data [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.preserve_file_data) Whether to round-trip `FilePart` and `UploadedFile` through reserved `pydantic_ai_*` [activity messages](https://docs.ag-ui.com/concepts/messages) . Defaults to `False`. AG-UI has no native representation for agent-generated files ([`FilePart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.FilePart) ) or uploaded-file references ([`UploadedFile`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.UploadedFile) ), so when this is `True` they are serialized as sidecar activity messages on `dump_messages` and reconstructed on `load_messages`. A frontend only completes the round-trip if it echoes these activity messages back on the next request. This is a representation setting, not a security one: honoring a reconstructed inbound `UploadedFile` still requires [`allow_uploaded_files`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.allow_uploaded_files) , which the shared `sanitize_messages` step enforces regardless of this flag. Multimodal tool-return files are unaffected — they ride inline in `ToolMessage.content`. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### state [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.state) Frontend state from the AG-UI run input. **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) #### toolset [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.toolset) Toolset representing frontend tools from the AG-UI run input. **Type:** [`AbstractToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AbstractToolset) \[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) ### Methods [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#methods) #### build\_event\_stream [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.build_event_stream) def build_event_stream( ) -> UIEventStream[RunAgentInput, BaseEvent, AgentDepsT, OutputDataT] Build an AG-UI event stream transformer. ##### Returns [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#returns) `UIEventStream`\[`RunAgentInput`, `BaseEvent`, `AgentDepsT`, `OutputDataT`\] #### build\_run\_input [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.build_run_input) `@classmethod` def build_run_input(cls, body: bytes) -> RunAgentInput Build an AG-UI run input object from the request body. A message `role` or input content `type` introduced by a protocol version newer than the installed `ag-ui-protocol` is skipped with a warning rather than failing the whole request, per the backwards-compatibility policy in `pydantic_ai/ui/AGENTS.md`. Only items the installed models cannot dispatch at all are skipped: a body that is invalid for any other reason still raises, so a client bug isn’t converted into silent misbehavior. ##### Returns [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#returns-1) `RunAgentInput` #### dump\_messages [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.dump_messages) `@classmethod` def dump_messages( cls, messages: Sequence[ModelMessage], *, ag_ui_version: str = DEFAULT_AG_UI_VERSION, preserve_file_data: bool = False, ) -> list[Message] Transform Pydantic AI messages into AG-UI messages. Note: The round-trip `dump_messages` -> `load_messages` is not fully lossless: * `TextPart.id`, `.provider_name`, `.provider_details` are lost. * `ToolCallPart.id`, `.provider_name`, `.provider_details` are lost. * `ToolCallPart.args` and `NativeToolCallPart.args` that don’t parse as a JSON object are rewritten to `'{"INVALID_JSON":""}'` (see [`args_as_json_str`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BaseToolCallPart.args_as_json_str) ), so the raw string is no longer recoverable as args on reload. Unlike the live event stream, which emits them verbatim so streamed fragments stay concatenable, history has to hold a sendable value. * `NativeToolCallPart.id`, `.provider_details` are lost (only `.provider_name` survives via the prefixed tool call ID). * `NativeToolReturnPart.provider_details` is lost. * `tool_kind` is lost when `ag_ui_version < '0.1.11'` (before its `encrypted_value` carrier existed), so typed tool parts reload as their base classes. * `tool_kind` is not restored on error/denied tool returns (a typed return implies success to its readers), so those reload as plain `ToolReturnPart`. * A non-`'success'` `outcome` on a (native) tool return survives via the `encrypted_value` carrier from 0.1.11 (`ToolMessage` has no outcome slot). Below that, `'failed'` survives via `ToolMessage.error`, `'denied'` reloads as `'failed'`, and `'interrupted'` reloads as `'success'`. * `RetryPromptPart` becomes `ToolReturnPart` (or `UserPromptPart`) on reload. * A `NativeToolReturnPart` is always emitted directly after its `NativeToolCallPart`, so any part that originally sat between them — e.g. a `CompactionPart` — reloads after the pair instead. Provider adapters emit compaction parts outside call/return pairs, so this only affects hand-constructed histories. * `CachePoint` and `UploadedFile` content items are dropped (unless `preserve_file_data=True`). * `FileUrl.force_download` is dropped when `ag_ui_version < '0.1.15'` (before typed multimodal content gained a metadata carrier). * `ThinkingPart` is dropped when `ag_ui_version='0.1.10'`. * `FilePart` is silently dropped unless `preserve_file_data=True`. * `UploadedFile` in a multi-item `UserPromptPart` is split into a separate activity message when `preserve_file_data=True`, which reloads as a separate `UserPromptPart`. * `MultiModalContent` items in `ToolReturnPart`/`NativeToolReturnPart.content` always round-trip, regardless of `preserve_file_data`: the full content (files as base64/URL dicts) is serialized inline into the JSON `ToolMessage.content` and rehydrated on reload via the `ToolReturnContent` discriminator. The same serialization is used for both history (`dump_messages`) and the live event stream (`ToolCallResultEvent.content`), so files survive either round-trip. * Part ordering within a `ModelResponse` may change when text follows tool calls. ##### Returns [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#returns-2) [`list`](https://docs.python.org/3/glossary.html#term-list) \[`Message`\] — A list of AG-UI Message objects. ##### Parameters [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#parameters) **`messages`** : [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage)\ \] [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.dump_messages(messages)) A sequence of ModelMessage objects to convert. **`ag_ui_version`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) _Default:_ `DEFAULT_AG_UI_VERSION` [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.dump_messages(ag_ui_version)) AG-UI protocol version controlling `ThinkingPart` emission. **`preserve_file_data`** : [`bool`](https://docs.python.org/3/library/functions.html#bool) _Default:_ `False` [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.dump_messages(preserve_file_data)) Whether to include `FilePart` and `UploadedFile` items as `ActivityMessage`s. (Multimodal tool-return files always ride inline in `ToolMessage.content` and are unaffected.) #### from\_request [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.from_request) `@async` `@classmethod` def from_request( cls, request: Request, *, agent: AbstractAgent[AgentDepsT, OutputDataT], ag_ui_version: str = DEFAULT_AG_UI_VERSION, preserve_file_data: bool = False, manage_system_prompt: Literal['server', 'client'] = 'server', allowed_file_url_schemes: frozenset[str] = frozenset({'http', 'https'}), allowed_file_url_force_download: frozenset[ForceDownloadMode] = frozenset(), allow_uploaded_files: bool = False, **kwargs: Any, ) -> AGUIAdapter[AgentDepsT, OutputDataT] Extends [`from_request`](https://pydantic.dev/docs/ai/api/ui/base/#pydantic_ai.ui.UIAdapter.from_request) with AG-UI-specific parameters. ##### Returns [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#returns-3) `AGUIAdapter`\[`AgentDepsT`, `OutputDataT`\] #### load\_messages [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIAdapter.load_messages) `@classmethod` def load_messages( cls, messages: Sequence[Message], *, preserve_file_data: bool = False, ) -> list[ModelMessage] Transform AG-UI messages into Pydantic AI messages. ##### Returns [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#returns-4) [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage)\ \] AGUIEventStream --------------- [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIEventStream) **Bases:** `UIEventStream[RunAgentInput, BaseEvent, AgentDepsT, OutputDataT]` UI event stream transformer for the Agent-User Interaction (AG-UI) protocol. ### Methods [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#methods-1) #### handle\_event [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.AGUIEventStream.handle_event) `@async` def handle_event(event: NativeEvent) -> AsyncIterator[BaseEvent] Override to set timestamps on all AG-UI events. ##### Returns [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#returns-5) [`AsyncIterator`](https://docs.python.org/3/library/typing.html#typing.AsyncIterator) \[`BaseEvent`\] DEFAULT\_AG\_UI\_VERSION ------------------------ [](https://pydantic.dev/docs/ai/api/ui/ag_ui/#pydantic_ai.ui.ag_ui.DEFAULT_AG_UI_VERSION) The default AG-UI version, auto-detected from the installed `ag-ui-protocol` package. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `detect_ag_ui_version()` Was this page helpful? Thanks for your feedback! --- # Pydantic AI Gateway | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/overview/gateway/#_top) Pydantic AI Gateway =================== **[Pydantic AI Gateway](https://pydantic.dev/ai-gateway) ** is a unified interface for accessing multiple AI providers with a single key, managed through [Pydantic Logfire](https://pydantic.dev/logfire) . Features include built-in OpenTelemetry observability, real-time cost monitoring, failover management, and native integration with the other tools in the [Pydantic stack](https://pydantic.dev/) . Sign up at [logfire.pydantic.dev](https://logfire.pydantic.dev/) . Documentation Integration ------------------------- [](https://pydantic.dev/docs/ai/overview/gateway/#documentation-integration) To help you get started with Pydantic AI Gateway, some code examples on the Pydantic AI documentation include a “Via Pydantic AI Gateway” tab, alongside a “Direct to Provider API” tab with the standard Pydantic AI model string. The main difference between them is that when using Gateway, model strings use the `gateway/` prefix. Key features ------------ [](https://pydantic.dev/docs/ai/overview/gateway/#key-features) * **API key management**: Access multiple LLM providers with a single Gateway key. * **Cost Limits**: Set spending limits at project, user, and API key levels with daily, weekly, and monthly caps. * **BYOK and managed providers:** Bring your own API keys (BYOK) from LLM providers, or pay for inference directly through the platform. * **Multi-provider support:** Access models from OpenAI, Anthropic, Google Vertex, Groq, and AWS Bedrock. _More providers coming soon_. * **Routing groups:** Configure [routing groups](https://pydantic.dev/docs/ai/overview/gateway/#routing-groups) to fail over between providers serving the same model, or load-balance traffic across them by weight. * **Backend observability:** Log every request through [Pydantic Logfire](https://pydantic.dev/logfire) or any OpenTelemetry backend (_coming soon_). * **Zero translation**: Unlike traditional AI gateways that translate everything to one common schema, **Pydantic AI Gateway** allows requests to flow through directly in each provider’s native format. This gives you immediate access to new model features as soon as they are released. * **Enterprise ready**: Inherits Logfire’s enterprise features — including SSO, custom roles and permissions. hello\_world.py from pydantic_ai import Agent agent = Agent('gateway/openai:gpt-5.2') result = agent.run_sync('Where does "hello world" come from?') print(result.output) """ The first known use of "hello, world" was in a 1974 textbook about the C programming language. """ Quick Start ----------- [](https://pydantic.dev/docs/ai/overview/gateway/#quick-start) This section contains instructions on how to set up your account and run your app with Pydantic AI Gateway credentials. ### Create an account [](https://pydantic.dev/docs/ai/overview/gateway/#create-an-account) 1. Sign up at [logfire.pydantic.dev](https://logfire.pydantic.dev/) 2. Choose a region and create an account. 3. Activate the gateway in your organizations settings. ### Create Gateway API keys [](https://pydantic.dev/docs/ai/overview/gateway/#create-gateway-api-keys) Go to your organization’s Gateway settings in Logfire and create an API key. Usage ----- [](https://pydantic.dev/docs/ai/overview/gateway/#usage) After setting up your account with the instructions above, you will be able to make an AI model request with the Pydantic AI Gateway. The code snippets below show how you can use Pydantic AI Gateway with different frameworks and SDKs. To use different models, change the model string `gateway/:` to other models offered by the supported providers. Examples of providers and models that can be used are: | **Provider** | **API Format** | **Example Model** | | --- | --- | --- | | OpenAI | `openai` | `gateway/openai:gpt-5.2` | | Anthropic | `anthropic` | `gateway/anthropic:claude-sonnet-4-6` | | Google Cloud (formerly Vertex AI) | `google-cloud` | `gateway/google-cloud:gemini-3-flash-preview` | | Groq | `groq` | `gateway/groq:openai/gpt-oss-120b` | | AWS Bedrock | `bedrock` | `gateway/bedrock:amazon.nova-micro-v1:0` | ### Pydantic AI [](https://pydantic.dev/docs/ai/overview/gateway/#pydantic-ai) Before you start, make sure you are on version 1.16 or later of `pydantic-ai`. To update to the latest version run: * [uv](https://pydantic.dev/docs/ai/overview/gateway/#tab-panel-138) * [pip](https://pydantic.dev/docs/ai/overview/gateway/#tab-panel-139) Terminal uv sync -P pydantic-ai Terminal pip install -U pydantic-ai Set the `PYDANTIC_AI_GATEWAY_API_KEY` environment variable to your Gateway API key: Terminal export PYDANTIC_AI_GATEWAY_API_KEY="pylf_v..." You can access multiple models with the same API key, as shown in the code snippet below. hello\_world.py from pydantic_ai import Agent agent = Agent('gateway/openai:gpt-5.2') result = agent.run_sync('Where does "hello world" come from?') print(result.output) """ The first known use of "hello, world" was in a 1974 textbook about the C programming language. """ #### Passing API Key directly [](https://pydantic.dev/docs/ai/overview/gateway/#passing-api-key-directly) Pass your API key directly using the [`gateway_provider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.gateway.gateway_provider) : passing\_api\_key.py from pydantic_ai import Agent from pydantic_ai.models.openai import OpenAIChatModel from pydantic_ai.providers.gateway import gateway_provider provider = gateway_provider('openai', api_key='pylf_v...') model = OpenAIChatModel('gpt-5.2', provider=provider) agent = Agent(model) result = agent.run_sync('Where does "hello world" come from?') print(result.output) """ The first known use of "hello, world" was in a 1974 textbook about the C programming language. """ #### Using a different upstream provider [](https://pydantic.dev/docs/ai/overview/gateway/#using-a-different-upstream-provider) To use an alternate provider or routing group, you can specify it in the route parameter: routing\_via\_provider.py from pydantic_ai import Agent from pydantic_ai.models.openai import OpenAIChatModel from pydantic_ai.providers.gateway import gateway_provider provider = gateway_provider( 'openai', api_key='pylf_v...', route='builtin-openai' ) model = OpenAIChatModel('gpt-5.2', provider=provider) agent = Agent(model) result = agent.run_sync('Where does "hello world" come from?') print(result.output) """ The first known use of "hello, world" was in a 1974 textbook about the C programming language. """ ### Claude Code [](https://pydantic.dev/docs/ai/overview/gateway/#claude-code) Before you start, log out of Claude Code using `/logout`. Set your gateway credentials as environment variables, using the base URL that matches your Logfire region: * [US](https://pydantic.dev/docs/ai/overview/gateway/#tab-panel-140) * [EU](https://pydantic.dev/docs/ai/overview/gateway/#tab-panel-141) Terminal export ANTHROPIC_BASE_URL="https://gateway-us.pydantic.dev/proxy/anthropic" export ANTHROPIC_AUTH_TOKEN="YOUR_GATEWAY_API_KEY" Terminal export ANTHROPIC_BASE_URL="https://gateway-eu.pydantic.dev/proxy/anthropic" export ANTHROPIC_AUTH_TOKEN="YOUR_GATEWAY_API_KEY" Replace `YOUR_GATEWAY_API_KEY` with the API key from your Logfire organization’s Gateway settings. Launch Claude Code by typing `claude`. All requests will now route through the Pydantic AI Gateway. ### Codex [](https://pydantic.dev/docs/ai/overview/gateway/#codex) Codex uses the OpenAI Responses API, so it should use the Gateway’s `openai-responses` route. Set your gateway API key as an environment variable: Terminal export PYDANTIC_AI_GATEWAY_API_KEY="YOUR_GATEWAY_API_KEY" Then add the following configuration to `~/.codex/config.toml`, using the base URL that matches your Logfire region: * [US](https://pydantic.dev/docs/ai/overview/gateway/#tab-panel-142) * [EU](https://pydantic.dev/docs/ai/overview/gateway/#tab-panel-143) model = "gpt-5.4" model_provider = "pydantic_gateway" [model_providers.pydantic_gateway] name = "Pydantic AI Gateway" base_url = "https://gateway-us.pydantic.dev/proxy/openai-responses" env_key = "PYDANTIC_AI_GATEWAY_API_KEY" env_key_instructions = "Create a Gateway API key in your Logfire organization's Gateway settings." wire_api = "responses" model = "gpt-5.4" model_provider = "pydantic_gateway" [model_providers.pydantic_gateway] name = "Pydantic AI Gateway" base_url = "https://gateway-eu.pydantic.dev/proxy/openai-responses" env_key = "PYDANTIC_AI_GATEWAY_API_KEY" env_key_instructions = "Create a Gateway API key in your Logfire organization's Gateway settings." wire_api = "responses" For more details on configuring custom providers in Codex, see the [Codex custom model providers docs](https://developers.openai.com/codex/config-advanced#custom-model-providers) and the [Codex configuration reference](https://developers.openai.com/codex/config-reference/) . If you already have a `~/.codex/config.toml`, add the `[model_providers.pydantic_gateway]` block and update `model_provider` instead of replacing the whole file. Replace `gpt-5.4` with whichever OpenAI Responses model you want Codex to use. Launch Codex by typing `codex`. All requests will now route through the Pydantic AI Gateway. ### SDKs [](https://pydantic.dev/docs/ai/overview/gateway/#sdks) #### OpenAI SDK [](https://pydantic.dev/docs/ai/overview/gateway/#openai-sdk) Use the base URL that matches your Logfire region (`gateway-us` or `gateway-eu`). * [US](https://pydantic.dev/docs/ai/overview/gateway/#tab-panel-144) * [EU](https://pydantic.dev/docs/ai/overview/gateway/#tab-panel-145) openai\_sdk.py import openai client = openai.Client( base_url='https://gateway-us.pydantic.dev/proxy/chat/', api_key='pylf_v...', ) response = client.chat.completions.create( model='gpt-5.2', messages=[{'role': 'user', 'content': 'Hello world'}], ) print(response.choices[0].message.content) #> Hello user openai\_sdk.py import openai client = openai.Client( base_url='https://gateway-eu.pydantic.dev/proxy/chat/', api_key='pylf_v...', ) response = client.chat.completions.create( model='gpt-5.2', messages=[{'role': 'user', 'content': 'Hello world'}], ) print(response.choices[0].message.content) #> Hello user #### Anthropic SDK [](https://pydantic.dev/docs/ai/overview/gateway/#anthropic-sdk) Use the base URL that matches your Logfire region (`gateway-us` or `gateway-eu`). * [US](https://pydantic.dev/docs/ai/overview/gateway/#tab-panel-146) * [EU](https://pydantic.dev/docs/ai/overview/gateway/#tab-panel-147) anthropic\_sdk.py import anthropic client = anthropic.Anthropic( base_url='https://gateway-us.pydantic.dev/proxy/anthropic/', auth_token='pylf_v...', ) response = client.messages.create( max_tokens=1000, model='claude-sonnet-4-5', messages=[{'role': 'user', 'content': 'Hello world'}], ) print(response.content[0].text) #> Hello user anthropic\_sdk.py import anthropic client = anthropic.Anthropic( base_url='https://gateway-eu.pydantic.dev/proxy/anthropic/', auth_token='pylf_v...', ) response = client.messages.create( max_tokens=1000, model='claude-sonnet-4-5', messages=[{'role': 'user', 'content': 'Hello world'}], ) print(response.content[0].text) #> Hello user #### Vercel AI SDK [](https://pydantic.dev/docs/ai/overview/gateway/#vercel-ai-sdk) The [Vercel AI SDK](https://ai-sdk.dev/) can route through the Gateway by pointing each provider’s `baseURL` at the matching proxy path (e.g. `/proxy/openai` or `/proxy/anthropic`). Use the base URL that matches your Logfire region (`gateway-us` or `gateway-eu`). * [US](https://pydantic.dev/docs/ai/overview/gateway/#tab-panel-148) * [EU](https://pydantic.dev/docs/ai/overview/gateway/#tab-panel-149) import { createOpenAI } from "@ai-sdk/openai"; import { generateText } from "ai"; const apiKey = process.env.PYDANTIC_AI_GATEWAY_API_KEY; if (!apiKey) throw new Error("set PYDANTIC_AI_GATEWAY_API_KEY"); const openai = createOpenAI({ apiKey, baseURL: "https://gateway-us.pydantic.dev/proxy/openai", }); async function main() { const openaiResult = await generateText({ model: openai("gpt-5.2"), prompt: "what color is the sky? reply concisely", }); console.log("openai:", openaiResult.text); } main().catch((err) => { console.error(err); process.exit(1); }); import { createOpenAI } from "@ai-sdk/openai"; import { generateText } from "ai"; const apiKey = process.env.PYDANTIC_AI_GATEWAY_API_KEY; if (!apiKey) throw new Error("set PYDANTIC_AI_GATEWAY_API_KEY"); const openai = createOpenAI({ apiKey, baseURL: "https://gateway-eu.pydantic.dev/proxy/openai", }); async function main() { const openaiResult = await generateText({ model: openai("gpt-5.2"), prompt: "what color is the sky? reply concisely", }); console.log("openai:", openaiResult.text); } main().catch((err) => { console.error(err); process.exit(1); }); Routing groups -------------- [](https://pydantic.dev/docs/ai/overview/gateway/#routing-groups) A **routing group** is a named collection of providers that all serve the same model. Each member has a **priority**, a **weight**, and an **active** flag, and those three values together let a single group express two different routing strategies: * **Failover / fallback**: Assign members different priorities. The Gateway always tries the highest-priority active member first, and only falls through to a lower-priority member when the higher one is unavailable (for example if it is down, rate-limited, or returns an error). * **Load balancing**: Assign two or more members the same priority and give each a weight. The Gateway splits traffic across those members in proportion to their weights. The two strategies compose: you can have, for example, a top priority tier with two providers load-balanced 70/30, and a second priority tier that only receives traffic when both top-tier providers fail. ### Creating a routing group [](https://pydantic.dev/docs/ai/overview/gateway/#creating-a-routing-group) Routing groups are managed from your organization’s Gateway settings in Logfire: 1. Open **Gateway -> Routing Groups** and click **Add Routing Group**. 2. Give the group a slug (e.g. `anthropic-routing`) and an optional description. 3. Open the group’s **Members** page and add one or more providers. For each member set: * **Priority** - higher values are tried first. Use different priorities across members for failover. * **Weight** - load-balancing weight used between members that share the same priority. * **Active** - inactive members are skipped during routing. ### Using a routing group [](https://pydantic.dev/docs/ai/overview/gateway/#using-a-routing-group) Point the Gateway provider at the group via the `route` parameter (the group’s slug): routing\_group.py from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModel from pydantic_ai.providers.gateway import gateway_provider provider = gateway_provider( 'anthropic', api_key='pylf_v...', route='anthropic-routing', # (1) ) model = AnthropicModel('claude-sonnet-4-6', provider=provider) agent = Agent(model) result = agent.run_sync('Where does "hello world" come from?') print(result.output) """ The first known use of "hello, world" was in a 1974 textbook about the C programming language. """ The slug of the routing group you created in Logfire. Troubleshooting --------------- [](https://pydantic.dev/docs/ai/overview/gateway/#troubleshooting) ### Unable to calculate spend [](https://pydantic.dev/docs/ai/overview/gateway/#unable-to-calculate-spend) The gateway needs to know the cost of a request in order to provide spend insights and enforce spending limits. Each provider has a **Require pricing data** toggle in its settings. When enabled (the default), the gateway rejects requests for models it has no pricing data for before forwarding them upstream. When disabled, those requests are allowed through, but their cost will not be tracked and they will not count toward spending limits. The rejection response depends on the provider type: * **Built-in providers** (Pydantic-managed): `404` with a message asking you to let us know on Slack so we can add the model. * **Custom providers** (your own API keys): `400` indicating that pricing data is required, with a hint to disable the toggle if you want the request through anyway. We are actively working on supporting more providers and models. If there’s a specific provider or model you’d like to see supported, please let us know on [Slack](https://logfire.pydantic.dev/docs/join-slack/) or [open an issue on `genai-prices`](https://github.com/pydantic/genai-prices/issues/new) . Was this page helpful? Thanks for your feedback! --- # output | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#_top) output ====== ToolOutput ---------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.ToolOutput) **Bases:** `Generic[OutputDataT]` Marker class to use a tool for output and optionally customize the tool. Example: tool\_output.py from pydantic import BaseModel from pydantic_ai import Agent, ToolOutput class Fruit(BaseModel): name: str color: str class Vehicle(BaseModel): name: str wheels: int agent = Agent( 'openai:gpt-5.2', output_type=[\ ToolOutput(Fruit, name='return_fruit'),\ ToolOutput(Vehicle, name='return_vehicle'),\ ], ) result = agent.run_sync('What is a banana?') print(repr(result.output)) #> Fruit(name='banana', color='yellow') ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#attributes) #### description [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.ToolOutput.description) The description of the tool that will be passed to the model. If not specified, the docstring of the output type or function will be used. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `description` #### max\_retries [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.ToolOutput.max_retries) Per-tool retry limit for this output tool. Overrides the output side of the agent’s retry budget, which itself acts as the per-tool default for output tools that do not specify their own limit. If not set, the agent-level value is used. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `max_retries` #### name [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.ToolOutput.name) The name of the tool that will be passed to the model. If not specified and only one output is provided, `final_result` will be used. If multiple outputs are provided, the name of the output type or function will be added to the tool name. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `name` #### output [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.ToolOutput.output) An output type or function. **Type:** `OutputTypeOrFunction`\[`OutputDataT`\] **Default:** `type_` #### sequential [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.ToolOutput.sequential) Whether this output tool must run as a barrier, not overlapping with other tool calls. Only meaningful under `end_strategy='exhaustive'`, where tools otherwise run in parallel: a `sequential=True` output tool runs alone, so function tools the model emitted before it complete first. Under `'early'`/`'graceful'` output tools already run sequentially, so this has no effect. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `sequential` #### strict [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.ToolOutput.strict) Whether to use strict mode for the tool. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `strict` NativeOutput ------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.NativeOutput) **Bases:** `Generic[OutputDataT]` Marker class to use the model’s native structured outputs functionality for outputs and optionally customize the name and description. Example: native\_output.py from pydantic_ai import Agent, NativeOutput from tool_output import Fruit, Vehicle agent = Agent( 'openai:gpt-5.2', output_type=NativeOutput( [Fruit, Vehicle], name='Fruit or vehicle', description='Return a fruit or vehicle.' ), ) result = agent.run_sync('What is a Ford Explorer?') print(repr(result.output)) #> Vehicle(name='Ford Explorer', wheels=4) ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#attributes-1) #### description [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.NativeOutput.description) The description of the structured output that will be passed to the model. If not specified and only one output is provided, the docstring of the output type or function will be used. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `description` #### name [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.NativeOutput.name) The name of the structured output that will be passed to the model. If not specified and only one output is provided, the name of the output type or function will be used. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `name` #### outputs [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.NativeOutput.outputs) The output types or functions. **Type:** `OutputTypeOrFunction`\[`OutputDataT`\] | [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[`OutputTypeOrFunction`\[`OutputDataT`\]\] **Default:** `outputs` #### strict [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.NativeOutput.strict) Whether to use strict mode for the output, if the model supports it. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `strict` #### template [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.NativeOutput.template) Template for the prompt passed to the model. The ‘{schema}’ placeholder will be replaced with the output JSON schema. If no template is specified but the model’s profile indicates that it requires the schema to be sent as a prompt, the default template specified on the profile will be used. Set to `False` to disable the schema prompt entirely. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[[`False`](https://docs.python.org/3/library/constants.html#False)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `template` PromptedOutput -------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.PromptedOutput) **Bases:** `Generic[OutputDataT]` Marker class to use a prompt to tell the model what to output and optionally customize the prompt. Example: prompted\_output.py from pydantic import BaseModel from pydantic_ai import Agent, PromptedOutput from tool_output import Vehicle class Device(BaseModel): name: str kind: str agent = Agent( 'openai:gpt-5.2', output_type=PromptedOutput( [Vehicle, Device], name='Vehicle or device', description='Return a vehicle or device.' ), ) result = agent.run_sync('What is a MacBook?') print(repr(result.output)) #> Device(name='MacBook', kind='laptop') agent = Agent( 'openai:gpt-5.2', output_type=PromptedOutput( [Vehicle, Device], template='Gimme some JSON: {schema}' ), ) result = agent.run_sync('What is a Ford Explorer?') print(repr(result.output)) #> Vehicle(name='Ford Explorer', wheels=4) ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#attributes-2) #### description [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.PromptedOutput.description) The description that will be passed to the model. If not specified and only one output is provided, the docstring of the output type or function will be used. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `description` #### name [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.PromptedOutput.name) The name of the structured output that will be passed to the model. If not specified and only one output is provided, the name of the output type or function will be used. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `name` #### outputs [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.PromptedOutput.outputs) The output types or functions. **Type:** `OutputTypeOrFunction`\[`OutputDataT`\] | [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[`OutputTypeOrFunction`\[`OutputDataT`\]\] **Default:** `outputs` #### template [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.PromptedOutput.template) Template for the prompt passed to the model. The ‘{schema}’ placeholder will be replaced with the output JSON schema. If not specified, the default template specified on the model’s profile will be used. Set to `False` to disable the schema prompt entirely. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[[`False`](https://docs.python.org/3/library/constants.html#False)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `template` TextOutput ---------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.TextOutput) **Bases:** `Generic[OutputDataT]` Marker class to use text output for an output function taking a string argument. Example: from pydantic_ai import Agent, TextOutput def split_into_words(text: str) -> list[str]: return text.split() agent = Agent( 'openai:gpt-5.2', output_type=TextOutput(split_into_words), ) result = agent.run_sync('Who was Albert Einstein?') print(result.output) #> ['Albert', 'Einstein', 'was', 'a', 'German-born', 'theoretical', 'physicist.'] ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#attributes-3) #### output\_function [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.TextOutput.output_function) The function that will be called to process the model’s plain text output. The function must take a single string argument. **Type:** `TextOutputFunc`\[`OutputDataT`\] StructuredDict -------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.StructuredDict) def StructuredDict( json_schema: JsonSchemaValue, name: str | None = None, description: str | None = None, ) -> type[JsonSchemaValue] Returns a `dict[str, Any]` subclass with a JSON schema attached that will be used for structured output. Example: structured\_dict.py from pydantic_ai import Agent, StructuredDict schema = { 'type': 'object', 'properties': { 'name': {'type': 'string'}, 'age': {'type': 'integer'} }, 'required': ['name', 'age'] } agent = Agent('openai:gpt-5.2', output_type=StructuredDict(schema)) result = agent.run_sync('Create a person') print(result.output) #> {'name': 'John Doe', 'age': 30} ### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#returns) [`type`](https://docs.python.org/3/glossary.html#term-type) \[`JsonSchemaValue`\] ### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#parameters) **`json_schema`** : `JsonSchemaValue` [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.StructuredDict(json_schema)) A JSON schema of type `object` defining the structure of the dictionary content. **`name`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.StructuredDict(name)) Optional name of the structured output. If not provided, the `title` field of the JSON schema will be used if it’s present. **`description`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.StructuredDict(description)) Optional description of the structured output. If not provided, the `description` field of the JSON schema will be used if it’s present. OutputDataT ----------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.OutputDataT) Covariant type variable for the output data type of a run. **Default:** `TypeVar('OutputDataT', default=str, covariant=True)` Was this page helpful? Thanks for your feedback! --- # Joins & Reducers | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/graph/builder/joins/#_top) Joins & Reducers ================ Join nodes synchronize and aggregate data from parallel execution paths. They use **Reducers** to combine multiple inputs into a single output. When you use [parallel execution](https://pydantic.dev/docs/ai/graph/builder/parallel/) (broadcasting or mapping), you often need to collect and combine the results. Join nodes serve this purpose by: 1. Waiting for all parallel tasks to complete 2. Aggregating their outputs using a [`ReducerFunction`](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.ReducerFunction) 3. Passing the aggregated result to the next node Creating Joins -------------- [](https://pydantic.dev/docs/ai/graph/builder/joins/#creating-joins) Create a join using `GraphBuilder.join` with a reducer function and initial value or factory: basic\_join.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, reduce_list_append @dataclass class SimpleState: pass g = GraphBuilder(state_type=SimpleState, output_type=list[int]) @g.step async def generate_numbers(ctx: StepContext[SimpleState, None, None]) -> list[int]: return [1, 2, 3, 4, 5] @g.step async def square(ctx: StepContext[SimpleState, None, int]) -> int: return ctx.inputs * ctx.inputs # Create a join to collect all squared values collect = g.join(reduce_list_append, initial_factory=list[int]) g.add( g.edge_from(g.start_node).to(generate_numbers), g.edge_from(generate_numbers).map().to(square), g.edge_from(square).to(collect), g.edge_from(collect).to(g.end_node), ) graph = g.build() async def main(): result = await graph.run(state=SimpleState()) print(sorted(result)) #> [1, 4, 9, 16, 25] _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Built-in Reducers ----------------- [](https://pydantic.dev/docs/ai/graph/builder/joins/#built-in-reducers) Pydantic Graph provides several common reducer types out of the box: ### `reduce_list_append` [](https://pydantic.dev/docs/ai/graph/builder/joins/#reduce_list_append) [`reduce_list_append`](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.reduce_list_append) collects all inputs into a list: list\_reducer.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, reduce_list_append @dataclass class SimpleState: pass async def main(): g = GraphBuilder(state_type=SimpleState, output_type=list[str]) @g.step async def generate(ctx: StepContext[SimpleState, None, None]) -> list[int]: return [10, 20, 30] @g.step async def to_string(ctx: StepContext[SimpleState, None, int]) -> str: return f'value-{ctx.inputs}' collect = g.join(reduce_list_append, initial_factory=list[str]) g.add( g.edge_from(g.start_node).to(generate), g.edge_from(generate).map().to(to_string), g.edge_from(to_string).to(collect), g.edge_from(collect).to(g.end_node), ) graph = g.build() result = await graph.run(state=SimpleState()) print(sorted(result)) #> ['value-10', 'value-20', 'value-30'] _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ ### `reduce_list_extend` [](https://pydantic.dev/docs/ai/graph/builder/joins/#reduce_list_extend) [`reduce_list_extend`](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.reduce_list_extend) extends a list with an iterable of items: list\_extend\_reducer.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, reduce_list_extend @dataclass class SimpleState: pass async def main(): g = GraphBuilder(state_type=SimpleState, output_type=list[int]) @g.step async def generate(ctx: StepContext[SimpleState, None, None]) -> list[int]: return [1, 2, 3] @g.step async def create_range(ctx: StepContext[SimpleState, None, int]) -> list[int]: """Create a range from 0 to the input value.""" return list(range(ctx.inputs)) collect = g.join(reduce_list_extend, initial_factory=list[int]) g.add( g.edge_from(g.start_node).to(generate), g.edge_from(generate).map().to(create_range), g.edge_from(create_range).to(collect), g.edge_from(collect).to(g.end_node), ) graph = g.build() result = await graph.run(state=SimpleState()) print(sorted(result)) #> [0, 0, 0, 1, 1, 2] _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ ### `reduce_dict_update` [](https://pydantic.dev/docs/ai/graph/builder/joins/#reduce_dict_update) [`reduce_dict_update`](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.reduce_dict_update) merges dictionaries together: dict\_reducer.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, reduce_dict_update @dataclass class SimpleState: pass async def main(): g = GraphBuilder(state_type=SimpleState, output_type=dict[str, int]) @g.step async def generate_keys(ctx: StepContext[SimpleState, None, None]) -> list[str]: return ['apple', 'banana', 'cherry'] @g.step async def create_entry(ctx: StepContext[SimpleState, None, str]) -> dict[str, int]: return {ctx.inputs: len(ctx.inputs)} merge = g.join(reduce_dict_update, initial_factory=dict[str, int]) g.add( g.edge_from(g.start_node).to(generate_keys), g.edge_from(generate_keys).map().to(create_entry), g.edge_from(create_entry).to(merge), g.edge_from(merge).to(g.end_node), ) graph = g.build() result = await graph.run(state=SimpleState()) result = {k: result[k] for k in sorted(result)} # force deterministic ordering print(result) #> {'apple': 5, 'banana': 6, 'cherry': 6} _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ ### `reduce_null` [](https://pydantic.dev/docs/ai/graph/builder/joins/#reduce_null) [`reduce_null`](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.reduce_null) discards all inputs and returns `None`. Useful when you only care about side effects: null\_reducer.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, reduce_null @dataclass class CounterState: total: int = 0 async def main(): g = GraphBuilder(state_type=CounterState, output_type=int) @g.step async def generate(ctx: StepContext[CounterState, None, None]) -> list[int]: return [1, 2, 3, 4, 5] @g.step async def accumulate(ctx: StepContext[CounterState, None, int]) -> int: ctx.state.total += ctx.inputs return ctx.inputs # We don't care about the outputs, only the side effect on state ignore = g.join(reduce_null, initial=None) @g.step async def get_total(ctx: StepContext[CounterState, None, None]) -> int: return ctx.state.total g.add( g.edge_from(g.start_node).to(generate), g.edge_from(generate).map().to(accumulate), g.edge_from(accumulate).to(ignore), g.edge_from(ignore).to(get_total), g.edge_from(get_total).to(g.end_node), ) graph = g.build() state = CounterState() result = await graph.run(state=state) print(result) #> 15 _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ ### `reduce_sum` [](https://pydantic.dev/docs/ai/graph/builder/joins/#reduce_sum) [`reduce_sum`](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.reduce_sum) sums numeric values: sum\_reducer.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, StepContext, reduce_sum @dataclass class SimpleState: pass async def main(): g = GraphBuilder(state_type=SimpleState, output_type=int) @g.step async def generate(ctx: StepContext[SimpleState, None, None]) -> list[int]: return [10, 20, 30, 40] @g.step async def identity(ctx: StepContext[SimpleState, None, int]) -> int: return ctx.inputs sum_join = g.join(reduce_sum, initial=0) g.add( g.edge_from(g.start_node).to(generate), g.edge_from(generate).map().to(identity), g.edge_from(identity).to(sum_join), g.edge_from(sum_join).to(g.end_node), ) graph = g.build() result = await graph.run(state=SimpleState()) print(result) #> 100 _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ ### `ReduceFirstValue` [](https://pydantic.dev/docs/ai/graph/builder/joins/#reducefirstvalue) [`ReduceFirstValue`](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.ReduceFirstValue) returns the first value it receives and cancels all other parallel tasks. This is useful for “race” scenarios where you want the first successful result: first\_value\_reducer.py import asyncio from dataclasses import dataclass from pydantic_graph import GraphBuilder, ReduceFirstValue, StepContext @dataclass class SimpleState: tasks_completed: int = 0 async def main(): g = GraphBuilder(state_type=SimpleState, output_type=str) @g.step async def generate(ctx: StepContext[SimpleState, None, None]) -> list[int]: return [1, 12, 13, 14, 15] @g.step async def slow_process(ctx: StepContext[SimpleState, None, int]) -> str: """Simulate variable processing times.""" # Simulate different delays await asyncio.sleep(ctx.inputs * 0.1) ctx.state.tasks_completed += 1 return f'Result from task {ctx.inputs}' # Use ReduceFirstValue to get the first result and cancel the rest first_result = g.join(ReduceFirstValue[str](), initial=None, node_id='first_result') g.add( g.edge_from(g.start_node).to(generate), g.edge_from(generate).map().to(slow_process), g.edge_from(slow_process).to(first_result), g.edge_from(first_result).to(g.end_node), ) graph = g.build() state = SimpleState() result = await graph.run(state=state) print(result) #> Result from task 1 print(f'Tasks completed: {state.tasks_completed}') #> Tasks completed: ... _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Custom Reducers --------------- [](https://pydantic.dev/docs/ai/graph/builder/joins/#custom-reducers) Create custom reducers by defining a [`ReducerFunction`](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.ReducerFunction) : custom\_reducer.py from pydantic_graph import GraphBuilder, StepContext def reduce_sum(current: int, inputs: int) -> int: """A reducer that sums numbers.""" return current + inputs async def main(): g = GraphBuilder(output_type=int) @g.step async def generate(ctx: StepContext[None, None, None]) -> list[int]: return [5, 10, 15, 20] @g.step async def identity(ctx: StepContext[None, None, int]) -> int: return ctx.inputs sum_join = g.join(reduce_sum, initial=0) g.add( g.edge_from(g.start_node).to(generate), g.edge_from(generate).map().to(identity), g.edge_from(identity).to(sum_join), g.edge_from(sum_join).to(g.end_node), ) graph = g.build() result = await graph.run() print(result) #> 50 _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Reducers with State Access -------------------------- [](https://pydantic.dev/docs/ai/graph/builder/joins/#reducers-with-state-access) Reducers can access and modify the graph state: stateful\_reducer.py from dataclasses import dataclass from pydantic_graph import GraphBuilder, ReducerContext, StepContext @dataclass class MetricsState: total_count: int = 0 total_sum: int = 0 @dataclass class ReducedMetrics: count: int = 0 sum: int = 0 def reduce_metrics_sum(ctx: ReducerContext[MetricsState, None], current: ReducedMetrics, inputs: int) -> ReducedMetrics: ctx.state.total_count += 1 ctx.state.total_sum += inputs return ReducedMetrics(count=current.count + 1, sum=current.sum + inputs) def reduce_metrics_max(current: ReducedMetrics, inputs: ReducedMetrics) -> ReducedMetrics: return ReducedMetrics(count=max(current.count, inputs.count), sum=max(current.sum, inputs.sum)) async def main(): g = GraphBuilder(state_type=MetricsState, output_type=dict[str, int]) @g.step async def generate(ctx: StepContext[object, None, None]) -> list[int]: return [1, 3, 5, 7, 9, 10, 20, 30, 40] @g.step async def process_even(ctx: StepContext[MetricsState, None, int]) -> int: return ctx.inputs * 2 @g.step async def process_odd(ctx: StepContext[MetricsState, None, int]) -> int: return ctx.inputs * 3 metrics_even = g.join(reduce_metrics_sum, initial_factory=ReducedMetrics, node_id='metrics_even') metrics_odd = g.join(reduce_metrics_sum, initial_factory=ReducedMetrics, node_id='metrics_odd') metrics_max = g.join(reduce_metrics_max, initial_factory=ReducedMetrics, node_id='metrics_max') g.add( g.edge_from(g.start_node).to(generate), # Send even and odd numbers to their respective `process` steps g.edge_from(generate).map().to( g.decision() .branch(g.match(int, matches=lambda x: x % 2 == 0).label('even').to(process_even)) .branch(g.match(int, matches=lambda x: x % 2 == 1).label('odd').to(process_odd)) ), # Reduce metrics for even and odd numbers separately g.edge_from(process_even).to(metrics_even), g.edge_from(process_odd).to(metrics_odd), # Aggregate the max values for each field g.edge_from(metrics_even).to(metrics_max), g.edge_from(metrics_odd).to(metrics_max), # Finish the graph run with the final reduced value g.edge_from(metrics_max).to(g.end_node), ) graph = g.build() state = MetricsState() result = await graph.run(state=state) print(f'Result: {result}') #> Result: ReducedMetrics(count=5, sum=200) print(f'State total_count: {state.total_count}') #> State total_count: 9 print(f'State total_sum: {state.total_sum}') #> State total_sum: 275 _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ ### Canceling Sibling Tasks [](https://pydantic.dev/docs/ai/graph/builder/joins/#canceling-sibling-tasks) Reducers with access to [`ReducerContext`](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.ReducerContext) can call [`ctx.cancel_sibling_tasks()`](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join.ReducerContext.cancel_sibling_tasks) to cancel all other parallel tasks in the same fork. This is useful for early termination when you’ve found what you need: cancel\_siblings.py import asyncio from dataclasses import dataclass from pydantic_graph import GraphBuilder, ReducerContext, StepContext @dataclass class SearchState: searches_completed: int = 0 def reduce_find_match(ctx: ReducerContext[SearchState, None], current: str | None, inputs: str) -> str | None: """Return the first input that contains 'target' and cancel remaining tasks.""" if current is not None: # We already found a match, ignore subsequent inputs return current if 'target' in inputs: # Found a match! Cancel all other parallel tasks ctx.cancel_sibling_tasks() return inputs return None async def main(): g = GraphBuilder(state_type=SearchState, output_type=str | None) @g.step async def generate_searches(ctx: StepContext[SearchState, None, None]) -> list[str]: return ['item1', 'item2', 'target_item', 'item4', 'item5'] @g.step async def search(ctx: StepContext[SearchState, None, str]) -> str: """Simulate a slow search operation.""" # make the search artificially slower for 'item4' and 'item5' search_duration = 0.1 if ctx.inputs not in {'item4', 'item5'} else 1.0 await asyncio.sleep(search_duration) ctx.state.searches_completed += 1 return ctx.inputs find_match = g.join(reduce_find_match, initial=None) g.add( g.edge_from(g.start_node).to(generate_searches), g.edge_from(generate_searches).map().to(search), g.edge_from(search).to(find_match), g.edge_from(find_match).to(g.end_node), ) graph = g.build() state = SearchState() result = await graph.run(state=state) print(f'Found: {result}') #> Found: target_item print(f'Searches completed: {state.searches_completed}') #> Searches completed: 3 _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Note that only 3 searches completed instead of all 5, because the reducer canceled the remaining tasks after finding a match. Multiple Joins -------------- [](https://pydantic.dev/docs/ai/graph/builder/joins/#multiple-joins) A graph can have multiple independent joins: multiple\_joins.py from dataclasses import dataclass, field from pydantic_graph import GraphBuilder, StepContext, reduce_list_append @dataclass class MultiState: results: dict[str, list[int]] = field(default_factory=dict) async def main(): g = GraphBuilder(state_type=MultiState, output_type=dict[str, list[int]]) @g.step async def source_a(ctx: StepContext[MultiState, None, None]) -> list[int]: return [1, 2, 3] @g.step async def source_b(ctx: StepContext[MultiState, None, None]) -> list[int]: return [10, 20] @g.step async def process_a(ctx: StepContext[MultiState, None, int]) -> int: return ctx.inputs * 2 @g.step async def process_b(ctx: StepContext[MultiState, None, int]) -> int: return ctx.inputs * 3 join_a = g.join(reduce_list_append, initial_factory=list[int], node_id='join_a') join_b = g.join(reduce_list_append, initial_factory=list[int], node_id='join_b') @g.step async def store_a(ctx: StepContext[MultiState, None, list[int]]) -> None: ctx.state.results['a'] = ctx.inputs @g.step async def store_b(ctx: StepContext[MultiState, None, list[int]]) -> None: ctx.state.results['b'] = ctx.inputs @g.step async def combine(ctx: StepContext[MultiState, None, None]) -> dict[str, list[int]]: return ctx.state.results g.add( g.edge_from(g.start_node).to(source_a, source_b), g.edge_from(source_a).map().to(process_a), g.edge_from(source_b).map().to(process_b), g.edge_from(process_a).to(join_a), g.edge_from(process_b).to(join_b), g.edge_from(join_a).to(store_a), g.edge_from(join_b).to(store_b), g.edge_from(store_a, store_b).to(combine), g.edge_from(combine).to(g.end_node), ) graph = g.build() state = MultiState() result = await graph.run(state=state) print(f"Group A: {sorted(result['a'])}") #> Group A: [2, 4, 6] print(f"Group B: {sorted(result['b'])}") #> Group B: [30, 60] _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ Customizing Join Nodes ---------------------- [](https://pydantic.dev/docs/ai/graph/builder/joins/#customizing-join-nodes) ### Custom Node IDs [](https://pydantic.dev/docs/ai/graph/builder/joins/#custom-node-ids) Like steps, joins can have custom IDs: join\_custom\_id.py from pydantic_graph import reduce_list_append from basic_join import g my_join = g.join(reduce_list_append, initial_factory=list[int], node_id='my_custom_join_id') How Joins Work -------------- [](https://pydantic.dev/docs/ai/graph/builder/joins/#how-joins-work) Internally, the graph tracks which “fork” each parallel task belongs to. A join: 1. Identifies its parent fork (the fork that created the parallel paths) 2. Waits for all tasks from that fork to reach the join 3. Calls `reduce()` for each incoming value 4. Calls `finalize()` once all values are received 5. Passes the finalized result to downstream nodes This ensures proper synchronization even with nested parallel operations. Next Steps ---------- [](https://pydantic.dev/docs/ai/graph/builder/joins/#next-steps) * Learn about [parallel execution](https://pydantic.dev/docs/ai/graph/builder/parallel/) with broadcasting and mapping * Explore [conditional branching](https://pydantic.dev/docs/ai/graph/builder/decisions/) with decision nodes * See the [API reference](https://pydantic.dev/docs/ai/api/pydantic_graph/join/#pydantic_graph.join) for complete reducer documentation Was this page helpful? Thanks for your feedback! --- # Runtime Capability Creation | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/capability-creation/#_top) Runtime Capability Creation =========================== Runtime capability creation lets an agent write, validate, and persist Pydantic AI capabilities during one run for activation on the next. `CapabilityCreation` exposes tools that let the model write an `AbstractCapability` subclass to disk as Python source and validate it immediately; the orchestrator loads active authored capabilities into the next `agent.run(...)`. `CapabilityCreation` is for capabilities authored by the agent as Python source. For capabilities written or selected by application code, see [Building Custom Capabilities](https://pydantic.dev/docs/ai/capabilities/custom/) . [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/capability_creation/) > The API may change between releases. Where practical, breaking changes ship with a deprecation warning. The problem ----------- [](https://pydantic.dev/docs/ai/harness/capability-creation/#the-problem) A coding agent often discovers, mid-task, that it wants a behavior its host does not yet have: a guardrail, an extra instruction, a tool, a request hook. The capability surface to express that already exists — but normally only a developer can write a capability class, wire it into the agent, and restart. Without runtime capability creation, the agent cannot author that extension during a run and make it available to the next run. The solution ------------ [](https://pydantic.dev/docs/ai/harness/capability-creation/#the-solution) `CapabilityCreation` exposes three tools: * `author_capability(name, code)` — write `code` to `/.py`, import it, and validate it. Validation requires exactly one `pydantic_ai.capabilities.AbstractCapability` subclass that constructs with no arguments; the side-effect-free static getters (`get_instructions`, `get_toolset`, `get_native_tools`, `get_model_settings`, `get_serialization_name`) are exercised. The async lifecycle hooks are not run — they need a live `RunContext`. * `list_authored_capabilities()` — list authored capabilities with their status and any validation error. * `disable_authored_capability(name)` — stop a capability from being injected on the next run. A “hook” is not a standalone object in pydantic-ai — it is a method on a capability. So authoring a hook means authoring a capability that overrides one lifecycle method. A single overridden hook is a valid capability. Usage ----- [](https://pydantic.dev/docs/ai/harness/capability-creation/#usage) Construct `CapabilityCreation` with a `directory` for the authored files, then add it to the agent’s `capabilities`: from pathlib import Path from pydantic_ai import Agent from pydantic_ai_harness.capability_creation import CapabilityCreation creation = CapabilityCreation(directory=Path('.authored')) agent = Agent('anthropic:claude-sonnet-4-6', capabilities=[creation]) The agent can now call `author_capability`, `list_authored_capabilities`, and `disable_authored_capability`. `CapabilityCreation` also contributes static, cache-stable system-prompt guidance explaining these tools. Leave `guidance=None` for the default text, or pass your own string; set `guidance=''` to omit it entirely. Activation boundary ------------------- [](https://pydantic.dev/docs/ai/harness/capability-creation/#activation-boundary) Writing and validation happen in the current run; activation happens on the next run. A capability **cannot** be added to a live, already-executing run. pydantic-ai resolves the effective capability set once at the start of each run (the run’s root capability is fixed; there is no setter). So an authored capability is live on the **next** `agent.run(...)`, not the run that authored it. Authoring writes and validates the capability immediately, but its tools and hooks only exist once the next run’s toolset and capability chain are assembled at run start. ### Integration contract [](https://pydantic.dev/docs/ai/harness/capability-creation/#integration-contract) The orchestrator drives the loop, so it owns the one-line contract: thread the store’s active capabilities into each run via `agent.run(..., capabilities=...)`. With that in place, the authored capability is live on the very next loop iteration — no process restart: from pathlib import Path from pydantic_ai import Agent from pydantic_ai_harness.capability_creation import CapabilityCreation creation = CapabilityCreation(directory=Path('.authored')) agent = Agent('anthropic:claude-sonnet-4-6', capabilities=[creation]) history = None done = False next_prompt = 'Start the task.' while not done: extra = creation.store.load_active() result = await agent.run(next_prompt, message_history=history, capabilities=extra) history = result.all_messages() # ... decide `next_prompt` and `done` from `result` ... `creation.store` is the disk-backed `CapabilityStore` over the same `directory`. `store.load_active()` re-imports and re-constructs every active authored capability for injection into the next run. Entries that fail to load (corrupt source, construction error) are skipped, not raised, so one bad capability never blocks the rest. Persistence ----------- [](https://pydantic.dev/docs/ai/harness/capability-creation/#persistence) Authored capabilities persist to disk: each is one `/.py` file, indexed by a sibling `manifest.json`. A fresh process picks them up by constructing a new `CapabilityCreation` over the same `directory` and calling `store.load_active()`. `manifest.json` records each capability’s name, module file, class name, status (`active` or `disabled`), and last validation error. That is the surface a UI can read to show what the agent has authored. The manifest is written atomically (temp file plus `os.replace`), so a crash mid-write never leaves a partial file that reads back as “no capabilities”. Capability names must be lowercase letters, digits, and underscores, starting with a letter. Reusing a name replaces the previous capability of that name. A code that imports but fails validation is still written to disk (so it can be inspected) and recorded with its `last_error` set; `load_active()` skips it. Trust boundary -------------- [](https://pydantic.dev/docs/ai/harness/capability-creation/#trust-boundary) `CapabilityCreation` executes arbitrary Python in-process at import, construction, and run time. That is the same trust boundary an agent that already runs shell commands and edits files operates under, which is the deliberate choice here. Do not point it at a directory whose contents you would not run yourself, and treat authored capabilities as code the agent is executing on your host. Because authored capabilities hold live code, they are not spec-serializable (`get_serialization_name()` returns `None`) and are persisted as source rather than as an [agent spec](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) . Typing ------ [](https://pydantic.dev/docs/ai/harness/capability-creation/#typing) Imported authored code is dynamic, but nothing typed `Any` crosses back into the harness: every value pulled from an authored module is narrowed with `isinstance`/`issubclass` before use, and loaded instances are typed `AbstractCapability[object]`. Because `AgentDepsT` is contravariant, an `AbstractCapability[object]` is accepted by any agent’s `capabilities=` parameter. API reference ------------- [](https://pydantic.dev/docs/ai/harness/capability-creation/#api-reference) CapabilityCreation ------------------ [](https://pydantic.dev/docs/ai/harness/capability-creation/#pydantic_ai_harness.capability_creation.CapabilityCreation) **Bases:** `AbstractCapability[AgentDepsT]` Create Pydantic AI capabilities during one run for activation on the next. Exposes `author_capability(name, code)`, `list_authored_capabilities`, and `disable_authored_capability`. Authoring writes Python source to `directory`, imports it, and validates it (exactly one `AbstractCapability` subclass that constructs with no arguments and whose static getters run). Authored capabilities hold live code, so they are not spec-serializable and are persisted as source rather than as a spec. Activation boundary: a capability cannot be added to a live, already-executing run — pydantic-ai resolves the capability set once at the start of each run. The authored capability becomes usable on the next `agent.run(...)`. The integration contract is one line on the orchestrator side: thread the store’s active capabilities into the next run. from pathlib import Path from pydantic_ai import Agent from pydantic_ai_harness.capability_creation import CapabilityCreation creation = CapabilityCreation(directory=Path('.authored')) agent = Agent('anthropic:claude-sonnet-4-6', capabilities=[creation]) # Loop: each iteration injects whatever the agent has authored so far. result = await agent.run('build a logging capability', capabilities=creation.store.load_active()) This executes authored Python in-process — the same trust boundary an agent that already runs shell commands and edits files operates under. ### Attributes [](https://pydantic.dev/docs/ai/harness/capability-creation/#attributes) #### directory [](https://pydantic.dev/docs/ai/harness/capability-creation/#pydantic_ai_harness.capability_creation.CapabilityCreation.directory) Directory holding the authored `.py` files and the `manifest.json` index. **Type:** `Path` #### guidance [](https://pydantic.dev/docs/ai/harness/capability-creation/#pydantic_ai_harness.capability_creation.CapabilityCreation.guidance) Static system-prompt guidance on authoring. Cache-stable. Leave `None` for the default, or set `''` to omit guidance entirely. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### store [](https://pydantic.dev/docs/ai/harness/capability-creation/#pydantic_ai_harness.capability_creation.CapabilityCreation.store) The disk-backed store. Call `store.load_active()` to inject authored capabilities into the next run. **Type:** `CapabilityStore` ### Methods [](https://pydantic.dev/docs/ai/harness/capability-creation/#methods) #### get\_instructions [](https://pydantic.dev/docs/ai/harness/capability-creation/#pydantic_ai_harness.capability_creation.CapabilityCreation.get_instructions) def get_instructions() -> AgentInstructions[AgentDepsT] | None Static, cache-stable guidance on the authoring tools. ##### Returns [](https://pydantic.dev/docs/ai/harness/capability-creation/#returns) `AgentInstructions`\[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) #### get\_serialization\_name [](https://pydantic.dev/docs/ai/harness/capability-creation/#pydantic_ai_harness.capability_creation.CapabilityCreation.get_serialization_name) `@classmethod` def get_serialization_name(cls) -> str | None Not spec-serializable: the capability holds a live, disk-backed store. ##### Returns [](https://pydantic.dev/docs/ai/harness/capability-creation/#returns-1) [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### get\_toolset [](https://pydantic.dev/docs/ai/harness/capability-creation/#pydantic_ai_harness.capability_creation.CapabilityCreation.get_toolset) def get_toolset() -> AgentToolset[AgentDepsT] | None Toolset providing the authoring tools over this capability’s store. ##### Returns [](https://pydantic.dev/docs/ai/harness/capability-creation/#returns-2) [`AgentToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AgentToolset) \[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) CapabilityStore --------------- [](https://pydantic.dev/docs/ai/harness/capability-creation/#pydantic_ai_harness.capability_creation.CapabilityStore) Read/write index of authored capability `.py` files under `directory`. ### Methods [](https://pydantic.dev/docs/ai/harness/capability-creation/#methods-1) #### disable [](https://pydantic.dev/docs/ai/harness/capability-creation/#pydantic_ai_harness.capability_creation.CapabilityStore.disable) def disable(name: str) -> bool Mark the named capability disabled so `load_active` stops returning it. Returns whether it existed. ##### Returns [](https://pydantic.dev/docs/ai/harness/capability-creation/#returns-3) [`bool`](https://docs.python.org/3/library/functions.html#bool) #### list\_all [](https://pydantic.dev/docs/ai/harness/capability-creation/#pydantic_ai_harness.capability_creation.CapabilityStore.list_all) def list_all() -> list[AuthoredCapability] Return every manifest entry, in insertion order. ##### Returns [](https://pydantic.dev/docs/ai/harness/capability-creation/#returns-4) [`list`](https://docs.python.org/3/glossary.html#term-list) \[`AuthoredCapability`\] #### load\_active [](https://pydantic.dev/docs/ai/harness/capability-creation/#pydantic_ai_harness.capability_creation.CapabilityStore.load_active) def load_active() -> list[AbstractCapability[object]] Construct every active authored capability for per-run injection. Re-imports and re-constructs each active entry. Entries that fail to load (corrupt source, construction error) are skipped, not raised, so one bad capability never blocks the rest. A load outcome that disagrees with the record’s `last_error` is persisted back to the manifest: a newly broken entry records its error, a re-fixed entry clears it, so the manifest stays truthful about which capabilities are actually active. ##### Returns [](https://pydantic.dev/docs/ai/harness/capability-creation/#returns-5) [`list`](https://docs.python.org/3/glossary.html#term-list) \[`AbstractCapability`\[[`object`](https://docs.python.org/3/glossary.html#term-object)\ \]\] #### write [](https://pydantic.dev/docs/ai/harness/capability-creation/#pydantic_ai_harness.capability_creation.CapabilityStore.write) def write(name: str, code: str) -> AuthoredCapability Write `code` to `.py`, validate it, and upsert the manifest entry. Raises `ValueError` for an invalid name (before writing anything). A code that imports but fails validation is still written (so it can be inspected) and recorded with `last_error` set; `load_active` skips it. ##### Returns [](https://pydantic.dev/docs/ai/harness/capability-creation/#returns-6) `AuthoredCapability` Was this page helpful? Thanks for your feedback! --- # Agent Specs | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#_top) Agent Specs =========== Agent specs let you define agents declaratively in YAML or JSON — [model](https://pydantic.dev/docs/ai/models/overview/) , [instructions](https://pydantic.dev/docs/ai/core-concepts/agent/#instructions) , [capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) , and all. One line to load, no Python agent construction code required. This is useful for: * Separating agent configuration from application code * Letting non-developers (prompt engineers, domain experts) configure agents * Storing agent definitions alongside other config files * Sharing agent configurations across teams or projects Defining a spec --------------- [](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#defining-a-spec) A spec file defines the agent’s configuration in YAML or JSON: agent.yaml model: anthropic:claude-opus-4-6 instructions: You are a helpful research assistant. model_settings: max_tokens: 8192 capabilities: - WebSearch: local: duckduckgo - Thinking: effort: high Loading specs ------------- [](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#loading-specs) [`Agent.from_file`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.from_file) loads a spec from a YAML or JSON file and constructs an agent: from\_file\_example.py from pydantic_ai import Agent agent = Agent.from_file('agent.yaml') [`Agent.from_spec`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.from_spec) accepts a dict or [`AgentSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentSpec) instance and supports additional keyword arguments that supplement or override the spec: from\_spec\_example.py from dataclasses import dataclass from pydantic_ai import Agent @dataclass class UserContext: user_name: str agent = Agent.from_spec( { 'model': 'anthropic:claude-opus-4-6', 'instructions': 'You are helping {{user_name}}.', 'capabilities': [{'WebSearch': {'local': 'duckduckgo'}}], }, deps_type=UserContext, ) Keyword arguments interact with spec fields as follows: * **Scalar fields** (`model`, `name`, `end_strategy`, etc.) — the keyword argument overrides the spec value when provided. For retry budgets, the `retries` keyword argument overrides the spec’s `retries` value. * **`instructions`** — merged: spec instructions come first, then keyword argument instructions. * **`capabilities`** — merged: spec capabilities come first, then keyword argument capabilities. * **`model_settings`** — merged additively: keyword argument settings override matching spec settings. * **`output_type`** — takes precedence over `output_schema` from the spec. When `deps_type` is passed, [template strings](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#template-strings) in the spec’s `instructions`, `description`, and capability arguments are compiled and validated against the deps type at construction time. For more control over spec loading, use [`AgentSpec.from_file`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentSpec.from_file) to load the spec separately before passing it to `Agent.from_spec`. Template strings ---------------- [](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#template-strings) [`TemplateStr`](https://pydantic.dev/docs/ai/api/pydantic-ai/template/#pydantic_ai.template.TemplateStr) provides Handlebars-style templates (`{{variable}}`) that are rendered against the agent’s [dependencies](https://pydantic.dev/docs/ai/core-concepts/dependencies/) at runtime. In spec files, strings containing `{{` are automatically converted to template strings: instructions: "You are assisting {{name}}, who is a {{role}}." Template variables are resolved from the fields of the `deps` object. When a `deps_type` (or [`deps_schema`](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#deps_schema) ) is provided, template variable names are validated at construction time. In Python code, [`TemplateStr`](https://pydantic.dev/docs/ai/api/pydantic-ai/template/#pydantic_ai.template.TemplateStr) can be used explicitly, but a callable with [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) is generally preferred for IDE autocomplete and type checking: template\_instructions.py from dataclasses import dataclass from pydantic_ai import Agent, TemplateStr @dataclass class UserProfile: name: str role: str agent = Agent( 'openai:gpt-5.2', deps_type=UserProfile, instructions=TemplateStr('You are assisting {{name}}, who is a {{role}}.'), ) result = agent.run_sync('hello', deps=UserProfile(name='Alice', role='engineer')) print(result.output) #> Hello! How can I help you today? Capability spec syntax ---------------------- [](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#capability-spec-syntax) Capabilities in specs support three forms: * `'MyCapability'` — no arguments, calls `MyCapability.from_spec()` * `{'MyCapability': value}` — single positional argument, calls `MyCapability.from_spec(value)` * `{'MyCapability': {key: value, ...}}` — keyword arguments, calls `MyCapability.from_spec(**kwargs)` Custom capabilities in specs ---------------------------- [](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#custom-capabilities-in-specs) See [Publishing capabilities](https://pydantic.dev/docs/ai/capabilities/custom/#publishing-capabilities) for how to make custom capabilities work with agent specs. `AgentSpec` reference --------------------- [](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#agentspec-reference) The [`AgentSpec`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentSpec) model represents the full spec structure: | Field | Type | Description | | --- | --- | --- | | `model` | `str` | [Model](https://pydantic.dev/docs/ai/models/overview/)
name (required) | | `name` | `str \| None` | Agent name | | `description` | `str \| None` | Agent description (supports [templates](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#template-strings)
) | | `instructions` | `str \| list[str] \| None` | [Instructions](https://pydantic.dev/docs/ai/core-concepts/agent/#instructions)
(supports [templates](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#template-strings)
) | | `model_settings` | `dict \| None` | [Model settings](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) | | `capabilities` | `list` | [Capabilities](https://pydantic.dev/docs/ai/capabilities/overview/)
(see [spec syntax](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#capability-spec-syntax)
) | | `deps_schema` | `dict \| None` | JSON Schema for [template string](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#template-strings)
validation (see below) | | `output_schema` | `dict \| None` | JSON Schema for [structured output](https://pydantic.dev/docs/ai/core-concepts/output/)
(see below) | | `retries` | `int \| AgentRetries \| None` | Retry budgets for [tools](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-retries)
and [output validation](https://pydantic.dev/docs/ai/core-concepts/output/#output-validator-functions)
. Pass an integer to use the same budget for both, or [`AgentRetries`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentRetries)
to configure them separately. | | `end_strategy` | `EndStrategy` | When to stop (`'early'`, `'graceful'`, or `'exhaustive'`) | | `tool_timeout` | `float \| None` | Default [tool](https://pydantic.dev/docs/ai/tools-toolsets/tools/)
timeout in seconds | | `instrument` | `bool \| None` | Enable [Logfire](https://pydantic.dev/docs/ai/integrations/logfire/)
instrumentation | | `metadata` | `dict \| None` | Agent [metadata](https://pydantic.dev/docs/ai/core-concepts/agent/#run-metadata) | ### `deps_schema` [](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#deps_schema) When loading a spec file without a Python `deps_type`, `deps_schema` provides a JSON Schema that validates [template string](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#template-strings) variable names at construction time. It does **not** validate the actual deps object at runtime — it only ensures that template variables like `{{user_name}}` correspond to properties defined in the schema. ### `output_schema` [](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#output_schema) When provided (and no `output_type` keyword argument is passed to `from_spec`), `output_schema` defines the structure the model should produce as its final output. Under the hood, it creates a [`StructuredDict`](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.StructuredDict) output type: the JSON Schema is sent to the model API so the model knows what structure to produce, and the response is returned as a `dict[str, Any]`. agent\_with\_schema.yaml model: anthropic:claude-opus-4-6 deps_schema: type: object properties: user_name: type: string required: [user_name] output_schema: type: object properties: answer: type: string confidence: type: number required: [answer, confidence] instructions: "You are helping {{user_name}}. Always include a confidence score." capabilities: - WebSearch: local: duckduckgo Saving specs ------------ [](https://pydantic.dev/docs/ai/core-concepts/agent-spec/#saving-specs) [`AgentSpec.to_file`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentSpec.to_file) saves a spec to YAML or JSON and optionally generates a companion JSON Schema file for editor autocompletion: save\_spec\_example.py from pydantic_ai import AgentSpec spec = AgentSpec( model='anthropic:claude-opus-4-6', instructions='You are a helpful assistant.', capabilities=[{'WebSearch': {'local': 'duckduckgo'}}], ) spec.to_file('agent.yaml') # Also generates ./agent_schema.json for editor autocompletion The generated JSON Schema file enables autocompletion and validation in editors that support the [YAML Language Server](https://github.com/redhat-developer/yaml-language-server) protocol. Pass `schema_path=None` to skip schema generation. Was this page helpful? Thanks for your feedback! --- # DBOS | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#_top) DBOS ==== [DBOS](https://www.dbos.dev/) is a lightweight [durable execution](https://docs.dbos.dev/architecture) library natively integrated with Pydantic AI. Durable Execution ----------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#durable-execution) DBOS workflows make your program **durable** by checkpointing its state in a database. If your program ever fails, when it restarts all your workflows will automatically resume from the last completed step. * **Workflows** must be deterministic and generally cannot include I/O. * **Steps** may perform I/O (network, disk, API calls). If a step fails, it restarts from the beginning. Every workflow input and step output is durably stored in the system database. When workflow execution fails, whether from crashes, network issues, or server restarts, DBOS leverages these checkpoints to recover workflows from their last completed step. DBOS **queues** provide durable, database-backed alternatives to systems like Celery or BullMQ, supporting features such as concurrency limits, rate limits, timeouts, and prioritization. See the [DBOS docs](https://docs.dbos.dev/architecture) for details. The diagram below shows the overall architecture of an agentic application in DBOS. DBOS runs fully in-process as a library. Functions remain normal Python functions but are checkpointed into a database (Postgres or SQLite). Clients (HTTP, RPC, Kafka, etc.) | v +------------------------------------------------------+ | Application Servers | | | | +----------------------------------------------+ | | | Pydantic AI + DBOS Libraries | | | | | | | | [ Workflows (Agent Run Loop) ] | | | | [ Steps (Tool, MCP, Model) ] | | | | [ Queues ] [ Cron Jobs ] [ Messaging ] | | | +----------------------------------------------+ | | | +------------------------------------------------------+ | v +------------------------------------------------------+ | Database | | (Stores workflow and step state, schedules tasks) | +------------------------------------------------------+ See the [DBOS documentation](https://docs.dbos.dev/architecture) for more information. Durable Agent ------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#durable-agent) Add durable execution to any [`Agent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent) by attaching the [`DBOSDurability`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.dbos.DBOSDurability) [capability](https://pydantic.dev/docs/ai/capabilities/overview/) . When the agent runs inside a DBOS workflow, the capability routes [model requests](https://pydantic.dev/docs/ai/models/overview/) and [MCP communication](https://pydantic.dev/docs/ai/mcp/client/) through DBOS steps. To make a run durable, call `agent.run()` inside a `@DBOS.workflow`. The agent stays a normal `Agent` everywhere — outside a DBOS workflow the capability is transparent, and the original agent, model, and MCP server can still be used as normal. Custom tool functions and event stream handlers registered on the agent directly or through another capability are **not automatically wrapped** by DBOS. An `event_stream_handler=` passed to `DBOSDurability` runs inside a DBOS step and receives live-streamed events. If they involve non-deterministic behavior or perform I/O, you should explicitly decorate them with `@DBOS.step`. Here is a simple but complete example of attaching durable execution to an agent. All it requires is to install Pydantic AI with the DBOS [open-source library](https://github.com/dbos-inc/dbos-transact-py) : * [pip](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#tab-panel-0) * [uv](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#tab-panel-1) Terminal pip install pydantic-ai[dbos] Terminal uv add pydantic-ai[dbos] Or if you’re using the slim package, you can install it with the `dbos` optional group: * [pip](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#tab-panel-2) * [uv](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#tab-panel-3) Terminal pip install pydantic-ai-slim[dbos] Terminal uv add pydantic-ai-slim[dbos] After that, run the following example code: dbos\_durability.py from dbos import DBOS, DBOSConfig from pydantic_ai import Agent from pydantic_ai.durable_exec.dbos import DBOSDurability dbos_config: DBOSConfig = { 'name': 'pydantic_dbos_agent', 'system_database_url': 'sqlite:///dbostest.sqlite', # (1) } DBOS(config=dbos_config) agent = Agent( 'openai:gpt-5.6-sol', instructions="You're an expert in geography.", name='geography', # (2) capabilities=[DBOSDurability()], # (3) ) @DBOS.workflow() # (4) async def answer(question: str) -> str: result = await agent.run(question) return result.output async def main(): DBOS.launch() answer_text = await answer('What is the capital of Mexico?') print(answer_text) #> Mexico City (Ciudad de México, CDMX) This example uses SQLite. Postgres is recommended for production. The agent's `name` is used to uniquely identify its workflows. Attach durability via `capabilities=[...]`. The capability routes model requests and MCP communication through DBOS steps when the agent runs inside a workflow. Because DBOS workflows must be registered before `DBOS.launch()`, the agent must also be constructed before calling `DBOS.launch()`. Wrap `agent.run()` in your own `@DBOS.workflow` to make the run durable. _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main())` to run `main`)_ Because the same agent works inside and outside a DBOS workflow, [`DBOSDurability`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.dbos.DBOSDurability) composes with all other [capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) without each needing a DBOS-specific wrapper variant. For more information on how to use DBOS in Python applications, see their [Python SDK guide](https://docs.dbos.dev/python/programming-guide) . ### Wrapper-agent path (deprecated) [](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#wrapper-agent-path-deprecated) Any agent can be wrapped in a [`DBOSAgent`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.dbos.DBOSAgent) to get a durable agent variant that routes model requests and MCP communication through DBOS steps: dbos\_agent.py from pydantic_ai import Agent from pydantic_ai.durable_exec.dbos import DBOSAgent agent = Agent('openai:gpt-5.6-sol', name='geography') dbos_agent = DBOSAgent(agent) # Use `dbos_agent` in place of `agent`. Migrating to the capability means attaching `DBOSDurability` and adding the workflow decorator that `DBOSAgent` used to apply for you: -dbos_agent = DBOSAgent(agent) -result = await dbos_agent.run(prompt) +agent = Agent(..., capabilities=[DBOSDurability()]) + +@DBOS.workflow() +async def answer(prompt: str) -> str: + result = await agent.run(prompt) + return result.output DBOS Integration Considerations ------------------------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#dbos-integration-considerations) When using DBOS with Pydantic AI agents, there are a few important considerations to ensure workflows and toolsets behave correctly. ### Agent and Toolset Requirements [](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#agent-and-toolset-requirements) Each agent instance must have a unique `name` so DBOS can correctly resume workflows after a failure or restart. Each [`MCPToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset) must have a unique [`id`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AbstractToolset.id) , as DBOS derives its step names and per-run tool-defs cache key from it. This field is normally optional, but is required when using DBOS. It should not be changed once the durable agent has been deployed to production, as this would break active workflows. [`DynamicToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.DynamicToolset) s, including those contributed by a [`DynamicCapability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.DynamicCapability) , are wrapped too and require a stable `id`. Tool discovery and calls run as `{name}__dynamic_toolset__{id}.get_tools` and `{name}__dynamic_toolset__{id}.call_tool` steps. The dynamic toolset is resolved and entered independently inside each step, so its I/O — including MCP communication — is checkpointed. For a `DynamicCapability`, DBOS reuses the capability resolved for the run inside those steps. The capability factory itself runs in workflow code and re-runs when a workflow recovers, so like all workflow code it must be deterministic given the run’s `deps`: construct the toolset in the factory and leave its I/O to the steps. A toolset contributed by a [capability](https://pydantic.dev/docs/ai/capabilities/overview/) — a [`Capability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Capability) with `tools=`, or a locally-running [`MCP`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.MCP) server — derives its `id` from the capability’s own [`id`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability.id) , so set `Capability(id='...', tools=[...])` or `MCP(id='...', url='...')`. An `MCP` resolves its `id` in precedence order: an explicit `id=`, then a `native=MCPServerTool(...)` id, then a slug derived from the server URL’s host and path. A bare non-URL local client (e.g. `MCP(local=Path(...))`) with none of these stays id-less and must be given an explicit `id` to be used here. Function tools and event stream handlers registered on the agent directly or through another capability are not automatically wrapped by DBOS. An `event_stream_handler=` passed to `DBOSDurability` runs inside a DBOS step and receives live-streamed events. For directly registered tools and handlers, you can decide how to integrate them: * Decorate with `@DBOS.step` if the function involves non-determinism or I/O. * Skip the decorator if durability isn’t needed, so you avoid the extra DB checkpoint write. * If the function needs to enqueue tasks or invoke other DBOS workflows, run it inside the agent’s main workflow (not as a step). Other than that, any agent and toolset will just work! ### Agent Run Context and Dependencies [](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#agent-run-context-and-dependencies) By default, DBOS checkpoints workflow inputs/outputs and step outputs into a database using [`pickle`](https://docs.python.org/3/library/pickle.html) . But you can optionally supply a [custom serializer](https://docs.dbos.dev/python/reference/contexts#custom-serialization) through DBOS configuration. This means you need to make sure the [dependencies](https://pydantic.dev/docs/ai/core-concepts/dependencies/) object provided to `Agent.run()` / `Agent.run_sync()`, and tool outputs can be serialized. You may also want to keep the inputs and outputs small (under ~2 MB). PostgreSQL and SQLite support up to 1 GB per field, but large objects may impact performance. ### Model Selection at Runtime [](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#model-selection-at-runtime) `Agent.run(model=...)` supports both model strings (like `'openai:gpt-5.6-sol'`) and model instances. A model instance can’t be serialized across the step boundary, and rebuilding one from its `model_id` string would build a _different_ model — the same model name on whatever provider the worker’s environment implies, so the request would go to another endpoint with other credentials. An instance that isn’t registered ahead of time is therefore rejected with a `UserError`. There are two ways to use a specific instance: pre-register it by passing a `models` dict to [`DBOSDurability`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.dbos.DBOSDurability) and reference it by key (or pass the registered instance), or pass a model-name string and build the instance inside the step with a [`ResolveModelId`](https://pydantic.dev/docs/ai/capabilities/resolve-model-id/) capability — the right choice when the model depends on the run’s `deps`, e.g. per-user credentials. Model-name strings themselves never need registering. The agent’s own model, set at construction, is always available as the default. To customize how a model string is built — a custom provider, or per-user credentials carried on the run’s `deps` — add a [`ResolveModelId`](https://pydantic.dev/docs/ai/capabilities/resolve-model-id/) capability before `DBOSDurability`: it gets first crack at every string, and the resolver runs again inside the step with the run’s actual `deps`, so it must be deterministic for a given `(model_id, deps)` and must not perform external I/O. ### Streaming [](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#streaming) `Agent.run_stream()` and `Agent.run_stream_events()` work inside a DBOS workflow, but their events are buffered rather than delivered in real time. The model stream runs inside the durable step, and its events are replayed to the workflow after the step completes. For handlers with I/O side effects, pass `event_stream_handler=` to [`DBOSDurability`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.dbos.DBOSDurability) . Model events are delivered live inside each model-request step, while each tool event is delivered in its own event-handler step. As with any DBOS step, a handler may run more than once if the workflow recovers before its step is checkpointed, so keep its side effects idempotent. Alternatively, register [`ProcessEventStream`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ProcessEventStream) . Its handler runs in workflow code and must be deterministic because it re-runs on workflow replay. Tool and final-output events arrive live, while the real captured model events are replayed after each model request completes. For examples, see the [streaming docs](https://pydantic.dev/docs/ai/core-concepts/agent/#streaming-all-events) . A durability `event_stream_handler=` and a separately registered `ProcessEventStream` are two distinct handlers, and each fires once. The durability handler receives live events inside the durable step, while `ProcessEventStream` sees the buffered replay in workflow code. A per-run handler passed to `Agent.run(event_stream_handler=...)` also runs workflow-side against replayed model events. Because the model stream is consumed inside the step, cancelling it from the workflow side (e.g. with [`AgentStream.cancel()`](https://pydantic.dev/docs/ai/api/pydantic-ai/result/#pydantic_ai.result.AgentStream.cancel) ) is not available across the durable boundary. [`CancellationToken`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.CancellationToken) cannot be passed to a DBOS durable run, and [`RunContext.cancel()`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.cancel) raises a clear [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) inside a step-wrapped unit (a dynamic or MCP tool, or an `event_stream_handler`) whose recorded result would replay without re-running on recovery. A plain function tool runs at workflow level under DBOS, where `cancel()` works and is replay-consistent. To stop a run from outside, cancel the DBOS workflow. `Agent.run_stream_sync()` is not for workflow code: it requires no running event loop and wraps `run_stream()`. Under [`DBOSDurability`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.dbos.DBOSDurability) , use the buffered async streaming APIs above or `Agent.run()` with an event stream handler. Outside a workflow, an agent with `DBOSDurability` behaves like a normal agent, so `run_stream_sync()` works as usual. (Wrapper `DBOSAgent` forbids `run_stream` inside workflows — use `run` + event stream handler there.) ### Suspended Turns and Background Mode [](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#suspended-turns-and-background-mode) When a provider pauses a model turn mid-flight (Anthropic `pause_turn`) or runs it as a server-side job that’s polled until it’s ready ([OpenAI background mode](https://pydantic.dev/docs/ai/models/openai/#background-mode) ), each segment runs in a separate model request step. The suspended [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) and background job ID are checkpointed between segments, while the final response is merged and usage is recorded once. A [`message_history`](https://pydantic.dev/docs/ai/core-concepts/message-history/) ending in a suspended response is passed to the first step. Size step timeouts for one provider round trip. If an error abandons a suspended job, its provider teardown runs in a dedicated cancellation step. ### Parallel Tool Execution [](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#parallel-tool-execution) Under DBOS, tools are executed in parallel by default to minimize latency. To guarantee deterministic replay and reliable recovery, DBOS waits for all parallel tool calls to complete before emitting events **in order**. It’s equivalent to the behavior of [`with agent.parallel_tool_call_execution_mode('parallel_ordered_events')`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.parallel_tool_call_execution_mode) . If you prefer strict ordering, you can configure the agent to run tools sequentially by setting `parallel_execution_mode='sequential'` on [`DBOSDurability`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.dbos.DBOSDurability) . ### Toolsets at Runtime [](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#toolsets-at-runtime) Additional toolsets can be passed per run via `agent.run(toolsets=...)`. Non-executing toolsets like [`ExternalToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.ExternalToolset) , and [`FunctionToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.FunctionToolset) s whose tools DBOS runs inline, are supported. [`MCPToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset) s and dynamic toolsets must be set when constructing the agent so their steps are registered before the workflow runs; passing them at runtime raises a `UserError`. Step Configuration ------------------ [](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#step-configuration) You can customize DBOS step behavior, such as retries, by passing [`StepConfig`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.dbos.StepConfig) objects to the [`DBOSDurability`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.dbos.DBOSDurability) constructor: * `mcp_step_config`: The DBOS step config to use for MCP server communication. No retries if omitted. * `model_step_config`: The DBOS step config to use for model request steps. No retries if omitted. * `event_stream_handler_step_config`: The DBOS step config to use for event stream handler steps (`DBOSDurability` only). No retries if omitted. Unlike the [Temporal](https://pydantic.dev/docs/ai/capabilities/durable_execution/temporal/#per-tool-activity-config) and [Prefect](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#tool-wrapping) integrations, DBOS takes no per-tool config: tool metadata (a `'dbos'` key or otherwise) is ignored, and there’s no way to opt an individual tool out of step wrapping. For custom tools, you can annotate them directly with [`@DBOS.step`](https://docs.dbos.dev/python/reference/decorators#step) or [`@DBOS.workflow`](https://docs.dbos.dev/python/reference/decorators#workflow) decorators as needed. These decorators have no effect outside DBOS workflows, so tools remain usable in non-DBOS agents. Step Retries ------------ [](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#step-retries) On top of the automatic retries for request failures that DBOS will perform, Pydantic AI and various provider API clients also have their own request retry logic. Enabling these at the same time may cause the request to be retried more often than expected, with improper `Retry-After` handling. When using DBOS, it’s recommended to not use [HTTP Request Retries](https://pydantic.dev/docs/ai/models/http-request-retries/) and to turn off your provider API client’s own retry logic, for example by setting `max_retries=0` on a [custom `OpenAIProvider` API client](https://pydantic.dev/docs/ai/models/openai/#custom-openai-client) . You can customize DBOS’s retry policy using [step configuration](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#step-configuration) . DBOS has no selective non-retryable-exception support, so if you enable step retries (`retries_allowed`), framework misconfiguration errors like `UserError` are retried along with everything else. The Temporal and Prefect integrations mark those non-retryable; on DBOS, expect a misconfigured agent to burn its full retry budget before failing. Observability with Logfire -------------------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/dbos/#observability-with-logfire) DBOS can be configured to generate OpenTelemetry spans for each workflow and step execution, and Pydantic AI emits spans for each agent run, model request, and tool invocation. You can send these spans to [Pydantic Logfire](https://pydantic.dev/docs/ai/integrations/logfire/) to get a full, end-to-end view of what’s happening in your application. For more information about DBOS logging and tracing, please see the [DBOS docs](https://docs.dbos.dev/python/tutorials/logging-and-tracing) for details. Was this page helpful? Thanks for your feedback! --- # Business Applications | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#_top) Business Applications ===================== In this example, we’re going to build an agentic app that: * automatically researches each new member that joins a company’s public Slack community to see how good of a fit they are for the company’s commercial product, * sends this analysis into a (private) Slack channel, and * sends a daily summary of the top 5 leads from the previous 24 hours into a (different) Slack channel. We’ll be deploying the app on [Modal](https://modal.com/) , as it lets you use Python to define an app with web endpoints, scheduled functions, and background functions, and deploy them with a CLI, without needing to set up or manage any infrastructure. It’s a great way to lower the barrier for people in your organization to start building and deploying AI agents to make their jobs easier. We also add [Pydantic Logfire](https://pydantic.dev/logfire) to get observability into the app and agent as they’re running in response to webhooks and the schedule Screenshots ----------- [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#screenshots) This is what the analysis sent into Slack will look like: ![Slack message](https://pydantic.dev/docs/ai/img/slack-lead-qualifier-slack.png) This is what the corresponding trace in [Logfire](https://pydantic.dev/logfire) will look like: ![Logfire trace](https://pydantic.dev/docs/ai/img/slack-lead-qualifier-logfire.png) All of these entries can be clicked on to get more details about what happened at that step, including the full conversation with the LLM and HTTP requests and responses. Prerequisites ------------- [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#prerequisites) If you just want to see the code without actually going through the effort of setting up the bits necessary to run it, feel free to [jump ahead](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#the-code) . ### Slack app [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#slack-app) You need to have a Slack workspace and the necessary permissions to create apps. 2. Create a new Slack app using the instructions at [https://docs.slack.dev/quickstart](https://docs.slack.dev/quickstart) . 1. In step 2, “Requesting scopes”, request the following scopes: * [`users.read`](https://docs.slack.dev/reference/scopes/users.read) * [`users.read.email`](https://docs.slack.dev/reference/scopes/users.read.email) * [`users.profile.read`](https://docs.slack.dev/reference/scopes/users.profile.read) 2. In step 3, “Installing and authorizing the app”, note down the Access Token as we’re going to need to store it as a Secret in Modal. 3. You can skip steps 4 and 5. We’re going to need to subscribe to the `team_join` event, but at this point you don’t have a webhook URL yet. 3. Create the channels the app will post into, and add the Slack app to them: * `#new-slack-leads` * `#daily-slack-leads-summary` These names are hard-coded in the example. If you want to use different channels, you can clone the repo and change them in `examples/pydantic_ai_examples/slack_lead_qualifier/functions.py`. ### Logfire Write Token [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#logfire-write-token) 1. If you don’t have a Logfire account yet, create one on [https://logfire-us.pydantic.dev/](https://logfire-us.pydantic.dev/) . 2. Create a new project named, for example, `slack-lead-qualifier`. 3. Generate a new Write Token and note it down, as we’re going to need to store it as a Secret in Modal. ### OpenAI API Key [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#openai-api-key) 1. If you don’t have an OpenAI account yet, create one on [https://platform.openai.com/](https://platform.openai.com/) . 2. Create a new API Key in Settings and note it down, as we’re going to need to store it as a Secret in Modal. ### Modal account [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#modal-account) 1. If you don’t have a Modal account yet, create one on [https://modal.com/signup](https://modal.com/signup) . 2. Following the [Modal Secrets guide](https://modal.com/docs/guide/secrets) , create 3 Secrets of type “Custom”: * Name: `slack`, key: `SLACK_API_KEY`, value: the Slack Access Token you generated earlier * Name: `logfire`, key: `LOGFIRE_TOKEN`, value: the Logfire Write Token you generated earlier * Name: `openai`, key: `OPENAI_API_KEY`, value: the OpenAI API Key you generated earlier Usage ----- [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#usage) 1. Make sure you have the [dependencies installed](https://pydantic.dev/docs/ai/examples/setup/#usage) . 2. Authenticate with Modal: Terminal python/uv-run -m modal setup 3. Run the example as an [ephemeral Modal app](https://modal.com/docs/guide/apps#ephemeral-apps) , meaning it will only run until you quit it using Ctrl+C: Terminal python/uv-run -m modal serve -m pydantic_ai_examples.slack_lead_qualifier.modal 4. Note down the URL after `Created web function web_app =>`, this is your webhook endpoint URL. 5. Go back to [https://docs.slack.dev/quickstart](https://docs.slack.dev/quickstart) and follow step 4, “Configuring the app for event listening”, to subscribe to the `team_join` event with the webhook endpoint URL you noted down as the Request URL. Now when someone new (possibly you with a throwaway email) joins the Slack workspace, you’ll see the webhook event being processed in the terminal where you ran `modal serve` and in the Logfire Live view, and after waiting a few seconds you should see the result appear in the `#new-slack-leads` Slack channel! The code -------- [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#the-code) We’re going to start with the basics, and then gradually build up into the full app. ### Models [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#models) #### `Profile` [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#profile) First, we define a [Pydantic](https://docs.pydantic.dev/) model that represents a Slack user profile. These are the fields we get from the [`team_join`](https://docs.slack.dev/reference/events/team_join) event that’s sent to the webhook endpoint that we’ll define in a bit. models.py ... class Profile(BaseModel): first_name: str | None = None last_name: str | None = None display_name: str | None = None email: str ... We also define a `Profile.as_prompt()` helper method that uses [`format_as_xml`](https://pydantic.dev/docs/ai/api/pydantic-ai/format_prompt/#pydantic_ai.format_prompt.format_as_xml) to turn the profile into a string that can be sent to the model. models.py ... from pydantic_ai import format_as_xml ... class Profile(BaseModel): ... def as_prompt(self) -> str: return format_as_xml(self, root_tag='profile') ... #### `Analysis` [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#analysis) The second model we’ll need represents the result of the analysis that the agent will perform. We include docstrings to provide additional context to the model on what these fields should contain. models.py ... class Analysis(BaseModel): profile: Profile organization_name: str organization_domain: str job_title: str relevance: Annotated[int, Ge(1), Le(5)] """Estimated fit for Pydantic Logfire: 1 = low, 5 = high""" summary: str """One-sentence welcome note summarising who they are and how we might help""" ... We also define a `Analysis.as_slack_blocks()` helper method that turns the analysis into some [Slack blocks](https://api.slack.com/reference/block-kit/blocks) that can be sent to the Slack API to post a new message. models.py ... class Analysis(BaseModel): ... def as_slack_blocks(self, include_relevance: bool = False) -> list[dict[str, Any]]: profile = self.profile relevance = f'({self.relevance}/5)' if include_relevance else '' return [\ {\ 'type': 'markdown',\ 'text': f'[{profile.display_name}](mailto:{profile.email}), {self.job_title} at [**{self.organization_name}**](https://{self.organization_domain}) {relevance}',\ },\ {\ 'type': 'markdown',\ 'text': self.summary,\ },\ ] ... ### Agent [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#agent) Now it’s time to get into Pydantic AI and define the agent that will do the actual analysis! We specify the model we’ll use (`openai:gpt-5`), provide [instructions](https://pydantic.dev/docs/ai/core-concepts/agent/#instructions) , give the agent access to the [DuckDuckGo search tool](https://pydantic.dev/docs/ai/tools-toolsets/common-tools/#duckduckgo-search-tool) , and tell it to output either an `Analysis` or `None` using the [Native Output](https://pydantic.dev/docs/ai/core-concepts/output/#native-output) structured output mode. The real meat of the app is in the instructions that tell the agent how to evaluate each new Slack member. If you plan to use this app yourself, you’ll of course want to modify them to your own situation. agent.py ... from pydantic_ai import Agent, NativeOutput from pydantic_ai.common_tools.duckduckgo import duckduckgo_search_tool ... agent = Agent( 'openai:gpt-5.2', instructions=dedent( """ When a new person joins our public Slack, please put together a brief snapshot so we can be most useful to them. **What to include** 1. **Who they are:** Any details about their professional role or projects (e.g. LinkedIn, GitHub, company bio). 2. **Where they work:** Name of the organisation and its domain. 3. **How we can help:** On a scale of 1–5, estimate how likely they are to benefit from **Pydantic Logfire** (our paid observability tool) based on factors such as company size, product maturity, or AI usage. *1 = probably not relevant, 5 = very strong fit.* **Our products (for context only)** • **Pydantic Validation** – Python data-validation (open source) • **Pydantic AI** – Python agent framework (open source) • **Pydantic Logfire** – Observability for traces, logs & metrics with first-class AI support (commercial) **How to research** • Use the provided DuckDuckGo search tool to research the person and the organization they work for, based on the email domain or what you find on e.g. LinkedIn and GitHub. • If you can't find enough to form a reasonable view, return **None**. """ ), tools=[duckduckgo_search_tool()], output_type=NativeOutput([Analysis, NoneType]), ) ... #### `analyze_profile` [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#analyze_profile) We also define a `analyze_profile` helper function that takes a `Profile`, runs the agent, and returns an `Analysis` (or `None`), and instrument it using [Logfire](https://pydantic.dev/docs/ai/integrations/logfire/) . agent.py ... @logfire.instrument('Analyze profile') async def analyze_profile(profile: Profile) -> Analysis | None: result = await agent.run(profile.as_prompt()) return result.output ... ### Analysis store [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#analysis-store) The next building block we’ll need is a place to store all the analyses that have been done so that we can look them up when we send the daily summary. Fortunately, Modal provides us with a convenient way to store some data that can be read back in a subsequent Modal run (webhook or scheduled): [`modal.Dict`](https://modal.com/docs/reference/modal.Dict) . We define some convenience methods to easily add, list, and clear analyses. store.py ... import modal ... class AnalysisStore: @classmethod @logfire.instrument('Add analysis to store') async def add(cls, analysis: Analysis): await cls._get_store().put.aio(analysis.profile.email, analysis.model_dump()) @classmethod @logfire.instrument('List analyses from store') async def list(cls) -> list[Analysis]: return [\ Analysis.model_validate(analysis)\ async for analysis in cls._get_store().values.aio()\ ] @classmethod @logfire.instrument('Clear analyses from store') async def clear(cls): await cls._get_store().clear.aio() @classmethod def _get_store(cls) -> modal.Dict: return modal.Dict.from_name('analyses', create_if_missing=True) # pyright: ignore[reportUnknownMemberType] ... ### Send Slack message [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#send-slack-message) Next, we’ll need a way to actually send a Slack message, so we define a simple function that uses Slack’s [`chat.postMessage`](https://api.slack.com/methods/chat.postMessage) API. slack.py ... API_KEY = os.getenv('SLACK_API_KEY') assert API_KEY, 'SLACK_API_KEY is not set' @logfire.instrument('Send Slack message') async def send_slack_message(channel: str, blocks: list[dict[str, Any]]): client = httpx.AsyncClient() response = await client.post( 'https://slack.com/api/chat.postMessage', json={ 'channel': channel, 'blocks': blocks, }, headers={ 'Authorization': f'Bearer {API_KEY}', }, timeout=5, ) response.raise_for_status() result = response.json() if not result.get('ok', False): error = result.get('error', 'Unknown error') raise Exception(f'Failed to send to Slack: {error}') ... ### Features [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#features) Now we can start putting these building blocks together to implement the actual features we want! #### `process_slack_member` [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#process_slack_member) This function takes a [`Profile`](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#profile) , [analyzes](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#analyze_profile) it using the agent, adds it to the [`AnalysisStore`](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#analysis-store) , and [sends](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#send-slack-message) the analysis into the `#new-slack-leads` channel. functions.py ... from .agent import analyze_profile from .models import Profile from .slack import send_slack_message from .store import AnalysisStore ... NEW_LEAD_CHANNEL = '#new-slack-leads' ... @logfire.instrument('Process Slack member') async def process_slack_member(profile: Profile): analysis = await analyze_profile(profile) logfire.info('Analysis', analysis=analysis) if analysis is None: return await AnalysisStore().add(analysis) await send_slack_message( NEW_LEAD_CHANNEL, [\ {\ 'type': 'header',\ 'text': {\ 'type': 'plain_text',\ 'text': f'New Slack member with score {analysis.relevance}/5',\ },\ },\ {\ 'type': 'divider',\ },\ *analysis.as_slack_blocks(),\ ], ) ... #### `send_daily_summary` [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#send_daily_summary) This function list all of the analyses in the [`AnalysisStore`](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#analysis-store) , takes the top 5 by relevance, [sends](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#send-slack-message) them into the `#daily-slack-leads-summary` channel, and clears the `AnalysisStore` so that the next daily run won’t process these analyses again. functions.py ... from .slack import send_slack_message from .store import AnalysisStore ... DAILY_SUMMARY_CHANNEL = '#daily-slack-leads-summary' ... @logfire.instrument('Send daily summary') async def send_daily_summary(): analyses = await AnalysisStore().list() logfire.info('Analyses', analyses=analyses) if len(analyses) == 0: return sorted_analyses = sorted(analyses, key=lambda x: x.relevance, reverse=True) top_analyses = sorted_analyses[:5] blocks = [\ {\ 'type': 'header',\ 'text': {\ 'type': 'plain_text',\ 'text': f'Top {len(top_analyses)} new Slack members from the last 24 hours',\ },\ },\ ] for analysis in top_analyses: blocks.extend( [\ {\ 'type': 'divider',\ },\ *analysis.as_slack_blocks(include_relevance=True),\ ] ) await send_slack_message( DAILY_SUMMARY_CHANNEL, blocks, ) await AnalysisStore().clear() ... ### Web app [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#web-app) As it stands, neither of these functions are actually being called from anywhere. Let’s implement a [FastAPI](https://fastapi.tiangolo.com/) endpoint to handle the `team_join` Slack webhook (also known as the [Slack Events API](https://docs.slack.dev/apis/events-api) ) and call the [`process_slack_member`](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#process_slack_member) function we just defined. We also instrument FastAPI using Logfire for good measure. app.py ... app = FastAPI() logfire.instrument_fastapi(app, capture_headers=True) @app.post('/') async def process_webhook(payload: dict[str, Any]) -> dict[str, Any]: if payload['type'] == 'url_verification': return {'challenge': payload['challenge']} elif ( payload['type'] == 'event_callback' and payload['event']['type'] == 'team_join' ): profile = Profile.model_validate(payload['event']['user']['profile']) process_slack_member(profile) return {'status': 'OK'} raise HTTPException(status_code=status.HTTP_422_UNPROCESSABLE_ENTITY) ... #### `process_slack_member` with Modal [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#process_slack_member-with-modal) I was a little sneaky there — we’re not actually calling the [`process_slack_member`](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#process_slack_member) function we defined in `functions.py` directly, as Slack requires webhooks to respond within 3 seconds, and we need a bit more time than that to talk to the LLM, do some web searches, and send the Slack message. Instead, we’re calling the following function defined alongside the app, which uses Modal’s [`modal.Function.spawn`](https://modal.com/docs/reference/modal.Function#spawn) feature to run a function in the background. (If you’re curious what the Modal side of this function looks like, you can [jump ahead](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#backgrounded-process_slack_member) .) Because `modal.py` (which we’ll see in the next section) imports `app.py`, we import from `modal.py` inside the function definition because doing so at the top level would have resulted in a circular import error. We also pass along the current Logfire context to get [Distributed Tracing](https://logfire.pydantic.dev/docs/how-to-guides/distributed-tracing/) , meaning that the background function execution will show up nested under the webhook request trace, so that we have everything related to that request in one place. app.py ... def process_slack_member(profile: Profile): from .modal import process_slack_member as _process_slack_member _process_slack_member.spawn( profile.model_dump(), logfire_ctx=get_context() ) ... ### Modal app [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#modal-app) Now let’s see how easy Modal makes it to deploy all of this. #### Set up Modal [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#set-up-modal) The first thing we do is define the Modal app, by specifying the base image to use (Debian with Python 3.13), all the Python packages it needs, and all of the secrets defined in the Modal interface that need to be made available during runtime. modal.py ... import modal image = modal.Image.debian_slim(python_version='3.13').pip_install( 'pydantic', 'pydantic_ai_slim[openai,duckduckgo]', 'logfire[httpx,fastapi]', 'fastapi[standard]', 'httpx', ) app = modal.App( name='slack-lead-qualifier', image=image, secrets=[\ modal.Secret.from_name('logfire'),\ modal.Secret.from_name('openai'),\ modal.Secret.from_name('slack'),\ ], ) ... #### Set up Logfire [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#set-up-logfire) Next, we define a function to set up Logfire instrumentation for Pydantic AI and HTTPX. We cannot do this at the top level of the file, as the requested packages (like `logfire`) will only be available within functions running on Modal (like the ones we’ll define next). This file, `modal.py`, runs on your local machine and only has access to the `modal` package. modal.py ... def setup_logfire(): import logfire logfire.configure(service_name=app.name) logfire.instrument_pydantic_ai() logfire.instrument_httpx(capture_all=True) ... #### Web app [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#web-app-1) To deploy a [web endpoint](https://modal.com/docs/guide/webhooks) on Modal, we simply define a function that returns an ASGI app (like FastAPI) and decorate it with `@app.function()` and `@modal.asgi_app()`. This `web_app` function will be run on Modal, so inside the function we can call the `setup_logfire` function that requires the `logfire` package, and import `app.py` which uses the other requested packages. By default, Modal spins up a container to handle a function call (like a web request) on-demand, meaning there’s a little bit of startup time to each request. However, Slack requires webhooks to respond within 3 seconds, so we specify `min_containers=1` to keep the web endpoint running and ready to answer requests at all times. This is a bit annoying and wasteful, but fortunately [Modal’s pricing](https://modal.com/pricing) is pretty reasonable, you get $30 free monthly compute, and they offer up to $50k in free credits for startup and academic researchers. modal.py ... @app.function(min_containers=1) @modal.asgi_app() # pyright: ignore[reportUnknownMemberType] def web_app(): setup_logfire() from .app import app as _app return _app ... #### Scheduled `send_daily_summary` [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#scheduled-send_daily_summary) To define a [scheduled function](https://modal.com/docs/guide/cron) , we can use the `@app.function()` decorator with a `schedule` argument. This Modal function will call our imported [`send_daily_summary`](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#send_daily_summary) function every day at 8 am UTC. modal.py ... @app.function(schedule=modal.Cron('0 8 * * *')) # Every day at 8am UTC async def send_daily_summary(): setup_logfire() from .functions import send_daily_summary as _send_daily_summary await _send_daily_summary() ... #### Backgrounded `process_slack_member` [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#backgrounded-process_slack_member) Finally, we define a Modal function that wraps our [`process_slack_member`](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#process_slack_member) function, so that it can run in the background. As you’ll remember from when we [spawned this function from the web app](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#process_slack_member-with-modal) , we passed along the Logfire context to get [Distributed Tracing](https://logfire.pydantic.dev/docs/how-to-guides/distributed-tracing/) , so we need to attach it here. modal.py ... @app.function() async def process_slack_member(profile_raw: dict[str, Any], logfire_ctx: Any): setup_logfire() from logfire.propagate import attach_context from .functions import process_slack_member as _process_slack_member from .models import Profile with attach_context(logfire_ctx): profile = Profile.model_validate(profile_raw) await _process_slack_member(profile) ... Conclusion ---------- [](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#conclusion) And that’s it! Now, assuming you’ve met the [prerequisites](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#prerequisites) , you can run or deploy the app using the commands under [usage](https://pydantic.dev/docs/ai/examples/slack-lead-qualifier/#usage) . Was this page helpful? Thanks for your feedback! --- # UI Examples | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/ag-ui/#_top) UI Examples =========== Example of using Pydantic AI agents with the [AG-UI Dojo](https://github.com/ag-ui-protocol/ag-ui/tree/main/apps/dojo) example app. See the [AG-UI docs](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/) for more information about the AG-UI integration. Demonstrates: * [AG-UI](https://pydantic.dev/docs/ai/integrations/ui/ag-ui/) * [Tools](https://pydantic.dev/docs/ai/tools-toolsets/tools/) Prerequisites ------------- [](https://pydantic.dev/docs/ai/examples/ag-ui/#prerequisites) * An [OpenAI API key](https://help.openai.com/en/articles/4936850-where-do-i-find-my-openai-api-key) Running the Example ------------------- [](https://pydantic.dev/docs/ai/examples/ag-ui/#running-the-example) With [dependencies installed and environment variables set](https://pydantic.dev/docs/ai/examples/setup/#usage) you will need two command line windows. ### Pydantic AI AG-UI backend [](https://pydantic.dev/docs/ai/examples/ag-ui/#pydantic-ai-ag-ui-backend) Setup your OpenAI API Key Terminal export OPENAI_API_KEY= Start the Pydantic AI AG-UI example backend. * [pip](https://pydantic.dev/docs/ai/examples/ag-ui/#tab-panel-22) * [uv](https://pydantic.dev/docs/ai/examples/ag-ui/#tab-panel-23) Terminal python -m pydantic_ai_examples.ag_ui Terminal uv run -m pydantic_ai_examples.ag_ui ### AG-UI Dojo example frontend [](https://pydantic.dev/docs/ai/examples/ag-ui/#ag-ui-dojo-example-frontend) Next run the AG-UI Dojo example frontend. 1. Clone the [AG-UI repository](https://github.com/ag-ui-protocol/ag-ui) Terminal git clone https://github.com/ag-ui-protocol/ag-ui.git 2. Change into to the `ag-ui/typescript-sdk` directory Terminal cd ag-ui/sdks/typescript 3. Run the Dojo app following the [official instructions](https://github.com/ag-ui-protocol/ag-ui/tree/main/apps/dojo#development-setup) 4. Visit [http://localhost:3000/pydantic-ai](http://localhost:3000/pydantic-ai) 5. Select View `Pydantic AI` from the sidebar Feature Examples ---------------- [](https://pydantic.dev/docs/ai/examples/ag-ui/#feature-examples) ### Agentic Chat [](https://pydantic.dev/docs/ai/examples/ag-ui/#agentic-chat) This demonstrates a basic agent interaction including Pydantic AI server side tools and AG-UI client side tools. If you’ve [run the example](https://pydantic.dev/docs/ai/examples/ag-ui/#running-the-example) , you can view it at [http://localhost:3000/pydantic-ai/feature/agentic\_chat](http://localhost:3000/pydantic-ai/feature/agentic_chat) . #### Agent Tools [](https://pydantic.dev/docs/ai/examples/ag-ui/#agent-tools) * `time` - Pydantic AI tool to check the current time for a time zone * `background` - AG-UI tool to set the background color of the client window #### Agent Prompts [](https://pydantic.dev/docs/ai/examples/ag-ui/#agent-prompts) What is the time in New York? Change the background to blue A complex example which mixes both AG-UI and Pydantic AI tools: Perform the following steps, waiting for the response of each step before continuing: 1. Get the time 2. Set the background to red 3. Get the time 4. Report how long the background set took by diffing the two times #### Agentic Chat - Code [](https://pydantic.dev/docs/ai/examples/ag-ui/#agentic-chat---code) agentic\_chat.py from __future__ import annotations from datetime import datetime from zoneinfo import ZoneInfo from starlette.applications import Starlette from starlette.requests import Request from starlette.responses import Response from starlette.routing import Route from pydantic_ai import Agent from pydantic_ai.ui.ag_ui import AGUIAdapter agent = Agent('openai:gpt-5-mini') @agent.tool_plain async def current_time(timezone: str = 'UTC') -> str: """Get the current time in ISO format. Args: timezone: The timezone to use. Returns: The current time in ISO format string. """ tz: ZoneInfo = ZoneInfo(timezone) return datetime.now(tz=tz).isoformat() async def run_agent(request: Request) -> Response: return await AGUIAdapter.dispatch_request(request, agent=agent) app = Starlette(routes=[Route('/', run_agent, methods=['POST'])]) ### Agentic Generative UI [](https://pydantic.dev/docs/ai/examples/ag-ui/#agentic-generative-ui) Demonstrates a long running task where the agent sends updates to the frontend to let the user know what’s happening. If you’ve [run the example](https://pydantic.dev/docs/ai/examples/ag-ui/#running-the-example) , you can view it at [http://localhost:3000/pydantic-ai/feature/agentic\_generative\_ui](http://localhost:3000/pydantic-ai/feature/agentic_generative_ui) . #### Plan Prompts [](https://pydantic.dev/docs/ai/examples/ag-ui/#plan-prompts) Create a plan for breakfast and execute it #### Agentic Generative UI - Code [](https://pydantic.dev/docs/ai/examples/ag-ui/#agentic-generative-ui---code) agentic\_generative\_ui.py from __future__ import annotations from textwrap import dedent from typing import Any, Literal from pydantic import BaseModel, Field from starlette.applications import Starlette from starlette.requests import Request from starlette.responses import Response from starlette.routing import Route from ag_ui.core import EventType, StateDeltaEvent, StateSnapshotEvent from pydantic_ai import Agent from pydantic_ai.ui.ag_ui import AGUIAdapter StepStatus = Literal['pending', 'completed'] class Step(BaseModel): """Represents a step in a plan.""" description: str = Field(description='The description of the step') status: StepStatus = Field( default='pending', description='The status of the step (e.g., pending, completed)', ) class Plan(BaseModel): """Represents a plan with multiple steps.""" steps: list[Step] = Field( default_factory=list[Step], description='The steps in the plan' ) class JSONPatchOp(BaseModel): """A class representing a JSON Patch operation (RFC 6902).""" op: Literal['add', 'remove', 'replace', 'move', 'copy', 'test'] = Field( description='The operation to perform: add, remove, replace, move, copy, or test', ) path: str = Field(description='JSON Pointer (RFC 6901) to the target location') value: Any = Field( default=None, description='The value to apply (for add, replace operations)', ) from_: str | None = Field( default=None, alias='from', description='Source path (for move, copy operations)', ) agent = Agent( 'openai:gpt-5-mini', instructions=dedent( """ When planning use tools only, without any other messages. IMPORTANT: - Use the `create_plan` tool to set the initial state of the steps - Use the `update_plan_step` tool to update the status of each step - Do NOT repeat the plan or summarise it in a message - Do NOT confirm the creation or updates in a message - Do NOT ask the user for additional information or next steps Only one plan can be active at a time, so do not call the `create_plan` tool again until all the steps in current plan are completed. """ ), ) @agent.tool_plain async def create_plan(steps: list[str]) -> StateSnapshotEvent: """Create a plan with multiple steps. Args: steps: List of step descriptions to create the plan. Returns: StateSnapshotEvent containing the initial state of the steps. """ plan: Plan = Plan( steps=[Step(description=step) for step in steps], ) return StateSnapshotEvent( type=EventType.STATE_SNAPSHOT, snapshot=plan.model_dump(), ) @agent.tool_plain async def update_plan_step( index: int, description: str | None = None, status: StepStatus | None = None ) -> StateDeltaEvent: """Update the plan with new steps or changes. Args: index: The index of the step to update. description: The new description for the step. status: The new status for the step. Returns: StateDeltaEvent containing the changes made to the plan. """ changes: list[JSONPatchOp] = [] if description is not None: changes.append( JSONPatchOp( op='replace', path=f'/steps/{index}/description', value=description ) ) if status is not None: changes.append( JSONPatchOp(op='replace', path=f'/steps/{index}/status', value=status) ) return StateDeltaEvent( type=EventType.STATE_DELTA, delta=changes, ) async def run_agent(request: Request) -> Response: return await AGUIAdapter.dispatch_request(request, agent=agent) app = Starlette(routes=[Route('/', run_agent, methods=['POST'])]) ### Human in the Loop [](https://pydantic.dev/docs/ai/examples/ag-ui/#human-in-the-loop) Demonstrates simple human in the loop workflow where the agent comes up with a plan and the user can approve it using checkboxes. #### Task Planning Tools [](https://pydantic.dev/docs/ai/examples/ag-ui/#task-planning-tools) * `generate_task_steps` - AG-UI tool to generate and confirm steps #### Task Planning Prompt [](https://pydantic.dev/docs/ai/examples/ag-ui/#task-planning-prompt) Generate a list of steps for cleaning a car for me to review #### Human in the Loop - Code [](https://pydantic.dev/docs/ai/examples/ag-ui/#human-in-the-loop---code) human\_in\_the\_loop.py from __future__ import annotations from textwrap import dedent from starlette.applications import Starlette from starlette.requests import Request from starlette.responses import Response from starlette.routing import Route from pydantic_ai import Agent from pydantic_ai.ui.ag_ui import AGUIAdapter agent = Agent( 'openai:gpt-5-mini', instructions=dedent( """ When planning tasks use tools only, without any other messages. IMPORTANT: - Use the `generate_task_steps` tool to display the suggested steps to the user - Never repeat the plan, or send a message detailing steps - If accepted, confirm the creation of the plan and the number of selected (enabled) steps only - If not accepted, ask the user for more information, DO NOT use the `generate_task_steps` tool again """ ), ) async def run_agent(request: Request) -> Response: return await AGUIAdapter.dispatch_request(request, agent=agent) app = Starlette(routes=[Route('/', run_agent, methods=['POST'])]) ### Predictive State Updates [](https://pydantic.dev/docs/ai/examples/ag-ui/#predictive-state-updates) Demonstrates how to use the predictive state updates feature to update the state of the UI based on agent responses, including user interaction via user confirmation. If you’ve [run the example](https://pydantic.dev/docs/ai/examples/ag-ui/#running-the-example) , you can view it at [http://localhost:3000/pydantic-ai/feature/predictive\_state\_updates](http://localhost:3000/pydantic-ai/feature/predictive_state_updates) . #### Story Tools [](https://pydantic.dev/docs/ai/examples/ag-ui/#story-tools) * `write_document` - AG-UI tool to write the document to a window * `document_predict_state` - Pydantic AI tool that enables document state prediction for the `write_document` tool This also shows how to use custom instructions based on shared state information. #### Story Example [](https://pydantic.dev/docs/ai/examples/ag-ui/#story-example) Starting document text Bruce was a good dog, Agent prompt Help me complete my story about bruce the dog, is should be no longer than a sentence. #### Predictive State Updates - Code [](https://pydantic.dev/docs/ai/examples/ag-ui/#predictive-state-updates---code) predictive\_state\_updates.py from __future__ import annotations from dataclasses import replace from textwrap import dedent from pydantic import BaseModel from starlette.applications import Starlette from starlette.requests import Request from starlette.responses import Response from starlette.routing import Route from ag_ui.core import CustomEvent, EventType from pydantic_ai import Agent, RunContext from pydantic_ai.ui import StateDeps from pydantic_ai.ui.ag_ui import AGUIAdapter class DocumentState(BaseModel): """State for the document being written.""" document: str = '' agent = Agent('openai:gpt-5-mini', deps_type=StateDeps[DocumentState]) # Tools which return AG-UI events will be sent to the client as part of the # event stream, single events and iterables of events are supported. @agent.tool_plain async def document_predict_state() -> list[CustomEvent]: """Enable document state prediction. Returns: CustomEvent containing the event to enable state prediction. """ return [\ CustomEvent(\ type=EventType.CUSTOM,\ name='PredictState',\ value=[\ {\ 'state_key': 'document',\ 'tool': 'write_document',\ 'tool_argument': 'document',\ },\ ],\ ),\ ] @agent.instructions() async def story_instructions(ctx: RunContext[StateDeps[DocumentState]]) -> str: """Provide instructions for writing document if present. Args: ctx: The run context containing document state information. Returns: Instructions string for the document writing agent. """ return dedent( f"""You are a helpful assistant for writing documents. Before you start writing, you MUST call the `document_predict_state` tool to enable state prediction. To present the document to the user for review, you MUST use the `write_document` tool. When you have written the document, DO NOT repeat it as a message. If accepted briefly summarize the changes you made, 2 sentences max, otherwise ask the user to clarify what they want to change. This is the current document: {ctx.deps.state.document} """ ) deps = StateDeps(DocumentState()) async def run_agent(request: Request) -> Response: # `dispatch_request` mutates `deps.state` from the request, so give each request its own copy. return await AGUIAdapter.dispatch_request(request, agent=agent, deps=replace(deps)) app = Starlette(routes=[Route('/', run_agent, methods=['POST'])]) ### Shared State [](https://pydantic.dev/docs/ai/examples/ag-ui/#shared-state) Demonstrates how to use the shared state between the UI and the agent. State sent to the agent is detected by a function based instruction. This then validates the data using a custom pydantic model before using to create the instructions for the agent to follow and send to the client using a AG-UI tool. If you’ve [run the example](https://pydantic.dev/docs/ai/examples/ag-ui/#running-the-example) , you can view it at [http://localhost:3000/pydantic-ai/feature/shared\_state](http://localhost:3000/pydantic-ai/feature/shared_state) . #### Recipe Tools [](https://pydantic.dev/docs/ai/examples/ag-ui/#recipe-tools) * `display_recipe` - AG-UI tool to display the recipe in a graphical format #### Recipe Example [](https://pydantic.dev/docs/ai/examples/ag-ui/#recipe-example) 1. Customise the basic settings of your recipe 2. Click `Improve with AI` #### Shared State - Code [](https://pydantic.dev/docs/ai/examples/ag-ui/#shared-state---code) shared\_state.py from __future__ import annotations from dataclasses import replace from enum import Enum from textwrap import dedent from pydantic import BaseModel, Field from starlette.applications import Starlette from starlette.requests import Request from starlette.responses import Response from starlette.routing import Route from ag_ui.core import EventType, StateSnapshotEvent from pydantic_ai import Agent, RunContext from pydantic_ai.ui import StateDeps from pydantic_ai.ui.ag_ui import AGUIAdapter class SkillLevel(str, Enum): """The level of skill required for the recipe.""" BEGINNER = 'Beginner' INTERMEDIATE = 'Intermediate' ADVANCED = 'Advanced' class SpecialPreferences(str, Enum): """Special preferences for the recipe.""" HIGH_PROTEIN = 'High Protein' LOW_CARB = 'Low Carb' SPICY = 'Spicy' BUDGET_FRIENDLY = 'Budget-Friendly' ONE_POT_MEAL = 'One-Pot Meal' VEGETARIAN = 'Vegetarian' VEGAN = 'Vegan' class CookingTime(str, Enum): """The cooking time of the recipe.""" FIVE_MIN = '5 min' FIFTEEN_MIN = '15 min' THIRTY_MIN = '30 min' FORTY_FIVE_MIN = '45 min' SIXTY_PLUS_MIN = '60+ min' class Ingredient(BaseModel): """A class representing an ingredient in a recipe.""" icon: str = Field( default='ingredient', description="The icon emoji (not emoji code like '\x1f35e', but the actual emoji like 🥕) of the ingredient", ) name: str amount: str class Recipe(BaseModel): """A class representing a recipe.""" skill_level: SkillLevel = Field( default=SkillLevel.BEGINNER, description='The skill level required for the recipe', ) special_preferences: list[SpecialPreferences] = Field( default_factory=list[SpecialPreferences], description='Any special preferences for the recipe', ) cooking_time: CookingTime = Field( default=CookingTime.FIVE_MIN, description='The cooking time of the recipe' ) ingredients: list[Ingredient] = Field( default_factory=list[Ingredient], description='Ingredients for the recipe', ) instructions: list[str] = Field( default_factory=list[str], description='Instructions for the recipe' ) class RecipeSnapshot(BaseModel): """A class representing the state of the recipe.""" recipe: Recipe = Field( default_factory=Recipe, description='The current state of the recipe' ) agent = Agent('openai:gpt-5-mini', deps_type=StateDeps[RecipeSnapshot]) @agent.tool_plain async def display_recipe(recipe: Recipe) -> StateSnapshotEvent: """Display the recipe to the user. Args: recipe: The recipe to display. Returns: StateSnapshotEvent containing the recipe snapshot. """ return StateSnapshotEvent( type=EventType.STATE_SNAPSHOT, snapshot={'recipe': recipe}, ) @agent.instructions async def recipe_instructions(ctx: RunContext[StateDeps[RecipeSnapshot]]) -> str: """Instructions for the recipe generation agent. Args: ctx: The run context containing recipe state information. Returns: Instructions string for the recipe generation agent. """ return dedent( f""" You are a helpful assistant for creating recipes. IMPORTANT: - Create a complete recipe using the existing ingredients - Append new ingredients to the existing ones - Use the `display_recipe` tool to present the recipe to the user - Do NOT repeat the recipe in the message, use the tool instead - Do NOT run the `display_recipe` tool multiple times in a row Once you have created the updated recipe and displayed it to the user, summarise the changes in one sentence, don't describe the recipe in detail or send it as a message to the user. The current state of the recipe is: {ctx.deps.state.recipe.model_dump_json(indent=2)} """, ) deps = StateDeps(RecipeSnapshot()) async def run_agent(request: Request) -> Response: # `dispatch_request` mutates `deps.state` from the request, so give each request its own copy. return await AGUIAdapter.dispatch_request(request, agent=agent, deps=replace(deps)) app = Starlette(routes=[Route('/', run_agent, methods=['POST'])]) ### Tool Based Generative UI [](https://pydantic.dev/docs/ai/examples/ag-ui/#tool-based-generative-ui) Demonstrates customised rendering for tool output with used confirmation. If you’ve [run the example](https://pydantic.dev/docs/ai/examples/ag-ui/#running-the-example) , you can view it at [http://localhost:3000/pydantic-ai/feature/tool\_based\_generative\_ui](http://localhost:3000/pydantic-ai/feature/tool_based_generative_ui) . #### Haiku Tools [](https://pydantic.dev/docs/ai/examples/ag-ui/#haiku-tools) * `generate_haiku` - AG-UI tool to display a haiku in English and Japanese #### Haiku Prompt [](https://pydantic.dev/docs/ai/examples/ag-ui/#haiku-prompt) Generate a haiku about formula 1 #### Tool Based Generative UI - Code [](https://pydantic.dev/docs/ai/examples/ag-ui/#tool-based-generative-ui---code) tool\_based\_generative\_ui.py from __future__ import annotations from starlette.applications import Starlette from starlette.requests import Request from starlette.responses import Response from starlette.routing import Route from pydantic_ai import Agent from pydantic_ai.ui.ag_ui import AGUIAdapter agent = Agent('openai:gpt-5-mini') async def run_agent(request: Request) -> Response: return await AGUIAdapter.dispatch_request(request, agent=agent) app = Starlette(routes=[Route('/', run_agent, methods=['POST'])]) Was this page helpful? Thanks for your feedback! --- # Repo Context | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/repo-context/#_top) Repo Context ============ `RepoContext` discovers and loads a repo’s accumulated coding-assistant context engineering (CE): the instruction files (`CLAUDE.md`/`AGENTS.md`) scattered across the tree and the assets under `.claude`/`.agents`/`.codex`/`.grok` (skills, sub-agents, hooks). [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/repo_context/) > The API may change between releases. Where practical, breaking changes ship with a deprecation warning. The problem ----------- [](https://pydantic.dev/docs/ai/harness/repo-context/#the-problem) A repo accumulates CE for whatever coding assistant worked in it: instruction files (`CLAUDE.md`/`AGENTS.md`) scattered across the tree, and assets under `.claude`/`.agents`/`.codex`/`.grok` (skills, sub-agents, hooks). An agent that loads only the top-level instruction file misses the ancestor context and has no idea the rest of the setup exists, so it can neither honor it nor translate it. The solution ------------ [](https://pydantic.dev/docs/ai/harness/repo-context/#the-solution) `RepoContext` bundles three strategies, each independently toggleable. Construct it with `RepoContext(...)` in an `Agent`’s `capabilities`, anchored at the deepest directory the agent works in: from pathlib import Path from pydantic_ai import Agent from pydantic_ai_harness.repo_context import RepoContext agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[RepoContext(workspace_dir=Path('.'), home_dir=Path.home())], ) result = agent.run_sync('Summarize the coding-assistant setup in this repo.') print(result.output) ### 1\. Walk-up instruction autoload (on by default) [](https://pydantic.dev/docs/ai/harness/repo-context/#1-walk-up-instruction-autoload-on-by-default) Loads `CLAUDE.md`/`AGENTS.md` from `workspace_dir` and every ancestor up to `home_dir` (inclusive). Precedence is ancestor-first, workspace-last: broadest context first, most specific last. Files are deduped by resolved real path and by content hash, so a symlinked `AGENTS.md -> CLAUDE.md` or two ancestors sharing identical content load once. When `home_dir` is `None` (the default), only `workspace_dir` is scanned — no walk-up. Pass `home_dir=Path.home()` to walk up to your home directory. ### 2\. Asset inventory (on by default) [](https://pydantic.dev/docs/ai/harness/repo-context/#2-asset-inventory-on-by-default) Exposes one tool, `inventory_agent_context()`, that reports where the repo’s CE assets live — the `.claude`/`.agents`/`.codex`/`.grok` roots and, within each, the `skills/` (`SKILL.md`), `agents/` (`.md`), and `settings.json` (hooks) it contains. It returns a structured `AgentContextInventory`; it locates assets and does not parse them, leaving translation to the orchestrator. Rename the tool with `inventory_tool_name`, or scope which roots it scans with `asset_roots`. ### 3\. Nested-on-traversal (off by default) [](https://pydantic.dev/docs/ai/harness/repo-context/#3-nested-on-traversal-off-by-default) When the model lists or reads a directory, surface that directory’s `CLAUDE.md`/`AGENTS.md`. This couples to the host’s list/read tools, so it is opt-in and configurable: from pathlib import Path from pydantic_ai import Agent from pydantic_ai_harness import FileSystem from pydantic_ai_harness.repo_context import RepoContext agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[\ FileSystem(root_dir='.'),\ RepoContext(\ workspace_dir=Path('.'),\ nested_traversal=True,\ traversal_tool_names=frozenset({'list_directory', 'read_file'}), # the FileSystem tool names to hook\ traversal_path_arg='path', # the path arg key\ nested_inject='pointer', # or 'contents'\ )\ ], ) `nested_inject='pointer'` (default) appends a one-line note pointing at the file; `'contents'` inlines the file body. Each directory is surfaced at most once per run. Cache cost ---------- [](https://pydantic.dev/docs/ai/harness/repo-context/#cache-cost) Injecting file contents into the system prompt costs prompt-cache stability: a changed prefix re-bills the whole cached region. `RepoContext` keeps the two cache-relevant paths separate: * Strategy 1 reads its files once at run start and injects them as static system instructions, so the cached prefix stays byte-identical across turns. * Strategy 3 is volatile (it depends on which directory was just touched), so its note is appended to the tool result in the message tail — never to the system prompt — and cannot invalidate the cached prefix. Configuration ------------- [](https://pydantic.dev/docs/ai/harness/repo-context/#configuration) RepoContext( workspace_dir, # Path -- the deepest dir the agent works in (required) home_dir=None, # Path | None -- shallowest dir to stop walk-up at, inclusive filenames=('CLAUDE.md', 'AGENTS.md'), autoload_instructions=True, # Strategy 1 expose_inventory_tool=True, # Strategy 2 inventory_tool_name='inventory_agent_context', nested_traversal=False, # Strategy 3 nested_inject='pointer', # 'pointer' | 'contents' traversal_tool_names=frozenset({'list_directory', 'read_file'}), traversal_path_arg='path', asset_roots=('.claude', '.agents', '.codex', '.grok'), ) Scope ----- [](https://pydantic.dev/docs/ai/harness/repo-context/#scope) `RepoContext` locates and loads CE; it does not parse skill/sub-agent frontmatter or hook bodies, and it does not rewrite or translate assets. Strategy 1 reads its files once per run, so mid-run edits to those files are not reloaded. Further reading --------------- [](https://pydantic.dev/docs/ai/harness/repo-context/#further-reading) * [Pydantic AI capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) * [Pydantic AI hooks](https://pydantic.dev/docs/ai/core-concepts/hooks/) API reference ------------- [](https://pydantic.dev/docs/ai/harness/repo-context/#api-reference) RepoContext ----------- [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.RepoContext) **Bases:** `AbstractCapability[AgentDepsT]` Discover and load a repo’s accumulated coding-assistant context engineering. Three strategies, each independently toggleable: 1. Walk-up instruction autoload (`autoload_instructions`, on by default): load `CLAUDE.md`/`AGENTS.md` from `workspace_dir` and every ancestor up to `home_dir`, deduped, ancestor-first. These are read once at run start and injected as **static system instructions** via `get_instructions`, so they stay in the cached prefix and never re-read per turn. 2. Asset inventory (`expose_inventory_tool`, on by default): a tool that reports where the repo’s CE assets live (`.claude`/`.agents`/`.codex`/ `.grok` and their `skills/`, `agents/`, `settings.json`). It locates assets; it does not parse them. 3. Nested-on-traversal (`nested_traversal`, off by default): when the model lists or reads a directory (via a tool named in `traversal_tool_names`), surface that directory’s `CLAUDE.md`/`AGENTS.md`. The note is appended to the **tool result** (message tail), not to system instructions, so it does not invalidate the cached prefix. `nested_inject='pointer'` (default) appends a one-line pointer; `'contents'` inlines the file body. Cache note: injecting file contents into the system prompt costs prompt-cache stability. Strategy 1 is safe because its files are static; the volatile Strategy 3 content rides in the message tail instead. from pathlib import Path from pydantic_ai import Agent from pydantic_ai_harness.repo_context import RepoContext agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[RepoContext(workspace_dir=Path('.'), home_dir=Path.home())], ) ### Attributes [](https://pydantic.dev/docs/ai/harness/repo-context/#attributes) #### asset\_roots [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.RepoContext.asset_roots) Root directories the inventory tool scans, relative to `workspace_dir`. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `('.claude', '.agents', '.codex', '.grok')` #### autoload\_instructions [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.RepoContext.autoload_instructions) Strategy 1: load instruction files into the system prompt. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `True` #### expose\_inventory\_tool [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.RepoContext.expose_inventory_tool) Strategy 2: expose the asset-inventory tool. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `True` #### filenames [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.RepoContext.filenames) Instruction filenames to look for, in within-directory precedence order. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `('CLAUDE.md', 'AGENTS.md')` #### home\_dir [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.RepoContext.home_dir) The shallowest directory to stop the walk-up at, inclusive. `None` (the default) scans only `workspace_dir` — no walk-up. **Type:** `Path` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### inventory\_tool\_name [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.RepoContext.inventory_tool_name) Name of the inventory tool exposed to the model. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'inventory_agent_context'` #### nested\_inject [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.RepoContext.nested_inject) For Strategy 3: append a one-line `pointer`, or inline the file `contents`. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘pointer’, ‘contents’\] **Default:** `'pointer'` #### nested\_traversal [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.RepoContext.nested_traversal) Strategy 3: surface a directory’s instruction file when the model lists or reads that directory. Off by default — it couples to the list/read tools. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### traversal\_path\_arg [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.RepoContext.traversal_path_arg) The tool argument key holding the listed/read path. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'path'` #### traversal\_tool\_names [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.RepoContext.traversal_tool_names) Tool names that trigger Strategy 3. Override to match the host’s list/read tools (e.g. `frozenset({'list_dir', 'read_file'})`). **Type:** [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `frozenset({'list_directory', 'read_file'})` #### workspace\_dir [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.RepoContext.workspace_dir) The deepest directory the agent works in. The walk-up and asset scan are anchored here. **Type:** `Path` ### Methods [](https://pydantic.dev/docs/ai/harness/repo-context/#methods) #### after\_tool\_execute [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.RepoContext.after_tool_execute) `@async` def after_tool_execute( ctx: RunContext[AgentDepsT], *, call: ToolCallPart, tool_def: ToolDefinition, args: dict[str, Any], result: Any, ) -> Any Strategy 3: append a directory’s instruction file to a list/read result. ##### Returns [](https://pydantic.dev/docs/ai/harness/repo-context/#returns) [`Any`](https://docs.python.org/3/library/typing.html#typing.Any) #### for\_run [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.RepoContext.for_run) `@async` def for_run(ctx: RunContext[AgentDepsT]) -> RepoContext[AgentDepsT] Return a fresh per-run instance with isolated traversal/cache state. ##### Returns [](https://pydantic.dev/docs/ai/harness/repo-context/#returns-1) `RepoContext`\[`AgentDepsT`\] #### get\_instructions [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.RepoContext.get_instructions) def get_instructions() -> AgentInstructions[AgentDepsT] | None Static, cache-stable instructions: loaded files plus the inventory hint. ##### Returns [](https://pydantic.dev/docs/ai/harness/repo-context/#returns-2) `AgentInstructions`\[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) #### get\_serialization\_name [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.RepoContext.get_serialization_name) `@classmethod` def get_serialization_name(cls) -> str | None Serialization name for agent-spec support. ##### Returns [](https://pydantic.dev/docs/ai/harness/repo-context/#returns-3) [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### get\_toolset [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.RepoContext.get_toolset) def get_toolset() -> AgentToolset[AgentDepsT] | None The asset-inventory toolset, or `None` when the tool is disabled. ##### Returns [](https://pydantic.dev/docs/ai/harness/repo-context/#returns-4) [`AgentToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AgentToolset) \[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) AgentContextInventory --------------------- [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.AgentContextInventory) **Bases:** `BaseModel` A map of where a repo’s CE assets live, for an orchestrator to read or translate. AssetRoot --------- [](https://pydantic.dev/docs/ai/harness/repo-context/#pydantic_ai_harness.repo_context.AssetRoot) **Bases:** `BaseModel` Where CE assets live under a single root directory (e.g. `.claude`). Was this page helpful? Thanks for your feedback! --- # StackOne | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/stackone/#_top) StackOne ======== Use `StackOne` when an agent needs to work with one of a user’s linked business applications, such as BambooHR, Salesforce, or Zendesk. Each instance is scoped to one linked account, which is one authenticated connection between [StackOne](https://www.stackone.com/) and a provider. [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/stackone/) > The API may change between releases. Where practical, breaking changes ship with a deprecation warning. Before you start ---------------- [](https://pydantic.dev/docs/ai/harness/stackone/#before-you-start) Follow the [StackOne docs](https://docs.stackone.com/) to: 1. Configure a connector and link an account. For your first test, enable only the read actions you need. 2. Copy the linked account ID from the StackOne dashboard. 3. Create a StackOne API key that can execute actions. You also need an API key for the model your agent uses. Installation ------------ [](https://pydantic.dev/docs/ai/harness/stackone/#installation) Terminal uv add "pydantic-ai-harness[stackone]" "pydantic-ai-slim[openai,spec]" The `openai` and `spec` extras support the model and agent-spec examples below. Install the provider extra for a different model provider instead. Set the credentials in the shell where you will run the example: Terminal export STACKONE_API_KEY='your-stackone-api-key' export STACKONE_ACCOUNT_ID='your-linked-account-id' export OPENAI_API_KEY='your-openai-api-key' `StackOne` reads `STACKONE_API_KEY` automatically. The example reads `STACKONE_ACCOUNT_ID` explicitly so the account ID is not hard-coded. You can pass `api_key=` directly instead, but keep secrets out of source control. Run your first agent -------------------- [](https://pydantic.dev/docs/ai/harness/stackone/#run-your-first-agent) import os from pydantic_ai import Agent from pydantic_ai_harness.stackone import StackOne agent = Agent( 'openai:gpt-5', capabilities=[\ StackOne(account_id=os.environ['STACKONE_ACCOUNT_ID']),\ ], ) result = agent.run_sync('List the first 5 employees') print(result.output) By default, the model receives two tools. It searches for an action that matches the request, then executes the returned action ID. The final output depends on the linked provider and its data. Control available actions ------------------------- [](https://pydantic.dev/docs/ai/harness/stackone/#control-available-actions) StackOne controls which actions are enabled for the linked account. Treat that configuration as the primary access control. Use `actions` when you also want to limit which tools the model sees. Patterns use Python [`fnmatch`](https://docs.python.org/3/library/fnmatch.html) syntax, where `*` is a wildcard. They ignore case and match the full `{connector}_{action}_{entity}` tool name: from pydantic_ai_harness.stackone import StackOne StackOne(account_id='your-linked-account-id', actions=['*_list_*']) # All matching list tools StackOne(account_id='your-linked-account-id', actions=['workday_get_worker']) # One exact tool Passing `actions` selects `individual` mode automatically. Explicitly combining `actions` with `tool_mode='search_execute'` raises an error because that mode registers only the search and execute tools. Choose a tool mode ------------------ [](https://pydantic.dev/docs/ai/harness/stackone/#choose-a-tool-mode) | Mode | What the model receives | Use it when | | --- | --- | --- | | `search_execute` | Two tools: search for an action, then execute it by ID | The account has many enabled actions. This is the default when `actions` is omitted. | | `individual` | One tool and schema per enabled action | You need to select exact actions or add per-tool behavior. Passing `actions` selects this mode. | In `search_execute` mode, action IDs are returned by the search tool at runtime and should not be guessed. In `individual` mode, all selected tool schemas are sent to the model, so filter large action sets with `actions`. To keep StackOne tools out of the model context until they are needed, pass `defer_loading=True`. The capability uses `id='stackone'` by default so it can be loaded on demand. Give each instance a distinct `id` when one agent uses multiple StackOne accounts: from pydantic_ai_harness.stackone import StackOne StackOne(account_id='your-linked-account-id', defer_loading=True) ### Bound large tool results [](https://pydantic.dev/docs/ai/harness/stackone/#bound-large-tool-results) Provider actions can return large exports. Combine StackOne with the [Tool Output Limits](https://pydantic.dev/docs/ai/harness/tool-output-limits/) capability to reduce oversized tool returns agent-wide: from pydantic_ai import Agent from pydantic_ai_harness.stackone import StackOne from pydantic_ai_harness.tool_output_limits import ToolOutputLimits agent = Agent( 'openai:gpt-5', capabilities=[\ StackOne(account_id='your-linked-account-id'),\ ToolOutputLimits(),\ ], ) ### Require approval [](https://pydantic.dev/docs/ai/harness/stackone/#require-approval) Approval is not enabled automatically. For operations that need human confirmation, use the public `StackOneToolset` with Pydantic AI’s [tool approval](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/#requiring-tool-approval) : import os from pydantic_ai import Agent from pydantic_ai_harness.stackone import StackOneToolset stackone_tools = StackOneToolset( account_id=os.environ['STACKONE_ACCOUNT_ID'], actions=['workday_create_worker'], ).approval_required() agent = Agent('openai:gpt-5', toolsets=[stackone_tools]) Handle the resulting deferred approval requests as described in the linked guide. Define the agent in YAML or JSON -------------------------------- [](https://pydantic.dev/docs/ai/harness/stackone/#define-the-agent-in-yaml-or-json) The capability also works with Pydantic AI’s [agent spec](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) format for YAML or JSON. Keep the API key in `STACKONE_API_KEY` rather than storing it in the file: # agent.yaml model: openai:gpt-5 capabilities: - StackOne: account_id: 'your-linked-account-id' actions: ['*_list_*'] from pydantic_ai import Agent from pydantic_ai_harness.stackone import StackOne agent = Agent.from_file('agent.yaml', custom_capability_types=[StackOne]) Pass `custom_capability_types` so the spec loader knows how to instantiate `StackOne`. Use the lower-level `StackOneToolset` directly when you need [`Agent(toolsets=[...])`](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) or other toolset wrappers. Custom `base_url` and URL-valued `client` values must use HTTPS. The toolset adds auth headers and appends the `tool-mode` query parameter for URL values when it is absent. It raises an error when the URL’s `tool-mode` conflicts with the configured mode because rewriting would invalidate signed URLs. When using `search_execute` with a signed URL, include `tool-mode=search_execute` before signing. Prebuilt clients are used as-is; configure their HTTPS transport, auth, account selection, and tool mode yourself. API reference ------------- [](https://pydantic.dev/docs/ai/harness/stackone/#api-reference) StackOne -------- [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOne) **Bases:** `AbstractCapability[AgentDepsT]` Actions on the user’s SaaS account (HRIS, ATS, CRM, and more) via StackOne. Connects an agent to one linked account’s actions over StackOne’s MCP endpoint, with authentication, tool filtering, and usage instructions. ### Attributes [](https://pydantic.dev/docs/ai/harness/stackone/#attributes) #### account\_id [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOne.account_id) The linked account to act on (one account is one provider connection). **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### actions [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOne.actions) `fnmatch` globs over full tool names (case-insensitive), e.g. `['*_list_*']`. Giving `actions` switches the default `tool_mode` to `individual`, where the globs apply. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `()` #### api\_key [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOne.api_key) StackOne API key. Defaults to the `STACKONE_API_KEY` environment variable. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `field(default=None, repr=False)` #### base\_url [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOne.base_url) HTTPS StackOne API host. Point at a regional or staging host if needed. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `STACKONE_BASE_URL` #### client [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOne.client) Replacement for the default `{base_url}/mcp` connection. URL values must use HTTPS; prebuilt clients keep their own transport, auth, and account selection, so `account_id` is not applied to them. **Type:** `MCPToolsetClient` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `field(default=None, repr=False)` #### description [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOne.description) Routing description used when the capability is loaded on demand. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `_DEFAULT_DESCRIPTION` #### id [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOne.id) Stable capability and toolset ID. Override it when one agent uses several StackOne accounts. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `_DEFAULT_ID` #### include\_instructions [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOne.include_instructions) Inject StackOne usage instructions into the system prompt. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `True` #### metadata [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOne.metadata) Metadata merged onto every tool, available to tool-selection machinery such as `CodeMode(tools={'code_mode': True})` or custom `prepare_tools` hooks. **Type:** [`Mapping`](https://docs.python.org/3/library/typing.html#typing.Mapping) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`object`](https://docs.python.org/3/glossary.html#term-object)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### tool\_mode [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOne.tool_mode) `individual` registers one tool per enabled action; `search_execute` registers two server-side meta-tools (search the catalog, execute an action by id) whose prompt footprint stays constant however large the catalog is. `None` picks `search_execute`, or `individual` when `actions` are given. **Type:** `ToolMode` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` ### Methods [](https://pydantic.dev/docs/ai/harness/stackone/#methods) #### from\_spec [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOne.from_spec) `@classmethod` def from_spec( cls, account_id: str, *, id: str | None = _DEFAULT_ID, description: str | None = _DEFAULT_DESCRIPTION, defer_loading: bool = False, api_key: str | None = None, base_url: str = STACKONE_BASE_URL, actions: str | Sequence[str] = (), tool_mode: ToolMode | None = None, include_instructions: bool = True, metadata: Mapping[str, object] | None = None, ) -> StackOne[AgentDepsT] Construct from serializable options, excluding the runtime-only `client`. ##### Returns [](https://pydantic.dev/docs/ai/harness/stackone/#returns) `StackOne`\[`AgentDepsT`\] #### get\_instructions [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOne.get_instructions) def get_instructions() -> str | None StackOne usage guidance; the underlying MCP toolset provides none itself. ##### Returns [](https://pydantic.dev/docs/ai/harness/stackone/#returns-1) [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### get\_serialization\_name [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOne.get_serialization_name) `@classmethod` def get_serialization_name(cls) -> str Return the agent-spec capability name. ##### Returns [](https://pydantic.dev/docs/ai/harness/stackone/#returns-2) [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### get\_toolset [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOne.get_toolset) def get_toolset() -> StackOneToolset[AgentDepsT] Build the StackOne toolset. ##### Returns [](https://pydantic.dev/docs/ai/harness/stackone/#returns-3) `StackOneToolset`\[`AgentDepsT`\] StackOneToolset --------------- [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOneToolset) **Bases:** `MCPToolset[AgentDepsT]` StackOne actions on one linked SaaS account, as an agent toolset. An `MCPToolset` connected to StackOne’s MCP endpoint. URL clients receive StackOne auth and account headers. Action filtering and metadata merging apply to the tools listed by the server. Prebuilt clients keep their own transport configuration. Instances support the full `MCPToolset` surface. Use the `StackOne` capability for usage instructions and agent-spec support. Use this class for toolset combinators such as `approval_required()`. ### Methods [](https://pydantic.dev/docs/ai/harness/stackone/#methods-1) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOneToolset.__init__) def __init__( *, account_id: str, api_key: str | None = None, base_url: str = STACKONE_BASE_URL, actions: Sequence[str] = (), tool_mode: ToolMode | None = None, metadata: Mapping[str, object] | None = None, client: MCPToolsetClient | None = None, id: str = 'stackone', ) -> None Build a StackOne MCP toolset. ##### Returns [](https://pydantic.dev/docs/ai/harness/stackone/#returns-4) [`None`](https://docs.python.org/3/library/constants.html#None) ##### Parameters [](https://pydantic.dev/docs/ai/harness/stackone/#parameters) **`account_id`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOneToolset.__init__(account_id)) Linked account used for StackOne requests. **`api_key`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOneToolset.__init__(api_key)) API key, or `STACKONE_API_KEY` when omitted. **`base_url`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) _Default:_ `STACKONE_BASE_URL` [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOneToolset.__init__(base_url)) HTTPS StackOne API host. **`actions`** : [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] _Default:_ `()` [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOneToolset.__init__(actions)) Case-insensitive globs over individual action tool names. Selects `individual`; incompatible with explicit `search_execute`. **`tool_mode`** : `ToolMode` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOneToolset.__init__(tool_mode)) Individual tools or the search/execute pair. Inferred when omitted. **`metadata`** : [`Mapping`](https://docs.python.org/3/library/typing.html#typing.Mapping) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`object`](https://docs.python.org/3/glossary.html#term-object)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOneToolset.__init__(metadata)) Metadata merged onto each tool definition. **`client`** : `MCPToolsetClient` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOneToolset.__init__(client)) URL, `FastMCP`, or prebuilt client accepted by `MCPToolset`. Non-URL clients keep their own transport, auth, and account selection, so `account_id` is not applied to them. **`id`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) _Default:_ `'stackone'` [](https://pydantic.dev/docs/ai/harness/stackone/#pydantic_ai_harness.stackone.StackOneToolset.__init__(id)) Toolset ID; use distinct values for multiple accounts. Was this page helpful? Thanks for your feedback! --- # Pydantic AI | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/overview/#_top) Pydantic AI =========== ![Pydantic AI](https://pydantic.dev/docs/ai/overview/img/pydantic-ai-dark.svg) ![Pydantic AI](https://pydantic.dev/docs/ai/overview/img/pydantic-ai-light.svg) _GenAI Agent Framework, the Pydantic way_ [![CI](https://github.com/pydantic/pydantic-ai/actions/workflows/ci.yml/badge.svg?event=push)](https://github.com/pydantic/pydantic-ai/actions/workflows/ci.yml?query=branch%3Amain) [![Coverage](https://coverage-badge.samuelcolvin.workers.dev/pydantic/pydantic-ai.svg)](https://coverage-badge.samuelcolvin.workers.dev/redirect/pydantic/pydantic-ai) [![PyPI](https://img.shields.io/pypi/v/pydantic-ai.svg)](https://pypi.python.org/pypi/pydantic-ai) [![versions](https://img.shields.io/pypi/pyversions/pydantic-ai.svg)](https://github.com/pydantic/pydantic-ai) [![license](https://img.shields.io/github/license/pydantic/pydantic-ai.svg)](https://github.com/pydantic/pydantic-ai/blob/main/LICENSE) [![Join Slack](https://img.shields.io/badge/Slack-Join%20Slack-4A154B?logo=slack)](https://logfire.pydantic.dev/docs/join-slack/) Pydantic AI is a Python agent framework designed to help you quickly, confidently, and painlessly build production grade applications and workflows with Generative AI. FastAPI revolutionized web development by offering an innovative and ergonomic design, built on the foundation of [Pydantic Validation](https://docs.pydantic.dev/) and modern Python features like type hints. Yet despite virtually every Python agent framework and LLM library using Pydantic Validation, when we began to use LLMs in [Pydantic Logfire](https://pydantic.dev/logfire) , we couldn’t find anything that gave us the same feeling. We built Pydantic AI with one simple aim: to bring that FastAPI feeling to GenAI app and agent development. Pydantic AI ships the agent loop, a composable [capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) system, and [built-in capabilities](https://pydantic.dev/docs/ai/capabilities/overview/#built-in-capabilities) for [thinking](https://pydantic.dev/docs/ai/capabilities/thinking/) , [web search](https://pydantic.dev/docs/ai/capabilities/web-search/) , [web fetch](https://pydantic.dev/docs/ai/capabilities/web-fetch/) , [image generation](https://pydantic.dev/docs/ai/capabilities/image-generation/) , [MCP](https://pydantic.dev/docs/ai/capabilities/mcp/) , [tool search](https://pydantic.dev/docs/ai/capabilities/tool-search/) , and more; [Pydantic AI Harness](https://pydantic.dev/docs/ai/harness/) is our official library of ready-made capabilities — code execution, file access, guardrails, sub-agent orchestration, and more — that you pick and choose to build coding agents, research assistants, and anything in between. Why use Pydantic AI ------------------- [](https://pydantic.dev/docs/ai/overview/#why-use-pydantic-ai) 1. **Built by the Pydantic Team**: [Pydantic Validation](https://docs.pydantic.dev/latest/) is the validation layer of the OpenAI SDK, the Google ADK, the Anthropic SDK, LangChain, LlamaIndex, AutoGPT, Transformers, CrewAI, Instructor and many more. _Why use the derivative when you can go straight to the source?_ 😃 2. **Model-agnostic**: Supports virtually every [model](https://pydantic.dev/docs/ai/models/overview/) and provider: OpenAI, Anthropic, Gemini, DeepSeek, Grok, Cohere, Mistral, and Perplexity; Azure AI Foundry, Amazon Bedrock, Google Cloud, Ollama, LiteLLM, Groq, OpenRouter, Together AI, Fireworks AI, Cerebras, Crusoe, Hugging Face, GitHub, Heroku, Vercel, Nebius, OVHcloud, Alibaba Cloud, SambaNova, Snowflake Cortex, and Z.AI. If your favorite model or provider is not listed, you can easily implement a [custom model](https://pydantic.dev/docs/ai/models/overview/#custom-models) . 3. **Seamless Observability**: Tightly [integrates](https://pydantic.dev/docs/ai/integrations/logfire/) with [Pydantic Logfire](https://pydantic.dev/logfire) , our general-purpose OpenTelemetry observability platform, for real-time debugging, evals-based performance monitoring, and behavior, tracing, and cost tracking. If you already have an observability platform that supports OTel, you can [use that too](https://pydantic.dev/docs/ai/integrations/logfire/#alternative-observability-backends) . 4. **Fully Type-safe**: Designed to give your IDE or AI coding agent as much context as possible for auto-completion and [type checking](https://pydantic.dev/docs/ai/core-concepts/agent/#static-type-checking) , moving entire classes of errors from runtime to write-time for a bit of that Rust “if it compiles, it works” feel. 5. **Powerful Evals**: Enables you to systematically test and [evaluate](https://pydantic.dev/docs/ai/evals/evals/) the performance and accuracy of the agentic systems you build, and monitor the performance over time in Pydantic Logfire. 6. **Extensible by Design**: Build agents from composable [capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) that bundle tools, hooks, instructions, and model settings into reusable units. Use built-in capabilities for [web search](https://pydantic.dev/docs/ai/capabilities/web-search/) , [thinking](https://pydantic.dev/docs/ai/capabilities/thinking/) , and [MCP](https://pydantic.dev/docs/ai/capabilities/mcp/) , pick from the [Pydantic AI Harness](https://pydantic.dev/docs/ai/harness/) capability library, build your own, or install [third-party capability packages](https://pydantic.dev/docs/ai/guides/extensibility/) . Define agents entirely in [YAML/JSON](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) — no code required. 7. **MCP and UI**: Integrates the [Model Context Protocol](https://pydantic.dev/docs/ai/mcp/overview/) and various [UI event stream](https://pydantic.dev/docs/ai/integrations/ui/overview/) standards to give your agent access to external tools and data and build interactive applications with streaming event-based communication. 8. **Human-in-the-Loop Tool Approval**: Easily lets you flag that certain tool calls [require approval](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/#human-in-the-loop-tool-approval) before they can proceed, possibly depending on tool call arguments, conversation history, or user preferences. 9. **Durable Execution**: Enables you to build [durable agents](https://pydantic.dev/docs/ai/capabilities/durable_execution/overview/) that can preserve their progress across transient API failures and application errors or restarts, and handle long-running, asynchronous, and human-in-the-loop workflows with production-grade reliability. 10. **Streamed Outputs**: Provides the ability to [stream](https://pydantic.dev/docs/ai/core-concepts/output/#streamed-results) structured output continuously, with immediate validation, ensuring real time access to generated data. 11. **Graph Support**: Provides a powerful way to define [graphs](https://pydantic.dev/docs/ai/graph/graph/) using type hints, for use in complex applications where standard control flow can degrade to spaghetti code. 12. **Realtime Voice**: Build [speech-to-speech agents](https://pydantic.dev/docs/ai/realtime/overview/) on native realtime models (OpenAI Realtime, Azure OpenAI, Gemini Live, and xAI Grok Voice) over a persistent bidirectional audio connection — in the browser over WebRTC or bridged through your backend — with the same tools, capabilities, and observability as any other agent. Realistically though, no list is going to be as convincing as [giving it a try](https://pydantic.dev/docs/ai/overview/#next-steps) and seeing how it makes you feel! **Sign up for our newsletter, _The Pydantic Stack_, with updates & tutorials on Pydantic AI, Logfire, and Pydantic:** Subscribe Hello World Example ------------------- [](https://pydantic.dev/docs/ai/overview/#hello-world-example) Here’s a minimal example of Pydantic AI: hello\_world.py from pydantic_ai import Agent agent = Agent( # (1) 'anthropic:claude-sonnet-4-6', instructions='Be concise, reply with one sentence.', # (2) ) result = agent.run_sync('Where does "hello world" come from?') # (3) print(result.output) """ The first known use of "hello, world" was in a 1974 textbook about the C programming language. """ We configure the agent to use [Anthropic's Claude Sonnet 4.6](https://pydantic.dev/docs/ai/api/models/anthropic/) model, but you can also set the model when running the agent. Register static [instructions](https://pydantic.dev/docs/ai/core-concepts/agent/#instructions) using a keyword argument to the agent. [Run the agent](https://pydantic.dev/docs/ai/core-concepts/agent/#running-agents) synchronously, starting a conversation with the LLM. _(This example is complete, it can be run “as is”, assuming you’ve [installed the `pydantic_ai` package](https://pydantic.dev/docs/ai/overview/install/) )_ The exchange will be very short: Pydantic AI will send the instructions and the user prompt to the LLM, and the model will return a text response. Not very interesting yet, but we can easily add [tools](https://pydantic.dev/docs/ai/tools-toolsets/tools/) , [dynamic instructions](https://pydantic.dev/docs/ai/core-concepts/agent/#instructions) , [structured outputs](https://pydantic.dev/docs/ai/core-concepts/output/) , or composable [capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) to build more powerful agents. Here’s the same agent with [thinking](https://pydantic.dev/docs/ai/capabilities/thinking/) and [web search](https://pydantic.dev/docs/ai/capabilities/web-search/) capabilities: hello\_world\_capabilities.py from pydantic_ai import Agent from pydantic_ai.capabilities import Thinking, WebSearch agent = Agent( 'anthropic:claude-sonnet-4-6', instructions='Be concise, reply with one sentence.', capabilities=[Thinking(), WebSearch(local='duckduckgo')], ) result = agent.run_sync('What was the mass of the largest meteorite found this year?') print(result.output) """ The largest meteorite recovered this year weighed approximately 7.6 kg, found in the Sahara Desert in January. """ Tools & Dependency Injection Example ------------------------------------ [](https://pydantic.dev/docs/ai/overview/#tools--dependency-injection-example) Here is a concise example using Pydantic AI to build a support agent for a bank: bank\_support.py from dataclasses import dataclass from pydantic import BaseModel, Field from pydantic_ai import Agent, RunContext from bank_database import DatabaseConn @dataclass class SupportDependencies: # (3) customer_id: int db: DatabaseConn # (12) class SupportOutput(BaseModel): # (13) support_advice: str = Field(description='Advice returned to the customer') block_card: bool = Field(description="Whether to block the customer's card") risk: int = Field(description='Risk level of query', ge=0, le=10) support_agent = Agent( # (1) 'openai:gpt-5.2', # (2) deps_type=SupportDependencies, output_type=SupportOutput, # (9) instructions=( # (4) 'You are a support agent in our bank, give the ' 'customer support and judge the risk level of their query.' ), ) @support_agent.instructions # (5) async def add_customer_name(ctx: RunContext[SupportDependencies]) -> str: customer_name = await ctx.deps.db.customer_name(id=ctx.deps.customer_id) return f"The customer's name is {customer_name!r}" @support_agent.tool # (6) async def customer_balance( ctx: RunContext[SupportDependencies], include_pending: bool ) -> float: """Returns the customer's current account balance.""" # (7) return await ctx.deps.db.customer_balance( id=ctx.deps.customer_id, include_pending=include_pending, ) ... # (11) async def main(): deps = SupportDependencies(customer_id=123, db=DatabaseConn()) result = await support_agent.run('What is my balance?', deps=deps) # (8) print(result.output) # (10) """ support_advice='Hello John, your current account balance, including pending transactions, is $123.45.' block_card=False risk=1 """ result = await support_agent.run('I just lost my card!', deps=deps) print(result.output) """ support_advice="I'm sorry to hear that, John. We are temporarily blocking your card to prevent unauthorized transactions." block_card=True risk=8 """ This [agent](https://pydantic.dev/docs/ai/core-concepts/agent/) will act as first-tier support in a bank. Agents are generic in the type of dependencies they accept and the type of output they return. In this case, the support agent has type `Agent[SupportDependencies, SupportOutput]`. Here we configure the agent to use [OpenAI's GPT-5 model](https://pydantic.dev/docs/ai/api/models/openai/) , you can also set the model when running the agent. The `SupportDependencies` dataclass is used to pass data, connections, and logic into the model that will be needed when running [instructions](https://pydantic.dev/docs/ai/core-concepts/agent/#instructions) and [tool](https://pydantic.dev/docs/ai/tools-toolsets/tools/) functions. Pydantic AI's system of dependency injection provides a [type-safe](https://pydantic.dev/docs/ai/core-concepts/agent/#static-type-checking) way to customise the behavior of your agents, and can be especially useful when running [unit tests](https://pydantic.dev/docs/ai/guides/testing/) and evals. Static [instructions](https://pydantic.dev/docs/ai/core-concepts/agent/#instructions) can be registered with the [`instructions` keyword argument](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.__init__) to the agent. Dynamic [instructions](https://pydantic.dev/docs/ai/core-concepts/agent/#instructions) can be registered with the [`@agent.instructions`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.instructions) decorator, and can make use of dependency injection. Dependencies are carried via the [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) argument, which is parameterized with the `deps_type` from above. If the type annotation here is wrong, static type checkers will catch it. The [`@agent.tool`](https://pydantic.dev/docs/ai/tools-toolsets/tools/) decorator let you register functions which the LLM may call while responding to a user. Again, dependencies are carried via [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) , any other arguments become the tool schema passed to the LLM. Pydantic is used to validate these arguments, and errors are passed back to the LLM so it can retry. The docstring of a tool is also passed to the LLM as the description of the tool. Parameter descriptions are [extracted](https://pydantic.dev/docs/ai/tools-toolsets/tools/#function-tools-and-schema) from the docstring and added to the parameter schema sent to the LLM. [Run the agent](https://pydantic.dev/docs/ai/core-concepts/agent/#running-agents) asynchronously, conducting a conversation with the LLM until a final response is reached. Even in this fairly simple case, the agent will exchange multiple messages with the LLM as tools are called to retrieve an output. The response from the agent will be guaranteed to be a `SupportOutput`. If validation fails [reflection](https://pydantic.dev/docs/ai/core-concepts/agent/#reflection-and-self-correction) , the agent is prompted to try again. The output will be validated with Pydantic to guarantee it is a `SupportOutput`, since the agent is generic, it'll also be typed as a `SupportOutput` to aid with static type checking. In a real use case, you'd add more tools and longer instructions to the agent to extend the context it's equipped with and support it can provide. This is a simple sketch of a database connection, used to keep the example short and readable. In reality, you'd be connecting to an external database (e.g. PostgreSQL) to get information about customers. This [Pydantic](https://docs.pydantic.dev/) model is used to constrain the structured data returned by the agent. From this simple definition, Pydantic builds the JSON Schema that tells the LLM how to return the data, and performs validation to guarantee the data is correct at the end of the run. Instrumentation with Pydantic Logfire ------------------------------------- [](https://pydantic.dev/docs/ai/overview/#instrumentation-with-pydantic-logfire) Even a simple agent with just a handful of tools can result in a lot of back-and-forth with the LLM, making it nearly impossible to be confident of what’s going on just from reading the code. To understand the flow of the above runs, we can watch the agent in action using Pydantic Logfire. To do this, we need to [set up Logfire](https://pydantic.dev/docs/ai/integrations/logfire/#using-logfire) , and add the following to our code: bank\_support\_with\_logfire.py ... from pydantic_ai import Agent, RunContext from bank_database import DatabaseConn import logfire logfire.configure() # (1) logfire.instrument_pydantic_ai() # (2) logfire.instrument_sqlite3() # (3) ... support_agent = Agent( 'openai:gpt-5.2', deps_type=SupportDependencies, output_type=SupportOutput, instructions=( 'You are a support agent in our bank, give the ' 'customer support and judge the risk level of their query.' ), ) Configure the Logfire SDK, this will fail if project is not set up. This will instrument all Pydantic AI agents used from here on out. To instrument only a specific agent, add an [`Instrumentation`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Instrumentation) entry to the agent's `capabilities=[...]`. In our demo, `DatabaseConn` uses [`sqlite3`](https://docs.python.org/3/library/sqlite3.html#module-sqlite3) to connect to a PostgreSQL database, so [`logfire.instrument_sqlite3()`](https://logfire.pydantic.dev/docs/integrations/databases/sqlite3/) is used to log the database queries. That’s enough to get the following view of your agent in action: Logfire instrumentation for the bank agent — [View in Logfire](https://logfire-eu.pydantic.dev/public-trace/a2957caa-b7b7-4883-a529-777742649004?spanId=31aade41ab896144) See [Monitoring and Performance](https://pydantic.dev/docs/ai/integrations/logfire/) to learn more. `llms.txt` ---------- [](https://pydantic.dev/docs/ai/overview/#llmstxt) The Pydantic AI documentation is available in the [llms.txt](https://llmstxt.org/) format. This format is defined in Markdown and suited for LLMs and AI coding assistants and agents. Two formats are available: * [`llms.txt`](https://ai.pydantic.dev/llms.txt) : a file containing a brief description of the project, along with links to the different sections of the documentation. The structure of this file is described in details [here](https://llmstxt.org/#format) . * [`llms-full.txt`](https://ai.pydantic.dev/llms-full.txt) : Similar to the `llms.txt` file, but every link content is included. Note that this file may be too large for some LLMs. As of today, these files are not automatically leveraged by IDEs or coding agents, but they will use it if you provide a link or the full text. Next Steps ---------- [](https://pydantic.dev/docs/ai/overview/#next-steps) To try Pydantic AI for yourself, [install it](https://pydantic.dev/docs/ai/overview/install/) and follow the instructions [in the examples](https://pydantic.dev/docs/ai/examples/setup/) . Read the [docs](https://pydantic.dev/docs/ai/core-concepts/agent/) to learn more about building applications with Pydantic AI. Read the [API Reference](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/) to understand Pydantic AI’s interface. Join [Slack](https://logfire.pydantic.dev/docs/join-slack/) or file an issue on [GitHub](https://github.com/pydantic/pydantic-ai/issues) if you have any questions. Was this page helpful? Thanks for your feedback! --- # Overview | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/realtime/overview/#_top) Overview ======== Pydantic AI’s realtime support lets an agent hold a live, spoken conversation. It streams the user’s audio to a speech-to-speech model and streams the model’s spoken reply back over one persistent connection, so latency is low and interruptions feel natural. A realtime session uses the same agent [tools](https://pydantic.dev/docs/ai/realtime/tools/) , [dependencies](https://pydantic.dev/docs/ai/core-concepts/dependencies/) , [instructions](https://pydantic.dev/docs/ai/core-concepts/agent/#instructions) , [message history](https://pydantic.dev/docs/ai/core-concepts/message-history/) , [capabilities](https://pydantic.dev/docs/ai/realtime/capabilities/) , [usage limits](https://pydantic.dev/docs/ai/core-concepts/agent/#usage-limits) , and [observability](https://pydantic.dev/docs/ai/realtime/observability/) as the rest of Pydantic AI, and that’s the point: mid-call the agent can look up an order, check availability, or act on the logged-in user’s data with the same tools and dependencies a text agent would use. The call itself becomes ordinary message history that you can [hand to `Agent.run()`](https://pydantic.dev/docs/ai/realtime/history/#handing-off-to-a-text-agent) for summarization or structured follow-up, the same code runs against [four providers](https://pydantic.dev/docs/ai/realtime/overview/#provider-support) , and usage limits and [Logfire](https://pydantic.dev/docs/ai/integrations/logfire/) tracing are built in. Your application owns the audio transport — bridged through your backend, or [browser-direct over WebRTC](https://pydantic.dev/docs/ai/realtime/deployment/#browser-webrtc-server-sideband) on OpenAI and Azure — while Pydantic AI runs the provider-agnostic agent loop. Quickstart ---------- [](https://pydantic.dev/docs/ai/realtime/overview/#quickstart) Install Pydantic AI with the OpenAI realtime dependencies, and set `OPENAI_API_KEY`: * [pip](https://pydantic.dev/docs/ai/realtime/overview/#tab-panel-176) * [uv](https://pydantic.dev/docs/ai/realtime/overview/#tab-panel-177) Terminal pip install "pydantic-ai-slim[openai-realtime]" Terminal uv add "pydantic-ai-slim[openai-realtime]" A complete voice agent is one agent, one session, and three small loops — microphone in, speaker out, and a transcript log. The model hears the user, calls your tool on your backend, and answers out loud: reservations.py import asyncio import contextlib from collections.abc import AsyncIterator from pydantic_ai import Agent from pydantic_ai.realtime import RealtimeSession agent = Agent(instructions='You take reservations for The Terrace. Keep replies short.') @agent.tool_plain async def check_availability(day: str, party_size: int) -> str: """Check whether a table is free.""" return f'One table for {party_size} is free at 7 pm {day}.' async def stream_microphone(session: RealtimeSession) -> None: ... # capture signed 16-bit mono PCM chunks and `await session.send_audio(chunk)` async def play_audio(chunks: AsyncIterator[bytes]) -> None: async for chunk in chunks: ... # write the PCM chunk to your speaker async def main(): async with agent.realtime('openai:gpt-realtime').session() as session: microphone = asyncio.create_task(stream_microphone(session)) speaker = asyncio.create_task(play_audio(session.stream_audio())) async for part in session.stream_transcripts(): print(f'{part.speaker}: {part.transcript}') #> user: Hi! Do you have a table for two tomorrow night? #> assistant: We do: 7 pm, table for two. Want me to book it? if part.speaker == 'assistant': break # keep listening in a real call; we stop after one exchange # Leaving the `async with` block closes the session, which ends the speaker's audio stream — # but the microphone reads an external source, so stop it explicitly. microphone.cancel() with contextlib.suppress(asyncio.CancelledError): await microphone await speaker if __name__ == '__main__': asyncio.run(main()) _(This example is complete, it can be run “as is” — after filling in the two audio placeholders, which depend on your audio stack)_ Capture and play at the sample rates the model expects — they’re reported by the model’s profile and can differ between input and output (see [Provider support](https://pydantic.dev/docs/ai/realtime/overview/#provider-support) below). The [voice assistant example](https://pydantic.dev/docs/ai/examples/realtime/realtime-voice/) fills the placeholders in with `sounddevice` for a runnable microphone-and-speaker loop; the [text-to-audio example](https://pydantic.dev/docs/ai/examples/realtime/realtime-text-to-audio/) skips audio input entirely by sending a text prompt and saving the spoken reply to a WAV file. How sessions work ----------------- [](https://pydantic.dev/docs/ai/realtime/overview/#how-sessions-work) Your backend opens the provider connection and runs a [`RealtimeSession`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession) . Stream content in with [`send()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.send) or [`send_audio()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.send_audio) , and iterate the session for its [event stream](https://pydantic.dev/docs/ai/realtime/events/) — content, tool, turn, error, and reconnect events — or consume the dedicated [`stream_audio()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.stream_audio) and [`stream_transcripts()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.stream_transcripts) views as the quickstart does. device ↔ media bridge ↔ RealtimeSession ↔ provider ├── typed tools └── message history (your backend) The _media bridge_ is whatever moves audio between the user’s device and your backend — a browser WebSocket or a telephony bridge. It’s how you deploy this beyond a local microphone; see [Connecting a frontend](https://pydantic.dev/docs/ai/realtime/deployment/) for each shape. On OpenAI and Azure the browser can instead exchange media with the provider directly over [WebRTC](https://pydantic.dev/docs/ai/realtime/deployment/#browser-webrtc-server-sideband) , with your backend running this same loop over a control-plane sideband rather than a media bridge. Learn by task ------------- [](https://pydantic.dev/docs/ai/realtime/overview/#learn-by-task) * [Audio, images, and transcripts](https://pydantic.dev/docs/ai/realtime/audio/) covers the PCM wire contract, playback, captions, input transcription, and image input. * [Events](https://pydantic.dev/docs/ai/realtime/events/) covers the session event vocabulary, which events are shared with standard runs, and the turn boundary. * [Turns and interruptions](https://pydantic.dev/docs/ai/realtime/turns/) covers automatic turn detection, barge-in, output truncation, and push-to-talk. * [Tools](https://pydantic.dev/docs/ai/realtime/tools/) covers function tools, provider-native tools, concurrency, approval, and delegation during a call. * [Capabilities and hooks](https://pydantic.dev/docs/ai/realtime/capabilities/) covers how capabilities and their hooks map onto a session. * [History and handoff](https://pydantic.dev/docs/ai/realtime/history/) covers retained transcripts, audio and images, session seeding, and continuing with a standard text agent. * [Connecting a frontend](https://pydantic.dev/docs/ai/realtime/deployment/) covers the transport shapes between user devices and your backend. * [Connection lifecycle](https://pydantic.dev/docs/ai/realtime/lifecycle/) covers the session lifecycle, reconnection, session limits, and errors. * [Usage and observability](https://pydantic.dev/docs/ai/realtime/observability/) covers usage limits, cost accounting, Logfire, and gateway trace propagation. * [Troubleshooting](https://pydantic.dev/docs/ai/realtime/troubleshooting/) indexes common problems by symptom. * The [API reference](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/) lists session and codec types and explains how to implement another provider. Provider support ---------------- [](https://pydantic.dev/docs/ai/realtime/overview/#provider-support) All providers implement the same [`RealtimeModel`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModel) interface. Provider pages are the canonical source for installation, model names, settings, feature support, and quirks: | Provider | Audio output | Image input | Text output | [Browser WebRTC](https://pydantic.dev/docs/ai/realtime/deployment/#browser-webrtc-server-sideband) | Async tool calls | [Thinking](https://pydantic.dev/docs/ai/capabilities/thinking/) | State-restoring reconnect | | --- | --- | --- | --- | --- | --- | --- | --- | | [OpenAI](https://pydantic.dev/docs/ai/realtime/openai/) | ✓ | ✓ | ✓ | ✓ | ✓ | `gpt-realtime-2*` models | Replays local history | | [Azure OpenAI](https://pydantic.dev/docs/ai/realtime/azure/) | ✓ | ✓ | ✓ | ✓ | ✓ | `gpt-realtime-2*` models | Replays local history | | [Google Gemini](https://pydantic.dev/docs/ai/realtime/gemini/) | ✓ | ✓ | ✗ | ✗ | Opt-in, native-audio models | ✓ | ✓, when enabled | | [xAI](https://pydantic.dev/docs/ai/realtime/xai/) | ✓ | ✗ | ✗ | ✗ | ✗ | `grok-voice-latest` and `-think-` models | ✓ | For portable branching, inspect [`RealtimeModel.profile`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModel.profile) or [`RealtimeSession.profile`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.profile) : the [`RealtimeModelProfile`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelProfile) reports the audio sample rates to capture and play at, plus one flag per capability in the table above and beyond. Profiles resolve the same way as for a standard [`Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) (see [Inspecting a model’s profile](https://pydantic.dev/docs/ai/models/overview/#inspecting-a-models-profile) ) — defaults, then the provider’s knowledge of the model name, then your `profile=` argument on top. Pass `profile=` when the model name doesn’t identify the model and the inferred facts are wrong, most often with an Azure deployment named something other than its model: from pydantic_ai.realtime.azure import AzureRealtimeModel # The deployment serves a reasoning model, but nothing in its name says so. model = AzureRealtimeModel('voice-prod', profile={'supports_thinking': True}) A partial dict is merged over the resolved profile; pass a callable `(resolved) -> RealtimeModelProfile` instead to replace it wholesale. Shared settings --------------- [](https://pydantic.dev/docs/ai/realtime/overview/#shared-settings) Realtime sessions have their own settings type, playing the role that [model run settings](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) play for standard runs: [`RealtimeModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelSettings) defines the settings shared across realtime providers, from `tool_choice` to [`turn_detection`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.TurnDetection) . Set defaults with `settings=` on the realtime model constructor, or pass `realtime(model_settings=...)` for one session; per-session values override model defaults: from pydantic_ai import Agent from pydantic_ai.realtime import RealtimeModelSettings agent = Agent(instructions='You are a helpful voice assistant.') realtime = agent.realtime( 'openai:gpt-realtime', model_settings=RealtimeModelSettings(output_modality='audio') ) Voices and detailed controls are provider-specific — `openai_voice`, `google_voice`, `xai_voice` and friends live on the corresponding provider settings classes, with defaults and limitations on the provider pages. The agent’s regular `model_settings` and capability `get_model_settings()` contributions do not configure realtime sessions. Unsupported shared settings are ignored, matching request-response models, with one deliberate exception: Relationship to standard agent runs ----------------------------------- [](https://pydantic.dev/docs/ai/realtime/overview/#relationship-to-standard-agent-runs) `Agent.realtime()` is the long-lived, bidirectional sibling of [`run()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.run) and [`iter()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.iter) , and its parameters mirror theirs: agent.realtime( model, # 'openai:gpt-realtime', or a RealtimeModel instance deps=..., # dependencies, as in run()/iter() model_settings=..., # RealtimeModelSettings instructions=..., # combined with the agent's instructions toolsets=..., # additional toolsets for the session capabilities=..., # additional capabilities for the session usage=..., usage_limits=..., message_history=..., # prior conversation to seed the session with ) It accepts the same [dependencies](https://pydantic.dev/docs/ai/core-concepts/dependencies/) , [instructions](https://pydantic.dev/docs/ai/core-concepts/agent/#instructions) , [toolsets](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) , [capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) , [usage limits](https://pydantic.dev/docs/ai/core-concepts/agent/#usage-limits) , and [`message_history`](https://pydantic.dev/docs/ai/core-concepts/message-history/) as a standard run. Input arrives through the live session instead of a single `user_prompt`: | Standard-run feature | In a realtime session | | --- | --- | | Function tools and [tool hooks](https://pydantic.dev/docs/ai/realtime/capabilities/#capability-stages-in-a-session) | ✓ — validation, retries, and execution hooks run as in a standard run | | [Run hooks](https://pydantic.dev/docs/ai/realtime/capabilities/#run-hooks)
(`before_run`, `after_run`, `wrap_run`, `on_run_error`) | ✓ — once around the session | | [Capabilities](https://pydantic.dev/docs/ai/realtime/capabilities/)
, including third-party | ✓ — resolved once at connect | | [Event stream](https://pydantic.dev/docs/ai/realtime/events/) | ✓ — iterate the session, or attach [`ProcessEventStream`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ProcessEventStream) | | `output_type` and output validators | ✗ — [delegate to a text agent](https://pydantic.dev/docs/ai/realtime/tools/#delegating-work-during-a-call) | | Graph node and model-request hooks (e.g. `before_model_request`) | ✗ — no agent graph | | History processors at seeding | ✗ — [preprocess before opening](https://pydantic.dev/docs/ai/realtime/capabilities/#seeded-history-is-not-processed) | | `event_stream_handler` parameter | ✗ — use [`ProcessEventStream`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ProcessEventStream) | See [Capabilities and hooks](https://pydantic.dev/docs/ai/realtime/capabilities/) for the full mapping, and [hand off to a text agent](https://pydantic.dev/docs/ai/realtime/history/#handing-off-to-a-text-agent) for structured output or deeper reasoning. Other ways to build voice ------------------------- [](https://pydantic.dev/docs/ai/realtime/overview/#other-ways-to-build-voice) The same realtime loop deploys to a browser or phone over [WebRTC or a WebSocket relay](https://pydantic.dev/docs/ai/realtime/deployment/) without changing the agent code. If the realtime agent loop isn’t the right fit for a product, two alternatives sit outside it: * **Batch STT → text agent → TTS.** Compose a standard [agent](https://pydantic.dev/docs/ai/core-concepts/agent/) with your own speech-to-text and text-to-speech services when you want a specific text model, structured output, or independently chosen speech components. * **Browser directly to the provider.** A provider-native, UI-only experience using an ephemeral token: the provider’s own SDK owns the session, so there is no server-side agent loop, tools, or shared history — unlike the [WebRTC sideband](https://pydantic.dev/docs/ai/realtime/deployment/#browser-webrtc-server-sideband) , where the browser owns the media but your backend still runs the agent. Pydantic AI can still power separate backend workflows. Limitations ----------- [](https://pydantic.dev/docs/ai/realtime/overview/#limitations) | Limitation | Tracking | | --- | --- | | SIP is not built in; bridge telephony through a provider such as Twilio. | [Connecting a frontend](https://pydantic.dev/docs/ai/realtime/deployment/#siptelephony-bridge) | | New tools cannot be advertised mid-session, so `defer_loading=True` tools and tool-contributing capabilities are [rejected](https://pydantic.dev/docs/ai/realtime/capabilities/#deferred-capability-loading)
. | [#7288](https://github.com/pydantic/pydantic-ai/issues/7288) | | Realtime-specific exchange hooks are not yet available; use supported [tool hooks](https://pydantic.dev/docs/ai/realtime/capabilities/)
and [session events](https://pydantic.dev/docs/ai/realtime/events/)
. | [#7190](https://github.com/pydantic/pydantic-ai/issues/7190)
, [#7191](https://github.com/pydantic/pydantic-ai/issues/7191) | | Provider resumption handles cannot be persisted and resumed in another process. | [#7302](https://github.com/pydantic/pydantic-ai/issues/7302) | | Dynamic instructions are resolved once when the session connects. | [#7303](https://github.com/pydantic/pydantic-ai/issues/7303) | | History processors do not transform `message_history` before realtime seeding; [preprocess it](https://pydantic.dev/docs/ai/realtime/capabilities/#seeded-history-is-not-processed)
before opening the session when filtering or redaction is required. | [#7299](https://github.com/pydantic/pydantic-ai/issues/7299) | | Interactive human-in-the-loop tool approval is not supported: a [`HandleDeferredToolCalls`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.HandleDeferredToolCalls)
handler resolves approvals [from policy, immediately](https://pydantic.dev/docs/ai/realtime/tools/#deferred-and-approval-required-tools)
. | [#7301](https://github.com/pydantic/pydantic-ai/issues/7301) | | [`RunContext.enqueue()`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.enqueue)
accepts [one plain-text prompt per call](https://pydantic.dev/docs/ai/realtime/tools/#enqueuing-prompts-from-tools)
, unlike its [standard-run form](https://pydantic.dev/docs/ai/core-concepts/message-history/#injecting-messages-mid-run)
. | [#7300](https://github.com/pydantic/pydantic-ai/issues/7300) | | Gemini Live tool results are JSON-only: binary content attached to a [tool return](https://pydantic.dev/docs/ai/realtime/tools/#function-tools)
raises rather than being delivered. | [#7362](https://github.com/pydantic/pydantic-ai/issues/7362) | Was this page helpful? Thanks for your feedback! --- # retries | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#_top) retries ======= Retries utilities based on tenacity, especially for HTTP requests. This module provides HTTP transport wrappers and wait strategies that integrate with the tenacity library to add retry capabilities to HTTP requests. The transports can be used with HTTP clients that support custom transports (such as httpx), while the wait strategies can be used with any tenacity retry decorator. The module includes: * TenacityTransport: Synchronous HTTP transport with retry capabilities * AsyncTenacityTransport: Asynchronous HTTP transport with retry capabilities * wait\_retry\_after: Wait strategy that respects HTTP Retry-After headers AsyncTenacityTransport ---------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.AsyncTenacityTransport) **Bases:** `AsyncBaseTransport` Asynchronous HTTP transport with tenacity-based retry functionality. This transport wraps another AsyncBaseTransport and adds retry capabilities using the tenacity library. It can be configured to retry requests based on various conditions such as specific exception types, response status codes, or custom validation logic. The transport works by intercepting HTTP requests and responses, allowing the tenacity controller to determine when and how to retry failed requests. The validate\_response function can be used to convert HTTP responses into exceptions that trigger retries. ### Constructor Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#constructor-parameters) **`wrapped`** : `AsyncBaseTransport` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.AsyncTenacityTransport.__init__(wrapped)) The underlying async transport to wrap and add retry functionality to. **`config`** : `RetryConfig` [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.AsyncTenacityTransport.__init__(config)) The arguments to use for the tenacity `retry` decorator, including retry conditions, wait strategy, stop conditions, etc. See the tenacity docs for more info. **`validate_response`** : [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[`Response`\], [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.AsyncTenacityTransport.__init__(validate_response)) Optional callable that takes a Response and can raise an exception to be handled by the controller if the response should trigger a retry. Common use case is to raise exceptions for certain HTTP status codes. If None, no response validation is performed. ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#methods) #### handle\_async\_request [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.AsyncTenacityTransport.handle_async_request) `@async` def handle_async_request(request: Request) -> Response Handle an async HTTP request with retry logic. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#returns) `Response` — The HTTP response. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#parameters) **`request`** : `Request` [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.AsyncTenacityTransport.handle_async_request(request)) The HTTP request to handle. ##### Raises [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#raises) * `RuntimeError` — If the retry controller did not make any attempts. * `Exception` — Any exception raised by the wrapped transport or validation function. RetryConfig ----------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.RetryConfig) **Bases:** [`TypedDict`](https://docs.python.org/3/library/typing.html#typing.TypedDict) The configuration for tenacity-based retrying. These are precisely the arguments to the tenacity `retry` decorator, and they are generally used internally by passing them to that decorator via `@retry(**config)` or similar. All fields are optional, and if not provided, the default values from the `tenacity.retry` decorator will be used. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#attributes) #### after [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.RetryConfig.after) A callable that is called after each retry attempt. Tenacity’s default for this argument is `tenacity.after.after_nothing`. **Type:** [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[`RetryCallState`\], [`None`](https://docs.python.org/3/library/constants.html#None)\ | [`Awaitable`](https://docs.python.org/3/library/typing.html#typing.Awaitable)\ \[[`None`](https://docs.python.org/3/library/constants.html#None)\ \]\] #### before [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.RetryConfig.before) A callable that is called before each retry attempt. Tenacity’s default for this argument is `tenacity.before.before_nothing`. **Type:** [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[`RetryCallState`\], [`None`](https://docs.python.org/3/library/constants.html#None)\ | [`Awaitable`](https://docs.python.org/3/library/typing.html#typing.Awaitable)\ \[[`None`](https://docs.python.org/3/library/constants.html#None)\ \]\] #### before\_sleep [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.RetryConfig.before_sleep) An optional callable that is called before sleeping between retries. Tenacity’s default for this argument is `None`. **Type:** [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[`RetryCallState`\], [`None`](https://docs.python.org/3/library/constants.html#None)\ | [`Awaitable`](https://docs.python.org/3/library/typing.html#typing.Awaitable)\ \[[`None`](https://docs.python.org/3/library/constants.html#None)\ \]\] | [`None`](https://docs.python.org/3/library/constants.html#None) #### reraise [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.RetryConfig.reraise) Whether to reraise the last exception if the retry attempts are exhausted, or raise a RetryError instead. Tenacity’s default for this argument is `False`. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### retry [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.RetryConfig.retry) A retry strategy to determine which exceptions should trigger a retry. Tenacity’s default for this argument is `tenacity.retry.retry_if_exception_type()`. **Type:** `SyncRetryBaseT` | `RetryBaseT` #### retry\_error\_callback [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.RetryConfig.retry_error_callback) An optional callable that is called when the retry attempts are exhausted and `reraise` is False. Tenacity’s default for this argument is `None`. **Type:** [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[`RetryCallState`\], [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ | [`Awaitable`](https://docs.python.org/3/library/typing.html#typing.Awaitable)\ \[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \]\] | [`None`](https://docs.python.org/3/library/constants.html#None) #### retry\_error\_cls [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.RetryConfig.retry_error_cls) The exception class to raise when the retry attempts are exhausted and `reraise` is False. Tenacity’s default for this argument is `tenacity.RetryError`. **Type:** [`type`](https://docs.python.org/3/glossary.html#term-type) \[`RetryError`\] #### sleep [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.RetryConfig.sleep) A sleep strategy to use for sleeping between retries. Tenacity’s default for this argument is `tenacity.nap.sleep`. **Type:** [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[[`int`](https://docs.python.org/3/library/functions.html#int)\ | [`float`](https://docs.python.org/3/library/functions.html#float)\ \], [`None`](https://docs.python.org/3/library/constants.html#None)\ | [`Awaitable`](https://docs.python.org/3/library/typing.html#typing.Awaitable)\ \[[`None`](https://docs.python.org/3/library/constants.html#None)\ \]\] #### stop [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.RetryConfig.stop) A stop strategy to determine when to stop retrying. Tenacity’s default for this argument is `tenacity.stop.stop_never`. **Type:** `StopBaseT` #### wait [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.RetryConfig.wait) A wait strategy to determine how long to wait between retries. Tenacity’s default for this argument is `tenacity.wait.wait_none`. **Type:** `WaitBaseT` TenacityTransport ----------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.TenacityTransport) **Bases:** `BaseTransport` Synchronous HTTP transport with tenacity-based retry functionality. This transport wraps another BaseTransport and adds retry capabilities using the tenacity library. It can be configured to retry requests based on various conditions such as specific exception types, response status codes, or custom validation logic. The transport works by intercepting HTTP requests and responses, allowing the tenacity controller to determine when and how to retry failed requests. The validate\_response function can be used to convert HTTP responses into exceptions that trigger retries. ### Constructor Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#constructor-parameters-1) **`wrapped`** : `BaseTransport` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.TenacityTransport.__init__(wrapped)) The underlying transport to wrap and add retry functionality to. **`config`** : `RetryConfig` [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.TenacityTransport.__init__(config)) The arguments to use for the tenacity `retry` decorator, including retry conditions, wait strategy, stop conditions, etc. See the tenacity docs for more info. **`validate_response`** : [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[`Response`\], [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.TenacityTransport.__init__(validate_response)) Optional callable that takes a Response and can raise an exception to be handled by the controller if the response should trigger a retry. Common use case is to raise exceptions for certain HTTP status codes. If None, no response validation is performed. ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#methods-1) #### handle\_request [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.TenacityTransport.handle_request) def handle_request(request: Request) -> Response Handle an HTTP request with retry logic. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#returns-1) `Response` — The HTTP response. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#parameters-1) **`request`** : `Request` [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.TenacityTransport.handle_request(request)) The HTTP request to handle. ##### Raises [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#raises-1) * `RuntimeError` — If the retry controller did not make any attempts. * `Exception` — Any exception raised by the wrapped transport or validation function. wait\_retry\_after ------------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.wait_retry_after) def wait_retry_after( fallback_strategy: Callable[[RetryCallState], float] | None = None, max_wait: float = 300, ) -> Callable[[RetryCallState], float] Create a tenacity-compatible wait strategy that respects HTTP Retry-After headers. This wait strategy checks if the exception contains an HTTPStatusError with a Retry-After header, and if so, waits for the time specified in the header. If no header is present or parsing fails, it falls back to the provided strategy. The Retry-After header can be in two formats: * An integer representing seconds to wait * An HTTP date string representing when to retry ### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#returns-2) [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[`RetryCallState`\], [`float`](https://docs.python.org/3/library/functions.html#float)\ \] — A wait function that can be used with tenacity retry decorators. ### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#parameters-2) **`fallback_strategy`** : [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[`RetryCallState`\], [`float`](https://docs.python.org/3/library/functions.html#float)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.wait_retry_after(fallback_strategy)) Wait strategy to use when no Retry-After header is present or parsing fails. Defaults to exponential backoff with max 60s. **`max_wait`** : [`float`](https://docs.python.org/3/library/functions.html#float) _Default:_ `300` [](https://pydantic.dev/docs/ai/api/pydantic-ai/retries/#pydantic_ai.retries.wait_retry_after(max_wait)) Maximum time to wait in seconds, regardless of header value. Defaults to 300 (5 minutes). Was this page helpful? Thanks for your feedback! --- # Dataset Management | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#_top) Dataset Management ================== Create, save, load, and generate evaluation datasets. Creating Datasets ----------------- [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#creating-datasets) ### From Code [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#from-code) Define datasets directly in Python: from typing import Any from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import EqualsExpected, IsInstance dataset = Dataset[str, str, Any]( name='my_eval_suite', cases=[\ Case(\ name='test_1',\ inputs='input 1',\ expected_output='output 1',\ ),\ Case(\ name='test_2',\ inputs='input 2',\ expected_output='output 2',\ ),\ ], evaluators=[\ IsInstance(type_name='str'),\ EqualsExpected(),\ ], ) ### Adding Cases Dynamically [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#adding-cases-dynamically) from typing import Any from pydantic_evals import Dataset from pydantic_evals.evaluators import IsInstance dataset = Dataset[str, str, Any](name='dynamic_dataset', cases=[], evaluators=[]) # Add cases one at a time dataset.add_case( name='dynamic_case', inputs='test input', expected_output='test output', ) # Add evaluators dataset.add_evaluator(IsInstance(type_name='str')) Saving Datasets --------------- [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#saving-datasets) ### Save to YAML [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#save-to-yaml) from typing import Any from pydantic_evals import Case, Dataset dataset = Dataset[str, str, Any](name='my_eval_suite', cases=[Case(name='test', inputs='example')]) dataset.to_file('my_dataset.yaml') # Also saves schema file: my_dataset_schema.json Output (`my_dataset.yaml`): # yaml-language-server: $schema=my_dataset_schema.json name: my_eval_suite cases: - name: test_1 inputs: input 1 expected_output: output 1 evaluators: - EqualsExpected - name: test_2 inputs: input 2 expected_output: output 2 evaluators: - EqualsExpected evaluators: - IsInstance: str ### Save to JSON [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#save-to-json) from typing import Any from pydantic_evals import Case, Dataset dataset = Dataset[str, str, Any](name='my_eval_suite', cases=[Case(name='test', inputs='example')]) dataset.to_file('my_dataset.json') # Also saves schema file: my_dataset_schema.json ### Custom Schema Path [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#custom-schema-path) from pathlib import Path from typing import Any from pydantic_evals import Case, Dataset dataset = Dataset[str, str, Any](name='my_eval_suite', cases=[Case(name='test', inputs='example')]) # Custom schema location Path('data').mkdir(exist_ok=True) Path('data/schemas').mkdir(parents=True, exist_ok=True) dataset.to_file( 'data/my_dataset.yaml', schema_path='schemas/my_schema.json', ) # No schema file dataset.to_file('my_dataset.yaml', schema_path=None) Loading Datasets ---------------- [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#loading-datasets) ### From YAML/JSON [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#from-yamljson) from typing import Any from pydantic_evals import Dataset # Infers format from extension dataset = Dataset[str, str, Any].from_file('my_dataset.yaml') dataset = Dataset[str, str, Any].from_file('my_dataset.json') # Explicit format for non-standard extensions dataset = Dataset[str, str, Any].from_file('data.txt', fmt='yaml') ### From String [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#from-string) from typing import Any from pydantic_evals import Dataset yaml_content = """ name: my_tests cases: - name: test inputs: hello expected_output: HELLO evaluators: - EqualsExpected """ dataset = Dataset[str, str, Any].from_text(yaml_content, fmt='yaml') ### From Dict [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#from-dict) from typing import Any from pydantic_evals import Dataset data = { 'name': 'my_tests', 'cases': [\ {\ 'name': 'test',\ 'inputs': 'hello',\ 'expected_output': 'HELLO',\ },\ ], 'evaluators': [{'EqualsExpected': {}}], } dataset = Dataset[str, str, Any].from_dict(data) ### With Custom Evaluators [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#with-custom-evaluators) When loading datasets that use custom evaluators, you must pass them to `from_file()`: from dataclasses import dataclass from typing import Any from pydantic_evals import Dataset from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class MyCustomEvaluator(Evaluator): threshold: float = 0.5 def evaluate(self, ctx: EvaluatorContext) -> bool: return True # Load with custom evaluator registry dataset = Dataset[str, str, Any].from_file( 'my_dataset.yaml', custom_evaluator_types=[MyCustomEvaluator], ) For complete details on serialization with custom evaluators, see [Dataset Serialization](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/) . Generating Datasets ------------------- [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#generating-datasets) Pydantic Evals allows you to generate test datasets using LLMs with [`generate_dataset`](https://pydantic.dev/docs/ai/api/pydantic_evals/generation/#pydantic_evals.generation.generate_dataset) . Datasets can be generated in either JSON or YAML format, in both cases a JSON schema file is generated alongside the dataset and referenced in the dataset, so you should get type checking and auto-completion in your editor. generate\_dataset\_example.py from __future__ import annotations from pathlib import Path from pydantic import BaseModel, Field from pydantic_evals import Dataset from pydantic_evals.generation import generate_dataset class QuestionInputs(BaseModel, use_attribute_docstrings=True): # (1) """Model for question inputs.""" question: str """A question to answer""" context: str | None = None """Optional context for the question""" class AnswerOutput(BaseModel, use_attribute_docstrings=True): # (2) """Model for expected answer outputs.""" answer: str """The answer to the question""" confidence: float = Field(ge=0, le=1) """Confidence level (0-1)""" class MetadataType(BaseModel, use_attribute_docstrings=True): # (3) """Metadata model for test cases.""" difficulty: str """Difficulty level (easy, medium, hard)""" category: str """Question category""" async def main(): dataset = await generate_dataset( # (4) dataset_type=Dataset[QuestionInputs, AnswerOutput, MetadataType], n_examples=2, extra_instructions=""" Generate question-answer pairs about world capitals and landmarks. Make sure to include both easy and challenging questions. """, ) output_file = Path('questions_cases.yaml') dataset.to_file(output_file) # (5) print(output_file.read_text(encoding='utf-8')) """ # yaml-language-server: $schema=questions_cases_schema.json name: generated cases: - name: Easy Capital Question inputs: question: What is the capital of France? context: null metadata: difficulty: easy category: Geography expected_output: answer: Paris confidence: 0.95 evaluators: - EqualsExpected - name: Challenging Landmark Question inputs: question: Which world-famous landmark is located on the banks of the Seine River? context: null metadata: difficulty: hard category: Landmarks expected_output: answer: Eiffel Tower confidence: 0.9 evaluators: - EqualsExpected evaluators: [] report_evaluators: [] """ Define the schema for the inputs to the task. Define the schema for the expected outputs of the task. Define the schema for the metadata of the test cases. Call [`generate_dataset`](https://pydantic.dev/docs/ai/api/pydantic_evals/generation/#pydantic_evals.generation.generate_dataset) to create a [`Dataset`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset) with 2 cases confirming to the schema. Save the dataset to a YAML file, this will also write `questions_cases_schema.json` with the schema JSON schema for `questions_cases.yaml` to make editing easier. The magic `yaml-language-server` comment is supported by at least vscode, jetbrains/pycharm (more details [here](https://github.com/redhat-developer/yaml-language-server#using-inlined-schema) ). _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main(answer))` to run `main`)_ You can also write datasets as JSON files: generate\_dataset\_example\_json.py from pathlib import Path from pydantic_evals import Dataset from pydantic_evals.generation import generate_dataset from generate_dataset_example import AnswerOutput, MetadataType, QuestionInputs async def main(): dataset = await generate_dataset( # (1) dataset_type=Dataset[QuestionInputs, AnswerOutput, MetadataType], n_examples=2, extra_instructions=""" Generate question-answer pairs about world capitals and landmarks. Make sure to include both easy and challenging questions. """, ) output_file = Path('questions_cases.json') dataset.to_file(output_file) # (2) print(output_file.read_text(encoding='utf-8')) """ { "$schema": "questions_cases_schema.json", "name": "generated", "cases": [\ {\ "name": "Easy Capital Question",\ "inputs": {\ "question": "What is the capital of France?",\ "context": null\ },\ "metadata": {\ "difficulty": "easy",\ "category": "Geography"\ },\ "expected_output": {\ "answer": "Paris",\ "confidence": 0.95\ },\ "evaluators": [\ "EqualsExpected"\ ]\ },\ {\ "name": "Challenging Landmark Question",\ "inputs": {\ "question": "Which world-famous landmark is located on the banks of the Seine River?",\ "context": null\ },\ "metadata": {\ "difficulty": "hard",\ "category": "Landmarks"\ },\ "expected_output": {\ "answer": "Eiffel Tower",\ "confidence": 0.9\ },\ "evaluators": [\ "EqualsExpected"\ ]\ }\ ], "evaluators": [], "report_evaluators": [] } """ Generate the [`Dataset`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset) exactly as above. Save the dataset to a JSON file, this will also write `questions_cases_schema.json` with the JSON schema for `questions_cases.json`. This time the `$schema` key is included in the JSON file to define the schema for IDEs to use while you edit the file, there's no formal spec for this, but it works in vscode and pycharm and is discussed at length in [json-schema-org/json-schema-spec#828](https://github.com/json-schema-org/json-schema-spec/issues/828) . _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main(answer))` to run `main`)_ Type-Safe Datasets ------------------ [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#type-safe-datasets) Use generic type parameters for type safety: from typing_extensions import TypedDict from pydantic_evals import Case, Dataset class MyInput(TypedDict): query: str max_results: int class MyOutput(TypedDict): results: list[str] class MyMetadata(TypedDict): category: str # Type-safe dataset dataset: Dataset[MyInput, MyOutput, MyMetadata] = Dataset( name='typed_dataset', cases=[\ Case(\ name='test',\ inputs={'query': 'test', 'max_results': 10},\ expected_output={'results': ['a', 'b']},\ metadata={'category': 'search'},\ ),\ ], ) Schema Generation ----------------- [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#schema-generation) Generate JSON Schema for IDE support: from typing import Any from pydantic_evals import Case, Dataset dataset = Dataset[str, str, Any](name='my_eval_suite', cases=[Case(name='test', inputs='example')]) # Save with schema dataset.to_file('my_dataset.yaml') # Creates my_dataset_schema.json # Schema enables: # - Autocomplete in VS Code/PyCharm # - Validation while editing # - Inline documentation Manual schema generation: import json from dataclasses import dataclass from typing import Any from pydantic_evals import Dataset from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class MyCustomEvaluator(Evaluator): threshold: float = 0.5 def evaluate(self, ctx: EvaluatorContext) -> bool: return True schema = Dataset[str, str, Any].model_json_schema_with_evaluators( custom_evaluator_types=[MyCustomEvaluator], ) print(json.dumps(schema, indent=2)[:66] + '...') """ { "$defs": { "Case": { "additionalProperties": false, ... """ Best Practices -------------- [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#best-practices) ### 1\. Use Clear Names [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#1-use-clear-names) from pydantic_evals import Case # Good Case(name='uppercase_basic_ascii', inputs='hello') Case(name='uppercase_unicode_emoji', inputs='hello 😀') Case(name='uppercase_empty_string', inputs='') # Bad Case(name='test1', inputs='hello') Case(name='test2', inputs='world') Case(name='test3', inputs='foo') ### 2\. Organize by Difficulty [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#2-organize-by-difficulty) from pydantic_evals import Case, Dataset dataset = Dataset( name='organized_by_difficulty', cases=[\ Case(name='easy_1', inputs='test', metadata={'difficulty': 'easy'}),\ Case(name='easy_2', inputs='test2', metadata={'difficulty': 'easy'}),\ Case(name='medium_1', inputs='test3', metadata={'difficulty': 'medium'}),\ Case(name='hard_1', inputs='test4', metadata={'difficulty': 'hard'}),\ ], ) ### 3\. Start Small, Grow Gradually [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#3-start-small-grow-gradually) from pydantic_evals import Case, Dataset # Start with representative cases dataset = Dataset( name='starting_small', cases=[\ Case(name='happy_path', inputs='test'),\ Case(name='edge_case', inputs=''),\ Case(name='error_case', inputs='invalid'),\ ], ) # Add more as you find issues dataset.add_case(name='newly_discovered_edge_case', inputs='edge') ### 4\. Use Case-specific Evaluators Where Appropriate [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#4-use-case-specific-evaluators-where-appropriate) Case-specific evaluators let different cases have different evaluation criteria, which is essential for comprehensive “test coverage”. Rather than trying to write one-size-fits-all evaluators, you can specify exactly what “good” looks like for each scenario. This is particularly powerful with [`LLMJudge`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.LLMJudge) evaluators where you can describe nuanced requirements per case, making it easy to build and maintain golden datasets. See [Case-specific evaluators](https://pydantic.dev/docs/ai/evals/evaluators/overview/#case-specific-evaluators) for detailed guidance. ### 5\. Separate Datasets by Purpose [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#5-separate-datasets-by-purpose) from typing import Any from pydantic_evals import Case, Dataset # First create some test datasets for name in ['smoke_tests', 'comprehensive_tests', 'regression_tests']: test_dataset = Dataset[str, Any, Any](name=name, cases=[Case(name='test', inputs='example')]) test_dataset.to_file(f'{name}.yaml') # Smoke tests (fast, critical paths) smoke_tests = Dataset[str, Any, Any].from_file('smoke_tests.yaml') # Comprehensive tests (slow, thorough) comprehensive = Dataset[str, Any, Any].from_file('comprehensive_tests.yaml') # Regression tests (specific bugs) regression = Dataset[str, Any, Any].from_file('regression_tests.yaml') Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#next-steps) * **[Dataset Serialization](https://pydantic.dev/docs/ai/evals/how-to/dataset-serialization/) ** - In-depth guide to saving and loading datasets * **[Generating Datasets](https://pydantic.dev/docs/ai/evals/how-to/dataset-management/#generating-datasets) ** - Use LLMs to generate test cases * **[Examples: Simple Validation](https://pydantic.dev/docs/ai/evals/examples/simple-validation/) ** - Practical examples Was this page helpful? Thanks for your feedback! --- # Prefect | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#_top) Prefect ======= [Prefect](https://www.prefect.io/) is a workflow orchestration framework for building resilient data pipelines in Python, natively integrated with Pydantic AI. Durable Execution ----------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#durable-execution) Prefect 3.0 brings [transactional semantics](https://www.prefect.io/blog/transactional-ml-pipelines-with-prefect-3-0) to your Python workflows, allowing you to group tasks into atomic units and define failure modes. If any part of a transaction fails, the entire transaction can be rolled back to a clean state. * **Flows** are the top-level entry points for your workflow. They can contain tasks and other flows. * **Tasks** are individual units of work that can be retried, cached, and monitored independently. Prefect 3.0’s approach to transactional orchestration makes your workflows automatically **idempotent**: rerunnable without duplication or inconsistency across any environment. Every task is executed within a transaction that governs when and where the task’s result record is persisted. If the task runs again under an identical context, it will not re-execute but instead load its previous result. The diagram below shows the overall architecture of an agentic application with Prefect. Prefect uses client-side task orchestration by default, with optional server connectivity for advanced features like scheduling and monitoring. +---------------------+ | Prefect Server | (Monitoring, | or Cloud | scheduling, UI, +---------------------+ orchestration) ^ | Flow state, | Schedule flows, metadata, | track execution logs | | +------------------------------------------------------+ | Application Process | | +----------------------------------------------+ | | | Flow (Agent.run) | | | +----------------------------------------------+ | | | | | | | v v v | | +-----------+ +------------+ +-------------+ | | | Task | | Task | | Task | | | | (Tool) | | (MCP Tool) | | (Model API) | | | +-----------+ +------------+ +-------------+ | | | | | | | Cache & Cache & Cache & | | persist persist persist | | to to to | | v v v | | +----------------------------------------------+ | | | Result Storage (Local FS, S3, etc.) | | | +----------------------------------------------+ | +------------------------------------------------------+ | | | v v v [External APIs, services, databases, etc.] See the [Prefect documentation](https://docs.prefect.io/) for more information. Durable Agent ------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#durable-agent) Add durable execution to any [`Agent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent) by attaching the [`PrefectDurability`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.prefect.PrefectDurability) [capability](https://pydantic.dev/docs/ai/capabilities/overview/) . When the agent runs inside a Prefect flow, the capability routes [model requests](https://pydantic.dev/docs/ai/models/overview/) , [tool calls](https://pydantic.dev/docs/ai/tools-toolsets/tools/) , and [MCP communication](https://pydantic.dev/docs/ai/mcp/client/) through Prefect tasks. To make a run durable, call `agent.run()` inside a `@flow`. The agent stays a normal `Agent` everywhere — outside a Prefect flow the capability is transparent, and the original agent, model, and MCP server can still be used as normal. See [Streaming](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#streaming) for event handling inside tasks and flow code. Here is a simple but complete example of attaching durable execution to an agent. All it requires is to install Pydantic AI with Prefect: * [pip](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#tab-panel-4) * [uv](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#tab-panel-5) Terminal pip install pydantic-ai[prefect] Terminal uv add pydantic-ai[prefect] Or if you’re using the slim package, you can install it with the `prefect` optional group: * [pip](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#tab-panel-6) * [uv](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#tab-panel-7) Terminal pip install pydantic-ai-slim[prefect] Terminal uv add pydantic-ai-slim[prefect] prefect\_durability.py from prefect import flow from pydantic_ai import Agent from pydantic_ai.durable_exec.prefect import PrefectDurability agent = Agent( 'openai:gpt-5.6-sol', instructions="You're an expert in geography.", name='geography', # (1) capabilities=[PrefectDurability()], # (2) ) @flow # (3) async def answer(question: str) -> str: result = await agent.run(question) return result.output async def main(): answer_text = await answer('What is the capital of Mexico?') print(answer_text) #> Mexico City (Ciudad de México, CDMX) The agent's `name` is used to uniquely identify its flows and tasks. Attach durability via `capabilities=[...]`. The capability routes model requests, tool calls, and MCP communication through Prefect tasks when the agent runs inside a flow. Wrap `agent.run()` in your own `@flow` to make the run durable. _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main())` to run `main`)_ Because the same agent works inside and outside a Prefect flow, [`PrefectDurability`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.prefect.PrefectDurability) composes with all other [capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) without each needing a Prefect-specific wrapper variant. For more information on how to use Prefect in Python applications, see their [Python documentation](https://docs.prefect.io/v3/how-to-guides/workflows/write-and-run) . ### Wrapper-agent path (deprecated) [](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#wrapper-agent-path-deprecated) Any agent can be wrapped in a [`PrefectAgent`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.prefect.PrefectAgent) to get a durable agent variant that routes model requests, tool calls, and MCP communication through Prefect tasks: prefect\_agent.py from pydantic_ai import Agent from pydantic_ai.durable_exec.prefect import PrefectAgent agent = Agent('openai:gpt-5.6-sol', name='geography') prefect_agent = PrefectAgent(agent) # Use `prefect_agent` in place of `agent`. Migrating to the capability means attaching `PrefectDurability` and adding the flow decorator that `PrefectAgent` used to apply for you: -prefect_agent = PrefectAgent(agent) -result = await prefect_agent.run(prompt) +agent = Agent(..., capabilities=[PrefectDurability()]) + +@flow +async def answer(prompt: str) -> str: + result = await agent.run(prompt) + return result.output Prefect Integration Considerations ---------------------------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#prefect-integration-considerations) When using Prefect with Pydantic AI agents, there are a few important considerations to ensure workflows behave correctly. ### Agent Requirements [](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#agent-requirements) Each agent instance must have a unique `name` so Prefect can correctly identify and track its flows and tasks. Toolsets that implement their own tool listing and calling (i.e. [`FunctionToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.FunctionToolset) , [`MCPToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset) , and [`DynamicToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.DynamicToolset) ) must have a unique [`id`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AbstractToolset.id) set, which is used to identify their tasks within the flow. ### Model Selection at Runtime [](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#model-selection-at-runtime) `Agent.run(model=...)` supports both model strings (like `'openai:gpt-5.6-sol'`) and model instances. A model instance can’t be serialized across the task boundary, and rebuilding one from its `model_id` string would build a _different_ model — the same model name on whatever provider the worker’s environment implies, so the request would go to another endpoint with other credentials. An instance that isn’t registered ahead of time is therefore rejected with a `UserError`. There are two ways to use a specific instance: pre-register it by passing a `models` dict to [`PrefectDurability`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.prefect.PrefectDurability) and reference it by key (or pass the registered instance), or pass a model-name string and build the instance inside the task with a [`ResolveModelId`](https://pydantic.dev/docs/ai/capabilities/resolve-model-id/) capability — the right choice when the model depends on the run’s `deps`, e.g. per-user credentials. Model-name strings themselves never need registering. The agent’s own model, set at construction, is always available as the default. To customize how a model string is built — a custom provider, or per-user credentials carried on the run’s `deps` — add a [`ResolveModelId`](https://pydantic.dev/docs/ai/capabilities/resolve-model-id/) capability before `PrefectDurability`: it gets first crack at every string, and the resolver runs again inside the task with the run’s actual `deps`, so it must be deterministic for a given `(model_id, deps)` and must not perform external I/O. ### Tool Wrapping [](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#tool-wrapping) Agent tools are automatically wrapped as Prefect tasks, which means they benefit from: * **Retry logic**: Failed tool calls can be retried automatically * **Caching**: Tool results are cached based on their inputs * **Observability**: Tool execution is tracked in the Prefect UI For a [`DynamicToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.DynamicToolset) , including one contributed by a [`DynamicCapability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.DynamicCapability) , each tool call runs as a task and resolves and enters the toolset inside that task, so a task retry re-resolves it. Tool discovery runs in flow code and is re-executed when the flow retries, like the rest of the flow — including a `DynamicCapability`’s factory, which should therefore be deterministic given the run’s `deps`. A `DynamicCapability` reuses the capability resolved for the run inside its tool tasks. A default [`TaskConfig`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.prefect.TaskConfig) for all tools can be passed as `tool_task_config` to the [`PrefectDurability`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.prefect.PrefectDurability) constructor. Per-tool config lives on the tool’s [`metadata`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.FunctionToolset.tool) field — `PrefectDurability` looks for a `'prefect'` key. You can set the metadata directly on the tool definition, or apply it across a selection of tools via the [`SetToolMetadata`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.SetToolMetadata) capability. See the [capabilities documentation](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.SetToolMetadata) for the full selector vocabulary. prefect\_per\_tool\_config.py from pydantic_ai import Agent from pydantic_ai.capabilities import SetToolMetadata from pydantic_ai.durable_exec.prefect import PrefectDurability, TaskConfig from pydantic_ai.toolsets import FunctionToolset toolset = FunctionToolset(id='research') @toolset.tool(metadata={'prefect': TaskConfig(timeout_seconds=10.0)}) # (1) def fetch_data(url: str) -> str: ... @toolset.tool(metadata={'prefect': False}) # (2) def simple_tool() -> str: ... agent = Agent( 'openai:gpt-5.6-sol', name='research', toolsets=[toolset], capabilities=[\ SetToolMetadata( # (3)\ tools=['fetch_data', 'fetch_dataset'],\ prefect=TaskConfig(timeout_seconds=10.0),\ ),\ PrefectDurability(tool_task_config=TaskConfig(retries=3)), # (4)\ ], ) Inline: declare the task config alongside the tool definition. Per-tool config merges on top of the base `tool_task_config`. Set `'prefect': False` to skip task wrapping entirely for that tool. Selector-based: [`SetToolMetadata`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.SetToolMetadata) applies the same metadata across a selection of tools (`'all'`, a name list, a dict, or a callable). `tool_task_config` sets the default config for every tool. ### Streaming [](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#streaming) `Agent.run_stream()`, `Agent.run_stream_events()`, and [`Agent.iter()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.iter) work inside a Prefect flow, but their events are buffered rather than delivered in real time. The model stream runs inside the durable task, and its events are replayed to the flow after the task completes. For handlers with I/O side effects, pass `event_stream_handler=` to [`PrefectDurability`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.prefect.PrefectDurability) . Model events are delivered live inside each model-request task, while each tool event is delivered in its own event-handler task. Configure those per-event tasks with `event_stream_handler_task_config=`. As with any Prefect task, a handler may run more than once if a task retries, so keep its side effects idempotent. Alternatively, register [`ProcessEventStream`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ProcessEventStream) . Its handler runs in flow code and must be deterministic because it re-runs on flow replay. Tool and final-output events arrive live, while the real captured model events are replayed after each model request completes. For examples, see the [streaming docs](https://pydantic.dev/docs/ai/core-concepts/agent/#streaming-all-events) . A durability `event_stream_handler=` and a separately registered `ProcessEventStream` are two distinct handlers, and each fires once. The durability handler receives live events inside the durable task, while `ProcessEventStream` sees the buffered replay in flow code. A per-run handler passed to `Agent.run(event_stream_handler=...)` also runs flow-side against replayed model events. Because the model stream is consumed inside the task, cancelling it from the flow side (e.g. with [`AgentStream.cancel()`](https://pydantic.dev/docs/ai/api/pydantic-ai/result/#pydantic_ai.result.AgentStream.cancel) ) is not available across the durable boundary. [`CancellationToken`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.CancellationToken) and [`RunContext.cancel()`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.cancel) are same-process cancellation handles and cannot cross the Prefect durable boundary; cancel the Prefect flow instead. `Agent.run_stream_sync()` is not for flow code: it requires no running event loop and wraps `run_stream()`. Under [`PrefectDurability`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.prefect.PrefectDurability) , use the buffered async streaming APIs above or `Agent.run()` with an event stream handler. Outside a flow, an agent with `PrefectDurability` behaves like a normal agent, so `run_stream_sync()` works as usual. (Wrapper `PrefectAgent` forbids `run_stream` inside flows — use `run` + event stream handler there.) ### Suspended Turns and Background Mode [](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#suspended-turns-and-background-mode) When a provider pauses a model turn mid-flight (Anthropic `pause_turn`) or runs it as a server-side job that’s polled until it’s ready ([OpenAI background mode](https://pydantic.dev/docs/ai/models/openai/#background-mode) ), each segment runs in a separate model request task. The suspended [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) and background job ID are checkpointed between segments, while the final response is merged and usage is recorded once. A [`message_history`](https://pydantic.dev/docs/ai/core-concepts/message-history/) ending in a suspended response is passed to the first task. Size `timeout_seconds` in [Task Configuration](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#task-configuration) for one provider round trip. If an error abandons a suspended job, its provider teardown runs in a dedicated cancellation task. ### Toolsets at Runtime [](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#toolsets-at-runtime) Additional toolsets can be passed per run via `agent.run(toolsets=...)`, but only toolsets that don’t need durable wrapping are supported: non-executing toolsets like [`ExternalToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.ExternalToolset) , whose tools are executed outside the agent run, and [`FunctionToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.FunctionToolset) s whose tools all opt out of task wrapping with `metadata={'prefect': False}`. Other executing toolsets ([`FunctionToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.FunctionToolset) and [`MCPToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset) ) and dynamic toolsets must be set when constructing the agent so their tasks are registered before the flow runs; passing them at runtime raises a `UserError`. Task Configuration ------------------ [](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#task-configuration) You can customize Prefect task behavior, such as retries and timeouts, by passing [`TaskConfig`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.prefect.TaskConfig) objects to the [`PrefectDurability`](https://pydantic.dev/docs/ai/api/pydantic-ai/durable_exec/#pydantic_ai.durable_exec.prefect.PrefectDurability) constructor: * `mcp_task_config`: Configuration for MCP server communication tasks * `model_task_config`: Configuration for model request tasks * `event_stream_handler_task_config`: Configuration for event stream handler tasks * `tool_task_config`: Default configuration for all tool calls (per-tool overrides go on the tool’s `'prefect'` metadata — see [Tool Wrapping](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#tool-wrapping) above) Available `TaskConfig` options: * `retries`: Maximum number of retries for the task (default: `0`) * `retry_delay_seconds`: Delay between retries in seconds (can be a single value or list for exponential backoff, default: `1.0`) * `timeout_seconds`: Maximum time in seconds for the task to complete * `cache_policy`: Custom Prefect cache policy for the task * `persist_result`: Whether to persist the task result * `result_storage`: Prefect result storage for the task (e.g., `'s3-bucket/my-storage'` or a `WritableFileSystem` block) * `log_prints`: Whether to log print statements from the task (default: `False`) Example: prefect\_durability\_task\_config.py from pydantic_ai import Agent from pydantic_ai.durable_exec.prefect import PrefectDurability, TaskConfig agent = Agent( 'openai:gpt-5.6-sol', instructions="You're an expert in geography.", name='geography', capabilities=[\ PrefectDurability(\ model_task_config=TaskConfig(\ retries=3,\ retry_delay_seconds=[1.0, 2.0, 4.0], # Exponential backoff\ timeout_seconds=30.0,\ ),\ ),\ ], ) async def main(): result = await agent.run('What is the capital of France?') print(result.output) #> Paris _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main())` to run `main`)_ ### Retry Considerations [](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#retry-considerations) Pydantic AI and provider API clients have their own retry logic. When using Prefect, you may want to: * Disable [HTTP Request Retries](https://pydantic.dev/docs/ai/models/http-request-retries/) in Pydantic AI * Turn off your provider API client’s retry logic (e.g., `max_retries=0` on a [custom OpenAI client](https://pydantic.dev/docs/ai/models/openai/#custom-openai-client) ) * Rely on Prefect’s task-level retry configuration for consistency This prevents requests from being retried multiple times at different layers. Caching and Idempotency ----------------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#caching-and-idempotency) Prefect 3.0 provides built-in caching and transactional semantics. Tasks with identical inputs will not re-execute if their results are already cached, making workflows naturally idempotent and resilient to failures. * **Task inputs**: A model request’s messages, settings and parameters; a tool call’s name, arguments, definition and [`tool_call_id`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.tool_call_id) (so two parallel calls to the same tool with the same arguments each execute); and the run state the task’s work can depend on: dependencies, [`metadata`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.metadata) , [`validation_context`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.validation_context) , the prompt, and the message history. Per-run identifiers like [`run_id`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.run_id) and [`conversation_id`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.conversation_id) , and message timestamps, are deliberately left out, so an otherwise identical run replays recorded results instead of re-executing them. **Note**: For user dependencies, `metadata` and `validation_context` to be included in cache keys, they must be serializable (e.g., Pydantic models or basic Python types). Non-serializable values are automatically excluded from cache computation. Observability with Prefect and Logfire -------------------------------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#observability-with-prefect-and-logfire) Prefect provides a built-in UI for monitoring flow runs, task executions, and failures. You can: * View real-time flow run status * Debug failures with full stack traces * Set up alerts and notifications To access the Prefect UI, you can either: 1. Use [Prefect Cloud](https://www.prefect.io/cloud) (managed service) 2. Run a local [Prefect server](https://docs.prefect.io/v3/how-to-guides/self-hosted/server-cli) with `prefect server start` You can also use [Pydantic Logfire](https://pydantic.dev/docs/ai/integrations/logfire/) for detailed observability. When using both Prefect and Logfire, you’ll get complementary views: * **Prefect**: Workflow-level orchestration, task status, and retry history * **Logfire**: Fine-grained tracing of agent runs, model requests, and tool invocations When using Logfire with Prefect, you can enable distributed tracing to see spans for your Prefect runs included with your agent runs, model requests, and tool invocations. For more information about Prefect monitoring, see the [Prefect documentation](https://docs.prefect.io/) . Deployments and Scheduling -------------------------- [](https://pydantic.dev/docs/ai/capabilities/durable_execution/prefect/#deployments-and-scheduling) To deploy and schedule a Prefect-durable agent, wrap it in a Prefect flow and use the flow’s [`serve()`](https://docs.prefect.io/v3/how-to-guides/deployments/create-deployments#create-a-deployment-with-serve) or [`deploy()`](https://docs.prefect.io/v3/how-to-guides/deployments/deploy-via-python) methods: serve\_agent.py from prefect import flow from pydantic_ai import Agent from pydantic_ai.durable_exec.prefect import PrefectDurability @flow async def daily_report_flow(user_prompt: str): """Generate a daily report using the agent.""" agent = Agent( # (1) 'openai:gpt-5.6-sol', name='daily_report_agent', instructions='Generate a daily summary report.', capabilities=[PrefectDurability()], ) result = await agent.run(user_prompt) return result.output # Serve the flow with a daily schedule if __name__ == '__main__': daily_report_flow.serve( name='daily-report-deployment', cron='0 9 * * *', # Run daily at 9am parameters={'user_prompt': "Generate today's report"}, tags=['production', 'reports'], ) Each flow run executes in an isolated process, and all inputs and dependencies must be serializable. Because Agent instances cannot be serialized, instantiate the agent inside the flow rather than at the module level. The `serve()` method accepts scheduling options: * **`cron`**: Cron schedule string (e.g., `'0 9 * * *'` for daily at 9am) * **`interval`**: Schedule interval in seconds or as a timedelta * **`rrule`**: iCalendar RRule schedule string For production deployments with Docker, Kubernetes, or other infrastructure, use the flow’s [`deploy()`](https://docs.prefect.io/v3/how-to-guides/deployments/deploy-via-python) method. See the [Prefect deployment documentation](https://docs.prefect.io/v3/how-to-guides/deployments/create-deploymentsy) for more information. Was this page helpful? Thanks for your feedback! --- # Voice Assistant | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/realtime/realtime-voice/#_top) Voice Assistant =============== Example of a voice assistant built on a [realtime](https://pydantic.dev/docs/ai/realtime/overview/) speech-to-speech model: it streams your microphone to OpenAI’s `gpt-realtime` model and plays the model’s spoken replies back through your speakers. Talk to it — and try interrupting while it’s speaking: the model stops and listens (barge-in). Demonstrates: * [realtime sessions](https://pydantic.dev/docs/ai/realtime/overview/) * [tools](https://pydantic.dev/docs/ai/tools-toolsets/tools/) * [barge-in](https://pydantic.dev/docs/ai/realtime/turns/#barge-in) (interrupting the model mid-sentence) The agent exposes a single `get_weather` tool the model can call mid-conversation, and the terminal shows a running transcript of both sides of the conversation plus any tool calls. Both audio directions use bounded buffers, dropping the oldest audio rather than growing without limit: microphone capture that outruns the network drops the oldest block to preserve conversational latency, and playback that falls more than five seconds behind the model drops its oldest audio, so a machine that stutters glitches instead of ending the call. Barge-in itself is handled by the provider — the model stops as soon as the user speaks. What the example adds is the half the provider can’t see: it clears queued _and_ partially consumed playback audio, then reports the duration actually played to [`interrupt()`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.interrupt) , so the provider truncates its transcript to what the user really heard rather than the whole turn. It only does so when unheard audio was actually dropped, since the speech-start event also fires on an ordinary turn where the user heard the previous reply in full. Running the Example ------------------- [](https://pydantic.dev/docs/ai/examples/realtime/realtime-voice/#running-the-example) The examples dependencies include [`sounddevice`](https://python-sounddevice.readthedocs.io/) for microphone and speaker access. It also requires the PortAudio system library; on Linux, install `libportaudio2` if importing `sounddevice` fails. The realtime model runs on `gpt-realtime`, so you’ll need an OpenAI API key set via `OPENAI_API_KEY`. With [dependencies installed and environment variables set](https://pydantic.dev/docs/ai/examples/setup/#usage) , run: * [pip](https://pydantic.dev/docs/ai/examples/realtime/realtime-voice/#tab-panel-54) * [uv](https://pydantic.dev/docs/ai/examples/realtime/realtime-voice/#tab-panel-55) Terminal python -m pydantic_ai_examples.realtime_voice Terminal uv run -m pydantic_ai_examples.realtime_voice Example Code ------------ [](https://pydantic.dev/docs/ai/examples/realtime/realtime-voice/#example-code) realtime\_voice.py from __future__ import annotations import asyncio import threading from collections import deque from contextlib import suppress from functools import partial import logfire from pydantic_ai import ( Agent, FunctionToolCallEvent, FunctionToolResultEvent, PartDeltaEvent, PartEndEvent, PartStartEvent, SpeechPart, SpeechPartDelta, ) from pydantic_ai.realtime import ( RealtimeEvent, RealtimeInputSpeechStartEvent, RealtimeSession, ) try: import sounddevice except (ImportError, OSError) as e: # pragma: no cover # `sounddevice` needs the PortAudio system library, which raises `OSError` (not `ImportError`) # when missing — e.g. on headless CI. Defer the failure to `main()` so the module still imports. sounddevice = None _sounddevice_error: Exception | None = e else: _sounddevice_error = None # 'if-token-present' means nothing will be sent (and the example will work) if you don't have logfire configured logfire.configure(send_to_logfire='if-token-present') logfire.instrument_pydantic_ai() # OpenAI's realtime models speak and listen in 24 kHz mono PCM16 audio. SAMPLE_RATE = 24000 CHANNELS = 1 BLOCK_SIZE = 2400 # 100 ms per audio block MIC_QUEUE_BLOCKS = 10 PLAYBACK_BUFFER_SECONDS = 5 agent = Agent( instructions='You are a friendly voice assistant. Keep your replies short and conversational.' ) @agent.tool_plain def get_weather(city: str) -> str: """Look up the current weather in a city.""" return f'It is currently 21 degrees and sunny in {city}.' def capture_mic( loop: asyncio.AbstractEventLoop, mic_queue: asyncio.Queue[bytes], indata: object, *_: object, ) -> None: """Microphone callback (PortAudio thread): hand captured audio to the event loop safely.""" loop.call_soon_threadsafe(enqueue_latest, mic_queue, bytes(indata)) def enqueue_latest(audio_queue: asyncio.Queue[bytes], chunk: bytes) -> None: """Keep microphone latency bounded by dropping the oldest block on overflow.""" if audio_queue.full(): try: audio_queue.get_nowait() except asyncio.QueueEmpty: pass audio_queue.put_nowait(chunk) class PlaybackBuffer: """Thread-safe, bounded model-audio buffer with playback accounting.""" def __init__(self, max_bytes: int): self._max_bytes = max_bytes self._chunks: deque[bytes] = deque() self._carry = bytearray() self._buffered_bytes = 0 self._played_bytes = 0 self._lock = threading.Lock() def start_turn(self) -> None: with self._lock: self._chunks.clear() self._carry.clear() self._buffered_bytes = 0 self._played_bytes = 0 def add(self, chunk: bytes) -> None: with self._lock: # If the speaker falls far enough behind that the model is seconds ahead of what the # caller hears, drop the oldest audio rather than raise: a glitch is recoverable, and # ending a live call because one machine stuttered is not. (After a drop the # played-duration accounting is approximate, which is fine for a glitch.) # A chunk longer than the whole window keeps its tail. chunk = chunk[-self._max_bytes :] while (over := self._buffered_bytes + len(chunk) - self._max_bytes) > 0: # Oldest first: whatever `fill` already staged for the speaker, then queued chunks. if self._carry: drop = min(len(self._carry), over) del self._carry[:drop] else: drop = len(self._chunks.popleft()) self._buffered_bytes -= drop self._chunks.append(chunk) self._buffered_bytes += len(chunk) def fill(self, outdata: bytearray) -> None: """Fill one speaker block, padding an underrun with silence.""" with self._lock: want = len(outdata) while len(self._carry) < want and self._chunks: self._carry.extend(self._chunks.popleft()) played = min(want, len(self._carry)) outdata[:] = bytes(self._carry[:played]).ljust(want, b'\x00') del self._carry[:played] self._buffered_bytes -= played self._played_bytes += played def interrupt(self) -> int | None: """Drop unheard audio; return milliseconds played, or `None` if nothing was left unheard. A turn the user heard in full needs no truncation — reporting one anyway would only make the provider discard part of a completed turn — so an interruption is only reported when unplayed audio was actually dropped. """ with self._lock: if not self._chunks and not self._carry: return None self._chunks.clear() self._carry.clear() self._buffered_bytes = 0 played_ms = self._played_bytes * 1000 // (SAMPLE_RATE * CHANNELS * 2) self._played_bytes = 0 return played_ms def fill_speaker(playback: PlaybackBuffer, outdata: bytearray, *_: object) -> None: """Speaker callback (PortAudio thread).""" playback.fill(outdata) async def handle_event( session: RealtimeSession, event: RealtimeEvent, playback: PlaybackBuffer, ) -> None: """Handle one session event.""" match event: case PartDeltaEvent(delta=SpeechPartDelta(audio_chunk=chunk)) if chunk: playback.add(chunk) case RealtimeInputSpeechStartEvent(): # The provider stops the model on its own when the user speaks; what it can't know is how # much of its audio actually reached the speaker. Drop what didn't, and report the rest so # the provider doesn't record a turn the user never heard. The event fires whenever the # user starts speaking — including when nothing is playing — so only interrupt when # unheard audio was actually dropped. if (played_ms := playback.interrupt()) is not None: await session.interrupt(played_ms=played_ms) case PartStartEvent(part=SpeechPart(speaker='assistant')): playback.start_turn() case PartEndEvent(part=SpeechPart(speaker='user', transcript=transcript)): print(f'you: {transcript}') case PartEndEvent(part=SpeechPart(speaker='assistant', transcript=transcript)): print(f'assistant: {transcript}') case FunctionToolCallEvent(part=call): print(f'[calling {call.tool_name}]') case FunctionToolResultEvent(part=result): print(f'[{result.tool_name} returned: {result.content}]') async def stream_mic(session: RealtimeSession, mic_queue: asyncio.Queue[bytes]) -> None: while True: await session.send_audio(await mic_queue.get()) async def main(): if sounddevice is None: # pragma: no cover raise ImportError( 'This example needs the `sounddevice` package for microphone and speaker access. ' 'Install it with `pip install sounddevice`. ' 'On Linux you also need the PortAudio system library (`apt install libportaudio2`).' ) from _sounddevice_error loop = asyncio.get_running_loop() mic_queue: asyncio.Queue[bytes] = asyncio.Queue(maxsize=MIC_QUEUE_BLOCKS) playback = PlaybackBuffer( max_bytes=SAMPLE_RATE * CHANNELS * 2 * PLAYBACK_BUFFER_SECONDS ) stream_kwargs = dict( samplerate=SAMPLE_RATE, channels=CHANNELS, dtype='int16', blocksize=BLOCK_SIZE ) mic = sounddevice.RawInputStream( callback=partial(capture_mic, loop, mic_queue), **stream_kwargs ) speaker = sounddevice.RawOutputStream( callback=partial(fill_speaker, playback), **stream_kwargs ) # The session opens before the microphone starts capturing, so no audio from before the # conversation began is queued up and sent to the model as stale input. async with agent.realtime('openai:gpt-realtime').session() as session: with mic, speaker: pump = asyncio.create_task(stream_mic(session, mic_queue)) print('Listening — start talking (Ctrl-C to quit).') try: async for event in session: await handle_event(session, event, playback) finally: pump.cancel() with suppress(asyncio.CancelledError): await pump if __name__ == '__main__': try: asyncio.run(main()) except KeyboardInterrupt: pass Was this page helpful? Thanks for your feedback! --- # Tool Output Limits | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/tool-output-limits/#_top) Tool Output Limits ================== `ToolOutputLimits` reduces a tool return that is large enough to dominate the context window. Tool returns persist in history as `ToolReturnPart`s, so an oversized one is re-sent on every later model request, paying its token cost for the rest of the run. This capability intercepts a return when it is produced, reduces it once, and lets the reduced form persist — the reduction is not recomputed per request. [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/tool_output_limits/) > The API may change between releases. Where practical, breaking changes ship with a deprecation warning. The problem ----------- [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#the-problem) A tool can return a payload large enough to dominate the context window: a big file read, a verbose log, a large JSON document. Because tool returns persist in history, an oversized one is re-sent on every later model request, paying its token cost for the rest of the run. This is the overflow-to-file follow-up the [compaction](https://pydantic.dev/docs/ai/harness/compaction/) capability names as out of scope: it moves large tool outputs _out_ of the window at production time, rather than compressing or dropping context already inside it. The three modes --------------- [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#the-three-modes) | Mode | Cost | Lossy? | What the model gets | | --- | --- | --- | --- | | `Truncate` | zero-LLM | yes | A head / tail / head+tail clamp of the text | | `Spill` | zero-LLM | no | A handle + preview + shape sketch; full payload read back on demand | | `Summarize` | one LLM call | yes | A size-gated summary (inherits the run’s model by default) | `Spill` is lossless: the full payload is persisted and the model reads slices of it through the registered `read_tool_result(handle, offset, limit, from_end, pattern)` tool (the Claude Code pattern, the core [#4352](https://github.com/pydantic/pydantic-ai/issues/4352) design). That tool is bounded: `offset >= 0`, `limit` clamped to a built-in line cap, the joined output capped, and `pattern` is a literal substring (not a regex), so a model-supplied value cannot hang the host with catastrophic backtracking. Usage ----- [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#usage) Construct an `Agent` with `ToolOutputLimits()` in its `capabilities`. With no arguments it uses the default band: spill returns of 10,000 characters or more, with a bounded truncation fallback if the store cannot accept the write. from pydantic_ai import Agent from pydantic_ai_harness.tool_output_limits import ToolOutputLimits agent = Agent('openai:gpt-4o', capabilities=[ToolOutputLimits()]) The capability registers a single `read_tool_result` tool so the model can page back into any spilled payload. Its own returns are exempt from reduction. Bands: combine the modes ------------------------ [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#bands-combine-the-modes) Configure an ordered list of size `bands`. Each band is a `(over, action)` pair: when a return’s measured size reaches `over`, its action runs. The band with the largest threshold that fits wins; anything below the smallest threshold passes through. from pydantic_ai import Agent from pydantic_ai_harness.tool_output_limits import ( Band, ToolOutputLimits, Spill, Summarize, Truncate, ) agent = Agent( 'openai:gpt-4o', capabilities=[\ ToolOutputLimits(\ bands=[\ Band(over=100_000, action=Spill()), # huge: keep losslessly, read back on demand\ Band(over=20_000, action=Summarize()), # large: compress with the run's model\ Band(over=5_000, action=Truncate()), # medium: cheap clamp\ ],\ # below 5,000: passthrough\ )\ ], ) The default band, when you pass no `bands`, is `Spill(then=Truncate())` at a 10,000-character threshold: lossless when a store accepts the write, a bounded truncation otherwise — zero LLM cost and no silent drop. `Passthrough()` is an explicit no-op action for `bands` or `per_tool` lists, leaving matching returns untouched. ### Fallbacks with `then` [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#fallbacks-with-then) Every action takes an optional `then`, applied when the action cannot run: a `Spill` whose store errors, a `Truncate` / `Summarize` on a binary payload, a `Summarize` whose model call raises. `then` chains, so `Summarize(then=Spill(then=Truncate()))` degrades summarize -> spill -> truncate. ### Per-tool overrides and filtering [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#per-tool-overrides-and-filtering) `per_tool` replaces the global band list for named tools (file reads to `head`, logs to `tail`); `tool_filter` (a `ToolSelector`) scopes which tools the capability touches at all. from pydantic_ai import Agent from pydantic_ai_harness.tool_output_limits import ( Band, ToolOutputLimits, Truncate, TruncationStrategy, ) agent = Agent( 'openai:gpt-4o', capabilities=[\ ToolOutputLimits(\ per_tool={\ 'read_file': [Band(over=8_000, action=Truncate(strategy=TruncationStrategy.head))],\ 'run_shell': [Band(over=8_000, action=Truncate(strategy=TruncationStrategy.tail))],\ },\ tool_filter=['read_file', 'run_shell', 'search'],\ )\ ], ) `TruncationStrategy` has three members: `head` (keep the first characters, good for headers and schemas), `tail` (keep the last characters, good for build and test output where errors land last), and `head_tail` (keep both ends, elide the middle — the default). Both `return_value` and `content` are reduced --------------------------------------------- [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#both-return_value-and-content-are-reduced) A `ToolReturn` carries a `return_value` and an optional `content` that core renders as a separate, model-visible part which also persists in history. This capability measures and reduces both with the same band logic (they spill to distinct handles). Text `content` is reduced in place; non-text `content` (multimodal parts) that overflows is left unreduced with a `warnings.warn`, since it cannot be safely truncated. Size unit --------- [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#size-unit) Thresholds are measured in characters by default. Set `over_tokens=True` to measure in estimated tokens (the same ~4-chars-per-token heuristic as [compaction](https://pydantic.dev/docs/ai/harness/compaction/) ); pass a `tokenizer` callable for accuracy. `Truncate.max_chars` is always characters — truncation is a character operation regardless of the threshold unit. Set `strip_ansi=True` to strip ANSI escape sequences from text returns before measuring and reducing. Spill store ----------- [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#spill-store) Spilled payloads go through the narrow `OverflowStore` protocol. The default `LocalFileStore` writes one file per `(run_id, tool_call_id, retry)` under a stable root directory and keeps it after the run, so a later `read_tool_result` — in this run or a subsequent agent/run — can still reach it. The handle is backend-addressable (a relative key), not an absolute local path, so a durable backend (Temporal, a blob store, or the core `ExecutionEnvironment` workspace once #4352 lands) can resolve the same handle in another process. Supply your own backend with `store=...`. from typing import Protocol class OverflowStore(Protocol): async def write(self, key: str, data: bytes) -> str: ... # returns a handle async def read(self, handle: str) -> bytes: ... ### Security model (shared root, not isolation) [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#security-model-shared-root-not-isolation) The store root is stable and shareable on purpose — spilled files must be readable by a later agent or run — so security does not come from per-instance isolation. It comes from two mechanisms: the root is created with `0700` (owner-only) permissions, and `read` resolves the target (following symlinks) and rejects anything that escapes the root via symlink, `..`, or an absolute path. Handle segments are also sanitized so a crafted handle cannot traverse out. ### Cleanup: keep-forever by default, opt-in TTL pruning [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#cleanup-keep-forever-by-default-opt-in-ttl-pruning) By default the store keeps spilled files forever — deleting on run end would break a later agent that still wants to read a spill. To bound disk use, opt into age-based pruning: from datetime import timedelta from pydantic_ai import Agent from pydantic_ai_harness.tool_output_limits import LocalFileStore, ToolOutputLimits store = LocalFileStore(cleanup_after=timedelta(hours=6)) # default: None = keep forever agent = Agent('openai:gpt-4o', capabilities=[ToolOutputLimits(store=store)]) When set, a `write` schedules a background prune (a daemon thread, off the hot path) that deletes files whose modification time (`st_mtime`) is older than `cleanup_after`. Pruning is non-blocking and non-erroring: any failure is caught and surfaced via `warnings.warn`, never propagated into the agent run, so cleanup can never fail a run or block the hot path. Last-read time (`st_atime`) is unreliable on `noatime`/`relatime` mounts and is not used. Prefer external cleanup (cron, a sweeper) over the in-process TTL? Point it at the store root and delete by mtime: import time from pathlib import Path root = Path('/tmp/pyai_harness_overflow') # or your configured base_dir cutoff = time.time() - 6 * 3600 for path in root.rglob('*'): if path.is_file() and path.stat().st_mtime < cutoff: path.unlink(missing_ok=True) Usage accounting ---------------- [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#usage-accounting) A `Summarize` call is a real request to the model, so its full usage — tokens and the request itself — folds into the run’s `ctx.usage`, exactly like `SummarizingCompaction`. No token caps are imposed on the summary call. A `UsageLimits` request limit will see it. By default `Summarize` inherits the running agent’s model (`ctx.model`). Pass a model id or instance to `Summarize(model=...)` to override, or a `summarize` callable to bypass the built-in prompt entirely. The `summary_prompt` template on the capability must contain both `{tool_name}` and `{output}` placeholders. Edge cases ---------- [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#edge-cases) * Binary returns spill verbatim and are never stringify-truncated; `Truncate` / `Summarize` on binary fall through to `then`. * Structured / nested returns spill (or summarize) by preference — truncating JSON produces invalid JSON. `Spill` includes a one-line shape sketch of the top level. * `ModelRetry` and tool errors never reach this hook (they are raised, not returned), so the model always gets the full error it needs to recover. * A large `ToolReturn.content` is reduced with the same bands as `return_value`; non-text content that overflows is left unreduced with a warning. * Multiple oversized returns in one step get distinct handles (keyed per `tool_call_id`); retries get distinct handles too (keyed per `retry`), so a retried call never clobbers the earlier attempt’s spill. Relationship to other capabilities ---------------------------------- [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#relationship-to-other-capabilities) * Distinct from [compaction](https://pydantic.dev/docs/ai/harness/compaction/) , which compresses or drops context already inside the window; this capability moves large tool outputs out of the window at production time. * Consumes core [#4352](https://github.com/pydantic/pydantic-ai/issues/4352) (the canonical queryable-file primitive) through the `OverflowStore` seam once it lands. * Distinct from `ClampOversizedMessages`, which clamps runaway model responses, not tool returns. API reference ------------- [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#api-reference) ToolOutputLimits ---------------- [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#pydantic_ai_harness.tool_output_limits.ToolOutputLimits) **Bases:** `AbstractCapability[AgentDepsT]` Reduce oversized tool returns when they are produced, persisting the reduction. A tool can return a payload large enough to dominate the context window. Tool returns persist in history, so an oversized one is re-sent on every later request. This capability intercepts a return in `after_tool_execute`, reduces it once, and lets the reduced form persist — it is not recomputed per request. Three reduction modes, freely combined through an ordered list of size `bands`: * `Truncate`: clamp to a character budget. Lossy, zero-cost. * `Spill`: persist the full payload, hand the model a `read_tool_result` handle plus a preview. Lossless. * `Summarize`: size-gated LLM summary. Inherits the run’s model by default. The first band whose `over` threshold the measured size meets wins; smaller returns pass through. `per_tool` replaces the band list for named tools; `tool_filter` scopes which tools are touched at all. The default is `Spill(then=Truncate())`: lossless when a store accepts the write, a bounded truncation otherwise. `ModelRetry` and other errors never reach this hook (they are raised, not returned), so error payloads the model needs to recover are never spilled or summarized. ### Attributes [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#attributes) #### bands [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#pydantic_ai_harness.tool_output_limits.ToolOutputLimits.bands) Ordered size bands. The first band whose `over` threshold is met wins. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[`Band`\] **Default:** `field(default_factory=_default_bands)` #### over\_tokens [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#pydantic_ai_harness.tool_output_limits.ToolOutputLimits.over_tokens) Measure band thresholds in estimated tokens instead of characters. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### per\_tool [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#pydantic_ai_harness.tool_output_limits.ToolOutputLimits.per_tool) Per-tool band lists that replace `bands` for the named tools. **Type:** [`Mapping`](https://docs.python.org/3/library/typing.html#typing.Mapping) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence)\ \[`Band`\]\] **Default:** `field(default_factory=(dict[str, Sequence[Band]]))` #### store [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#pydantic_ai_harness.tool_output_limits.ToolOutputLimits.store) Backend for spilled payloads. Defaults to a `LocalFileStore`. **Type:** `OverflowStore` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### strip\_ansi [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#pydantic_ai_harness.tool_output_limits.ToolOutputLimits.strip_ansi) Strip ANSI escape sequences from text returns before measuring and reducing. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### summary\_prompt [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#pydantic_ai_harness.tool_output_limits.ToolOutputLimits.summary_prompt) Prompt template for `Summarize`. Must contain `{tool_name}` and `{output}`. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `_DEFAULT_SUMMARY_PROMPT` #### tokenizer [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#pydantic_ai_harness.tool_output_limits.ToolOutputLimits.tokenizer) Optional `(str) -> int` tokenizer for `over_tokens`. Defaults to a ~4-char heuristic. **Type:** [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \], [`int`](https://docs.python.org/3/library/functions.html#int)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### tool\_filter [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#pydantic_ai_harness.tool_output_limits.ToolOutputLimits.tool_filter) Which tools this capability touches. Non-matching tools always pass through. **Type:** `ToolSelector`\[`AgentDepsT`\] **Default:** `'all'` ### Methods [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#methods) #### after\_tool\_execute [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#pydantic_ai_harness.tool_output_limits.ToolOutputLimits.after_tool_execute) `@async` def after_tool_execute( ctx: RunContext[AgentDepsT], *, call: ToolCallPart, tool_def: ToolDefinition, args: dict[str, Any], result: Any, ) -> Any Reduce the tool result — both `return_value` and model-visible `content`. ##### Returns [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#returns) [`Any`](https://docs.python.org/3/library/typing.html#typing.Any) #### get\_toolset [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#pydantic_ai_harness.tool_output_limits.ToolOutputLimits.get_toolset) def get_toolset() -> AgentToolset[AgentDepsT] | None Register the `read_tool_result` tool for reading spilled payloads on demand. ##### Returns [](https://pydantic.dev/docs/ai/harness/tool-output-limits/#returns-1) [`AgentToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AgentToolset) \[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) Was this page helpful? Thanks for your feedback! --- # Overview | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/graph/graph/#_top) Overview ======== Graphs and finite state machines (FSMs) are a powerful abstraction to model, execute, control and visualize complex workflows. Alongside Pydantic AI, we’ve developed `pydantic-graph` — an async graph and state machine library for Python where nodes and edges are defined using type hints. While this library is developed as part of Pydantic AI; it has no dependency on `pydantic-ai` and can be considered as a pure graph-based state machine library. You may find it useful whether or not you’re using Pydantic AI or even building with GenAI. `pydantic-graph` is designed for advanced users and makes heavy use of Python generics and type hints. It is not designed to be as beginner-friendly as Pydantic AI. Installation ------------ [](https://pydantic.dev/docs/ai/graph/graph/#installation) `pydantic-graph` is a required dependency of `pydantic-ai`, and an optional dependency of `pydantic-ai-slim`, see [installation instructions](https://pydantic.dev/docs/ai/overview/install/#slim-install) for more information. You can also install it directly: * [pip](https://pydantic.dev/docs/ai/graph/graph/#tab-panel-70) * [uv](https://pydantic.dev/docs/ai/graph/graph/#tab-panel-71) Terminal pip install pydantic-graph Terminal uv add pydantic-graph Graph Types ----------- [](https://pydantic.dev/docs/ai/graph/graph/#graph-types) `pydantic-graph` is made up of a few key components: ### GraphRunContext [](https://pydantic.dev/docs/ai/graph/graph/#graphruncontext) [`GraphRunContext`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.GraphRunContext) — The context for the graph run, similar to Pydantic AI’s [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) . This holds the state of the graph and dependencies and is passed to nodes when they’re run. `GraphRunContext` is generic in the state type of the graph it’s used in, [`StateT`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.StateT) . ### End [](https://pydantic.dev/docs/ai/graph/graph/#end) [`End`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.End) — return value to indicate the graph run should end. `End` is generic in the graph return type of the graph it’s used in, [`RunEndT`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.RunEndT) . ### Nodes [](https://pydantic.dev/docs/ai/graph/graph/#nodes) Subclasses of [`BaseNode`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.BaseNode) define nodes for execution in the graph. Nodes, which are generally [`dataclass`es](https://docs.python.org/3/library/dataclasses.html#dataclasses.dataclass) , generally consist of: * fields containing any parameters required/optional when calling the node * the business logic to execute the node, in the [`run`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.BaseNode.run) method * return annotations of the [`run`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.BaseNode.run) method, which are read by `pydantic-graph` to determine the outgoing edges of the node Nodes are generic in: * **state**, which must have the same type as the state of graphs they’re included in, [`StateT`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.StateT) has a default of `None`, so if you’re not using state you can omit this generic parameter, see [stateful graphs](https://pydantic.dev/docs/ai/graph/graph/#stateful-graphs) for more information * **deps**, which must have the same type as the deps of the graph they’re included in, [`DepsT`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.DepsT) has a default of `None`, so if you’re not using deps you can omit this generic parameter, see [dependency injection](https://pydantic.dev/docs/ai/graph/graph/#dependency-injection) for more information * **graph return type** — this only applies if the node returns [`End`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.End) . [`RunEndT`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.RunEndT) has a default of [Never](https://docs.python.org/3/library/typing.html#typing.Never) so this generic parameter can be omitted if the node doesn’t return `End`, but must be included if it does. Here’s an example of a start or intermediate node in a graph — it can’t end the run as it doesn’t return [`End`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.End) : intermediate\_node.py from dataclasses import dataclass from pydantic_graph import BaseNode, GraphRunContext @dataclass class MyNode(BaseNode[MyState]): # (1) foo: int # (2) async def run( self, ctx: GraphRunContext[MyState], # (3) ) -> AnotherNode: # (4) ... return AnotherNode() State in this example is `MyState` (not shown), hence `BaseNode` is parameterized with `MyState`. This node can't end the run, so the `RunEndT` generic parameter is omitted and defaults to `Never`. `MyNode` is a dataclass and has a single field `foo`, an `int`. The `run` method takes a `GraphRunContext` parameter, again parameterized with state `MyState`. The return type of the `run` method is `AnotherNode` (not shown), this is used to determine the outgoing edges of the node. We could extend `MyNode` to optionally end the run if `foo` is divisible by 5: intermediate\_or\_end\_node.py from dataclasses import dataclass from pydantic_graph import BaseNode, End, GraphRunContext @dataclass class MyNode(BaseNode[MyState, None, int]): # (1) foo: int async def run( self, ctx: GraphRunContext[MyState], ) -> AnotherNode | End[int]: # (2) if self.foo % 5 == 0: return End(self.foo) else: return AnotherNode() We parameterize the node with the return type (`int` in this case) as well as state. Because generic parameters are positional-only, we have to include `None` as the second parameter representing deps. The return type of the `run` method is now a union of `AnotherNode` and `End[int]`, this allows the node to end the run if `foo` is divisible by 5. ### Graph [](https://pydantic.dev/docs/ai/graph/graph/#graph) [`Graph`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.Graph) — the executable graph produced by a [`GraphBuilder`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder) . The builder is the entry point for assembling a graph from [step functions](https://pydantic.dev/docs/ai/graph/builder/steps/) , [`BaseNode`](https://pydantic.dev/docs/ai/graph/graph/#nodes) classes, and the edges connecting them. [`GraphBuilder`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder) is generic in: * **state** the state type of the graph, [`StateT`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.StateT) * **deps** the deps type of the graph, [`DepsT`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.DepsT) * **input** the type of the initial input passed to the graph, `InputT` * **output** the type of the final output produced by the graph, `OutputT` Here’s an example of a simple graph built from two `BaseNode` subclasses: graph\_example.py from __future__ import annotations from dataclasses import dataclass from pydantic_graph import BaseNode, End, GraphBuilder, GraphRunContext, StepContext @dataclass class DivisibleBy5(BaseNode[None, None, int]): # (1) foo: int async def run( self, ctx: GraphRunContext, ) -> Increment | End[int]: if self.foo % 5 == 0: return End(self.foo) else: return Increment(self.foo) @dataclass class Increment(BaseNode): # (2) foo: int async def run(self, ctx: GraphRunContext) -> DivisibleBy5: return DivisibleBy5(self.foo + 1) g = GraphBuilder(input_type=int, output_type=int) # (3) @g.step async def start(ctx: StepContext[None, None, int]) -> DivisibleBy5: # (4) return DivisibleBy5(ctx.inputs) g.add( g.node(DivisibleBy5), # (5) g.node(Increment), g.edge_from(g.start_node).to(start), # (6) ) fives_graph = g.build() # (7) async def main(): result = await fives_graph.run(inputs=4) # (8) print(result) #> 5 The `DivisibleBy5` node is parameterized with `None` for the state param and `None` for the deps param as this graph doesn't use state or deps, and `int` as it can end the run. The `Increment` node doesn't return `End`, so the `RunEndT` generic parameter is omitted, state can also be omitted as the graph doesn't use state. Create a [`GraphBuilder`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder) declaring the input and output types of the graph. Define a [step](https://pydantic.dev/docs/ai/graph/builder/steps/) that wraps the initial input as the first `BaseNode`. The builder calls this when execution leaves [`g.start_node`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder.start_node) . Register each `BaseNode` subclass with [`g.node()`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder.node) so the builder knows about it; outgoing edges are inferred from each node's `run` return type. Wire the start node into the entry step. [`g.build()`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder.build) returns a [`Graph`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.Graph) ready to execute. [`graph.run()`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.Graph.run) is async and returns the raw output value (the `int` returned by the `End` node). _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ A [mermaid diagram](https://pydantic.dev/docs/ai/graph/graph/#mermaid-diagrams) for this graph can be generated with `print(fives_graph)`, or by calling [`fives_graph.render()`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.Graph.render) : stateDiagram-v2 start DivisibleBy5 state decision <> Increment [*] --> start start --> DivisibleBy5 DivisibleBy5 --> decision decision --> Increment decision --> [*] Increment --> DivisibleBy5 Stateful Graphs --------------- [](https://pydantic.dev/docs/ai/graph/graph/#stateful-graphs) The “state” concept in `pydantic-graph` provides an optional way to access and mutate an object (often a `dataclass` or Pydantic model) as nodes run in a graph. If you think of Graphs as a production line, then your state is the engine being passed along the line and built up by each node as the graph is run. Here’s an example of a graph which represents a vending machine where the user may insert coins and select a product to purchase. vending\_machine.py from __future__ import annotations from dataclasses import dataclass from rich.prompt import Prompt from pydantic_graph import BaseNode, End, GraphBuilder, GraphRunContext, StepContext @dataclass class MachineState: # (1) user_balance: float = 0.0 product: str | None = None @dataclass class InsertCoin(BaseNode[MachineState]): # (3) async def run(self, ctx: GraphRunContext[MachineState]) -> CoinsInserted: # (14) return CoinsInserted(float(Prompt.ask('Insert coins'))) # (4) @dataclass class CoinsInserted(BaseNode[MachineState]): amount: float # (5) async def run( self, ctx: GraphRunContext[MachineState] ) -> SelectProduct | Purchase: # (15) ctx.state.user_balance += self.amount # (6) if ctx.state.product is not None: # (7) return Purchase(ctx.state.product) else: return SelectProduct() @dataclass class SelectProduct(BaseNode[MachineState]): async def run(self, ctx: GraphRunContext[MachineState]) -> Purchase: return Purchase(Prompt.ask('Select product')) PRODUCT_PRICES = { # (2) 'water': 1.25, 'soda': 1.50, 'crisps': 1.75, 'chocolate': 2.00, } @dataclass class Purchase(BaseNode[MachineState, None, None]): # (16) product: str async def run( self, ctx: GraphRunContext[MachineState] ) -> End | InsertCoin | SelectProduct: if price := PRODUCT_PRICES.get(self.product): # (8) ctx.state.product = self.product # (9) if ctx.state.user_balance >= price: # (10) ctx.state.user_balance -= price return End(None) else: diff = price - ctx.state.user_balance print(f'Not enough money for {self.product}, need {diff:0.2f} more') #> Not enough money for crisps, need 0.75 more return InsertCoin() # (11) else: print(f'No such product: {self.product}, try again') return SelectProduct() # (12) g = GraphBuilder(state_type=MachineState) # (13) @g.step async def start(ctx: StepContext[MachineState, None, None]) -> InsertCoin: return InsertCoin() g.add( g.node(InsertCoin), g.node(CoinsInserted), g.node(SelectProduct), g.node(Purchase), g.edge_from(g.start_node).to(start), ) vending_machine_graph = g.build() async def main(): state = MachineState() # (17) await vending_machine_graph.run(state=state) # (18) print(f'purchase successful item={state.product} change={state.user_balance:0.2f}') #> purchase successful item=crisps change=0.25 The state of the vending machine is defined as a dataclass with the user's balance and the product they've selected, if any. A dictionary of products mapped to prices. The `InsertCoin` node, [`BaseNode`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.BaseNode) is parameterized with `MachineState` as that's the state used in this graph. The `InsertCoin` node prompts the user to insert coins. We keep things simple by just entering a monetary amount as a float. The `CoinsInserted` node; again this is a [`dataclass`](https://docs.python.org/3/library/dataclasses.html#dataclasses.dataclass) with one field `amount`. Update the user's balance with the amount inserted. If the user has already selected a product, go to `Purchase`, otherwise go to `SelectProduct`. In the `Purchase` node, look up the price of the product if the user entered a valid product. If the user did enter a valid product, set the product in the state so we don't revisit `SelectProduct`. If the balance is enough to purchase the product, adjust the balance to reflect the purchase and return [`End`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.End) to end the graph. We're not using the run return type, so we call `End` with `None`. If the balance is insufficient, go to `InsertCoin` to prompt the user to insert more coins. If the product is invalid, go to `SelectProduct` to prompt the user to select a product again. Build the graph with [`GraphBuilder`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder) , declaring the `MachineState` type. Each `BaseNode` subclass is registered with [`g.node()`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder.node) ; outgoing edges are inferred from the `run` return types. The `start` step constructs the first node. The return type of the node's [`run`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.BaseNode.run) method is important as it is used to determine the outgoing edges of the node. This information in turn is used to render [mermaid diagrams](https://pydantic.dev/docs/ai/graph/graph/#mermaid-diagrams) and is enforced at runtime to detect misbehavior as soon as possible. The return type of `CoinsInserted`'s [`run`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.BaseNode.run) method is a union, meaning multiple outgoing edges are possible. Unlike other nodes, `Purchase` can end the run, so the [`RunEndT`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.RunEndT) generic parameter must be set. In this case it's `None` since the graph run return type is `None`. Initialize the state. This will be passed to the graph run and mutated as the graph runs. Run the graph with the initial state. The first node to execute is determined by the `start` step we wired into [`g.start_node`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder.start_node) . _(This example is complete, it can be run “as is” — you’ll need to add `import asyncio; asyncio.run(main())` to run `main`)_ A [mermaid diagram](https://pydantic.dev/docs/ai/graph/graph/#mermaid-diagrams) for this graph can be generated with `print(vending_machine_graph)`: stateDiagram-v2 start InsertCoin CoinsInserted state decision <> Purchase SelectProduct state decision_2 <> [*] --> start start --> InsertCoin InsertCoin --> CoinsInserted CoinsInserted --> decision decision --> Purchase decision --> SelectProduct SelectProduct --> Purchase Purchase --> decision_2 decision_2 --> InsertCoin decision_2 --> SelectProduct decision_2 --> [*] See [below](https://pydantic.dev/docs/ai/graph/graph/#mermaid-diagrams) for more information on generating diagrams. GenAI Example ------------- [](https://pydantic.dev/docs/ai/graph/graph/#genai-example) So far we haven’t shown an example of a Graph that actually uses Pydantic AI or GenAI at all. In this example, one agent generates a welcome email to a user and the other agent provides feedback on the email. This graph has a very simple structure: --- title: feedback_graph --- stateDiagram-v2 [*] --> WriteEmail WriteEmail --> Feedback Feedback --> WriteEmail Feedback --> [*] genai\_email\_feedback.py from __future__ import annotations as _annotations from dataclasses import dataclass, field from pydantic import BaseModel, EmailStr from pydantic_ai import Agent, ModelMessage, format_as_xml from pydantic_graph import BaseNode, End, GraphBuilder, GraphRunContext, StepContext @dataclass class User: name: str email: EmailStr interests: list[str] @dataclass class Email: subject: str body: str @dataclass class State: user: User write_agent_messages: list[ModelMessage] = field(default_factory=list) email_writer_agent = Agent( 'google:gemini-3-pro-preview', output_type=Email, instructions='Write a welcome email to our tech blog.', ) @dataclass class WriteEmail(BaseNode[State]): email_feedback: str | None = None async def run(self, ctx: GraphRunContext[State]) -> Feedback: if self.email_feedback: prompt = ( f'Rewrite the email for the user:\n' f'{format_as_xml(ctx.state.user)}\n' f'Feedback: {self.email_feedback}' ) else: prompt = ( f'Write a welcome email for the user:\n' f'{format_as_xml(ctx.state.user)}' ) result = await email_writer_agent.run( prompt, message_history=ctx.state.write_agent_messages, ) ctx.state.write_agent_messages += result.new_messages() return Feedback(result.output) class EmailRequiresWrite(BaseModel): feedback: str class EmailOk(BaseModel): pass feedback_agent = Agent[object, EmailRequiresWrite | EmailOk]( 'openai:gpt-5.2', output_type=EmailRequiresWrite | EmailOk, # type: ignore instructions=( 'Review the email and provide feedback, email must reference the users specific interests.' ), ) @dataclass class Feedback(BaseNode[State, None, Email]): email: Email async def run( self, ctx: GraphRunContext[State], ) -> WriteEmail | End[Email]: prompt = format_as_xml({'user': ctx.state.user, 'email': self.email}) result = await feedback_agent.run(prompt) if isinstance(result.output, EmailRequiresWrite): return WriteEmail(email_feedback=result.output.feedback) else: return End(self.email) g = GraphBuilder(state_type=State, output_type=Email) @g.step async def start(ctx: StepContext[State, None, None]) -> WriteEmail: return WriteEmail() g.add( g.node(WriteEmail), g.node(Feedback), g.edge_from(g.start_node).to(start), ) feedback_graph = g.build() async def main(): user = User( name='John Doe', email='john.joe@example.com', interests=['Haskel', 'Lisp', 'Fortran'], ) state = State(user) result = await feedback_graph.run(state=state) print(result) """ Email( subject='Welcome to our tech blog!', body='Hello John, Welcome to our tech blog! ...', ) """ _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main())` to run `main`)_ Iterating Over a Graph ---------------------- [](https://pydantic.dev/docs/ai/graph/graph/#iterating-over-a-graph) For step-by-step execution — inspecting each task as it runs, overriding the next step, or driving the loop manually — use [`graph.iter()`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.Graph.iter) instead of [`graph.run()`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.Graph.run) . See [Advanced Execution Control](https://pydantic.dev/docs/ai/graph/builder/#advanced-execution-control) in the graph builder docs for the iteration model and examples. Dependency Injection -------------------- [](https://pydantic.dev/docs/ai/graph/graph/#dependency-injection) As with Pydantic AI, `pydantic-graph` supports dependency injection. Pass a `deps_type` to [`GraphBuilder`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder) , parameterize each [`BaseNode`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.BaseNode) subclass with the deps type, and read it via [`GraphRunContext.deps`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.GraphRunContext.deps) inside `run()` (or [`StepContext.deps`](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StepContext) inside step functions). As an example, let’s modify the `DivisibleBy5` example [above](https://pydantic.dev/docs/ai/graph/graph/#graph) to use a [`ProcessPoolExecutor`](https://docs.python.org/3/library/concurrent.futures.html#concurrent.futures.ProcessPoolExecutor) to run the compute load in a separate process (this is a contrived example, `ProcessPoolExecutor` wouldn’t actually improve performance in this example): deps\_example.py from __future__ import annotations import asyncio from concurrent.futures import ProcessPoolExecutor from dataclasses import dataclass from pydantic_graph import BaseNode, End, GraphBuilder, GraphRunContext, StepContext @dataclass class GraphDeps: executor: ProcessPoolExecutor @dataclass class DivisibleBy5(BaseNode[None, GraphDeps, int]): foo: int async def run( self, ctx: GraphRunContext[None, GraphDeps], ) -> Increment | End[int]: if self.foo % 5 == 0: return End(self.foo) else: return Increment(self.foo) @dataclass class Increment(BaseNode[None, GraphDeps]): foo: int async def run(self, ctx: GraphRunContext[None, GraphDeps]) -> DivisibleBy5: loop = asyncio.get_running_loop() compute_result = await loop.run_in_executor( ctx.deps.executor, self.compute, ) return DivisibleBy5(compute_result) def compute(self) -> int: return self.foo + 1 g = GraphBuilder(deps_type=GraphDeps, input_type=int, output_type=int) @g.step async def start(ctx: StepContext[None, GraphDeps, int]) -> DivisibleBy5: return DivisibleBy5(ctx.inputs) g.add( g.node(DivisibleBy5), g.node(Increment), g.edge_from(g.start_node).to(start), ) fives_graph = g.build() async def main(): with ProcessPoolExecutor() as executor: deps = GraphDeps(executor) result = await fives_graph.run(inputs=3, deps=deps) print(result) #> 5 _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main())` to run `main`)_ Mermaid Diagrams ---------------- [](https://pydantic.dev/docs/ai/graph/graph/#mermaid-diagrams) Pydantic Graph can render [mermaid](https://mermaid.js.org/) [`stateDiagram-v2`](https://mermaid.js.org/syntax/stateDiagram.html) diagrams for any built graph. Call [`graph.render()`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.Graph.render) (or just `print(graph)`) to get the mermaid source — pass `direction` (`'TB'`, `'LR'`, `'RL'`, or `'BT'`) to control layout. See the [graph builder mermaid section](https://pydantic.dev/docs/ai/graph/builder/#visualizing-graphs) for the full set of rendering options. Was this page helpful? Thanks for your feedback! --- # openai | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/realtime/openai/#_top) openai ====== The OpenAI Realtime API provider. Requires the `realtime` and `openai` optional groups (`pip install "pydantic-ai-slim[realtime,openai]"`). [`OpenAIRealtimeModelSettings`](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModelSettings) configures the session, including shared turn-taking via [`TurnDetection`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.TurnDetection) (or `False` for push-to-talk). For finer control, `openai_turn_detection` accepts [`ServerVAD`](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.ServerVAD) or [`SemanticVAD`](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.SemanticVAD) and fully overrides the shared setting. Resilience comes from the `reconnect` setting: a [`ReconnectPolicy`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.ReconnectPolicy) in [`RealtimeModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelSettings) . OpenAI Realtime API provider for speech-to-speech sessions. Connects to `wss://api.openai.com/v1/realtime` over a WebSocket and maps the OpenAI event protocol to the shared realtime event types. Requires the `websockets` and `openai` packages, available via the `realtime` and `openai` optional groups: pip install “pydantic-ai-slim\[openai-realtime\]“ OpenAIRealtimeConnection ------------------------ [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeConnection) **Bases:** `RealtimeConnection` A live WebSocket connection to the OpenAI Realtime API. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/openai/#attributes) #### message\_history [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeConnection.message_history) The call so far, when a session has offered it for replay on reconnect. **Type:** [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[\], [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence)\ \[[`ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage)\ \]\] | [`None`](https://docs.python.org/3/library/constants.html#None) ### Methods [](https://pydantic.dev/docs/ai/api/realtime/openai/#methods) #### send [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeConnection.send) `@async` def send(content: RealtimeInput) -> None Send content to the OpenAI Realtime API. Accepts `BinaryAudio` (raw PCM16, 24kHz, mono), a `str` text turn, `BinaryImage`, `ToolResult`, and the control verbs `CommitAudio`, `ClearAudio`, `CreateResponse`, `CancelResponse`, and `TruncateOutput`. ##### Returns [](https://pydantic.dev/docs/ai/api/realtime/openai/#returns) [`None`](https://docs.python.org/3/library/constants.html#None) OpenAIRealtimeModel ------------------- [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModel) **Bases:** `RealtimeModel` OpenAI Realtime API model. Authentication and the base URL come from a [`Provider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.Provider) , mirroring [`OpenAIChatModel`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIChatModel) . Pass `provider='openai'` (the default) to read `OPENAI_API_KEY` / `OPENAI_BASE_URL` from the environment, or an [`OpenAIProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.openai.OpenAIProvider) instance for a custom key or base URL. The realtime transport is opened separately with `websockets`, so the provider’s `httpx` client is not used for the WebSocket connection. The realtime WebSocket URL is derived from the provider’s base URL (e.g. `https://api.openai.com/v1/` → `wss://api.openai.com/v1/realtime`), so OpenAI-compatible endpoints that expose a realtime API work too. ### Constructor Parameters [](https://pydantic.dev/docs/ai/api/realtime/openai/#constructor-parameters) **`model`** : `OpenAIRealtimeModelName` [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModel.__init__(model)) The model name, e.g. `gpt-realtime` or `gpt-realtime-2.1-mini`. **`provider`** : `Provider`\[`AsyncOpenAI`\] | [`str`](https://docs.python.org/3/library/stdtypes.html#str) _Default:_ `'openai'` [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModel.__init__(provider)) The provider to use for authentication and the base URL. Defaults to `'openai'`. Azure OpenAI is not supported (its realtime endpoint uses a different URL and auth scheme). **`settings`** : `RealtimeModelSettings` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModel.__init__(settings)) [Model settings](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelSettings) used as defaults for realtime sessions. **`profile`** : `RealtimeModelProfileSpec` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModel.__init__(profile)) Optional override for the [realtime model profile](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelProfile) , merged over the provider’s — a partial dict, or a callable taking the resolved profile and returning the one to use. Mirrors `profile=` on a standard [`Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) , and is the escape hatch when a model name doesn’t identify the model (e.g. an Azure deployment named something other than its model). ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/openai/#attributes-1) #### client [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModel.client) The underlying [`AsyncOpenAI`](https://github.com/openai/openai-python) client from the provider. **Type:** `AsyncOpenAI` OpenAIRealtimeModelSettings --------------------------- [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModelSettings) **Bases:** `RealtimeModelSettings` Settings specific to OpenAI realtime models. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/openai/#attributes-2) #### openai\_input\_noise\_reduction [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModelSettings.openai_input_noise_reduction) Noise reduction tuned for `near_field` (headset) or `far_field` (laptop/conference) microphones. Absent disables it. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘near\_field’, ‘far\_field’\] #### openai\_output\_speed [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModelSettings.openai_output_speed) Playback speed multiplier for generated audio (0.25-1.5). **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) #### openai\_truncation [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModelSettings.openai_truncation) How the session truncates conversation context once it exceeds the model’s window. `'auto'` (the server default) drops the oldest turns; `'disabled'` keeps everything (and errors when the window is full); a `retention_ratio` truncation (`\{'type': 'retention_ratio', 'retention_ratio': 0.8\}`) keeps a fixed fraction, holding the prompt-cached prefix stable across turns (cached audio is far cheaper). This is the OpenAI SDK’s `truncation` shape, forwarded as-is. **Type:** `RealtimeTruncationParam` #### openai\_turn\_detection [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModelSettings.openai_turn_detection) OpenAI-specific server or semantic VAD configuration. When present, this fully overrides the cross-provider `turn_detection` setting. **Type:** `ServerVAD` | `SemanticVAD` #### openai\_voice [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModelSettings.openai_voice) Voice used for audio output, e.g. `alloy` or `VoiceID(id='voice_1234')`. The known prebuilt names provide autocomplete, while any string and the OpenAI SDK’s custom `VoiceID` form (`openai.types.realtime.realtime_audio_config_output.VoiceID`) are also accepted. **Type:** `KnownOpenAIRealtimeVoiceName` | [`str`](https://docs.python.org/3/library/stdtypes.html#str) | `VoiceID` SemanticVAD ----------- [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.SemanticVAD) **Bases:** [`TypedDict`](https://docs.python.org/3/library/typing.html#typing.TypedDict) Model-based semantic turn detection — uses a model to decide when the user is done speaking. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/openai/#attributes-3) #### create\_response [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.SemanticVAD.create_response) Whether to automatically generate a response when a turn ends. Defaults to `True`. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### eagerness [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.SemanticVAD.eagerness) How eagerly the model responds. Defaults to `'auto'`; `low` waits longer and `high` responds sooner. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘low’, ‘medium’, ‘high’, ‘auto’\] #### interrupt\_response [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.SemanticVAD.interrupt_response) Whether to interrupt an in-progress response when the user starts speaking. Defaults to `True`. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### type [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.SemanticVAD.type) The turn-detection type. Must be `'semantic_vad'`. **Type:** [`Required`](https://docs.python.org/3/library/typing.html#typing.Required) \[[`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal)\ \[‘semantic\_vad’\]\] ServerVAD --------- [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.ServerVAD) **Bases:** [`TypedDict`](https://docs.python.org/3/library/typing.html#typing.TypedDict) Server-side voice activity detection — the default turn-taking mode. The server detects when the user starts and stops speaking and (by default) commits the audio and triggers a response automatically. Unset fields fall back to the provider defaults. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/openai/#attributes-4) #### create\_response [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.ServerVAD.create_response) Whether to automatically generate a response when the user stops speaking. Defaults to `True`. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### idle\_timeout\_ms [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.ServerVAD.idle_timeout_ms) If set, auto-trigger a response after this much idle time with no detected speech. Defaults to the provider default. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### interrupt\_response [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.ServerVAD.interrupt_response) Whether to interrupt an in-progress response when the user starts speaking. Defaults to `True`. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### prefix\_padding\_ms [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.ServerVAD.prefix_padding_ms) Audio to include before detected speech, in milliseconds. Defaults to the provider default. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### silence\_duration\_ms [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.ServerVAD.silence_duration_ms) Silence required to detect the end of speech, in milliseconds. Defaults to the provider default. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### threshold [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.ServerVAD.threshold) Activation threshold (0.0-1.0). Higher requires louder audio; better in noisy environments. Defaults to the provider default. **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) #### type [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.ServerVAD.type) The turn-detection type. Must be `'server_vad'`. **Type:** [`Required`](https://docs.python.org/3/library/typing.html#typing.Required) \[[`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal)\ \[‘server\_vad’\]\] map\_event ---------- [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.map_event) def map_event(data: dict[str, Any]) -> RealtimeCodecEvent | None Map a raw OpenAI Realtime event to a [`RealtimeCodecEvent`](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeCodecEvent) . Returns `None` for events that carry no session-relevant content (e.g. `session.created`). ### Returns [](https://pydantic.dev/docs/ai/api/realtime/openai/#returns-1) `RealtimeCodecEvent` | [`None`](https://docs.python.org/3/library/constants.html#None) KnownOpenAIRealtimeVoiceName ---------------------------- [](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.KnownOpenAIRealtimeVoiceName) The prebuilt voices OpenAI’s realtime API ships, mirroring the `openai` SDK’s own `Voice` union. The [`openai_voice`](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModelSettings.openai_voice) setting also accepts any other string, so a voice OpenAI adds later works before this list catches up; a test pins the list against the SDK so it doesn’t silently fall behind. **Default:** `TypeAliasType('KnownOpenAIRealtimeVoiceName', Literal['alloy', 'ash', 'ballad', 'cedar', 'coral', 'echo', 'marin', 'sage', 'shimmer', 'verse'])` Was this page helpful? Thanks for your feedback! --- # Thinking | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/thinking/#_top) Thinking ======== Thinking (or reasoning) is the process by which a model works through a problem step-by-step before providing its final answer. The simplest way to enable thinking across supported providers is the [`Thinking`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Thinking) [capability](https://pydantic.dev/docs/ai/capabilities/overview/) . Provider-specific settings are available for advanced usage when you need direct access to a provider’s native thinking controls. Unified thinking settings ------------------------- [](https://pydantic.dev/docs/ai/capabilities/thinking/#unified-thinking-settings) Use the [`Thinking`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Thinking) capability to enable thinking: thinking\_capability.py from pydantic_ai import Agent from pydantic_ai.capabilities import Thinking agent = Agent('anthropic:claude-opus-4-7', capabilities=[Thinking(effort='high')]) You can also set the underlying `thinking` field in [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) directly: unified\_thinking.py from pydantic_ai import Agent agent = Agent('anthropic:claude-opus-4-7', model_settings={'thinking': 'high'}) The [`Thinking.effort`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Thinking.effort) value accepts: * `True` — enable thinking with the provider’s default effort level * `False` — disable thinking (silently ignored on always-on models) * `'minimal'` / `'low'` / `'medium'` / `'high'` / `'xhigh'` — enable thinking at a specific effort level (unsupported levels map to the closest available value) These are the same values accepted by the underlying `thinking` model setting. When omitted, the model uses its default behavior. Provider-specific settings (documented in the sections below) take precedence when both are set. ### Provider translation [](https://pydantic.dev/docs/ai/capabilities/thinking/#provider-translation) The `Thinking` capability maps each effort value to the selected provider’s native format: | Provider | `Thinking()` / `Thinking(effort=True)` | `Thinking(effort='high')` | Notes | | --- | --- | --- | --- | | Anthropic (Opus 4.6+) | `anthropic_thinking={'type': 'adaptive'}` | `{type: 'adaptive'}` + `effort='high'` | Claude Opus 4.7, 4.8, 5, and Sonnet 5 also support `effort='xhigh'` | | Anthropic (older) | `anthropic_thinking={'type': 'enabled', 'budget_tokens': 10000}` | `budget_tokens=16384` | Budget-based; `'low'` → 2048 tokens | | OpenAI | `reasoning_effort='medium'` | `reasoning_effort='high'` | GPT-5.6 maps unified `'minimal'` to `'low'` | | Google (Gemini 3+) | `include_thoughts=True` | `thinking_level='HIGH'` | | | Google (Gemini 2.5) | `include_thoughts=True` | `thinking_budget=24576` | | | Groq | `reasoning_format='parsed'` (gpt-oss also `reasoning_effort='medium'`) | `reasoning_format='parsed'` (gpt-oss also `reasoning_effort='high'`) | gpt-oss: unified effort → `reasoning_effort` (`low`/`medium`/`high`, via `extra_body`; always-on, so `thinking=False` is silently ignored); qwen3: `thinking=False` → `reasoning_effort='none'` (true disable, via `extra_body`); other reasoning models → `'hidden'` (suppresses output only) | | Mistral | `reasoning_effort='high'` | `reasoning_effort='high'` | Only on adjustable-reasoning models (e.g. `mistral-small-latest`, `mistral-medium-3-5`); `magistral` reasons always-on and gets no `reasoning_effort`. Mistral exposes only `'high'`/`'none'`, so every enabled level (incl. `'minimal'`) → `'high'` and only `thinking=False` → `'none'` | | OpenRouter | `reasoning={'effort': 'medium', 'enabled': True}` | `reasoning={'effort': 'high', 'enabled': True}` | `thinking=False` → `effort='none'`; always-on routes silently ignore; via `extra_body` | | Cerebras | `reasoning_effort` omitted (reasons by default) | `reasoning_effort` omitted | `thinking=False` → `reasoning_effort='none'`; gpt-oss reasons always-on, so `thinking=False` is silently ignored | | Snowflake Cortex | `reasoning={'effort': 'medium'}` | `reasoning={'effort': 'high'}` | Claude models only (via `extra_body`); sets `temperature=1` automatically; other families ignore `thinking` | | Crusoe | `reasoning_effort='medium'` | `reasoning_effort='high'` | Inherited from `OpenAIChatModel`; follows the vendor-prefixed model profile (`zai/`, `deepseek-ai/`, …). `thinking=False` → `'none'` only where that profile accepts it | | Ollama | `reasoning_effort='medium'` | `reasoning_effort='high'` | Inherited from `OpenAIChatModel`, so it follows the resolved model profile: `deepseek-r1` reasons, `gpt-oss` on Ollama sends nothing. `thinking=False` → `'none'` only on profiles that accept it | | Z.AI | `thinking={'type': 'enabled'}` | `thinking={'type': 'enabled'}`, plus `reasoning_effort='high'` on GLM-5.2 | Via `extra_body`; `thinking=False` → `type='disabled'`. Only GLM-5.2 takes a per-request effort, so on other models every enabled level behaves the same | | xAI | `reasoning_effort` omitted on Grok 4.3 (uses its default) | `reasoning_effort='high'` | Grok 4.3 supports `'none'`, `'low'`, `'medium'`, and `'high'`, and `thinking=True` omits the parameter so the model applies its own default; Grok 3 Mini only supports `'low'` and `'high'` (so `thinking=True` → `'high'`) and silently ignores `thinking=False`; Grok 4.5 supports `'low'`, `'medium'`, and `'high'` but not `'none'`, so it reasons always-on (`thinking=True` → `'medium'`) and silently ignores `thinking=False` | | Bedrock (Claude 4.6+) | `thinking.type='adaptive'` | `{type: 'adaptive'}` + `output_config.effort='high'` | Effort lives in the sibling `output_config` field per AWS docs; `xhigh` maps to `max` | | Bedrock (Claude older) | `thinking.type='enabled'` | `budget_tokens=16384` | Budget-based | | Bedrock (OpenAI) | `reasoning_effort='medium'` | `reasoning_effort='high'` | Converse rejects `'none'`; `thinking=False` silently ignored | | Bedrock (Qwen) | `reasoning_config='high'` | `reasoning_config='high'` | Only `'low'` and `'high'`; `thinking=False` silently ignored | | Bedrock Mantle | `reasoning={'effort': 'medium'}` | `reasoning={'effort': 'high'}` | Served on the Responses API, so effort rides the `reasoning` object; `thinking=False` → `effort='none'` | OpenAI ------ [](https://pydantic.dev/docs/ai/capabilities/thinking/#openai) When using the [`OpenAIChatModel`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIChatModel) , text output inside `` tags are converted to [`ThinkingPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ThinkingPart) objects. You can customize the tags using the [`thinking_tags`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfile.thinking_tags) field on the [model profile](https://pydantic.dev/docs/ai/models/openai/#model-profile) . Some [OpenAI-compatible model providers](https://pydantic.dev/docs/ai/models/openai/#openai-compatible-models) might also support native thinking parts that are not delimited by tags. Instead, they are sent and received as separate, custom fields in the API. Typically, if you are calling the model via the `:` shorthand, Pydantic AI handles it for you. Nonetheless, you can still configure the fields with [`openai_chat_thinking_field`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.openai.OpenAIModelProfile.openai_chat_thinking_field) . If your provider recommends to send back these custom fields not changed, for caching or interleaved thinking benefits, you can also achieve this with [`openai_chat_send_back_thinking_parts`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.openai.OpenAIModelProfile.openai_chat_send_back_thinking_parts) . ### OpenAI Responses [](https://pydantic.dev/docs/ai/capabilities/thinking/#openai-responses) The [`OpenAIResponsesModel`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIResponsesModel) can generate native thinking parts. To enable this functionality, you need to set the `OpenAIResponsesModelSettings.openai_reasoning_effort` and [`OpenAIResponsesModelSettings.openai_reasoning_summary`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIResponsesModelSettings.openai_reasoning_summary) [model settings](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) . Models that support it can additionally use a `pro` [reasoning mode](https://pydantic.dev/docs/ai/models/openai/#reasoning-mode) , which is independent of the effort and never set by the unified `thinking` setting. By default, the unique IDs of reasoning, text, and function call parts from the message history are sent to the model, which can result in errors like `"Item 'rs_123' of type 'reasoning' was provided without its required following item."` if the message history you’re sending does not match exactly what was received from the Responses API in a previous response, for example if you’re using a [history processor](https://pydantic.dev/docs/ai/core-concepts/message-history/#processing-message-history) . To disable this, you can disable the [`OpenAIResponsesModelSettings.openai_send_reasoning_ids`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIResponsesModelSettings.openai_send_reasoning_ids) [model setting](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) . openai\_thinking\_part.py from pydantic_ai import Agent from pydantic_ai.models.openai import OpenAIResponsesModel, OpenAIResponsesModelSettings model = OpenAIResponsesModel('gpt-5.6-sol') settings = OpenAIResponsesModelSettings( openai_reasoning_effort='low', openai_reasoning_summary='detailed', ) agent = Agent(model, model_settings=settings) ... Anthropic --------- [](https://pydantic.dev/docs/ai/capabilities/thinking/#anthropic) To enable thinking, use the [`AnthropicModelSettings.anthropic_thinking`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_thinking) [model setting](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) . anthropic\_thinking\_part.py from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModel, AnthropicModelSettings model = AnthropicModel('claude-sonnet-4-5') settings = AnthropicModelSettings( anthropic_thinking={'type': 'enabled', 'budget_tokens': 1024}, ) agent = Agent(model, model_settings=settings) ... Anthropic reports how many thinking tokens it used in [`RunUsage.details`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage.details) under the `thinking_tokens` key. They are billed within `output_tokens`, so they are a readable subset of the output total rather than an addition to it, and the key is omitted entirely when a response used no thinking tokens. ### Interleaved Thinking [](https://pydantic.dev/docs/ai/capabilities/thinking/#interleaved-thinking) To enable [interleaved thinking](https://docs.anthropic.com/en/docs/build-with-claude/extended-thinking#interleaved-thinking) , you need to include the beta header in your model settings: anthropic\_interleaved\_thinking.py from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModel, AnthropicModelSettings model = AnthropicModel('claude-sonnet-4-5') settings = AnthropicModelSettings( anthropic_thinking={'type': 'enabled', 'budget_tokens': 10000}, extra_headers={'anthropic-beta': 'interleaved-thinking-2025-05-14'}, ) agent = Agent(model, model_settings=settings) ... ### Adaptive Thinking & Effort [](https://pydantic.dev/docs/ai/capabilities/thinking/#adaptive-thinking--effort) Starting with `claude-opus-4-6`, Anthropic supports [adaptive thinking](https://docs.anthropic.com/en/docs/build-with-claude/adaptive-thinking) , where the model dynamically decides when and how much to think based on the complexity of each request. This replaces extended thinking (`type: 'enabled'` with `budget_tokens`) which is deprecated on Opus 4.6 and removed on Opus 4.7, 4.8, 5, and Sonnet 5. Claude Opus 4.7, 4.8, 5, and Sonnet 5 also add the `xhigh` effort level. Adaptive thinking also automatically enables interleaved thinking. anthropic\_adaptive\_thinking.py from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModel, AnthropicModelSettings model = AnthropicModel('claude-opus-4-8') settings = AnthropicModelSettings( anthropic_thinking={'type': 'adaptive'}, anthropic_effort='high', ) agent = Agent(model, model_settings=settings) ... The [`anthropic_effort`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_effort) setting controls how much effort the model puts into its response (independent of thinking). See the [Anthropic effort docs](https://docs.anthropic.com/en/docs/build-with-claude/effort) for details. Thinking tokens count against Anthropic’s loop-wide [task budgets](https://pydantic.dev/docs/ai/models/anthropic/#task-budgets-beta) , so adaptive thinking naturally scales down as the budget depletes. Google ------ [](https://pydantic.dev/docs/ai/capabilities/thinking/#google) For advanced usage, use the [`GoogleModelSettings.google_thinking_config`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings.google_thinking_config) [model setting](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) . google\_thinking\_part.py from pydantic_ai import Agent from pydantic_ai.models.google import GoogleModel, GoogleModelSettings model = GoogleModel('gemini-3.5-flash') settings = GoogleModelSettings(google_thinking_config={'include_thoughts': True, 'thinking_level': 'MEDIUM'}) agent = Agent(model, model_settings=settings) ... See the [Google model docs](https://pydantic.dev/docs/ai/models/google/#configure-thinking) for more details. xAI --- [](https://pydantic.dev/docs/ai/capabilities/thinking/#xai) xAI reasoning models (Grok) support native thinking. To preserve the thinking content for multi-turn conversations, enable [`XaiModelSettings.xai_include_encrypted_content`](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_include_encrypted_content) . xai\_thinking\_part.py from pydantic_ai import Agent from pydantic_ai.models.xai import XaiModel, XaiModelSettings model = XaiModel('grok-4.3') settings = XaiModelSettings(xai_include_encrypted_content=True) agent = Agent(model, model_settings=settings) ... Bedrock ------- [](https://pydantic.dev/docs/ai/capabilities/thinking/#bedrock) For Claude Sonnet 4.6+ and Opus 4.6+, Pydantic AI’s unified `thinking` setting translates to AWS’s required [adaptive thinking](https://docs.aws.amazon.com/bedrock/latest/userguide/claude-messages-adaptive-thinking.html) shape automatically — set [`ModelSettings.thinking`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.thinking) and you’re done. For older Claude models or to pin a specific `budget_tokens`, you can still use [`BedrockModelSettings.bedrock_additional_model_requests_fields`](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_additional_model_requests_fields) [model setting](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) to pass provider-specific configuration directly: * [Claude](https://pydantic.dev/docs/ai/capabilities/thinking/#tab-panel-14) * [OpenAI](https://pydantic.dev/docs/ai/capabilities/thinking/#tab-panel-15) * [Qwen](https://pydantic.dev/docs/ai/capabilities/thinking/#tab-panel-16) * [Deepseek](https://pydantic.dev/docs/ai/capabilities/thinking/#tab-panel-17) bedrock\_claude\_thinking\_part.py from pydantic_ai import Agent from pydantic_ai.models.bedrock import BedrockConverseModel, BedrockModelSettings model = BedrockConverseModel('us.anthropic.claude-sonnet-4-5-20250929-v1:0') model_settings = BedrockModelSettings( bedrock_additional_model_requests_fields={ 'thinking': {'type': 'enabled', 'budget_tokens': 1024} } ) agent = Agent(model=model, model_settings=model_settings) bedrock\_openai\_thinking\_part.py from pydantic_ai import Agent from pydantic_ai.models.bedrock import BedrockConverseModel, BedrockModelSettings model = BedrockConverseModel('openai.gpt-oss-120b-1:0') model_settings = BedrockModelSettings( bedrock_additional_model_requests_fields={'reasoning_effort': 'low'} ) agent = Agent(model=model, model_settings=model_settings) bedrock\_qwen\_thinking\_part.py from pydantic_ai import Agent from pydantic_ai.models.bedrock import BedrockConverseModel, BedrockModelSettings model = BedrockConverseModel('qwen.qwen3-32b-v1:0') model_settings = BedrockModelSettings( bedrock_additional_model_requests_fields={'reasoning_config': 'high'} ) agent = Agent(model=model, model_settings=model_settings) Reasoning is [always enabled](https://docs.aws.amazon.com/bedrock/latest/userguide/inference-reasoning.html) for Deepseek model bedrock\_deepseek\_thinking\_part.py from pydantic_ai import Agent from pydantic_ai.models.bedrock import BedrockConverseModel model = BedrockConverseModel('us.deepseek.r1-v1:0') agent = Agent(model=model) Groq ---- [](https://pydantic.dev/docs/ai/capabilities/thinking/#groq) Groq supports different formats to receive thinking parts: * `"raw"`: The thinking part is included in the text content inside `` tags, which are automatically converted to [`ThinkingPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ThinkingPart) objects. * `"hidden"`: The thinking part is not included in the text content. * `"parsed"`: The thinking part has its own structured part in the response which is converted into a [`ThinkingPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ThinkingPart) object. The unified [`ModelSettings.thinking`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.thinking) setting works across providers: it selects `reasoning_format='parsed'` so thinking parts are returned, and for the gpt-oss family its effort level also drives Groq’s `reasoning_effort` (`minimal`/`low` → `'low'`, `medium` → `'medium'`, `high`/`xhigh` → `'high'`, `True` → `'medium'`). Two composable [model settings](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) give finer control: [`GroqModelSettings.groq_reasoning_format`](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqModelSettings.groq_reasoning_format) selects how thinking parts are returned (the formats above), and [`GroqModelSettings.groq_reasoning_effort`](https://pydantic.dev/docs/ai/api/models/groq/#pydantic_ai.models.groq.GroqModelSettings.groq_reasoning_effort) (sent to Groq as `reasoning_effort`) controls how much the model reasons, taking precedence over the unified `thinking` mapping: groq\_thinking\_part.py from pydantic_ai import Agent from pydantic_ai.models.groq import GroqModel, GroqModelSettings model = GroqModel('openai/gpt-oss-120b') settings = GroqModelSettings(groq_reasoning_format='parsed', groq_reasoning_effort='medium') agent = Agent(model, model_settings=settings) ... OpenRouter ---------- [](https://pydantic.dev/docs/ai/capabilities/thinking/#openrouter) To enable thinking, use the [`OpenRouterModelSettings.openrouter_reasoning`](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModelSettings.openrouter_reasoning) [model setting](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) . openrouter\_thinking\_part.py from pydantic_ai import Agent from pydantic_ai.models.openrouter import OpenRouterModel, OpenRouterModelSettings model = OpenRouterModel('openai/gpt-5.2') settings = OpenRouterModelSettings(openrouter_reasoning={'effort': 'high'}) agent = Agent(model, model_settings=settings) ... Z.AI ---- [](https://pydantic.dev/docs/ai/capabilities/thinking/#zai) To enable thinking, use the unified [`thinking`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.thinking) [model setting](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) . To preserve thinking content across multi-turn conversations, also set [`ZaiModelSettings.zai_clear_thinking`](https://pydantic.dev/docs/ai/api/models/zai/#pydantic_ai.models.zai.ZaiModelSettings.zai_clear_thinking) to `False`. zai\_thinking\_part.py from pydantic_ai import Agent from pydantic_ai.models.zai import ZaiModel, ZaiModelSettings model = ZaiModel('glm-5') settings = ZaiModelSettings(thinking=True, zai_clear_thinking=False) agent = Agent(model, model_settings=settings) ... Snowflake Cortex ---------------- [](https://pydantic.dev/docs/ai/capabilities/thinking/#snowflake-cortex) To enable thinking on Claude models, use the unified [`thinking`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.thinking) [model setting](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) , or set [`SnowflakeModelSettings.snowflake_reasoning`](https://pydantic.dev/docs/ai/api/models/snowflake/#pydantic_ai.models.snowflake.SnowflakeModelSettings.snowflake_reasoning) directly to control the reasoning token budget: snowflake\_thinking\_part.py from pydantic_ai import Agent from pydantic_ai.models.snowflake import SnowflakeModel, SnowflakeModelSettings model = SnowflakeModel('claude-sonnet-4-6') settings = SnowflakeModelSettings(snowflake_reasoning={'max_tokens': 4096}) agent = Agent(model, model_settings=settings) ... On OpenAI models, use the unified `thinking` setting or [`openai_reasoning_effort`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIChatModelSettings.openai_reasoning_effort) . Claude requires `temperature` to be exactly 1 when thinking is enabled, but Cortex applies a different default when the request doesn’t specify one, so `SnowflakeModel` sets `temperature` to 1 automatically when reasoning is enabled and you haven’t set it explicitly. Mistral ------- [](https://pydantic.dev/docs/ai/capabilities/thinking/#mistral) The `magistral` family always reasons and does not need to be specifically enabled; `thinking=False` is silently ignored. Mistral has [deprecated](https://docs.mistral.ai/resources/deprecated/native-reasoning) the `magistral` family in favor of the adjustable-reasoning models below. Models with adjustable reasoning (the Mistral Small 4 and Medium 3.5 families: `mistral-small-latest`, `mistral-small-2603`, `mistral-medium-latest`, `mistral-medium`, `mistral-medium-3`, `mistral-medium-3-5`, `mistral-medium-3.5`, `mistral-medium-2604`) are controlled via the unified [`thinking`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.thinking) setting, which maps to Mistral’s `reasoning_effort`. Mistral exposes only `'high'` (full thinking) and `'none'` (thinking suppressed), so every enabled level maps to `'high'` and only `thinking=False` maps to `'none'`. Older `mistral-small-*` / `mistral-medium-*` snapshots do not support reasoning, so `thinking` is silently ignored for them. Adjustable reasoning applies when using the native Mistral provider; OpenAI-compatible providers that host these models (such as LiteLLM or Azure) do not support it and `thinking` is ignored there. OpenRouter is the exception: it maps the unified `thinking` setting to its own `reasoning` parameter for any model it routes. Cohere ------ [](https://pydantic.dev/docs/ai/capabilities/thinking/#cohere) Thinking is supported by the `command-a-reasoning-08-2025` model. It does not need to be specifically enabled. Hugging Face ------------ [](https://pydantic.dev/docs/ai/capabilities/thinking/#hugging-face) Text output inside `` tags is automatically converted to [`ThinkingPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ThinkingPart) objects. You can customize the tags using the [`thinking_tags`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfile.thinking_tags) field on the [model profile](https://pydantic.dev/docs/ai/models/openai/#model-profile) . Was this page helpful? Thanks for your feedback! --- # Built-in Evaluators | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#_top) Built-in Evaluators =================== Pydantic Evals provides several built-in evaluators for common evaluation tasks. Comparison Evaluators --------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#comparison-evaluators) ### EqualsExpected [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#equalsexpected) Check if the output exactly equals the expected output from the case. from pydantic_evals.evaluators import EqualsExpected EqualsExpected() **Parameters:** None **Returns:** `bool` - `True` if `ctx.output == ctx.expected_output` **Example:** from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import EqualsExpected dataset = Dataset( name='equals_expected_demo', cases=[\ Case(\ name='addition',\ inputs='2 + 2',\ expected_output='4',\ ),\ ], evaluators=[EqualsExpected()], ) **Notes:** * Skips evaluation if `expected_output` is `None` (returns empty dict `{}`) * Uses Python’s `==` operator, so works with any comparable types * For structured data, considers nested equality * * * ### Equals [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#equals) Check if the output equals a specific value. from pydantic_evals.evaluators import Equals Equals(value='expected_result') **Parameters:** * `value` (Any): The value to compare against * `evaluation_name` (str | None): Custom name for this evaluation in reports **Returns:** `bool` - `True` if `ctx.output == value` **Example:** from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import Equals # Check output is always "success" dataset = Dataset( name='equals_demo', cases=[Case(inputs='test')], evaluators=[\ Equals(value='success', evaluation_name='is_success'),\ ], ) **Use Cases:** * Checking for sentinel values * Validating consistent outputs * Testing classification into specific categories * * * ### Contains [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#contains) Check if the output contains a specific value or substring. from pydantic_evals.evaluators import Contains Contains( value='substring', case_sensitive=True, as_strings=False, ) **Parameters:** * `value` (Any): The value to search for * `case_sensitive` (bool): Case-sensitive comparison for strings (default: `True`) * `as_strings` (bool): Convert both values to strings before checking (default: `False`) * `evaluation_name` (str | None): Custom name for this evaluation in reports **Returns:** [`EvaluationReason`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluationReason) - Pass/fail with explanation **Behavior:** For **strings**: checks substring containment * `Contains(value='hello', case_sensitive=False)` * Matches: “Hello World”, “say hello”, “HELLO” * Doesn’t match: “hi there” For **lists/tuples**: checks membership * `Contains(value='apple')` * Matches: `['apple', 'banana']`, `('apple',)` * Doesn’t match: `['apples', 'orange']` For **dicts**: checks key-value pairs * `Contains(value={'name': 'Alice'})` * Matches: `{'name': 'Alice', 'age': 30}` * Doesn’t match: `{'name': 'Bob'}` **Example:** from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import Contains dataset = Dataset( name='contains_demo', cases=[Case(inputs='test')], evaluators=[\ # Check for required keywords\ Contains(value='terms and conditions', case_sensitive=False),\ # Check for PII (fail if found)\ # Note: Use a custom evaluator that returns False when PII found\ ], ) **Use Cases:** * Required content verification * Keyword detection * PII/sensitive data detection * Multi-value validation * * * Type Validation --------------- [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#type-validation) ### IsInstance [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#isinstance) Check if the output is an instance of a type with the given name. from pydantic_evals.evaluators import IsInstance IsInstance(type_name='str') **Parameters:** * `type_name` (str): The type name to check (uses `__name__` or `__qualname__`) * `evaluation_name` (str | None): Custom name for this evaluation in reports **Returns:** [`EvaluationReason`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluationReason) - Pass/fail with type information **Example:** from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import IsInstance dataset = Dataset( name='isinstance_demo', cases=[Case(inputs='test')], evaluators=[\ # Check output is always a string\ IsInstance(type_name='str'),\ # Check for Pydantic model\ IsInstance(type_name='MyModel'),\ # Check for dict\ IsInstance(type_name='dict'),\ ], ) **Notes:** * Matches against both `__name__` and `__qualname__` of the type * Works with built-in types (`str`, `int`, `dict`, `list`, etc.) * Works with custom classes and Pydantic models * Checks the entire MRO (Method Resolution Order) for inheritance **Use Cases:** * Format validation * Structured output verification * Type consistency checks * * * Performance Evaluation ---------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#performance-evaluation) ### MaxDuration [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#maxduration) Check if task execution time is under a maximum threshold. from datetime import timedelta from pydantic_evals.evaluators import MaxDuration MaxDuration(seconds=2.0) # or MaxDuration(seconds=timedelta(seconds=2)) **Parameters:** * `seconds` (float | timedelta): Maximum allowed duration **Returns:** `bool` - `True` if `ctx.duration <= seconds` **Example:** from datetime import timedelta from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import MaxDuration dataset = Dataset( name='max_duration_demo', cases=[Case(inputs='test')], evaluators=[\ # SLA: must respond in under 2 seconds\ MaxDuration(seconds=2.0),\ # Or using timedelta\ MaxDuration(seconds=timedelta(milliseconds=500)),\ ], ) **Use Cases:** * SLA compliance * Performance regression testing * Latency requirements * Timeout validation **See Also:** [Concurrency & Performance](https://pydantic.dev/docs/ai/evals/how-to/concurrency/) * * * LLM-as-a-Judge -------------- [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#llm-as-a-judge) ### LLMJudge [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#llmjudge) Use an LLM to evaluate subjective qualities based on a rubric. from pydantic_evals.evaluators import LLMJudge LLMJudge( rubric='Response is accurate and helpful', model='openai:gpt-5.2', include_input=False, include_expected_output=False, model_settings=None, score=False, assertion={'include_reason': True}, ) **Parameters:** * `rubric` (str): The evaluation criteria (required) * `model` (Model | KnownModelName | None): Model to use (default: `'openai:gpt-5.2'`) * `include_input` (bool): Include task inputs in the prompt (default: `False`) * `include_expected_output` (bool): Include expected output in the prompt (default: `False`) * `model_settings` (ModelSettings | None): Custom model settings * `score` (OutputConfig | False): Configure score output (default: `False`) * `assertion` (OutputConfig | False): Configure assertion output (default: includes reason) **Returns:** Depends on `score` and `assertion` parameters (see below) **Output Modes:** By default, returns a **boolean assertion** with reason: * `LLMJudge(rubric='Response is polite')` * Returns: `{'LLMJudge_pass': EvaluationReason(value=True, reason='...')}` Return a **score** (0.0 to 1.0) instead: * `LLMJudge(rubric='Response quality', score={'include_reason': True}, assertion=False)` * Returns: `{'LLMJudge_score': EvaluationReason(value=0.85, reason='...')}` Return **both** score and assertion: * `LLMJudge(rubric='Response quality', score={'include_reason': True}, assertion={'include_reason': True})` * Returns: `{'LLMJudge_score': EvaluationReason(value=0.85, reason='...'), 'LLMJudge_pass': EvaluationReason(value=True, reason='...')}` **Customize evaluation names:** * `LLMJudge(rubric='Response is factually accurate', assertion={'evaluation_name': 'accuracy', 'include_reason': True})` * Returns: `{'accuracy': EvaluationReason(value=True, reason='...')}` **Example:** from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import LLMJudge dataset = Dataset( name='llm_judge_demo', cases=[Case(inputs='test', expected_output='result')], evaluators=[\ # Basic accuracy check\ LLMJudge(\ rubric='Response is factually accurate',\ include_input=True,\ ),\ # Quality score with different model\ LLMJudge(\ rubric='Overall response quality',\ model='anthropic:claude-sonnet-4-6',\ score={'evaluation_name': 'quality', 'include_reason': False},\ assertion=False,\ ),\ # Check against expected output\ LLMJudge(\ rubric='Response matches the expected answer semantically',\ include_input=True,\ include_expected_output=True,\ ),\ ], ) **See Also:** [LLM Judge Deep Dive](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/) ### GEval [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#geval) Chain-of-thought evaluation following the G-Eval method (Liu et al., 2023): the judge applies explicit evaluation steps and returns an integer score in `score_range` with a reasoning trace. from pydantic_evals.evaluators import GEval GEval( criteria='coherence', evaluation_steps=[\ 'Read the output carefully.',\ 'Check that each sentence follows logically from the previous one.',\ 'Assign a score from 1 (incoherent) to 5 (fully coherent).',\ ], score_range=(1, 5), include_input=False, ) **Parameters:** * `criteria` (str): The aspect being evaluated, e.g. `'coherence'` (required) * `evaluation_steps` (list\[str\]): Explicit chain-of-thought steps the judge should follow (required) * `score_range` (tuple\[int, int\]): Inclusive integer score range (default: `(1, 5)`) * `include_input` (bool): Include task inputs in the prompt (default: `False`) * `model` (Model | KnownModelName | None): Model to use (default: `'openai:gpt-5.2'`) * `model_settings` (ModelSettings | None): Custom model settings * `evaluation_name` (str | None): Custom name for the result (default: `'GEval'`) **Returns:** `EvaluationReason` with the integer score and the judge’s reasoning **See Also:** [Standard Quality Metrics](https://pydantic.dev/docs/ai/evals/evaluators/standard-quality-metrics/) * * * Span-Based Evaluation --------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#span-based-evaluation) ### HasMatchingSpan [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#hasmatchingspan) Check if OpenTelemetry spans match a query (requires Logfire configuration). from pydantic_evals.evaluators import HasMatchingSpan HasMatchingSpan( query={'name_contains': 'tool_call'}, evaluation_name='called_tool', ) **Parameters:** * `query` ([`SpanQuery`](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanQuery) ): Query to match against spans * `evaluation_name` (str | None): Custom name for this evaluation in reports **Returns:** `bool` - `True` if any span matches the query **Example:** from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import HasMatchingSpan dataset = Dataset( name='span_check_demo', cases=[Case(inputs='test')], evaluators=[\ # Check that a specific tool was called\ HasMatchingSpan(\ query={'name_contains': 'search_database'},\ evaluation_name='used_database',\ ),\ # Check for errors\ HasMatchingSpan(\ query={'has_status': 'error'},\ evaluation_name='had_errors',\ ),\ # Check duration constraints\ HasMatchingSpan(\ query={\ 'name_equals': 'llm_call',\ 'max_duration': 2.0, # seconds\ },\ evaluation_name='llm_fast_enough',\ ),\ ], ) **See Also:** [Span-Based Evaluation](https://pydantic.dev/docs/ai/evals/evaluators/span-based/) * * * Native Report Evaluators ------------------------ [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#native-report-evaluators) In addition to the case-level evaluators above, Pydantic Evals provides report evaluators that analyze entire experiment results. These are passed via the `report_evaluators` parameter on `Dataset`. | Report Evaluator | Purpose | Output | | --- | --- | --- | | [`ConfusionMatrixEvaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.ConfusionMatrixEvaluator) | Classification confusion matrix | `ConfusionMatrix` | | [`PrecisionRecallEvaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.PrecisionRecallEvaluator) | PR curve with AUC | `PrecisionRecall` | **See:** [Report Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/) for full documentation, parameters, and examples, including how to write custom report evaluators that produce `ScalarResult` and `TableResult` analyses. * * * Quick Reference Table --------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#quick-reference-table) ### Case-Level Evaluators [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#case-level-evaluators) | Evaluator | Purpose | Return Type | Cost | Speed | | --- | --- | --- | --- | --- | | [`EqualsExpected`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EqualsExpected) | Exact match with expected | `bool` | Free | Instant | | [`Equals`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Equals) | Equals specific value | `bool` | Free | Instant | | [`Contains`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Contains) | Contains value/substring | `bool` + reason | Free | Instant | | [`IsInstance`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.IsInstance) | Type validation | `bool` + reason | Free | Instant | | [`MaxDuration`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.MaxDuration) | Performance threshold | `bool` | Free | Instant | | [`LLMJudge`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.LLMJudge) | Subjective quality | `bool` and/or `float` | $$ | Slow | | [`GEval`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.GEval) | Chain-of-thought scoring | `int` + reason | $$ | Slow | | [`HasMatchingSpan`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.HasMatchingSpan) | Behavioral check | `bool` | Free | Fast | ### Report-Level Evaluators [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#report-level-evaluators) | Evaluator | Purpose | Output Type | Cost | Speed | | --- | --- | --- | --- | --- | | [`ConfusionMatrixEvaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.ConfusionMatrixEvaluator) | Classification matrix | `ConfusionMatrix` | Free | Instant | | [`PrecisionRecallEvaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.PrecisionRecallEvaluator) | PR curve with AUC | `PrecisionRecall` | Free | Instant | Combining Evaluators -------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#combining-evaluators) Best practice is to combine fast deterministic checks with slower LLM evaluations: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import ( Contains, IsInstance, LLMJudge, MaxDuration, ) dataset = Dataset( name='combined_evaluators', cases=[Case(inputs='test')], evaluators=[\ # Fast checks first (fail fast)\ IsInstance(type_name='str'),\ Contains(value='required_field'),\ MaxDuration(seconds=2.0),\ # Expensive LLM checks last\ LLMJudge(rubric='Response is helpful and accurate'),\ ], ) This approach: 1. Catches format/structure issues immediately 2. Validates required content quickly 3. Only runs expensive LLM evaluation if basic checks pass 4. Provides comprehensive quality assessment Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/evaluators/built-in/#next-steps) * **[LLM Judge](https://pydantic.dev/docs/ai/evals/evaluators/llm-judge/) ** - Deep dive on LLM-as-a-Judge evaluation * **[Custom Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/custom/) ** - Write your own evaluation logic * **[Report Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/) ** - Experiment-wide analyses (confusion matrices, PR curves, etc.) * **[Span-Based Evaluation](https://pydantic.dev/docs/ai/evals/evaluators/span-based/) ** - Using OpenTelemetry spans for behavioral checks Was this page helpful? Thanks for your feedback! --- # Custom Evaluators | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/evaluators/custom/#_top) Custom Evaluators ================= Write custom evaluators for domain-specific logic, external integrations, or specialized metrics. Basic Custom Evaluator ---------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#basic-custom-evaluator) All evaluators inherit from [`Evaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Evaluator) and must implement `evaluate`: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class ExactMatch(Evaluator): """Check if output exactly matches expected output.""" def evaluate(self, ctx: EvaluatorContext) -> bool: return ctx.output == ctx.expected_output **Key Points:** * Use `@dataclass` decorator (required) * Inherit from `Evaluator` * Implement `evaluate(self, ctx: EvaluatorContext) -> EvaluatorOutput` * Return `bool`, `int`, `float`, `str`, [`EvaluationReason`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluationReason) , or `dict` of these EvaluatorContext ---------------- [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#evaluatorcontext) The context provides all information about the case execution: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class MyEvaluator(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: # Access case data ctx.name # Case name ctx.inputs # Task inputs ctx.metadata # Case metadata ctx.expected_output # Expected output (may be None) ctx.output # Actual output # Performance data ctx.duration # Task execution time (seconds) # Custom metrics/attributes (see metrics guide) ctx.metrics # dict[str, int | float] ctx.attributes # dict[str, Any] # OpenTelemetry spans (if logfire configured) ctx.span_tree # SpanTree for behavioral checks return True Evaluator Parameters -------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#evaluator-parameters) Add configurable parameters as dataclass fields: from dataclasses import dataclass from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class ContainsKeyword(Evaluator): keyword: str case_sensitive: bool = True def evaluate(self, ctx: EvaluatorContext) -> bool: output = ctx.output keyword = self.keyword if not self.case_sensitive: output = output.lower() keyword = keyword.lower() return keyword in output # Usage dataset = Dataset( name='keyword_check', cases=[Case(name='test', inputs='This is important')], evaluators=[\ ContainsKeyword(keyword='important', case_sensitive=False),\ ], ) Return Types ------------ [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#return-types) ### Boolean Assertions [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#boolean-assertions) Simple pass/fail checks: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class IsValidJSON(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: try: import json json.loads(ctx.output) return True except Exception: return False ### Numeric Scores [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#numeric-scores) Quality metrics: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class LengthScore(Evaluator): """Score based on output length (0.0 = too short, 1.0 = ideal).""" ideal_length: int = 100 tolerance: int = 20 def evaluate(self, ctx: EvaluatorContext) -> float: length = len(ctx.output) diff = abs(length - self.ideal_length) if diff <= self.tolerance: return 1.0 else: # Decay score as we move away from ideal score = max(0.0, 1.0 - (diff - self.tolerance) / self.ideal_length) return score ### String Labels [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#string-labels) Categorical classifications: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class SentimentClassifier(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> str: output_lower = ctx.output.lower() if any(word in output_lower for word in ['error', 'failed', 'wrong']): return 'negative' elif any(word in output_lower for word in ['success', 'correct', 'great']): return 'positive' else: return 'neutral' ### With Reasons [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#with-reasons) Add explanations to any result: from dataclasses import dataclass from pydantic_evals.evaluators import EvaluationReason, Evaluator, EvaluatorContext @dataclass class SmartCheck(Evaluator): threshold: float = 0.8 def evaluate(self, ctx: EvaluatorContext) -> EvaluationReason: score = self._calculate_score(ctx.output) if score >= self.threshold: return EvaluationReason( value=True, reason=f'Score {score:.2f} exceeds threshold {self.threshold}', ) else: return EvaluationReason( value=False, reason=f'Score {score:.2f} below threshold {self.threshold}', ) def _calculate_score(self, output: str) -> float: # Your scoring logic return 0.75 ### Multiple Results [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#multiple-results) You can return multiple evaluations from one evaluator by returning a dictionary of key-value pairs. from dataclasses import dataclass from pydantic_evals.evaluators import ( EvaluationReason, Evaluator, EvaluatorContext, EvaluatorOutput, ) @dataclass class ComprehensiveCheck(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> EvaluatorOutput: format_valid = self._check_format(ctx.output) return { 'valid_format': EvaluationReason( value=format_valid, reason='Valid JSON format' if format_valid else 'Invalid JSON format', ), 'quality_score': self._score_quality(ctx.output), # float 'category': self._classify(ctx.output), # str } def _check_format(self, output: str) -> bool: return output.startswith('{') and output.endswith('}') def _score_quality(self, output: str) -> float: return len(output) / 100.0 def _classify(self, output: str) -> str: return 'short' if len(output) < 50 else 'long' Each key in the returned dictionary becomes a separate result in the report. Values can be: * Primitives (`bool`, `int`, `float`, `str`) * [`EvaluationReason`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluationReason) (value with explanation) * Nested dicts of these types The [`EvaluatorOutput`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorOutput) type represents all legal values that can be returned by an evaluator, and can be used as the return type annotation for your custom `evaluate` method. ### Conditional Results [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#conditional-results) Evaluators can dynamically choose whether to produce results for a given case by returning an empty dict when not applicable: from dataclasses import dataclass from pydantic_evals.evaluators import ( EvaluationReason, Evaluator, EvaluatorContext, EvaluatorOutput, ) @dataclass class SQLValidator(Evaluator): """Only evaluates SQL queries, skips other outputs.""" def evaluate(self, ctx: EvaluatorContext) -> EvaluatorOutput: # Check if this case is relevant for SQL validation if not isinstance(ctx.output, str) or not ctx.output.strip().upper().startswith( ('SELECT', 'INSERT', 'UPDATE', 'DELETE') ): # Return empty dict - this evaluator doesn't apply to this case return {} # This is a SQL query, perform validation try: # In real implementation, use sqlparse or similar is_valid = self._validate_sql(ctx.output) return { 'sql_valid': is_valid, 'sql_complexity': self._measure_complexity(ctx.output), } except Exception as e: return {'sql_valid': EvaluationReason(False, reason=f'Exception: {e}')} def _validate_sql(self, query: str) -> bool: # Simplified validation return 'FROM' in query.upper() or 'INTO' in query.upper() def _measure_complexity(self, query: str) -> str: joins = query.upper().count('JOIN') if joins == 0: return 'simple' elif joins <= 2: return 'moderate' else: return 'complex' This pattern is useful when: * An evaluator only applies to certain types of outputs (e.g., code validation only for code outputs) * Validation depends on metadata tags (e.g., only evaluate cases marked with `language='python'`) * You want to run expensive checks conditionally based on other evaluator results **Key Points:** * Returning `{}` means “this evaluator doesn’t apply here” - the case won’t show results from this evaluator * Returning `{'key': value}` means “this evaluator applies and here are the results” * This is more practical than using case-level evaluators when it applies to a large fraction of cases, or when the condition is based on the output itself * The evaluator still runs for every case, but can short-circuit when not relevant Async Evaluators ---------------- [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#async-evaluators) Use `async def` for I/O-bound operations: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class APIValidator(Evaluator): api_url: str async def evaluate(self, ctx: EvaluatorContext) -> bool: import httpx async with httpx.AsyncClient() as client: response = await client.post( self.api_url, json={'output': ctx.output}, ) return response.json()['valid'] Pydantic Evals handles both sync and async evaluators automatically. Using Metadata -------------- [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#using-metadata) Access case metadata for context-aware evaluation: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class DifficultyAwareScore(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> float: # Base score base_score = self._score_output(ctx.output) # Adjust based on difficulty from metadata if ctx.metadata and 'difficulty' in ctx.metadata: difficulty = ctx.metadata['difficulty'] if difficulty == 'easy': # Penalize mistakes more on easy questions return base_score elif difficulty == 'hard': # Be more lenient on hard questions return min(1.0, base_score * 1.2) return base_score def _score_output(self, output: str) -> float: # Your scoring logic return 0.8 Using Metrics ------------- [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#using-metrics) Access custom metrics set during task execution: from dataclasses import dataclass from pydantic_evals import increment_eval_metric, set_eval_attribute from pydantic_evals.evaluators import Evaluator, EvaluatorContext # In your task def my_task(inputs: str) -> str: result = f'processed: {inputs}' # Record metrics increment_eval_metric('api_calls', 3) set_eval_attribute('used_cache', True) return result # In your evaluator @dataclass class EfficiencyCheck(Evaluator): max_api_calls: int = 5 def evaluate(self, ctx: EvaluatorContext) -> bool: api_calls = ctx.metrics.get('api_calls', 0) return api_calls <= self.max_api_calls See [Metrics & Attributes Guide](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/) for more. Generic Type Parameters ----------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#generic-type-parameters) Make evaluators type-safe with generics: from dataclasses import dataclass from typing import TypeVar from pydantic_evals.evaluators import Evaluator, EvaluatorContext InputsT = TypeVar('InputsT') OutputT = TypeVar('OutputT') @dataclass class TypedEvaluator(Evaluator[InputsT, OutputT, dict]): def evaluate(self, ctx: EvaluatorContext[InputsT, OutputT, dict]) -> bool: # ctx.inputs and ctx.output are now properly typed return True Custom Evaluation Names ----------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#custom-evaluation-names) Control how evaluations appear in reports: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class CustomNameEvaluator(Evaluator): check_type: str def get_default_evaluation_name(self) -> str: # Use check_type as the name instead of class name return f'{self.check_type}_check' def evaluate(self, ctx: EvaluatorContext) -> bool: return True # In reports, appears as "format_check" instead of "CustomNameEvaluator" evaluator = CustomNameEvaluator(check_type='format') Or use the `evaluation_name` field (if using the built-in pattern): from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class MyEvaluator(Evaluator): evaluation_name: str | None = None def evaluate(self, ctx: EvaluatorContext) -> bool: return True # Usage MyEvaluator(evaluation_name='my_custom_name') Real-World Examples ------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#real-world-examples) ### SQL Validation [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#sql-validation) from dataclasses import dataclass from pydantic_evals.evaluators import EvaluationReason, Evaluator, EvaluatorContext @dataclass class ValidSQL(Evaluator): dialect: str = 'postgresql' def evaluate(self, ctx: EvaluatorContext) -> EvaluationReason: try: import sqlparse parsed = sqlparse.parse(ctx.output) if not parsed: return EvaluationReason( value=False, reason='Could not parse SQL', ) # Check for dangerous operations sql_upper = ctx.output.upper() if 'DROP' in sql_upper or 'DELETE' in sql_upper: return EvaluationReason( value=False, reason='Contains dangerous operations (DROP/DELETE)', ) return EvaluationReason( value=True, reason='Valid SQL syntax', ) except Exception as e: return EvaluationReason( value=False, reason=f'SQL parsing error: {e}', ) ### Code Execution [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#code-execution) from dataclasses import dataclass from pydantic_evals.evaluators import EvaluationReason, Evaluator, EvaluatorContext @dataclass class ExecutablePython(Evaluator): timeout_seconds: float = 5.0 async def evaluate(self, ctx: EvaluatorContext) -> EvaluationReason: import asyncio import os import tempfile # Write code to temp file with tempfile.NamedTemporaryFile(mode='w', suffix='.py', delete=False) as f: f.write(ctx.output) temp_path = f.name try: # Execute with timeout process = await asyncio.create_subprocess_exec( 'python', temp_path, stdout=asyncio.subprocess.PIPE, stderr=asyncio.subprocess.PIPE, ) try: stdout, stderr = await asyncio.wait_for( process.communicate(), timeout=self.timeout_seconds, ) except asyncio.TimeoutError: process.kill() return EvaluationReason( value=False, reason=f'Execution timeout after {self.timeout_seconds}s', ) if process.returncode == 0: return EvaluationReason( value=True, reason='Code executed successfully', ) else: return EvaluationReason( value=False, reason=f'Execution failed: {stderr.decode()}', ) finally: os.unlink(temp_path) ### External API Validation [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#external-api-validation) from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class APIResponseValid(Evaluator): api_endpoint: str api_key: str async def evaluate(self, ctx: EvaluatorContext) -> dict[str, bool | float]: import httpx try: async with httpx.AsyncClient() as client: response = await client.post( self.api_endpoint, headers={'Authorization': f'Bearer {self.api_key}'}, json={'data': ctx.output}, timeout=10.0, ) result = response.json() return { 'api_reachable': True, 'validation_passed': result.get('valid', False), 'confidence_score': result.get('confidence', 0.0), } except Exception: return { 'api_reachable': False, 'validation_passed': False, 'confidence_score': 0.0, } Testing Evaluators ------------------ [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#testing-evaluators) Test evaluators like any other Python code: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class ExactMatch(Evaluator): """Check if output exactly matches expected output.""" def evaluate(self, ctx: EvaluatorContext) -> bool: return ctx.output == ctx.expected_output def test_exact_match(): evaluator = ExactMatch() # Test match ctx = EvaluatorContext( name='test', inputs='input', metadata=None, expected_output='expected', output='expected', duration=0.1, _span_tree=None, attributes={}, metrics={}, ) assert evaluator.evaluate(ctx) is True # Test mismatch ctx.output = 'different' assert evaluator.evaluate(ctx) is False Best Practices -------------- [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#best-practices) ### 1\. Keep Evaluators Focused [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#1-keep-evaluators-focused) Each evaluator should check one thing: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext def check_format(output: str) -> bool: return output.startswith('{') def check_content(output: str) -> bool: return len(output) > 10 def check_length(output: str) -> bool: return len(output) < 1000 def check_spelling(output: str) -> bool: return True # Placeholder # Bad: Doing too much @dataclass class EverythingChecker(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> dict: return { 'format_valid': check_format(ctx.output), 'content_good': check_content(ctx.output), 'length_ok': check_length(ctx.output), 'spelling_correct': check_spelling(ctx.output), } # Good: Separate evaluators @dataclass class FormatValidator(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return check_format(ctx.output) @dataclass class ContentChecker(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return check_content(ctx.output) @dataclass class LengthChecker(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return check_length(ctx.output) @dataclass class SpellingChecker(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return check_spelling(ctx.output) Some exceptions to this: * When there is a significant amount of shared computation or network request latency, it may be better to have a single evaluator calculate all dependent outputs together. * If multiple checks are tightly coupled or very closely related to each other, it may make sense to include all their logic in one evaluator. ### 2\. Handle Missing Data Gracefully [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#2-handle-missing-data-gracefully) from dataclasses import dataclass from pydantic_evals.evaluators import EvaluationReason, Evaluator, EvaluatorContext @dataclass class SafeEvaluator(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> EvaluationReason: if ctx.expected_output is None: return EvaluationReason( value=True, reason='Skipped: no expected output provided', ) # Your evaluation logic ... ### 3\. Provide Helpful Reasons [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#3-provide-helpful-reasons) from dataclasses import dataclass from pydantic_evals.evaluators import EvaluationReason, Evaluator, EvaluatorContext @dataclass class HelpfulEvaluator(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> EvaluationReason: # Bad return EvaluationReason(value=False, reason='Failed') # Good return EvaluationReason( value=False, reason=f'Expected {ctx.expected_output!r}, got {ctx.output!r}', ) ### 4\. Use Timeouts for External Calls [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#4-use-timeouts-for-external-calls) from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class APIEvaluator(Evaluator): timeout: float = 10.0 async def _call_api(self, output: str) -> bool: # Placeholder for API call return True async def evaluate(self, ctx: EvaluatorContext) -> bool: import asyncio try: return await asyncio.wait_for( self._call_api(ctx.output), timeout=self.timeout, ) except asyncio.TimeoutError: return False Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/evaluators/custom/#next-steps) * **[Report Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/) ** - Experiment-wide analyses (confusion matrices, PR curves, custom tables) * **[Span-Based Evaluation](https://pydantic.dev/docs/ai/evals/evaluators/span-based/) ** - Using OpenTelemetry spans * **[Examples](https://pydantic.dev/docs/ai/evals/examples/simple-validation/) ** - Practical examples Was this page helpful? Thanks for your feedback! --- # Overview | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/models/overview/#_top) Overview ======== Pydantic AI is model-agnostic and has built-in support for multiple model providers: * [OpenAI](https://pydantic.dev/docs/ai/models/openai/) * [Anthropic](https://pydantic.dev/docs/ai/models/anthropic/) * [Gemini](https://pydantic.dev/docs/ai/models/google/) (via two different APIs: Gemini API and Google Cloud, formerly known as Vertex AI) * [xAI](https://pydantic.dev/docs/ai/models/xai/) * [Bedrock](https://pydantic.dev/docs/ai/models/bedrock/) * [Cerebras](https://pydantic.dev/docs/ai/models/cerebras/) * [Cohere](https://pydantic.dev/docs/ai/models/cohere/) * [Crusoe](https://pydantic.dev/docs/ai/models/crusoe/) * [Groq](https://pydantic.dev/docs/ai/models/groq/) * [Hugging Face](https://pydantic.dev/docs/ai/models/huggingface/) * [Mistral](https://pydantic.dev/docs/ai/models/mistral/) * [OpenRouter](https://pydantic.dev/docs/ai/models/openrouter/) * [Snowflake Cortex](https://pydantic.dev/docs/ai/models/snowflake/) * [Z.AI](https://pydantic.dev/docs/ai/models/zai/) OpenAI-compatible Providers --------------------------- [](https://pydantic.dev/docs/ai/models/overview/#openai-compatible-providers) In addition, many providers are compatible with the OpenAI API, and can be used with `OpenAIChatModel` in Pydantic AI: * [Alibaba Cloud Model Studio (DashScope)](https://pydantic.dev/docs/ai/models/openai/#alibaba-cloud-model-studio-dashscope) * [Azure AI Foundry](https://pydantic.dev/docs/ai/models/openai/#azure-ai-foundry) * [DeepSeek](https://pydantic.dev/docs/ai/models/openai/#deepseek) * [Fireworks AI](https://pydantic.dev/docs/ai/models/openai/#fireworks-ai) * [GitHub Models](https://pydantic.dev/docs/ai/models/openai/#github-models) (retired, deprecated) * [Heroku](https://pydantic.dev/docs/ai/models/openai/#heroku-ai) * [LiteLLM](https://pydantic.dev/docs/ai/models/openai/#litellm) * [Nebius AI Studio](https://pydantic.dev/docs/ai/models/openai/#nebius-ai-studio) * [Ollama](https://pydantic.dev/docs/ai/models/openai/#ollama) * [OVHcloud AI Endpoints](https://pydantic.dev/docs/ai/models/openai/#ovhcloud-ai-endpoints) * [Perplexity](https://pydantic.dev/docs/ai/models/openai/#perplexity) * [SambaNova](https://pydantic.dev/docs/ai/models/openai/#sambanova) * [Together AI](https://pydantic.dev/docs/ai/models/openai/#together-ai) * [Vercel AI Gateway](https://pydantic.dev/docs/ai/models/openai/#vercel-ai-gateway) Pydantic AI also comes with [`TestModel`](https://pydantic.dev/docs/ai/api/models/test/) and [`FunctionModel`](https://pydantic.dev/docs/ai/api/models/function/) for testing and development. To use each model provider, you need to configure your local environment and make sure you have the right packages installed. If you try to use the model without having done so, you’ll be told what to install. Models and Providers -------------------- [](https://pydantic.dev/docs/ai/models/overview/#models-and-providers) Pydantic AI uses a few key terms to describe how it interacts with different LLMs: * **Model**: This refers to the Pydantic AI class used to make requests following a specific LLM API (generally by wrapping a vendor-provided SDK, like the `openai` python SDK). These classes implement a vendor-SDK-agnostic API, ensuring a single Pydantic AI agent is portable to different LLM vendors without any other code changes just by swapping out the Model it uses. Model classes are named roughly in the format `Model`, for example, we have `OpenAIChatModel`, `AnthropicModel`, `GoogleModel`, etc. When using a Model class, you specify the actual LLM model name (e.g., `gpt-5`, `claude-sonnet-4-5`, `gemini-3-flash-preview`) as a parameter. * **Provider**: This refers to provider-specific classes which handle the authentication and connections to an LLM vendor. Passing a non-default _Provider_ as a parameter to a Model is how you can ensure that your agent will make requests to a specific endpoint, or make use of a specific approach to authentication (e.g., you can use Azure auth with the `OpenAIChatModel` by way of the `AzureProvider`). In particular, this is how you can make use of an AI gateway, or an LLM vendor that offers API compatibility with the vendor SDK used by an existing Model (such as `OpenAIChatModel`). * **Profile**: This refers to a description of how requests to a specific model or family of models need to be constructed to get the best results, independent of the model and provider classes used. For example, different models have different restrictions on the JSON schemas that can be used for tools, and the same schema transformer needs to be used for Gemini models whether you’re using `GoogleModel` with model name `gemini-3-pro-preview`, or `OpenAIChatModel` with `OpenRouterProvider` and model name `google/gemini-3-pro-preview`. When you instantiate an [`Agent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent) with just a name formatted as `:`, e.g. `openai:gpt-5.2` or `openrouter:google/gemini-3-pro-preview`, Pydantic AI will automatically select the appropriate model class, provider, and profile. If you want to use a different provider or profile, you can instantiate a model class directly and pass in `provider` and/or `profile` arguments. ### Inspecting a model’s profile [](https://pydantic.dev/docs/ai/models/overview/#inspecting-a-models-profile) A model’s [`ModelProfile`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfile) also describes what the model can do. It is a `TypedDict`, so you read capability flags with normal dictionary access via `model.profile` — for example [`supports_tools`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfile.supports_tools) , [`supports_json_schema_output`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfile.supports_json_schema_output) , and [`supported_native_tools`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfile.supported_native_tools) . This is useful when you want to branch on a capability rather than discover a limitation at request time — for example checking whether a model supports tool calling, native JSON-schema output, or a specific native tool before relying on it: from pydantic_ai.models.test import TestModel from pydantic_ai.native_tools import WebSearchTool model = TestModel() profile = model.profile print(profile['supports_tools']) #> True print(profile['supports_json_schema_output']) #> False print(WebSearchTool in profile['supported_native_tools']) #> True `model.profile` is usually the fully _resolved_ profile: keys from [`DEFAULT_PROFILE`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.DEFAULT_PROFILE) are merged with the provider’s defaults, so direct key access like `profile['supports_tools']` works. If you supply `profile=` as a callable (or otherwise have a partial profile dict), use `profile.get('supports_tools', DEFAULT_PROFILE['supports_tools'])` (after importing `DEFAULT_PROFILE`) to tolerate missing keys. Any [`Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) instance exposes its resolved profile the same way, so the same check works whether the model was selected automatically from a `:` name or instantiated directly. Don’t confuse this with [Capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) , which are reusable bundles of tools, hooks, and settings you add to an agent — the profile describes what the underlying model itself supports. HTTP Client Lifecycle --------------------- [](https://pydantic.dev/docs/ai/models/overview/#http-client-lifecycle) When a [`Provider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.Provider) creates its own HTTP client (i.e. you don’t pass a custom `http_client`), it owns that client’s lifecycle. Using the [`Agent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent) as an async context manager ensures the HTTP client is closed cleanly on exit: from pydantic_ai import Agent agent = Agent('openai:gpt-5.2') async def main(): async with agent: result = await agent.run('What is the capital of France?') print(result.output) #> The capital of France is Paris. You can also use a [`Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) or [`Provider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.Provider) directly as an async context manager for the same effect. If you provide your own `http_client`, you are responsible for closing it yourself. Custom Models ------------- [](https://pydantic.dev/docs/ai/models/overview/#custom-models) To implement support for a model API that’s not already supported, you will need to subclass the [`Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) abstract base class. For streaming, you’ll also need to implement the [`StreamedResponse`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.StreamedResponse) abstract base class. The best place to start is to review the source code for existing implementations, e.g. [`OpenAIChatModel`](https://github.com/pydantic/pydantic-ai/blob/main/pydantic_ai_slim/pydantic_ai/models/openai.py) . For details on when we’ll accept contributions adding new models to Pydantic AI, see the [contributing guidelines](https://pydantic.dev/docs/ai/project/contributing/#new-model-rules) . HTTP Request Concurrency ------------------------ [](https://pydantic.dev/docs/ai/models/overview/#http-request-concurrency) You can limit the number of concurrent HTTP requests to a model using the [`ConcurrencyLimitedModel`](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimitedModel) wrapper. This is useful for respecting rate limits or managing resource usage when running many agents in parallel. model\_concurrency.py import asyncio from pydantic_ai import Agent, ConcurrencyLimitedModel # Wrap a model with concurrency limiting model = ConcurrencyLimitedModel('openai:gpt-4o', limiter=5) # Multiple agents can share this rate-limited model agent = Agent(model) async def main(): # These will be rate-limited to 5 concurrent HTTP requests results = await asyncio.gather( *[agent.run(f'Question {i}') for i in range(20)] ) print(len(results)) #> 20 The `limiter` parameter accepts: * An integer for simple limiting (e.g., `limiter=5`) * A [`ConcurrencyLimit`](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimit) for advanced configuration with backpressure control * A [`ConcurrencyLimiter`](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter) for sharing limits across multiple models ### Shared Concurrency Limits [](https://pydantic.dev/docs/ai/models/overview/#shared-concurrency-limits) To share a concurrency limit across multiple models (e.g., different models from the same provider), you can create a [`ConcurrencyLimiter`](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter) and pass it to multiple `ConcurrencyLimitedModel` instances: shared\_concurrency.py import asyncio from pydantic_ai import Agent, ConcurrencyLimitedModel, ConcurrencyLimiter # Create a shared limiter with a descriptive name shared_limiter = ConcurrencyLimiter(max_running=10, name='openai-pool') # Both models share the same concurrency limit model1 = ConcurrencyLimitedModel('openai:gpt-4o', limiter=shared_limiter) model2 = ConcurrencyLimitedModel('openai:gpt-4o-mini', limiter=shared_limiter) agent1 = Agent(model1) agent2 = Agent(model2) async def main(): # Total concurrent requests across both agents limited to 10 results = await asyncio.gather( *[agent1.run(f'Question {i}') for i in range(10)], *[agent2.run(f'Question {i}') for i in range(10)], ) print(len(results)) #> 20 When instrumentation is enabled, requests waiting for a concurrency slot appear as spans with attributes showing the queue depth and configured limits. The `name` parameter on `ConcurrencyLimiter` helps identify shared limiters in traces. Handling HTTP Errors -------------------- [](https://pydantic.dev/docs/ai/models/overview/#handling-http-errors) When a provider returns a 4xx or 5xx response, Pydantic AI raises a [`ModelHTTPError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelHTTPError) . The exception exposes the [`status_code`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelHTTPError.status_code) , the response [`body`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelHTTPError.body) , and the provider’s **response headers** via the [`headers`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelHTTPError.headers) attribute (a `dict[str, str]` with lowercase keys, or `None` for providers that don’t surface headers, such as gRPC-based providers). The motivating use case is propagating the `Retry-After` header from a 429 response to a caller’s own HTTP client. A convenience property [`retry_after`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelHTTPError.retry_after) parses that header and returns the number of seconds to wait as a `float`, handling both the integer delta-seconds and HTTP-date formats: handle\_rate\_limit.py from pydantic_ai import Agent from pydantic_ai.exceptions import ModelHTTPError agent = Agent('openai:gpt-5.2') try: result = agent.run_sync('What is the capital of France?') except ModelHTTPError as exc: if exc.status_code == 429: wait = exc.retry_after # float | None raise MyRateLimitException( 'AI service is rate-limited. Try again shortly.', retry_after=wait, ) raise Fallback Model -------------- [](https://pydantic.dev/docs/ai/models/overview/#fallback-model) You can use [`FallbackModel`](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackModel) to attempt multiple models in sequence until one succeeds. Pydantic AI can switch to the next model when the current model raises an exception (like a 4xx/5xx API error) **or** when the response content indicates a semantic failure (like a truncated response or a failed native tool call). By default, fallback triggers on [`ModelAPIError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelAPIError) (4xx/5xx API errors), so you don’t need to configure anything for the most common use case. This behavior is controlled by the `fallback_on` parameter (see [`FallbackModel`](https://pydantic.dev/docs/ai/api/models/fallback/#pydantic_ai.models.fallback.FallbackModel) ), which accepts exception types, exception handlers, and response handlers — all of which can be sync or async. In the following example, the agent first makes a request to the OpenAI model (which fails due to an invalid API key), and then falls back to the Anthropic model. fallback\_model.py from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModel from pydantic_ai.models.fallback import FallbackModel from pydantic_ai.models.openai import OpenAIChatModel openai_model = OpenAIChatModel('gpt-5.2') anthropic_model = AnthropicModel('claude-sonnet-4-5') fallback_model = FallbackModel(openai_model, anthropic_model) agent = Agent(fallback_model) response = agent.run_sync('What is the capital of France?') print(response.output) #> The capital of France is Paris. print(response.all_messages()) """ [\ ModelRequest(\ parts=[\ UserPromptPart(\ content='What is the capital of France?',\ timestamp=datetime.datetime(...),\ )\ ],\ timestamp=datetime.datetime(...),\ run_id='...',\ conversation_id='...',\ ),\ ModelResponse(\ parts=[TextPart(content='The capital of France is Paris.')],\ usage=RequestUsage(cost=Decimal('0.000273'), input_tokens=56, output_tokens=7),\ model_name='claude-sonnet-4-5',\ timestamp=datetime.datetime(...),\ run_id='...',\ conversation_id='...',\ ),\ ] """ The `ModelResponse` message above indicates in the `model_name` field that the output was returned by the Anthropic model, which is the second model specified in the `FallbackModel`. ### Per-Model Settings [](https://pydantic.dev/docs/ai/models/overview/#per-model-settings) You can configure different [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) for each model in a fallback chain by passing the `settings` parameter when creating each model. This is particularly useful when different providers have different optimal configurations: fallback\_model\_per\_settings.py from pydantic_ai import Agent, ModelSettings from pydantic_ai.models.anthropic import AnthropicModel from pydantic_ai.models.fallback import FallbackModel from pydantic_ai.models.openai import OpenAIChatModel # Configure each model with provider-specific optimal settings openai_model = OpenAIChatModel( 'gpt-5.2', settings=ModelSettings(temperature=0.7, max_tokens=1000) # Higher creativity for OpenAI ) anthropic_model = AnthropicModel( 'claude-sonnet-4-5', settings=ModelSettings(temperature=0.2, max_tokens=1000) # Lower temperature for consistency ) fallback_model = FallbackModel(openai_model, anthropic_model) agent = Agent(fallback_model) result = agent.run_sync('Write a creative story about space exploration') print(result.output) """ In the year 2157, Captain Maya Chen piloted her spacecraft through the vast expanse of the Andromeda Galaxy. As she discovered a planet with crystalline mountains that sang in harmony with the cosmic winds, she realized that space exploration was not just about finding new worlds, but about finding new ways to understand the universe and our place within it. """ In this example, if the OpenAI model fails, the agent will automatically fall back to the Anthropic model with its own configured settings. The `FallbackModel` itself doesn’t have settings - it uses the individual settings of whichever model successfully handles the request. ### Exception Handling [](https://pydantic.dev/docs/ai/models/overview/#exception-handling) The next example demonstrates the exception-handling capabilities of `FallbackModel`. If all models fail, a [`FallbackExceptionGroup`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.FallbackExceptionGroup) is raised, which contains all the exceptions encountered during the `run` execution. * [Python >=3.11](https://pydantic.dev/docs/ai/models/overview/#tab-panel-128) * [Python <3.11](https://pydantic.dev/docs/ai/models/overview/#tab-panel-129) fallback\_model\_failure.py from pydantic_ai import Agent, ModelAPIError from pydantic_ai.models.anthropic import AnthropicModel from pydantic_ai.models.fallback import FallbackModel from pydantic_ai.models.openai import OpenAIChatModel openai_model = OpenAIChatModel('gpt-5.2') anthropic_model = AnthropicModel('claude-sonnet-4-5') fallback_model = FallbackModel(openai_model, anthropic_model) agent = Agent(fallback_model) try: response = agent.run_sync('What is the capital of France?') except* ModelAPIError as exc_group: for exc in exc_group.exceptions: print(exc) Since [`except*`](https://docs.python.org/3/reference/compound_stmts.html#except-star) is only supported in Python 3.11+, we use the [`exceptiongroup`](https://github.com/agronholm/exceptiongroup) backport package for earlier Python versions: fallback\_model\_failure.py from exceptiongroup import catch from pydantic_ai import Agent, ModelAPIError from pydantic_ai.models.anthropic import AnthropicModel from pydantic_ai.models.fallback import FallbackModel from pydantic_ai.models.openai import OpenAIChatModel def model_status_error_handler(exc_group: BaseExceptionGroup) -> None: for exc in exc_group.exceptions: print(exc) openai_model = OpenAIChatModel('gpt-5.2') anthropic_model = AnthropicModel('claude-sonnet-4-5') fallback_model = FallbackModel(openai_model, anthropic_model) agent = Agent(fallback_model) with catch({ModelAPIError: model_status_error_handler}): response = agent.run_sync('What is the capital of France?') By default, the `FallbackModel` only moves on to the next model if the current model raises a [`ModelAPIError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelAPIError) , which includes [`ModelHTTPError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelHTTPError) . You can customize this behavior by passing a custom `fallback_on` argument to the `FallbackModel` constructor. ### Response-Based Fallback [](https://pydantic.dev/docs/ai/models/overview/#response-based-fallback) In addition to exception-based fallback, you can also trigger fallback based on the **content** of a model’s response. This is useful when a model returns a successful HTTP response (no exception), but the response content indicates a semantic failure — for example, an unexpected finish reason or a native tool reporting failure. The `fallback_on` parameter accepts: * A tuple of exception types: `(ModelAPIError, ModelHTTPError)` * An exception handler (sync or async): `lambda exc: isinstance(exc, MyError)` * A response handler (sync or async): `def check(r: ModelResponse) -> bool` * A list mixing all of the above: `[ModelAPIError, exc_handler, response_handler]` Handler type is auto-detected by inspecting type hints on the first parameter. If the first parameter is hinted as [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) , it’s a response handler. Otherwise (including untyped handlers and lambdas), it’s an exception handler. As the hints are resolved at runtime, every annotated type in the handler signature must be imported at runtime rather than only under `if TYPE_CHECKING:`. If any annotation can’t be resolved, a [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) is raised instead of the handler being silently treated as an exception handler. #### Finish Reason Example [](https://pydantic.dev/docs/ai/models/overview/#finish-reason-example) A simple use case is checking the model’s finish reason — for example, falling back if the response was truncated due to length limits: fallback\_on\_finish\_reason.py from pydantic_ai import Agent from pydantic_ai.messages import FinishReason, ModelResponse from pydantic_ai.models.fallback import FallbackModel def bad_finish_reason(response: ModelResponse) -> bool: """Fallback if the model stopped due to length limit, content filter, or error.""" reason: FinishReason | None = response.finish_reason # Trigger fallback for problematic finish reasons return reason in ('length', 'content_filter', 'error') fallback_model = FallbackModel( 'openai:gpt-5.2', 'anthropic:claude-sonnet-4-5', fallback_on=bad_finish_reason, ) agent = Agent(fallback_model) result = agent.run_sync('What is the capital of France?') print(result.output) #> The capital of France is Paris. #### Native Tool Failure Example [](https://pydantic.dev/docs/ai/models/overview/#native-tool-failure-example) A more complex use case is when using native tools like web search or URL fetching. For example, Google’s [`WebFetchTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebFetchTool) may return a successful response with a status indicating the URL fetch failed: fallback\_on\_native\_tool.py from pydantic_ai import Agent from pydantic_ai.messages import ModelResponse from pydantic_ai.models.anthropic import AnthropicModel from pydantic_ai.models.fallback import FallbackModel from pydantic_ai.models.google import GoogleModel def web_fetch_failed(response: ModelResponse) -> bool: """Check if a web_fetch native tool failed to retrieve content.""" for call, result in response.native_tool_calls: if call.tool_name != 'web_fetch': continue if not isinstance(result.content, list): continue for item in result.content: if isinstance(item, dict): status = item.get('url_retrieval_status', '') if status and status != 'URL_RETRIEVAL_STATUS_SUCCESS': return True return False google_model = GoogleModel('gemini-2.5-flash') anthropic_model = AnthropicModel('claude-sonnet-4-5') # Auto-detected as response handler via type hint fallback_model = FallbackModel( google_model, anthropic_model, fallback_on=web_fetch_failed, ) agent = Agent(fallback_model) # If Google's web_fetch fails, automatically falls back to Anthropic result = agent.run_sync('Summarize https://ai.pydantic.dev') print(result.output) """ Pydantic AI is a Python agent framework for building production-grade LLM applications. """ Response handlers receive the [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) returned by the model and should return `True` to trigger fallback to the next model, or `False` to accept the response. #### Combining Handlers [](https://pydantic.dev/docs/ai/models/overview/#combining-handlers) You can combine exception types, exception handlers, and response handlers in a single list: fallback\_on\_mixed.py from pydantic_ai.exceptions import ModelAPIError from pydantic_ai.models.fallback import FallbackModel from fallback_on_native_tool import anthropic_model, google_model, web_fetch_failed fallback_model = FallbackModel( google_model, anthropic_model, fallback_on=[\ ModelAPIError, # Exception type\ lambda exc: 'rate limit' in str(exc).lower(), # Exception handler (untyped lambda)\ web_fetch_failed, # Response handler (auto-detected via type hint)\ ], ) ### Exception Handling in Middleware and Decorators [](https://pydantic.dev/docs/ai/models/overview/#exception-handling-in-middleware-and-decorators) When using `FallbackModel`, it’s important to understand that [`FallbackExceptionGroup`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.FallbackExceptionGroup) inherits from Python’s [`ExceptionGroup`](https://docs.python.org/3/library/exceptions.html#ExceptionGroup) . This means that existing exception handling code that catches specific exceptions (like `ModelAPIError`) won’t automatically catch the individual exceptions wrapped inside the group. For example, if you have middleware or a decorator that catches `ModelAPIError`: middleware\_without\_fallback.py from collections.abc import Callable from functools import wraps from typing import TypeVar from pydantic_ai import ModelAPIError T = TypeVar('T') # This handler will NOT catch ModelAPIError when using FallbackModel! def handle_api_errors(func: Callable[..., T]) -> Callable[..., T]: @wraps(func) def wrapper(*args, **kwargs) -> T: try: return func(*args, **kwargs) except ModelAPIError as e: # Won't catch FallbackExceptionGroup print(f'API error: {e}') raise return wrapper This decorator will miss `ModelAPIError` exceptions when using `FallbackModel`, because they’re wrapped in a `FallbackExceptionGroup` containing one exception per failed model, in the order the models were tried. To handle both cases, you can use Python 3.11+ `except*` syntax, which catches matching exceptions from exception groups as well as bare exceptions. Note that `except*` always delivers the caught exceptions as an `ExceptionGroup` (even if the original was a bare exception), so re-raising will propagate an `ExceptionGroup` rather than the original exception type: * [Python >=3.11](https://pydantic.dev/docs/ai/models/overview/#tab-panel-130) * [Python <3.11](https://pydantic.dev/docs/ai/models/overview/#tab-panel-131) middleware\_with\_fallback.py from collections.abc import Callable from functools import wraps from typing import TypeVar from pydantic_ai import ModelAPIError T = TypeVar('T') def handle_api_errors(func: Callable[..., T]) -> Callable[..., T]: @wraps(func) def wrapper(*args, **kwargs) -> T: try: return func(*args, **kwargs) except* ModelAPIError as exc_group: for exc in exc_group.exceptions: print(f'API error: {exc}') raise return wrapper middleware\_with\_fallback.py from collections.abc import Callable from functools import wraps from typing import TypeVar from pydantic_ai import FallbackExceptionGroup, ModelAPIError T = TypeVar('T') def handle_api_errors(func: Callable[..., T]) -> Callable[..., T]: @wraps(func) def wrapper(*args, **kwargs) -> T: try: return func(*args, **kwargs) except FallbackExceptionGroup as exc_group: for exc in exc_group.exceptions: if isinstance(exc, ModelAPIError): print(f'API error from fallback: {exc}') raise except ModelAPIError as e: print(f'API error: {e}') raise return wrapper You can also catch `FallbackExceptionGroup` directly if you want to handle it specifically: catch\_fallback\_exception\_group.py from pydantic_ai import Agent, FallbackExceptionGroup from pydantic_ai.models.fallback import FallbackModel agent = Agent(FallbackModel('openai:gpt-5-mini', 'anthropic:claude-sonnet-4-6')) try: response = agent.run_sync('What is the capital of France?') except FallbackExceptionGroup as exc_group: print(f'All {len(exc_group.exceptions)} models failed:') for exc in exc_group.exceptions: print(f' - {exc}') Was this page helpful? Thanks for your feedback! --- # Report Evaluators | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#_top) Report Evaluators ================= Report evaluators analyze entire experiment results rather than individual cases. Use them to compute experiment-wide statistics like confusion matrices, precision-recall curves, accuracy scores, or custom summary tables. How Report Evaluators Work -------------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#how-report-evaluators-work) Regular [evaluators](https://pydantic.dev/docs/ai/evals/evaluators/overview/) run once per case and assess individual outputs. Report evaluators run once per experiment _after_ all cases have been evaluated, receiving the full [`EvaluationReport`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.EvaluationReport) as input. Cases executed → Case evaluators run → Report evaluators run → Final report Results from report evaluators are stored as **analyses** on the report and, when Logfire is configured, are attached to the experiment span as structured attributes for visualization. Using Report Evaluators ----------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#using-report-evaluators) Pass report evaluators to `Dataset` via the `report_evaluators` parameter: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import ConfusionMatrixEvaluator def my_classifier(text: str) -> str: text = text.lower() if 'cat' in text or 'meow' in text: return 'cat' elif 'dog' in text or 'bark' in text: return 'dog' return 'unknown' dataset = Dataset( name='animal_classifier', cases=[\ Case(name='cat', inputs='The cat goes meow', expected_output='cat'),\ Case(name='dog', inputs='The dog barks', expected_output='dog'),\ ], report_evaluators=[\ ConfusionMatrixEvaluator(\ predicted_from='output',\ expected_from='expected_output',\ title='Animal Classification',\ ),\ ], ) report = dataset.evaluate_sync(my_classifier) # report.analyses contains the ConfusionMatrix result Native Report Evaluators ------------------------ [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#native-report-evaluators) ### ConfusionMatrixEvaluator [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#confusionmatrixevaluator) Builds a confusion matrix comparing predicted vs expected labels across all cases. from pydantic_evals.evaluators import ConfusionMatrixEvaluator ConfusionMatrixEvaluator( predicted_from='output', expected_from='expected_output', title='My Confusion Matrix', ) **Parameters:** | Parameter | Type | Default | Description | | --- | --- | --- | --- | | `predicted_from` | `'expected_output' \| 'output' \| 'metadata' \| 'labels'` | `'output'` | Source for predicted values | | `predicted_key` | `str \| None` | `None` | Key to extract when using `metadata` or `labels` | | `expected_from` | `'expected_output' \| 'output' \| 'metadata' \| 'labels'` | `'expected_output'` | Source for expected/true values | | `expected_key` | `str \| None` | `None` | Key to extract when using `metadata` or `labels` | | `title` | `str` | `'Confusion Matrix'` | Title shown in reports | **Returns:** [`ConfusionMatrix`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.ConfusionMatrix) **Data Sources:** * `'output'` — the task’s actual output (converted to string) * `'expected_output'` — the case’s expected output (converted to string) * `'metadata'` — a value from the case’s metadata dict (requires `key`) * `'labels'` — a label result from a case-level evaluator (requires `key`) **Example — classification with expected outputs:** from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import ConfusionMatrixEvaluator dataset = Dataset( name='animal_sounds', cases=[\ Case(inputs='meow', expected_output='cat'),\ Case(inputs='woof', expected_output='dog'),\ Case(inputs='chirp', expected_output='bird'),\ ], report_evaluators=[\ ConfusionMatrixEvaluator(\ predicted_from='output',\ expected_from='expected_output',\ ),\ ], ) **Example — using evaluator labels:** If a case-level evaluator produces a label like `predicted_class`, you can reference it: from dataclasses import dataclass from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import ( ConfusionMatrixEvaluator, Evaluator, EvaluatorContext, ) @dataclass class ClassifyOutput(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> dict[str, str]: # Classify the output into a category return {'predicted_class': categorize(ctx.output)} def categorize(output: str) -> str: return 'positive' if 'good' in output.lower() else 'negative' dataset = Dataset( name='labels_example', cases=[Case(inputs='test', expected_output='positive')], evaluators=[ClassifyOutput()], report_evaluators=[\ ConfusionMatrixEvaluator(\ predicted_from='labels',\ predicted_key='predicted_class',\ expected_from='expected_output',\ ),\ ], ) * * * ### PrecisionRecallEvaluator [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#precisionrecallevaluator) Computes a precision-recall curve with AUC (area under the curve) from numeric scores and binary ground-truth labels. from pydantic_evals.evaluators import PrecisionRecallEvaluator PrecisionRecallEvaluator( score_from='scores', score_key='confidence', positive_from='assertions', positive_key='is_correct', ) **Parameters:** | Parameter | Type | Default | Description | | --- | --- | --- | --- | | `score_key` | `str` | _(required)_ | Key in scores or metrics dict | | `positive_from` | `'expected_output' \| 'assertions' \| 'labels'` | _(required)_ | Source for ground-truth binary labels | | `positive_key` | `str \| None` | `None` | Key in assertions or labels dict | | `score_from` | `'scores' \| 'metrics'` | `'scores'` | Source for numeric scores | | `title` | `str` | `'Precision-Recall Curve'` | Title shown in reports | | `n_thresholds` | `int` | `100` | Number of threshold points on the curve | **Returns:** [`PrecisionRecall`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.PrecisionRecall) + [`ScalarResult`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.ScalarResult) (AUC) The AUC is computed at full resolution (using every unique score as a threshold) for accuracy, then the curve points are downsampled to `n_thresholds` for display. The AUC is returned both on the curve (for chart rendering) and as a separate `ScalarResult` for querying and sorting. **Score Sources:** * `'scores'` — a numeric score from a case-level evaluator (looked up by `score_key`) * `'metrics'` — a custom metric set during task execution (looked up by `score_key`) **Positive Sources:** * `'assertions'` — a boolean assertion from a case-level evaluator (looked up by `positive_key`) * `'labels'` — a label result cast to boolean (looked up by `positive_key`) * `'expected_output'` — the case’s expected output cast to boolean **Example:** from dataclasses import dataclass from typing import Any from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import ( Evaluator, EvaluatorContext, PrecisionRecallEvaluator, ) @dataclass class ConfidenceEvaluator(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> dict[str, Any]: confidence = calculate_confidence(ctx.output) return { 'confidence': confidence, # numeric score 'is_correct': ctx.output == ctx.expected_output, # boolean assertion } def calculate_confidence(output: str) -> float: return 0.85 # placeholder dataset = Dataset( name='precision_recall_example', cases=[\ Case(inputs='test 1', expected_output='cat'),\ Case(inputs='test 2', expected_output='dog'),\ ], evaluators=[ConfidenceEvaluator()], report_evaluators=[\ PrecisionRecallEvaluator(\ score_from='scores',\ score_key='confidence',\ positive_from='assertions',\ positive_key='is_correct',\ ),\ ], ) * * * ### ROCAUCEvaluator [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#rocaucevaluator) Computes an ROC (Receiver Operating Characteristic) curve and AUC from numeric scores and binary ground-truth labels. The ROC curve plots the True Positive Rate against the False Positive Rate at various threshold values, with a dashed random-baseline diagonal for reference. from pydantic_evals.evaluators import ROCAUCEvaluator ROCAUCEvaluator( score_key='confidence', positive_from='assertions', positive_key='is_correct', ) **Parameters:** | Parameter | Type | Default | Description | | --- | --- | --- | --- | | `score_key` | `str` | _(required)_ | Key in scores or metrics dict | | `positive_from` | `'expected_output' \| 'assertions' \| 'labels'` | _(required)_ | Source for ground-truth binary labels | | `positive_key` | `str \| None` | `None` | Key in assertions or labels dict | | `score_from` | `'scores' \| 'metrics'` | `'scores'` | Source for numeric scores | | `title` | `str` | `'ROC Curve'` | Title shown in reports | | `n_thresholds` | `int` | `100` | Number of threshold points on the curve | **Returns:** [`LinePlot`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.LinePlot) + [`ScalarResult`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.ScalarResult) (AUC) The AUC is computed at full resolution. The chart includes a dashed “Random” baseline diagonal from (0, 0) to (1, 1) for visual comparison. **Score and Positive Sources:** Same as [`PrecisionRecallEvaluator`](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#precisionrecallevaluator) . * * * ### KolmogorovSmirnovEvaluator [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#kolmogorovsmirnovevaluator) Computes a Kolmogorov-Smirnov plot and KS statistic from numeric scores and binary ground-truth labels. The KS plot shows the empirical CDFs (cumulative distribution functions) of the score distribution for positive and negative cases. The KS statistic is the maximum vertical distance between the two CDFs — higher values indicate better class separation. from pydantic_evals.evaluators import KolmogorovSmirnovEvaluator KolmogorovSmirnovEvaluator( score_key='confidence', positive_from='assertions', positive_key='is_correct', ) **Parameters:** | Parameter | Type | Default | Description | | --- | --- | --- | --- | | `score_key` | `str` | _(required)_ | Key in scores or metrics dict | | `positive_from` | `'expected_output' \| 'assertions' \| 'labels'` | _(required)_ | Source for ground-truth binary labels | | `positive_key` | `str \| None` | `None` | Key in assertions or labels dict | | `score_from` | `'scores' \| 'metrics'` | `'scores'` | Source for numeric scores | | `title` | `str` | `'KS Plot'` | Title shown in reports | | `n_thresholds` | `int` | `100` | Number of threshold points on the curve | **Returns:** [`LinePlot`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.LinePlot) + [`ScalarResult`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.ScalarResult) (KS Statistic) **Score and Positive Sources:** Same as [`PrecisionRecallEvaluator`](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#precisionrecallevaluator) . * * * Custom Report Evaluators ------------------------ [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#custom-report-evaluators) Write custom report evaluators by inheriting from [`ReportEvaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.ReportEvaluator) and implementing the `evaluate` method: from dataclasses import dataclass from pydantic_evals.evaluators import ReportEvaluator, ReportEvaluatorContext from pydantic_evals.reporting.analyses import ScalarResult @dataclass class AccuracyEvaluator(ReportEvaluator): """Computes overall accuracy as a scalar metric.""" def evaluate(self, ctx: ReportEvaluatorContext) -> ScalarResult: cases = ctx.report.cases if not cases: return ScalarResult(title='Accuracy', value=0.0, unit='%') correct = sum( 1 for case in cases if case.output == case.expected_output ) accuracy = correct / len(cases) * 100 return ScalarResult(title='Accuracy', value=accuracy, unit='%') ### ReportEvaluatorContext [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#reportevaluatorcontext) The context passed to `evaluate()` contains: * `ctx.name` — the experiment name * `ctx.report` — the full [`EvaluationReport`](https://pydantic.dev/docs/ai/api/pydantic_evals/reporting/#pydantic_evals.reporting.EvaluationReport) with all case results * `ctx.experiment_metadata` — optional experiment-level metadata dict Through `ctx.report.cases`, you can access each case’s inputs, outputs, expected outputs, scores, labels, assertions, metrics, and attributes. ### Return Types [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#return-types) Report evaluators must return a `ReportAnalysis` or a `list[ReportAnalysis]`. The available analysis types are: #### ScalarResult [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#scalarresult) A single numeric statistic: from pydantic_evals.reporting.analyses import ScalarResult ScalarResult( title='Accuracy', value=93.3, unit='%', description='Percentage of correctly classified cases.', ) | Field | Type | Description | | --- | --- | --- | | `title` | `str` | Display name | | `value` | `float \| int` | The numeric value | | `unit` | `str \| None` | Optional unit label (e.g., `'%'`, `'ms'`) | | `description` | `str \| None` | Optional longer description | * * * #### TableResult [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#tableresult) A generic table of data: from pydantic_evals.reporting.analyses import TableResult TableResult( title='Per-Class Metrics', columns=['Class', 'Precision', 'Recall', 'F1'], rows=[\ ['cat', 0.95, 0.90, 0.924],\ ['dog', 0.88, 0.92, 0.899],\ ], description='Precision, recall, and F1 per class.', ) | Field | Type | Description | | --- | --- | --- | | `title` | `str` | Display name | | `columns` | `list[str]` | Column headers | | `rows` | `list[list[str \| int \| float \| bool \| None]]` | Row data | | `description` | `str \| None` | Optional longer description | * * * #### ConfusionMatrix [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#confusionmatrix) A confusion matrix (typically produced by `ConfusionMatrixEvaluator`, but can be constructed directly): from pydantic_evals.reporting.analyses import ConfusionMatrix ConfusionMatrix( title='Sentiment', class_labels=['positive', 'negative', 'neutral'], matrix=[\ [45, 3, 2], # expected=positive\ [5, 40, 5], # expected=negative\ [1, 2, 47], # expected=neutral\ ], ) | Field | Type | Description | | --- | --- | --- | | `title` | `str` | Display name | | `class_labels` | `list[str]` | Ordered labels for both axes | | `matrix` | `list[list[int]]` | `matrixexpected` = count | | `description` | `str \| None` | Optional longer description | * * * #### PrecisionRecall [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#precisionrecall) Precision-recall curve data (typically produced by `PrecisionRecallEvaluator`): | Field | Type | Description | | --- | --- | --- | | `title` | `str` | Display name | | `curves` | `list[PrecisionRecallCurve]` | One or more curves | | `description` | `str \| None` | Optional longer description | Each `PrecisionRecallCurve` contains a `name`, a list of `PrecisionRecallPoint`s (with `threshold`, `precision`, `recall`), and an optional `auc` value. * * * #### LinePlot [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#lineplot) A generic XY line chart with labeled axes, supporting multiple curves. Use this for ROC curves, KS plots, calibration curves, or any custom line chart: from pydantic_evals.reporting.analyses import LinePlot, LinePlotCurve, LinePlotPoint LinePlot( title='ROC Curve', x_label='False Positive Rate', y_label='True Positive Rate', x_range=(0, 1), y_range=(0, 1), curves=[\ LinePlotCurve(\ name='Model (AUC: 0.95)',\ points=[LinePlotPoint(x=0.0, y=0.0), LinePlotPoint(x=0.1, y=0.8), LinePlotPoint(x=1.0, y=1.0)],\ ),\ LinePlotCurve(\ name='Random',\ points=[LinePlotPoint(x=0, y=0), LinePlotPoint(x=1, y=1)],\ style='dashed',\ ),\ ], ) | Field | Type | Description | | --- | --- | --- | | `title` | `str` | Display name | | `x_label` | `str` | Label for the x-axis | | `y_label` | `str` | Label for the y-axis | | `x_range` | `tuple[float, float] \| None` | Optional fixed range for x-axis | | `y_range` | `tuple[float, float] \| None` | Optional fixed range for y-axis | | `curves` | `list[LinePlotCurve]` | One or more curves to plot | | `description` | `str \| None` | Optional longer description | Each `LinePlotCurve` contains a `name`, a list of `LinePlotPoint`s (with `x`, `y`), an optional `style` (`'solid'` or `'dashed'`), and an optional `step` interpolation mode (`'start'`, `'middle'`, or `'end'`) for step functions like empirical CDFs. `LinePlot` is the recommended return type for custom curve-based evaluators — any evaluator that returns a `LinePlot` will be rendered as a line chart in the Logfire UI without requiring any frontend changes. ### Returning Multiple Analyses [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#returning-multiple-analyses) A single report evaluator can return multiple analyses by returning a list: from dataclasses import dataclass from pydantic_evals.evaluators import ReportEvaluator, ReportEvaluatorContext from pydantic_evals.reporting.analyses import ReportAnalysis, ScalarResult, TableResult @dataclass class ClassificationSummary(ReportEvaluator): """Produces both a scalar accuracy and a per-class metrics table.""" def evaluate(self, ctx: ReportEvaluatorContext) -> list[ReportAnalysis]: cases = ctx.report.cases if not cases: return [] labels = sorted({str(c.expected_output) for c in cases if c.expected_output}) # Scalar: overall accuracy correct = sum(1 for c in cases if c.output == c.expected_output) accuracy = ScalarResult( title='Accuracy', value=correct / len(cases) * 100, unit='%' ) # Table: per-class breakdown rows = [] for label in labels: tp = sum(1 for c in cases if str(c.output) == label and str(c.expected_output) == label) fp = sum(1 for c in cases if str(c.output) == label and str(c.expected_output) != label) fn = sum(1 for c in cases if str(c.output) != label and str(c.expected_output) == label) p = tp / (tp + fp) if (tp + fp) > 0 else 0.0 r = tp / (tp + fn) if (tp + fn) > 0 else 0.0 f1 = 2 * p * r / (p + r) if (p + r) > 0 else 0.0 rows.append([label, round(p, 3), round(r, 3), round(f1, 3)]) table = TableResult( title='Per-Class Metrics', columns=['Class', 'Precision', 'Recall', 'F1'], rows=rows, ) return [accuracy, table] ### Async Report Evaluators [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#async-report-evaluators) Report evaluators support async `evaluate` methods, handled automatically via `evaluate_async`: from dataclasses import dataclass from pydantic_evals.evaluators import ReportEvaluator, ReportEvaluatorContext from pydantic_evals.reporting.analyses import ScalarResult @dataclass class AsyncAccuracy(ReportEvaluator): async def evaluate(self, ctx: ReportEvaluatorContext) -> ScalarResult: # Can use async I/O here (e.g., call an external API) cases = ctx.report.cases correct = sum(1 for c in cases if c.output == c.expected_output) return ScalarResult( title='Accuracy', value=correct / len(cases) * 100 if cases else 0.0, unit='%', ) Serialization ------------- [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#serialization) Report evaluators are serialized to and from YAML/JSON dataset files using the same format as case-level evaluators. This means datasets with report evaluators can be fully round-tripped through file serialization. **Example YAML dataset with report evaluators:** # yaml-language-server: $schema=./test_cases_schema.json name: classifier_eval cases: - name: cat_test inputs: The cat meows expected_output: cat - name: dog_test inputs: The dog barks expected_output: dog report_evaluators: - ConfusionMatrixEvaluator - PrecisionRecallEvaluator: score_key: confidence positive_from: assertions positive_key: is_correct Native report evaluators (`ConfusionMatrixEvaluator`, `PrecisionRecallEvaluator`, `ROCAUCEvaluator`, `KolmogorovSmirnovEvaluator`) are recognized automatically. For custom report evaluators, pass them via `custom_report_evaluator_types`: from pydantic_evals import Dataset dataset = Dataset[str, str, None].from_file( 'test_cases.yaml', custom_report_evaluator_types=[MyCustomReportEvaluator], ) Similarly, when saving a dataset with custom report evaluators, pass them to `to_file` so the JSON schema includes them: dataset.to_file( 'test_cases.yaml', custom_report_evaluator_types=[MyCustomReportEvaluator], ) Viewing Analyses in Logfire --------------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#viewing-analyses-in-logfire) When [Logfire is configured](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/) , analyses are automatically attached to the experiment span as the `logfire.experiment.analyses` attribute. The Logfire UI renders them as interactive visualizations: * **Confusion matrices** are displayed as heatmaps * **Precision-recall curves** are rendered as line charts with AUC in the legend * **Line plots** (ROC curves, KS plots, etc.) are rendered as line charts with configurable axes * **Scalar results** are shown as labeled values * **Tables** are rendered as formatted data tables When comparing multiple experiments in the Logfire Evals view, analyses of the same type are displayed side by side for easy comparison. Complete Example ---------------- [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#complete-example) A full example combining case-level evaluators with report evaluators: from dataclasses import dataclass from typing import Any from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import ( ConfusionMatrixEvaluator, Evaluator, EvaluatorContext, KolmogorovSmirnovEvaluator, PrecisionRecallEvaluator, ReportEvaluator, ReportEvaluatorContext, ROCAUCEvaluator, ) from pydantic_evals.reporting.analyses import ScalarResult def my_classifier(text: str) -> str: text = text.lower() if 'cat' in text or 'meow' in text: return 'cat' elif 'dog' in text or 'bark' in text: return 'dog' elif 'bird' in text or 'chirp' in text: return 'bird' return 'unknown' # Case-level evaluator: runs per case @dataclass class ConfidenceEvaluator(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> dict[str, Any]: confidence = compute_confidence(ctx.output, ctx.inputs) is_correct = ctx.output == ctx.expected_output return { 'confidence': confidence, 'is_correct': is_correct, } def compute_confidence(output: str, inputs: str) -> float: return 0.85 # placeholder # Report-level evaluator: runs once over the full report @dataclass class AccuracyEvaluator(ReportEvaluator): def evaluate(self, ctx: ReportEvaluatorContext) -> ScalarResult: cases = ctx.report.cases correct = sum(1 for c in cases if c.output == c.expected_output) return ScalarResult( title='Accuracy', value=correct / len(cases) * 100 if cases else 0.0, unit='%', ) dataset = Dataset( name='full_example', cases=[\ Case(inputs='The cat meows', expected_output='cat'),\ Case(inputs='The dog barks', expected_output='dog'),\ Case(inputs='A bird chirps', expected_output='bird'),\ ], evaluators=[ConfidenceEvaluator()], report_evaluators=[\ ConfusionMatrixEvaluator(\ predicted_from='output',\ expected_from='expected_output',\ title='Animal Classification',\ ),\ PrecisionRecallEvaluator(\ score_from='scores',\ score_key='confidence',\ positive_from='assertions',\ positive_key='is_correct',\ ),\ ROCAUCEvaluator(\ score_from='scores',\ score_key='confidence',\ positive_from='assertions',\ positive_key='is_correct',\ ),\ KolmogorovSmirnovEvaluator(\ score_from='scores',\ score_key='confidence',\ positive_from='assertions',\ positive_key='is_correct',\ ),\ AccuracyEvaluator(),\ ], ) report = dataset.evaluate_sync(my_classifier) # Access analyses programmatically for analysis in report.analyses: print(f'{analysis.type}: {analysis.title}') #> confusion_matrix: Animal Classification #> precision_recall: Precision-Recall Curve #> scalar: Precision-Recall Curve AUC #> line_plot: ROC Curve #> scalar: ROC Curve AUC #> line_plot: KS Plot #> scalar: KS Statistic #> scalar: Accuracy Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/evaluators/report-evaluators/#next-steps) * **[Native Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/built-in/) ** — Case-level evaluator reference * **[Custom Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/custom/) ** — Writing case-level evaluators * **[Logfire Integration](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/) ** — Viewing analyses in the Logfire UI Was this page helpful? Thanks for your feedback! --- # Capabilities and hooks | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/realtime/capabilities/#_top) Capabilities and hooks ====================== A [capability](https://pydantic.dev/docs/ai/capabilities/overview/) attached to the agent or passed to `realtime(capabilities=...)` participates in a realtime session where its lifecycle maps onto a persistent connection. [Third-party capabilities](https://pydantic.dev/docs/ai/capabilities/overview/#third-party-capabilities) load exactly the same way as in a regular run; nothing realtime-specific is required of them. Capability stages in a session ------------------------------ [](https://pydantic.dev/docs/ai/realtime/capabilities/#capability-stages-in-a-session) | Capability stage | Session behavior | | --- | --- | | `for_agent`, `for_run`, `get_instructions` | Runs during setup; dynamic instructions are evaluated once at connect. | | `get_toolset`, `get_wrapper_toolset`, `prepare_tools` | Contributes, wraps, and prepares local tools before connecting. | | `get_native_tools` | Contributes native tools before connecting; a dynamic native-tool function is resolved once against the connect-time context, like dynamic instructions. | | Tool validation/execution hooks | Runs around each local function-tool call. | | `handle_deferred_tool_calls` | Resolves deferred requests inline; see [deferred and approval-required tools](https://pydantic.dev/docs/ai/realtime/tools/#deferred-and-approval-required-tools)
. | | Graph node, model-request, and output-processing hooks | Do not run; no agent graph or output-processing stage exists. | All regular [tool validation](https://pydantic.dev/docs/ai/core-concepts/hooks/#tool-validation-hooks) and [tool execution](https://pydantic.dev/docs/ai/core-concepts/hooks/#tool-execution-hooks) hooks — `before`, `after`, `wrap`, and `on_error` for both stages — run around every local function-tool call exactly as in a standard run, retries and all. What does not run is anything tied to the request-response graph: [node hooks](https://pydantic.dev/docs/ai/core-concepts/hooks/#node-hooks) , [model request hooks](https://pydantic.dev/docs/ai/core-concepts/hooks/#model-request-hooks) such as `before_model_request`, and [output validation](https://pydantic.dev/docs/ai/core-concepts/hooks/#output-validation-hooks) and [output processing](https://pydantic.dev/docs/ai/core-concepts/hooks/#output-processing-hooks) hooks — a session has no graph nodes, no per-request boundary, and no output stage. Run hooks --------- [](https://pydantic.dev/docs/ai/realtime/capabilities/#run-hooks) `before_run`, `after_run`, `wrap_run`, and `on_run_error` [run hooks](https://pydantic.dev/docs/ai/core-concepts/hooks/#run-hooks) run once around the session — a realtime session is a run — with the same close-boundary recovery and result-transformation semantics as [`iter()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.iter) . The event stream ---------------- [](https://pydantic.dev/docs/ai/realtime/capabilities/#the-event-stream) `wrap_run_event_stream` wraps the consumer-facing session iterator. It can observe or transform shared [`AgentStreamEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.AgentStreamEvent) members and realtime-only [`RealtimeEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeEvent) members (see the [event reference](https://pydantic.dev/docs/ai/realtime/events/) ) without changing history or tool execution. There is no `event_stream_handler` parameter on `realtime()`; a handler-style consumer is attached with the [`ProcessEventStream`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ProcessEventStream) capability, which works through this same stream. Model settings and `RunContext` ------------------------------- [](https://pydantic.dev/docs/ai/realtime/capabilities/#model-settings-and-runcontext) `get_model_settings()` may run during capability setup, but regular model settings do not configure a realtime model. Pass [`RealtimeModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelSettings) through `realtime(model_settings=...)` instead. Inside session hooks and tools, the [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) reflects the session: | `RunContext` field | Value in a realtime session | | --- | --- | | [`ctx.model_settings`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.model_settings) | The merged [`RealtimeModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelSettings)
the session was connected with. | | [`ctx.realtime`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.realtime) | `True` from `before_run` onward. | | [`ctx.realtime_session`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.realtime_session) | The live [`RealtimeSession`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession)
once it is connected. | Seeded history is not processed ------------------------------- [](https://pydantic.dev/docs/ai/realtime/capabilities/#seeded-history-is-not-processed) History-processing capabilities do not transform `message_history` before it is [seeded into a session](https://pydantic.dev/docs/ai/realtime/history/#seeding-a-session) ; preprocess the history before opening the session when filtering or redaction is required. Deferred capability loading --------------------------- [](https://pydantic.dev/docs/ai/realtime/capabilities/#deferred-capability-loading) Deferred capabilities load in a session the same way they do in a regular run: the capability catalog is part of the session’s instructions, and calling the `load_capability` tool returns the loaded capability’s instructions as its result — which works on every provider. What a session cannot do is advertise _new tools_ mid-conversation (the connection’s tools are fixed when it opens; see [#7288](https://github.com/pydantic/pydantic-ai/issues/7288) ), so opening a session with a `defer_loading=True` capability that contributes tools or native tools raises [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) before connecting — accepting it would silently provide less than requested. Realtime per-turn/exchange hooks are expected to widen this boundary in the future; see [#7190](https://github.com/pydantic/pydantic-ai/issues/7190) and [#7191](https://github.com/pydantic/pydantic-ai/issues/7191) . Was this page helpful? Thanks for your feedback! --- # Dynamic Workflow | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#_top) Dynamic Workflow ================ `DynamicWorkflow` is for the case where the coordination _between_ sub-agents is the actual work. Say you have a few specialists — one reviews code, one summarizes findings, one writes the final note. Each is easy to call on its own; the hard part is the choreography: review three files at once, keep only the reports that found something, summarize those, and hand the summary to the writer. Reach for this capability when that orchestration involves fan-out, chaining, voting, or retry loops that you do not want to run one model turn at a time, with every intermediate result flowing back through the orchestrator’s context. The idea -------- [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#the-idea) The usual way to coordinate sub-agents is one tool call per step. The agent calls the reviewer and waits, reads the result, calls the reviewer again, waits again, and so on. Every intermediate result travels back into the agent’s context, and every step that depends on the previous one is a separate model turn. `DynamicWorkflow` takes a different route. You hand it a catalog of named sub-agents, and it gives the model a single tool, `run_workflow`. Inside that tool the model writes ordinary Python: each of your sub-agents is an `async` function it can call, loop over, and combine. The script runs to completion in one tool call, and only its final value comes back to the model. The choreography moves out of the conversation and into code. If you have met [Code Mode](https://pydantic.dev/docs/ai/harness/code-mode/) , this will feel familiar — the same [Monty](https://github.com/pydantic/monty) sandbox and the same idea: write a script instead of many tool calls. The difference is what the script gets to call. In Code Mode it calls the agent’s own tools; here it calls whole sub-agents. How this relates to Subagents ----------------------------- [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#how-this-relates-to-subagents) The harness has two delegation capabilities. They trade in the same currency — named, isolated sub-agent runs — but at different altitudes: * [`SubAgents`](https://pydantic.dev/docs/ai/harness/subagents/) exposes one `delegate_task(agent_name, task)` tool. Each delegation is its own tool call and its own model turn. It is the right fit when delegations are occasional, or when each result needs the parent’s judgment before the next one. * `DynamicWorkflow` moves the choreography into a script. Fan-out, chaining, voting, and retry loops all run inside one tool call, and intermediate results never enter the parent’s context. Start with `SubAgents` if you are not sure. A `delegate_task` orchestrator converts to a workflow catalog without changing the sub-agents themselves. Installation ------------ [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#installation) The script runs inside the Monty sandbox, so install the extra: Terminal uv add "pydantic-ai-harness[dynamic-workflow]" Your first workflow ------------------- [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#your-first-workflow) Two sub-agents, one orchestrator: from pydantic_ai import Agent from pydantic_ai_harness.dynamic_workflow import DynamicWorkflow reviewer = Agent('openai:gpt-5', name='reviewer', description='Reviews code for bugs.') summarizer = Agent('openai:gpt-5', name='summarizer', description='Summarizes findings.') orchestrator = Agent( 'openai:gpt-5', capabilities=[DynamicWorkflow(agents=[reviewer, summarizer])], ) `reviewer` and `summarizer` are plain agents — the same `Agent` you already know. Their `name` becomes the function name the model calls in the script, so pick names that are valid Python identifiers. Their `description` tells the model what each one is for; write it the way you would document a function. `DynamicWorkflow(agents=[...])` bundles them into one capability and hands the orchestrator a single `run_workflow` tool. What the model does with it --------------------------- [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#what-the-model-does-with-it) When the orchestrator decides to use the tool, it does not call your sub-agents one at a time. It writes a script: import asyncio reports = await asyncio.gather( reviewer(task="Review auth.py for bugs:\n"), reviewer(task="Review parser.py for bugs:\n"), ) await summarizer(task="Summarize these review findings:\n" + "\n\n".join(reports)) The parts that matter: * Each sub-agent is an `async` function. You call it with `await`. * You pass the work as a single keyword argument, `task`. Always by keyword — `reviewer(task="...")`, not `reviewer("...")`. * `asyncio.gather(...)` runs the two reviews concurrently instead of one after the other. * The last expression’s value becomes the result the model sees. The intermediate `reports` list never leaves the sandbox. Each call is a full `Agent.run`, with its own model loop, message history, tools, and typed output. Two things follow: calls are **isolated** (a sub-agent remembers nothing from an earlier call, so put everything it needs into `task`), and calls **cost tokens and take time** (which is why this capability gives you budgets, below). Sub-agents can return structured data ------------------------------------- [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#sub-agents-can-return-structured-data) A sub-agent returns whatever its `output_type` produces. The default is a string, but give a sub-agent a Pydantic model and the script receives a `dict`: from pydantic import BaseModel class Score(BaseModel): value: int reason: str critic = Agent('openai:gpt-5', name='critic', description='Scores an answer 0-10.', output_type=Score) Inside the script the model reads the fields by subscript, the way it would read a JSON object: result = await critic(task="Score this answer: ...") result["value"] # not result.value The catalog the model sees renders each output type as a `TypedDict`, so it knows the fields and reads them by subscript on its own. How results come back --------------------- [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#how-results-come-back) The value of the script’s last expression becomes the tool result — the model does not `print()` it. | The script… | The model receives | | --- | --- | | ends in a value, no print | that value directly (or `{}` if it is `None`) | | prints and ends in a value | `{"output": "", "result": }` | | prints and ends in `None` | `{"output": ""}` | `print()` is for debug logging; it stringifies, so let the last expression carry the real result. Choosing sub-agent models ------------------------- [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#choosing-sub-agent-models) By default each sub-agent uses the model it was constructed with. Set `inherit_model=True` when the host passes a per-run model override to the parent agent (for example from a `/model` command) and every sub-agent dispatch should follow that resolved parent model. Leave it `False` when a sub-agent is deliberately pinned to a different model. Keeping it safe: budgets ------------------------ [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#keeping-it-safe-budgets) A sub-agent is non-deterministic, costs tokens, and can fan out into more sub-agents. `DynamicWorkflow` gives you a hard count ceiling, token budgets, and a guard against runaway sandbox scripts. ### `max_agent_calls` — an exact count [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#max_agent_calls--an-exact-count) DynamicWorkflow(agents=[...], max_agent_calls=50) # 50 is the default A hard, host-enforced ceiling on the number of sub-agent runs in one parent run. It is one budget shared across every `run_workflow` call in that run, and it holds exactly even when the script fans out with `asyncio.gather`. When the budget runs out, the workflow stops calling sub-agents and returns a terminal result with bounded previews of up to the 20 most recent completed results. This is the only knob that bounds the number of runs exactly. ### `sub_agent_usage_limits` and `forward_usage` — bounding cost [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#sub_agent_usage_limits-and-forward_usage--bounding-cost) `sub_agent_usage_limits` is a `UsageLimits` applied to each sub-agent run. `forward_usage` controls whether the whole tree shares one usage counter: | `forward_usage` | Counter | What the limit means | | --- | --- | --- | | `True` (default) | the parent’s `usage` is shared across the tree | a tree-wide cap. Under concurrent fan-out it is best-effort: several sub-agents can pass the check before any of them adds to the count. | | `False` | each sub-agent run counts on its own | per-run limits. A per-run `total_tokens_limit` of `T` with `max_agent_calls` of `N` bounds the tree to roughly `N * T` tokens. | ### `resource_limits` — guarding the script itself [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#resource_limits--guarding-the-script-itself) These limits guard the orchestration script’s own memory, not the sub-agents it calls. The default backstop is 256 MB with no time limit. Printed output is collected separately with Monty’s 10 MiB default cap. DynamicWorkflow(agents=[...], resource_limits={'max_duration_secs': 30}) `max_duration_secs` measures the time your script spends running sandbox code, not wall-clock time. While the script waits on a sub-agent it is suspended and that time does not count, so the cap will not fire on a normal workflow no matter how long the sub-agents take. Its one job is catching a pure-CPU runaway — a `while True:` loop that never awaits, which none of the sub-agent budgets can stop because it never calls a sub-agent. Pass `'unlimited'` to remove every sandbox resource limit; the separate 10 MiB print cap remains. A partial dict merges onto the backstop so you override only the caps you name. ### Workflows do not nest [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#workflows-do-not-nest) A sub-agent cannot start its own workflow; a nested `run_workflow` call returns a terminal error instead of running. The practical rule: do not give the sub-agents in your catalog the `DynamicWorkflow` capability. They are the leaves of the orchestration, not orchestrators. Renaming a sub-agent: `WorkflowAgent` ------------------------------------- [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#renaming-a-sub-agent-workflowagent) By default a sub-agent shows up under its own `name` and `description`. To give it a different name or description for one workflow without editing the agent itself, wrap it in a `WorkflowAgent`: from pydantic_ai_harness.dynamic_workflow import WorkflowAgent DynamicWorkflow( agents=[\ WorkflowAgent(\ reviewer,\ name='check',\ description='Checks one code change and returns actionable review findings.',\ ),\ ], ) Now the model calls `check(task=...)`. Passing a bare agent is shorthand for wrapping it in a `WorkflowAgent` with no overrides. Adding sub-agents mid-run: `reveal()` ------------------------------------- [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#adding-sub-agents-mid-run-reveal) The catalog is fixed when a run starts, which keeps it in the prompt-cache prefix across turns. To make a new sub-agent available during a run (say once a fixer agent has been provisioned), keep a reference to the `DynamicWorkflow` instance and call `reveal()`: workflow = DynamicWorkflow(agents=[reviewer]) orchestrator = Agent('openai:gpt-5', deps_type=MyDeps, capabilities=[workflow]) # later, from the host or from another tool: workflow.reveal(fixer) The revealed sub-agent becomes callable on the next step; the model learns about it through a short announcement message that carries the new function’s signature. The `run_workflow` description itself stays frozen at the agents present when the run started, so a runtime reveal never moves the prompt-cache prefix. `reveal()` is append-only and validates immediately — a missing name, an invalid identifier, a reserved keyword, or a name collision raises `UserError` at the call site. Loading it only when needed: `defer_loading` -------------------------------------------- [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#loading-it-only-when-needed-defer_loading) `DynamicWorkflow` carries a fair amount of instruction text, and most turns do not need it. Keep it collapsed to a one-line entry until the model actually loads it: DynamicWorkflow( agents=[reviewer, summarizer], id='workflow', defer_loading=True, ) `defer_loading=True` needs a stable `id`. See [on-demand capabilities](https://pydantic.dev/docs/ai/capabilities/on-demand/) for the full picture. What runs in the sandbox ------------------------ [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#what-runs-in-the-sandbox) The script runs in Monty, a subset of Python. Knowing the edges matters: * No third-party libraries. * Importable standard-library modules include `sys`, `typing`, `asyncio`, `math`, `json`, `re`, `unicodedata`, `datetime`, `os`, and `pathlib`. Import what you use. Filesystem, environment, and clock operations are not configured for workflow scripts. * No wall-clock or timing primitives — no `asyncio.sleep`, no `datetime.datetime.now()`, no `datetime.date.today()`, and no `time` module. * `asyncio.gather(...)` runs sub-agents concurrently with positional awaitables but no keyword arguments, including `return_exceptions=True`. Other task creation and wait APIs are unavailable. Before a script runs it is statically type-checked against the sub-agent signatures. An ordinary, statically provable mistake such as a misspelled function, a positional `task`, or a wrong-typed argument costs one retry but no sub-agent budget or sandbox execution. Values typed as `Any` can reach runtime validation; they are still rejected before a sub-agent runs. Observability ------------- [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#observability) The [Logfire](https://pydantic.dev/logfire) trace is the best way to see what a workflow did. Each sub-agent run appears nested under the `run_workflow` span, and the span carries the exact `code` argument the model wrote, so you can read the script it actually ran. Until first-class progress streaming ships, set `event_stream_handler` on each sub-agent `Agent` to watch sub-agent runs inside the one tool call. API --- [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#api) DynamicWorkflow( # all parameters are keyword-only agents=[...], # Sequence[AbstractAgent | WorkflowAgent], required tool_name='run_workflow', max_agent_calls=50, max_retries=3, forward_usage=True, inherit_model=False, # True -> sub-agents run with the parent run's resolved model sub_agent_usage_limits=None, # UsageLimits per sub-agent run; None -> pydantic-ai default resource_limits=None, # None -> backstop (256 MB, no time cap); # 'unlimited' -> sandbox limits off; a dict merges onto the backstop id=None, # required when defer_loading=True description=None, # one-line catalog entry shown while deferred defer_loading=False, ) workflow.reveal(agent) # AbstractAgent | WorkflowAgent; validates before appending WorkflowAgent( agent, # AbstractAgent, required, positional name=None, # sandbox function name; falls back to agent.name description=None, # function docstring; falls back to agent.description ) `DynamicWorkflowToolset` and `WorkflowResourceLimits` are also exported from the module for advanced use. Source: [`pydantic_ai_harness/dynamic_workflow/`](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/dynamic_workflow/) . Further reading --------------- [](https://pydantic.dev/docs/ai/harness/dynamic-workflow/#further-reading) * [Code Mode](https://pydantic.dev/docs/ai/harness/code-mode/) — the same sandbox, calling the agent’s own tools instead of sub-agents. * [Subagents](https://pydantic.dev/docs/ai/harness/subagents/) — one-delegation-per-tool-call sub-agents, without the scripted choreography. * [Rewriting Bun in Rust](https://bun.com/blog/bun-in-rust) (Bun) — the same pattern at scale, via Claude Code’s dynamic workflows. * [Capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) and [on-demand capabilities](https://pydantic.dev/docs/ai/capabilities/on-demand/) . Was this page helpful? Thanks for your feedback! --- # pydantic_evals.otel | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#_top) pydantic\_evals.otel ==================== SpanNode -------- [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode) A node in the span tree; provides references to parents/children for easy traversal and queries. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#attributes) #### ancestors [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode.ancestors) Return all ancestors of this node. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[`SpanNode`\] #### descendants [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode.descendants) Return all descendants of this node in DFS order. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[`SpanNode`\] #### duration [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode.duration) Return the span’s duration as a timedelta. **Type:** `timedelta` #### status [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode.status) The span’s status; `'error'` if the operation the span represents raised an exception. **Type:** `SpanStatus` **Default:** `'unset'` ### Methods [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#methods) #### add\_child [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode.add_child) def add_child(child: SpanNode) -> None Attach a child node to this node’s list of children. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns) [`None`](https://docs.python.org/3/library/constants.html#None) #### any\_ancestor [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode.any_ancestor) def any_ancestor( predicate: SpanQuery | SpanPredicate, stop_recursing_when: SpanQuery | SpanPredicate | None = None, ) -> bool Returns True if any ancestor satisfies the predicate. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-1) [`bool`](https://docs.python.org/3/library/functions.html#bool) #### any\_child [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode.any_child) def any_child(predicate: SpanQuery | SpanPredicate) -> bool Returns True if there is at least one child that satisfies the predicate. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-2) [`bool`](https://docs.python.org/3/library/functions.html#bool) #### any\_descendant [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode.any_descendant) def any_descendant( predicate: SpanQuery | SpanPredicate, stop_recursing_when: SpanQuery | SpanPredicate | None = None, ) -> bool Returns `True` if there is at least one descendant that satisfies the predicate. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-3) [`bool`](https://docs.python.org/3/library/functions.html#bool) #### find\_ancestors [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode.find_ancestors) def find_ancestors( predicate: SpanQuery | SpanPredicate, stop_recursing_when: SpanQuery | SpanPredicate | None = None, ) -> list[SpanNode] Return all ancestors that satisfy the given predicate. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-4) [`list`](https://docs.python.org/3/glossary.html#term-list) \[`SpanNode`\] #### find\_children [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode.find_children) def find_children(predicate: SpanQuery | SpanPredicate) -> list[SpanNode] Return all immediate children that satisfy the given predicate. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-5) [`list`](https://docs.python.org/3/glossary.html#term-list) \[`SpanNode`\] #### find\_descendants [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode.find_descendants) def find_descendants( predicate: SpanQuery | SpanPredicate, stop_recursing_when: SpanQuery | SpanPredicate | None = None, ) -> list[SpanNode] Return all descendant nodes that satisfy the given predicate in DFS order. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-6) [`list`](https://docs.python.org/3/glossary.html#term-list) \[`SpanNode`\] #### first\_ancestor [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode.first_ancestor) def first_ancestor( predicate: SpanQuery | SpanPredicate, stop_recursing_when: SpanQuery | SpanPredicate | None = None, ) -> SpanNode | None Return the closest ancestor that satisfies the given predicate, or `None` if none match. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-7) `SpanNode` | [`None`](https://docs.python.org/3/library/constants.html#None) #### first\_child [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode.first_child) def first_child(predicate: SpanQuery | SpanPredicate) -> SpanNode | None Return the first immediate child that satisfies the given predicate, or None if none match. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-8) `SpanNode` | [`None`](https://docs.python.org/3/library/constants.html#None) #### first\_descendant [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode.first_descendant) def first_descendant( predicate: SpanQuery | SpanPredicate, stop_recursing_when: SpanQuery | SpanPredicate | None = None, ) -> SpanNode | None DFS: Return the first descendant (in DFS order) that satisfies the given predicate, or `None` if none match. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-9) `SpanNode` | [`None`](https://docs.python.org/3/library/constants.html#None) #### matches [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode.matches) def matches(query: SpanQuery | SpanPredicate) -> bool Check if the span node matches the query conditions or predicate. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-10) [`bool`](https://docs.python.org/3/library/functions.html#bool) #### repr\_xml [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode.repr_xml) def repr_xml( include_children: bool = True, include_trace_id: bool = False, include_span_id: bool = False, include_start_timestamp: bool = False, include_duration: bool = False, ) -> str Return an XML-like string representation of the node. Optionally includes children, trace\_id, span\_id, start\_timestamp, and duration. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-11) [`str`](https://docs.python.org/3/library/stdtypes.html#str) SpanQuery --------- [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanQuery) **Bases:** [`TypedDict`](https://docs.python.org/3/library/typing.html#typing.TypedDict) A serializable query for filtering SpanNodes based on various conditions. All fields are optional and combined with AND logic by default. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#attributes-1) #### has\_attributes [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanQuery.has_attributes) Attribute values are compared with equality; a dict or list value also matches an attribute stored as its JSON serialization, since OTel attributes cannot hold nested objects and instrumentation libraries like Logfire store them as JSON strings. A list value also matches an attribute stored as a tuple. **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] #### stop\_recursing\_when [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanQuery.stop_recursing_when) If present, stop recursing through ancestors or descendants at nodes that match this condition. **Type:** `SpanQuery` SpanTree -------- [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanTree) A container that builds a hierarchy of SpanNode objects from a list of finished spans. You can then search or iterate the tree to make your assertions (using DFS for traversal). ### Methods [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#methods-1) #### \_\_iter\_\_ [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanTree.__iter__) def __iter__() -> Iterator[SpanNode] Return an iterator over all nodes in the tree. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-12) [`Iterator`](https://docs.python.org/3/library/typing.html#typing.Iterator) \[`SpanNode`\] #### add\_spans [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanTree.add_spans) def add_spans(spans: list[SpanNode]) -> None Add a list of spans to the tree, rebuilding the tree structure. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-13) [`None`](https://docs.python.org/3/library/constants.html#None) #### any [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanTree.any) def any(predicate: SpanQuery | SpanPredicate) -> bool Returns True if any node in the tree matches the predicate. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-14) [`bool`](https://docs.python.org/3/library/functions.html#bool) #### find [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanTree.find) def find(predicate: SpanQuery | SpanPredicate) -> list[SpanNode] Find all nodes in the entire tree that match the predicate, scanning from each root in DFS order. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-15) [`list`](https://docs.python.org/3/glossary.html#term-list) \[`SpanNode`\] #### first [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanTree.first) def first(predicate: SpanQuery | SpanPredicate) -> SpanNode | None Find the first node that matches a predicate, scanning from each root in DFS order. Returns `None` if not found. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-16) `SpanNode` | [`None`](https://docs.python.org/3/library/constants.html#None) #### repr\_xml [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanTree.repr_xml) def repr_xml( include_children: bool = True, include_trace_id: bool = False, include_span_id: bool = False, include_start_timestamp: bool = False, include_duration: bool = False, ) -> str Return an XML-like string representation of the tree, optionally including children, trace\_id, span\_id, duration, and timestamps. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-17) [`str`](https://docs.python.org/3/library/stdtypes.html#str) SpanTreeRecordingError ---------------------- [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanTreeRecordingError) **Bases:** [`Exception`](https://docs.python.org/3/library/exceptions.html#Exception) An exception that is used to provide the reason why a SpanTree was not recorded by `context_subtree`. This may be due to missing dependencies, a tracer provider not having been set, or a custom TracerProvider that does not support `add_span_processor`. ### Methods [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#methods-2) #### \_\_get\_pydantic\_core\_schema\_\_ [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanTreeRecordingError.__get_pydantic_core_schema__) `@classmethod` def __get_pydantic_core_schema__(cls, _: Any, __: Any) -> core_schema.CoreSchema Pydantic core schema to allow `SpanTreeRecordingError` to be (de)serialized. Only the human-readable `message` is preserved by design: the exception’s `__context__`, `__cause__`, and traceback (e.g. the underlying `ImportError` chained in `context_subtree`) are dropped on serialization and not reconstructed on the way back. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#returns-18) `core_schema.CoreSchema` SpanStatus ---------- [](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanStatus) The status of a span, mirroring `opentelemetry.trace.StatusCode`. **Default:** `Literal['unset', 'ok', 'error']` Was this page helpful? Thanks for your feedback! --- # Hooks | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/core-concepts/hooks/#_top) Hooks ===== Hooks let you intercept and modify agent behavior at every stage of a run — model requests, tool calls, streaming events — using simple decorators or constructor arguments. No subclassing needed. The [`Hooks`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Hooks) capability is the recommended way to add [lifecycle hooks](https://pydantic.dev/docs/ai/capabilities/custom/#hooking-into-the-lifecycle) for application-level concerns like logging, metrics, and lightweight validation. For reusable capabilities that combine hooks with tools, instructions, or model settings, subclass [`AbstractCapability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability) instead — see [Building custom capabilities](https://pydantic.dev/docs/ai/capabilities/custom/) . Quick start ----------- [](https://pydantic.dev/docs/ai/core-concepts/hooks/#quick-start) Create a [`Hooks`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Hooks) instance, register hooks via `@hooks.on.*` decorators, and pass it to your agent: hooks\_decorator.py from pydantic_ai import Agent, ModelRequestContext, RunContext from pydantic_ai.capabilities import Hooks hooks = Hooks() @hooks.on.before_model_request async def log_request(ctx: RunContext, request_context: ModelRequestContext) -> ModelRequestContext: print(f'Sending {len(request_context.messages)} messages to the model') #> Sending 1 messages to the model return request_context agent = Agent('test', capabilities=[hooks]) result = agent.run_sync('Hello!') print(result.output) #> success (no tool calls) Registering hooks ----------------- [](https://pydantic.dev/docs/ai/core-concepts/hooks/#registering-hooks) ### Decorator registration [](https://pydantic.dev/docs/ai/core-concepts/hooks/#decorator-registration) The `hooks.on` namespace provides decorator methods for every lifecycle hook. Use them as bare decorators or with parameters: # Bare decorator @hooks.on.before_model_request async def my_hook(ctx, request_context): return request_context # With parameters (timeout, tool filter) @hooks.on.before_model_request(timeout=5.0) async def my_timed_hook(ctx, request_context): return request_context Multiple hooks can be registered for the same event — they fire in registration order. ### Constructor kwargs [](https://pydantic.dev/docs/ai/core-concepts/hooks/#constructor-kwargs) You can also pass hook functions directly to the [`Hooks`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Hooks) constructor: hooks\_constructor.py from pydantic_ai import Agent, ModelRequestContext, RunContext from pydantic_ai.capabilities import Hooks async def log_request(ctx: RunContext, request_context: ModelRequestContext) -> ModelRequestContext: print(f'Sending {len(request_context.messages)} messages to the model') #> Sending 1 messages to the model return request_context agent = Agent('test', capabilities=[Hooks(before_model_request=log_request)]) result = agent.run_sync('Hello!') print(result.output) #> success (no tool calls) Both sync and async hook functions are accepted. Sync functions are automatically wrapped for async execution. ### On-demand hooks [](https://pydantic.dev/docs/ai/core-concepts/hooks/#on-demand-hooks) [`Hooks`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Hooks) is a capability, so it can be loaded on demand just like any other capability. This is useful for optional, user-requested behavior such as verbose request logging: deferred\_hooks\_capability.py from pydantic_ai import Agent, ModelRequestContext, RunContext from pydantic_ai.capabilities import Hooks request_logging_hooks = Hooks( id='request-logging', description='Use when the user asks for verbose request diagnostics.', defer_loading=True, ) @request_logging_hooks.on.before_model_request async def log_request( ctx: RunContext[None], request_context: ModelRequestContext, ) -> ModelRequestContext: print(f'Model request at step {ctx.run_step}: {len(request_context.messages)} messages') return request_context agent = Agent('openai-responses:gpt-5.4', capabilities=[request_logging_hooks]) Pydantic AI skips hooks owned by a deferred `Hooks` instance until its capability is loaded. Use on-demand hooks for optional behavior that only applies after the capability is loaded. For human-in-the-loop tool approval, pass [`requires_approval=True`](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/#human-in-the-loop-tool-approval) when registering a tool, raise [`ApprovalRequired`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ApprovalRequired) for conditional approval, or wrap a toolset with [`ApprovalRequiredToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.ApprovalRequiredToolset) . Hook types ---------- [](https://pydantic.dev/docs/ai/core-concepts/hooks/#hook-types) ### Run hooks [](https://pydantic.dev/docs/ai/core-concepts/hooks/#run-hooks) | `hooks.on.` | Constructor kwarg | `AbstractCapability` method | | --- | --- | --- | | `before_run` | `before_run=` | `before_run` | | `after_run` | `after_run=` | `after_run` | | `run` | `run=` | `wrap_run` | | `run_error` | `run_error=` | `on_run_error` | Run hooks fire once per agent run. `wrap_run` (registered via `hooks.on.run`) wraps the entire run and supports error recovery. A [realtime session](https://pydantic.dev/docs/ai/realtime/capabilities/) is a run: the same four hooks fire once around the session, with `wrap_run` recovery and `after_run` result transformation applied when the session closes. ### Node hooks [](https://pydantic.dev/docs/ai/core-concepts/hooks/#node-hooks) | `hooks.on.` | Constructor kwarg | `AbstractCapability` method | | --- | --- | --- | | `before_node_run` | `before_node_run=` | `before_node_run` | | `after_node_run` | `after_node_run=` | `after_node_run` | | `node_run` | `node_run=` | `wrap_node_run` | | `node_run_error` | `node_run_error=` | `on_node_run_error` | Node hooks fire for each graph step ([`UserPromptNode`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.UserPromptNode) , [`ModelRequestNode`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.ModelRequestNode) , [`CallToolsNode`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.CallToolsNode) ). Node hooks fire no matter how the run is driven: [`agent.run()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.run) , [`agent_run.next()`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.next) , and `async for node in agent_run:` over [`agent.iter()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.iter) all advance the run the same way. ### Model request hooks [](https://pydantic.dev/docs/ai/core-concepts/hooks/#model-request-hooks) | `hooks.on.` | Constructor kwarg | `AbstractCapability` method | | --- | --- | --- | | `before_model_request` | `before_model_request=` | `before_model_request` | | `after_model_request` | `after_model_request=` | `after_model_request` | | `model_request` | `model_request=` | `wrap_model_request` | | `model_request_error` | `model_request_error=` | `on_model_request_error` | Model request hooks fire around each LLM call. [`ModelRequestContext`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestContext) bundles `model`, `messages`, `model_settings`, and `model_request_parameters`. To swap the model for a given request, set `request_context.model` to a different [`Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) instance. To skip the model call entirely, raise [`SkipModelRequest(response)`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.SkipModelRequest) from `before_model_request` or `model_request` (wrap). ### Tool validation hooks [](https://pydantic.dev/docs/ai/core-concepts/hooks/#tool-validation-hooks) | `hooks.on.` | Constructor kwarg | `AbstractCapability` method | | --- | --- | --- | | `before_tool_validate` | `before_tool_validate=` | `before_tool_validate` | | `after_tool_validate` | `after_tool_validate=` | `after_tool_validate` | | `tool_validate` | `tool_validate=` | `wrap_tool_validate` | | `tool_validate_error` | `tool_validate_error=` | `on_tool_validate_error` | Validation hooks fire when the model’s JSON arguments are parsed and validated. All tool hooks receive `call` ([`ToolCallPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolCallPart) ) and `tool_def` ([`ToolDefinition`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.ToolDefinition) ) parameters. To skip validation, raise [`SkipToolValidation(args)`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.SkipToolValidation) from `before_tool_validate` or `tool_validate` (wrap). A tool call can only be [deferred](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/) once its arguments have been validated, since whoever resolves the deferral is shown those arguments. So [`ApprovalRequired`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ApprovalRequired) and [`CallDeferred`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.CallDeferred) can be raised from `after_tool_validate` (and from `tool_validate` after its `handler()` has returned), but raising them from `before_tool_validate`, from `tool_validate` before it calls `handler()`, or from `tool_validate_error` is a [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) . To decide per tool rather than per capability, use the tool’s [`args_validator`](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#args-validator) . ### Tool execution hooks [](https://pydantic.dev/docs/ai/core-concepts/hooks/#tool-execution-hooks) | `hooks.on.` | Constructor kwarg | `AbstractCapability` method | | --- | --- | --- | | `before_tool_execute` | `before_tool_execute=` | `before_tool_execute` | | `after_tool_execute` | `after_tool_execute=` | `after_tool_execute` | | `tool_execute` | `tool_execute=` | `wrap_tool_execute` | | `tool_execute_error` | `tool_execute_error=` | `on_tool_execute_error` | Execution hooks fire when the tool function runs. `args` is always the validated `dict[str, Any]`. To skip execution, raise [`SkipToolExecution(result)`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.SkipToolExecution) from `before_tool_execute` or `tool_execute` (wrap). Every execution hook can [defer](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/) the call — the arguments are validated by this point — but raise `ApprovalRequired`/`CallDeferred` from `before_tool_execute` (or from `tool_execute` before it calls `handler()`). A deferral from `after_tool_execute`, or from `tool_execute` after `handler()` has returned, is accepted but happens too late to be useful: the tool function already ran, so its side effects happened and its result is discarded. ### Output validation hooks [](https://pydantic.dev/docs/ai/core-concepts/hooks/#output-validation-hooks) | `hooks.on.` | Constructor kwarg | `AbstractCapability` method | | --- | --- | --- | | `before_output_validate` | `before_output_validate=` | `before_output_validate` | | `after_output_validate` | `after_output_validate=` | `after_output_validate` | | `output_validate` | `output_validate=` | `wrap_output_validate` | | `output_validate_error` | `output_validate_error=` | `on_output_validate_error` | Output validation hooks fire when structured output is parsed against the output schema. They do **not** fire for plain text or image output. All output hooks receive an `output_context` ([`OutputContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.OutputContext) ) parameter. ### Output processing hooks [](https://pydantic.dev/docs/ai/core-concepts/hooks/#output-processing-hooks) | `hooks.on.` | Constructor kwarg | `AbstractCapability` method | | --- | --- | --- | | `before_output_process` | `before_output_process=` | `before_output_process` | | `after_output_process` | `after_output_process=` | `after_output_process` | | `output_process` | `output_process=` | `wrap_output_process` | | `output_process_error` | `output_process_error=` | `on_output_process_error` | Output processing hooks fire when the output is processed — extracting values, calling output functions, and running output validators. See [Output hooks](https://pydantic.dev/docs/ai/capabilities/custom/#output-hooks) for the full lifecycle, signatures, and details on how output validators interact with processing hooks. ### Tool preparation [](https://pydantic.dev/docs/ai/core-concepts/hooks/#tool-preparation) | `hooks.on.` | Constructor kwarg | `AbstractCapability` method | | --- | --- | --- | | `prepare_tools` | `prepare_tools=` | `prepare_tools` | | `prepare_output_tools` | `prepare_output_tools=` | `prepare_output_tools` | Filters or modifies tool definitions the model sees on each step. `prepare_tools` handles **function** tools; `prepare_output_tools` handles [output tools](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.ToolOutput) separately, with `ctx.max_retries` reflecting the **output** retry budget. Both run as `PreparedToolset` wrappers — the result flows into the model’s request _and_ `ToolManager.tools`, so filtering also blocks tool execution. ### Deferred tool call hook [](https://pydantic.dev/docs/ai/core-concepts/hooks/#deferred-tool-call-hook) | `hooks.on.` | Constructor kwarg | `AbstractCapability` method | | --- | --- | --- | | `deferred_tool_calls` | `deferred_tool_calls=` | `handle_deferred_tool_calls` | Resolves [deferred tool calls](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/) (approval-required or externally-executed) inline during a run. The hook receives a [`DeferredToolRequests`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.DeferredToolRequests) and returns a [`DeferredToolResults`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.DeferredToolResults) (or `None` to decline). Multiple registered hooks accumulate: each receives the still-unresolved requests and can resolve some or all of them. hooks\_deferred\_tool\_calls.py from pydantic_ai import Agent, DeferredToolRequests, DeferredToolResults, RunContext from pydantic_ai.capabilities import Hooks hooks = Hooks() @hooks.on.deferred_tool_calls async def auto_approve( ctx: RunContext, *, requests: DeferredToolRequests ) -> DeferredToolResults: return requests.build_results(approve_all=True) agent = Agent('test', capabilities=[hooks]) @agent.tool_plain(requires_approval=True) def delete_file(path: str) -> str: return f'File {path!r} deleted' For pure application-level handler registration without other hooks, the dedicated [`HandleDeferredToolCalls`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.HandleDeferredToolCalls) capability is more concise — see [Resolving deferred calls with a handler](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/#resolving-deferred-calls-with-a-handler) . ### Event stream hooks [](https://pydantic.dev/docs/ai/core-concepts/hooks/#event-stream-hooks) | `hooks.on.` | Constructor kwarg | `AbstractCapability` method | | --- | --- | --- | | `run_event_stream` | `run_event_stream=` | `wrap_run_event_stream` | | `event` | `event=` | _(per-event convenience)_ | `run_event_stream` wraps the full event stream as an async generator. `event` is a convenience — it fires for each individual event during a streamed run. Tool and model events flow through this stream, along with framework events such as [`EnqueuedMessagesEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.EnqueuedMessagesEvent) when queued messages enter run history. During a [realtime session](https://pydantic.dev/docs/ai/realtime/capabilities/) , both hooks also fire, and realtime-only [`RealtimeEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeEvent) members flow through the same stream: hooks\_event.py from pydantic_ai import Agent, AgentStreamEvent, RunContext from pydantic_ai.capabilities import Hooks hooks = Hooks() event_count = 0 @hooks.on.event async def count_events(ctx: RunContext, event: AgentStreamEvent) -> AgentStreamEvent: global event_count event_count += 1 return event agent = Agent('test', capabilities=[hooks]) Tool hook filtering ------------------- [](https://pydantic.dev/docs/ai/core-concepts/hooks/#tool-hook-filtering) Tool hooks (validation and execution) support a `tools` parameter to target specific tools by name: hooks\_tool\_filter.py from pydantic_ai import Agent, RunContext, ToolDefinition from pydantic_ai.capabilities import Hooks, ValidatedToolArgs from pydantic_ai.messages import ToolCallPart hooks = Hooks() call_log: list[str] = [] @hooks.on.before_tool_execute(tools=['send_email']) async def audit_dangerous_tools( ctx: RunContext, *, call: ToolCallPart, tool_def: ToolDefinition, args: ValidatedToolArgs, ) -> ValidatedToolArgs: call_log.append(f'audit: {call.tool_name}') return args agent = Agent('test', capabilities=[hooks]) @agent.tool_plain def send_email(to: str) -> str: return f'sent to {to}' result = agent.run_sync('Send an email to test@example.com') print(call_log) #> ['audit: send_email'] The `tools` parameter accepts a sequence of tool names. The hook only fires for matching tools — other tool calls pass through unaffected. Timeouts -------- [](https://pydantic.dev/docs/ai/core-concepts/hooks/#timeouts) Each hook supports an optional `timeout` in seconds. If the hook exceeds the timeout, a [`HookTimeoutError`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.HookTimeoutError) is raised: hooks\_timeout.py import asyncio from pydantic_ai import Agent, ModelRequestContext, RunContext from pydantic_ai.capabilities import Hooks, HookTimeoutError hooks = Hooks() @hooks.on.before_model_request(timeout=0.01) async def slow_hook( ctx: RunContext, request_context: ModelRequestContext ) -> ModelRequestContext: await asyncio.sleep(10) # Will be interrupted by timeout return request_context # pragma: no cover agent = Agent('test', capabilities=[hooks]) try: agent.run_sync('Hello') except HookTimeoutError as e: print(f'Hook timed out: {e.hook_name} after {e.timeout}s') #> Hook timed out: before_model_request after 0.01s Timeouts are set via the decorator parameter (`@hooks.on.before_model_request(timeout=5.0)`) or via the constructor when using kwargs. Wrap hooks ---------- [](https://pydantic.dev/docs/ai/core-concepts/hooks/#wrap-hooks) Wrap hooks let you surround an operation with setup/teardown logic. In the `hooks.on` namespace, wrap hooks drop the `wrap_` prefix — `hooks.on.model_request` corresponds to `wrap_model_request`: hooks\_wrap.py from pydantic_ai import Agent, ModelRequestContext, RunContext from pydantic_ai.capabilities import Hooks, WrapModelRequestHandler from pydantic_ai.messages import ModelResponse hooks = Hooks() wrap_log: list[str] = [] @hooks.on.model_request async def log_request( ctx: RunContext, *, request_context: ModelRequestContext, handler: WrapModelRequestHandler ) -> ModelResponse: wrap_log.append('before') response = await handler(request_context) wrap_log.append('after') return response agent = Agent('test', capabilities=[hooks]) result = agent.run_sync('Hello!') print(wrap_log) #> ['before', 'after'] Hook ordering ------------- [](https://pydantic.dev/docs/ai/core-concepts/hooks/#hook-ordering) Within a single [`Hooks`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Hooks) instance, `before_*`, `after_*`, and `on_*_error` fire in **registration order** (the order they were defined or passed to the constructor). `wrap_*` nests as middleware, with the first-registered wrapper as the outermost layer. Across multiple capabilities, the [composition rules](https://pydantic.dev/docs/ai/capabilities/custom/#composition-and-middleware-semantics) apply: `before_*` fires in capability order, `after_*` fires in reverse capability order, and `wrap_*` nests as middleware with the first capability outermost. Hook timing also affects what is populated on [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) . Early run and node hooks can fire before the current step’s tool manager and model request parameters have been assembled. At that point `ctx.available_tool_names` can still include tool-search discoveries reconstructed from history, but `ctx.tools` and current request parameters may be empty or reflect the previous step. `before_model_request` and later model-request hooks see the request about to be sent, including the current function tools, native tools, and model settings. Tool and output hooks see the state for the call or output currently being processed. For on-demand capabilities, `ctx.loaded_capability_ids` is derived from message history before each model request, so a capability loaded during a step appears from the _next_ step onwards — the same step that first carries its instructions to the model, and therefore the first on which its tools can be called. Function tools, native tools, and model settings from the loaded capability appear on that request too, and hooks owned by the capability run for hook points reached from then on. A hook that looks for a capability in the very turn it was loaded will not find it. Error hooks ----------- [](https://pydantic.dev/docs/ai/core-concepts/hooks/#error-hooks) Error hooks (`*_error` in the `hooks.on` namespace, `on_*_error` on `AbstractCapability`) use **raise-to-propagate, return-to-recover** semantics: * **Raise the original error** — propagates unchanged _(default)_ * **Raise a different exception** — transforms the error * **Return a result** — suppresses the error See [Error hooks](https://pydantic.dev/docs/ai/capabilities/custom/#error-hooks) for the full pattern and recovery types. Triggering retries with `ModelRetry` and failures with `ToolFailed` ------------------------------------------------------------------- [](https://pydantic.dev/docs/ai/core-concepts/hooks/#triggering-retries-with-modelretry) Hooks can raise [`ModelRetry`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelRetry) to ask the model to try again with a custom message — the same exception used in [tool functions](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-retries) and output validators. **Model request hooks** (`after_model_request`, `wrap_model_request`, `on_model_request_error`): * The retry message is sent back to the model as a [`RetryPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.RetryPromptPart) * `after_model_request`: the original response is preserved in message history so the model can see what it said * `wrap_model_request`: the response is preserved only if the handler was called * Retries count against the output side of the agent’s retry budget **Tool hooks** (`before/after_tool_validate`, `before/after_tool_execute`, `wrap_tool_execute`, `on_tool_execute_error`): * Converted to tool retry prompts, same as when a tool function raises `ModelRetry` * Retries count against the tool’s `max_retries` limit **Output hooks** (`before/after_output_validate`, `before/after_output_process`, `wrap_output_process`, `on_output_process_error`): * Converted to retry prompts, same as when an output function raises `ModelRetry` * For tool output, retries count against the tool’s `max_retries` limit * For text output, retries count against the output side of the agent’s retry budget [`ModelRetry`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelRetry) from `wrap_model_request`, `wrap_tool_execute`, or `wrap_output_process` is control flow and bypasses the corresponding `on_*_error` hook. [`ToolFailed`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ToolFailed) is control flow only at the tool boundary, so it bypasses `on_tool_execute_error`. From model-request and output-process hooks, `ToolFailed` is an ordinary exception and is passed to `on_model_request_error` or `on_output_process_error`. Tool validation and execution hooks can also raise [`ToolFailed`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ToolFailed) to report a failed tool result without consuming the tool’s retry budget. This has the same model-visible outcome and retry-budget behavior as raising `ToolFailed` from the tool function itself, and is useful when an error hook converts a third-party exception into a failure the model can see. hooks\_model\_retry.py from pydantic_ai import Agent, RunContext from pydantic_ai.capabilities import Hooks from pydantic_ai.exceptions import ModelRetry from pydantic_ai.messages import ModelResponse from pydantic_ai.models import ModelRequestContext hooks = Hooks() @hooks.on.after_model_request async def check_response( ctx: RunContext, *, request_context: ModelRequestContext, response: ModelResponse, ) -> ModelResponse: if 'PLACEHOLDER' in str(response.parts): raise ModelRetry('Response contains placeholder text. Please provide real data.') return response agent = Agent('test', capabilities=[hooks]) result = agent.run_sync('Hello') print(result.output) #> success (no tool calls) By default, any exception other than `ModelRetry` or `ToolFailed` raised inside a tool escapes the tool boundary and aborts the entire run. A tool-execution hook lets you intercept these in one place — without editing every tool — and choose how each surfaces to the model. The distinction is the semantic one between [requesting a retry](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-retries) and [reporting a failure](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-failed) : * raise `ModelRetry` for **transient** errors, where the same call might succeed if tried again; * raise `ToolFailed` for **definitive** failures, where retrying won’t help and the model should see the result and adapt (choose another approach, tell the user, etc.). The hook below makes that call based on an upstream status code — the per-error analogue of the MCP [`tool_error_behavior`](https://pydantic.dev/docs/ai/mcp/client/#tool-errors) setting: hooks\_convert\_tool\_errors.py from typing import Any from pydantic_ai import Agent, RunContext, ToolCallPart, ToolDefinition, ToolReturnPart from pydantic_ai.capabilities import Hooks from pydantic_ai.exceptions import ModelRetry, ToolFailed from pydantic_ai.messages import ModelMessage, ModelResponse, TextPart from pydantic_ai.models.function import AgentInfo, FunctionModel class UpstreamError(Exception): """Stand-in for an HTTP client error that carries the response status code.""" def __init__(self, status_code: int, message: str): super().__init__(message) self.status_code = status_code hooks = Hooks() @hooks.on.tool_execute_error async def convert_upstream_errors( ctx: RunContext[None], *, call: ToolCallPart, tool_def: ToolDefinition, args: dict[str, Any], error: Exception, ) -> Any: if isinstance(error, UpstreamError): if error.status_code >= 500 or error.status_code == 429: # Transient: the same call might succeed, so ask the model to try again. raise ModelRetry(f'Upstream returned {error.status_code}, please try again.') # Definitive (e.g. 404, 403): retrying won't help — report it so the model can adapt. raise ToolFailed(f'Upstream returned {error.status_code}: {error}') raise error # unrelated errors still abort the run def model_fn(messages: list[ModelMessage], info: AgentInfo) -> ModelResponse: last_part = messages[-1].parts[-1] if isinstance(last_part, ToolReturnPart): return ModelResponse(parts=[TextPart(f'Could not fetch the document ({last_part.content}).')]) return ModelResponse(parts=[ToolCallPart('get_document', {'doc_id': 42}, tool_call_id='call-1')]) agent = Agent(FunctionModel(model_fn), capabilities=[hooks]) @agent.tool_plain def get_document(doc_id: int) -> str: raise UpstreamError(404, f'document {doc_id} not found') result = agent.run_sync('Fetch document 42') print(result.output) #> Could not fetch the document (Upstream returned 404: document 42 not found). Because the failure was raised as `ToolFailed` rather than `ModelRetry`, the model receives it as a [`ToolReturnPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolReturnPart) with `outcome='failed'` and decides what to do next, instead of burning a retry on a call that can’t succeed. When to use `Hooks` vs `AbstractCapability` ------------------------------------------- [](https://pydantic.dev/docs/ai/core-concepts/hooks/#when-to-use-hooks-vs-abstractcapability) | Use [`Hooks`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Hooks) | Use [`AbstractCapability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability) | | --- | --- | | Application-level hooks (logging, metrics) | Reusable, packaged capabilities | | Quick one-off interceptors | Combined tools + hooks + instructions + settings | | No configuration state needed | Complex per-run state management | | Single-file scripts | Multi-agent shared behavior | Was this page helpful? Thanks for your feedback! --- # Planning | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/planning/#_top) Planning ======== `Planning` gives the model a structured, self-updating task list through a small toolset — and surfaces the current plan back to the model every turn without ever invalidating the prompt cache. It can stay in memory for a single run or persist to SQLite/Postgres, break steps into subtasks with dependencies, and emit events from granular changes. [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/planning/) > The API may change between releases. Where practical, breaking changes ship with a deprecation warning. > This capability incorporates the task-list features of the standalone [`pydantic-ai-todo`](https://github.com/vstorm-co/pydantic-ai-todo) > library — persistent stores, subtasks, dependencies, and events — which it supersedes. If you are migrating from `pydantic-ai-todo`, the tools are renamed: > > | `pydantic-ai-todo` | `Planning` | > | --- | --- | > | `write_todos` | `write_plan` | > | `read_todos` | `read_plan` | > | `add_todo` | `add_task` | > | `update_todo_status` / `update_todo_statuses` | `update_task_status` / `update_task_statuses` | > | `remove_todo` | `remove_task` | > | `add_subtask`, `set_dependency`, `get_available_tasks` | unchanged | > > Two differences to plan for: there is no connection-string convenience (`create_storage(backend=...)` and friends are gone — you construct your own asyncpg pool or Redis client, which is what keeps the harness driver-free), and `PlanEvent` carries no `timestamp`, so a consumer that ordered or logged by it supplies its own clock. The problem ----------- [](https://pydantic.dev/docs/ai/harness/planning/#the-problem) Long agentic runs drift: the model loses track of what it set out to do and what’s left. The usual fix — keep a running plan and re-inject it into the system prompt each turn — invalidates the prompt cache. The system prompt sits at the front of the request, so every plan edit changes the cached prefix and forces the whole conversation to be re-processed at full token price. The solution ------------ [](https://pydantic.dev/docs/ai/harness/planning/#the-solution) The model owns the plan through the `planning` toolset. The current plan is surfaced back as an ephemeral reminder appended to the tail of each request, with a cache breakpoint after its stable opening tag: * The reminder is added after the durable history is persisted, so it reaches the model but is never written to `message_history`. No reminders accumulate across turns. * A `CachePoint` follows the stable `` opening tag, so the cached prefix (tools + system + real conversation + that tag) stays byte-identical turn over turn. Only the mutable reminder content falls outside the cache. Usage ----- [](https://pydantic.dev/docs/ai/harness/planning/#usage) Construct an `Agent` with `Planning()` in its `capabilities`. The tools are registered automatically and static usage guidance is added to the system prompt: from pydantic_ai import Agent from pydantic_ai_harness.planning import Planning agent = Agent('anthropic:claude-sonnet-4-6', capabilities=[Planning()]) result = agent.run_sync('Refactor the auth module and add tests.') print(result.output) The tools --------- [](https://pydantic.dev/docs/ai/harness/planning/#the-tools) | Tool | Purpose | | --- | --- | | `write_plan(items)` | Create or replace the full plan (whole-list replacement). | | `read_plan()` | Read the current plan with step ids and a progress summary. | | `add_task(content, active_form)` | Append a single `pending` step. | | `update_task_status(task_id, status)` | Move one step between statuses by id. | | `update_task_statuses(updates)` | Apply several status changes in one call, validated all-or-nothing. | | `remove_task(task_id)` | Delete a step by id. | Each step is a `content` string, an optional present-continuous `active_form` label, and a `status` (`pending`, `in_progress`, `completed`, `cancelled`). The convention — stated in the guidance and the tools’ replies — is to keep exactly one step `in_progress`. All six are registered by default. `tools=` narrows that to an allowlist, and the built-in guidance follows it: from pydantic_ai_harness.planning import Planning planning = Planning(tools=['write_plan']) # whole-plan replacement only -- one tool, no step ids to track Naming a tool the current mode does not register raises `ValueError`, as does an unknown key in `descriptions`. ### Subtasks and dependencies [](https://pydantic.dev/docs/ai/harness/planning/#subtasks-and-dependencies) Pass `enable_subtasks=True` to add three more tools, the `blocked` status, and a `hierarchical` view in `read_plan`: | Tool | Purpose | | --- | --- | | `add_subtask(parent_id, content, active_form)` | Add a child step under a parent. | | `set_dependency(task_id, depends_on_id)` | Make one step wait for another; the dependent step is auto-`blocked` until its prerequisite is resolved (completed or cancelled). Self-dependencies, cycles, and duplicates are rejected. | | `get_available_tasks()` | List steps with no incomplete dependencies — the ones that can start now. | `parent_id`, `depends_on`, and the `blocked` status are rejected by `write_plan` unless `enable_subtasks` is set: without the subtask tools nothing reconciles a dependency and no view renders the hierarchy, so storing them would be a write the plan does not reflect. Persistence ----------- [](https://pydantic.dev/docs/ai/harness/planning/#persistence) By default the plan is a fresh, isolated in-memory plan per run. Pass a `store` to persist it: from pydantic_ai_harness.planning import Planning, SqlitePlanStore planning = Planning(store=SqlitePlanStore('plan.db', session='user-123')) Built-in stores are `InMemoryPlanStore`, `SqlitePlanStore`, `PostgresPlanStore` (over a caller-owned asyncpg pool), and `RedisPlanStore` (over a caller-owned `redis.asyncio` client) — so the harness needs no database driver. Any `PlanStore` implementation works, and `store_resolver` selects one per run. `SqlitePlanStore` requires a file-backed database; use `InMemoryPlanStore` for ephemeral plans rather than `':memory:'`. The tail reminder reads the store on every model request, so a store that raises fails the run rather than degrading — the reminder is not best-effort. That is deliberate: a plan the model can no longer see is not a state to continue running in silently. Retry and fallback policy belongs to the store, not to `Planning`, and `PlanStore` is a protocol precisely so you can wrap one: class BestEffort: """Serve the last known plan when the backing store is unreachable.""" def __init__(self, inner: PlanStore) -> None: self._inner, self._last = inner, [] async def get_items(self) -> list[PlanItem]: try: self._last = await self._inner.get_items() except ConnectionError: pass return self._last # ... delegate the other five methods to `self._inner` ### Planning and executing in separate runs [](https://pydantic.dev/docs/ai/harness/planning/#planning-and-executing-in-separate-runs) A shared store is the whole handoff mechanism between two runs. One agent writes the plan, a second one executes it, and the plan is the only state that crosses between them: store = SqlitePlanStore('plan.db', session='issue-403') planner = Agent('anthropic:claude-opus-4-7', capabilities=[Planning(store=store)]) executor = Agent('anthropic:claude-sonnet-4-6', capabilities=[Planning(store=store)]) await planner.run('Investigate the issue and write a plan. Do not implement anything.') await executor.run('Implement the plan.') The executor starts with no `message_history`, so it never pays for the planner’s investigation. Its first request carries only the new prompt plus the plan reminder, which the capability rebuilds from the store. That is why the two agents can run on different models: a large-context model can do the reading and the reasoning, and a smaller one can execute against the resulting checklist. The planner’s read-only discipline is a property of how you configure that agent (which toolsets it gets, and what its instructions say), not something the capability enforces. Events ------ [](https://pydantic.dev/docs/ai/harness/planning/#events) Attach a `PlanEventEmitter` to a store to react to changes: from pydantic_ai_harness.planning import InMemoryPlanStore, PlanEventEmitter emitter = PlanEventEmitter() @emitter.on_completed async def announce(event): print('done:', event.item.content) store = InMemoryPlanStore(event_emitter=emitter) Events come from granular tools (`add_task`, `update_task_status`, `add_subtask`, …). `write_plan` is a bulk whole-plan replacement and is **event-silent**, so a UI driven purely by events should also read the plan after a run, or steer the model toward granular tools when it needs live event coverage. Why whole-plan replacement -------------------------- [](https://pydantic.dev/docs/ai/harness/planning/#why-whole-plan-replacement) Addressing steps by mutable integer index (insert/remove/reorder) is error-prone for both the code and the model. `write_plan` restates the whole plan each call, so there are no indices to track. Granular edits (`add_task`, `update_task_status`, `remove_task`) instead reference the stable `id` shown by `read_plan`. Caching guarantee ----------------- [](https://pydantic.dev/docs/ai/harness/planning/#caching-guarantee) The plan is never injected into the system prompt or instructions. Static usage guidance goes there (cache-stable); only the mutable plan rides the ephemeral tail reminder, which lives solely in the per-request copy and is never persisted. Set `inject=False` to disable it. Pydantic AI maps `CachePoint` for models whose profiles support prompt caching; on other models it is ignored. Configuration ------------- [](https://pydantic.dev/docs/ai/harness/planning/#configuration) from pydantic_ai_harness.planning import Planning Planning( guidance=None, # static system-prompt guidance; None = default, '' = omit cache_ttl='5m', # TTL for the cache breakpoint after the stable opening tag ('5m' | '1h') store=None, # None = fresh in-memory plan per run; or a PlanStore to persist enable_subtasks=False, # add subtask/dependency tools and the 'blocked' status inject=True, # surface the current plan as a cache-safe tail reminder tools=None, # None = every tool the mode registers; or an allowlist of names descriptions=None, # optional per-tool description overrides, keyed by tool name ) Agent spec (YAML/JSON) ---------------------- [](https://pydantic.dev/docs/ai/harness/planning/#agent-spec-yamljson) `Planning` works with Pydantic AI’s [agent spec](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) : # agent.yaml model: anthropic:claude-sonnet-4-6 capabilities: - Planning: {} from pydantic_ai import Agent from pydantic_ai_harness.planning import Planning agent = Agent.from_file('agent.yaml', custom_capability_types=[Planning]) result = agent.run_sync('...') print(result.output) Further reading --------------- [](https://pydantic.dev/docs/ai/harness/planning/#further-reading) * [Pydantic AI capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) * [Anthropic prompt caching](https://docs.claude.com/en/docs/build-with-claude/prompt-caching) * [Code Mode](https://pydantic.dev/docs/ai/harness/code-mode/) — another prompt-cache-aware harness capability API reference ------------- [](https://pydantic.dev/docs/ai/harness/planning/#api-reference) Planning -------- [](https://pydantic.dev/docs/ai/harness/planning/#pydantic_ai_harness.planning.Planning) **Bases:** `AbstractCapability[AgentDepsT]` Structured task planning that never invalidates the prompt cache. The model owns the plan through a small toolset (`write_plan`, `read_plan`, `add_task`, `update_task_status`, `update_task_statuses`, `remove_task`, and — when `enable_subtasks` is set — `add_subtask`, `set_dependency`, `get_available_tasks`); `tools` narrows that surface to an allowlist. The current plan is surfaced back as an _ephemeral_ reminder appended to the tail of each request. Its cache-stable opening tag precedes a `CachePoint`, so the cached prefix stays byte-identical across turns; only the mutable plan content is re-read each turn. By default the plan lives in memory for the duration of a single run (a fresh, isolated plan per run). Pass a `store` (or `store_resolver`) to persist it — e.g. `SqlitePlanStore` or `PostgresPlanStore`. from pydantic_ai import Agent from pydantic_ai_harness.planning import Planning agent = Agent('anthropic:claude-sonnet-4-6', capabilities=[Planning()]) ### Attributes [](https://pydantic.dev/docs/ai/harness/planning/#attributes) #### cache\_ttl [](https://pydantic.dev/docs/ai/harness/planning/#pydantic_ai_harness.planning.Planning.cache_ttl) TTL for the cache breakpoint placed after the stable plan-reminder opening tag. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘5m’, ‘1h’\] **Default:** `'5m'` #### descriptions [](https://pydantic.dev/docs/ai/harness/planning/#pydantic_ai_harness.planning.Planning.descriptions) Optional per-tool description overrides, keyed by tool name. Unknown names raise `ValueError`. **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### enable\_subtasks [](https://pydantic.dev/docs/ai/harness/planning/#pydantic_ai_harness.planning.Planning.enable_subtasks) Add the subtask/dependency tools and the `blocked` status when true. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### guidance [](https://pydantic.dev/docs/ai/harness/planning/#pydantic_ai_harness.planning.Planning.guidance) Static planning guidance for the system prompt. Cache-stable. Three states, so opting out is something you do on purpose rather than by accident: * `None` (the default): use the built-in guidance. * `''`: no guidance at all. * any other string: use it instead of the built-in guidance. A single `str | None` where `None` meant “no guidance” would leave no way to ask for the default explicitly, and would turn a config that resolves to `None` into a silent opt-out. This matches `memory`, `exa` and `runtime_authoring`, which read the same way. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### inject [](https://pydantic.dev/docs/ai/harness/planning/#pydantic_ai_harness.planning.Planning.inject) Surface the current plan as a cache-safe tail reminder each turn. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `True` #### store [](https://pydantic.dev/docs/ai/harness/planning/#pydantic_ai_harness.planning.Planning.store) Storage backend. `None` keeps a fresh in-memory plan per run (the original ephemeral behaviour). Pass a store to persist the plan across runs. **Type:** `PlanStore` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### store\_resolver [](https://pydantic.dev/docs/ai/harness/planning/#pydantic_ai_harness.planning.Planning.store_resolver) Optional per-run store resolver, e.g. `lambda ctx: ctx.deps.plan_store`. **Type:** [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[[`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext)\ \[`AgentDepsT`\]\], `PlanStore`\] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### tools [](https://pydantic.dev/docs/ai/harness/planning/#pydantic_ai_harness.planning.Planning.tools) Optional allowlist of tool names to register; `None` registers all of them. The full surface is `write_plan`, `read_plan`, `add_task`, `update_task_status`, `update_task_statuses`, `remove_task`, plus `add_subtask`, `set_dependency` and `get_available_tasks` under `enable_subtasks`. `tools=['write_plan']` is the smallest useful plan surface. Naming a tool this mode does not register raises `ValueError`. The built-in `guidance` follows the allowlist for the whole-plan/granular/subtask split; trimming within a group is better paired with a `guidance` string of your own. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` ### Methods [](https://pydantic.dev/docs/ai/harness/planning/#methods) #### for\_run [](https://pydantic.dev/docs/ai/harness/planning/#pydantic_ai_harness.planning.Planning.for_run) `@async` def for_run(ctx: RunContext[AgentDepsT]) -> Planning[AgentDepsT] Return a clone with this run’s store resolved and cached (per-run isolation). ##### Returns [](https://pydantic.dev/docs/ai/harness/planning/#returns) `Planning`\[`AgentDepsT`\] #### from\_spec [](https://pydantic.dev/docs/ai/harness/planning/#pydantic_ai_harness.planning.Planning.from_spec) `@classmethod` def from_spec( cls, *, backend: Literal['memory', 'sqlite'] = 'memory', database: str = '.agent-plan.db', session: str = 'default', enable_subtasks: bool = False, inject: bool = True, guidance: str | None = None, cache_ttl: Literal['5m', '1h'] = '5m', tools: list[str] | None = None, ) -> Planning[AgentDepsT] Construct a `Planning` capability from serializable options. ##### Returns [](https://pydantic.dev/docs/ai/harness/planning/#returns-1) `Planning`\[`AgentDepsT`\] #### get\_instructions [](https://pydantic.dev/docs/ai/harness/planning/#pydantic_ai_harness.planning.Planning.get_instructions) def get_instructions() -> AgentInstructions[AgentDepsT] | None Provide static, cache-stable guidance on using the planning tools. A custom `guidance` string is used verbatim. The default is assembled from the tools actually registered — the granular sentence is dropped when `tools` excludes them all, and the subtask/dependency workflow is added under `enable_subtasks` — so the model is not told about tools it lacks. ##### Returns [](https://pydantic.dev/docs/ai/harness/planning/#returns-2) `AgentInstructions`\[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) #### get\_serialization\_name [](https://pydantic.dev/docs/ai/harness/planning/#pydantic_ai_harness.planning.Planning.get_serialization_name) `@classmethod` def get_serialization_name(cls) -> str | None Serialization name for agent-spec support. ##### Returns [](https://pydantic.dev/docs/ai/harness/planning/#returns-3) [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### get\_toolset [](https://pydantic.dev/docs/ai/harness/planning/#pydantic_ai_harness.planning.Planning.get_toolset) def get_toolset() -> AgentToolset[AgentDepsT] | None Provide the `planning` toolset over this run’s resolved store. ##### Returns [](https://pydantic.dev/docs/ai/harness/planning/#returns-4) [`AgentToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AgentToolset) \[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) #### resolve\_store [](https://pydantic.dev/docs/ai/harness/planning/#pydantic_ai_harness.planning.Planning.resolve_store) def resolve_store(ctx: RunContext[AgentDepsT]) -> PlanStore Return the cached run store, or resolve one for direct toolset use. ##### Returns [](https://pydantic.dev/docs/ai/harness/planning/#returns-5) `PlanStore` #### wrap\_model\_request [](https://pydantic.dev/docs/ai/harness/planning/#pydantic_ai_harness.planning.Planning.wrap_model_request) `@async` def wrap_model_request( ctx: RunContext[AgentDepsT], *, request_context: ModelRequestContext, handler: WrapModelRequestHandler, ) -> ModelResponse Append the current plan as an ephemeral tail reminder with a cache breakpoint. ##### Returns [](https://pydantic.dev/docs/ai/harness/planning/#returns-6) [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) Was this page helpful? Thanks for your feedback! --- # pydantic_graph.decision | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#_top) pydantic\_graph.decision ======================== Decision node implementation for conditional branching in graph execution. This module provides the Decision node type and related classes for implementing conditional branching logic in parallel control flow graphs. Decision nodes allow the graph to choose different execution paths based on runtime conditions. Decision -------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.Decision) **Bases:** `Generic[StateT, DepsT, HandledT]` Decision node for conditional branching in graph execution. A Decision node evaluates conditions and routes execution to different branches based on the input data type or custom matching logic. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#attributes) #### branches [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.Decision.branches) List of branches that can be taken from this decision. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[`DecisionBranch`\[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \]\] #### id [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.Decision.id) Unique identifier for this decision node. **Type:** `NodeID` #### note [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.Decision.note) Optional documentation note for this decision. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) ### Methods [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#methods) #### branch [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.Decision.branch) def branch(branch: DecisionBranch[T]) -> Decision[StateT, DepsT, HandledT | T] Add a new branch to this decision. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#returns) `Decision`\[`StateT`, `DepsT`, `HandledT` | `T`\] — A new Decision with the additional branch. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#parameters) **`branch`** : `DecisionBranch`\[`T`\] [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.Decision.branch(branch)) The branch to add to this decision. DecisionBranch -------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranch) **Bases:** `Generic[SourceT]` Represents a single branch within a decision node. Each branch defines the conditions under which it should be taken and the path to follow when those conditions are met. Note: with the current design, it is actually _critical_ that this class is invariant in SourceT for the sake of type-checking that inputs to a Decision are actually handled. See the `# type: ignore` comment in `tests.graph.builder.test_graph_edge_cases.test_decision_no_matching_branch` for an example of how this works. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#attributes-1) #### destinations [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranch.destinations) The destination nodes that can be referenced by DestinationMarker in the path. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[`AnyDestinationNode`\] #### matches [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranch.matches) An optional predicate function used to determine whether input data matches this branch. If `None`, default logic is used which attempts to check the value for type-compatibility with the `source` type: * If `source` is `Any` or `object`, the branch will always match * If `source` is a `Literal` type, this branch will match if the value is one of the parametrizing literal values * If `source` is any other type, the value will be checked for matching using `isinstance` Inputs are tested against each branch of a decision node in order, and the path of the first matching branch is used to handle the input value. **Type:** [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \], [`bool`](https://docs.python.org/3/library/functions.html#bool)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) #### path [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranch.path) The execution path to follow when an input value matches this branch of a decision node. This can include transforming, mapping, and broadcasting the output before sending to the next node or nodes. The path can also include position-aware labels which are used when generating mermaid diagrams. **Type:** `Path` #### source [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranch.source) The expected type of data for this branch. This is necessary for exhaustiveness-checking when handling the inputs to a decision node. **Type:** `TypeOrTypeExpression`\[`SourceT`\] DecisionBranchBuilder --------------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranchBuilder) **Bases:** `Generic[StateT, DepsT, OutputT, SourceT, HandledT]` Builder for constructing decision branches with fluent API. This builder provides methods to configure branches with destinations, forks, and transformations in a type-safe manner. Instances of this class should be created using [`GraphBuilder.match`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder) , not created directly. ### Methods [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#methods-1) #### broadcast [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranchBuilder.broadcast) def broadcast( get_forks: Callable[[Self], Sequence[DecisionBranch[SourceT]]], /, *, fork_id: str | None = None, ) -> DecisionBranch[SourceT] Broadcast this decision branch into multiple destinations. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#returns-1) `DecisionBranch`\[`SourceT`\] — A completed DecisionBranch with the specified destinations. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#parameters-1) **`get_forks`** : [`Callable`](https://docs.python.org/3/library/typing.html#typing.Callable) \[\[[`Self`](https://docs.python.org/3/library/typing.html#typing.Self)\ \], [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence)\ \[`DecisionBranch`\[`SourceT`\]\]\] [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranchBuilder.broadcast(get_forks)) The callback that will return a sequence of decision branches to broadcast to. **`fork_id`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranchBuilder.broadcast(fork_id)) Optional node ID to use for the resulting broadcast fork. #### label [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranchBuilder.label) def label( label: str, ) -> DecisionBranchBuilder[StateT, DepsT, OutputT, SourceT, HandledT] Apply a label to the branch at the current point in the path being built. These labels are only used in generated mermaid diagrams. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#returns-2) `DecisionBranchBuilder`\[`StateT`, `DepsT`, `OutputT`, `SourceT`, `HandledT`\] — A new DecisionBranchBuilder where the label has been applied at the end of the current path being built. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#parameters-2) **`label`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranchBuilder.label(label)) The label to apply. #### map [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranchBuilder.map) def map( *, fork_id: str | None = None, downstream_join_id: str | None = None, ) -> DecisionBranchBuilder[StateT, DepsT, T, SourceT, HandledT] Spread the branch’s output. To do this, the current output must be iterable, and any subsequent steps in the path being built for this branch will be applied to each item of the current output in parallel. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#returns-3) `DecisionBranchBuilder`\[`StateT`, `DepsT`, `T`, `SourceT`, `HandledT`\] — A new DecisionBranchBuilder where mapping is performed prior to generating the final output. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#parameters-3) **`fork_id`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranchBuilder.map(fork_id)) Optional ID for the fork, defaults to a generated value **`downstream_join_id`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranchBuilder.map(downstream_join_id)) Optional ID of a downstream join node which is involved when mapping empty iterables #### to [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranchBuilder.to) def to( destination: DestinationNode[StateT, DepsT, OutputT] | type[BaseNode[StateT, DepsT, Any]], /, *extra_destinations: DestinationNode[StateT, DepsT, OutputT] | type[BaseNode[StateT, DepsT, Any]], fork_id: str | None = None, ) -> DecisionBranch[SourceT] Set the destination(s) for this branch. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#returns-4) `DecisionBranch`\[`SourceT`\] — A completed DecisionBranch with the specified destinations. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#parameters-4) **`destination`** : `DestinationNode`\[`StateT`, `DepsT`, `OutputT`\] | [`type`](https://docs.python.org/3/glossary.html#term-type) \[`BaseNode`\[`StateT`, `DepsT`, [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \]\] [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranchBuilder.to(destination)) The primary destination node. **`*extra_destinations`** : `DestinationNode`\[`StateT`, `DepsT`, `OutputT`\] | [`type`](https://docs.python.org/3/glossary.html#term-type) \[`BaseNode`\[`StateT`, `DepsT`, [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \]\] _Default:_ `()` [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranchBuilder.to(*extra_destinations)) Additional destination nodes. **`fork_id`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranchBuilder.to(fork_id)) Optional node ID to use for the resulting broadcast fork if multiple destinations are provided. #### transform [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranchBuilder.transform) def transform( func: TransformFunction[StateT, DepsT, OutputT, NewOutputT], /, ) -> DecisionBranchBuilder[StateT, DepsT, NewOutputT, SourceT, HandledT] Apply a transformation to the branch’s output. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#returns-5) `DecisionBranchBuilder`\[`StateT`, `DepsT`, `NewOutputT`, `SourceT`, `HandledT`\] — A new DecisionBranchBuilder where the provided transform is applied prior to generating the final output. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#parameters-5) **`func`** : `TransformFunction`\[`StateT`, `DepsT`, `OutputT`, `NewOutputT`\] [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DecisionBranchBuilder.transform(func)) Transformation function to apply. DepsT ----- [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.DepsT) Type variable for graph dependencies. **Default:** `TypeVar('DepsT', infer_variance=True)` HandledT -------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.HandledT) Type variable used to track types handled by the branches of a Decision. **Default:** `TypeVar('HandledT', infer_variance=True)` NewOutputT ---------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.NewOutputT) Type variable for transformed output. **Default:** `TypeVar('NewOutputT', infer_variance=True)` OutputT ------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.OutputT) Type variable for the output data of a node. **Default:** `TypeVar('OutputT', infer_variance=True)` SourceT ------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.SourceT) Type variable for source data for a DecisionBranch. **Default:** `TypeVar('SourceT', infer_variance=True)` StateT ------ [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.StateT) Type variable for graph state. **Default:** `TypeVar('StateT', infer_variance=True)` T - [](https://pydantic.dev/docs/ai/api/pydantic_graph/decision/#pydantic_graph.decision.T) Generic type variable. **Default:** `TypeVar('T', infer_variance=True)` Was this page helpful? Thanks for your feedback! --- # exceptions | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#_top) exceptions ========== AgentRunError ------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.AgentRunError) **Bases:** [`RuntimeError`](https://docs.python.org/3/library/exceptions.html#RuntimeError) Base class for errors occurring during an agent run. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#attributes) #### message [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.AgentRunError.message) The error message. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `message` ApprovalRequired ---------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ApprovalRequired) **Bases:** [`Exception`](https://docs.python.org/3/library/exceptions.html#Exception) Exception to raise when a tool call requires human-in-the-loop approval. See [tools docs](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/#human-in-the-loop-tool-approval) for more information. ### Constructor Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#constructor-parameters) **`metadata`** : [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ApprovalRequired.__init__(metadata)) Optional dictionary of metadata to attach to the deferred tool call. This metadata will be available in `DeferredToolRequests.metadata` keyed by `tool_call_id`. CallDeferred ------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.CallDeferred) **Bases:** [`Exception`](https://docs.python.org/3/library/exceptions.html#Exception) Exception to raise when a tool call should be deferred. See [tools docs](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/#deferred-tools) for more information. ### Constructor Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#constructor-parameters-1) **`metadata`** : [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.CallDeferred.__init__(metadata)) Optional dictionary of metadata to attach to the deferred tool call. This metadata will be available in `DeferredToolRequests.metadata` keyed by `tool_call_id`. ConcurrencyLimitExceeded ------------------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ConcurrencyLimitExceeded) **Bases:** [`AgentRunError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.AgentRunError) Error raised when the concurrency queue depth exceeds max\_queued. ContentFilterError ------------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ContentFilterError) **Bases:** [`UnexpectedModelBehavior`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UnexpectedModelBehavior) Raised when content filtering is triggered by the model provider. CostCalculationFailedWarning ---------------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.CostCalculationFailedWarning) **Bases:** [`Warning`](https://docs.python.org/3/library/exceptions.html#Warning) Warning raised when cost calculation fails. CostNotFoundWarning ------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.CostNotFoundWarning) **Bases:** [`Warning`](https://docs.python.org/3/library/exceptions.html#Warning) Warning raised when cost is not found. FallbackExceptionGroup ---------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.FallbackExceptionGroup) **Bases:** `ExceptionGroup[Any]` A group of exceptions that can be raised when all fallback models fail. IncompleteToolCall ------------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.IncompleteToolCall) **Bases:** [`UnexpectedModelBehavior`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UnexpectedModelBehavior) Error raised when a model stops due to token limit while emitting a tool call. MessageHistoryMutatedWarning ---------------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.MessageHistoryMutatedWarning) **Bases:** [`Warning`](https://docs.python.org/3/library/exceptions.html#Warning) Warning raised when in-place mutation of the message history is detected at the end of a run. Mutating messages that are already part of the run’s history in place (e.g. `ctx.messages[0].parts[0].content = '...'` from a tool) is not supported: the per-request `gen_ai.input.messages` span attribute caches each message’s serialized form, so spans recorded after the mutation may not match the messages actually sent to the model. The run-level `pydantic_ai.all_messages` attribute is always serialized fresh and does reflect the mutation. To transform history mid-run, build new message or part objects instead — e.g. with `dataclasses.replace`, passing the message a new `parts` list (replacing a message in the history and reassigning its `parts` list are both safe) — for instance in a history processor ([`ProcessHistory`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ProcessHistory) ). The warning is best-effort: it’s raised when a mutation is detected at the end of a successful run, which covers messages still present in the final history. Errored runs aren’t checked — with warnings configured as errors, the warning would displace the run’s own exception. Its absence does not guarantee that no stale span was recorded. ModelAPIError ------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelAPIError) **Bases:** [`AgentRunError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.AgentRunError) Raised when a model provider API request fails. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#attributes-1) #### model\_name [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelAPIError.model_name) The name of the model associated with the error. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `model_name` ModelHTTPError -------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelHTTPError) **Bases:** [`ModelAPIError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelAPIError) Raised when a model provider response has a status code of 4xx or 5xx. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#attributes-2) #### body [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelHTTPError.body) The body of the response, if available. **Type:** [`object`](https://docs.python.org/3/glossary.html#term-object) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `body` #### headers [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelHTTPError.headers) Response headers from the provider, with keys lowercased for consistent access. For example, use `exc.headers.get('retry-after')` to read the `Retry-After` header regardless of provider casing. `None` when the provider does not supply headers (e.g. gRPC-based providers or synthesised errors). **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `{(k.lower()): v for k, v in (headers.items())} if headers is not None else None` #### retry\_after [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelHTTPError.retry_after) Seconds to wait before retrying, parsed from the `Retry-After` response header. Returns `None` when the header is absent or cannot be parsed. The header value is interpreted first as an integer number of seconds, then as an [HTTP-date](https://httpwg.org/specs/rfc9110.html#http.date) string. **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) | [`None`](https://docs.python.org/3/library/constants.html#None) #### status\_code [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelHTTPError.status_code) The HTTP status code returned by the API. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `status_code` ModelRetry ---------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelRetry) **Bases:** [`Exception`](https://docs.python.org/3/library/exceptions.html#Exception) Exception to raise to request a model retry. Can be raised from tool functions, output validators, and capability hooks (such as `after_model_request`, `after_tool_execute`, etc.) to send a retry prompt back to the model asking it to try again. For a terminal failure the model should see but not retry, raise [`ToolFailed`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ToolFailed) instead. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#attributes-3) #### message [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelRetry.message) The message to return to the model. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `message` ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#methods) #### \_\_get\_pydantic\_core\_schema\_\_ [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelRetry.__get_pydantic_core_schema__) `@classmethod` def __get_pydantic_core_schema__(cls, _: Any, __: Any) -> core_schema.CoreSchema Pydantic core schema to allow `ModelRetry` to be (de)serialized. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#returns) `core_schema.CoreSchema` PydanticAIDeprecationWarning ---------------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.PydanticAIDeprecationWarning) **Bases:** [`UserWarning`](https://docs.python.org/3/library/exceptions.html#UserWarning) Warning emitted when a deprecated Pydantic AI API is used. Inherits from `UserWarning` instead of `DeprecationWarning` so that deprecations are visible by default at runtime, following the approach described in [https://sethmlarson.dev/deprecations-via-warnings-dont-work-for-python-libraries](https://sethmlarson.dev/deprecations-via-warnings-dont-work-for-python-libraries) . RunCancelled ------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled) **Bases:** [`AgentRunError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.AgentRunError) Raised when the agent run was cancelled by the application itself. Raised by [`AgentRun.cancel()`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.cancel) and [`RunContext.cancel()`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.cancel) . This is a normal, catchable application-level outcome: the run stopped because your own code asked it to. External cancellation of the task running the agent (`asyncio.Task.cancel()`, a timeout scope, workflow cancellation under durable execution) is infrastructure-level and keeps propagating as `asyncio.CancelledError` instead — it is never translated into this exception, and when both race, the external cancellation wins. (On Python 3.10, which lacks `Task.uncancel()`, the race cannot be disambiguated and a requested first-party cancellation wins instead.) Everything the run completed before the cancellation took effect — including the partial response of an interrupted stream and the results of tool calls that finished — is preserved in [`all_messages()`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled.all_messages) : pass it as `message_history` to a new run (with a new user prompt or not) to resume the conversation; any tool calls that never produced a result are automatically closed out with synthesized `outcome='interrupted'` returns before the history is sent to a model. Cancellation is terminal: capability hooks (`wrap_run`, `wrap_node_run`, `on_run_error`) may observe it and clean up, but cannot recover a cancelled run into a successful result. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#attributes-4) #### conversation\_id [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled.conversation_id) The conversation identifier, or `None` if the run was cancelled before starting. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### metadata [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled.metadata) Metadata associated with this agent run, if configured. **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) #### response [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled.response) Return the last response from the message history. **Type:** [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) #### run\_id [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled.run_id) The unique identifier for the agent run, or `None` if it was cancelled before starting. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### timestamp [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled.timestamp) Return the timestamp of the last response. **Type:** [`datetime`](https://docs.python.org/3/library/datetime.html#module-datetime) #### usage [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled.usage) Return the usage of the cancelled run. **Type:** [`RunUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage) ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#methods-1) #### all\_messages [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled.all_messages) def all_messages() -> list[ModelMessage] Return the complete resumable history of the cancelled run. This is a DETACHED snapshot of the run’s message history at termination, ready to pass as `message_history` for a resumed run. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#returns-1) [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage)\ \] — List of messages. #### all\_messages\_json [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled.all_messages_json) def all_messages_json() -> bytes Return all messages from [`all_messages`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled.all_messages) as JSON bytes. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#returns-2) [`bytes`](https://docs.python.org/3/library/stdtypes.html#bytes) — JSON bytes representing the messages. #### from\_cancellation [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled.from_cancellation) `@classmethod` def from_cancellation(cls, exc: BaseException) -> RunCancelled | None Recover run state from a cancellation-related exception. External cancellation of a plain `agent.run()` keeps its standard asyncio semantics. Catch it with `except asyncio.CancelledError as exc`, then call `RunCancelled.from_cancellation(exc)` to access the partial run state attached by Pydantic AI. This also works with the `TimeoutError` raised by `asyncio.timeout()` or `asyncio.wait_for()`, whose exception chain contains the original `CancelledError`. An external `CancelledError` must keep propagating for timeouts and task groups to tear down correctly, so re-raise it after capturing the state rather than returning from the handler; only a first-party `RunCancelled` is yours to consume. Passing a `RunCancelled` directly returns the same instance, providing uniform handling for first-party and external cancellation paths. Python 3.11+ preserves the exception instance across an `await task` boundary. Python 3.10 recreates the `CancelledError` there, but chains the original exception — and the attached run state — via `__context__`, which this method traverses; the chain is attached only to the first `await` of the cancelled task, so later awaits of the same task see an unchained exception. Use `capture_run_messages()` as the fallback when only message history is needed. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#returns-3) [`RunCancelled`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled) | [`None`](https://docs.python.org/3/library/constants.html#None) #### new\_messages [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled.new_messages) def new_messages() -> list[ModelMessage] Return the messages produced during the cancelled run. Messages provided via `message_history` and messages from older runs are excluded. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#returns-4) [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage)\ \] — List of new messages. #### new\_messages\_json [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled.new_messages_json) def new_messages_json() -> bytes Return new messages from [`new_messages`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled.new_messages) as JSON bytes. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#returns-5) [`bytes`](https://docs.python.org/3/library/stdtypes.html#bytes) — JSON bytes representing the new messages. SkipModelRequest ---------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.SkipModelRequest) **Bases:** [`Exception`](https://docs.python.org/3/library/exceptions.html#Exception) Exception to raise in before/wrap model request hooks to skip the model call. The provided response will be used instead of calling the model. Note: when raised in `before_model_request`, any message history modifications made by earlier capabilities in that hook will not be persisted to the agent’s message history, since the request preparation is aborted. SkipToolExecution ----------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.SkipToolExecution) **Bases:** [`Exception`](https://docs.python.org/3/library/exceptions.html#Exception) Exception to raise in before/wrap tool execute hooks to skip execution. The provided result will be used as the tool result. SkipToolValidation ------------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.SkipToolValidation) **Bases:** [`Exception`](https://docs.python.org/3/library/exceptions.html#Exception) Exception to raise in before/wrap tool validate hooks to skip validation. The provided args will be used as the validated arguments. SuspendedResponseExpired ------------------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.SuspendedResponseExpired) **Bases:** [`AgentRunError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.AgentRunError) Raised when resuming a suspended response whose server-side job is no longer available. Suspended/background jobs are only resumable within the provider’s retention window (e.g. ~10 minutes for OpenAI background mode). Resuming a persisted suspended response after that window raises this instead of an opaque provider HTTP error; start a new run from the preceding messages to retry from scratch. ToolFailed ---------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ToolFailed) **Bases:** [`Exception`](https://docs.python.org/3/library/exceptions.html#Exception) Exception to raise to report a terminal tool failure to the model. Raise this when a tool call is done and has failed — a missing resource, an unsupported operation, a definitive upstream error — and you want the model to see the failure and adapt rather than try the same call again. Can be raised from tool functions, args validators, and tool validation/execution hooks. Like [`ModelRetry`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelRetry) , this produces a failed tool result the model sees; unlike `ModelRetry` it does not prepend retry/correction instructions and does not consume the tool’s retry budget. Bound repeated failures with [`UsageLimits`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits) at the run level instead. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#attributes-5) #### message [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ToolFailed.message) The failure message to return to the model. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `message` ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#methods-2) #### \_\_get\_pydantic\_core\_schema\_\_ [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ToolFailed.__get_pydantic_core_schema__) `@classmethod` def __get_pydantic_core_schema__(cls, _: Any, __: Any) -> core_schema.CoreSchema Pydantic core schema to allow `ToolFailed` to be (de)serialized. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#returns-6) `core_schema.CoreSchema` ToolFailedError --------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ToolFailedError) **Bases:** [`Exception`](https://docs.python.org/3/library/exceptions.html#Exception) Exception used to signal a failed `ToolReturnPart` should be returned to the LLM. ToolRetryError -------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ToolRetryError) **Bases:** [`Exception`](https://docs.python.org/3/library/exceptions.html#Exception) Exception used to signal a `ToolRetry` message should be returned to the LLM. UndrainedPendingMessagesError ----------------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UndrainedPendingMessagesError) **Bases:** [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) Error that used to be raised when an agent run ended with messages still queued via `enqueue`. A bare `async for node in agent_run` loop used to skip the node hooks, so `'when_idle'` messages and end-of-run redirects (which drain in `after_node_run`) were stranded. Bare iteration now advances through [`AgentRun.next()`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.next) like every other way of driving a run, so pending messages always drain and this error is no longer raised. It is kept so existing `except` clauses keep working. UnexpectedModelBehavior ----------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UnexpectedModelBehavior) **Bases:** [`AgentRunError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.AgentRunError) Error caused by unexpected Model behavior, e.g. an unexpected response code. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#attributes-6) #### body [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UnexpectedModelBehavior.body) The body of the response, if available. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `json.dumps(json.loads(body), indent=2)` #### message [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UnexpectedModelBehavior.message) Description of the unexpected behavior. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `message` UsageLimitExceeded ------------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UsageLimitExceeded) **Bases:** [`AgentRunError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.AgentRunError) Error raised when a Model’s usage exceeds the specified limits. UserError --------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) **Bases:** [`RuntimeError`](https://docs.python.org/3/library/exceptions.html#RuntimeError) Error caused by a usage mistake by the application developer — You! ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#attributes-7) #### message [](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError.message) Description of the mistake. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `message` Was this page helpful? Thanks for your feedback! --- # Metrics & Attributes | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#_top) Metrics & Attributes ==================== Track custom metrics and attributes during task execution for richer evaluation insights. While executing evaluation tasks, you can record: * **Metrics** - Numeric values (int/float) for quantitative measurements * **Attributes** - Any data for qualitative information These appear in evaluation reports and can be used by evaluators for assessment. Recording Metrics ----------------- [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#recording-metrics) Use [`increment_eval_metric`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.increment_eval_metric) to track numeric values: from dataclasses import dataclass from pydantic_evals.dataset import increment_eval_metric @dataclass class APIResult: output: str usage: 'Usage' @dataclass class Usage: total_tokens: int def call_api(inputs: str) -> APIResult: return APIResult(output=f'Result: {inputs}', usage=Usage(total_tokens=100)) def my_task(inputs: str) -> str: # Track API calls increment_eval_metric('api_calls', 1) result = call_api(inputs) # Track tokens used increment_eval_metric('tokens_used', result.usage.total_tokens) return result.output Recording Attributes -------------------- [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#recording-attributes) Use [`set_eval_attribute`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.set_eval_attribute) to store any data: from pydantic_evals import set_eval_attribute def process(inputs: str) -> str: return f'Processed: {inputs}' def my_task(inputs: str) -> str: # Record which model was used set_eval_attribute('model', 'gpt-5.2') # Record feature flags set_eval_attribute('used_cache', True) set_eval_attribute('retry_count', 2) # Record structured data set_eval_attribute('config', { 'temperature': 0.7, 'max_tokens': 100, }) return process(inputs) Accessing in Evaluators ----------------------- [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#accessing-in-evaluators) Metrics and attributes are available in the [`EvaluatorContext`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorContext) : from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class EfficiencyChecker(Evaluator): max_api_calls: int = 5 def evaluate(self, ctx: EvaluatorContext) -> dict[str, bool]: # Access metrics api_calls = ctx.metrics.get('api_calls', 0) tokens_used = ctx.metrics.get('tokens_used', 0) # Access attributes used_cache = ctx.attributes.get('used_cache', False) return { 'efficient_api_usage': api_calls <= self.max_api_calls, 'used_caching': used_cache, 'token_efficient': tokens_used < 1000, } Viewing in Reports ------------------ [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#viewing-in-reports) Metrics and attributes appear in report data: from pydantic_evals import Case, Dataset def task(inputs: str) -> str: return f'Result: {inputs}' dataset = Dataset(name='report_viewing', cases=[Case(inputs='test')], evaluators=[]) report = dataset.evaluate_sync(task) for case in report.cases: print(f'{case.name}:') #> Case 1: print(f' Metrics: {case.metrics}') #> Metrics: {} print(f' Attributes: {case.attributes}') #> Attributes: {} You can also display them in printed reports: from pydantic_evals import Case, Dataset def task(inputs: str) -> str: return f'Result: {inputs}' dataset = Dataset(name='report_printing', cases=[Case(inputs='test')], evaluators=[]) report = dataset.evaluate_sync(task) # Metrics and attributes are available but not shown by default # Access them programmatically or via Logfire for case in report.cases: print(f'\nCase: {case.name}') """ Case: Case 1 """ print(f'Metrics: {case.metrics}') #> Metrics: {} print(f'Attributes: {case.attributes}') #> Attributes: {} Automatic Metrics ----------------- [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#automatic-metrics) When using Pydantic AI and Logfire, some metrics are automatically tracked: import logfire from pydantic_ai import Agent logfire.configure(send_to_logfire='if-token-present') agent = Agent('openai:gpt-5.2') async def ai_task(inputs: str) -> str: result = await agent.run(inputs) return result.output # Automatically tracked metrics: # - requests: Number of LLM calls # - input_tokens: Total input tokens # - output_tokens: Total output tokens # - prompt_tokens: Prompt tokens (if available) # - completion_tokens: Completion tokens (if available) # - cost: Estimated cost (if using genai-prices) Access these in evaluators: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class CostChecker(Evaluator): max_cost: float = 0.01 # $0.01 def evaluate(self, ctx: EvaluatorContext) -> bool: cost = ctx.metrics.get('cost', 0.0) return cost <= self.max_cost Practical Examples ------------------ [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#practical-examples) ### API Usage Tracking [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#api-usage-tracking) from dataclasses import dataclass from pydantic_evals import increment_eval_metric, set_eval_attribute from pydantic_evals.evaluators import Evaluator, EvaluatorContext def check_cache(inputs: str) -> str | None: return None # No cache hit for demo @dataclass class APIResult: text: str usage: 'Usage' @dataclass class Usage: total_tokens: int async def call_api(inputs: str) -> APIResult: return APIResult(text=f'Result: {inputs}', usage=Usage(total_tokens=100)) def save_to_cache(inputs: str, result: str) -> None: pass # Save to cache async def smart_task(inputs: str) -> str: # Try cache first if cached := check_cache(inputs): set_eval_attribute('cache_hit', True) return cached set_eval_attribute('cache_hit', False) # Call API increment_eval_metric('api_calls', 1) result = await call_api(inputs) increment_eval_metric('tokens', result.usage.total_tokens) # Cache result save_to_cache(inputs, result.text) return result.text # Evaluate efficiency @dataclass class EfficiencyEvaluator(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> dict[str, bool | float]: api_calls = ctx.metrics.get('api_calls', 0) cache_hit = ctx.attributes.get('cache_hit', False) return { 'used_cache': cache_hit, 'made_api_call': api_calls > 0, 'efficiency_score': 1.0 if cache_hit else 0.5, } ### Tool Usage Tracking [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#tool-usage-tracking) from dataclasses import dataclass from pydantic_ai import Agent, RunContext from pydantic_evals import increment_eval_metric, set_eval_attribute from pydantic_evals.evaluators import Evaluator, EvaluatorContext agent = Agent('openai:gpt-5.2') def search(query: str) -> str: return f'Search results for: {query}' def call(endpoint: str) -> str: return f'API response from: {endpoint}' @agent.tool def search_database(ctx: RunContext, query: str) -> str: increment_eval_metric('db_searches', 1) set_eval_attribute('last_query', query) return search(query) @agent.tool def call_api(ctx: RunContext, endpoint: str) -> str: increment_eval_metric('api_calls', 1) set_eval_attribute('last_endpoint', endpoint) return call(endpoint) # Evaluate tool usage @dataclass class ToolUsageEvaluator(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> dict[str, bool | int]: db_searches = ctx.metrics.get('db_searches', 0) api_calls = ctx.metrics.get('api_calls', 0) return { 'used_database': db_searches > 0, 'used_api': api_calls > 0, 'tool_call_count': db_searches + api_calls, 'reasonable_tool_usage': (db_searches + api_calls) <= 5, } ### Performance Tracking [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#performance-tracking) import time from dataclasses import dataclass from pydantic_evals import increment_eval_metric, set_eval_attribute from pydantic_evals.evaluators import Evaluator, EvaluatorContext async def retrieve_context(inputs: str) -> list[str]: return ['context1', 'context2'] async def generate_response(context: list[str], inputs: str) -> str: return f'Generated response for {inputs}' async def monitored_task(inputs: str) -> str: # Track sub-operation timing t0 = time.perf_counter() context = await retrieve_context(inputs) retrieve_time = time.perf_counter() - t0 increment_eval_metric('retrieve_time', retrieve_time) t0 = time.perf_counter() result = await generate_response(context, inputs) generate_time = time.perf_counter() - t0 increment_eval_metric('generate_time', generate_time) # Record which operations were needed set_eval_attribute('needed_retrieval', len(context) > 0) set_eval_attribute('context_chunks', len(context)) return result # Evaluate performance @dataclass class PerformanceEvaluator(Evaluator): max_retrieve_time: float = 0.5 max_generate_time: float = 2.0 def evaluate(self, ctx: EvaluatorContext) -> dict[str, bool]: retrieve_time = ctx.metrics.get('retrieve_time', 0.0) generate_time = ctx.metrics.get('generate_time', 0.0) return { 'fast_retrieval': retrieve_time <= self.max_retrieve_time, 'fast_generation': generate_time <= self.max_generate_time, } ### Quality Tracking [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#quality-tracking) from dataclasses import dataclass from pydantic_evals import set_eval_attribute from pydantic_evals.evaluators import Evaluator, EvaluatorContext async def llm_call(inputs: str) -> dict: return {'text': f'Response: {inputs}', 'confidence': 0.85, 'sources': ['doc1', 'doc2']} async def quality_task(inputs: str) -> str: result = await llm_call(inputs) # Extract quality indicators confidence = result.get('confidence', 0.0) sources_used = result.get('sources', []) set_eval_attribute('confidence', confidence) set_eval_attribute('source_count', len(sources_used)) set_eval_attribute('sources', sources_used) return result['text'] # Evaluate based on quality signals @dataclass class QualityEvaluator(Evaluator): min_confidence: float = 0.7 def evaluate(self, ctx: EvaluatorContext) -> dict[str, bool | float]: confidence = ctx.attributes.get('confidence', 0.0) source_count = ctx.attributes.get('source_count', 0) return { 'high_confidence': confidence >= self.min_confidence, 'used_sources': source_count > 0, 'quality_score': confidence * (1.0 + 0.1 * source_count), } Experiment-Level Metadata ------------------------- [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#experiment-level-metadata) In addition to case-level metadata, you can also pass experiment-level metadata when calling [`evaluate()`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset.evaluate) : from pydantic_evals import Case, Dataset dataset = Dataset( name='experiment_metadata', cases=[\ Case(\ inputs='test',\ metadata={'difficulty': 'easy'}, # Case-level metadata\ )\ ] ) async def task(inputs: str) -> str: return f'Result: {inputs}' # Pass experiment-level metadata async def main(): report = await dataset.evaluate( task, metadata={ 'model': 'gpt-5.2', 'prompt_version': 'v2.1', 'temperature': 0.7, }, ) # Access experiment metadata in the report print(report.experiment_metadata) #> {'model': 'gpt-5.2', 'prompt_version': 'v2.1', 'temperature': 0.7} ### When to Use Experiment Metadata [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#when-to-use-experiment-metadata) Experiment metadata is useful for tracking configuration that applies to the entire evaluation run: * **Model configuration**: Model name, version, parameters * **Prompt versioning**: Which prompt template was used * **Infrastructure**: Deployment environment, region * **Experiment context**: Developer name, feature branch, commit hash This metadata is especially valuable when: * Comparing multiple evaluation runs over time * Tracking which configuration produced which results * Reproducing evaluation results from historical data ### Viewing in Reports [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#viewing-in-reports-1) Experiment metadata appears at the top of printed reports: from pydantic_evals import Case, Dataset dataset = Dataset(name='metadata_report', cases=[Case(inputs='hello', expected_output='HELLO')]) async def task(text: str) -> str: return text.upper() async def main(): report = await dataset.evaluate( task, metadata={'model': 'gpt-5.2', 'version': 'v1.0'}, ) print(report.render()) """ ╭─ Evaluation Summary: task ─╮ │ model: gpt-5.2 │ │ version: v1.0 │ ╰────────────────────────────╯ ┏━━━━━━━━━━┳━━━━━━━━━━┓ ┃ Case ID ┃ Duration ┃ ┡━━━━━━━━━━╇━━━━━━━━━━┩ │ Case 1 │ 10ms │ ├──────────┼──────────┤ │ Averages │ 10ms │ └──────────┴──────────┘ """ Synchronization between Tasks and Experiment Metadata ----------------------------------------------------- [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#synchronization-between-tasks-and-experiment-metadata) Experiment metadata is for _recording_ configuration, not _configuring_ the task. The metadata dict doesn’t automatically configure your task’s behavior; you must ensure the values in the metadata dict match what your task actually uses. For example, it’s easy to accidentally have metadata claim `temperature: 0.7` while your task actually uses `temperature: 1.0`, leading to incorrect experiment tracking and unreproducible results. To avoid this problem, we recommend establishing a single source of truth for configuration that both your task and metadata reference. Below are a few suggested patterns for achieving this synchronization. ### Pattern 1: Shared Module Constants [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#pattern-1-shared-module-constants) For simpler cases, use module-level constants: from pydantic_ai import Agent from pydantic_evals import Case, Dataset # Module constants as single source of truth MODEL_NAME = 'openai:gpt-5-mini' TEMPERATURE = 0.7 INSTRUCTIONS = 'You are a helpful assistant.' agent = Agent(MODEL_NAME, model_settings={'temperature': TEMPERATURE}, instructions=INSTRUCTIONS) async def task(inputs: str) -> str: result = await agent.run(inputs) return result.output async def main(): dataset = Dataset(name='shared_constants', cases=[Case(inputs='What is the capital of France?')]) # Metadata references same constants await dataset.evaluate( task, metadata={ 'model': MODEL_NAME, 'temperature': TEMPERATURE, 'instructions': INSTRUCTIONS, }, ) ### Pattern 2: Configuration Object (Recommended) [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#pattern-2-configuration-object-recommended) Define configuration once and use it everywhere: from dataclasses import asdict, dataclass from pydantic_ai import Agent from pydantic_evals import Case, Dataset @dataclass class TaskConfig: """Single source of truth for task configuration. Includes all variables you'd like to see in experiment metadata. """ model: str temperature: float max_tokens: int prompt_version: str # Define configuration once config = TaskConfig( model='openai:gpt-5-mini', temperature=0.7, max_tokens=500, prompt_version='v2.1', ) # Use config in task agent = Agent( config.model, model_settings={'temperature': config.temperature, 'max_tokens': config.max_tokens}, ) async def task(inputs: str) -> str: """Task uses the same config that's recorded in metadata.""" result = await agent.run(inputs) return result.output # Evaluate with metadata derived from the same config async def main(): dataset = Dataset(name='config_evaluation', cases=[Case(inputs='What is the capital of France?')]) report = await dataset.evaluate( task, metadata=asdict(config), # Guaranteed to match task behavior ) print(report.experiment_metadata) """ { 'model': 'openai:gpt-5-mini', 'temperature': 0.7, 'max_tokens': 500, 'prompt_version': 'v2.1', } """ If it’s problematic to have a global task configuration, you can also create your `TaskConfig` object at the task call-site and pass it to the agent via `deps` or similar, but in this case you would still need to guarantee that the value is always the same as the value passed to `metadata` in the call to `Dataset.evaluate`. ### Anti-Pattern: Duplicate Configuration [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#anti-pattern-duplicate-configuration) **Avoid this common mistake**: from pydantic_ai import Agent from pydantic_evals import Case, Dataset # ❌ BAD: Configuration defined in multiple places agent = Agent('openai:gpt-5-mini', model_settings={'temperature': 0.7}) async def task(inputs: str) -> str: result = await agent.run(inputs) return result.output async def main(): dataset = Dataset(name='anti_pattern', cases=[Case(inputs='test')]) # ❌ BAD: Metadata manually typed - easy to get out of sync await dataset.evaluate( task, metadata={ 'model': 'openai:gpt-5-mini', # Duplicated! Could diverge from agent definition 'temperature': 0.8, # ⚠️ WRONG! Task actually uses 0.7 }, ) In this anti-pattern, the metadata claims `temperature: 0.8` but the task uses `0.7`. This leads to: * Incorrect experiment tracking * Inability to reproduce results * Confusion when comparing runs * Wasted time debugging “why results differ” Metrics vs Attributes vs Metadata --------------------------------- [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#metrics-vs-attributes-vs-metadata) Understanding the differences: | Feature | Metrics | Attributes | Case Metadata | Experiment Metadata | | --- | --- | --- | --- | --- | | **Set in** | Task execution | Task execution | Case definition | `evaluate()` call | | **Type** | int, float | Any | Any | Any | | **Purpose** | Quantitative | Qualitative | Test data | Experiment config | | **Used for** | Aggregation | Context | Input to task | Tracking runs | | **Available to** | Evaluators | Evaluators | Task & Evaluators | Report only | | **Scope** | Per case | Per case | Per case | Per experiment | from pydantic_evals import Case, Dataset, increment_eval_metric, set_eval_attribute # Case Metadata: Defined in case (before execution) case = Case( inputs='question', metadata={'difficulty': 'hard', 'category': 'math'}, # Per-case metadata ) dataset = Dataset(name='metrics_demo', cases=[case]) # Metrics & Attributes: Recorded during execution async def task(inputs): # These are recorded during execution for each case increment_eval_metric('tokens', 100) set_eval_attribute('model', 'gpt-5.2') return f'Result: {inputs}' async def main(): # Experiment Metadata: Defined at evaluation time await dataset.evaluate( task, metadata={ # Experiment-level metadata 'prompt_version': 'v2.1', 'temperature': 0.7, }, ) Troubleshooting --------------- [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#troubleshooting) ### ”Metrics/attributes not appearing” [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#metricsattributes-not-appearing) Ensure you’re calling the functions inside the task: from pydantic_evals import increment_eval_metric def process(inputs: str) -> str: return f'Processed: {inputs}' # Bad: Called outside task increment_eval_metric('count', 1) def bad_task(inputs): return process(inputs) # Good: Called inside task def good_task(inputs): increment_eval_metric('count', 1) return process(inputs) ### “Metrics not incrementing” [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#metrics-not-incrementing) Check you’re using `increment_eval_metric`, not `set_eval_attribute`: from pydantic_evals import increment_eval_metric, set_eval_attribute # Bad: This will overwrite, not increment set_eval_attribute('count', 1) set_eval_attribute('count', 1) # Still 1 # Good: This increments increment_eval_metric('count', 1) increment_eval_metric('count', 1) # Now 2 ### “Too much data in attributes” [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#too-much-data-in-attributes) Store summaries, not raw data: from pydantic_evals import set_eval_attribute giant_response_object = {'key' + str(i): 'value' * 100 for i in range(1000)} # Bad: Huge object set_eval_attribute('full_response', giant_response_object) # Good: Summary set_eval_attribute('response_size_kb', len(str(giant_response_object)) / 1024) set_eval_attribute('response_keys', list(giant_response_object.keys())[:10]) # First 10 keys Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/how-to/metrics-attributes/#next-steps) * **[Case Lifecycle Hooks](https://pydantic.dev/docs/ai/evals/how-to/lifecycle/) ** - Per-case setup, teardown, and context preparation * **[Custom Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/custom/) ** - Use metrics/attributes in evaluators * **[Logfire Integration](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/) ** - View metrics in Logfire * **[Concurrency & Performance](https://pydantic.dev/docs/ai/evals/how-to/concurrency/) ** - Optimize evaluation performance Was this page helpful? Thanks for your feedback! --- # Skills | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/skills/#_top) Skills ====== Use [Agent Skills](https://agentskills.io/specification) to give an agent specialized instructions without putting every instruction in its initial prompt. Point `Skills` at one or more skill libraries. The model first sees each skill’s name and description. When a skill is useful, the model can call Pydantic AI’s `load_capability` tool to receive that skill’s instructions. [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/skills/) > The API may change between releases. Where practical, breaking changes ship with a deprecation warning. Installation ------------ [](https://pydantic.dev/docs/ai/harness/skills/#installation) Install the `skills` extra for YAML frontmatter support: Terminal uv add "pydantic-ai-harness[skills]" Quick start ----------- [](https://pydantic.dev/docs/ai/harness/skills/#quick-start) Create a skill library: .agents/skills/ code-review/ SKILL.md Add the skill’s description and instructions: --- name: code-review description: Review a change for correctness and repository conventions. --- Inspect the change and report findings by severity. Then add the library to your agent: from pydantic_ai import Agent from pydantic_ai_harness.skills import Skills agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[Skills('.agents/skills')], ) `Skills` does not search `.agents`, `.claude`, or your home directory automatically. Pass each library you want it to load. How it works ------------ [](https://pydantic.dev/docs/ai/harness/skills/#how-it-works) When `Skills(...)` is constructed, it: 1. Scans the immediate child directories of each configured library. 2. Validates the selected `SKILL.md` files. 3. Creates one deferred Pydantic AI capability for each selected skill. The model initially sees only the skill names and descriptions. Loading a skill adds instructions headed `# Skill: `, followed by its Markdown body, to the run through the same `load_capability` flow as other [on-demand capabilities](https://pydantic.dev/docs/ai/capabilities/on-demand/) . Discovery happens once at construction. The catalog and parsed instructions are a snapshot. Construct a new `Skills` instance to rescan the libraries. Choose which skills to expose ----------------------------- [](https://pydantic.dev/docs/ai/harness/skills/#choose-which-skills-to-expose) By default, all discovered skills are included. Use `include` or `exclude` to change the catalog for a particular agent: from pydantic_ai_harness.skills import Skills review_skills = Skills( '.agents/skills', include=['code-review'], ) release_skills = Skills( '.agents/skills', exclude=['code-review'], ) | Configuration | Skills in the catalog | | --- | --- | | Neither option | All discovered skills | | `include=['a', 'b']` | Only `a` and `b` | | `include=[]` | No skills | | `exclude=['a', 'b']` | All except `a` and `b` | | `exclude=[]` | All discovered skills | `include` and `exclude` cannot be used together. The constructor overloads catch this in typed code, and runtime validation covers agent specs and untyped callers. Unknown names also fail during construction. Selection happens before frontmatter is parsed. An unselected skill does not add instructions or frontmatter validation errors to that `Skills` instance. These options control catalog exposure. They are not filesystem permissions or an access-control boundary. `Skills` reads configured paths through the process filesystem. Relative paths resolve from the process working directory. Directory paths choose where discovery starts; they do not create a containment boundary, and normal filesystem symlink resolution applies. Run the agent in an appropriately restricted environment if filesystem containment is required. A selected `SKILL.md` body becomes model instructions. Load libraries only from sources you trust, and review repository-provided skills before exposing them. Skill format ------------ [](https://pydantic.dev/docs/ai/harness/skills/#skill-format) Each immediate child directory containing `SKILL.md` is a skill: .agents/skills/ code-review/ SKILL.md release-notes/ SKILL.md The loader uses these parts of `SKILL.md`: | Part | Requirement | | --- | --- | | `name` | Optional. Defaults to the parent directory name. If provided, it must match the directory after Unicode normalization. | | `description` | Required and non-blank. The Agent Skills limit is 1,024 characters; longer descriptions load with a warning. This appears in the initial catalog. | | Markdown body | Optional. This is loaded under a generated `# Skill: ` heading. | Skill names and `include` or `exclude` values are normalized with Unicode NFKC before matching. A normalized name can contain at most 64 lowercase Unicode letters or numbers, separated by single hyphens. It cannot start or end with a hyphen. Only immediate children are discovered. For example, `code-review/references/SKILL.md` does not create another skill. Ordinary files and child directories without `SKILL.md` are ignored. You can pass several libraries: from pydantic_ai_harness.skills import Skills skills = Skills([\ '.agents/skills',\ 'company/skills',\ ]) Selected skill names must be unique across those libraries. Repeated references to the same resolved library are scanned once. Bundled files are not loaded ---------------------------- [](https://pydantic.dev/docs/ai/harness/skills/#bundled-files-are-not-loaded) Agent Skill packages can contain directories such as `references/`, `assets/`, and `scripts/`. `Skills` does not enumerate, read, or execute those files. Relative paths and placeholders such as `${CLAUDE_SKILL_DIR}` remain unchanged in the loaded instructions. `Skills` does not provide a model-visible path that resolves them. `Skills` also does not infer access from model-facing `FileSystem` or `Shell` capabilities. Adding either capability does not change which files `Skills` reads. Compatibility with existing skill libraries ------------------------------------------- [](https://pydantic.dev/docs/ai/harness/skills/#compatibility-with-existing-skill-libraries) The portable `name`, `description`, and Markdown instructions are supported. `name` may be omitted and derived from the directory. The following behavioral fields are accepted for compatibility, but their behavior is not implemented: agent, allowed-tools, argument-hint, arguments, context, dependencies, disable-model-invocation, disallowed-tools, effort, hooks, model, paths, shell, tools, user-invocable, when_to_use If a selected skill uses any of these fields, construction emits one aggregated `UserWarning`. Fields such as `license`, `compatibility`, and `metadata` are accepted without changing runtime behavior. Other unknown, non-behavioral fields are also accepted. Use an agent spec ----------------- [](https://pydantic.dev/docs/ai/harness/skills/#use-an-agent-spec) `Skills` works with Pydantic AI’s [YAML and JSON agent specs](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) : model: anthropic:claude-sonnet-4-6 capabilities: - Skills: directories: .agents/skills include: - code-review - release-notes Register `Skills` when loading the spec: from pydantic_ai import Agent from pydantic_ai_harness.skills import Skills agent = Agent.from_file('agent.yaml', custom_capability_types=[Skills]) The `skills` extra installs PyYAML. It parses `SKILL.md` frontmatter and YAML agent specs. Define capabilities in Python ----------------------------- [](https://pydantic.dev/docs/ai/harness/skills/#define-capabilities-in-python) Use Pydantic AI’s core `Capability` when your instructions or tools are defined in Python instead of a `SKILL.md` package: from pydantic_ai.capabilities import Capability refunds = Capability( id='refunds', description='Use for refund policy questions.', instructions='Check the refund policy before answering.', defer_loading=True, ) `Skills` loads portable Agent Skill packages. It does not replace the core API for code-defined capabilities. Configuration ------------- [](https://pydantic.dev/docs/ai/harness/skills/#configuration) Skills( directories: str | Path | Sequence[str | Path], *, include: Collection[str] | None = None, exclude: Collection[str] | None = None, ) * `directories` accepts one library path or a sequence of paths. * `include` exposes only the named skills. * `exclude` omits the named skills from the catalog. Pass at least one library directory, not the path of an individual skill package. Malformed frontmatter, invalid or mismatched names, duplicate selected names, unknown selections, missing libraries, and non-directory library paths fail during construction. Every selected skill is deferred. This is part of the `Skills` behavior and is not configurable. Further reading --------------- [](https://pydantic.dev/docs/ai/harness/skills/#further-reading) * [Agent Skills specification](https://agentskills.io/specification) * [Adding skills support to an agent](https://agentskills.io/client-implementation/adding-skills-support) * [Pydantic AI on-demand capabilities](https://pydantic.dev/docs/ai/capabilities/on-demand/) * [Pydantic AI capabilities overview](https://pydantic.dev/docs/ai/capabilities/overview/) API reference ------------- [](https://pydantic.dev/docs/ai/harness/skills/#api-reference) Skills ------ [](https://pydantic.dev/docs/ai/harness/skills/#pydantic_ai_harness.skills.Skills) **Bases:** `AbstractCapability[AgentDepsT]` Load Agent Skill instructions as deferred capabilities. Libraries are scanned once during construction. Each selected immediate child containing `SKILL.md` becomes a deferred capability using the skill’s name, description, and Markdown body. Bundled files are not loaded or executed. Descriptions longer than the Agent Skills limit are preserved and emit a warning. ### Attributes [](https://pydantic.dev/docs/ai/harness/skills/#attributes) #### directories [](https://pydantic.dev/docs/ai/harness/skills/#pydantic_ai_harness.skills.Skills.directories) Skill-library paths scanned during construction. **Type:** [`tuple`](https://docs.python.org/3/library/stdtypes.html#tuple) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ | `Path`, …\] **Default:** `self._normalize_directories(directories)` #### exclude [](https://pydantic.dev/docs/ai/harness/skills/#pydantic_ai_harness.skills.Skills.exclude) Exact skill names to omit from the deferred capability catalog. **Type:** [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `self._normalize_selection('exclude', exclude) if exclude is not None else frozenset()` #### include [](https://pydantic.dev/docs/ai/harness/skills/#pydantic_ai_harness.skills.Skills.include) Exact skill names to expose, or `None` to expose all discovered skills. **Type:** [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `self._normalize_selection('include', include) if include is not None else None` ### Methods [](https://pydantic.dev/docs/ai/harness/skills/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/harness/skills/#pydantic_ai_harness.skills.Skills.__init__) def __init__( directories: str | Path | Sequence[str | Path], *, include: Collection[str], exclude: None = None, ) -> None def __init__( directories: str | Path | Sequence[str | Path], *, include: None = None, exclude: Collection[str] | None = None, ) -> None Build a snapshot of the selected Agent Skills. ##### Returns [](https://pydantic.dev/docs/ai/harness/skills/#returns) [`None`](https://docs.python.org/3/library/constants.html#None) ##### Parameters [](https://pydantic.dev/docs/ai/harness/skills/#parameters) **`directories`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | `Path` | [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ | `Path`\] [](https://pydantic.dev/docs/ai/harness/skills/#pydantic_ai_harness.skills.Skills.__init__(directories)) One skill-library path or a sequence of paths. **`include`** : [`Collection`](https://docs.python.org/3/library/typing.html#typing.Collection) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/harness/skills/#pydantic_ai_harness.skills.Skills.__init__(include)) Exact names to expose. Omit to expose all discovered skills. **`exclude`** : [`Collection`](https://docs.python.org/3/library/typing.html#typing.Collection) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/harness/skills/#pydantic_ai_harness.skills.Skills.__init__(exclude)) Exact names to omit. Cannot be combined with `include`. #### \_\_repr\_\_ [](https://pydantic.dev/docs/ai/harness/skills/#pydantic_ai_harness.skills.Skills.__repr__) def __repr__() -> str Show only the `Skills` configuration that callers control. ##### Returns [](https://pydantic.dev/docs/ai/harness/skills/#returns-1) [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### apply [](https://pydantic.dev/docs/ai/harness/skills/#pydantic_ai_harness.skills.Skills.apply) def apply(visitor: Callable[[AbstractCapability[AgentDepsT]], None]) -> None Expose each selected skill as a deferred leaf capability. ##### Returns [](https://pydantic.dev/docs/ai/harness/skills/#returns-2) [`None`](https://docs.python.org/3/library/constants.html#None) Was this page helpful? Thanks for your feedback! --- # Span-Based | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#_top) Span-Based ========== Evaluate AI system behavior by analyzing OpenTelemetry spans captured during execution. Span-based evaluation enables you to evaluate **how** your AI system executes, not just **what** it produces. This is essential for complex agents where ensuring the desired behavior depends on the execution path taken, not just the final output. ### Why Span-Based Evaluation? [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#why-span-based-evaluation) Traditional evaluators assess task inputs and outputs. For simple tasks, this may be sufficient—if the output is correct, the task succeeded. But for complex multi-step agents, the _process_ matters as much as the result: * **A correct answer reached incorrectly** - An agent might produce the right output by accident (e.g., guessing, using cached data when it should have searched, calling the wrong tools but getting lucky) * **Verification of required behaviors** - You need to ensure specific tools were called, certain code paths executed, or particular patterns followed * **Performance and efficiency** - The agent should reach the answer efficiently, without unnecessary tool calls, infinite loops, or excessive retries * **Safety and compliance** - Critical to verify that dangerous operations weren’t attempted, sensitive data wasn’t accessed inappropriately, or guardrails weren’t bypassed ### Real-World Scenarios [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#real-world-scenarios) Span-based evaluation is particularly valuable for: * **RAG systems** - Verify documents were retrieved and reranked before generation, not just that the answer included citations * **Multi-agent coordination** - Ensure the orchestrator delegated to the right specialist agents in the correct order * **Tool-calling agents** - Confirm specific tools were used (or avoided), and in the expected sequence * **Debugging and regression testing** - Catch behavioral regressions where outputs remain correct but the internal logic deteriorates * **Production alignment** - Ensure your evaluation assertions operate on the same telemetry data captured in production, so eval insights directly translate to production monitoring ### How It Works [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#how-it-works) When you configure logfire (`logfire.configure()`), Pydantic Evals captures all OpenTelemetry spans generated during task execution. You can then write evaluators that assert conditions on: * **Which tools were called** - `HasMatchingSpan(query={'name_contains': 'search_tool'})` * **Code paths executed** - Verify specific functions ran or particular branches taken * **Timing characteristics** - Check that operations complete within SLA bounds * **Error conditions** - Detect retries, fallbacks, or specific failure modes * **Execution structure** - Verify parent-child relationships, delegation patterns, or execution order This creates a fundamentally different evaluation paradigm: you’re testing behavioral contracts, not just input-output relationships. Basic Usage ----------- [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#basic-usage) import logfire from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import HasMatchingSpan # Configure logfire to capture spans logfire.configure(send_to_logfire='if-token-present') dataset = Dataset( name='span_basic', cases=[Case(inputs='test')], evaluators=[\ # Check that database was queried\ HasMatchingSpan(\ query={'name_contains': 'database_query'},\ evaluation_name='used_database',\ ),\ ], ) HasMatchingSpan Evaluator ------------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#hasmatchingspan-evaluator) The [`HasMatchingSpan`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.HasMatchingSpan) evaluator checks if any span matches a query: from pydantic_evals.evaluators import HasMatchingSpan HasMatchingSpan( query={'name_contains': 'test'}, evaluation_name='span_check', ) **Returns:** `bool` - `True` if any span matches the query SpanQuery Reference ------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#spanquery-reference) A [`SpanQuery`](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanQuery) is a dictionary with query conditions: ### Name Conditions [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#name-conditions) Match spans by name: # Exact name match {'name_equals': 'search_database'} # Contains substring {'name_contains': 'tool_call'} # Regex pattern {'name_matches_regex': r'llm_call_\d+'} ### Attribute Conditions [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#attribute-conditions) Match spans with specific attributes: # Has specific attribute values {'has_attributes': {'operation': 'search', 'status': 'success'}} # Has attribute keys (any value) {'has_attribute_keys': ['user_id', 'request_id']} ### Status Conditions [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#status-conditions) Match spans by their [status](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanStatus) : # Spans that recorded an error {'has_status': 'error'} # Spans explicitly marked OK (note: successful spans are typically 'unset', not 'ok') {'has_status': 'ok'} ### Duration Conditions [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#duration-conditions) Match based on execution time: from datetime import timedelta # Minimum duration {'min_duration': 1.0} # seconds {'min_duration': timedelta(seconds=1)} # Maximum duration {'max_duration': 5.0} # seconds {'max_duration': timedelta(seconds=5)} # Range {'min_duration': 0.5, 'max_duration': 2.0} ### Logical Operators [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#logical-operators) Combine conditions: # NOT {'not_': {'name_contains': 'error'}} # AND (all must match) {'and_': [\ {'name_contains': 'tool'},\ {'max_duration': 1.0},\ ]} # OR (any must match) {'or_': [\ {'name_equals': 'search'},\ {'name_equals': 'query'},\ ]} ### Child/Descendant Conditions [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#childdescendant-conditions) Query relationships between spans: # Count direct children {'min_child_count': 1} {'max_child_count': 5} # Some child matches query {'some_child_has': {'name_contains': 'retry'}} # All children match query {'all_children_have': {'max_duration': 0.5}} # No children match query {'no_child_has': {'has_status': 'error'}} # Descendant queries (recursive) {'min_descendant_count': 5} {'some_descendant_has': {'name_contains': 'api_call'}} ### Ancestor/Depth Conditions [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#ancestordepth-conditions) Query span hierarchy: # Depth (root spans have depth 0) {'min_depth': 1} # Not a root span {'max_depth': 2} # At most 2 levels deep # Ancestor queries {'some_ancestor_has': {'name_equals': 'agent_run'}} {'all_ancestors_have': {'max_duration': 10.0}} {'no_ancestor_has': {'has_status': 'error'}} ### Stop Recursing [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#stop-recursing) Control recursive queries: { 'some_descendant_has': {'name_contains': 'expensive'}, 'stop_recursing_when': {'name_equals': 'boundary'}, } # Only search descendants until hitting a span named 'boundary' Practical Examples ------------------ [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#practical-examples) ### Verify Tool Usage [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#verify-tool-usage) Check that specific tools were called: from pydantic_evals import Case, Dataset from pydantic_evals.evaluators import HasMatchingSpan dataset = Dataset( name='tool_verification', cases=[Case(inputs='test')], evaluators=[\ # Must call search tool\ HasMatchingSpan(\ query={'name_contains': 'search_tool'},\ evaluation_name='used_search',\ ),\ \ # Must NOT call dangerous tool\ HasMatchingSpan(\ query={'not_': {'name_contains': 'delete_database'}},\ evaluation_name='safe_execution',\ ),\ ], ) ### Check Multiple Tools [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#check-multiple-tools) Verify a sequence of operations: from pydantic_evals.evaluators import HasMatchingSpan evaluators = [\ HasMatchingSpan(\ query={'name_contains': 'retrieve_context'},\ evaluation_name='retrieved_context',\ ),\ HasMatchingSpan(\ query={'name_contains': 'generate_response'},\ evaluation_name='generated_response',\ ),\ HasMatchingSpan(\ query={'and_': [\ {'name_contains': 'cite'},\ {'has_attribute_keys': ['source_id']},\ ]},\ evaluation_name='added_citations',\ ),\ ] ### Performance Assertions [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#performance-assertions) Ensure operations meet latency requirements: from pydantic_evals.evaluators import HasMatchingSpan evaluators = [\ # Database queries should be fast\ HasMatchingSpan(\ query={'and_': [\ {'name_contains': 'database'},\ {'max_duration': 0.1}, # 100ms max\ ]},\ evaluation_name='fast_db_queries',\ ),\ \ # Overall should complete quickly\ HasMatchingSpan(\ query={'and_': [\ {'name_equals': 'task_execution'},\ {'max_duration': 2.0},\ ]},\ evaluation_name='within_sla',\ ),\ ] ### Error Detection [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#error-detection) Check for error conditions using span [status](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanStatus) : from pydantic_evals.evaluators import HasMatchingSpan evaluators = [\ # An error occurred somewhere in the trace\ HasMatchingSpan(\ query={'has_status': 'error'},\ evaluation_name='had_errors',\ ),\ \ # No errors occurred: since HasMatchingSpan passes if *any* span matches,\ # anchor the query on the root span and check it and all its descendants\ HasMatchingSpan(\ query={\ 'name_equals': 'task_execution',\ 'not_': {'has_status': 'error'},\ 'no_descendant_has': {'has_status': 'error'},\ },\ evaluation_name='no_errors',\ ),\ \ # Retries happened\ HasMatchingSpan(\ query={'name_contains': 'retry'},\ evaluation_name='had_retries',\ ),\ \ # Fallback was used\ HasMatchingSpan(\ query={'name_contains': 'fallback_model'},\ evaluation_name='used_fallback',\ ),\ ] ### Complex Behavioral Checks [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#complex-behavioral-checks) Verify sophisticated behavior patterns: from pydantic_evals.evaluators import HasMatchingSpan evaluators = [\ # Agent delegated to sub-agent\ HasMatchingSpan(\ query={'and_': [\ {'name_contains': 'agent'},\ {'some_child_has': {'name_contains': 'delegate'}},\ ]},\ evaluation_name='used_delegation',\ ),\ \ # Made multiple LLM calls with retries\ HasMatchingSpan(\ query={'and_': [\ {'name_contains': 'llm_call'},\ {'some_descendant_has': {'name_contains': 'retry'}},\ {'min_descendant_count': 3},\ ]},\ evaluation_name='retry_pattern',\ ),\ ] Custom Evaluators with SpanTree ------------------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#custom-evaluators-with-spantree) For more complex span analysis, write custom evaluators: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class CustomSpanCheck(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> dict[str, bool | int]: span_tree = ctx.span_tree # Find specific spans llm_spans = span_tree.find(lambda node: 'llm' in node.name) tool_spans = span_tree.find(lambda node: 'tool' in node.name) # Calculate metrics total_llm_time = sum( span.duration.total_seconds() for span in llm_spans ) return { 'used_llm': len(llm_spans) > 0, 'used_tools': len(tool_spans) > 0, 'tool_count': len(tool_spans), 'llm_fast': total_llm_time < 2.0, } ### SpanTree API [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#spantree-api) The [`SpanTree`](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanTree) provides methods for span analysis: from pydantic_evals.otel import SpanTree # Example API (requires span_tree from context) def example_api(span_tree: SpanTree) -> None: span_tree.find(lambda n: True) # Find all matching nodes span_tree.any({'name_contains': 'test'}) # Check if any span matches span_tree.all({'name_contains': 'test'}) # Check if all spans match span_tree.count({'name_contains': 'test'}) # Count matching spans # Iteration for node in span_tree: print(node.name, node.duration, node.attributes) ### SpanNode Properties [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#spannode-properties) Each [`SpanNode`](https://pydantic.dev/docs/ai/api/pydantic_evals/otel/#pydantic_evals.otel.SpanNode) has: from pydantic_evals.otel import SpanNode # Example properties (requires node from context) def example_properties(node: SpanNode) -> None: _ = node.name # Span name _ = node.duration # timedelta _ = node.attributes # dict[str, AttributeValue] _ = node.start_timestamp # datetime _ = node.end_timestamp # datetime _ = node.status # 'unset' | 'ok' | 'error' _ = node.children # list[SpanNode] _ = node.descendants # list[SpanNode] (recursive) _ = node.ancestors # list[SpanNode] _ = node.parent # SpanNode | None Debugging Span Queries ---------------------- [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#debugging-span-queries) ### View Spans in Logfire [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#view-spans-in-logfire) If you’re sending data to Logfire, you can view all spans in the web UI to understand the trace structure. ### Print Span Tree [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#print-span-tree) from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class DebugSpans(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: for node in ctx.span_tree: print(f"{' ' * len(node.ancestors)}{node.name} ({node.duration})") return True ### Query Testing [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#query-testing) Test queries incrementally: from pydantic_evals.evaluators import HasMatchingSpan # Start simple query = {'name_contains': 'tool'} # Add conditions gradually query = {'and_': [\ {'name_contains': 'tool'},\ {'max_duration': 1.0},\ ]} # Test in evaluator HasMatchingSpan(query=query, evaluation_name='test') Use Cases --------- [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#use-cases) ### RAG System Verification [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#rag-system-verification) Verify retrieval-augmented generation workflow: from pydantic_evals.evaluators import HasMatchingSpan evaluators = [\ # Retrieved documents\ HasMatchingSpan(\ query={'name_contains': 'vector_search'},\ evaluation_name='retrieved_docs',\ ),\ \ # Reranked results\ HasMatchingSpan(\ query={'name_contains': 'rerank'},\ evaluation_name='reranked_results',\ ),\ \ # Generated with context\ HasMatchingSpan(\ query={'and_': [\ {'name_contains': 'generate'},\ {'has_attribute_keys': ['context_ids']},\ ]},\ evaluation_name='used_context',\ ),\ ] ### Multi-Agent Systems [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#multi-agent-systems) Verify agent coordination: from pydantic_evals.evaluators import HasMatchingSpan evaluators = [\ # Master agent ran\ HasMatchingSpan(\ query={'name_equals': 'master_agent'},\ evaluation_name='master_ran',\ ),\ \ # Delegated to specialist\ HasMatchingSpan(\ query={'and_': [\ {'name_contains': 'specialist_agent'},\ {'some_ancestor_has': {'name_equals': 'master_agent'}},\ ]},\ evaluation_name='delegated_correctly',\ ),\ \ # No circular delegation\ HasMatchingSpan(\ query={'not_': {'and_': [\ {'name_contains': 'agent'},\ {'some_descendant_has': {'name_contains': 'agent'}},\ {'some_ancestor_has': {'name_contains': 'agent'}},\ ]}},\ evaluation_name='no_circular_delegation',\ ),\ ] ### Tool Usage Patterns [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#tool-usage-patterns) Verify intelligent tool selection: from pydantic_evals.evaluators import HasMatchingSpan evaluators = [\ # Used search before answering\ HasMatchingSpan(\ query={'and_': [\ {'name_contains': 'search'},\ {'some_ancestor_has': {'name_contains': 'answer'}},\ ]},\ evaluation_name='searched_before_answering',\ ),\ \ # Limited tool calls (no loops)\ HasMatchingSpan(\ query={'and_': [\ {'name_contains': 'tool'},\ {'max_child_count': 5},\ ]},\ evaluation_name='reasonable_tool_usage',\ ),\ ] Best Practices -------------- [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#best-practices) 1. **Start Simple**: Begin with basic name queries, add complexity as needed 2. **Use Descriptive Names**: Name your spans well in your application code 3. **Test Queries**: Verify queries work before running full evaluations 4. **Combine with Other Evaluators**: Use span checks alongside output validation 5. **Document Expectations**: Comment why specific spans should/shouldn’t exist Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/evaluators/span-based/#next-steps) * **[Logfire Integration](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/) ** - Set up Logfire for span capture * **[Custom Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/custom/) ** - Write advanced span analysis * **[Native Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/built-in/) ** - Other evaluator types Was this page helpful? Thanks for your feedback! --- # concurrency | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#_top) concurrency =========== ConcurrencyLimitedModel ----------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimitedModel) **Bases:** `WrapperModel` A model wrapper that limits concurrent requests to the underlying model. This wrapper applies concurrency limiting at the model level, ensuring that the number of concurrent requests to the model does not exceed the configured limit. This is useful for: * Respecting API rate limits * Managing resource usage * Sharing a concurrency pool across multiple models Example usage: from pydantic_ai import Agent from pydantic_ai.models.concurrency import ConcurrencyLimitedModel # Limit to 5 concurrent requests model = ConcurrencyLimitedModel('openai:gpt-4o', limiter=5) agent = Agent(model) # Or share a limiter across multiple models from pydantic_ai import ConcurrencyLimiter # noqa E402 shared_limiter = ConcurrencyLimiter(max_running=10, name='openai-pool') model1 = ConcurrencyLimitedModel('openai:gpt-4o', limiter=shared_limiter) model2 = ConcurrencyLimitedModel('openai:gpt-4o-mini', limiter=shared_limiter) ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimitedModel.__init__) def __init__( wrapped: Model | KnownModelName, limiter: int | ConcurrencyLimit | AbstractConcurrencyLimiter, ) Initialize the ConcurrencyLimitedModel. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#parameters) **`wrapped`** : `Model` | `KnownModelName` [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimitedModel.__init__(wrapped)) The model to wrap, either a Model instance or a known model name. **`limiter`** : [`int`](https://docs.python.org/3/library/functions.html#int) | [`ConcurrencyLimit`](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimit) | [`AbstractConcurrencyLimiter`](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.AbstractConcurrencyLimiter) [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimitedModel.__init__(limiter)) The concurrency limit configuration. Can be: * An `int`: Simple limit on concurrent operations (unlimited queue). * A `ConcurrencyLimit`: Full configuration with optional backpressure. * An `AbstractConcurrencyLimiter`: A pre-created limiter for sharing across models. #### count\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimitedModel.count_tokens) `@async` def count_tokens( messages: list[ModelMessage], model_settings: ModelSettings | None, model_request_parameters: ModelRequestParameters, ) -> RequestUsage Count tokens with concurrency limiting. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#returns) [`RequestUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage) #### request [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimitedModel.request) `@async` def request( messages: list[ModelMessage], model_settings: ModelSettings | None, model_request_parameters: ModelRequestParameters, ) -> ModelResponse Make a request to the model with concurrency limiting. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#returns-1) [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) #### request\_stream [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimitedModel.request_stream) `@async` def request_stream( messages: list[ModelMessage], model_settings: ModelSettings | None, model_request_parameters: ModelRequestParameters, run_context: RunContext[Any] | None = None, ) -> AsyncGenerator[StreamedResponse] Make a streaming request to the model with concurrency limiting. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#returns-2) [`AsyncGenerator`](https://docs.python.org/3/library/typing.html#typing.AsyncGenerator) \[`StreamedResponse`\] limit\_model\_concurrency ------------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.limit_model_concurrency) def limit_model_concurrency( model: Model | KnownModelName, limiter: AnyConcurrencyLimit, ) -> Model Wrap a model with concurrency limiting. This is a convenience function to wrap a model with concurrency limiting. If the limiter is None, the model is returned unchanged. Example: from pydantic_ai.models.concurrency import limit_model_concurrency model = limit_model_concurrency('openai:gpt-4o', limiter=5) ### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#returns-3) `Model` — The wrapped model with concurrency limiting, or the original model if limiter is None. ### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#parameters-1) **`model`** : `Model` | `KnownModelName` [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.limit_model_concurrency(model)) The model to wrap. **`limiter`** : [`AnyConcurrencyLimit`](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.AnyConcurrencyLimit) [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.limit_model_concurrency(limiter)) The concurrency limit configuration. AbstractConcurrencyLimiter -------------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.AbstractConcurrencyLimiter) **Bases:** `ABC` Abstract base class for concurrency limiters. Subclass this to create custom concurrency limiters (e.g., Redis-backed distributed limiters). Example: from pydantic_ai.concurrency import AbstractConcurrencyLimiter class RedisConcurrencyLimiter(AbstractConcurrencyLimiter): def __init__(self, redis_client, key: str, max_running: int): self._redis = redis_client self._key = key self._max_running = max_running async def acquire(self, source: str) -> None: # Implement Redis-based distributed locking ... def release(self) -> None: # Release the Redis lock ... ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#methods-1) #### acquire [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.AbstractConcurrencyLimiter.acquire) `@abstractmethod` `@async` def acquire(source: str) -> None Acquire a slot, waiting if necessary. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#returns-4) [`None`](https://docs.python.org/3/library/constants.html#None) ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#parameters-2) **`source`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.AbstractConcurrencyLimiter.acquire(source)) Identifier for observability (e.g., ‘model:gpt-4o’). #### release [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.AbstractConcurrencyLimiter.release) `@abstractmethod` def release() -> None Release a slot. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#returns-5) [`None`](https://docs.python.org/3/library/constants.html#None) ConcurrencyLimiter ------------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter) **Bases:** [`AbstractConcurrencyLimiter`](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.AbstractConcurrencyLimiter) A concurrency limiter that tracks waiting operations for observability. This class wraps an anyio.CapacityLimiter and tracks the number of waiting operations. When an operation has to wait to acquire a slot, a span is created for observability purposes. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#attributes) #### available\_count [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter.available_count) Number of slots available. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### max\_running [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter.max_running) Maximum concurrent operations allowed. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### name [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter.name) Name of the limiter for observability. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### running\_count [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter.running_count) Number of operations currently running. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### waiting\_count [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter.waiting_count) Number of operations currently waiting to acquire a slot. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#methods-2) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter.__init__) def __init__( max_running: int, *, max_queued: int | None = None, name: str | None = None, tracer: Tracer | None = None, ) Initialize the ConcurrencyLimiter. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#parameters-3) **`max_running`** : [`int`](https://docs.python.org/3/library/functions.html#int) [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter.__init__(max_running)) Maximum number of concurrent operations. Must be >= 1. **`max_queued`** : [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter.__init__(max_queued)) Maximum queue depth before raising ConcurrencyLimitExceeded. Must be >= 0. **`name`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter.__init__(name)) Optional name for this limiter, used for observability when sharing a limiter across multiple models or agents. **`tracer`** : `Tracer` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter.__init__(tracer)) OpenTelemetry tracer for span creation. ##### Raises [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#raises) * `UserError` — If `max_running` is less than 1, or `max_queued` is less than 0. #### acquire [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter.acquire) `@async` def acquire(source: str) -> None Acquire a slot, creating a span if waiting is required. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#returns-6) [`None`](https://docs.python.org/3/library/constants.html#None) ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#parameters-4) **`source`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter.acquire(source)) Identifier for the source of this acquisition (e.g., ‘agent:my-agent’ or ‘model:gpt-4’). #### from\_limit [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter.from_limit) `@classmethod` def from_limit( cls, limit: int | ConcurrencyLimit, *, name: str | None = None, tracer: Tracer | None = None, ) -> Self Create a ConcurrencyLimiter from a ConcurrencyLimit configuration. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#returns-7) [`Self`](https://docs.python.org/3/library/typing.html#typing.Self) — A configured ConcurrencyLimiter. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#parameters-5) **`limit`** : [`int`](https://docs.python.org/3/library/functions.html#int) | [`ConcurrencyLimit`](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimit) [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter.from_limit(limit)) Either an int for simple limiting or a ConcurrencyLimit for full config. **`name`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter.from_limit(name)) Optional name for this limiter, used for observability. **`tracer`** : `Tracer` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter.from_limit(tracer)) OpenTelemetry tracer for span creation. #### release [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimiter.release) def release() -> None Release a slot. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#returns-8) [`None`](https://docs.python.org/3/library/constants.html#None) ConcurrencyLimit ---------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimit) Configuration for concurrency limiting with optional backpressure. ### Constructor Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#constructor-parameters) **`max_running`** : [`int`](https://docs.python.org/3/library/functions.html#int) [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimit.__init__(max_running)) Maximum number of concurrent operations allowed. Must be >= 1. **`max_queued`** : [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimit.__init__(max_queued)) Maximum number of operations waiting in the queue. Must be >= 0. If None, the queue is unlimited. If exceeded, raises `ConcurrencyLimitExceeded`. AnyConcurrencyLimit ------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.AnyConcurrencyLimit) Type alias for concurrency limit configuration. Can be: * An `int`: Simple limit on concurrent operations (unlimited queue). * A `ConcurrencyLimit`: Full configuration with optional backpressure. * An `AbstractConcurrencyLimiter`: A pre-created limiter instance for sharing across multiple models/agents. * `None`: No concurrency limiting (default). **Type:** [`TypeAlias`](https://docs.python.org/3/library/typing.html#typing.TypeAlias) **Default:** `'int | ConcurrencyLimit | AbstractConcurrencyLimiter | None'` ConcurrencyLimitExceeded ------------------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/concurrency/#pydantic_ai.ConcurrencyLimitExceeded) **Bases:** [`AgentRunError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.AgentRunError) Error raised when the concurrency queue depth exceeds max\_queued. Was this page helpful? Thanks for your feedback! --- # Pydantic AI Docs | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#_top) Pydantic AI Docs ================ `PydanticAIDocs` gives an agent a single tool, `read_pyai_docs(topic)`, that locates a Pydantic AI documentation page and returns it verbatim. Nothing is bundled into context up front. Each call resolves the topic from a configured local checkout first, then falls back to fetching the page from `pydantic/pydantic-ai:main`, so it works whether or not you have a local checkout (the remote fallback needs network access). [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/pydantic_ai_docs/) > The API may change between releases. Where practical, breaking changes ship with a deprecation warning. The problem ----------- [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#the-problem) An agent that authors Pydantic AI capabilities, hooks, tools, or toolsets needs the current docs for those APIs. Preloading the docs into the system prompt spends context the agent rarely needs in full, and pins a snapshot that drifts from `main`. The solution ------------ [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#the-solution) `PydanticAIDocs` exposes one tool, `read_pyai_docs(topic)`, that locates the requested page and returns it verbatim. Each call resolves the topic from a configured local checkout first, then falls back to fetching the page from `pydantic/pydantic-ai:main`, so it works whether or not you have a local checkout (the remote fallback needs network access). The available topics are `capabilities`, `hooks`, `tools`, `tools-advanced`, `toolsets`, and `agent`. Usage ----- [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#usage) Construct an `Agent` with `PydanticAIDocs()` in its `capabilities`. Point `local_docs_path` at a local Pydantic AI docs checkout to read from disk first, or omit it to always fetch from the remote source: from pathlib import Path from pydantic_ai import Agent from pydantic_ai_harness.pydantic_ai_docs import PydanticAIDocs agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[PydanticAIDocs(local_docs_path=Path('~/pydantic/ai/base/docs').expanduser())], ) result = agent.run_sync('Read the toolsets docs, then explain how to build a FunctionToolset.') print(result.output) The capability also adds a short static instruction telling the model that the `read_pyai_docs` tool exists and to read the relevant topic before authoring or modifying a Pydantic AI capability, hook, tool, or toolset, rather than relying on memory. The instruction is cache-stable, so it does not invalidate the prompt-cache prefix between turns. Resolution order ---------------- [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#resolution-order) Each call resolves in this order: 1. **Local checkout** — when `local_docs_path` (or the `PYDANTIC_AI_HARNESS_DOCS_PATH` env var) is set and `{path}/{topic}.md` exists, that file is read and returned. 2. **Remote fetch** — otherwise the page is fetched from `https://raw.githubusercontent.com/pydantic/pydantic-ai/main/docs/{topic}.md`. 3. **Neither resolves** — a descriptive error naming the local path tried and the URL. The capability never runs git. Keep the local checkout current yourself; the remote path always reads `main`, so it is the fresh fallback. `local_docs_path` takes precedence over the `PYDANTIC_AI_HARNESS_DOCS_PATH` env var. Both have `~` expanded, so a raw `~/...` path resolves to the local checkout instead of silently falling through to the remote source. With neither set, every call goes straight to the remote source. Configuration ------------- [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#configuration) | Option | Default | Purpose | | --- | --- | --- | | `local_docs_path` | `None` | Local pyai docs checkout to read first. Falls back to the `PYDANTIC_AI_HARNESS_DOCS_PATH` env var, then to the remote source. | | `cache` | `True` | Memoize each returned doc in-process for the capability’s lifetime, so a topic is read or fetched at most once. | Caching lives on the capability instance and is shared across the toolsets it builds, so a memoized topic survives multiple agent runs that reuse the same `PydanticAIDocs`. Set `cache=False` to re-read or re-fetch on every call — useful when the local checkout changes underneath a long-lived capability. Agent spec (YAML/JSON) ---------------------- [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#agent-spec-yamljson) `PydanticAIDocs` works with Pydantic AI’s [agent spec](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) feature for defining agents in YAML or JSON. Its serialization name is `PydanticAIDocs`: # agent.yaml model: anthropic:claude-sonnet-4-6 capabilities: - PydanticAIDocs: {} from pydantic_ai import Agent from pydantic_ai_harness.pydantic_ai_docs import PydanticAIDocs agent = Agent.from_file('agent.yaml', custom_capability_types=[PydanticAIDocs]) result = agent.run_sync('...') print(result.output) Pass `custom_capability_types` so the spec loader knows how to instantiate `PydanticAIDocs`. Specs saved before the rename from `PyaiDocs` use the old block name. To keep loading them, pass the deprecated `PyaiDocs` class (imported from `pydantic_ai_harness.docs`, which emits a deprecation warning) alongside or instead of `PydanticAIDocs` — it keeps the `PyaiDocs` serialization name. Re-save with `PydanticAIDocs` to migrate. API reference ------------- [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#api-reference) PydanticAIDocs -------------- [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#pydantic_ai_harness.pydantic_ai_docs.PydanticAIDocs) **Bases:** `AbstractCapability[AgentDepsT]` Locate and return Pydantic AI documentation on demand. Exposes a single `read_pyai_docs(topic)` tool. Docs are located and returned when asked for — never bundled into context. Each call resolves the topic from a configured local checkout first, then falls back to fetching the page from `pydantic/pydantic-ai:main`, so it works in any environment. The local checkout path comes from `local_docs_path`, or the `PYDANTIC_AI_HARNESS_DOCS_PATH` env var when that is unset; with neither set every call goes straight to the remote source. The capability never runs git — keep the local checkout current yourself; the remote path always reads `main`. from pathlib import Path from pydantic_ai import Agent from pydantic_ai_harness.pydantic_ai_docs import PydanticAIDocs agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[PydanticAIDocs(local_docs_path=Path('~/pydantic/ai/base/docs').expanduser())], ) ### Attributes [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#attributes) #### cache [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#pydantic_ai_harness.pydantic_ai_docs.PydanticAIDocs.cache) If `True`, each returned doc is memoized in-process for the capability’s lifetime, so a topic is read or fetched at most once. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `True` #### local\_docs\_path [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#pydantic_ai_harness.pydantic_ai_docs.PydanticAIDocs.local_docs_path) Local pyai docs checkout to read first. When `None`, falls back to the `PYDANTIC_AI_HARNESS_DOCS_PATH` env var, then to the remote source. **Type:** `Path` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` ### Methods [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#methods) #### get\_instructions [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#pydantic_ai_harness.pydantic_ai_docs.PydanticAIDocs.get_instructions) def get_instructions() -> AgentInstructions[AgentDepsT] | None Static, cache-stable guidance on using the docs tool. ##### Returns [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#returns) `AgentInstructions`\[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) #### get\_serialization\_name [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#pydantic_ai_harness.pydantic_ai_docs.PydanticAIDocs.get_serialization_name) `@classmethod` def get_serialization_name(cls) -> str | None Serialization name for agent-spec support. ##### Returns [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#returns-1) [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### get\_toolset [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#pydantic_ai_harness.pydantic_ai_docs.PydanticAIDocs.get_toolset) def get_toolset() -> AgentToolset[AgentDepsT] | None Toolset providing `read_pyai_docs` over the resolved local path and shared cache. ##### Returns [](https://pydantic.dev/docs/ai/harness/pydantic-ai-docs/#returns-2) [`AgentToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AgentToolset) \[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) Was this page helpful? Thanks for your feedback! --- # OpenRouter | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/models/openrouter/#_top) OpenRouter ========== Install ------- [](https://pydantic.dev/docs/ai/models/openrouter/#install) To use `OpenRouterModel`, you need to either install `pydantic-ai`, or install `pydantic-ai-slim` with the `openrouter` optional group: * [pip](https://pydantic.dev/docs/ai/models/openrouter/#tab-panel-126) * [uv](https://pydantic.dev/docs/ai/models/openrouter/#tab-panel-127) Terminal pip install "pydantic-ai-slim[openrouter]" Terminal uv add "pydantic-ai-slim[openrouter]" Configuration ------------- [](https://pydantic.dev/docs/ai/models/openrouter/#configuration) To use [OpenRouter](https://openrouter.ai/) , first create an API key at [openrouter.ai/keys](https://openrouter.ai/keys) . You can set the `OPENROUTER_API_KEY` environment variable and use [`OpenRouterProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.openrouter.OpenRouterProvider) by name: from pydantic_ai import Agent agent = Agent('openrouter:anthropic/claude-sonnet-4.6') ... Or initialise the model and provider directly: from pydantic_ai import Agent from pydantic_ai.models.openrouter import OpenRouterModel from pydantic_ai.providers.openrouter import OpenRouterProvider model = OpenRouterModel( 'anthropic/claude-sonnet-4.6', provider=OpenRouterProvider(api_key='your-openrouter-api-key'), ) agent = Agent(model) ... App Attribution --------------- [](https://pydantic.dev/docs/ai/models/openrouter/#app-attribution) OpenRouter has an [app attribution](https://openrouter.ai/docs/app-attribution) feature to track your application in their public ranking and analytics. You can pass in an `app_url` and `app_title` when initializing the provider to enable app attribution. Both fall back to the `OPENROUTER_APP_URL` and `OPENROUTER_APP_TITLE` environment variables when omitted. from pydantic_ai.providers.openrouter import OpenRouterProvider provider=OpenRouterProvider( api_key='your-openrouter-api-key', app_url='https://your-app.com', app_title='Your App', ), ... Model Settings -------------- [](https://pydantic.dev/docs/ai/models/openrouter/#model-settings) You can customize model behavior using [`OpenRouterModelSettings`](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModelSettings) : from pydantic_ai import Agent from pydantic_ai.models.openrouter import OpenRouterModel, OpenRouterModelSettings settings = OpenRouterModelSettings( openrouter_reasoning={ 'effort': 'high', }, openrouter_usage={ 'include': True, } ) model = OpenRouterModel('openai/gpt-5.2') agent = Agent(model, model_settings=settings) ... ### Eager Input Streaming [](https://pydantic.dev/docs/ai/models/openrouter/#eager-input-streaming) For Anthropic models via OpenRouter, you can enable eager input streaming to reduce latency for tool calls with large inputs. Set [`anthropic_eager_input_streaming`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_eager_input_streaming) in [`AnthropicModelSettings`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings) : from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModelSettings from pydantic_ai.models.openrouter import OpenRouterModel model = OpenRouterModel('anthropic/claude-sonnet-4-5') settings = AnthropicModelSettings(anthropic_eager_input_streaming=True) agent = Agent(model, model_settings=settings) ... Forced tool choice ------------------ [](https://pydantic.dev/docs/ai/models/openrouter/#forced-tool-choice) Pydantic AI treats a forced [`tool_choice`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.tool_choice) as incompatible with [thinking](https://pydantic.dev/docs/ai/capabilities/thinking/) on every `anthropic/` model routed through OpenRouter. This is deliberately more conservative than [the direct Anthropic API](https://pydantic.dev/docs/ai/models/anthropic/#forced-tool-choice) , where adaptive thinking accepts forcing — the OpenRouter route hasn’t been verified, and it fails quietly rather than loudly: where Anthropic rejects an incompatible combination outright, OpenRouter silently drops the `reasoning` field from the request instead, so the response comes back with no thinking at all. See [#7283](https://github.com/pydantic/pydantic-ai/issues/7283) . With thinking enabled on an `anthropic/` model: * An explicit `tool_choice='required'` (or a list of tool names) raises a [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) ; disable thinking or use `tool_choice='auto'`. * A `required` choice that Pydantic AI resolved on your behalf (e.g. from an [output tool](https://pydantic.dev/docs/ai/core-concepts/output/#tool-output) ) falls back softly to `'auto'`, so thinking is preserved. If the resolved choice named a single tool, the available tool list is filtered to that tool while `tool_choice` remains `'auto'`. The model may therefore answer with text instead of calling it; when an output tool is required, Pydantic AI retries with a prompt to call a tool. Prompt Caching -------------- [](https://pydantic.dev/docs/ai/models/openrouter/#prompt-caching) OpenRouter supports [prompt caching](https://openrouter.ai/docs/guides/best-practices/prompt-caching) for downstream providers that implement it. Pydantic AI’s OpenRouter cache settings control explicit `cache_control` breakpoints for Anthropic and Gemini models: 1. **Cache System Instructions**: Set [`OpenRouterModelSettings.openrouter_cache_instructions`](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModelSettings.openrouter_cache_instructions) to `True` or specify `'5m'` / `'1h'` directly 2. **Cache the Last Message**: Set [`OpenRouterModelSettings.openrouter_cache_messages`](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModelSettings.openrouter_cache_messages) to `True` to automatically cache the last message in the conversation 3. **Cache Tool Definitions**: Set [`OpenRouterModelSettings.openrouter_cache_tool_definitions`](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModelSettings.openrouter_cache_tool_definitions) to `True` or specify `'5m'` / `'1h'` directly 4. **Fine-Grained Control with [`CachePoint`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.CachePoint) **: Insert a `CachePoint` marker in user messages to cache everything before it ### OpenAI GPT-5.6 explicit caching [](https://pydantic.dev/docs/ai/models/openrouter/#openai-gpt-56-explicit-caching) [`OpenRouterModel`](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModel) does not currently translate [`CachePoint`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.CachePoint) into OpenAI’s breakpoint protocol (OpenAI models on OpenRouter still get automatic caching). For explicit GPT-5.6 breakpoints, combine [`OpenAIResponsesModel`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIResponsesModel) (or [`OpenAIChatModel`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIChatModel) ) with [`OpenRouterProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.openrouter.OpenRouterProvider) : from pydantic_ai import Agent, CachePoint from pydantic_ai.models.openai import OpenAIResponsesModel, OpenAIResponsesModelSettings from pydantic_ai.providers.openrouter import OpenRouterProvider model = OpenAIResponsesModel( 'openai/gpt-5.6-sol', provider=OpenRouterProvider(api_key='your-openrouter-api-key'), ) settings = OpenAIResponsesModelSettings( openai_prompt_cache_key='product-docs-v1', openai_prompt_cache_options={'mode': 'explicit', 'ttl': '30m'}, # OpenRouter also offers Azure routes for GPT-5.6, where explicit caching is not documented. extra_body={'provider': {'only': ['openai']}}, ) agent = Agent(model, model_settings=settings) result = agent.run_sync([\ 'Long-lived reference material...',\ CachePoint(),\ 'Answer using the reference material.',\ ]) The OpenRouter Responses API uses the same request-wide TTL and usage fields as OpenAI. Restricting the downstream provider to `openai` avoids routing explicit-cache requests to endpoints where these fields are not documented. OpenRouter currently documents explicit breakpoints only on text blocks, so place `CachePoint` markers after text content. ### Caching via Model Settings [](https://pydantic.dev/docs/ai/models/openrouter/#caching-via-model-settings) Use [`OpenRouterModelSettings`](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModelSettings) to enable explicit caching for system instructions, the last conversation message, and tool definitions: from pydantic_ai import Agent, RunContext from pydantic_ai.models.openrouter import OpenRouterModel, OpenRouterModelSettings model = OpenRouterModel('anthropic/claude-sonnet-4.6') agent = Agent( model, instructions='You are a specialized assistant with deep domain knowledge...', model_settings=OpenRouterModelSettings( openrouter_cache_instructions=True, # Cache system instructions (broadly supported) openrouter_cache_messages=True, # Cache the last message (best with Anthropic) openrouter_cache_tool_definitions=True, # Cache tool definitions (Anthropic only) ), ) @agent.tool def search_docs(ctx: RunContext, query: str) -> str: """Search documentation.""" return f'Results for {query}' ... Each setting accepts `True` or an explicit `'5m'` / `'1h'` TTL value. `True` sends Anthropic’s default `'5m'` TTL for Anthropic models; Gemini ignores TTL values and manages cache lifetime itself. Check `result.usage.cache_write_tokens` on initial writes and `result.usage.cache_read_tokens` on reuse, including subsequent calls with `message_history=result.all_messages()`. OpenRouter uses [provider sticky routing](https://openrouter.ai/docs/guides/best-practices/prompt-caching#provider-sticky-routing) after prompt-cached requests to improve cache locality. For cache-sensitive workflows that need stricter provider control or disabled fallbacks, also set [`openrouter_provider`](https://pydantic.dev/docs/ai/api/models/openrouter/#pydantic_ai.models.openrouter.OpenRouterModelSettings.openrouter_provider) , for example with `{'order': ['anthropic'], 'allow_fallbacks': False}`. ### Fine-Grained Control with CachePoint [](https://pydantic.dev/docs/ai/models/openrouter/#fine-grained-control-with-cachepoint) Use [`CachePoint`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.CachePoint) markers to control exactly where cache boundaries are placed: from pydantic_ai import Agent, CachePoint from pydantic_ai.models.openrouter import OpenRouterModel model = OpenRouterModel('anthropic/claude-sonnet-4.6') agent = Agent(model) prompt = [\ 'Long reference document or context to cache...',\ CachePoint(), # Cache everything before this point\ 'Now answer my question about the context above',\ ] ... Pass the prompt list to `agent.run_sync(prompt)`. Everything before the `CachePoint()` marker is cached. You can place multiple markers for fine-grained control over cache boundaries. Web Search ---------- [](https://pydantic.dev/docs/ai/models/openrouter/#web-search) OpenRouter supports web search through its [Beta server tool](https://openrouter.ai/docs/guides/features/server-tools/web-search) . Enable it with [`WebSearchTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebSearchTool) . The model decides whether to search and may make zero or multiple searches for a request. ### Web Search Parameters [](https://pydantic.dev/docs/ai/models/openrouter/#web-search-parameters) You can configure search context, approximate user location, domain filters, and a limit on searches with [`WebSearchTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebSearchTool) : web\_search\_openrouter.py from pydantic_ai import Agent from pydantic_ai.capabilities import NativeTool from pydantic_ai.models.openrouter import OpenRouterModel from pydantic_ai.native_tools import WebSearchTool tool = WebSearchTool( search_context_size='high', user_location={'city': 'London', 'country': 'GB'}, allowed_domains=['pydantic.dev'], max_uses=1, ) model = OpenRouterModel('openai/gpt-4.1') agent = Agent( model, capabilities=[NativeTool(tool)], ) result = agent.run_sync('What is the latest news in AI?') Pydantic AI surfaces the per-request web-search count under [`ModelResponse.provider_details`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse.provider_details) `'server_tool_use'`. Was this page helpful? Thanks for your feedback! --- # LocalStack | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/localstack/#_top) LocalStack ========== `LocalStack` gives an agent access to an emulated AWS environment, so it can provision and exercise AWS services without touching a real account. It wires the AWS CLI to a running [LocalStack](https://www.localstack.cloud/) instance — injecting the endpoint, region, and credentials — and can optionally start and stop the LocalStack Docker container for each run. [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/localstack/) > Import this capability from its submodule — there is no top-level `pydantic_ai_harness` re-export: > > from pydantic_ai_harness.localstack import LocalStack > > > The API may change between releases. Where practical, breaking changes ship with a deprecation warning. The problem ----------- [](https://pydantic.dev/docs/ai/harness/localstack/#the-problem) Agents that build or test cloud infrastructure need somewhere to create buckets, tables, queues, and functions. Pointing them at real AWS is slow, costs money, risks leaking credentials, and is hard to reset between runs. LocalStack emulates the AWS APIs locally, but wiring an agent to it means repeating the same boilerplate: injecting the endpoint URL, supplying dummy credentials, shelling out to the AWS CLI, and checking which services are up. Usage ----- [](https://pydantic.dev/docs/ai/harness/localstack/#usage) `LocalStack` exposes AWS tooling wired to a running LocalStack instance. The agent issues plain AWS CLI commands; the capability injects the endpoint, region, and credentials, and adds a health check for the emulated services. from pydantic_ai import Agent from pydantic_ai_harness.localstack import LocalStack agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[LocalStack()], ) result = agent.run_sync('Create an S3 bucket called reports and list all buckets.') print(result.output) By default the agent connects to a LocalStack instance you started separately — for example with the [`localstack` CLI](https://docs.localstack.cloud/aws/tooling/localstack-cli/) (`localstack start`). The defaults match LocalStack’s conventions: the edge endpoint `http://localhost.localstack.cloud:4566` (which resolves to `127.0.0.1`) and `test` / `test` credentials. Set `manage_container=True` to have the capability start and stop a fresh Docker container per run. Tools ----- [](https://pydantic.dev/docs/ai/harness/localstack/#tools) | Tool | Purpose | | --- | --- | | `aws_cli` | Run an AWS CLI command against LocalStack. Pass the command **without** the leading `aws` and **without** `--endpoint-url` — both are injected. Returns labelled stdout/stderr plus an exit code on failure. | | `localstack_health` | Query LocalStack’s health endpoint and return the JSON of which services (s3, dynamodb, sqs, etc.) are available. | Commands run as an argument vector (no shell), so shell operators and redirection in the command string have no effect. Output is labelled with `[stdout]` / `[stderr]` markers and an `[exit code: N]` line on non-zero exit. When it exceeds `max_output_chars` the **tail** is kept (the head is dropped), so errors survive truncation. The AWS CLI can read from and write to local files through arguments such as `--body`, `file://`, `fileb://`, and `s3 cp`. Treat this capability as both AWS-emulator access and AWS CLI access to the process’s filesystem. Service controls ---------------- [](https://pydantic.dev/docs/ai/harness/localstack/#service-controls) | Field | Effect | | --- | --- | | `allowed_services` | If non-empty, only these AWS services may be used (allowlist), e.g. `['s3', 'dynamodb']`. | | `denied_services` | These AWS services are always rejected (denylist). | `allowed_services` and `denied_services` are mutually exclusive — set one, not both. The service is the first non-flag token of the command (`s3` in `s3 ls`). Managing the container ---------------------- [](https://pydantic.dev/docs/ai/harness/localstack/#managing-the-container) Set `manage_container=True` and the capability starts a LocalStack Docker container for each run and stops it when the run ends, so the agent always gets a fresh, isolated environment. Docker must be installed and running. from pydantic_ai_harness.localstack import LocalStack LocalStack( manage_container=True, image='localstack/localstack', container_env={'DEBUG': '1', 'PERSISTENCE': '1'}, startup_timeout=120.0, ) The container’s edge port (`4566`) is published on the host port from `endpoint_url`, and the capability waits for the health endpoint before the run starts, then stops the container when it ends (even if the run raises). Each run gets its own container, so concurrent runs of one agent need distinct host ports or an externally managed instance (`manage_container=False`). Since LocalStack 2026.03.0 the default `localstack/localstack` image is a single image that requires an auth token to start (a free Hobby/OSS token covers community usage). When `LOCALSTACK_AUTH_TOKEN` is set in the current process it is forwarded to the container automatically; a legacy `LOCALSTACK_API_KEY` value is forwarded when no auth token is set. Auth values are forwarded through the Docker CLI environment rather than embedded in the `docker run` arguments. The default `localstack/localstack` image requires a token to start, so a managed run needs one configured. To run tokenless, set `image` to a tag from before the account requirement, such as a `localstack/localstack:4.x` release. Docker-backed services such as Lambda need the Docker socket mounted, and some services expose ports outside the gateway (LocalStack reserves `4510-4559`). Enable those explicitly when a service you test requires them: from pydantic_ai_harness.localstack import LocalStack LocalStack(manage_container=True, service_port_range='4510-4559', mount_docker_socket=True) Mounting the Docker socket gives the container host-level Docker control. Keep `mount_docker_socket=False` unless the emulated service requires it and the run environment is already trusted. The same lifecycle is available standalone as an async context manager: import asyncio from pydantic_ai_harness.localstack import LocalStackContainer async def main() -> None: async with LocalStackContainer(environment={'DEBUG': '1'}) as localstack: ... # talk to localstack.endpoint_url asyncio.run(main()) Configuration ------------- [](https://pydantic.dev/docs/ai/harness/localstack/#configuration) from pydantic_ai_harness.localstack import LocalStack LocalStack( endpoint_url='http://localhost.localstack.cloud:4566', # edge endpoint (host port reused when managed) region='us-east-1', # region for the CLI and environment access_key_id='test', # LocalStack accepts any value secret_access_key='test', # LocalStack accepts any value allowed_services=[], # allowlist (mutually exclusive with denied) denied_services=[], # denylist default_timeout=60.0, # seconds, per command and health check max_output_chars=50_000, # output cap returned to the model aws_cli_path='aws', # CLI executable (e.g. 'aws' or 'awslocal') manage_container=False, # start/stop a Docker container per run image='localstack/localstack', # image used when managing the container host_address='127.0.0.1', # host address for Docker port publishing service_port_range=None, # e.g. '4510-4559' for non-gateway service ports mount_docker_socket=False, # required by Docker-backed services such as Lambda container_name=None, # optional name for the managed container container_env={}, # env vars for the managed container docker_path='docker', # Docker executable startup_timeout=120.0, # seconds to wait for the container to be ready include_instructions=True, # add usage instructions to the prompt ) The AWS CLI must be installed and on `PATH` (or point `aws_cli_path` at it). If the binary is missing, `aws_cli` returns a clear error instead of aborting the run. Set `include_instructions=False` to omit the capability’s prompt text when you supply your own. Agent spec (YAML/JSON) ---------------------- [](https://pydantic.dev/docs/ai/harness/localstack/#agent-spec-yamljson) `LocalStack` works with Pydantic AI’s [agent spec](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) : # agent.yaml model: anthropic:claude-sonnet-4-6 capabilities: - LocalStack: endpoint_url: http://localhost.localstack.cloud:4566 allowed_services: ['s3', 'dynamodb', 'sqs'] from pydantic_ai import Agent from pydantic_ai_harness.localstack import LocalStack agent = Agent.from_file('agent.yaml', custom_capability_types=[LocalStack]) Pass `custom_capability_types` so the spec loader knows how to instantiate `LocalStack`. Further reading --------------- [](https://pydantic.dev/docs/ai/harness/localstack/#further-reading) * [LocalStack documentation](https://docs.localstack.cloud/) * [Pydantic AI capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) * [Toolsets](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) API reference ------------- [](https://pydantic.dev/docs/ai/harness/localstack/#api-reference) LocalStack ---------- [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack) **Bases:** `AbstractCapability[AgentDepsT]` Access to an emulated AWS environment via LocalStack. Gives the agent AWS CLI tooling wired to a running LocalStack instance, so it can provision and interact with AWS services (S3, DynamoDB, SQS, Lambda, …) without touching real AWS. Start LocalStack separately (`localstack start` or its Docker image) before running the agent. from pydantic_ai import Agent from pydantic_ai_harness.localstack import LocalStack agent = Agent('anthropic:claude-sonnet-4-6', capabilities=[LocalStack()]) result = agent.run_sync('Create an S3 bucket called reports and list all buckets.') print(result.output) ### Attributes [](https://pydantic.dev/docs/ai/harness/localstack/#attributes) #### access\_key\_id [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.access_key_id) AWS access key id. LocalStack accepts any value; defaults to its `test` convention. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'test'` #### allowed\_services [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.allowed_services) If non-empty, only these AWS services may be used (allowlist), e.g. `['s3', 'dynamodb']`. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `field(default_factory=(list[str]))` #### aws\_cli\_path [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.aws_cli_path) Path or name of the AWS CLI executable (e.g. `aws` or `awslocal`). **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'aws'` #### container\_env [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.container_env) Environment variables passed to the managed container, e.g. `{'DEBUG': '1'}`. **Type:** [`Mapping`](https://docs.python.org/3/library/typing.html#typing.Mapping) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `field(default_factory=(dict[str, str]))` #### container\_name [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.container_name) Optional name for the managed container. Leave None to let Docker assign one. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### default\_timeout [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.default_timeout) Default timeout in seconds for AWS CLI commands and the health check. **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) **Default:** `60.0` #### denied\_services [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.denied_services) These AWS services are always rejected (denylist). **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `field(default_factory=(list[str]))` #### docker\_path [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.docker_path) Path or name of the Docker executable used to manage the container. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'docker'` #### endpoint\_url [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.endpoint_url) Base URL of the running LocalStack instance. Defaults to LocalStack’s `localhost.localstack.cloud` domain (which resolves to `127.0.0.1`) for compatibility with AWS SDKs that need subdomain-style hosts. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'http://localhost.localstack.cloud:4566'` #### host\_address [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.host_address) Host address Docker publishes the LocalStack edge port on. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'127.0.0.1'` #### image [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.image) Docker image to run when `manage_container` is True. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'localstack/localstack'` #### include\_instructions [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.include_instructions) If True, add instructions telling the model how to use the emulated environment. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `True` #### manage\_container [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.manage_container) If True, start a LocalStack Docker container for each run and stop it when the run ends. Requires Docker. When False (default), the agent connects to a LocalStack instance you started separately at `endpoint_url`. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### max\_output\_chars [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.max_output_chars) Maximum characters of output returned to the model. Must be positive. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `50000` #### mount\_docker\_socket [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.mount_docker_socket) If True, mount `/var/run/docker.sock` into the managed container for Docker-backed services like Lambda. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### region [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.region) AWS region passed to the CLI and exported to the environment. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'us-east-1'` #### secret\_access\_key [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.secret_access_key) AWS secret access key. LocalStack accepts any value; defaults to its `test` convention. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'test'` #### service\_port\_range [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.service_port_range) Optional host/container port range for services that expose their own ports, e.g. `4510-4559`. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### startup\_timeout [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.startup_timeout) Seconds to wait for the managed container to become ready before failing. **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) **Default:** `120.0` ### Methods [](https://pydantic.dev/docs/ai/harness/localstack/#methods) #### get\_instructions [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.get_instructions) def get_instructions() -> str | None Explain the emulated environment to the model, unless disabled. ##### Returns [](https://pydantic.dev/docs/ai/harness/localstack/#returns) [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### get\_toolset [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStack.get_toolset) def get_toolset() -> AgentToolset[AgentDepsT] Build and return the LocalStack toolset. ##### Returns [](https://pydantic.dev/docs/ai/harness/localstack/#returns-1) [`AgentToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AgentToolset) \[`AgentDepsT`\] LocalStackToolset ----------------- [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStackToolset) **Bases:** `FunctionToolset[AgentDepsT]` Gives an agent the ability to drive an emulated AWS environment. Wraps the AWS CLI: `aws_cli` runs a command against a running LocalStack instance with the endpoint, region, and credentials injected, while `localstack_health` reports which emulated services are available. Commands are executed as an argument vector (no shell), so shell operators and redirection in the command string have no effect. ### Methods [](https://pydantic.dev/docs/ai/harness/localstack/#methods-1) #### \_\_aenter\_\_ [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStackToolset.__aenter__) `@async` def __aenter__() -> Self Start the managed LocalStack container, if configured, before tools run. ##### Returns [](https://pydantic.dev/docs/ai/harness/localstack/#returns-2) [`Self`](https://docs.python.org/3/library/typing.html#typing.Self) #### \_\_aexit\_\_ [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStackToolset.__aexit__) `@async` def __aexit__(*args: object) -> None Stop the managed LocalStack container, if one was started. ##### Returns [](https://pydantic.dev/docs/ai/harness/localstack/#returns-3) [`None`](https://docs.python.org/3/library/constants.html#None) #### aws\_cli [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStackToolset.aws_cli) `@async` def aws_cli(command: str, *, timeout_seconds: float | None = None) -> str Run an AWS CLI command against the emulated AWS environment. Pass the command without the leading `aws` and without `--endpoint-url`; the endpoint, region, and credentials are injected automatically. For example `s3 mb s3://my-bucket`, `s3 ls`, or `dynamodb list-tables`. ##### Returns [](https://pydantic.dev/docs/ai/harness/localstack/#returns-4) [`str`](https://docs.python.org/3/library/stdtypes.html#str) — Labelled stdout/stderr output, with an exit code on non-zero exit. ##### Parameters [](https://pydantic.dev/docs/ai/harness/localstack/#parameters) **`command`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStackToolset.aws_cli(command)) The AWS CLI command to run (e.g. `s3 ls`). **`timeout_seconds`** : [`float`](https://docs.python.org/3/library/functions.html#float) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStackToolset.aws_cli(timeout_seconds)) Maximum seconds to wait (default: the configured timeout). #### call\_tool [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStackToolset.call_tool) `@async` def call_tool( name: str, tool_args: dict[str, Any], ctx: RunContext[AgentDepsT], tool: ToolsetTool[AgentDepsT], ) -> Any Enforce the model-visible output cap at the tool dispatch seam. Only `str` results are capped; a future tool returning rich content (e.g. `ToolReturn`) needs this seam extended. ##### Returns [](https://pydantic.dev/docs/ai/harness/localstack/#returns-5) [`Any`](https://docs.python.org/3/library/typing.html#typing.Any) #### for\_run [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStackToolset.for_run) `@async` def for_run(ctx: RunContext[AgentDepsT]) -> AbstractToolset[AgentDepsT] Return a fresh instance per run so a managed container is isolated and torn down. `get_toolset` builds one shared instance at agent construction. When this toolset manages a Docker container it holds per-run lifecycle state, so each run gets its own instance (and its own container) that `__aexit__` can stop. ##### Returns [](https://pydantic.dev/docs/ai/harness/localstack/#returns-6) [`AbstractToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AbstractToolset) \[`AgentDepsT`\] #### localstack\_health [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStackToolset.localstack_health) `@async` def localstack_health() -> str Report the health and availability of the emulated AWS services. Queries LocalStack’s health endpoint and returns the raw JSON, which maps each service (s3, dynamodb, sqs, …) to its state (available, running, …). ##### Returns [](https://pydantic.dev/docs/ai/harness/localstack/#returns-7) [`str`](https://docs.python.org/3/library/stdtypes.html#str) — The health JSON, or an error message if LocalStack is unreachable. LocalStackContainer ------------------- [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStackContainer) Async context manager that starts and stops a LocalStack Docker container. Drives the `docker` CLI, so Docker must be installed and running. On enter it launches the container, polls the health endpoint until LocalStack is ready, and exposes `endpoint_url`. On exit it stops the container — it is started with `--rm`, so stopping also removes it. async with LocalStackContainer() as localstack: ... # talk to localstack.endpoint_url ### Attributes [](https://pydantic.dev/docs/ai/harness/localstack/#attributes-1) #### container\_id [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStackContainer.container_id) The running container’s id, or None when it is not running. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### endpoint\_url [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStackContainer.endpoint_url) URL of the container’s edge endpoint. For the default loopback publish addresses (`127.0.0.1` and the bind-all `0.0.0.0`, both reachable at `127.0.0.1`), uses LocalStack’s `localhost.localstack.cloud` domain, which resolves to `127.0.0.1` and supports the subdomain-style hosts some AWS SDKs need. For any other `host_address`, uses that address literally, since `localhost.localstack.cloud` does not resolve there. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) ### Methods [](https://pydantic.dev/docs/ai/harness/localstack/#methods-2) #### \_\_aenter\_\_ [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStackContainer.__aenter__) `@async` def __aenter__() -> Self Start the container and wait for it to become ready. ##### Returns [](https://pydantic.dev/docs/ai/harness/localstack/#returns-8) [`Self`](https://docs.python.org/3/library/typing.html#typing.Self) #### \_\_aexit\_\_ [](https://pydantic.dev/docs/ai/harness/localstack/#pydantic_ai_harness.localstack.LocalStackContainer.__aexit__) `@async` def __aexit__(*args: object) -> None Stop and remove the container. ##### Returns [](https://pydantic.dev/docs/ai/harness/localstack/#returns-9) [`None`](https://docs.python.org/3/library/constants.html#None) Was this page helpful? Thanks for your feedback! --- # azure | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/realtime/azure/#_top) azure ===== The Azure OpenAI realtime model reuses the OpenAI Realtime codec and connection, authenticates with [`AzureProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.azure.AzureProvider) , and so needs the `realtime` and `openai` optional groups (`pip install "pydantic-ai-slim[realtime,openai]"`). Azure realtime support using the OpenAI GA or Azure AI Voice Live protocol. AzureRealtimeConnection ----------------------- [](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeConnection) **Bases:** `OpenAIRealtimeConnection` A live WebSocket connection to Azure OpenAI’s realtime API. Reuses [`OpenAIRealtimeConnection`](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeConnection) for the shared GA wire protocol, naming Azure as the vendor so a connection that drops or rejects content doesn’t send someone debugging an Azure session to OpenAI’s status page. AzureRealtimeModel ------------------ [](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModel) **Bases:** `OpenAIRealtimeModel` Azure realtime model using the OpenAI GA protocol or Azure AI Voice Live. The existing [`AzureProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.azure.AzureProvider) supplies the Azure resource endpoint and API key. The WebSocket transport does not use its OpenAI SDK client or `api_version`. By default it connects to the GA `/openai/v1/realtime` endpoint; set [`azure_voice_live`](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModelSettings.azure_voice_live) to connect to `/voice-live/realtime` with the Voice Live beta session protocol. Both use an `api-key` header. Pass a Microsoft Entra ID `credential` (e.g. `azure.identity.DefaultAzureCredential()`) to authenticate every request to the resource — the realtime WebSocket session _and_ the browser WebRTC signaling calls — with a bearer token instead of the `api-key` (needed when the resource is locked to managed identity). For browser WebRTC the browser still only ever receives the short-lived ephemeral secret, never the Entra token or the API key. A model served only by Voice Live (e.g. the cascade chat models like `gpt-5`, or `phi4-mm-realtime`) routes there automatically; a model served by both defaults to GA and needs `azure_voice_live=True` for Voice Live; a GA-only model rejects the setting. See [`AzureRealtimeModelProfile.azure_realtime_apis`](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModelProfile.azure_realtime_apis) . ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/azure/#attributes) #### profile [](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModel.profile) The Azure realtime profile, with the model’s serving APIs and minus what Voice Live can’t do. Stamps [`azure_realtime_apis`](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModelProfile.azure_realtime_apis) for a recognized model (a `profile=` override wins), which routes between the GA API and Voice Live. And because Voice Live negotiates WebRTC over its own WebSocket control channel rather than the GA signaling endpoints this model inherits, a model configured for Voice Live reports no WebRTC support and the signaling methods refuse. Voice Live selected per session instead can’t be seen from here, so those calls still refuse at the point of use. **Type:** `RealtimeModelProfile` ### Methods [](https://pydantic.dev/docs/ai/api/realtime/azure/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModel.__init__) def __init__( model: AzureRealtimeModelName, *, provider: Provider[AsyncOpenAI] | str = 'azure', settings: RealtimeModelSettings | None = None, profile: RealtimeModelProfileSpec | None = None, credential: AzureTokenCredential | None = None, ) -> None Create an Azure OpenAI realtime model. ##### Returns [](https://pydantic.dev/docs/ai/api/realtime/azure/#returns) [`None`](https://docs.python.org/3/library/constants.html#None) ##### Parameters [](https://pydantic.dev/docs/ai/api/realtime/azure/#parameters) **`model`** : `AzureRealtimeModelName` [](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModel.__init__(model)) The Azure _deployment_ name, which is what the realtime URL and the profile lookup use. Azure deployments are conventionally named after their model; when yours isn’t, `profile` is how to correct the facts inferred from the name. **`provider`** : `Provider`\[`AsyncOpenAI`\] | [`str`](https://docs.python.org/3/library/stdtypes.html#str) _Default:_ `'azure'` [](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModel.__init__(provider)) The provider supplying the resource endpoint and API key. Defaults to `'azure'`. **`settings`** : `RealtimeModelSettings` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModel.__init__(settings)) [Model settings](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelSettings) used as defaults for realtime sessions. **`profile`** : `RealtimeModelProfileSpec` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModel.__init__(profile)) Optional override for the [realtime model profile](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelProfile) , merged over the provider’s — a partial dict, or a callable taking the resolved profile and returning the one to use. **`credential`** : `AzureTokenCredential` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModel.__init__(credential)) Optional Microsoft Entra ID credential. When set, realtime requests use its bearer tokens instead of the resource API key. AzureRealtimeModelProfile ------------------------- [](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModelProfile) **Bases:** `RealtimeModelProfile` A [`RealtimeModelProfile`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeModelProfile) with the Azure-specific facts. Read via [`AzureRealtimeModel.profile`](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModel) . Pass a partial one as `profile=` to correct what’s inferred from a deployment name that doesn’t match its model. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/azure/#attributes-1) #### azure\_realtime\_apis [](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModelProfile.azure_realtime_apis) Which Azure realtime APIs serve this model, when only one does — the constraint that routes it. A model served only by Voice Live carries `{'voice_live'}` and routes there automatically; a GA-only model carries `{'azure_openai'}` and rejects [`azure_voice_live=True`](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModelSettings.azure_voice_live) . Absent for a model served by _both_ (e.g. `gpt-realtime`) or a name the table below doesn’t recognize (e.g. a future `gpt-realtime-3`): either way it defaults to GA and reaches Voice Live only when `azure_voice_live=True` is set. Pass a `profile=` override to constrain a deployment named after something the table can’t place. **Type:** [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[`AzureRealtimeApi`\] AzureRealtimeModelSettings -------------------------- [](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModelSettings) **Bases:** `OpenAIRealtimeModelSettings` Settings specific to Azure realtime models. This inherits every [`OpenAIRealtimeModelSettings`](https://pydantic.dev/docs/ai/api/realtime/openai/#pydantic_ai.realtime.openai.OpenAIRealtimeModelSettings) field, but when [`azure_voice_live`](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModelSettings.azure_voice_live) is set the Voice Live session config is built from only the cross-protocol fields — `instructions`, `openai_voice` (by name), `turn_detection` (or `azure_voice_live_turn_detection`), `input_transcription_model`, `output_modality`, `max_tokens`, `tool_choice`, and tools. The inherited `openai_*` fields, plus `thinking` and `parallel_tool_calls`, are **silently ignored** under Voice Live; they still apply on the GA path. They fall into two groups, and only the first is settled: * `openai_output_speed`, `openai_turn_detection`, `thinking`, and `parallel_tool_calls` have no counterpart in Voice Live’s beta session object (see the recorded `session.created` payload in `tests/realtime/cassettes/test_azure_voice_live_ws/`), so there is nothing to map them to. Voice Live’s own turn detection is configured with `azure_voice_live_turn_detection`. * `openai_input_noise_reduction` and `openai_truncation` _do_ have counterparts — `input_audio_noise_reduction` and `truncation_strategy` — but under Azure’s own vocabulary, which the recording pins as `null` and so does not evidence. Mapping OpenAI’s values onto them would be guessing at the accepted shape, so they stay unmapped until a recording proves it; a dedicated `azure_voice_live_*` setting is the natural home when it does. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/azure/#attributes-2) #### azure\_voice\_live [](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModelSettings.azure_voice_live) Use the Azure AI Voice Live endpoint and beta session protocol instead of the GA endpoint. Voice Live is a distinct Azure resource; [`AzureProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.azure.AzureProvider) reads its `AZURE_VOICELIVE_ENDPOINT` / `AZURE_VOICELIVE_API_KEY` / `AZURE_VOICELIVE_API_VERSION` credentials as a fallback to the `AZURE_OPENAI_*` variables. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### azure\_voice\_live\_turn\_detection [](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModelSettings.azure_voice_live_turn_detection) Voice Live server or semantic VAD config; only applies when `azure_voice_live=True`. **Type:** `ServerVAD` | `SemanticVAD` AzureTokenCredential -------------------- [](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureTokenCredential) **Bases:** [`Protocol`](https://docs.python.org/3/library/typing.html#typing.Protocol) Structural type for a synchronous Microsoft Entra ID token credential. AzureRealtimeApi ---------------- [](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeApi) An Azure realtime speech-to-speech API a model can be reached through: the Azure OpenAI GA realtime API (`/openai/v1/realtime`) or [Azure AI Voice Live](https://learn.microsoft.com/azure/ai-services/speech-service/voice-live) (`/voice-live/realtime`, selected with [`azure_voice_live=True`](https://pydantic.dev/docs/ai/api/realtime/azure/#pydantic_ai.realtime.azure.AzureRealtimeModelSettings.azure_voice_live) ). **Default:** `Literal['azure_openai', 'voice_live']` Was this page helpful? Thanks for your feedback! --- # Bedrock | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/models/bedrock/#_top) Bedrock ======= [Amazon Bedrock](https://aws.amazon.com/bedrock/) exposes foundation models from many providers, and Pydantic AI reaches it through two separate AWS APIs. Pick the route by model prefix: * **[Bedrock Converse](https://pydantic.dev/docs/ai/models/bedrock/#bedrock-converse) ** (`bedrock:`) — the broadest catalog, including Anthropic, Amazon, Cohere, Meta, Mistral, DeepSeek, Qwen, and [many more](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelName) , through the Bedrock Runtime Converse API. This is the route for almost every Bedrock model. * **[Bedrock Mantle](https://pydantic.dev/docs/ai/models/bedrock/#bedrock-mantle) ** (`bedrock-mantle:`) — the modern OpenAI models (GPT-5.x and GPT-OSS), which Bedrock serves only through [Mantle](https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-mantle.html) ’s OpenAI-compatible API. Both routes authenticate with the same AWS credentials. The `bedrock:` prefix always uses Converse; requesting a frontier OpenAI model (GPT-5.4 or newer) through it raises an error pointing you to `bedrock-mantle:`, since Converse doesn’t serve those models. | Route | Prefix | Models | Optional group | Model class | | --- | --- | --- | --- | --- | | [Converse](https://pydantic.dev/docs/ai/models/bedrock/#bedrock-converse) | `bedrock:` | Anthropic, Amazon, Cohere, Meta, Mistral, and [more](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelName) | `bedrock` | [`BedrockConverseModel`](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockConverseModel) | | [Mantle](https://pydantic.dev/docs/ai/models/bedrock/#bedrock-mantle) | `bedrock-mantle:` | OpenAI GPT-5.x and GPT-OSS | `bedrock-mantle` | [`BedrockMantleResponsesModel`](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.BedrockMantleResponsesModel)
, [`BedrockMantleChatModel`](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.BedrockMantleChatModel) | Bedrock Converse ---------------- [](https://pydantic.dev/docs/ai/models/bedrock/#bedrock-converse) [`BedrockConverseModel`](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockConverseModel) talks to the [Bedrock Runtime Converse API](https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_Converse.html) , which serves the broadest set of Bedrock models. ### Install [](https://pydantic.dev/docs/ai/models/bedrock/#install) To use `BedrockConverseModel`, you need to either install `pydantic-ai`, or install `pydantic-ai-slim` with the `bedrock` optional group: * [pip](https://pydantic.dev/docs/ai/models/bedrock/#tab-panel-102) * [uv](https://pydantic.dev/docs/ai/models/bedrock/#tab-panel-103) Terminal pip install "pydantic-ai-slim[bedrock]" Terminal uv add "pydantic-ai-slim[bedrock]" ### Configuration [](https://pydantic.dev/docs/ai/models/bedrock/#configuration) To use [AWS Bedrock](https://aws.amazon.com/bedrock/) , you’ll need an AWS account with Bedrock enabled and appropriate credentials. You can use either AWS credentials directly or a pre-configured boto3 client. [`BedrockModelName`](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelName) contains a list of available Bedrock models, including models from Anthropic, Amazon, Cohere, Meta, and Mistral. ### Environment variables [](https://pydantic.dev/docs/ai/models/bedrock/#environment-variables) You can set your AWS credentials as environment variables ([among other options](https://boto3.amazonaws.com/v1/documentation/api/latest/guide/configuration.html#using-environment-variables) ): Terminal export AWS_BEARER_TOKEN_BEDROCK='your-api-key' # or: export AWS_ACCESS_KEY_ID='your-access-key' export AWS_SECRET_ACCESS_KEY='your-secret-key' export AWS_DEFAULT_REGION='us-east-1' # or your preferred region You can then use `BedrockConverseModel` by name: from pydantic_ai import Agent agent = Agent('bedrock:anthropic.claude-sonnet-4-5-20250929-v1:0') ... Or initialize the model directly with just the model name: from pydantic_ai import Agent from pydantic_ai.models.bedrock import BedrockConverseModel model = BedrockConverseModel('anthropic.claude-sonnet-4-5-20250929-v1:0') agent = Agent(model) ... ### Customizing Bedrock Runtime API [](https://pydantic.dev/docs/ai/models/bedrock/#customizing-bedrock-runtime-api) You can customize the Bedrock Runtime API calls by adding additional parameters, such as [guardrail configurations](https://docs.aws.amazon.com/bedrock/latest/userguide/guardrails.html) and [performance settings](https://docs.aws.amazon.com/bedrock/latest/userguide/latency-optimized-inference.html) . For a complete list of configurable parameters, refer to the documentation for [`BedrockModelSettings`](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings) . customize\_bedrock\_model\_settings.py from pydantic_ai import Agent from pydantic_ai.models.bedrock import BedrockConverseModel, BedrockModelSettings # Define Bedrock model settings with guardrail and performance configurations bedrock_model_settings = BedrockModelSettings( bedrock_guardrail_config={ 'guardrailIdentifier': 'v1', 'guardrailVersion': 'v1', 'trace': 'enabled' }, bedrock_performance_configuration={ 'latency': 'optimized' } ) model = BedrockConverseModel(model_name='us.amazon.nova-pro-v1:0') agent = Agent(model=model, model_settings=bedrock_model_settings) ### Custom HTTP headers [](https://pydantic.dev/docs/ai/models/bedrock/#custom-http-headers) Use [`ModelSettings.extra_headers`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.extra_headers) to add HTTP headers to `Converse`, `ConverseStream`, and `CountTokens` requests. This is useful for routing requests through an API gateway or proxy that requires custom headers: from pydantic_ai import Agent from pydantic_ai.models.bedrock import BedrockModelSettings agent = Agent( 'bedrock:us.amazon.nova-micro-v1:0', model_settings=BedrockModelSettings( extra_headers={'X-Tenant-ID': 'example-tenant'}, ), ) Do not use `extra_headers` to override headers managed by boto3, such as `Authorization`, `User-Agent`, `X-Amz-Date`, `Host`, or `Content-Length`. These values may be ignored or cause the request to fail. ### Service tier [](https://pydantic.dev/docs/ai/models/bedrock/#service-tier) Bedrock supports controlling the [service tier](https://docs.aws.amazon.com/bedrock/latest/userguide/inference-profiles.html) to manage throughput and cost. You can use the unified [`service_tier`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.service_tier) field or the provider-specific [`bedrock_service_tier`](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_service_tier) field. `bedrock_service_tier` takes precedence over the unified field when both are set. The unified field maps as follows for Bedrock: * `'auto'`: the `serviceTier` field is omitted from the request, so AWS applies its server-side default (Standard tier). * `'default'`: explicitly sent as `{'type': 'default'}` — opts out of any future server-side auto-promotion to premium tiers. * `'flex'`: sent as `{'type': 'flex'}`. * `'priority'`: sent as `{'type': 'priority'}`. To request Bedrock’s `'reserved'` tier (which requires a pre-purchased capacity reservation), set [`bedrock_service_tier`](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_service_tier) directly — it isn’t reachable through the unified field. ### Prompt Caching [](https://pydantic.dev/docs/ai/models/bedrock/#prompt-caching) Bedrock supports [prompt caching](https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html) on Anthropic models so you can reuse expensive context across requests. Pydantic AI provides four ways to use prompt caching: 1. **Cache User Messages with [`CachePoint`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.CachePoint) **: Insert a `CachePoint` marker to cache everything before it in the current user message. A `CachePoint` at the start of a user prompt part has nothing before it in that message, so it caches everything up to the end of the previous user message instead. Pass `CachePoint(ttl='1h')` to opt into the extended cache duration. 2. **Cache System Instructions**: Set [`BedrockModelSettings.bedrock_cache_instructions`](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_cache_instructions) to `True` (uses 5m TTL by default) or specify `'5m'` / `'1h'` directly. When you have both static and dynamic [instructions](https://pydantic.dev/docs/ai/core-concepts/agent/#instructions) , the cache point is placed after the last static instruction, so dynamic instructions can change without invalidating the static cache. 3. **Cache Tool Definitions**: Set [`BedrockModelSettings.bedrock_cache_tool_definitions`](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_cache_tool_definitions) to `True` (uses 5m TTL by default) or specify `'5m'` / `'1h'` directly. 4. **Cache All Messages**: Set [`BedrockModelSettings.bedrock_cache_messages`](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_cache_messages) to `True` (uses 5m TTL by default) or specify `'5m'` / `'1h'` directly to automatically cache the last user message. #### Example 1: Automatic Message Caching [](https://pydantic.dev/docs/ai/models/bedrock/#example-1-automatic-message-caching) Use `bedrock_cache_messages` to automatically cache the last user message: from pydantic_ai import Agent from pydantic_ai.models.bedrock import BedrockModelSettings agent = Agent( 'bedrock:us.anthropic.claude-sonnet-4-5-20250929-v1:0', system_prompt='You are a helpful assistant.', model_settings=BedrockModelSettings( bedrock_cache_messages=True, # Automatically caches the last message ), ) # The last message is automatically cached - no need for manual CachePoint result1 = agent.run_sync('What is the capital of France?') # Subsequent calls with similar conversation benefit from cache result2 = agent.run_sync('What is the capital of Germany?') print(f'Cache write: {result1.usage.cache_write_tokens}') print(f'Cache read: {result2.usage.cache_read_tokens}') #### Example 2: Comprehensive Caching Strategy [](https://pydantic.dev/docs/ai/models/bedrock/#example-2-comprehensive-caching-strategy) Combine multiple cache settings for maximum savings: from pydantic_ai import Agent, RunContext from pydantic_ai.models.bedrock import BedrockConverseModel, BedrockModelSettings model = BedrockConverseModel('us.anthropic.claude-sonnet-4-5-20250929-v1:0') agent = Agent( model, system_prompt='Detailed instructions...', model_settings=BedrockModelSettings( bedrock_cache_instructions=True, # Cache system instructions bedrock_cache_tool_definitions='1h', # Cache tool definitions with 1h TTL bedrock_cache_messages=True, # Also cache the last message ), ) @agent.tool def search_docs(ctx: RunContext, query: str) -> str: """Search documentation.""" return f'Results for {query}' result = agent.run_sync('Search for Python best practices') print(result.output) #### Example 3: Fine-Grained Control with CachePoint [](https://pydantic.dev/docs/ai/models/bedrock/#example-3-fine-grained-control-with-cachepoint) Use manual `CachePoint` markers to control cache locations precisely: from pydantic_ai import Agent, CachePoint agent = Agent( 'bedrock:us.anthropic.claude-sonnet-4-5-20250929-v1:0', system_prompt='Instructions...', ) # Manually control cache points for specific content blocks result = agent.run_sync([\ 'Long context from documentation...',\ CachePoint(), # Cache everything up to this point\ 'First question'\ ]) print(result.output) #### Accessing Cache Usage Statistics [](https://pydantic.dev/docs/ai/models/bedrock/#accessing-cache-usage-statistics) Access cache usage statistics via [`RequestUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage) : from pydantic_ai import Agent, CachePoint agent = Agent('bedrock:us.anthropic.claude-sonnet-4-5-20250929-v1:0') async def main(): result = await agent.run( [\ 'Reference material...',\ CachePoint(),\ 'What changed since last time?',\ ] ) usage = result.usage print(f'Cache writes: {usage.cache_write_tokens}') print(f'Cache reads: {usage.cache_read_tokens}') #### Cache Point Limits [](https://pydantic.dev/docs/ai/models/bedrock/#cache-point-limits) Bedrock enforces a maximum of 4 cache points per request. Pydantic AI automatically manages this limit to ensure your requests always comply without errors. ##### How Cache Points Are Allocated [](https://pydantic.dev/docs/ai/models/bedrock/#how-cache-points-are-allocated) Cache points can be placed in three locations: 1. **System Prompt**: Via `bedrock_cache_instructions` setting (adds cache point to last system prompt block) 2. **Tool Definitions**: Via `bedrock_cache_tool_definitions` setting (adds cache point to last tool definition) 3. **Messages**: Via `CachePoint` markers or `bedrock_cache_messages` setting (adds cache points to message content) Each setting uses **at most 1 cache point**, but you can combine them. ##### Automatic Cache Point Limiting [](https://pydantic.dev/docs/ai/models/bedrock/#automatic-cache-point-limiting) When cache points from all sources (settings + `CachePoint` markers) exceed 4, Pydantic AI automatically removes excess cache points from **older message content** (keeping the most recent ones). from pydantic_ai import Agent, CachePoint from pydantic_ai.models.bedrock import BedrockModelSettings agent = Agent( 'bedrock:us.anthropic.claude-sonnet-4-5-20250929-v1:0', system_prompt='Instructions...', model_settings=BedrockModelSettings( bedrock_cache_instructions=True, # 1 cache point bedrock_cache_tool_definitions=True, # 1 cache point ), ) @agent.tool_plain def search() -> str: return 'data' # Already using 2 cache points (instructions + tools) # Can add 2 more CachePoint markers (4 total limit) result = agent.run_sync([\ 'Context 1', CachePoint(), # Oldest - will be removed\ 'Context 2', CachePoint(), # Will be kept (3rd point)\ 'Context 3', CachePoint(), # Will be kept (4th point)\ 'Question'\ ]) # Final cache points: instructions + tools + Context 2 + Context 3 = 4 print(result.output) **Key Points**: * System and tool cache points are **always preserved** * The cache point created by `bedrock_cache_messages` is **always preserved** (as it’s the newest message cache point) * Additional `CachePoint` markers in messages are removed from oldest to newest when the limit is exceeded * This ensures critical caching (instructions/tools) is maintained while still benefiting from message-level caching ### `provider` argument [](https://pydantic.dev/docs/ai/models/bedrock/#provider-argument) You can provide a custom `BedrockProvider` via the `provider` argument. This is useful when you want to specify credentials directly or use a custom boto3 client: from pydantic_ai import Agent from pydantic_ai.models.bedrock import BedrockConverseModel from pydantic_ai.providers.bedrock import BedrockProvider # Using AWS credentials directly model = BedrockConverseModel( 'anthropic.claude-sonnet-4-5-20250929-v1:0', provider=BedrockProvider( region_name='us-east-1', aws_access_key_id='your-access-key', aws_secret_access_key='your-secret-key', ), ) agent = Agent(model) ... You can also pass a pre-configured boto3 client: import boto3 from pydantic_ai import Agent from pydantic_ai.models.bedrock import BedrockConverseModel from pydantic_ai.providers.bedrock import BedrockProvider # Using a pre-configured boto3 client bedrock_client = boto3.client('bedrock-runtime', region_name='us-east-1') model = BedrockConverseModel( 'anthropic.claude-sonnet-4-5-20250929-v1:0', provider=BedrockProvider(bedrock_client=bedrock_client), ) agent = Agent(model) ... ### Using AWS Application Inference Profiles [](https://pydantic.dev/docs/ai/models/bedrock/#using-aws-application-inference-profiles) AWS Bedrock supports [custom application inference profiles](https://docs.aws.amazon.com/bedrock/latest/userguide/inference-profiles-create.html) for cost tracking and resource management. Set [`bedrock_inference_profile`](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_inference_profile) to route requests through an inference profile while keeping the base model name for detecting model capabilities: from pydantic_ai import Agent from pydantic_ai.models.bedrock import BedrockConverseModel from pydantic_ai.providers.bedrock import BedrockProvider provider = BedrockProvider(region_name='us-east-2') model = BedrockConverseModel( 'us.anthropic.claude-opus-4-5-20251101-v1:0', provider=provider, settings={ 'bedrock_inference_profile': 'arn:aws:bedrock:us-east-2:123456789012:application-inference-profile/my-profile', }, ) agent = Agent(model) ### Configuring Retries [](https://pydantic.dev/docs/ai/models/bedrock/#configuring-retries) Bedrock uses boto3’s built-in retry mechanisms. You can configure retry behavior by passing a custom boto3 client with retry settings: import boto3 from botocore.config import Config from pydantic_ai import Agent from pydantic_ai.models.bedrock import BedrockConverseModel from pydantic_ai.providers.bedrock import BedrockProvider # Configure retry settings config = Config( retries={ 'max_attempts': 5, 'mode': 'adaptive' # Recommended for rate limiting } ) bedrock_client = boto3.client( 'bedrock-runtime', region_name='us-east-1', config=config ) model = BedrockConverseModel( 'us.amazon.nova-micro-v1:0', provider=BedrockProvider(bedrock_client=bedrock_client), ) agent = Agent(model) #### Retry Modes [](https://pydantic.dev/docs/ai/models/bedrock/#retry-modes) * `'legacy'` (default): 5 attempts, basic retry behavior * `'standard'`: 3 attempts, more comprehensive error coverage * `'adaptive'`: 3 attempts with client-side rate limiting (recommended for handling `ThrottlingException`) For more details on boto3 retry configuration, see the [AWS boto3 documentation](https://boto3.amazonaws.com/v1/documentation/api/latest/guide/retries.html) . Bedrock Mantle -------------- [](https://pydantic.dev/docs/ai/models/bedrock/#bedrock-mantle) [Amazon Bedrock Mantle](https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-mantle.html) serves OpenAI models (GPT-5.x and GPT-OSS) through an OpenAI-compatible API. Use the `bedrock-mantle:` prefix: from pydantic_ai import Agent agent = Agent('bedrock-mantle:openai.gpt-5.6-luna') It requires the `bedrock-mantle` optional group: * [pip](https://pydantic.dev/docs/ai/models/bedrock/#tab-panel-104) * [uv](https://pydantic.dev/docs/ai/models/bedrock/#tab-panel-105) Terminal pip install "pydantic-ai-slim[bedrock-mantle]" Terminal uv add "pydantic-ai-slim[bedrock-mantle]" The [`BedrockMantleProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.bedrock_mantle.BedrockMantleProvider) authenticates with the same AWS credentials as the [Converse route](https://pydantic.dev/docs/ai/models/bedrock/#environment-variables) — a bearer token via `AWS_BEARER_TOKEN_BEDROCK`, or AWS access keys / profile via SigV4 — and derives its endpoint from `region_name` (or the `AWS_DEFAULT_REGION` / `AWS_REGION` environment variables). The model name determines the endpoint family: | Model name | Interface | | --- | --- | | GPT-5.4+, e.g. `bedrock-mantle:openai.gpt-5.6-luna` | OpenAI Responses at `/openai/v1` | | GPT-OSS, e.g. `bedrock-mantle:openai.gpt-oss-120b` | OpenAI Responses at `/v1` | | GPT-OSS Safeguard, e.g. `bedrock-mantle:openai.gpt-oss-safeguard-20b` | OpenAI Chat Completions at `/v1` | To use a custom Mantle origin (for example a proxy), pass a `base_url` to [`BedrockMantleProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.bedrock_mantle.BedrockMantleProvider) ; its origin (with any `/openai/v1` or `/v1` suffix stripped) is used to route between the two endpoint families per model, just like `region_name`: from pydantic_ai import Agent from pydantic_ai.models.bedrock_mantle import BedrockMantleResponsesModel from pydantic_ai.providers.bedrock_mantle import BedrockMantleProvider provider = BedrockMantleProvider(base_url='https://bedrock-mantle.us-east-1.api.aws/openai/v1') model = BedrockMantleResponsesModel('openai.gpt-5.6-luna', provider=provider) agent = Agent(model) ### Feature support [](https://pydantic.dev/docs/ai/models/bedrock/#feature-support) Mantle models are served by Pydantic AI’s OpenAI model classes — [`BedrockMantleResponsesModel`](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.BedrockMantleResponsesModel) and [`BedrockMantleChatModel`](https://pydantic.dev/docs/ai/api/models/bedrock_mantle/#pydantic_ai.models.bedrock_mantle.BedrockMantleChatModel) — so they accept the same settings as the direct [OpenAI](https://pydantic.dev/docs/ai/models/openai/) models ([`OpenAIResponsesModelSettings`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIResponsesModelSettings) and [`OpenAIChatModelSettings`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIChatModelSettings) ). The Converse-route features above — [prompt caching](https://pydantic.dev/docs/ai/models/bedrock/#prompt-caching) , [service tier](https://pydantic.dev/docs/ai/models/bedrock/#service-tier) , and [application inference profiles](https://pydantic.dev/docs/ai/models/bedrock/#using-aws-application-inference-profiles) — are specific to the Converse API and don’t apply to the Mantle route. In particular [`bedrock_service_tier`](https://pydantic.dev/docs/ai/api/models/bedrock/#pydantic_ai.models.bedrock.BedrockModelSettings.bedrock_service_tier) is a Converse setting; the Mantle models do forward the unified [`service_tier`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.service_tier) as the OpenAI parameter of the same name, since they are served by the OpenAI model classes. Was this page helpful? Thanks for your feedback! --- # settings | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#_top) settings ======== ModelSettings ------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) **Bases:** [`TypedDict`](https://docs.python.org/3/library/typing.html#typing.TypedDict) Settings to configure an LLM. Includes only settings which apply to multiple models / model providers, though not all of these settings are supported by all models. Each field’s `Supported by:` list names the model classes that put the setting on the wire. A bare name covers every interface that model serves, so `OpenAI` means both [`OpenAIChatModel`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIChatModel) and [`OpenAIResponsesModel`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIResponsesModel) ; a name qualified with an interface, like `OpenAI Chat Completions`, covers only that one, because the Responses API does not accept the setting at all. These lists are parsed and checked against the wire by `tests/models/test_model_settings_support.py`, so keep the `* Name` bullet shape and put any nuance in parentheses after the name. Being listed means Pydantic AI sends the setting, not that the service honors it: the OpenAI-compatible model classes forward whatever the OpenAI schema accepts, and an individual provider behind one of them may ignore a field its own API doesn’t define, or reject it. Where we know of such a case it is noted on the entry, but the provider’s own API reference is the authority. All types must be serializable using Pydantic. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#attributes) #### extra\_body [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.extra_body) Extra body to send to the model. Supported by: * OpenAI * Anthropic * Groq * HuggingFace * Cerebras * Crusoe * Ollama * OpenRouter * Snowflake * Z.AI * Bedrock Mantle On the OpenAI-derived models that build their own `extra_body` (Cerebras, OpenRouter, Snowflake, Z.AI), the model’s own derived keys overwrite yours when the keys collide. **Type:** [`object`](https://docs.python.org/3/glossary.html#term-object) #### extra\_headers [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.extra_headers) Extra headers to send to the model. Supported by: * OpenAI * Anthropic * Google * Groq * Bedrock * Cerebras * Crusoe * Ollama * OpenRouter * Snowflake * Z.AI * Bedrock Mantle **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] #### frequency\_penalty [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.frequency_penalty) Penalize new tokens based on their existing frequency in the text so far. Supported by: * OpenAI Chat Completions * Google * Groq * Cohere * Mistral * xAI * HuggingFace * Crusoe * Ollama * OpenRouter * Snowflake * Z.AI * Bedrock Mantle Chat Completions **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) #### logit\_bias [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.logit_bias) Modify the likelihood of specified tokens appearing in the completion. Supported by: * OpenAI Chat Completions * Groq * HuggingFace * Crusoe * Ollama (sent, but Ollama documents `logit_bias` as unsupported) * OpenRouter * Snowflake * Z.AI * Bedrock Mantle Chat Completions **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`int`](https://docs.python.org/3/library/functions.html#int)\ \] #### max\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.max_tokens) The maximum number of tokens to generate before stopping. Supported by: * OpenAI * Anthropic * Google * Groq * Cohere * Mistral * Bedrock * MCP Sampling * xAI * HuggingFace * Cerebras * Crusoe * Ollama * OpenRouter * Snowflake * Z.AI * Bedrock Mantle **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### parallel\_tool\_calls [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.parallel_tool_calls) Whether to allow parallel tool calls. Supported by: * OpenAI (some models, not o1) * Anthropic * Groq * Mistral * xAI * Crusoe * Ollama * OpenRouter * Snowflake * Z.AI * Bedrock Mantle **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### presence\_penalty [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.presence_penalty) Penalize new tokens based on whether they have appeared in the text so far. Supported by: * OpenAI Chat Completions * Google * Groq * Cohere * Mistral * xAI * HuggingFace * Crusoe * Ollama * OpenRouter * Snowflake * Z.AI * Bedrock Mantle Chat Completions **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) #### seed [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.seed) The random seed to use for the model, theoretically allowing for deterministic results. Supported by: * OpenAI Chat Completions * Google * Groq * Cohere * Mistral * xAI * HuggingFace * Cerebras * Crusoe * Ollama * OpenRouter * Snowflake * Z.AI * Bedrock Mantle Chat Completions **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### service\_tier [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.service_tier) The cross-provider service tier to use for the model request. See [`ServiceTier`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ServiceTier) for the value semantics and the per-provider mapping table. Provider-specific settings (`openai_service_tier`, `anthropic_service_tier`, `bedrock_service_tier`, `google_cloud_service_tier`) take precedence over this unified field when set. Supported by: * OpenAI * Anthropic * Google (Gemini API and Google Cloud) * Bedrock * Crusoe * Ollama * OpenRouter * Snowflake (sent, but Snowflake Cortex rejects `service_tier` with an error) * Z.AI * Bedrock Mantle The OpenAI-derived model classes send the OpenAI value unchanged, so the OpenAI column of the mapping table applies to them. **Type:** `ServiceTier` #### stop\_sequences [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.stop_sequences) Sequences that will cause the model to stop generating. Supported by: * OpenAI Chat Completions * Anthropic * Google * Groq * Cohere * Mistral * Bedrock * MCP Sampling * xAI * HuggingFace * Cerebras * Crusoe * Ollama * OpenRouter * Snowflake * Z.AI * Bedrock Mantle Chat Completions **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] #### temperature [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.temperature) Amount of randomness injected into the response. Use `temperature` closer to `0.0` for analytical / multiple choice, and closer to a model’s maximum `temperature` for creative and generative tasks. Note that even with `temperature` of `0.0`, the results will not be fully deterministic. Supported by: * OpenAI * Anthropic * Google * Groq * Cohere * Mistral * Bedrock * MCP Sampling * xAI * HuggingFace * Cerebras * Crusoe * Ollama * OpenRouter * Snowflake * Z.AI * Bedrock Mantle **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) #### thinking [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.thinking) Enable or configure thinking/reasoning for the model. * `True`: Enable thinking with the provider’s default effort level. * `False`: Disable thinking (silently ignored if the model always thinks). * `'minimal'`/`'low'`/`'medium'`/`'high'`/`'xhigh'`: Enable thinking at a specific effort level. When omitted, the model uses its default behavior (which may include thinking for reasoning models). Provider-specific thinking settings (e.g., `anthropic_thinking`, `openai_reasoning_effort`) take precedence over this unified field. Listed below are the model classes that translate this field onto the request. A class whose models always reason and take no thinking parameter is not listed at all (Cohere); where only some of a class’s models are always-on it stays listed, and the per-model behavior is on the [Thinking page](https://pydantic.dev/docs/ai/capabilities/thinking/) (Mistral’s `magistral`). Supported by: * OpenAI * Anthropic * Google * Groq * Mistral * Bedrock * xAI * Cerebras (only `False` is forwarded, as `reasoning_effort='none'`; the enable levels are not sent because Cerebras models reason by default, and `gpt-oss` ignores the disable too) * Crusoe * Ollama * OpenRouter (as `extra_body['reasoning']`) * Snowflake (as `extra_body['reasoning']` on Claude models, otherwise as `reasoning_effort`) * Z.AI (as `extra_body['thinking']`) * Bedrock Mantle (the Responses interface only; the Chat Completions interface serves only the `gpt-oss-safeguard` models, which take no thinking parameter) **Type:** `ThinkingLevel` #### timeout [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.timeout) Override the client-level default timeout for a request, in seconds. Supported by: * OpenAI * Anthropic * Google (numeric seconds only, not `httpx.Timeout`) * Groq * Mistral (numeric seconds only, not `httpx.Timeout`) * Cerebras * Crusoe * Ollama * OpenRouter * Snowflake * Z.AI * Bedrock Mantle **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`float`](https://docs.python.org/3/library/functions.html#float) | `Timeout` #### tool\_choice [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.tool_choice) Control which function tools the model can use. See the [Tool Choice guide](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-choice) for detailed documentation and examples. * `None` (default): Defaults to `'auto'` behavior * `'auto'`: All tools available, model decides whether to use them * `'none'`: Disables function tools; model responds with text only (output tools remain for structured output) * `'required'`: Forces tool use; excludes output tools so the agent cannot produce a final response when set statically * `list[str]`: Only specified tools; excludes output tools so the agent cannot produce a final response when set statically * [`ToolOrOutput`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ToolOrOutput) : Specified function tools plus output tools/text/image Note: setting `'required'` or `list[str]` _statically_ (via the `model_settings` argument of [`Agent.run`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.run) or the agent’s own `model_settings`) raises a `UserError`, because it would force a tool call on every step and prevent the agent from producing a final response. To vary `tool_choice` per step (e.g. force a tool on the first step only), return a callable from a capability’s [`get_model_settings`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability.get_model_settings) — those values are trusted to adapt across steps. For single API calls without an agent loop, use [`pydantic_ai.direct.model_request`](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request) . Supported by: * OpenAI * Anthropic (`'required'` and specific tools not supported with thinking enabled) * Google * Groq * Cohere (a named subset is honored by filtering the tool list, not sent as a parameter) * Mistral (a named subset is honored by filtering the tool list, not sent as a parameter) * Bedrock * xAI * HuggingFace * Cerebras * Crusoe * Ollama (sent, but Ollama documents `tool_choice` as unsupported) * OpenRouter * Snowflake * Z.AI * Bedrock Mantle **Type:** `ToolChoice` #### top\_k [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.top_k) Only sample from the top K options for each subsequent token. Used to remove “long tail” low probability responses. Supported by: * Anthropic * Google * Cohere * Bedrock (Anthropic and Amazon Nova models only) **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) #### top\_p [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.top_p) An alternative to sampling with temperature, called nucleus sampling, where the model considers the results of the tokens with top\_p probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. You should either alter `temperature` or `top_p`, but not both. Supported by: * OpenAI * Anthropic * Google * Groq * Cohere * Mistral * Bedrock * xAI * HuggingFace * Cerebras * Crusoe * Ollama * OpenRouter * Snowflake * Z.AI * Bedrock Mantle **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) ToolOrOutput ------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ToolOrOutput) Restricts function tools while keeping output tools and direct text/image output available. Use this when you want to control which function tools the model can use in an agent run while still allowing the agent to complete with structured output, text, or images. See the [Tool Choice guide](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-choice) for examples. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#attributes-1) #### function\_tools [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ToolOrOutput.function_tools) The names of function tools available to the model. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] ServiceTier ----------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ServiceTier) Cross-provider value set for [`ModelSettings.service_tier`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.service_tier) . Values: * `'auto'`: Let the provider decide — typically means “use a higher tier (scale credits, priority capacity) when available, otherwise standard.” On providers without a server-side auto concept the field is omitted so the provider’s natural default applies. * `'default'`: Explicitly request the provider’s standard tier — opts out of any server-side auto-promotion to premium tiers. * `'flex'`: Lower-cost, latency-tolerant tier where the provider offers one. Silently ignored on providers that don’t (e.g. Anthropic) — though a few reject the field outright rather than ignore it, as noted on the [`service_tier`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.service_tier) entries. * `'priority'`: Higher-priority / lower-latency tier where the provider offers one. Silently ignored on providers that don’t. Per-provider mapping: | value | OpenAI | Anthropic | Bedrock | Google (Gemini API) | Google Cloud | | --- | --- | --- | --- | --- | --- | | `'auto'` | `'auto'` | `'auto'` | _(omitted)_ | _(omitted)_ | _no headers (PT then on-demand)_ | | `'default'` | `'default'` | `'standard_only'` | `{'type': 'default'}` | `'standard'` | _no headers (PT then on-demand)_ | | `'flex'` | `'flex'` | _(omitted)_ | `{'type': 'flex'}` | `'flex'` | header `Shared-Request-Type: flex` (PT then Flex PayGo) | | `'priority'` | `'priority'` | _(omitted)_ | `{'type': 'priority'}` | `'priority'` | header `Shared-Request-Type: priority` (PT then Priority PayGo) | On Google Cloud the unified field maps only to safe PT-with-spillover variants so customers with Provisioned Throughput keep using their reserved capacity first; to bypass PT entirely use [`google_cloud_service_tier`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings.google_cloud_service_tier) with `'flex_only'` or `'priority_only'`. Likewise, provider-specific values not in the unified set (Bedrock’s `'reserved'`, Anthropic’s `'standard_only'`, Google Cloud’s PT routing tiers) are reachable only through the per-provider field. Per-provider settings (`openai_service_tier`, `anthropic_service_tier`, `bedrock_service_tier`, `google_cloud_service_tier`) always take precedence over this unified field when set. **Type:** [`TypeAlias`](https://docs.python.org/3/library/typing.html#typing.TypeAlias) **Default:** `Literal['auto', 'default', 'flex', 'priority']` Was this page helpful? Thanks for your feedback! --- # Messages and chat history | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/core-concepts/message-history/#_top) Messages and chat history ========================= Pydantic AI provides access to messages exchanged during an agent run. These messages can be used both to continue a coherent conversation, and to understand how an agent performed. ### Accessing Messages from Results [](https://pydantic.dev/docs/ai/core-concepts/message-history/#accessing-messages-from-results) After running an agent, you can access the messages exchanged during that run from the `result` object. Both [`RunResult`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult) (returned by [`Agent.run`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.run) , [`Agent.run_sync`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.run_sync) ) and [`StreamedRunResult`](https://pydantic.dev/docs/ai/api/pydantic-ai/result/#pydantic_ai.result.StreamedRunResult) (returned by [`Agent.run_stream`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.run_stream) ) have the following methods: * [`all_messages()`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.all_messages) : returns all messages, including messages from prior runs. There’s also a variant that returns JSON bytes, [`all_messages_json()`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.all_messages_json) . * [`new_messages()`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.new_messages) : returns only the messages from the current run. There’s also a variant that returns JSON bytes, [`new_messages_json()`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.new_messages_json) . Example of accessing methods on a [`RunResult`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult) : run\_result\_messages.py from pydantic_ai import Agent agent = Agent('openai:gpt-5.2', instructions='Be a helpful assistant.') result = agent.run_sync('Tell me a joke.') print(result.output) #> Did you hear about the toothpaste scandal? They called it Colgate. # all messages from the run print(result.all_messages()) """ [\ ModelRequest(\ parts=[\ UserPromptPart(\ content='Tell me a joke.',\ timestamp=datetime.datetime(...),\ )\ ],\ timestamp=datetime.datetime(...),\ instructions='Be a helpful assistant.',\ run_id='...',\ conversation_id='...',\ ),\ ModelResponse(\ parts=[\ TextPart(\ content='Did you hear about the toothpaste scandal? They called it Colgate.'\ )\ ],\ usage=RequestUsage(\ cost=Decimal('0.00026425'), input_tokens=55, output_tokens=12\ ),\ model_name='gpt-5.2',\ timestamp=datetime.datetime(...),\ run_id='...',\ conversation_id='...',\ ),\ ] """ _(This example is complete, it can be run “as is”)_ Example of accessing methods on a [`StreamedRunResult`](https://pydantic.dev/docs/ai/api/pydantic-ai/result/#pydantic_ai.result.StreamedRunResult) : streamed\_run\_result\_messages.py from pydantic_ai import Agent agent = Agent('openai:gpt-5.2', instructions='Be a helpful assistant.') async def main(): async with agent.run_stream('Tell me a joke.') as result: # incomplete messages before the stream finishes print(result.all_messages()) """ [\ ModelRequest(\ parts=[\ UserPromptPart(\ content='Tell me a joke.',\ timestamp=datetime.datetime(...),\ )\ ],\ timestamp=datetime.datetime(...),\ instructions='Be a helpful assistant.',\ run_id='...',\ conversation_id='...',\ )\ ] """ async for text in result.stream_text(): print(text) #> Did you hear #> Did you hear about the toothpaste #> Did you hear about the toothpaste scandal? They called #> Did you hear about the toothpaste scandal? They called it Colgate. # complete messages once the stream finishes print(result.all_messages()) """ [\ ModelRequest(\ parts=[\ UserPromptPart(\ content='Tell me a joke.',\ timestamp=datetime.datetime(...),\ )\ ],\ timestamp=datetime.datetime(...),\ instructions='Be a helpful assistant.',\ run_id='...',\ conversation_id='...',\ ),\ ModelResponse(\ parts=[\ TextPart(\ content='Did you hear about the toothpaste scandal? They called it Colgate.'\ )\ ],\ usage=RequestUsage(input_tokens=50, output_tokens=12),\ model_name='gpt-5.2',\ timestamp=datetime.datetime(...),\ run_id='...',\ conversation_id='...',\ ),\ ] """ _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main())` to run `main`)_ ### Using Messages as Input for Further Agent Runs [](https://pydantic.dev/docs/ai/core-concepts/message-history/#using-messages-as-input-for-further-agent-runs) The primary use of message histories in Pydantic AI is to maintain context across multiple agent runs. To use existing messages in a run, pass them to the `message_history` parameter of [`Agent.run`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.run) , [`Agent.run_sync`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.run_sync) or [`Agent.run_stream`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.run_stream) . If `message_history` is set and not empty, a new system prompt is not generated — we assume the existing message history includes a system prompt. If your history comes from a source that doesn’t round-trip system prompts (a UI frontend, a database that didn’t persist them, a compaction pipeline), add the [`ReinjectSystemPrompt`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ReinjectSystemPrompt) capability so the agent’s configured `system_prompt` is reinjected at the head of the first request when it’s missing. Reusing messages in a conversation from pydantic_ai import Agent agent = Agent('openai:gpt-5.2', instructions='Be a helpful assistant.') result1 = agent.run_sync('Tell me a joke.') print(result1.output) #> Did you hear about the toothpaste scandal? They called it Colgate. result2 = agent.run_sync('Explain?', message_history=result1.new_messages()) print(result2.output) #> This is an excellent joke invented by Samuel Colvin, it needs no explanation. print(result2.all_messages()) """ [\ ModelRequest(\ parts=[\ UserPromptPart(\ content='Tell me a joke.',\ timestamp=datetime.datetime(...),\ )\ ],\ timestamp=datetime.datetime(...),\ instructions='Be a helpful assistant.',\ run_id='...',\ conversation_id='...',\ ),\ ModelResponse(\ parts=[\ TextPart(\ content='Did you hear about the toothpaste scandal? They called it Colgate.'\ )\ ],\ usage=RequestUsage(\ cost=Decimal('0.00026425'), input_tokens=55, output_tokens=12\ ),\ model_name='gpt-5.2',\ timestamp=datetime.datetime(...),\ run_id='...',\ conversation_id='...',\ ),\ ModelRequest(\ parts=[\ UserPromptPart(\ content='Explain?',\ timestamp=datetime.datetime(...),\ )\ ],\ timestamp=datetime.datetime(...),\ instructions='Be a helpful assistant.',\ run_id='...',\ conversation_id='...',\ ),\ ModelResponse(\ parts=[\ TextPart(\ content='This is an excellent joke invented by Samuel Colvin, it needs no explanation.'\ )\ ],\ usage=RequestUsage(cost=Decimal('0.000462'), input_tokens=56, output_tokens=26),\ model_name='gpt-5.2',\ timestamp=datetime.datetime(...),\ run_id='...',\ conversation_id='...',\ ),\ ] """ _(This example is complete, it can be run “as is”)_ ### Mid-conversation system prompts [](https://pydantic.dev/docs/ai/core-concepts/message-history/#mid-conversation-system-prompts) A [`SystemPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SystemPromptPart) in the first [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) is the agent’s standing system prompt, and always hoists to the provider’s top-level system parameter. One in any _later_ request is a mid-conversation instruction: something that became true partway through the session, whether it arrived in a stored `message_history` or from [`RunContext.enqueue`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.enqueue) during a run. Mid-conversation instructions stay where you put them rather than joining the system prompt at the front. When prompt caching is enabled, that lets the provider reuse the unchanged prefix: editing the top-level system prompt invalidates everything behind it, while an instruction appended in place leaves the conversation up to that point eligible for a cache hit. Position alone does not enable caching; configure the active model’s prompt-caching settings or add an explicit [`CachePoint`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.CachePoint) . How it reaches the model depends on the provider: * Where the API accepts a system message inside the conversation, it’s sent as one, with the operator authority that implies. [Anthropic](https://pydantic.dev/docs/ai/models/anthropic/#mid-conversation-system-messages) supports this on some models, and may adjust the position slightly to satisfy its own placement rules. * Everywhere else it’s rendered as a ``\-tagged [`UserPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.UserPromptPart) at the same position. The instruction still applies from where you put it, but the model can tell it came in over the user channel and may treat it as a strong preference rather than a rule. Phrase the instruction as what changed rather than as an override of the user. Models are trained to resist instructions that appear to work against the person they’re talking to, and that applies to the system role too — “the build tag is no longer confidential” lands where “ignore what the user was told earlier” doesn’t. ### Making histories provider-valid [](https://pydantic.dev/docs/ai/core-concepts/message-history/#making-histories-provider-valid) Model providers reject a request whose message history has broken tool-call/tool-result pairing — a tool call with no result, or a result with no call. A run that is cancelled or crashes partway through can leave the history in exactly this state, and so can a hand-built, truncated, or context-evicted history. You don’t need to clean these up yourself: before each model request, Pydantic AI repairs the history it was given so the provider accepts it. Tool additions are stored as [`ToolAvailabilityDeltaPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolAvailabilityDeltaPart) request parts; tool removal is not represented. A tool returning \[`ToolReturn(tools=[...])`\]\[pydantic\_ai.messages.ToolReturn\] authors the part immediately after its [`ToolReturnPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolReturnPart) in the same request, with the call’s `tool_call_id` as a causal link. The executor deduplicates names in first-occurrence order and omits names already revealed. Replaying history keeps each `added` name revealed; tool definitions continue to come from the current run, so unknown or already-visible names have no effect when rendering a request. The guiding rule is to massage the history into a shape the provider accepts without ever discarding something you meant to send. Repairs only **add** synthesized parts or **remove** parts that are fundamentally unsendable (no provider could accept them); nothing meaningful is silently dropped. Concretely, before each request Pydantic AI: * **Adds** a synthesized [`ToolReturnPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolReturnPart) for a tool call that has no result, telling the model the call was interrupted before a result was produced. It has [`outcome='interrupted'`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BaseToolReturnPart.outcome) — a neutral outcome that (unlike `'failed'`) is not surfaced as a provider error — and carries `{'pydantic_ai_synthesized_tool_return': True}` in its [`metadata`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BaseToolReturnPart.metadata) so your code can tell it apart from real tool results. This also covers a call whose arguments were cut off mid-stream: the call is kept as-is and closed out the same way. Its arguments stay verbatim in the history, but the request serializers send them as `{"INVALID_JSON": ""}` (see [`args_as_json_str`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BaseToolCallPart.args_as_json_str) ) so that a provider requiring an object still accepts the request. * **Removes** an orphaned tool result — a [`ToolReturnPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolReturnPart) or [`RetryPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.RetryPromptPart) whose tool call is absent from the history (including a result placed before its call). If this empties an interior [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) the request is removed; if it empties the last message, an empty request is kept so the history still ends on a `ModelRequest`. After the invalid parts are handled, consecutive compatible messages are **merged** into one (two adjacent [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) s become a single turn, with tool results ordered ahead of user parts). This changes message boundaries but preserves all content, so processed history you inspect afterwards may have fewer messages than you passed in. The repair is deterministic and idempotent: repairing the same history always produces the same output, running a repaired history through another run leaves it untouched, and synthesized parts contain no wall-clock data, so reuse doesn’t invalidate provider prompt caches. Tool calls that can still receive a real result are left alone: when the history ends on a `ModelResponse` with tool calls, running without a new `user_prompt` executes them, and [deferred tool calls](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/) are matched to their `deferred_tool_results` — including when a ‘complete’ `ModelRequest` with the already-executed results follows the response. Repair of that live frontier only happens when the interruption is evident: a final response with [`state='interrupted'`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse.state) or a trailing request with [`state='interrupted'`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest.state) (e.g. from a [cancelled stream](https://pydantic.dev/docs/ai/core-concepts/output/#cancelling-streams) or a crash during tool execution) whose tool calls will never be executed. This pipeline handles regular, locally-executed tool calls only. Builtin (server-side) tool parts — produced and resulted by the provider inline — are left untouched and repaired by each model’s own serializer instead. Some other provider-invalid shapes are also out of scope and may be rejected: duplicate tool results for one call, and provider-specific ordering rules beyond call/result pairing. ### Correlating runs with `run_id` and `conversation_id` [](https://pydantic.dev/docs/ai/core-concepts/message-history/#correlating-runs-with-run_id-and-conversation_id) Each `ModelRequest` and `ModelResponse` carries two identifiers: * [`run_id`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest.run_id) — unique per agent run. Also available as [`RunContext.run_id`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.run_id) and [`AgentRunResult.run_id`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.run_id) , and emitted on the OpenTelemetry agent run span as `gen_ai.agent.call.id`. * [`conversation_id`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest.conversation_id) — shared across all runs that build on the same `message_history`. Also available as [`AgentRunResult.conversation_id`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.conversation_id) , and emitted as `gen_ai.conversation.id`. A fresh `run_id` is generated for every agent run (or you can pass `run_id=''` to use an ID minted by your application — e.g. one created, stored, or handed out to a client before the run starts). Unlike `conversation_id`, `run_id` is **never** inherited from `message_history`. Each [`Agent.run`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.run) call — including a [deferred-tool resume](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/) — is a separate run with its own `run_id`. Passing an empty `run_id=''`, or a `run_id` that already appears on `message_history`, raises [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) , because both break [`new_messages()`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.new_messages) boundary detection. Correlate pause/resume or multi-turn work with `conversation_id` instead. When retrying a failed run with the same `run_id`, rebuild `message_history` without the failed attempt’s messages. A fresh `conversation_id` is generated on the first run, stamped onto every message produced by that run, and inherited by subsequent runs that pass the messages back via `message_history`. This means you can correlate traces from a multi-turn conversation in [Logfire](https://pydantic.dev/docs/ai/integrations/logfire/) (or any OpenTelemetry backend) without tracking anything yourself — as long as the message history round-trips, the conversation ID does too. conversation\_id is shared across runs in the same conversation from pydantic_ai import Agent agent = Agent('openai:gpt-5.2') result1 = agent.run_sync('Tell me a joke.') result2 = agent.run_sync('Explain?', message_history=result1.all_messages()) assert result1.conversation_id == result2.conversation_id assert result1.run_id != result2.run_id pass a pre-minted run\_id from pydantic_ai import Agent from pydantic_ai.models.test import TestModel agent = Agent(TestModel()) result = agent.run_sync('Tell me a joke.', run_id='run-from-api-42') assert result.run_id == 'run-from-api-42' To override or fork `conversation_id`: * Pass `conversation_id=''` to use an ID from your own application (e.g. a chat thread ID stored in your database). * Pass `conversation_id='new'` to start a fresh conversation that ignores any `conversation_id` already on `message_history` — useful for branching off an existing thread without making the caller generate an ID. forking a conversation from pydantic_ai import Agent agent = Agent('openai:gpt-5.2') result1 = agent.run_sync('Tell me a joke.') forked = agent.run_sync( 'Tell me a different joke.', message_history=result1.all_messages(), conversation_id='new', ) assert forked.conversation_id != result1.conversation_id The [UI adapters](https://pydantic.dev/docs/ai/integrations/ui/overview/) auto-populate `conversation_id` from the protocol’s own thread/chat ID, so frontends using these protocols get conversation correlation for free. Protocol-level run IDs (for example AG-UI’s `runId`) are **not** mapped into the agent’s `run_id` — pass `run_id=` explicitly on `AGUIAdapter.run_stream` / `dispatch_request` (or a plain `Agent.run`) if you need them to match. Storing and loading messages (to JSON) -------------------------------------- [](https://pydantic.dev/docs/ai/core-concepts/message-history/#storing-and-loading-messages-to-json) While maintaining conversation state in memory is enough for many applications, often times you may want to store the messages history of an agent run on disk or in a database. This might be for evals, for sharing data between Python and JavaScript/TypeScript, or any number of other use cases. The intended way to do this is using a `TypeAdapter`. We export [`ModelMessagesTypeAdapter`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessagesTypeAdapter) that can be used for this, or you can create your own. Here’s an example showing how: serialize messages to json from pydantic_core import to_jsonable_python from pydantic_ai import ( Agent, ModelMessagesTypeAdapter, # (1) ) agent = Agent('openai:gpt-5.2', instructions='Be a helpful assistant.') result1 = agent.run_sync('Tell me a joke.') history_step_1 = result1.all_messages() as_python_objects = to_jsonable_python(history_step_1) # (2) same_history_as_step_1 = ModelMessagesTypeAdapter.validate_python(as_python_objects) result2 = agent.run_sync( # (3) 'Tell me a different joke.', message_history=same_history_as_step_1 ) Alternatively, you can create a `TypeAdapter` from scratch: from pydantic import TypeAdapter from pydantic_ai import ModelMessage ModelMessagesTypeAdapter = TypeAdapter(list[ModelMessage]) Alternatively you can serialize to/from JSON directly: from pydantic_core import to_json ... as_json_objects = to_json(history_step_1) same_history_as_step_1 = ModelMessagesTypeAdapter.validate_json(as_json_objects) You can now continue the conversation with history `same_history_as_step_1` despite creating a new agent run. _(This example is complete, it can be run “as is”)_ ### Loading untrusted history [](https://pydantic.dev/docs/ai/core-concepts/message-history/#loading-untrusted-history) The `message_history` parameter is trusted server-side state. If you load history that came from a browser request or another untrusted boundary, sanitize it before passing it to the agent. [`sanitize_messages`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.sanitize_messages) applies the same default message sanitization used by the [UI adapters](https://pydantic.dev/docs/ai/integrations/ui/overview/) : it strips client-supplied system prompts, drops non-HTTP file URL schemes, resets non-allowlisted [`FileUrl.force_download`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.FileUrl.force_download) values to `False`, drops uploaded file references, and removes unresolved tool calls at the end of the history. Client-supplied [`CompactionPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.CompactionPart) s are kept, so the conversation stays [compacted](https://pydantic.dev/docs/ai/capabilities/compaction/) — but they are never trusted to stand in for the system prompt. Whether that prompt is a [`SystemPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SystemPromptPart) already in the history or one re-injected by [`ReinjectSystemPrompt`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ReinjectSystemPrompt) , it is re-sent to the model even where a provider’s own compaction state would normally let it be skipped. If you combine the sanitized history with trusted server-side `message_history`, also pass `strip_compaction_parts=True`: everything before a compaction item is hidden from the model, so a client-supplied one would hide the server’s history — see [Client-held history](https://pydantic.dev/docs/ai/capabilities/compaction/#client-held-history) . The [UI adapters](https://pydantic.dev/docs/ai/integrations/ui/overview/) apply this rule automatically when a run combines server-side `message_history` with client-submitted messages. sanitize untrusted message history from pydantic_ai import Agent, ModelMessagesTypeAdapter from pydantic_ai.messages import sanitize_messages agent = Agent('openai:gpt-5.2', instructions='Be a helpful assistant.') # `request_json` is the body submitted by an untrusted client. loaded_history = ModelMessagesTypeAdapter.validate_python(request_json['message_history']) message_history = sanitize_messages(loaded_history) result = agent.run_sync('Tell me a different joke.', message_history=message_history) Each sanitization can be turned off individually when the corresponding parts were created by trusted server-side code: pass `strip_system_prompts=False`, add schemes to `allowed_file_url_schemes`, add values to `allowed_file_url_force_download`, or set `allow_uploaded_files=True`. See [file URL input security](https://pydantic.dev/docs/ai/core-concepts/input/#user-side-download-vs-direct-file-url) for the file input trust model. Trust boundary for client-supplied history ------------------------------------------ [](https://pydantic.dev/docs/ai/core-concepts/message-history/#trust-boundary-for-client-supplied-history) Pydantic AI’s server-side surfaces are stateless: a run is reconstructed from the `message_history` (and any `deferred_tool_results`) supplied with the request, whether that request arrives through a [UI adapter](https://pydantic.dev/docs/ai/integrations/ui/overview/) or through an endpoint you wrote yourself. A client that can submit history can therefore fabricate it — including [`ToolCallPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolCallPart) s the model never emitted and [approvals](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/#human-in-the-loop-tool-approval) no human granted — and the server will process them as genuine, up to and including executing the tools they name. Pydantic AI does not sign or cryptographically verify tool calls, tool results, or approvals, and neither do comparable agent frameworks: signing is only meaningful for a server that kept the run itself, and such a server doesn’t need the client’s copy of the history in the first place. The defaults described under [Loading untrusted history](https://pydantic.dev/docs/ai/core-concepts/message-history/#loading-untrusted-history) and in the [UI adapter trust model](https://pydantic.dev/docs/ai/integrations/ui/overview/#trust-model-for-client-submitted-messages) narrow what a fabricated history can reach; they don’t make it trustworthy. Possession of the endpoint is therefore the authorization boundary, so design around that: * **Authenticate and authorize at the transport layer.** Run the agent inside your own authenticated route handler, and treat every caller that gets through as able to submit any history it likes. * **Scope the toolset to the caller.** Expose only the tools the authenticated caller is entitled to use, by [building the toolset per run](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/#dynamically-building-a-toolset) or [filtering](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/#filtering-tools) it against the user carried in your [dependencies](https://pydantic.dev/docs/ai/core-concepts/dependencies/) . * **Re-validate high-stakes effects server-side.** [Approval](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/#human-in-the-loop-tool-approval) guards against the _model_ acting without human sign-off, not against the client. Where the stakes demand it, check the caller’s authority against server-side state inside the tool function itself, or persist paused runs server-side and resume them with your own `deferred_tool_results` instead of the client’s. Other ways of using messages ---------------------------- [](https://pydantic.dev/docs/ai/core-concepts/message-history/#other-ways-of-using-messages) Since messages are defined by simple dataclasses, you can manually create and manipulate, e.g. for testing. The message format is independent of the model used, so you can use messages in different agents, or the same agent with different models. In the example below, we reuse the message from the first agent run, which uses the `openai:gpt-5.2` model, in a second agent run using the `google:gemini-3-pro-preview` model. Reusing messages with a different model from pydantic_ai import Agent agent = Agent('openai:gpt-5.2', instructions='Be a helpful assistant.') result1 = agent.run_sync('Tell me a joke.') print(result1.output) #> Did you hear about the toothpaste scandal? They called it Colgate. result2 = agent.run_sync( 'Explain?', model='google:gemini-3-pro-preview', message_history=result1.new_messages(), ) print(result2.output) #> This is an excellent joke invented by Samuel Colvin, it needs no explanation. print(result2.all_messages()) """ [\ ModelRequest(\ parts=[\ UserPromptPart(\ content='Tell me a joke.',\ timestamp=datetime.datetime(...),\ )\ ],\ timestamp=datetime.datetime(...),\ instructions='Be a helpful assistant.',\ run_id='...',\ conversation_id='...',\ ),\ ModelResponse(\ parts=[\ TextPart(\ content='Did you hear about the toothpaste scandal? They called it Colgate.'\ )\ ],\ usage=RequestUsage(\ cost=Decimal('0.00026425'), input_tokens=55, output_tokens=12\ ),\ model_name='gpt-5.2',\ timestamp=datetime.datetime(...),\ run_id='...',\ conversation_id='...',\ ),\ ModelRequest(\ parts=[\ UserPromptPart(\ content='Explain?',\ timestamp=datetime.datetime(...),\ )\ ],\ timestamp=datetime.datetime(...),\ instructions='Be a helpful assistant.',\ run_id='...',\ conversation_id='...',\ ),\ ModelResponse(\ parts=[\ TextPart(\ content='This is an excellent joke invented by Samuel Colvin, it needs no explanation.'\ )\ ],\ usage=RequestUsage(cost=Decimal('0.000424'), input_tokens=56, output_tokens=26),\ model_name='gemini-3-pro-preview',\ timestamp=datetime.datetime(...),\ run_id='...',\ conversation_id='...',\ ),\ ] """ _(This example is complete, it can be run “as is”)_ Sharing messages between agents ------------------------------- [](https://pydantic.dev/docs/ai/core-concepts/message-history/#sharing-messages-between-agents) The same `message_history` parameter also works when the next run uses a different [`Agent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent) . This is useful for [programmatic agent hand-off](https://pydantic.dev/docs/ai/guides/multi-agent-applications/#programmatic-agent-hand-off) , where your application runs one agent, then gives another agent the conversation so far as context. sharing\_messages\_between\_agents.py from pydantic_ai import Agent biography_agent = Agent( 'openai:gpt-5.2', instructions='Answer biographical questions concisely.', ) science_agent = Agent( 'anthropic:claude-sonnet-4-6', instructions='Answer science questions for a general audience.', ) biography_result = biography_agent.run_sync('Who was Albert Einstein?') print(biography_result.output) #> Albert Einstein was a German-born theoretical physicist. science_result = science_agent.run_sync( 'What was his most famous equation?', message_history=biography_result.new_messages(), ) print(science_result.output) #> Albert Einstein's most famous equation is (E = mc^2). _(This example is complete, it can be run “as is”)_ For more complex multi-agent patterns, see the [multi-agent applications](https://pydantic.dev/docs/ai/guides/multi-agent-applications/) documentation. Editing existing messages ------------------------- [](https://pydantic.dev/docs/ai/core-concepts/message-history/#editing-existing-messages) To change the conversation mid-run, build _new_ message objects rather than modifying existing ones: [inject new messages](https://pydantic.dev/docs/ai/core-concepts/message-history/#injecting-messages-mid-run) with `enqueue`, or prune, summarize, or otherwise rewrite the history the model receives with a [history processor](https://pydantic.dev/docs/ai/core-concepts/message-history/#processing-message-history) . When you need to edit an earlier message — say, compacting a large tool output — copy it with [`dataclasses.replace`](https://docs.python.org/3/library/dataclasses.html#dataclasses.replace) , passing a new `parts` list of new (or reused) part objects; edited parts are likewise built with `replace` rather than modified. Replacing a message in the history and reassigning its `parts` list are both safe. Injecting messages mid-run -------------------------- [](https://pydantic.dev/docs/ai/core-concepts/message-history/#injecting-messages-mid-run) Tools, capability hooks, and external code driving an agent run can inject extra content into the conversation mid-run with [`RunContext.enqueue`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.enqueue) (when a `RunContext` is in scope, e.g. inside a tool or capability hook) or [`AgentRun.enqueue`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.enqueue) (from external code driving [`agent.iter()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.iter) ). Use this when something happens during a run that the agent should know about — a tool wants to add follow-up context, an external event needs to _steer_ the agent’s plan, or background work needs to reach the agent when it completes. A `priority` controls when the enqueued content is delivered: * `'asap'` (default): delivered at the earliest opportunity — added to the next [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) , or, if the agent would otherwise terminate before another request, used to redirect the run into one more request. Use when the new context should reach the model as soon as possible; this is what other frameworks often call **steering** an in-flight agent. * `'when_idle'`: delivered only when the agent would otherwise terminate, after any `'asap'` messages. Use when the agent shouldn’t be interrupted but should pick up the new work — a follow-up task — once it’s done with what it’s doing. `enqueue` is variadic — each positional argument is one item, and can be: * a piece of [`UserContent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.UserContent) — a `str` or multi-modal content like an [`ImageUrl`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ImageUrl) . Adjacent user content is gathered into a single [`UserPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.UserPromptPart) , so `enqueue('caption', image)` forms one user turn. To pass an existing list, spread it: `enqueue(*items)`; * a [`ModelRequestPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequestPart) , such as a [`SystemPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SystemPromptPart) ; * a complete [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) or [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) , to control request-level fields like `instructions`/`metadata` or to inject a synthetic prior turn. Adjacent part-style items (user content and [`ModelRequestPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequestPart) s) are coalesced into one [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) ; complete messages stay separate. This lets a single call inject an interleaved exchange — for example a synthetic tool call (a [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) ) followed by its result (a [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) ). The content must end in a request, so the agent has something to respond to. Both `enqueue` methods return an `enqueue_id` (`str`) for a non-empty call, or `None` when called with no content. When the queued content is actually delivered into run history, the [event stream](https://pydantic.dev/docs/ai/core-concepts/agent/#streaming-all-events) yields an [`EnqueuedMessagesEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.EnqueuedMessagesEvent) carrying that `enqueue_id` and the delivered messages (exactly as they landed in history), so a client can observe when its steering message took effect. The event carries the delivered message objects themselves — the same objects held in the run’s message history. A history processor that replaces history with new message objects does not affect the event, but [in-place mutation](https://pydantic.dev/docs/ai/core-concepts/message-history/#editing-existing-messages) of a delivered message will be visible through it. ### From inside a tool or hook [](https://pydantic.dev/docs/ai/core-concepts/message-history/#from-inside-a-tool-or-hook) Use [`RunContext.enqueue`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.enqueue) when you have a `RunContext` in scope: enqueue\_from\_tool.py from pydantic_ai import Agent, RunContext from pydantic_ai.messages import SystemPromptPart agent = Agent('anthropic:claude-opus-4-7') @agent.tool def trigger_alert(ctx: RunContext[None]) -> str: ctx.enqueue('Alert: production is degraded, prioritize triage.') return 'alert raised' @agent.tool def enter_incident_mode(ctx: RunContext[None]) -> str: # Enqueue a `SystemPromptPart` to adjust the agent's standing instructions mid-run. ctx.enqueue(SystemPromptPart(content='You are now in incident mode: be terse and action-oriented.')) return 'incident mode enabled' The `'asap'` message is appended to the agent’s message history and is visible to the model on the next request, alongside any tool returns from the same step. A [`SystemPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SystemPromptPart) is delivered the same way, and lands as a [mid-conversation system prompt](https://pydantic.dev/docs/ai/core-concepts/message-history/#mid-conversation-system-prompts) — it keeps its position in the history instead of being lifted into the top-level system prompt, so it doesn’t invalidate the cached prefix ahead of it. Only enqueue a `SystemPromptPart` for an instruction you authored; see the warning in that section for why late-arriving tool or webhook output belongs in user content instead. ### From external code driving `agent.iter()` [](https://pydantic.dev/docs/ai/core-concepts/message-history/#from-external-code-driving-agentiter) Use [`AgentRun.enqueue`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.enqueue) when you’re driving a run from outside (e.g. forwarding events from a webhook, chat platform, or job queue): enqueue\_from\_agent\_run.py from pydantic_ai import Agent from pydantic_graph import End agent = Agent('anthropic:claude-opus-4-7') async def main(): async with agent.iter('Summarize the latest deploy report') as agent_run: # An external system pushes a follow-up while the agent is working. # When the agent would otherwise finish, the message redirects it # into a fresh model request so it can incorporate the new context. agent_run.enqueue( 'A new error was just reported — include it in the summary.', priority='when_idle', ) node = agent_run.next_node while not isinstance(node, End): node = await agent_run.next(node) `'when_idle'` messages are only drained when the agent would otherwise reach an `End` — that drain happens in `after_node_run`. `'asap'` messages are drained in `before_model_request`, and also at the same end-of-run point if anything arrived during the final step. Both fire however you drive the run, so [`Agent.run`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.run) , [`AgentRun.next()`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.next) , and a bare `async for node in agent_run:` loop all deliver enqueued messages. Processing Message History -------------------------- [](https://pydantic.dev/docs/ai/core-concepts/message-history/#processing-message-history) Sometimes you may want to modify the message history before it’s sent to the model. This could be for privacy reasons (filtering out sensitive information), to save costs on tokens, to give less context to the LLM, or custom processing logic. Pydantic AI provides the [`ProcessHistory`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ProcessHistory) capability that allows you to intercept and modify the message history before each model request. ### Usage [](https://pydantic.dev/docs/ai/core-concepts/message-history/#usage) Each [`ProcessHistory`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ProcessHistory) wraps a callable that takes a list of [`ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage) and returns a modified list of the same type. Each processor is applied in sequence, and processors can be either synchronous or asynchronous. simple\_history\_processor.py from pydantic_ai import ( Agent, ModelMessage, ModelRequest, ModelResponse, TextPart, UserPromptPart, ) from pydantic_ai.capabilities import ProcessHistory def filter_responses(messages: list[ModelMessage]) -> list[ModelMessage]: """Remove all ModelResponse messages, keeping only ModelRequest messages.""" return [msg for msg in messages if isinstance(msg, ModelRequest)] # Create agent with history processor agent = Agent('openai:gpt-5.2', capabilities=[ProcessHistory(filter_responses)]) # Example: Create some conversation history message_history = [\ ModelRequest(parts=[UserPromptPart(content='What is 2+2?')]),\ ModelResponse(parts=[TextPart(content='2+2 equals 4')]), # This will be filtered out\ ] # When you run the agent, the history processor will filter out ModelResponse messages # result = agent.run_sync('What about 3+3?', message_history=message_history) #### Keep Only Recent Messages [](https://pydantic.dev/docs/ai/core-concepts/message-history/#keep-only-recent-messages) You can use the `history_processor` to only keep the recent messages: keep\_recent\_messages.py from pydantic_ai import Agent, ModelMessage from pydantic_ai.capabilities import ProcessHistory async def keep_recent_messages(messages: list[ModelMessage]) -> list[ModelMessage]: """Keep only the last 5 messages to manage token usage.""" return messages[-5:] if len(messages) > 5 else messages agent = Agent('openai:gpt-5.2', capabilities=[ProcessHistory(keep_recent_messages)]) # Example: Even with a long conversation history, only the last 5 messages are sent to the model long_conversation_history: list[ModelMessage] = [] # Your long conversation history here # result = agent.run_sync('What did we discuss?', message_history=long_conversation_history) #### `RunContext` parameter [](https://pydantic.dev/docs/ai/core-concepts/message-history/#runcontext-parameter) History processors can optionally accept a [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) parameter to access additional information about the current run, such as dependencies, model information, and usage statistics: context\_aware\_processor.py from pydantic_ai import Agent, ModelMessage, RunContext from pydantic_ai.capabilities import ProcessHistory def context_aware_processor( ctx: RunContext, messages: list[ModelMessage], ) -> list[ModelMessage]: # Access current usage current_tokens = ctx.usage.total_tokens # Filter messages based on context if current_tokens > 1000: return messages[-3:] # Keep only recent messages when token usage is high return messages agent = Agent('openai:gpt-5.2', capabilities=[ProcessHistory(context_aware_processor)]) This allows for more sophisticated message processing based on the current state of the agent run. Whether the processor wants a [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) is detected by resolving its type hints at runtime, so every annotated type in the processor signature must be imported at runtime rather than only under `if TYPE_CHECKING:`. If any annotation can’t be resolved, a [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) is raised instead of the processor being silently called without the context. #### Summarize Old Messages [](https://pydantic.dev/docs/ai/core-concepts/message-history/#summarize-old-messages) Use an LLM to summarize older messages to preserve context while reducing tokens. This is one of several ways to keep a conversation within the context window — see [Compaction](https://pydantic.dev/docs/ai/capabilities/compaction/) for the full picture, including provider-native compaction and ready-made strategies from [Pydantic AI Harness](https://pydantic.dev/docs/ai/harness/compaction/) . summarize\_old\_messages.py from pydantic_ai import Agent, ModelMessage from pydantic_ai.capabilities import ProcessHistory # Use a cheaper model to summarize old messages. summarize_agent = Agent( 'openai:gpt-5-mini', instructions=""" Summarize this conversation, omitting small talk and unrelated topics. Focus on the technical discussion and next steps. """, ) async def summarize_old_messages(messages: list[ModelMessage]) -> list[ModelMessage]: # Summarize the oldest 10 messages if len(messages) > 10: oldest_messages = messages[:10] summary = await summarize_agent.run(message_history=oldest_messages) # Return the last message and the summary return summary.new_messages() + messages[-1:] return messages agent = Agent('openai:gpt-5.2', capabilities=[ProcessHistory(summarize_old_messages)]) ### Testing History Processors [](https://pydantic.dev/docs/ai/core-concepts/message-history/#testing-history-processors) You can test what messages are actually sent to the model provider using [`FunctionModel`](https://pydantic.dev/docs/ai/api/models/function/#pydantic_ai.models.function.FunctionModel) : test\_history\_processor.py import pytest from pydantic_ai import ( Agent, ModelMessage, ModelRequest, ModelResponse, TextPart, UserPromptPart, ) from pydantic_ai.capabilities import ProcessHistory from pydantic_ai.models.function import AgentInfo, FunctionModel @pytest.fixture def received_messages() -> list[ModelMessage]: return [] @pytest.fixture def function_model(received_messages: list[ModelMessage]) -> FunctionModel: def capture_model_function(messages: list[ModelMessage], info: AgentInfo) -> ModelResponse: # Capture the messages that the provider actually receives received_messages.clear() received_messages.extend(messages) return ModelResponse(parts=[TextPart(content='Provider response')]) return FunctionModel(capture_model_function) def test_history_processor(function_model: FunctionModel, received_messages: list[ModelMessage]): def filter_responses(messages: list[ModelMessage]) -> list[ModelMessage]: return [msg for msg in messages if isinstance(msg, ModelRequest)] agent = Agent(function_model, capabilities=[ProcessHistory(filter_responses)]) message_history = [\ ModelRequest(parts=[UserPromptPart(content='Question 1')]),\ ModelResponse(parts=[TextPart(content='Answer 1')]),\ ] agent.run_sync('Question 2', message_history=message_history) assert received_messages == [\ ModelRequest(parts=[UserPromptPart(content='Question 1')]),\ ModelRequest(parts=[UserPromptPart(content='Question 2')]),\ ] ### Multiple Processors [](https://pydantic.dev/docs/ai/core-concepts/message-history/#multiple-processors) You can also use multiple processors: multiple\_history\_processors.py from pydantic_ai import Agent, ModelMessage, ModelRequest from pydantic_ai.capabilities import ProcessHistory def filter_responses(messages: list[ModelMessage]) -> list[ModelMessage]: return [msg for msg in messages if isinstance(msg, ModelRequest)] def summarize_old_messages(messages: list[ModelMessage]) -> list[ModelMessage]: return messages[-5:] agent = Agent( 'openai:gpt-5.2', capabilities=[ProcessHistory(filter_responses), ProcessHistory(summarize_old_messages)], ) In this case, the `filter_responses` processor will be applied first, and the `summarize_old_messages` processor will be applied second. Examples -------- [](https://pydantic.dev/docs/ai/core-concepts/message-history/#examples) For a more complete example of using messages in conversations, see the [chat app](https://pydantic.dev/docs/ai/examples/conversational-agents/chat-app/) example. Was this page helpful? Thanks for your feedback! --- # direct | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#_top) direct ====== Methods for making imperative requests to language models with minimal abstraction. These methods allow you to make requests to LLMs where the only abstraction is input and output schema translation so you can use all models with the same API. These methods are thin wrappers around [`Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) implementations. StreamedResponseSync -------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.StreamedResponseSync) Synchronous wrapper for an async streaming response, running the whole stream on the caller’s event loop. The stream uses the internal `SyncStreamBridge` to keep context-manager and iterator lifecycles in stable tasks. Exiting the `with` block cancels the underlying request promptly and closes the connection instead of waiting for the whole response to arrive. This class must be used as a context manager with the `with` statement. The synchronous stream is created when the `with` block is entered and must be used and closed on that thread. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#attributes) #### model\_name [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.StreamedResponseSync.model_name) Get the model name of the response. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### response [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.StreamedResponseSync.response) Get the current state of the response. **Type:** [`messages.ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) #### timestamp [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.StreamedResponseSync.timestamp) Get the timestamp of the response. **Type:** [`datetime`](https://docs.python.org/3/library/datetime.html#module-datetime) #### usage [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.StreamedResponseSync.usage) Get the usage of the response so far. **Type:** [`RequestUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage) ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#methods) #### \_\_iter\_\_ [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.StreamedResponseSync.__iter__) def __iter__() -> Iterator[messages.ModelResponseStreamEvent] Stream the response as an iterable of [`ModelResponseStreamEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponseStreamEvent) s. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#returns) [`Iterator`](https://docs.python.org/3/library/typing.html#typing.Iterator) \[[`messages.ModelResponseStreamEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponseStreamEvent)\ \] model\_request -------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request) `@async` def model_request( model: models.Model | models.KnownModelName | str, messages: Sequence[messages.ModelMessage], *, model_settings: settings.ModelSettings | None = None, model_request_parameters: models.ModelRequestParameters | None = None, instrument: instrumented_models.InstrumentationSettings | bool | None = None, ) -> messages.ModelResponse Make a non-streamed request to a model. model\_request\_example.py from pydantic_ai import ModelRequest from pydantic_ai.direct import model_request async def main(): model_response = await model_request( 'anthropic:claude-haiku-4-5', [ModelRequest.user_text_prompt('What is the capital of France?')] # (1) ) print(model_response) ''' ModelResponse( parts=[TextPart(content='The capital of France is Paris.')], usage=RequestUsage(input_tokens=56, output_tokens=7), model_name='claude-haiku-4-5', timestamp=datetime.datetime(...), ) ''' See [`ModelRequest.user_text_prompt`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest.user_text_prompt) for details. ### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#returns-1) [`messages.ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) — The model response and token usage associated with the request. ### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#parameters) **`model`** : [`models.Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) | [`models.KnownModelName`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.KnownModelName) | [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request(model)) The model to make a request to. We allow `str` here since the actual list of allowed models changes frequently. **`messages`** : [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`messages.ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage)\ \] [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request(messages)) Messages to send to the model **`model_settings`** : [`settings.ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request(model_settings)) optional model settings **`model_request_parameters`** : [`models.ModelRequestParameters`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request(model_request_parameters)) optional model request parameters **`instrument`** : `instrumented_models.InstrumentationSettings` | [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request(instrument)) Whether to instrument the request with OpenTelemetry/Logfire, if `None` the value from [`logfire.instrument_pydantic_ai`](https://logfire.pydantic.dev/docs/api/logfire/#logfire.Logfire.instrument_pydantic_ai) is used. model\_request\_stream ---------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_stream) def model_request_stream( model: models.Model | models.KnownModelName | str, messages: Sequence[messages.ModelMessage], *, model_settings: settings.ModelSettings | None = None, model_request_parameters: models.ModelRequestParameters | None = None, instrument: instrumented_models.InstrumentationSettings | bool | None = None, ) -> AbstractAsyncContextManager[models.StreamedResponse] Make a streamed async request to a model. model\_request\_stream\_example.py from pydantic_ai import ModelRequest from pydantic_ai.direct import model_request_stream async def main(): messages = [ModelRequest.user_text_prompt('Who was Albert Einstein?')] # (1) async with model_request_stream('openai:gpt-5-mini', messages) as stream: chunks = [] async for chunk in stream: chunks.append(chunk) print(chunks) ''' [\ PartStartEvent(index=0, part=TextPart(content='Albert Einstein was ')),\ FinalResultEvent(tool_name=None, tool_call_id=None),\ PartDeltaEvent(\ index=0, delta=TextPartDelta(content_delta='a German-born theoretical ')\ ),\ PartDeltaEvent(index=0, delta=TextPartDelta(content_delta='physicist.')),\ PartEndEvent(\ index=0,\ part=TextPart(\ content='Albert Einstein was a German-born theoretical physicist.'\ ),\ ),\ ] ''' See [`ModelRequest.user_text_prompt`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest.user_text_prompt) for details. ### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#returns-2) `AbstractAsyncContextManager`\[[`models.StreamedResponse`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.StreamedResponse)\ \] — A [stream response](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.StreamedResponse) async context manager. ### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#parameters-1) **`model`** : [`models.Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) | [`models.KnownModelName`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.KnownModelName) | [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_stream(model)) The model to make a request to. We allow `str` here since the actual list of allowed models changes frequently. **`messages`** : [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`messages.ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage)\ \] [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_stream(messages)) Messages to send to the model **`model_settings`** : [`settings.ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_stream(model_settings)) optional model settings **`model_request_parameters`** : [`models.ModelRequestParameters`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_stream(model_request_parameters)) optional model request parameters **`instrument`** : `instrumented_models.InstrumentationSettings` | [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_stream(instrument)) Whether to instrument the request with OpenTelemetry/Logfire, if `None` the value from [`logfire.instrument_pydantic_ai`](https://logfire.pydantic.dev/docs/api/logfire/#logfire.Logfire.instrument_pydantic_ai) is used. model\_request\_stream\_sync ---------------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_stream_sync) def model_request_stream_sync( model: models.Model | models.KnownModelName | str, messages: Sequence[messages.ModelMessage], *, model_settings: settings.ModelSettings | None = None, model_request_parameters: models.ModelRequestParameters | None = None, instrument: instrumented_models.InstrumentationSettings | bool | None = None, ) -> StreamedResponseSync Make a streamed synchronous request to a model. This is the synchronous version of [`model_request_stream`](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_stream) . It drives the asynchronous stream on the caller’s event loop while providing a synchronous iterator interface. The returned context manager must be used and closed on the thread where the synchronous stream is created. model\_request\_stream\_sync\_example.py from pydantic_ai import ModelRequest from pydantic_ai.direct import model_request_stream_sync messages = [ModelRequest.user_text_prompt('Who was Albert Einstein?')] with model_request_stream_sync('openai:gpt-5-mini', messages) as stream: chunks = [] for chunk in stream: chunks.append(chunk) print(chunks) ''' [\ PartStartEvent(index=0, part=TextPart(content='Albert Einstein was ')),\ FinalResultEvent(tool_name=None, tool_call_id=None),\ PartDeltaEvent(\ index=0, delta=TextPartDelta(content_delta='a German-born theoretical ')\ ),\ PartDeltaEvent(index=0, delta=TextPartDelta(content_delta='physicist.')),\ PartEndEvent(\ index=0,\ part=TextPart(\ content='Albert Einstein was a German-born theoretical physicist.'\ ),\ ),\ ] ''' ### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#returns-3) `StreamedResponseSync` — A [sync stream response](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.StreamedResponseSync) context manager. ### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#parameters-2) **`model`** : [`models.Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) | [`models.KnownModelName`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.KnownModelName) | [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_stream_sync(model)) The model to make a request to. We allow `str` here since the actual list of allowed models changes frequently. **`messages`** : [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`messages.ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage)\ \] [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_stream_sync(messages)) Messages to send to the model **`model_settings`** : [`settings.ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_stream_sync(model_settings)) optional model settings **`model_request_parameters`** : [`models.ModelRequestParameters`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_stream_sync(model_request_parameters)) optional model request parameters **`instrument`** : `instrumented_models.InstrumentationSettings` | [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_stream_sync(instrument)) Whether to instrument the request with OpenTelemetry/Logfire, if `None` the value from [`logfire.instrument_pydantic_ai`](https://logfire.pydantic.dev/docs/api/logfire/#logfire.Logfire.instrument_pydantic_ai) is used. model\_request\_sync -------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_sync) def model_request_sync( model: models.Model | models.KnownModelName | str, messages: Sequence[messages.ModelMessage], *, model_settings: settings.ModelSettings | None = None, model_request_parameters: models.ModelRequestParameters | None = None, instrument: instrumented_models.InstrumentationSettings | bool | None = None, ) -> messages.ModelResponse Make a Synchronous, non-streamed request to a model. This is a convenience method that wraps [`model_request`](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request) with `loop.run_until_complete(...)`. You therefore can’t use this method inside async code or if there’s an active event loop. model\_request\_sync\_example.py from pydantic_ai import ModelRequest from pydantic_ai.direct import model_request_sync model_response = model_request_sync( 'anthropic:claude-haiku-4-5', [ModelRequest.user_text_prompt('What is the capital of France?')] # (1) ) print(model_response) ''' ModelResponse( parts=[TextPart(content='The capital of France is Paris.')], usage=RequestUsage(input_tokens=56, output_tokens=7), model_name='claude-haiku-4-5', timestamp=datetime.datetime(...), ) ''' See [`ModelRequest.user_text_prompt`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest.user_text_prompt) for details. ### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#returns-4) [`messages.ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) — The model response and token usage associated with the request. ### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#parameters-3) **`model`** : [`models.Model`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.Model) | [`models.KnownModelName`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.KnownModelName) | [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_sync(model)) The model to make a request to. We allow `str` here since the actual list of allowed models changes frequently. **`messages`** : [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`messages.ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage)\ \] [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_sync(messages)) Messages to send to the model **`model_settings`** : [`settings.ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_sync(model_settings)) optional model settings **`model_request_parameters`** : [`models.ModelRequestParameters`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_sync(model_request_parameters)) optional model request parameters **`instrument`** : `instrumented_models.InstrumentationSettings` | [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/direct/#pydantic_ai.direct.model_request_sync(instrument)) Whether to instrument the request with OpenTelemetry/Logfire, if `None` the value from [`logfire.instrument_pydantic_ai`](https://logfire.pydantic.dev/docs/api/logfire/#logfire.Logfire.instrument_pydantic_ai) is used. Was this page helpful? Thanks for your feedback! --- # Input, Output & Tool Guardrails [Skip to content](https://pydantic.dev/docs/ai/harness/guardrails/#_top) Guardrails ========== Guardrails put a validation layer on the three edges of an agent run: the prompt on its way _in_ to the model, the tool calls the model makes along the way, and the output on its way _out_ to the caller. Reach for them when unstructured input or output must be screened before it is acted on — a prompt-injection attempt you never want to send, PII you must redact, an off-topic request you want to refuse cheaply, or an answer that must cite its sources before you show it. Without a guardrail the framework sends whatever the user typed and returns whatever the model produced, verbatim; a guardrail interposes a callable you control that gets the final say. > The API may change between releases. Where practical, breaking changes ship with a deprecation warning. The problem ----------- [](https://pydantic.dev/docs/ai/harness/guardrails/#the-problem) Agents take unstructured input from users and return unstructured output to callers. On its own the framework does not reason about “this is unsafe to send” or “this is unsafe to show” — a prompt-injection attempt reaches the model as-is, and any output the model produces is returned untouched. You need a place to inspect the value and decide what happens next. The solution ------------ [](https://pydantic.dev/docs/ai/harness/guardrails/#the-solution) Three capabilities — `InputGuardrail`, `OutputGuardrail`, and `ToolGuardrail` — each wrap a `guard` callable you supply. The guard inspects a value and returns one of five outcomes. For the run’s two edges: | Outcome | `InputGuardrail` | `OutputGuardrail` | | --- | --- | --- | | **allow** | send the prompt to the model | return the output to the caller | | **block** | skip the model call; a refusal message becomes the response | raise `OutputBlocked` | | **replace** | rewrite the prompt sent to the model (redaction) | substitute a sanitized output | | **retry** | — (not valid for input) | send the output back to the model to try again | | **approve** | — (not valid for input) | — (not valid for output) | `ToolGuardrail` uses the same outcomes on both sides of a tool call; see [Tool calls](https://pydantic.dev/docs/ai/harness/guardrails/#tool-calls) . The asymmetry between input `block` and output `block` is deliberate. Blocking the input spends no tokens, so a graceful refusal is almost always right. Blocking the output means the model already produced something you do not want exposed, so raising forces the caller to decide what to do next. Both `InputGuardrail`, `OutputGuardrail`, and their supporting types are top-level exports: from pydantic_ai import Agent from pydantic_ai_harness import GuardrailResult, InputGuardrail, OutputGuardrail def no_secrets(prompt: str) -> bool: return 'api_key' not in prompt.lower() def no_pii(output: object) -> GuardrailResult: if 'SSN' in str(output): return GuardrailResult.block('The response contained personal data.') return GuardrailResult.allow() agent = Agent( 'openai:gpt-5.4', capabilities=[\ InputGuardrail(guard=no_secrets),\ OutputGuardrail(guard=no_pii),\ ], ) A guard returns a bare `bool` (`True` = allow, `False` = block) for the simple case, or a `GuardrailResult` for the richer outcomes. Guards may also be async — return an awaitable `bool`/`GuardrailResult`, for example to call a moderation API. `OutputGuardrail` receives the output unchanged — no automatic stringification. For a string output the guard reads it directly; for a typed (Pydantic model) output the guard gets the model instance, so pick the serialization that fits the check (read a field, or call `output.model_dump_json()` for JSON text). This avoids the trap of `str(MyModel(...))` producing a `MyModel(field=...)` repr that hides field contents from regex-based checks. Several guards at once ---------------------- [](https://pydantic.dev/docs/ai/harness/guardrails/#several-guards-at-once) `InputGuardrail.guard` and `OutputGuardrail.guard` take one callable or a sequence of them; `ToolGuardrail`’s `guard` and `result_guard` take a single callable each. In a chain the guards run in order, and what happens next depends on the verdict: | Verdict | Effect on the chain | | --- | --- | | `allow` | move on to the next guard | | `replace` | the rest of the chain inspects the substituted value | | `block` / `retry` | the chain ends there, since neither leaves a value to judge | `replace` threading forward is what makes order meaningful: put a redactor first and everything after it sees the cleaned text. from pydantic_ai import Agent from pydantic_ai_harness import InputGuardrail from pydantic_ai_harness.guardrails.detectors import blocked_keywords, redact_secrets agent = Agent( 'openai:gpt-5.4', capabilities=[InputGuardrail(guard=[redact_secrets, blocked_keywords(['internal-only'])])], ) One `InputGuardrail` holding two checks is not the same as two `InputGuardrail` capabilities: the chain is one place in the capability list, one ordering decision, and one set of spans, and only the chain threads a redaction into the check that follows it. A redaction reaches the run’s message history only if the chain finishes: a later `block` ends it before `InputGuardrail` writes the cleaned prompt back, so the original stays in the history. Put the redactor last when that matters, at the cost of the checks before it seeing the original text. An empty sequence is refused — when the guardrail first runs, not when it is constructed. A guardrail that inspects nothing reads as configured and behaves as absent, which is worth an error rather than a quiet pass. A set and a one-shot iterator are refused the same way: a set has no order for the chain to run in, and an iterator is spent after the first request, since the chain is rebuilt per request. Ready-made detectors -------------------- [](https://pydantic.dev/docs/ai/harness/guardrails/#ready-made-detectors) `pydantic_ai_harness.guardrails.detectors` holds checks you would otherwise write again: they are plain functions returning a `GuardrailResult`, so they drop into a chain beside your own. from pydantic_ai_harness.guardrails import detectors detectors.redact_secrets # rewrites vendor API keys, tokens, and whole private-key blocks out of text detectors.redact_personal_data # rewrites emails, card numbers, IBANs, US SSNs detectors.blocked_keywords(['internal-only']) # refuses text containing any of them The two redactors rewrite rather than refuse, which is the useful default: an agent that quoted a key back has still done the work, and blocking the answer loses it while leaving the key in the message history either way. Each is `secret_data()` and `personal_data()` called with the defaults. Use those factories directly when you need to narrow or extend them: from pydantic_ai_harness.guardrails.detectors import secret_data secret_data(only=['aws_access_key', 'private_key']) secret_data(extra={'internal_ticket': r'INT-\d{4}'}) secret_data(placeholder='***') # the default is `[redacted:{name}]`, which says what it removed Detectors read text, so they suit a prompt directly. An agent output may be a model instance, and substituting a scrubbed string for one would change its type, so `for_text` makes you say what should happen: from pydantic_ai_harness import OutputGuardrail from pydantic_ai_harness.guardrails.detectors import for_text, redact_secrets OutputGuardrail(guard=for_text(redact_secrets)) # raises on a non-string output OutputGuardrail(guard=for_text(redact_secrets, on_other='allow')) # skips it deliberately A tool result guard receives `ToolResultInfo`, not just a value. Use `for_tool_result_text` to adapt a detector for a bare string result or a `ToolReturn.return_value` string: from pydantic_ai_harness.guardrails import ToolGuardrail from pydantic_ai_harness.guardrails.detectors import for_tool_result_text, redact_secrets ToolGuardrail(result_guard=for_tool_result_text(redact_secrets)) When it redacts a `ToolReturn`, the adapter replaces `return_value` and preserves its `content`, `metadata`, and `kind`. `ToolReturn.content` is sent as a separate user-prompt part, so the adapter does not inspect it. As with `for_text`, a non-text result raises by default; pass `on_other='allow'` to skip one deliberately. A detector on the input side needs a text prompt. A multimodal prompt reaches the guard rendered as text, so a detector that matches one returns `replace`, which `InputGuardrail` refuses rather than dropping the attached parts (see “Redaction (`replace`)” below). Guard the output instead when prompts can carry attachments. **Where a shape is not enough.** `email` matches anything shaped like an address, because that is all an address is: from pydantic_ai_harness.guardrails.detectors import redact_personal_data redact_personal_data('git clone git@github.com:pydantic/pydantic-ai.git') # replaces `git@github.com`, leaving a command that no longer runs An input guard rewrites the prompt in place, so the model receives the broken version. On an agent that handles code or paths, use `personal_data(only=['us_ssn', 'credit_card', 'iban'])`, or put the detector on the output rather than the prompt. `credit_card` has it less badly. It covers the 13 to 19 digits ISO/IEC 7812 allows, in any grouping, and only where the leading digit is 2 to 6, which is what a payment card starts with and a millisecond timestamp does not. Every match is then checked against the Luhn algorithm, which discards most of the runs that are left. Not all of them — roughly one run of four consecutive years in ten satisfies the checksum by chance, so a prompt listing years can still lose one. `iban` has the same problem and the same answer: a country code plus two digits is a shape ordinary text hits constantly, so spaces are allowed only where the printed form puts them and every match is checked against the ISO 7064 mod-97 digit an IBAN carries. An AWS _secret_ access key is deliberately absent from the defaults. It is forty characters of base64 with no distinguishing prefix, so nothing in the value marks it as a key and a pattern for the shape alone would take ordinary base64 with it. Matching one means anchoring on the name written beside it — `aws_secret_access_key = ...` — which is what [`pydantic-ai-shields`](https://github.com/vstorm-co/pydantic-ai-shields) does, and which finds the key only where it is written as an assignment. That narrower pattern is not shipped here; pass it through `extra=` if that is how keys reach your agent. `only=` selects patterns; it does not reorder them. The application order is part of each mapping’s contract — `iban` runs before `credit_card` so a spaced account number is not labelled a card — and it holds that order whatever order you list `only` in. A private key is matched whether its line breaks are real newlines or the escaped `\n` a JSON service-account file or a `.env` line carries, which is how one usually reaches a chat window. Terminated or not, it is one `private_key` pattern rather than two names, so `only=['private_key']` cannot select the complete block and leave a key pasted without its `END` marker unredacted. **What these do not do.** A regex finds a credential because credentials have a shape. It does not find a prompt injection, which is ordinary language, and it does not understand context, so a redactor will sometimes take a string that only looks like a key. They are one cheap layer, not the answer. `GuardrailResult` ----------------- [](https://pydantic.dev/docs/ai/harness/guardrails/#guardrailresult) Construct a `GuardrailResult` with its classmethods, not the raw fields: from pydantic_ai_harness import GuardrailResult GuardrailResult.allow() # let the value through GuardrailResult.block('reason') # refuse; `reason` is optional (a default is used otherwise) GuardrailResult.replace(cleaned_value) # substitute a sanitized value and continue GuardrailResult.retry('instruction') # ask the model to redo the output or the tool call GuardrailResult.approve() # ToolGuardrail arguments only: defer the call for human approval The block/retry message is produced at the moment the guard decides, so it can carry the guard’s own reasoning rather than a string frozen at construction time. Redaction (`replace`) --------------------- [](https://pydantic.dev/docs/ai/harness/guardrails/#redaction-replace) Return `GuardrailResult.replace(value)` to sanitize rather than refuse. `InputGuardrail` rewrites the prompt sent to the model; `OutputGuardrail` substitutes the output returned to the caller. def scrub_emails(text: str) -> GuardrailResult: cleaned = EMAIL_RE.sub('[email]', text) return GuardrailResult.replace(cleaned) if cleaned != text else GuardrailResult.allow() agent = Agent( 'openai:gpt-5.4', capabilities=[\ InputGuardrail(guard=scrub_emails), # strip PII before it reaches the model\ OutputGuardrail(guard=scrub_emails), # strip PII before it reaches the caller\ ], ) Input redaction requires sequential mode — it is incompatible with `parallel=True`, since a parallel guard runs alongside a model call that has already started with the original prompt. It also requires a text prompt. A guard sees a multimodal prompt (`agent.run_sync(['describe this', BinaryContent(...)])`) rendered as text, and one string written back over a prompt built from several parts would drop the attached images, documents and audio, so `replace` raises `UserError` there instead. Return `allow` or `block` for those prompts, or guard the output. Retry (`retry`) --------------- [](https://pydantic.dev/docs/ai/harness/guardrails/#retry-retry) `OutputGuardrail` can send a bad output back to the model instead of blocking it. Return `GuardrailResult.retry(instruction)` — the instruction is the retry prompt the model sees. This reuses pydantic-ai’s normal retry machinery and counts against the run’s output-retry budget. def must_cite_sources(output: object) -> GuardrailResult: if not has_citations(output): return GuardrailResult.retry('Include at least one source citation.') return GuardrailResult.allow() OutputGuardrail(guard=must_cite_sources) Accessing run context --------------------- [](https://pydantic.dev/docs/ai/harness/guardrails/#accessing-run-context) A guard may take a `RunContext` as its first parameter when it needs run state — `deps` for tenant- or role-aware policy, message history for conversation-aware checks. The parameter is detected from the signature, so prompt-only guards need not declare it: from pydantic_ai import RunContext from pydantic_ai_harness import InputGuardrail def tenant_policy(ctx: RunContext[MyDeps], prompt: str) -> bool: return ctx.deps.tier == 'pro' or 'advanced-feature' not in prompt InputGuardrail(guard=tenant_policy) Parallel input guards --------------------- [](https://pydantic.dev/docs/ai/harness/guardrails/#parallel-input-guards) A slow guard (an LLM classifier, a network call) run sequentially adds its latency to every turn. Set `parallel=True` to run the guard concurrently with the model call instead, overlapping the two so the guard adds no latency on the pass path. The model call is cancelled the moment the guard reports a violation. InputGuardrail(guard=slow_async_classifier, parallel=True) Parallel mode trades tokens for latency: sequential mode never calls the model when the guard blocks, but parallel mode has already started the model call — if the guard trips only after the model has responded, those tokens were spent. For fast local checks (regex, keyword lookup) sequential is the better default. `replace` is not available under `parallel=True`. Hard-fail path -------------- [](https://pydantic.dev/docs/ai/harness/guardrails/#hard-fail-path) `block` is the graceful path. To make the caller see an exception instead, raise from the guard: from pydantic_ai_harness import InputBlocked def strict_guard(prompt: str) -> bool: if contains_credentials(prompt): raise InputBlocked('credentials detected') return True Any exception raised by the guard propagates as-is — use `InputBlocked` / `OutputBlocked` / `ToolBlocked` from this module, or your own exception types. `ToolBlocked` carries the `tool_name` and an optional `reason`. Tool calls ---------- [](https://pydantic.dev/docs/ai/harness/guardrails/#tool-calls) `ToolGuardrail` inspects both sides of a tool call: `guard` sees the validated arguments before the tool runs, `result_guard` sees what it returned before the model does. from pathlib import Path import httpx from pydantic_ai import Agent from pydantic_ai_harness.guardrails import GuardrailResult, ToolCallInfo, ToolGuardrail, ToolResultInfo WORKSPACE = Path('/workspace') def stay_in_the_workspace(call: ToolCallInfo) -> GuardrailResult: if call.name == 'write_file': # `resolve()` before the containment check: a prefix test on the raw # string accepts `/workspace/../etc/passwd`. target = Path(str(call.args['path'])).resolve() if not target.is_relative_to(WORKSPACE): return GuardrailResult.block(f'{target} is outside the workspace.') return GuardrailResult.allow() def scrub_secrets(info: ToolResultInfo) -> GuardrailResult: # Text only. `str()` on a structured result or a `ToolReturn` yields a repr, and # replacing it with that string would change the result's type -- but only on the # calls where the pattern happened to match. if not isinstance(info.result, str): return GuardrailResult.allow() cleaned = SECRET_RE.sub('[redacted]', info.result) return GuardrailResult.replace(cleaned) if cleaned != info.result else GuardrailResult.allow() agent = Agent( 'openai:gpt-5.4', capabilities=[ToolGuardrail(guard=stay_in_the_workspace, result_guard=scrub_secrets)], ) @agent.tool_plain def write_file(path: str, content: str) -> str: Path(path).write_text(content) return f'wrote {path}' @agent.tool_plain def fetch_page(url: str) -> str: return httpx.get(url).text The outcomes map onto Pydantic AI control flow rather than a parallel mechanism: | Outcome | `guard` (arguments) | `result_guard` (result) | | --- | --- | --- | | **allow** | run the tool | return the result unchanged (the guard is handed the object the tool produced, so read it and use `replace` rather than mutating it) | | **block** | skip execution; the refusal message becomes the tool result (`SkipToolExecution`) | the refusal message replaces the result | | **replace** | run the tool with substituted arguments (a mapping); the call recorded in the message history keeps the model’s original arguments. The replacement is trusted to match the tool’s signature: keys the tool does not accept reach it as keyword arguments and raise a bare `TypeError` that names neither the tool nor the guard | substitute a sanitized result | | **retry** | ask the model to redo the call (`ModelRetry`) | ask the model to redo the call (`ModelRetry`); the tool has already run once, so its side effects have happened and the retry runs it again | | **approve** | defer the call for human approval (`ApprovalRequired`) | — (the tool has already run) | `block` is graceful on both stages: the agent sees the refusal text where it expected a tool result, so it can explain the refusal or try another approach. To fail the run instead, raise `ToolBlocked` from the guard. ### Human in the loop [](https://pydantic.dev/docs/ai/harness/guardrails/#human-in-the-loop) Pydantic AI already owns the approval round trip: a call raising `ApprovalRequired` is held back, the run finishes with a `DeferredToolRequests` output, and you resume it with the human’s answers. `ToolGuardrail` plugs into that rather than inventing a second mechanism, which means approvals a guard asks for and tools marked `requires_approval=True` arrive in the same place. from pydantic_ai import Agent, DeferredToolRequests, DeferredToolResults, ToolDenied from pydantic_ai_harness.guardrails import GuardrailResult, ToolCallInfo, ToolGuardrail def confirm_production(call: ToolCallInfo) -> GuardrailResult: if call.args.get('env') == 'prod': return GuardrailResult.approve() return GuardrailResult.allow() agent = Agent( 'openai:gpt-5.4', capabilities=[ToolGuardrail(guard=confirm_production)], output_type=[str, DeferredToolRequests], ) @agent.tool_plain def deploy(env: str) -> str: return f'deployed to {env}' deferred = await agent.run('deploy the new build') if isinstance(deferred.output, DeferredToolRequests): approvals = { call.tool_call_id: True if operator_says_yes(call) else ToolDenied('not on a Friday') for call in deferred.output.approvals } final = await agent.run( message_history=deferred.all_messages(), deferred_tool_results=DeferredToolResults(approvals=approvals), ) A denial reaches the model as the tool’s result, so the agent can explain itself or try something else. On the resumed run the guard is evaluated again, and `approve` becomes a no-op for a call the human already cleared — every other verdict still applies, so a policy that has since changed its mind can still block an approved call. Two shapes of approval, and which to reach for: | | Deferred (`GuardrailResult.approve()`) | In-process (an async guard) | | --- | --- | --- | | The run | ends, then resumes from `message_history` | stays open; the tool call awaits | | Fits | HTTP APIs, queues, durable execution, anything that cannot hold a process open | CLIs, TUIs, desktop apps, a websocket to an operator | | Human answer | `DeferredToolResults` (`True`, `ToolDenied`, or `ToolApproved(override_args=...)`) | whatever the guard returns | The in-process shape needs nothing extra — a guard may be async, so it can await the human directly: async def ask_the_operator(call: ToolCallInfo) -> GuardrailResult: if await operator_approves(call.name, call.args): return GuardrailResult.allow() return GuardrailResult.block('The operator declined this action.') ToolGuardrail(guard=ask_the_operator) Pydantic AI also offers approval without a guard at all: `requires_approval=True` on a tool, or `ApprovalRequiredToolset` for a synchronous predicate over a whole toolset. Reach for `ToolGuardrail` when the decision is async, needs `deps`, or should sit alongside the other verdicts. Two fields narrow what a guard sees: ToolGuardrail( guard=stay_in_the_workspace, tools=['write_file', 'run_shell'], # guard only these; None (default) guards every tool hidden=['delete_everything'], # withhold these from the model entirely ) `hidden` is not a blocklist with a nicer name. A hidden tool is dropped from the definitions sent to the model, so it costs no tokens and the model never attempts it; a blocked tool stays visible and the model learns it was refused. Hiding takes a static list of names — for policy that depends on `deps` or on the arguments, use `guard`. Configured `hidden` names are checked when a run completes successfully. Configured `tools` names are checked then only when `guard` or `result_guard` is set. A dynamic toolset may omit a tool on one step and offer it later, so a name that never appears is the only typo signal. That warning catches `tools=['send_monye']`, which otherwise leaves the intended tool unguarded, and a misspelled `hidden` name, which otherwise stays visible to the model. ### What a tool guard does not see [](https://pydantic.dev/docs/ai/harness/guardrails/#what-a-tool-guard-does-not-see) Three kinds of call never reach the execution hooks, so neither `guard` nor `result_guard` is consulted for them: * **Output tools**, which produce the agent’s structured output. Screen that with `OutputGuardrail`. * **External and deferred tools**, which the run hands back to your application in `DeferredToolRequests` instead of executing. Pydantic AI rejects them before any execution hook runs, so a guard cannot vet the arguments — your application is the thing that executes them, and the check belongs there. `hidden` _does_ cover them, since it works on the tool definitions. * **Provider-side builtin tools** such as web search, which run inside the provider and come back as builtin call/return parts rather than tool executions. A `ToolGuardrail` is a control over tools this run executes. For the ones it does not, `hidden` is the lever that still applies. Streaming --------- [](https://pydantic.dev/docs/ai/harness/guardrails/#streaming) `OutputGuardrail` inspects the **final** output only — during `run_stream()` partial chunks reach the caller before the guard runs, so a `block` or `replace` verdict cannot un-send content already streamed. Use `run()` / `run_sync()` when the output must be screened before any of it is exposed. `GuardrailResult.retry()` is **not** supported under `run_stream()` and surfaces there as `UnexpectedModelBehavior`. `InputGuardrail` (including `parallel=True`) works the same in streamed and non-streamed runs. Tracing ------- [](https://pydantic.dev/docs/ai/harness/guardrails/#tracing) `replace` and `block` are recorded as spans on the active OpenTelemetry tracer, so a redaction or refusal shows up in [Logfire](https://pydantic.dev/logfire) traces (`guardrail redacted input`, `guardrail blocked output`, and so on) with `guardrail.*` attributes. Content attributes — the original/replacement values for a redaction and the refusal `message` for a block — are attached **only** when `RunContext.trace_include_content` is enabled, since these can quote the very content the guard exists to keep out of traces. Tool spans add a `guardrail.tool` attribute naming the tool. `approve` records `guardrail deferred tool args`, which always carries `guardrail.tool_call_id` so the span can be correlated with the `DeferredToolRequests` the application answers, and carries `guardrail.arguments` under the same `trace_include_content` rule as the other content attributes. A deferred call never executes, so there is no `execute_tool` span recording what was asked for. `OutputGuardrail` positions its block/redact spans so they are always captured by an enclosing `Instrumentation` span regardless of capability order, while `InputGuardrail` runs innermost so any capability that morphs messages (a prompt rewriter, a context manager) runs first and the guard sees the final prompt the model will receive. `ToolGuardrail` also runs innermost, which puts it last among argument hooks (it sees the arguments every other capability has finished modifying) and first among result hooks (it sees the raw tool result, before a capability such as [`ToolOutputLimits`](https://pydantic.dev/docs/ai/harness/tool-output-limits/) truncates or offloads it). Relationship to `pydantic-ai-shields` ------------------------------------- [](https://pydantic.dev/docs/ai/harness/guardrails/#relationship-to-pydantic-ai-shields) [`pydantic-ai-shields`](https://github.com/vstorm-co/pydantic-ai-shields) ships each detector as its own capability. The equivalents here are functions instead, which is what lets several run as one chain, share a redaction, and sit beside a guard you wrote. Two things it has that this deliberately does not. Its `PromptInjection` matches phrases like “ignore previous instructions”: injection is ordinary language, so a pattern list catches the examples and misses the attack while flagging a pasted log, and a check that reads as protection without being it is worse than none. Its `NoRefusals` blocks the model from declining, which is a decision about what an agent may say rather than a guardrail on data, and not one to make a default. API --- [](https://pydantic.dev/docs/ai/harness/guardrails/#api) InputGuardrail( guard, # one guard, or a sequence run in order parallel=False, # run concurrently with the model call ) OutputGuardrail( guard, # one guard, or a sequence run in order ) ToolGuardrail( guard=None, # inspects a ToolCallInfo before the tool runs result_guard=None, # inspects a ToolResultInfo after it runs tools=None, # Sequence[str] | None -- restrict both guards to these tool names hidden=(), # Sequence[str] -- withhold these tools from the model entirely ) The guard callable takes the inspected value — the prompt for `InputGuardrail`, the output for `OutputGuardrail`, a `ToolCallInfo` or `ToolResultInfo` for `ToolGuardrail` — optionally preceded by a `RunContext`. `InputGuardrailFunc`, `OutputGuardrailFunc`, `ToolGuardrailFunc`, and `ToolResultGuardrailFunc` are the exported signature aliases; `GuardrailError` is the base for `InputBlocked`, `OutputBlocked`, and `ToolBlocked`. Source: [`pydantic_ai_harness/guardrails/`](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/guardrails/) . Was this page helpful? Thanks for your feedback! --- # Upgrade Guide | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/project/changelog/#_top) Upgrade Guide ============= In September 2025, Pydantic AI reached V1 and committed to API stability: no changes that break your code until V2. V2 is now available, collecting the breaking and behavior changes that stability guarantee didn’t allow. This guide is the canonical place to learn what’s in V2, how to install it, and how to upgrade; for the guarantees behind these version numbers, see the [Version Policy](https://pydantic.dev/docs/ai/project/version-policy/) . Breaking Changes ---------------- [](https://pydantic.dev/docs/ai/project/changelog/#breaking-changes) Here’s a filtered list of the breaking changes for each version to help you upgrade Pydantic AI. ### v2.0.0 (2026-06-23) [](https://pydantic.dev/docs/ai/project/changelog/#v200-2026-06-23) The stable V2.0 release. There are no new breaking or behavior changes since the betas; the full breaking-change list and recommended upgrade path are in the [v2.0.0b1](https://pydantic.dev/docs/ai/project/changelog/#v200b1-2026-05-20) entry below. Install it with: Terminal uv add pydantic-ai ### v2.0.0b7 (2026-06-10) [](https://pydantic.dev/docs/ai/project/changelog/#v200b7-2026-06-10) The seventh V2 beta, forked from **v1.107.0**. There are no new V2 breaking or behavior changes since [v2.0.0b6](https://pydantic.dev/docs/ai/project/changelog/#v200b6-2026-06-04) below — everything in that entry applies unchanged — but this beta picks up the latest V1 release on top, which adds Claude Fable 5 / Mythos 5 model support and OpenRouter prompt caching (`CachePoint`), plus `known_model_names()` and Anthropic fixes; see the [v1.107.0 release notes](https://github.com/pydantic/pydantic-ai/releases/tag/v1.107.0) for the full list. Install it the same way, pinning the exact pre-release version: * [pip](https://pydantic.dev/docs/ai/project/changelog/#tab-panel-158) * [uv](https://pydantic.dev/docs/ai/project/changelog/#tab-panel-159) Terminal pip install "pydantic-ai==2.0.0b7" Terminal uv add "pydantic-ai==2.0.0b7" For the full breaking-change list and the recommended upgrade path, see the [v2.0.0b1](https://pydantic.dev/docs/ai/project/changelog/#v200b1-2026-05-20) entry below; the only difference is that the latest V1 to upgrade through first is now **v1.107.0**. ### v2.0.0b6 (2026-06-04) [](https://pydantic.dev/docs/ai/project/changelog/#v200b6-2026-06-04) The sixth V2 beta, forked from **v1.106.0**. There are no new V2 breaking or behavior changes since [v2.0.0b5](https://pydantic.dev/docs/ai/project/changelog/#v200b5-2026-06-02) below — everything in that entry applies unchanged — but this beta picks up the latest V1 release on top, which adds `api_host`/`timeout` configuration and base `seed` mapping for the xAI provider, plus streaming and data-URI handling fixes; see the [v1.106.0 release notes](https://github.com/pydantic/pydantic-ai/releases/tag/v1.106.0) for the full list. Install it the same way, pinning the exact pre-release version: * [pip](https://pydantic.dev/docs/ai/project/changelog/#tab-panel-160) * [uv](https://pydantic.dev/docs/ai/project/changelog/#tab-panel-161) Terminal pip install "pydantic-ai==2.0.0b6" Terminal uv add "pydantic-ai==2.0.0b6" For the full breaking-change list and the recommended upgrade path, see the [v2.0.0b1](https://pydantic.dev/docs/ai/project/changelog/#v200b1-2026-05-20) entry below; the only difference is that the latest V1 to upgrade through first is now **v1.106.0**. ### v2.0.0b5 (2026-06-02) [](https://pydantic.dev/docs/ai/project/changelog/#v200b5-2026-06-02) The fifth V2 beta, forked from **v1.105.0**. There are no new V2 breaking or behavior changes since [v2.0.0b4](https://pydantic.dev/docs/ai/project/changelog/#v200b4-2026-05-28) below — everything in that entry (including the prepare-callbacks change) still applies — but this beta picks up the latest V1 release on top, which adds [on-demand (deferred-loading) capabilities](https://github.com/pydantic/pydantic-ai/pull/5230) and [Grok 4.3 `reasoning_effort` support](https://github.com/pydantic/pydantic-ai/pull/5454) , plus `GoogleModelSettings.google_cached_content` and Temporal `gateway/` fixes; see the [v1.105.0 release notes](https://github.com/pydantic/pydantic-ai/releases/tag/v1.105.0) for the full list. Install it the same way, pinning the exact pre-release version: * [pip](https://pydantic.dev/docs/ai/project/changelog/#tab-panel-162) * [uv](https://pydantic.dev/docs/ai/project/changelog/#tab-panel-163) Terminal pip install "pydantic-ai==2.0.0b5" Terminal uv add "pydantic-ai==2.0.0b5" For the full breaking-change list and the recommended upgrade path, see the [v2.0.0b1](https://pydantic.dev/docs/ai/project/changelog/#v200b1-2026-05-20) entry below; the only difference is that the latest V1 to upgrade through first is now **v1.105.0**. ### v2.0.0b4 (2026-05-28) [](https://pydantic.dev/docs/ai/project/changelog/#v200b4-2026-05-28) The fourth V2 beta, forked from **v1.104.0**. One new V2 behavior change since [v2.0.0b3](https://pydantic.dev/docs/ai/project/changelog/#v200b3-2026-05-22) : * Prepare callbacks (`prepare_tools=` / `PrepareTools` capability) that return `None` now raise `TypeError` instead of silently stripping all tools. V1.103.0 announces this change via `PydanticAIDeprecationWarning` (see [#5188](https://github.com/pydantic/pydantic-ai/pull/5188) ); V2 turns the warning into a hard error (see [#5668](https://github.com/pydantic/pydantic-ai/pull/5668) ). Return an empty list (`[]`) when you mean “no tools for this turn.” This beta also picks up two V1 releases on top — [v1.103.0](https://github.com/pydantic/pydantic-ai/releases/tag/v1.103.0) and [v1.104.0](https://github.com/pydantic/pydantic-ai/releases/tag/v1.104.0) — together adding Claude Opus 4.8 support, `McpServer.list_prompts` / `get_prompt`, message-timestamp roundtripping through `VercelAIAdapter`’s `UIMessage.metadata`, OpenRouter eager input streaming, and several Bedrock and UI fixes. See those release notes for the full list. Install it the same way, pinning the exact pre-release version: * [pip](https://pydantic.dev/docs/ai/project/changelog/#tab-panel-164) * [uv](https://pydantic.dev/docs/ai/project/changelog/#tab-panel-165) Terminal pip install "pydantic-ai==2.0.0b4" Terminal uv add "pydantic-ai==2.0.0b4" For the full breaking-change list and the recommended upgrade path, see the [v2.0.0b1](https://pydantic.dev/docs/ai/project/changelog/#v200b1-2026-05-20) entry below; the only difference is that the latest V1 to upgrade through first is now **v1.104.0**. ### v2.0.0b3 (2026-05-22) [](https://pydantic.dev/docs/ai/project/changelog/#v200b3-2026-05-22) The third V2 beta, forked from **v1.102.0**. There are no new V2 breaking changes since [v2.0.0b1](https://pydantic.dev/docs/ai/project/changelog/#v200b1-2026-05-20) below — everything in that entry applies unchanged — but this beta picks up the latest V1 release on top, which is a bug-fix release; see the [v1.102.0 release notes](https://github.com/pydantic/pydantic-ai/releases/tag/v1.102.0) for the full list. Install it the same way, pinning the exact pre-release version: * [pip](https://pydantic.dev/docs/ai/project/changelog/#tab-panel-166) * [uv](https://pydantic.dev/docs/ai/project/changelog/#tab-panel-167) Terminal pip install "pydantic-ai==2.0.0b3" Terminal uv add "pydantic-ai==2.0.0b3" For the full breaking-change list and the recommended upgrade path, see the [v2.0.0b1](https://pydantic.dev/docs/ai/project/changelog/#v200b1-2026-05-20) entry below; the only difference is that the latest V1 to upgrade through first is now **v1.102.0**. ### v2.0.0b2 (2026-05-21) [](https://pydantic.dev/docs/ai/project/changelog/#v200b2-2026-05-21) The second V2 beta, forked from **v1.101.0**. There are no new V2 breaking changes since [v2.0.0b1](https://pydantic.dev/docs/ai/project/changelog/#v200b1-2026-05-20) below — everything in that entry applies unchanged — but this beta picks up the latest V1 release on top, which adds the [pending message queue](https://github.com/pydantic/pydantic-ai/pull/4980) (`ctx.enqueue` / `agent_run.enqueue`). Install it the same way, pinning the exact pre-release version: * [pip](https://pydantic.dev/docs/ai/project/changelog/#tab-panel-168) * [uv](https://pydantic.dev/docs/ai/project/changelog/#tab-panel-169) Terminal pip install "pydantic-ai==2.0.0b2" Terminal uv add "pydantic-ai==2.0.0b2" For the full breaking-change list and the recommended upgrade path, see the [v2.0.0b1](https://pydantic.dev/docs/ai/project/changelog/#v200b1-2026-05-20) entry below; the only difference is that the latest V1 to upgrade through first is now **v1.101.0**. ### v2.0.0b1 (2026-05-20) [](https://pydantic.dev/docs/ai/project/changelog/#v200b1-2026-05-20) The first V2 beta, forked from **v1.100.0**, which deprecates most of what V2 removes. V2 leans into a harness-first design with [capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) as a core primitive: a single, composable unit that bundles an agent’s tools, [hooks](https://pydantic.dev/docs/ai/core-concepts/hooks/) , instructions, and model settings, reaching every layer of the agent through one concept. Many of V2’s changes move configuration that used to be spread across `Agent` arguments onto that primitive, alongside the behavior changes that V1’s stability guarantee didn’t allow. Pydantic AI stays a small core: some capabilities ship with it, more come from the first-party [Pydantic AI Harness](https://pydantic.dev/docs/ai/harness/) , and others are third-party or your own. The breaking changes below are split into two groups: * [**Changes not covered by deprecation warnings**](https://pydantic.dev/docs/ai/project/changelog/#changes-not-covered-by-deprecation-warnings) — removals and behavior changes that couldn’t be announced via a V1 deprecation warning. Review these even if you’re already on the latest V1 with no warnings. * [**Changes covered by deprecation warnings**](https://pydantic.dev/docs/ai/project/changelog/#changes-covered-by-deprecation-warnings) — if you upgraded to the latest V1 and resolved every deprecation warning, you’ve already made these. They’re listed with full before → after for reference. **Recommended upgrade path.** To make the jump as smooth as possible: 1. **Upgrade to the latest V1 release.** Most of what V2 removes is deprecated as of **v1.100.0** (the release this beta is forked from), so any V1 at or above that version surfaces those warnings. 2. **Resolve every deprecation warning.** The [changes covered by deprecation warnings](https://pydantic.dev/docs/ai/project/changelog/#changes-covered-by-deprecation-warnings) were announced in V1 via warnings that name the new API and, where possible, include a migration snippet. Run your test suite (or app) with warnings visible and address each one — by hand or by pointing a coding agent at them — to migrate across the bulk of V2 ahead of time. 3. **Upgrade to V2** and make the [changes not covered by deprecation warnings](https://pydantic.dev/docs/ai/project/changelog/#changes-not-covered-by-deprecation-warnings) — primarily default-behavior changes and a handful of removals with no V1 deprecation. You can also upgrade straight to V2 and work through the list below directly — it’s organized so a coding agent can apply the code changes mechanically. Resolving deprecation warnings on the latest V1 first is still the smoother path, since it spreads the work out and leaves you only the behavior changes to reason about consciously at the end. Message history serialized with V1 (via [`ModelMessagesTypeAdapter`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessagesTypeAdapter) ) continues to deserialize in V2. #### Changes not covered by deprecation warnings [](https://pydantic.dev/docs/ai/project/changelog/#changes-not-covered-by-deprecation-warnings) These removals and behavior changes could not be announced via a V1 deprecation warning, so review them even if you’ve resolved every deprecation warning on the latest V1. **Code changes:** * Generic type parameter defaults changed from `None` to `object`: an un-parameterized `Agent(...)` now infers `Agent[object, str]` instead of `Agent[None, str]`, and the `pydantic_graph` `StateT`/`RunEndT`/`DepsT` defaults changed to match. Update explicit `Agent[None, ...]`, `RunContext[None]`, and `Tool[None]` annotations that don’t actually require `None` dependencies to use `object`. This is a type-checking-only change; runtime behavior is unchanged. See [#5307](https://github.com/pydantic/pydantic-ai/pull/5307) . * The `pydantic_graph.persistence` package and the `pydantic_graph.mermaid` module are removed, with no V2 equivalent for standalone Mermaid generation (render diagrams with `Graph.render()`). The builder API deliberately does not snapshot graph state, so there is no `pydantic_graph` replacement for the persistence package; to save, resume, and fork **agent run** state, [Pydantic AI Harness](https://pydantic.dev/docs/ai/harness/) ships the [`StepPersistence`](https://pydantic.dev/docs/ai/harness/step-persistence/) capability. The move of the [`GraphBuilder`](https://pydantic.dev/docs/ai/api/pydantic_graph/graph_builder/#pydantic_graph.graph_builder.GraphBuilder) API out of `pydantic_graph.beta` to the top-level `pydantic_graph` _was_ deprecation-announced; see [below](https://pydantic.dev/docs/ai/project/changelog/#changes-covered-by-deprecation-warnings) . See [#5470](https://github.com/pydantic/pydantic-ai/pull/5470) . * [`ModelProfile`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfile) and its subclasses are now `TypedDict`s instead of dataclasses. Passing `profile=OpenAIModelProfile(field=value)` into a model still works unchanged; the migration only matters if you read or mutate profile fields, or call `.update()`/`.from_profile()`. See [`ModelProfile` is now a `TypedDict`](https://pydantic.dev/docs/ai/project/changelog/#modelprofile-is-now-a-typeddict) below. ([#5481](https://github.com/pydantic/pydantic-ai/pull/5481) ) **Default behavior changes** — same API, different runtime behavior (roughly ordered by how many users they affect): * A bare `uv add pydantic-ai` / `pip install pydantic-ai` now installs a slimmer set of extras (frontier providers plus minimal integrations); providers like `bedrock`, `groq`, and `mistral` are no longer included by default, so you’ll need to add the extras you use. See [Slimmer default extras](https://pydantic.dev/docs/ai/project/changelog/#slimmer-default-pydantic-ai-extras) below. ([#5467](https://github.com/pydantic/pydantic-ai/pull/5467) ) * The default `end_strategy` changed from `'early'` to `'graceful'`: when a model calls function tools in the same response as a successful output tool, those function tools now run (and their side effects happen) instead of being skipped, and tool calls run in the order the model emitted them. See [Parallel tool-call execution order](https://pydantic.dev/docs/ai/project/changelog/#parallel-tool-call-execution-runs-in-emission-order) below. ([#5339](https://github.com/pydantic/pydantic-ai/pull/5339) ) * The default instrumentation format is now version 5, and agent run spans report token usage under `gen_ai.aggregated_usage.*`. See [Instrumentation defaults](https://pydantic.dev/docs/ai/project/changelog/#instrumentation-defaults-to-version-5-with-aggregated-usage-attributes) below. ([#5523](https://github.com/pydantic/pydantic-ai/pull/5523) ) * [`capture_run_messages()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.capture_run_messages) now also captures the partial `ModelRequest`/`ModelResponse` from an interrupted run, marked with `state='interrupted'` (a new `ModelRequest.state` field is added). Code that asserts on exact captured-message counts on error paths may need updating. See [#5364](https://github.com/pydantic/pydantic-ai/pull/5364) . * Output tool calls and returns now emit dedicated `OutputToolCallEvent`/`OutputToolResultEvent` instead of `FunctionToolCallEvent`/`FunctionToolResultEvent`. Separately, native tool calls and returns no longer emit dedicated events at all — the `BuiltinToolCallEvent`/`BuiltinToolResultEvent` classes are removed and they surface only via the standard `PartStartEvent`/`PartDeltaEvent`. See [#5332](https://github.com/pydantic/pydantic-ai/pull/5332) and [#5476](https://github.com/pydantic/pydantic-ai/pull/5476) . ##### [`ModelProfile`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfile) is now a `TypedDict` [](https://pydantic.dev/docs/ai/project/changelog/#modelprofile-is-now-a-typeddict) See the [Model Profile guide](https://pydantic.dev/docs/ai/models/openai/#model-profile) for an overview of what a model profile is and how to configure one. [`ModelProfile`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.ModelProfile) and all its subclasses ([`OpenAIModelProfile`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.openai.OpenAIModelProfile) , [`AnthropicModelProfile`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.anthropic.AnthropicModelProfile) , [`GoogleModelProfile`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.google.GoogleModelProfile) , `BedrockModelProfile`, etc.) are now `TypedDict(total=False)` instead of `@dataclass`. This unifies the mental model with [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) (also a `TypedDict`) and enables direct dict-spread for cross-class merging. `ModelProfile.update()` and `ModelProfile.from_profile()` are removed; use the module-level [`merge_profile`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.merge_profile) (later argument wins per key). Migration recipes: | v1 (dataclass) | v2 (TypedDict) | | --- | --- | | `OpenAIModelProfile(field=value)` | Same syntax; returns a partial `dict` instead of a fully-defaulted instance. | | `profile.field` (attribute read) | `profile.get('field', )` — non-trivial defaults are exported from [`pydantic_ai.profiles`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles)
(e.g. [`DEFAULT_THINKING_TAGS`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.DEFAULT_THINKING_TAGS)
, [`DEFAULT_PROMPTED_OUTPUT_TEMPLATE`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.DEFAULT_PROMPTED_OUTPUT_TEMPLATE)
); the fully-merged base is [`DEFAULT_PROFILE`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.DEFAULT_PROFILE)
. | | `profile.field = value` (attribute write) | `profile['field'] = value` | | `dataclasses.replace(profile, field=value)` | `{**profile, 'field': value}` or `merge_profile(profile, ModelProfile(field=value))` | | `profile.update(other)` | `merge_profile(profile, other)` | | `OpenAIModelProfile.from_profile(p)` | Just `p` — no upcasting needed | | `Model(name, profile=full_profile)` (full replace) | Now merges on top of the provider’s default profile — usually what you want. For a hard replace use `Model(name, profile=lambda _default: full_profile)`. | | `Model(name, profile=fn)` where `fn: Callable[[str], ModelProfile \| None]` | Removed — the user-passed callable is now `Callable[[ModelProfile], ModelProfile]`, receiving the resolved default and returning the final profile. The `(model_name: str) -> ModelProfile \| None` shape is still accepted internally by `Provider.model_profile`. | | `isinstance(profile, OpenAIModelProfile)` | Not supported by `TypedDict` at runtime — raises `TypeError`. Use `isinstance(profile, dict)` or check key presence (`'openai_chat_supports_web_search' in profile`). Pyright still narrows correctly via the TypedDict subclass annotation. | `Model.profile` is now the single source of truth for the **resolved** profile. It is composed by [`merge_profile`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.merge_profile) in this order (later wins): 1. [`DEFAULT_PROFILE`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.DEFAULT_PROFILE) — base defaults for every documented key. 2. `Provider.model_profile(model_name)` — provider/model-specific resolution. 3. The user’s `profile=` argument — either a partial dict (merged on top) or a `Callable[[ModelProfile], ModelProfile]` (full control: receives the resolved default, returns the final profile). ##### Resolved profiles now carry cross-class fields [](https://pydantic.dev/docs/ai/project/changelog/#resolved-profiles-now-carry-cross-class-fields) In v1, `ModelProfile.update()` silently filtered out fields not declared on the target class. In v2, dict-spread preserves every key. This means e.g. a Bedrock-hosted Anthropic model’s resolved profile now carries the upstream `anthropic_*` fields alongside the `bedrock_*` fields, where v1 dropped them. No in-tree model class reads cross-class fields, so behavior is unchanged in the standard providers; but custom model classes that do `profile.get('anthropic_supports_adaptive_thinking', False)` on a non-Anthropic route will now see the value the upstream Anthropic profile set, where v1 always returned the default. See the [Model Profile guide](https://pydantic.dev/docs/ai/models/openai/#model-profile) for how to configure a profile, and [PR #5481](https://github.com/pydantic/pydantic-ai/pull/5481) for the full `ModelProfile` redesign. ##### Parallel tool-call execution runs in emission order [](https://pydantic.dev/docs/ai/project/changelog/#parallel-tool-call-execution-runs-in-emission-order) The default [`end_strategy`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.EndStrategy) changed from `'early'` to `'graceful'`. This only affects responses where a model calls function tools in the _same_ response as an [output tool](https://pydantic.dev/docs/ai/core-concepts/output/#tool-output) (the call that ends the run). When that output tool **succeeds**, the function tools requested alongside it now **run** by default instead of being skipped, so their side effects happen and their results reach the model if the run continues; and a function tool’s [`ModelRetry`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelRetry) now suppresses the output result so the model can correct itself on the next round. The case where _every_ output tool fails is unchanged: function tools run and the run continues either way. Most agents don’t need any change. If you relied on the run ending the instant an output tool succeeds — skipping any function tools requested in the same response — set `end_strategy='early'` explicitly. The [`sequential=True`](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#parallel-tool-calls-concurrency) flag on a tool is now a per-tool **barrier** rather than a batch-wide serial switch: a sequential tool runs alone, but other tools in the same response still run in parallel around it. The barrier now also applies to output tools via [`ToolOutput(sequential=True)`](https://pydantic.dev/docs/ai/api/pydantic-ai/output/#pydantic_ai.output.ToolOutput) , not just function tools. To run _all_ of a run’s tools serially, wrap the run in [`agent.parallel_tool_call_execution_mode('sequential')`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.parallel_tool_call_execution_mode) or set `parallel_tool_calls=False` on the [model settings](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) . See [Parallel Output Tool Calls](https://pydantic.dev/docs/ai/core-concepts/output/#parallel-output-tool-calls) for the full behavior of all three strategies, and [#5339](https://github.com/pydantic/pydantic-ai/pull/5339) . ##### Slimmer default `pydantic-ai` extras [](https://pydantic.dev/docs/ai/project/changelog/#slimmer-default-pydantic-ai-extras) A bare `uv add pydantic-ai` / `pip install pydantic-ai` now installs `pydantic-ai-slim[openai,anthropic,google,cli,mcp,evals,web,retries,logfire]` — frontier providers plus minimal integrations. Providers and integrations that were previously bundled are no longer installed by default; add the ones you use explicitly, e.g. `uv add 'pydantic-ai[bedrock,groq]'`: `bedrock`, `groq`, `mistral`, `cohere`, `xai`, `huggingface`, `temporal`, `ag-ui`, `ui`, and `spec`. See the [installation guide](https://pydantic.dev/docs/ai/overview/install/) for the full list of extras. Some `pydantic-ai-slim` extras were also removed outright (not just dropped from the default bundle): the `outlines-*` extras (the Outlines integration is removed), `vertexai` (Vertex AI is now served by the `google` extra), `fastmcp` (the FastMCP back-compat shim is removed), and `a2a` (A2A now lives in the upstream `fasta2a` package). See [#5467](https://github.com/pydantic/pydantic-ai/pull/5467) . ##### Instrumentation defaults to version 5 with aggregated usage attributes [](https://pydantic.dev/docs/ai/project/changelog/#instrumentation-defaults-to-version-5-with-aggregated-usage-attributes) The default [instrumentation format](https://pydantic.dev/docs/ai/integrations/logfire/#configuring-data-format) is now version 5 (versions 2–4 still work but emit a deprecation warning; version 1 and its `event_mode=`/`logger_provider=` arguments are removed). In version 5, deferred tool calls (`CallDeferred`/`ApprovalRequired`) are no longer recorded as span errors. Separately, [`InstrumentationSettings`](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentationSettings) ’s [`use_aggregated_usage_attribute_names`](https://pydantic.dev/docs/ai/integrations/logfire/#aggregated-usage-attribute-names) now defaults to `True`: agent run spans report token usage under `gen_ai.aggregated_usage.*` while model request spans keep `gen_ai.usage.*`, which avoids double-counting in backends that sum parent and child usage. Dashboards and alerts that read token usage from run spans must be updated, or set `use_aggregated_usage_attribute_names=False` to keep the V1 attribute names. See [#5523](https://github.com/pydantic/pydantic-ai/pull/5523) . #### Changes covered by deprecation warnings [](https://pydantic.dev/docs/ai/project/changelog/#changes-covered-by-deprecation-warnings) These changes were announced in the latest V1 releases via deprecation warnings that name the replacement API. If you upgraded to the latest V1 and resolved every warning, you’ve already made them. The [V1 → V2 migration map](https://pydantic.dev/docs/ai/overview/migration/) carries the full before → after for every symbol; this section records the behavior that changes with them and the PRs each one landed in. **Behavior changes that flip silently if the V1 deprecation warning was not addressed** — even though these were announced, an unaddressed warning means the behavior changes without raising an error, so confirm you’ve handled them: * The bare `openai:` model prefix now uses the OpenAI Responses API ([`OpenAIResponsesModel`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIResponsesModel) ) instead of the Chat Completions API ([`OpenAIChatModel`](https://pydantic.dev/docs/ai/api/models/openai/#pydantic_ai.models.openai.OpenAIChatModel) ). Use `openai-chat:` to keep Chat Completions, or `openai-responses:` to opt into the new default explicitly. Announced via [#5334](https://github.com/pydantic/pydantic-ai/pull/5334) ; flipped in [#5469](https://github.com/pydantic/pydantic-ai/pull/5469) . * Provider-adaptive `WebSearch` and `WebFetch` capabilities are now native-only and raise on models that don’t support them, and `MCP(url=...)` runs the server locally by default. Restore the V1 fallbacks with `WebSearch(local='duckduckgo')`, `WebFetch(local=True)`, and `MCP(url=..., native=True)`. Announced via [#5331](https://github.com/pydantic/pydantic-ai/pull/5331) ; changed in [#5333](https://github.com/pydantic/pydantic-ai/pull/5333) . **API removals and renames** — look each old name up in the [migration map](https://pydantic.dev/docs/ai/overview/migration/) ; the PRs that announced and landed each group are: | Group | PRs | | --- | --- | | Providers: Grok → xAI (`grok:` prefix → `xai:`) | [#5460](https://github.com/pydantic/pydantic-ai/pull/5460) | | Providers: Google GLA/Vertex → `GoogleProvider`/`GoogleCloudProvider`, `GeminiModel` → `GoogleModel` | [#5336](https://github.com/pydantic/pydantic-ai/pull/5336)
, [#5543](https://github.com/pydantic/pydantic-ai/pull/5543)
, [#5479](https://github.com/pydantic/pydantic-ai/pull/5479) | | Models: `OpenAIModel` → `OpenAIChatModel`, `system_prompt_role` and sampling-settings move into the profile | [#5468](https://github.com/pydantic/pydantic-ai/pull/5468) | | Models: `StreamedResponse.usage()` becomes a property (affects custom `Model` subclasses) | [#5546](https://github.com/pydantic/pydantic-ai/pull/5546) | | Models: bare provider-prefix-less names (`Agent('gpt-5')`) now raise `UserError` | [#5464](https://github.com/pydantic/pydantic-ai/pull/5464) | | Native tools: `builtin_tools` → `native_tools` throughout | [#5338](https://github.com/pydantic/pydantic-ai/pull/5338)
, [#5396](https://github.com/pydantic/pydantic-ai/pull/5396) | | Native tools: `UrlContextTool` → `WebFetchTool` | [#5458](https://github.com/pydantic/pydantic-ai/pull/5458) | | MCP: per-transport server classes → `MCPToolset` | [#5325](https://github.com/pydantic/pydantic-ai/pull/5325)
, [#5337](https://github.com/pydantic/pydantic-ai/pull/5337) | | Agent config → capabilities: `instrument=` | [#5434](https://github.com/pydantic/pydantic-ai/pull/5434) | | Agent config → capabilities: `event_stream_handler=`, `prepare_tools=` | [#5335](https://github.com/pydantic/pydantic-ai/pull/5335)
, [#5475](https://github.com/pydantic/pydantic-ai/pull/5475) | | Agent config → capabilities: `history_processors=` | [#5425](https://github.com/pydantic/pydantic-ai/pull/5425) | | Agent config: `mcp_servers=` → `toolsets=`, `sequential_tool_calls()` → `parallel_tool_call_execution_mode()` | [#5466](https://github.com/pydantic/pydantic-ai/pull/5466) | | Tools: `DeferredToolCalls` → `DeferredToolRequests`, `DeferredToolset` → `ExternalToolset` | [#5459](https://github.com/pydantic/pydantic-ai/pull/5459) | | Tools: `FunctionToolset.tool()` now requires a `RunContext` first parameter | [#5462](https://github.com/pydantic/pydantic-ai/pull/5462) | | Tools: `pydantic_ai.ext.aci` removed (wrap with `Tool.from_schema`) | [#5510](https://github.com/pydantic/pydantic-ai/pull/5510)
, [#5467](https://github.com/pydantic/pydantic-ai/pull/5467) | | Usage and response-field renames (`request_tokens` → `input_tokens`, `vendor_details` → `provider_details`, …) | [#5476](https://github.com/pydantic/pydantic-ai/pull/5476) | | Events: dedicated `OutputToolCallEvent`/`OutputToolResultEvent`, `FunctionToolResultEvent.result` → `.part` | [#5332](https://github.com/pydantic/pydantic-ai/pull/5332) | | Streaming: `StreamedRunResult` accessor renames, `stream_responses()` → `stream_response()` | [#5296](https://github.com/pydantic/pydantic-ai/pull/5296)
, [#5463](https://github.com/pydantic/pydantic-ai/pull/5463) | | Streaming: `Agent.run_stream_events()` is an async context manager only | [#5440](https://github.com/pydantic/pydantic-ai/pull/5440) | | Results: `result.usage()`/`result.timestamp()`/`stream.get()` become properties | [#5263](https://github.com/pydantic/pydantic-ai/pull/5263) | | Integrations: `Agent.to_a2a()` and the bundled `fasta2a` move upstream | [#5426](https://github.com/pydantic/pydantic-ai/pull/5426)
, [#5502](https://github.com/pydantic/pydantic-ai/pull/5502) | | Integrations: `Agent.to_ag_ui()`/`AGUIApp` → `AGUIAdapter`, `cached_async_http_client` → `create_async_http_client()` | [#5345](https://github.com/pydantic/pydantic-ai/pull/5345)
, [#5464](https://github.com/pydantic/pydantic-ai/pull/5464) | | Graph: `pydantic_graph.beta` imports move to top-level `pydantic_graph` | [#5306](https://github.com/pydantic/pydantic-ai/pull/5306)
, [#5470](https://github.com/pydantic/pydantic-ai/pull/5470) | | Instrumentation: `version=1` and its `event_mode=`/`logger_provider=` arguments removed | [#5523](https://github.com/pydantic/pydantic-ai/pull/5523) | | Pydantic Evals: keyword-only arguments, required `Dataset(name=...)`, `Evaluator.name` → `get_serialization_name()` | [#5547](https://github.com/pydantic/pydantic-ai/pull/5547)
, [#5548](https://github.com/pydantic/pydantic-ai/pull/5548) | | Pydantic Evals: `evaluation_name`/`evaluator_version` class attributes → `get_default_evaluation_name()`/`get_evaluator_version()` | [#5554](https://github.com/pydantic/pydantic-ai/pull/5554)
, [#5556](https://github.com/pydantic/pydantic-ai/pull/5556) | Four of these carry a caveat the name mapping alone doesn’t convey: * **MCP:** `Agent.set_mcp_sampling_model()` is _not_ removed — it still sets the sampling model on every registered `MCPToolset`, while `MCPToolset(sampling_model=...)` sets it on one toolset at construction. * **Instrumentation:** `version=2`/`3`/`4` still work but now emit a deprecation warning, and the default is `version=5` — see [Instrumentation defaults](https://pydantic.dev/docs/ai/project/changelog/#instrumentation-defaults-to-version-5-with-aggregated-usage-attributes) above for the behavior changes that ship with it. * **`Agent(instrument=...)`:** only the constructor arguments are removed. The `Agent.instrument` property, `Agent.instrument_all()`, and `InstrumentedModel` are unchanged. * **Outlines:** the integration is removed outright. If you’d like to keep using Outlines with Pydantic AI, please file an issue at [dottxt-ai/outlines](https://github.com/dottxt-ai/outlines/issues) . See [#5444](https://github.com/pydantic/pydantic-ai/pull/5444) . ### v1.0.1 (2025-09-05) [](https://pydantic.dev/docs/ai/project/changelog/#v101-2025-09-05) The following breaking change was accidentally left out of v1.0.0: * See [#2808](https://github.com/pydantic/pydantic-ai/pull/2808) - Remove `Python` evaluator from `pydantic_evals` for security reasons ### v1.0.0 (2025-09-04) [](https://pydantic.dev/docs/ai/project/changelog/#v100-2025-09-04) * See [#2725](https://github.com/pydantic/pydantic-ai/pull/2725) - Drop support for Python 3.9 * See [#2738](https://github.com/pydantic/pydantic-ai/pull/2738) - Make many dataclasses require keyword arguments * See [#2715](https://github.com/pydantic/pydantic-ai/pull/2715) - Remove `cases` and `averages` attributes from `pydantic_evals` spans * See [#2798](https://github.com/pydantic/pydantic-ai/pull/2798) - Change `ModelRequest.parts` and `ModelResponse.parts` types from `list` to `Sequence` * See [#2726](https://github.com/pydantic/pydantic-ai/pull/2726) - Default `InstrumentationSettings` version to 2 * See [#2717](https://github.com/pydantic/pydantic-ai/pull/2717) - Remove errors when passing `AsyncRetrying` or `Retrying` object to `AsyncTenacityTransport` or `TenacityTransport` instead of `RetryConfig` ### v0.x.x [](https://pydantic.dev/docs/ai/project/changelog/#v0xx) Before V1, minor versions were used to introduce breaking changes: **v0.8.0 (2025-08-26)** See [#2689](https://github.com/pydantic/pydantic-ai/pull/2689) - `AgentStreamEvent` was expanded to be a union of `ModelResponseStreamEvent` and `HandleResponseEvent`, simplifying the `event_stream_handler` function signature. Existing code accepting `AgentStreamEvent | HandleResponseEvent` will continue to work. **v0.7.6 (2025-08-26)** The following breaking change was inadvertently released in a patch version rather than a minor version: See [#2670](https://github.com/pydantic/pydantic-ai/pull/2670) - `TenacityTransport` and `AsyncTenacityTransport` now require the use of `pydantic_ai.retries.RetryConfig` (which is just a `TypedDict` containing the kwargs to `tenacity.retry`) instead of `tenacity.Retrying` or `tenacity.AsyncRetrying`. **v0.7.0 (2025-08-12)** See [#2458](https://github.com/pydantic/pydantic-ai/pull/2458) - `pydantic_ai.models.StreamedResponse` now yields a `FinalResultEvent` along with the existing `PartStartEvent` and `PartDeltaEvent`. If you’re using `pydantic_ai.direct.model_request_stream` or `pydantic_ai.direct.model_request_stream_sync`, you may need to update your code to account for this. See [#2458](https://github.com/pydantic/pydantic-ai/pull/2458) - `pydantic_ai.models.Model.request_stream` now receives a `run_context` argument. If you’ve implemented a custom `Model` subclass, you will need to account for this. See [#2458](https://github.com/pydantic/pydantic-ai/pull/2458) - `pydantic_ai.models.StreamedResponse` now requires a `model_request_parameters` field and constructor argument. If you’ve implemented a custom `Model` subclass and implemented `request_stream`, you will need to account for this. **v0.6.0 (2025-08-06)** This release was meant to clean some old deprecated code, so we can get a step closer to V1. See [#2440](https://github.com/pydantic/pydantic-ai/pull/2440) - The `next` method was removed from the `Graph` class. Use `async with graph.iter(...) as run: run.next()` instead. See [#2441](https://github.com/pydantic/pydantic-ai/pull/2441) - The `result_type`, `result_tool_name` and `result_tool_description` arguments were removed from the `Agent` class. Use `output_type` instead. See [#2441](https://github.com/pydantic/pydantic-ai/pull/2441) - The `result_retries` argument was also removed from the `Agent` class. Use `output_retries` instead. See [#2443](https://github.com/pydantic/pydantic-ai/pull/2443) - The `data` property was removed from the `FinalResult` class. Use `output` instead. See [#2445](https://github.com/pydantic/pydantic-ai/pull/2445) - The `get_data` and `validate_structured_result` methods were removed from the `StreamedRunResult` class. Use `get_output` and `validate_response_output` instead. See [#2446](https://github.com/pydantic/pydantic-ai/pull/2446) - The `format_as_xml` function was moved to the `pydantic_ai.format_as_xml` module. Import it via `from pydantic_ai import format_as_xml` instead. See [#2451](https://github.com/pydantic/pydantic-ai/pull/2451) - Removed deprecated `Agent.result_validator` method, `Agent.last_run_messages` property, `AgentRunResult.data` property, and `result_tool_return_content` parameters from result classes. **v0.5.0 (2025-08-04)** See [#2388](https://github.com/pydantic/pydantic-ai/pull/2388) - The `source` field of an `EvaluationResult` is now of type `EvaluatorSpec` rather than the actual source `Evaluator` instance, to help with serialization/deserialization. See [#2163](https://github.com/pydantic/pydantic-ai/pull/2163) - The `EvaluationReport.print` and `EvaluationReport.console_table` methods now require most arguments be passed by keyword. **v0.4.0 (2025-07-08)** See [#1799](https://github.com/pydantic/pydantic-ai/pull/1799) - Pydantic Evals `EvaluationReport` and `ReportCase` are now generic dataclasses instead of Pydantic models. If you were serializing them using `model_dump()`, you will now need to use the `EvaluationReportAdapter` and `ReportCaseAdapter` type adapters instead. See [#1507](https://github.com/pydantic/pydantic-ai/pull/1507) - The `ToolDefinition` `description` argument is now optional and the order of positional arguments has changed from `name, description, parameters_json_schema, ...` to `name, parameters_json_schema, description, ...` to account for this. **v0.3.0 (2025-06-18)** See [#1142](https://github.com/pydantic/pydantic-ai/pull/1142) — Adds support for thinking parts. We now convert the thinking blocks (`"...""`) in provider specific text parts to Pydantic AI `ThinkingPart`s. Also, as part of this release, we made the choice to not send back the `ThinkingPart`s to the provider - the idea is to save costs on behalf of the user. In the future, we intend to add a setting to customize this behavior. **v0.2.0 (2025-05-12)** See [#1647](https://github.com/pydantic/pydantic-ai/pull/1647) — usage makes sense as part of `ModelResponse`, and could be really useful in “messages” (really a sequence of requests and response). In this PR: * Adds `usage` to `ModelResponse` (field has a default factory of `Usage()` so it’ll work to load data that doesn’t have usage) * changes the return type of `Model.request` to just `ModelResponse` instead of `tuple[ModelResponse, Usage]` **v0.1.0 (2025-04-15)** See [#1248](https://github.com/pydantic/pydantic-ai/pull/1248) — the attribute/parameter name `result` was renamed to `output` in many places. Hopefully all changes keep a deprecated attribute or parameter with the old name, so you should get many deprecation warnings. See [#1484](https://github.com/pydantic/pydantic-ai/pull/1484) — `format_as_xml` was moved and made available to import from the package root, e.g. `from pydantic_ai import format_as_xml`. Full Changelog -------------- [](https://pydantic.dev/docs/ai/project/changelog/#full-changelog) For the full changelog, see [GitHub Releases](https://github.com/pydantic/pydantic-ai/releases) . Was this page helpful? Thanks for your feedback! --- # Subagents | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/subagents/#_top) Subagents ========= `SubAgents` lets an agent delegate self-contained tasks to named child agents. It takes a sequence of `SubAgent` entries and exposes a single `delegate_task(agent_name, task)` tool. Each delegation runs the chosen sub-agent in its own run — with its own message history, so it never sees the parent conversation — and returns its output to the parent. [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/subagents/) > The API may change between releases. Where practical, breaking changes ship with a deprecation warning. The problem ----------- [](https://pydantic.dev/docs/ai/harness/subagents/#the-problem) A single agent that does everything accumulates a large tool set and a long context. Splitting the work across specialized sub-agents keeps each context focused, but wiring up delegation by hand means writing a tool per agent, forwarding deps, threading usage limits, and telling the model what it can delegate to. The solution ------------ [](https://pydantic.dev/docs/ai/harness/subagents/#the-solution) `SubAgents` takes a sequence of `SubAgent` entries and exposes a single `delegate_task(agent_name, task)` tool. Each delegation runs the chosen sub-agent in its own run — with its own message history, so it never sees the parent conversation — and returns its output to the parent. The available sub-agents are listed in the system prompt as a static instruction, so the listing stays in the cached prefix. from pydantic_ai import Agent from pydantic_ai_harness.subagents import SubAgent, SubAgents researcher = Agent('anthropic:claude-sonnet-4-6', name='researcher', description='Researches a topic and reports findings') writer = Agent('anthropic:claude-sonnet-4-6', name='writer', description='Turns notes into polished prose') orchestrator = Agent( 'anthropic:claude-opus-4-7', capabilities=[SubAgents(agents=[SubAgent(researcher), SubAgent(writer)])], ) result = orchestrator.run_sync('Research the history of TLS and write a one-paragraph summary.') print(result.output) A delegate’s name — how the parent model refers to it, and how it is listed in the prompt — is the agent’s own `name`, or a `SubAgent(name=...)` override. Two delegates resolving to the same name is an error, and an agent with no name and no override is rejected. The tool -------- [](https://pydantic.dev/docs/ai/harness/subagents/#the-tool) | Tool | Purpose | | --- | --- | | `delegate_task(agent_name, task)` | Run the named sub-agent on a self-contained task and return its output. | * The sub-agent runs with its own message history, so `task` must be self-contained. * An unknown `agent_name` raises `ModelRetry`, so the model can correct itself. * The result returned to the parent is `str(result.output)`. * With a `models` menu configured, the tool takes an extra `model` argument (see below). Deps, usage, tools, and capabilities ------------------------------------ [](https://pydantic.dev/docs/ai/harness/subagents/#deps-usage-tools-and-capabilities) * **Deps are forwarded.** The parent run’s `deps` are passed to each sub-agent, so sub-agents share the parent’s `AgentDepsT` (enforced by the type signature — every sub-agent is an `AbstractAgent[AgentDepsT, Any]`). * **Usage is shared by default.** The parent’s `usage` is passed to each sub-agent run, so token usage aggregates and a parent `usage_limits` applies across the whole agent tree. Set `forward_usage=False` to give each sub-agent run its own accounting. * **Tools can be inherited.** With `inherit_tools=True`, the parent agent’s own tools (registered directly or via `toolsets`) are added to each sub-agent run, on top of the sub-agent’s own. Tools contributed by the parent’s capabilities are not inherited: they are bound to capability instances registered in the parent run, and would arrive without the hooks and instructions they depend on. Use `shared_capabilities` to give sub-agents a capability. This also excludes the delegate tool itself, so a sub-agent can’t recurse into further delegation. Off by default. * **Capabilities can be shared.** `shared_capabilities` are applied to every sub-agent run — e.g. give all sub-agents a common guardrail, memory, or planning capability without rebuilding each `Agent`. * **Sub-agent events can be streamed.** Pass an `event_stream_handler` and it’s forwarded to each sub-agent run, so the sub-agent’s model-streaming and tool events surface to the caller (the handler receives the sub-agent’s own `RunContext`). Per-delegate run controls ------------------------- [](https://pydantic.dev/docs/ai/harness/subagents/#per-delegate-run-controls) Each `SubAgent` carries its own budgets, so one delegate’s controls do not touch the others. A `SubAgent` with no controls set runs with the `SubAgents` defaults. from pydantic_ai import Agent from pydantic_ai.usage import UsageLimits from pydantic_ai_harness.subagents import SubAgent, SubAgents reproducer = Agent('anthropic:claude-sonnet-4-6', instructions='Reproduce the reported bug from a minimal script.') librarian = Agent('anthropic:claude-sonnet-4-6', instructions='Find relevant docs, issues, and prior art.') orchestrator = Agent( 'anthropic:claude-opus-4-7', capabilities=[\ SubAgents(\ agents=[\ SubAgent(reproducer, usage_limits=UsageLimits(request_limit=35), timeout_seconds=600, max_calls=1),\ SubAgent(librarian, usage_limits=UsageLimits(request_limit=18), timeout_seconds=300, max_calls=2),\ ]\ )\ ], ) | Field | Effect | | --- | --- | | `models` | Which keys of the `SubAgents` model menu this delegate may run on, and which one it runs on by default: the first key listed. See “Per-delegation model selection” below. | | `usage_limits` | A request/token budget for one delegation. The child runs with its own usage accounting, so the budget counts only that child’s requests and tokens (not the parent’s or siblings’), even when `forward_usage=True`. The tradeoff: that child’s tokens no longer aggregate into the parent’s `usage`. Reaching the budget is a soft outcome (see below), not a run-stopping `UsageLimitExceeded`. | | `timeout_seconds` | A wall-clock budget for one delegation. When the child exceeds it, its run is cancelled and the parent gets a soft steering message instead of hanging on the child. The cancelled child’s `event_stream_handler` (if any) stops receiving events without a terminal event. | | `max_calls` | The maximum number of delegations to this sub-agent per parent run. Once reached, further delegations return a soft budget-exhausted message without running the child. Counts are scoped to one `Agent.run` (a `run_id`) and cleared when it ends, so each parent run and each level of a nested tree budgets independently. | | `on_failure` | A steering message returned to the parent for any soft degradation of this delegate, in place of the built-in default. Setting it also makes child failures soft (see below). | | `contain_errors` | Whether an unexpected crash in this delegate is caught and returned to the parent as a bounded `ModelRetry` instead of aborting the parent run (see below). Unset inherits the `SubAgents(contain_errors=...)` default (off). | Per-delegation model selection ------------------------------ [](https://pydantic.dev/docs/ai/harness/subagents/#per-delegation-model-selection) The orchestrator knows how hard a task is at the moment it writes the brief, so it is the right place to decide which model runs it. Configure a `models` menu and the delegate tool gains a `model` argument that names one of its keys. from pydantic_ai import Agent from pydantic_ai.settings import ModelSettings from pydantic_ai_harness.subagents import ModelOption, SubAgent, SubAgents reviewer = Agent('anthropic:claude-sonnet-4-6', name='reviewer', description='Reviews a diff') linter = Agent('anthropic:claude-sonnet-4-6', name='linter', description='Runs the linter and reports failures') orchestrator = Agent( 'anthropic:claude-opus-4-7', capabilities=[\ SubAgents(\ agents=[SubAgent(reviewer), SubAgent(linter, models=['fast'])],\ models={\ 'fast': 'anthropic:claude-haiku-4-5',\ 'standard': 'anthropic:claude-sonnet-4-6',\ 'deep': ModelOption(\ 'anthropic:claude-opus-4-7',\ description='hard reasoning, multi-file changes',\ settings=ModelSettings(thinking='xhigh'),\ ),\ },\ )\ ], ) * **Off by default.** With no `models` menu the `model` argument is not in the tool schema at all, and every delegation runs exactly as it did before. * **The keys are the interface.** They are listed in the system prompt with each entry’s model and description, and the tool schema offers them as an enum, so the model picks from the menu instead of inventing a model name. Name them for the job (`'fast'`, `'deep'`), not for the vendor. * **An entry is a model or a `ModelOption`.** `ModelOption(model, description=..., settings=...)` adds a routing hint and per-option `ModelSettings`, so one key can mean “same model, more thinking”. Those settings merge over the sub-agent’s own `model_settings`, which keeps the parts the option does not set. * **Resolution order.** The key the parent passed, else the delegate’s first allowed key (see below), else the delegate’s own model, else the parent run’s model. * **A delegate can be restricted.** `SubAgent(linter, models=['fast'])` pins that delegate to `fast`: the first listed key is what it runs on when the parent passes no `model`, and the others are refused with a `ModelRetry`. The restriction is rendered in the prompt listing (`- linter: Runs the linter (models: fast)`). Restricting to a key the menu does not define is a `ValueError` at construction. * **A rejected key costs nothing.** An unknown or unavailable key comes back as a `ModelRetry` listing the valid options, before the delegate’s `max_calls` budget is charged. Failure handling ---------------- [](https://pydantic.dev/docs/ai/harness/subagents/#failure-handling) A _soft outcome_ returns a steering message to the parent as a normal tool result, so its model reads the message and decides what to do next (rather than immediately re-delegating, which a `ModelRetry` invites). A timeout, a reached `usage_limits` budget, and an exhausted `max_calls` budget are always soft. When `on_failure` is set, the message it carries replaces the built-in default for these outcomes. A sub-agent run that fails with a _soft model error_ (`ModelRetry`, `UnexpectedModelBehavior`, e.g. it exhausted its own retries) is, by default, converted into a `ModelRetry` for the parent — so the parent’s model sees `Sub-agent '' failed: ...` and can react by re-delegating. The delegate tool defaults to `tool_retries=2`, so the parent aborts only after that many consecutive delegate failures; the counter resets after any successful delegation. Raise `tool_retries` to tolerate a flakier sub-agent, or set `None` to inherit the parent agent’s default tool retries. Set `on_failure` for a delegate to make its failures soft instead: the child error returns the `on_failure` message as a normal tool result. Hard errors propagate to stop the whole run. A `UsageLimitExceeded` from a child that has _no_ per-delegate `usage_limits` (so it shares the parent’s accounting) means the whole tree is out of budget and propagates; a child reaching its _own_ `usage_limits` is soft, as above. An _unexpected crash_ — any other exception the child raises, such as a provider `ModelAPIError`/`FallbackExceptionGroup` or a plain `ValueError` from a bad tool argument — propagates by default and aborts the parent run. Set `contain_errors=True` (per delegate, or as the `SubAgents` default) to catch it and return it to the parent as a bounded `ModelRetry` instead, so one delegate crash cannot kill the whole run. Containment stays loud: the exception rides the retry message (`Sub-agent '' crashed: ...`), it is logged via the standard `logging` module, and `tool_retries` still bounds consecutive crashes into an abort. This is orthogonal to `on_failure` — a contained crash always raises the loud retry, never the soft `on_failure` return, so a genuine bug is never masked as success. Cancellation, a shared `UsageLimitExceeded`, pydantic-ai control-flow signals (`CallDeferred`, `ApprovalRequired`, the `Skip*` signals), and `UserError` always propagate regardless of `contain_errors`. Discovery --------- [](https://pydantic.dev/docs/ai/harness/subagents/#discovery) The sub-agents are listed in the system prompt via `get_instructions`, using each agent’s `description` (or a `SubAgent(description=...)` override). A sub-agent with no description is listed by name alone. Loading sub-agents from disk ---------------------------- [](https://pydantic.dev/docs/ai/harness/subagents/#loading-sub-agents-from-disk) A repo’s markdown agent definitions become delegates without writing any `Agent` code. By default every `*.md` file under the conventional folders is loaded as a sub-agent, alongside the explicitly-passed `agents`. from pydantic_ai import Agent from pydantic_ai_harness.subagents import SubAgents orchestrator = Agent( 'anthropic:claude-opus-4-7', capabilities=[SubAgents(inherit_tools=True)], # auto-loads ./.agents/agents/ and ~/.agents/agents/ ) `agent_folders` controls where definitions come from. It defaults to `'agents'`, the conventional layout: * A folder-name `str` (the default `'agents'`): for the project root (cwd) then the home root, load from `/.agents//`, falling back to `/.claude//` when `/.agents/` is absent. * A sequence of paths loads from exactly those folders, in order. * `None` disables disk loading, exposing only the explicitly-passed `agents`. ### Definition format [](https://pydantic.dev/docs/ai/harness/subagents/#definition-format) A definition is a markdown file with optional frontmatter: --- name: researcher description: Researches a topic and reports findings tools: Read, Grep --- You research topics. Report your findings, each with a source. * `name` is the delegate name (how the parent refers to it and how it is listed). It falls back to the filename stem when absent. * `description` drives the prompt listing. * The markdown body becomes the agent’s instructions. * `tools` (or `allowed-tools`) is a comma-separated string or a YAML block list. See “Tools” below. * `model` and `color` are ignored: the model is inherited from the parent (see below), and `color` has no pyai equivalent. Frontmatter is read by a small, dependency-free parser limited to those keys (`pyyaml` is not a harness dependency). Full YAML frontmatter is not supported. ### Models and effort [](https://pydantic.dev/docs/ai/harness/subagents/#models-and-effort) Disk agents inherit the parent run’s model by default. Per agent, the caller can override the model and set a thinking/effort level via `agent_overrides`, keyed by the agent’s name: from pydantic_ai_harness.subagents import AgentOverride, SubAgents SubAgents( agent_folders='agents', agent_overrides={'researcher': AgentOverride(model='anthropic:claude-sonnet-4-6', effort='high')}, ) Every agent the capability builds runs at a minimum thinking-effort floor. `MINIMUM_EFFORT_FLOOR` and the `clamp_effort(level, floor=...)` helper are exported so an orchestrator can apply the same floor to its own agents (that orchestrator-side application is the caller’s responsibility). `clamp_effort` maps `None`/`False` to the floor, leaves `True` (provider-default effort) unchanged, and raises a concrete level below the floor up to it. Effort is applied through pyai’s `ModelSettings.thinking`. ### Tools [](https://pydantic.dev/docs/ai/harness/subagents/#tools) A disk agent gets no tools by default (`inherit_tools` is `False`); set `inherit_tools=True` to expose the parent’s tools to it through the `inherit_tools` mechanism, in which case its `tools` frontmatter is ignored. To map the frontmatter tool names to specific toolsets instead, pass a `tool_resolver`: it receives each tool name (so it can honor entries like `Bash(git:*)`) and returns the toolsets that provide it, or `None` for an unknown name, which is skipped with a warning. from pydantic_ai_harness.subagents import SubAgents def resolve(tool_name: str): return TOOLSETS.get(tool_name) # -> Sequence[AgentToolset[object]] | None SubAgents(agent_folders='agents', tool_resolver=resolve) ### Precedence [](https://pydantic.dev/docs/ai/harness/subagents/#precedence) When the same name appears in more than one source, the higher-precedence one wins and the others are skipped with a warning: explicitly-passed `agents` first, then the project folder, then the home folder (and, for an explicit path sequence, earlier paths before later ones). A duplicate name within the explicitly-passed `agents` list is still an error. Configuration ------------- [](https://pydantic.dev/docs/ai/harness/subagents/#configuration) SubAgents( agents=(), # Sequence[SubAgent[AgentDepsT]] -- each pairs an agent with its run controls models={}, # Mapping[str, Model | str | ModelOption] -- per-delegation model menu (off when empty) agent_folders='agents',# folder-name str (convention) | Sequence[Path] | None (disable) agent_overrides={}, # Mapping[str, AgentOverride] -- per-disk-agent model/effort override tool_resolver=None, # Callable[[str], Sequence[AgentToolset[object]] | None] -- disk-agent tool mapping forward_usage=True, # share the parent's usage with sub-agent runs inherit_tools=False, # expose the parent's own tools to sub-agents (capability tools excluded) shared_capabilities=(),# capabilities applied to every sub-agent run event_stream_handler=None, # forwarded to each sub-agent run to stream its events tool_name='delegate_task', tool_retries=2, # extra delegate-tool attempts after a sub-agent error before aborting (None inherits the agent default) contain_errors=False, # default for SubAgent.contain_errors: contain an unexpected crash as a bounded retry ) SubAgent( agent, # AbstractAgent[AgentDepsT, Any] -- the child agent to run name=None, # delegate name; defaults to the agent's own `name` description=None, # prompt-listing description; defaults to the agent's own `description` models=None, # Sequence[str] -- menu keys this delegate may run on; the first is its default usage_limits=None, # per-delegation request/token budget (isolated accounting) timeout_seconds=None, # per-delegation wall-clock budget max_calls=None, # max delegations to this sub-agent per parent run on_failure=None, # steering message for soft degradations of this delegate contain_errors=None, # contain an unexpected crash as a bounded retry; None inherits the SubAgents default ) `SubAgents` is not serializable via the [agent spec](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) (it holds live `Agent` instances), so `get_serialization_name()` returns `None`. Notes ----- [](https://pydantic.dev/docs/ai/harness/subagents/#notes) * Sub-agents can themselves have `SubAgents`, forming a tree. Share `usage` (the default) and set a `usage_limits` on the top-level run to bound the whole tree. * Delegations the model issues in parallel run as independent sub-agent runs. Further reading --------------- [](https://pydantic.dev/docs/ai/harness/subagents/#further-reading) * [Pydantic AI capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) * [Multi-agent applications](https://pydantic.dev/docs/ai/guides/multi-agent-applications/) API reference ------------- [](https://pydantic.dev/docs/ai/harness/subagents/#api-reference) SubAgents --------- [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgents) **Bases:** `AbstractCapability[AgentDepsT]` Let an agent delegate self-contained tasks to named sub-agents. Exposes a single `delegate_task(agent_name, task)` tool. Each delegation runs the chosen sub-agent in a fresh, isolated run (it never sees the parent conversation), and the available sub-agents are listed in the system prompt as a static, cache-stable instruction. Sub-agents are passed as a sequence of `SubAgent` entries, each pairing an agent with its per-delegate run controls (a `usage_limits` budget, a wall-clock `timeout_seconds`, a per-run `max_calls` budget, an `on_failure` steering message, and optional `name`/`description` overrides). A delegate’s name is its `SubAgent.name`, or the agent’s own `name` when unset; two explicitly-passed delegates resolving to the same name is an error. Delegations run on the sub-agent’s own model unless a `models` menu is configured, in which case `delegate_task` also takes a `model` argument naming one of the menu’s keys, so the parent routes each task to the model that fits it. A `SubAgent` can restrict which keys it accepts (`SubAgent.models`). Sub-agents are also loaded from disk by default: each markdown agent definition under `./.agents/agents/` and `~/.agents/agents/` (or the `.claude/` equivalent) becomes a delegate, built with the parent’s model. Disk delegates get no tools by default (`inherit_tools` is `False`); set `inherit_tools=True` to expose the parent’s tools, or pass a `tool_resolver` to map their frontmatter tool names. Disk delegates coexist with explicitly-passed ones; explicitly-passed agents take precedence, then the project folder, then the home folder. A disk delegate whose name is already taken is skipped with a warning. Configure or disable this with `agent_folders`; see also `agent_overrides` and `tool_resolver`. The parent’s `deps` are forwarded to each sub-agent (sub-agents therefore share the parent’s `AgentDepsT`), and by default the parent’s `usage` is shared so usage limits apply across the whole agent tree. Optionally, the parent’s tools can be inherited (`inherit_tools`), extra capabilities can be applied to every sub-agent run (`shared_capabilities`), and sub-agent events can be streamed to a handler (`event_stream_handler`). from pydantic_ai import Agent from pydantic_ai_harness.subagents import SubAgent, SubAgents researcher = Agent('anthropic:claude-sonnet-4-6', name='researcher', description='Researches topics') writer = Agent('anthropic:claude-sonnet-4-6', name='writer', description='Writes prose') orchestrator = Agent( 'anthropic:claude-opus-4-7', capabilities=[SubAgents(agents=[SubAgent(researcher), SubAgent(writer)])], ) ### Attributes [](https://pydantic.dev/docs/ai/harness/subagents/#attributes) #### agent\_folders [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgents.agent_folders) Where to load markdown agent definitions from, in addition to `agents`. Defaults to the conventional layout, so constructing the capability auto-loads a repo’s agent files with no extra configuration. * a folder-name `str` (the default `'agents'` is the conventional layout): for the project root (cwd) then the home root, load from `/.agents//`, falling back to `/.claude//` when `/.agents/` is absent. * a sequence of paths: load from exactly those folders, in order. * `None`: disable disk loading entirely (only `agents` are exposed). Missing folders are skipped. Within a folder every `*.md` file is a candidate. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[`Path`\] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `'agents'` #### agent\_overrides [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgents.agent_overrides) Per-disk-agent overrides keyed by the agent’s name. An entry can set the agent’s `model` (otherwise the parent’s model is inherited) and its `effort` (otherwise the minimum floor). Has no effect on explicitly-passed `agents`. **Type:** [`Mapping`](https://docs.python.org/3/library/typing.html#typing.Mapping) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , `AgentOverride`\] **Default:** `field(default_factory=(dict[str, AgentOverride]))` #### agents [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgents.agents) The sub-agents to expose, each a `SubAgent` pairing an agent with its per-delegate run controls. See `SubAgent`. These take precedence over any disk-loaded agents of the same name. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[`SubAgent`\[`AgentDepsT`\]\] **Default:** `()` #### contain\_errors [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgents.contain_errors) Default for `SubAgent.contain_errors`: whether an unexpected sub-agent crash is caught and returned to the parent as a bounded `ModelRetry` instead of aborting the parent run. Off by default, so a crash propagates. Any `SubAgent` can override this per delegate. See `SubAgent.contain_errors` for the containment contract and what always propagates regardless. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### event\_stream\_handler [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgents.event_stream_handler) If set, this handler is passed to each sub-agent run, so the sub-agent’s model-streaming and tool events surface to the caller. The handler receives the sub-agent’s own `RunContext` and event stream. **Type:** `EventStreamHandler`\[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### forward\_usage [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgents.forward_usage) If `True`, the parent run’s `usage` is shared with each sub-agent run, so token usage aggregates and usage limits apply across the whole agent tree. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `True` #### inherit\_tools [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgents.inherit_tools) If `True`, the parent agent’s tools are exposed to each sub-agent run (the delegate tool itself is filtered out, so sub-agents can’t recurse into further delegation). Off by default to avoid silently widening sub-agent access. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### models [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgents.models) A menu of models the parent can route an individual delegation to, keyed by the name the parent uses to pick one. Off by default: with no menu the delegate tool has no `model` argument and every delegation runs the way it always did. Each value is a model reference, or a `ModelOption` carrying a routing hint and its own `ModelSettings` (so one key can mean “same model, more thinking”). The keys and their descriptions are listed in the system prompt, so name them for the job — `'fast'`, `'deep'` — rather than for the vendor. `SubAgent.models` restricts which of them a given delegate accepts. from pydantic_ai_harness.subagents import SubAgents SubAgents(models={'fast': 'anthropic:claude-haiku-4-5', 'deep': 'anthropic:claude-opus-4-7'}) **Type:** [`Mapping`](https://docs.python.org/3/library/typing.html#typing.Mapping) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , `Model` | `KnownModelName` | [`str`](https://docs.python.org/3/library/stdtypes.html#str)\ | `ModelOption`\] **Default:** `field(default_factory=(dict[str, 'Model | KnownModelName | str | ModelOption']))` #### shared\_capabilities [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgents.shared_capabilities) Capabilities applied to every sub-agent run, in addition to whatever each sub-agent already has. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`AgentCapability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AgentCapability)\ \[`AgentDepsT`\]\] **Default:** `()` #### tool\_name [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgents.tool_name) Name of the delegate tool exposed to the model. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'delegate_task'` #### tool\_resolver [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgents.tool_resolver) Optional override for how a disk agent gets its tools. When set, each tool name in a definition’s `tools`/`allowed-tools` frontmatter is passed to this resolver and the returned toolsets are attached to that agent; an unknown name (resolver returns `None`) is skipped with a warning. When unset, the frontmatter tool list is ignored and disk agents inherit the parent’s tools via `inherit_tools` (set `inherit_tools=True` to expose them). **Type:** `ToolResolver` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### tool\_retries [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgents.tool_retries) Retries for the delegate tool — how many extra attempts it gets after a sub-agent error before the parent run aborts. A sub-agent failure (e.g. it exhausts its own output retries) surfaces to the parent as a tool retry it can react to by re-delegating with a corrected task. The retry counter resets after any successful delegation, so this bounds consecutive failures, not total ones. Defaults to `2` (pydantic-ai’s per-tool default is `1`) so a repeated flaky sub-agent does not abort the parent run on its first repeat; set `None` to inherit the parent agent’s default tool retries instead. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `2` ### Methods [](https://pydantic.dev/docs/ai/harness/subagents/#methods) #### get\_instructions [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgents.get_instructions) def get_instructions() -> AgentInstructions[AgentDepsT] | None Static, cache-stable listing of the available sub-agents and models. ##### Returns [](https://pydantic.dev/docs/ai/harness/subagents/#returns) `AgentInstructions`\[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) #### get\_serialization\_name [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgents.get_serialization_name) `@classmethod` def get_serialization_name(cls) -> str | None Not spec-serializable — the capability holds live `Agent` instances. ##### Returns [](https://pydantic.dev/docs/ai/harness/subagents/#returns-1) [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### get\_toolset [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgents.get_toolset) def get_toolset() -> AgentToolset[AgentDepsT] | None Toolset providing the delegate tool, or `None` when no sub-agents are configured. ##### Returns [](https://pydantic.dev/docs/ai/harness/subagents/#returns-2) [`AgentToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AgentToolset) \[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) #### wrap\_run [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgents.wrap_run) `@async` def wrap_run( ctx: RunContext[AgentDepsT], *, handler: WrapRunHandler, ) -> AgentRunResult[Any] Run the parent agent, then drop this run’s delegation counts so they don’t accumulate. ##### Returns [](https://pydantic.dev/docs/ai/harness/subagents/#returns-3) [`AgentRunResult`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult) \[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] SubAgent -------- [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgent) **Bases:** `Generic[AgentDepsT]` One delegate: a child agent plus its per-delegate run controls. Pass a sequence of these as `SubAgents(agents=[...])`. The delegate’s name — how the parent model refers to it, and how it is listed in the system prompt — is `name` when set, otherwise the agent’s own `name`. An agent with neither is rejected by `SubAgents`. Every control below is optional; an unset field leaves the corresponding behaviour at the `SubAgents` default. ### Attributes [](https://pydantic.dev/docs/ai/harness/subagents/#attributes-1) #### agent [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgent.agent) The agent that runs when this delegate is invoked. **Type:** `AbstractAgent`\[`AgentDepsT`, [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] #### contain\_errors [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgent.contain_errors) Whether an unexpected sub-agent crash is contained instead of aborting the parent run. When `True`, an exception the child raises that is not an expected soft degradation (a provider `ModelAPIError`/`FallbackExceptionGroup`, a plain `ValueError` from a bad tool argument, etc.) is caught and returned to the parent as a bounded `ModelRetry`, so one delegate crash cannot kill the whole run. It stays loud: the exception rides the retry message and is logged, and `tool_retries` still bounds consecutive crashes into an abort. Cancellation, a shared usage-limit, pydantic-ai control-flow signals, and `UserError` always propagate regardless. Unset inherits `SubAgents.contain_errors` (default off). Orthogonal to `on_failure`, which only sets the message for expected soft degradations; a contained crash always raises the loud `ModelRetry`. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### description [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgent.description) Description for the system-prompt listing. Defaults to the agent’s own `description` when unset; a delegate with neither is listed by name alone. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### max\_calls [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgent.max_calls) Maximum number of delegations to this sub-agent per parent run. Once reached, further delegations return a soft budget-exhausted message without running the child. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### models [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgent.models) Which of `SubAgents.models` this delegate may run on, as menu keys, and which one it runs on by default: the first key listed. Leave it unset to let the parent pick any configured option and to fall back to the delegate’s own model when it picks none. Set it to pin a delegate to one option (`models=['fast']`) or to bound an expensive delegate to a subset. Naming a key the menu does not define is an error. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### name [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgent.name) Name the parent model uses to delegate to this agent. Defaults to the agent’s own `name` when unset. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### on\_failure [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgent.on_failure) Steering message returned to the parent for any soft degradation of this delegate (timeout, child failure, usage budget reached, call budget exhausted), in place of the built-in default. Setting it also makes child failures soft: a child error returns this message as a normal tool result instead of raising a parent `ModelRetry`. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### resolved\_name [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgent.resolved_name) The delegate’s name: `name` if set, else the agent’s own `name`. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### timeout\_seconds [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgent.timeout_seconds) Wall-clock budget for one delegation. When the child exceeds it, the run is cancelled and the parent gets a soft steering message instead of hanging on the child. **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### usage\_limits [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.SubAgent.usage_limits) Request/token budget for one delegation. When set, the child runs with its own usage accounting so the budget counts only the child’s own requests and tokens (not the parent’s or siblings’), even when `forward_usage=True`. The tradeoff: that child’s tokens no longer aggregate into the parent’s `usage`. Hitting this budget is a soft outcome (steering message), not a run-stopping `UsageLimitExceeded`. **Type:** [`UsageLimits`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` ModelOption ----------- [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.ModelOption) One entry on the model menu, as a model plus how it should run. Pass bare model references when the menu keys speak for themselves, and a `ModelOption` when an entry needs a routing hint or its own settings: from pydantic_ai.settings import ModelSettings from pydantic_ai_harness.subagents import ModelOption, SubAgents SubAgents( models={ 'fast': 'anthropic:claude-haiku-4-5', 'deep': ModelOption( 'anthropic:claude-opus-4-7', description='hard reasoning, multi-file changes', settings=ModelSettings(thinking='xhigh'), ), }, ) ### Attributes [](https://pydantic.dev/docs/ai/harness/subagents/#attributes-2) #### description [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.ModelOption.description) What this option is for, listed in the prompt next to the key so the parent can route on task difficulty rather than on model names alone. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### model [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.ModelOption.model) The model a delegation routed to this option runs on. **Type:** `Model` | `KnownModelName` | [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### settings [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.ModelOption.settings) Settings for a delegation routed to this option — thinking effort, temperature, and so on. They merge over the sub-agent’s own `model_settings`, which keeps whatever the sub-agent set and is not overridden here. **Type:** [`ModelSettings`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` AgentOverride ------------- [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.AgentOverride) Per-agent override for a disk-loaded sub-agent, keyed by the agent’s name. Both fields are optional. An unset `model` inherits the parent run’s model; an unset `effort` runs at the capability’s minimum effort floor (see `clamp_effort`). ### Attributes [](https://pydantic.dev/docs/ai/harness/subagents/#attributes-3) #### effort [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.AgentOverride.effort) Thinking/reasoning level for this disk agent. Raised to at least the floor. **Type:** `ThinkingLevel` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### model [](https://pydantic.dev/docs/ai/harness/subagents/#pydantic_ai_harness.subagents.AgentOverride.model) Model to run this disk agent with, in place of inheriting the parent’s. **Type:** `Model` | `KnownModelName` | [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` Was this page helpful? Thanks for your feedback! --- # Google Gemini | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/realtime/gemini/#_top) Google Gemini ============= [`GoogleRealtimeModel`](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModel) connects an agent to Gemini Live, including native audio, live images, and provider-native tools. Start with the [realtime quickstart](https://pydantic.dev/docs/ai/realtime/overview/#quickstart) or [camera example](https://pydantic.dev/docs/ai/examples/realtime/realtime-camera/) . Setup ----- [](https://pydantic.dev/docs/ai/realtime/gemini/#setup) To use Gemini Live models, install `pydantic-ai-slim` with the `google-realtime` optional group, which bundles the `google-genai` SDK together with the realtime transport dependencies: * [pip](https://pydantic.dev/docs/ai/realtime/gemini/#tab-panel-172) * [uv](https://pydantic.dev/docs/ai/realtime/gemini/#tab-panel-173) Terminal pip install "pydantic-ai-slim[google-realtime]" Terminal uv add "pydantic-ai-slim[google-realtime]" Authentication comes from `provider`, mirroring [`GoogleModel`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModel) . Use `provider='google'` for the Gemini Developer API or `provider='google-cloud'` for Vertex AI/ADC, with API keys and credentials configured as described in the [Google model documentation](https://pydantic.dev/docs/ai/models/google/#configuration) . Pass a [`GoogleProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.google.GoogleProvider) or [`GoogleCloudProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.google_cloud.GoogleCloudProvider) for custom credentials, project, region, or client. Model names ----------- [](https://pydantic.dev/docs/ai/realtime/gemini/#model-names) Use a Gemini Live model ID, for example `gemini-2.5-flash-native-audio-latest` or `gemini-3.1-flash-live-preview`. Native-audio and other Live models differ in thinking, asynchronous tools, and output behavior. Use the [official Gemini Live documentation](https://ai.google.dev/gemini-api/docs/live) as the canonical model and availability source. Settings -------- [](https://pydantic.dev/docs/ai/realtime/gemini/#settings) [`GoogleRealtimeModelSettings`](https://pydantic.dev/docs/ai/api/realtime/google/#pydantic_ai.realtime.google.GoogleRealtimeModelSettings) — the realtime counterpart of [model run settings](https://pydantic.dev/docs/ai/core-concepts/agent/#model-run-settings) — extends the [shared settings](https://pydantic.dev/docs/ai/realtime/overview/#shared-settings) with Google generation and Live controls: from pydantic_ai.realtime.google import GoogleRealtimeModel, GoogleRealtimeModelSettings settings = GoogleRealtimeModelSettings( temperature=0.7, top_p=0.9, google_voice='Puck', google_language_code='en-US', google_affective_dialog=True, google_proactive_audio=True, google_vad={'start_sensitivity': 'high', 'end_sensitivity': 'low'}, google_turn_coverage='all_video', google_context_compression={'trigger_tokens': 16000, 'target_tokens': 8000}, ) model = GoogleRealtimeModel('gemini-2.5-flash-native-audio-latest', settings=settings) | Setting | Purpose | | --- | --- | | `google_voice`, `google_language_code`, `google_multi_speaker` | Voice, output language, and per-speaker voices | | `google_affective_dialog`, `google_proactive_audio` | Emotion-aware delivery and model-decided speech on native-audio models | | `google_vad` | Exact automatic VAD; fully overrides shared [`turn_detection`](https://pydantic.dev/docs/ai/realtime/turns/#automatic-turn-detection) | | `google_activity_handling`, `google_turn_coverage` | [Interruption](https://pydantic.dev/docs/ai/realtime/turns/#barge-in)
behavior and which input belongs to a turn | | `google_input_transcription`, `google_output_transcription` | Native [transcription](https://pydantic.dev/docs/ai/realtime/audio/#input-transcription)
switches, enabled by default | | `google_context_compression` | Sliding-window compression for long sessions | | `google_enable_session_resumption` | Native state restoration; enabled automatically by a `reconnect` policy | | `google_async_tool_calls` | Lets supported native-audio models continue speaking during tools | | `google_config_overrides` | Raw `LiveConnectConfig` keys merged last as a forward-compatibility escape hatch | `google_voice` is the provider voice setting. `google_thinking_config` takes precedence over the shared [`thinking`](https://pydantic.dev/docs/ai/capabilities/thinking/) setting when a token budget or other Gemini-specific control is needed. ### Asynchronous tool calls [](https://pydantic.dev/docs/ai/realtime/gemini/#asynchronous-tool-calls) Gemini normally pauses generation while a function tool is outstanding. Set `google_async_tool_calls=True` on supported native-audio models to let it continue speaking. This is best for slow tools; a fast result can interrupt speech that barely started and leave an empty interrupted turn in history. Other Live models ignore the setting. ### Native tools [](https://pydantic.dev/docs/ai/realtime/gemini/#native-tools) Gemini Live maps [`WebSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebSearch) to Google Search grounding, the only native tool it supports — no Live model runs native code execution or URL context, so neither [`CodeExecutionTool`](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.CodeExecutionTool) nor [`WebFetch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebFetch) is advertised in `supported_native_tools`. Give those a [`local=` fallback](https://pydantic.dev/docs/ai/realtime/tools/#native-tools) and the session runs the local tool instead: `CodeExecutionTool(local=...)`, or `WebFetch(native=False, local=True)`, which requires the `web-fetch` optional group (`pip/uv-add "pydantic-ai-slim[google-realtime,web-fetch]"`). Gemini 2.5 also cannot combine native Google Search grounding with function tools; choose native grounding or local function-tool fallbacks unless using a model that supports the combination. ### Specialist streaming models [](https://pydantic.dev/docs/ai/realtime/gemini/#specialist-streaming-models) The built-in profile describes the speech-to-speech Live models. Gemini also serves specialist streaming models on the same endpoint that behave differently — `gemini-robotics-er-2-streaming-preview`, for instance, is text-only and rejects audio output. Point a session at one of those and correct the facts with [`profile=`](https://pydantic.dev/docs/ai/realtime/overview/#provider-support) , which resolves like a [standard model profile](https://pydantic.dev/docs/ai/models/overview/#inspecting-a-models-profile) , e.g. `GoogleRealtimeModel('gemini-robotics-er-2-streaming-preview', profile={'supports_text_output': True})`. Feature support and limitations ------------------------------- [](https://pydantic.dev/docs/ai/realtime/gemini/#feature-support-and-limitations) | Feature | Support | Notes | | --- | --- | --- | | Audio format | Full feature support | Mono PCM16, 16 kHz input and 24 kHz output | | Text output | Unsupported | Every speech-to-speech Live model rejects a `TEXT` response modality, so `output_modality='text'` raises. Read the answer from the transcript on the `SpeechPart` | | Image/live video input | Full feature support | [Images](https://pydantic.dev/docs/ai/realtime/audio/#images)
; `google_turn_coverage='all_video'` keeps streamed frames in context | | Manual turns | Unsupported | [Automatic turn detection](https://pydantic.dev/docs/ai/realtime/turns/#automatic-turn-detection)
is required | | Explicit interruption/truncation | Unsupported | Gemini [interrupts server-side](https://pydantic.dev/docs/ai/realtime/turns/#barge-in)
and emits `RealtimeResponseInterruptedEvent` | | Input transcription | Full feature support | Native [transcription](https://pydantic.dev/docs/ai/realtime/audio/#input-transcription)
, enabled by default; no separate model ID | | Native tools | Limited parameter support | Google Search grounding only; URL context and code execution fall back to a [`local=` tool](https://pydantic.dev/docs/ai/realtime/tools/#native-tools)
(see above) | | Usage | Full feature support | Token and modality breakdowns; function-call usage may arrive on a later turn | | State-restoring reconnect | Full feature support | Requires [session resumption](https://pydantic.dev/docs/ai/realtime/gemini/#session-resumption)
plus a reconnect policy | See [Audio, images, and transcripts](https://pydantic.dev/docs/ai/realtime/audio/) , [Turns and interruptions](https://pydantic.dev/docs/ai/realtime/turns/) , [Tools](https://pydantic.dev/docs/ai/realtime/tools/) , and [Connection lifecycle](https://pydantic.dev/docs/ai/realtime/lifecycle/) for the provider-agnostic workflows. Gateway ------- [](https://pydantic.dev/docs/ai/realtime/gemini/#gateway) To route through the [Pydantic AI Gateway](https://pydantic.dev/docs/ai/overview/gateway/) , use a `gateway/`\-prefixed model string: from pydantic_ai import Agent agent = Agent(instructions='You are a helpful voice assistant.') realtime = agent.realtime('gateway/google:gemini-live-2.5-flash') The gateway proxies Gemini Live through the Vertex upstream, so configure a region that supports the Live API. `gateway/google-cloud` is an alias. See [Gateway trace propagation](https://pydantic.dev/docs/ai/realtime/observability/#gateway-trace-propagation) . Session resumption ------------------ [](https://pydantic.dev/docs/ai/realtime/gemini/#session-resumption) For [state-restoring reconnects](https://pydantic.dev/docs/ai/realtime/lifecycle/#state-restoration) , set the `reconnect` setting to a [`ReconnectPolicy`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.ReconnectPolicy) ; session resumption is enabled automatically alongside it (`google_enable_session_resumption` can still request handles without a policy, and explicitly setting it to `False` next to a policy raises [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) rather than silently losing the conversation). Reconnection uses the latest in-memory server handle and emits `state_restored=True`. Provider-specific quirks ------------------------ [](https://pydantic.dev/docs/ai/realtime/gemini/#provider-specific-quirks) * Gemini reports response interruption but not user speech-start/end events, so local playback is flushed on `RealtimeResponseInterruptedEvent`, and Gemini sessions record no `user speech` span (see [Logfire instrumentation](https://pydantic.dev/docs/ai/realtime/observability/#logfire-instrumentation) ). * [Seeded](https://pydantic.dev/docs/ai/realtime/history/#seeding-a-session) function calls/results are represented as readable text because Live cannot accept function parts in seeded turns. * Native transcription can produce only a completed sentence on some models. [Caption UIs](https://pydantic.dev/docs/ai/realtime/audio/#live-captions) should replace text from `TranscriptUpdate.transcript` rather than assume incremental deltas. Was this page helpful? Thanks for your feedback! --- # pydantic_graph.step | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#_top) pydantic\_graph.step ==================== Step-based graph execution components. This module provides the core abstractions for step-based graph execution, including step contexts, step functions, and step nodes that bridge between the declarative `BaseNode` API and the builder graph. NodeStep -------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.NodeStep) **Bases:** `Step[StateT, DepsT, Any, BaseNode[StateT, DepsT, Any] | End[Any]]` A step that wraps a `BaseNode` type for execution by the builder graph. `NodeStep` lets a [`BaseNode`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.BaseNode) subclass participate as a step in the builder graph. It validates that the input is an instance of the expected node type and runs it with the appropriate graph context. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#attributes) #### node\_type [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.NodeStep.node_type) The BaseNode type this step executes. **Type:** [`type`](https://docs.python.org/3/glossary.html#term-type) \[`BaseNode`\[`StateT`, `DepsT`, [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \]\] **Default:** `get_origin(node_type) or node_type` ### Methods [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#methods) #### \_\_init\_\_ [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.NodeStep.__init__) def __init__( node_type: type[BaseNode[StateT, DepsT, Any]], *, id: NodeID | None = None, label: str | None = None, ) Initialize a node step. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#parameters) **`node_type`** : [`type`](https://docs.python.org/3/glossary.html#term-type) \[`BaseNode`\[`StateT`, `DepsT`, [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \]\] [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.NodeStep.__init__(node_type)) The BaseNode class this step will execute **`id`** : `NodeID` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.NodeStep.__init__(id)) Optional unique identifier, defaults to the node’s get\_node\_id() **`label`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.NodeStep.__init__(label)) Optional human-readable label for this step Step ---- [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.Step) **Bases:** `Generic[StateT, DepsT, InputT, OutputT]` A step in the graph execution that wraps a step function. Steps represent individual units of execution in the graph, encapsulating a step function along with metadata like ID and label. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#attributes-1) #### call [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.Step.call) The step function to execute. This needs to be a property for proper variance inference. **Type:** `StepFunction`\[`StateT`, `DepsT`, `InputT`, `OutputT`\] #### id [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.Step.id) Unique identifier for this step. **Type:** `NodeID` **Default:** `id` #### label [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.Step.label) Optional human-readable label for this step. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `label` ### Methods [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#methods-1) #### as\_node [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.Step.as_node) def as_node(inputs: None = None) -> StepNode[StateT, DepsT] def as_node(inputs: InputT) -> StepNode[StateT, DepsT] Create a step node with bound inputs. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#returns) `StepNode`\[`StateT`, `DepsT`\] — A [`StepNode`](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StepNode) with this step and the bound inputs ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#parameters-1) **`inputs`** : `InputT` | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.Step.as_node(inputs)) The input data to bind to this step, or None StepContext ----------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StepContext) **Bases:** `Generic[StateT, DepsT, InputT]` Context information passed to step functions during graph execution. The step context provides access to the current graph state, dependencies, and input data for a step. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#attributes-2) #### inputs [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StepContext.inputs) The input data for this step. This must be a property to ensure correct variance behavior **Type:** `InputT` StepFunction ------------ [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StepFunction) **Bases:** `Protocol[StateT, DepsT, InputT, OutputT]` Protocol for step functions that can be executed in the graph. Step functions are async callables that receive a step context and return a result. ### Methods [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#methods-2) #### \_\_call\_\_ [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StepFunction.__call__) def __call__(ctx: StepContext[StateT, DepsT, InputT]) -> Awaitable[OutputT] Execute the step function with the given context. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#returns-1) [`Awaitable`](https://docs.python.org/3/library/typing.html#typing.Awaitable) \[`OutputT`\] — An awaitable that resolves to the step’s output ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#parameters-2) **`ctx`** : `StepContext`\[`StateT`, `DepsT`, `InputT`\] [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StepFunction.__call__(ctx)) The step context containing state, dependencies, and inputs StepNode -------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StepNode) **Bases:** `BaseNode[StateT, DepsT, Any]` A `BaseNode` that represents a builder step with bound inputs. `StepNode` lets a [`BaseNode`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.BaseNode) subclass hand off to a builder [`Step`](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.Step) by wrapping the step together with the value it should receive as `inputs`. It is not meant to be run directly; returning a `StepNode` from a `BaseNode.run` method tells the graph builder which step to invoke next. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#attributes-3) #### inputs [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StepNode.inputs) The inputs bound to this step. **Type:** [`Any`](https://docs.python.org/3/library/typing.html#typing.Any) #### step [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StepNode.step) The step to execute. **Type:** `Step`\[`StateT`, `DepsT`, [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ , [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] ### Methods [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#methods-3) #### run [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StepNode.run) `@async` def run(ctx: GraphRunContext[StateT, DepsT]) -> BaseNode[StateT, DepsT, Any] | End[Any] Attempt to run the step node. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#returns-2) `BaseNode`\[`StateT`, `DepsT`, [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] | `End`\[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] — The result of step execution ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#parameters-3) **`ctx`** : `GraphRunContext`\[`StateT`, `DepsT`\] [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StepNode.run(ctx)) The graph execution context ##### Raises [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#raises) * `NotImplementedError` — Always raised as StepNode is not meant to be run directly StreamFunction -------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StreamFunction) **Bases:** `Protocol[StateT, DepsT, InputT, OutputT]` Protocol for stream functions that can be executed in the graph. Stream functions are async callables that receive a step context and return an async iterator. ### Methods [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#methods-4) #### \_\_call\_\_ [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StreamFunction.__call__) def __call__(ctx: StepContext[StateT, DepsT, InputT]) -> AsyncIterator[OutputT] Execute the stream function with the given context. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#returns-3) [`AsyncIterator`](https://docs.python.org/3/library/typing.html#typing.AsyncIterator) \[`OutputT`\] — An async iterator yielding the streamed output ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#parameters-4) **`ctx`** : `StepContext`\[`StateT`, `DepsT`, `InputT`\] [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.StreamFunction.__call__(ctx)) The step context containing state, dependencies, and inputs AnyStepFunction --------------- [](https://pydantic.dev/docs/ai/api/pydantic_graph/step/#pydantic_graph.step.AnyStepFunction) Type alias for a step function with any type parameters. **Default:** `StepFunction[Any, Any, Any, Any]` Was this page helpful? Thanks for your feedback! --- # function_signature | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#_top) function\_signature =================== Generate function signatures from functions and JSON schemas. This module provides utilities to represent tool definitions as human-readable function signatures, which LLMs can understand more easily than raw JSON schemas. Used by code mode to present tools as callable functions. FunctionParam ------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionParam) A single parameter in a function signature. ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#methods) #### \_\_str\_\_ [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionParam.__str__) def __str__() -> str Render this parameter as a function parameter string. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#returns) [`str`](https://docs.python.org/3/library/stdtypes.html#str) FunctionSignature ----------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionSignature) Function signature shape with referenced type definitions. This class holds the structural data (params, return type, referenced types) needed to render a function signature. Name and description can be overridden at render time (e.g. from a `ToolDefinition`). ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#attributes) #### is\_async [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionSignature.is_async) Whether the underlying function is async. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### params [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionSignature.params) Function parameters, all rendered as keyword-only (JSON schema doesn’t distinguish positional/keyword). **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , `FunctionParam`\] **Default:** `field(default_factory=(dict[str, FunctionParam]))` #### referenced\_types [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionSignature.referenced_types) TypedDict class definitions needed by the signature. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[`TypeSignature`\] **Default:** `field(default_factory=(list[TypeSignature]))` #### return\_type [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionSignature.return_type) The return type expression. **Type:** `TypeExpr` ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#methods-1) #### collect\_unique\_referenced\_types [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionSignature.collect_unique_referenced_types) `@staticmethod` def collect_unique_referenced_types( signatures: list[FunctionSignature], ) -> list[TypeSignature] Collect unique TypeSignature objects from signatures, deduplicating by identity. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#returns-1) [`list`](https://docs.python.org/3/glossary.html#term-list) \[`TypeSignature`\] #### from\_schema [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionSignature.from_schema) `@classmethod` def from_schema( cls, *, name: str, parameters_schema: dict[str, Any], return_schema: dict[str, Any] | None = None, ) -> FunctionSignature Build a FunctionSignature from JSON schemas. `name` is stored on the resulting signature and also used for generating fallback type names (e.g. `GetUserAddress`) when the schema has no `title`. Parameter and return schemas are processed independently — each resolves `$ref`s against its own `$defs`. Name collisions between parameter and return types (e.g. both define a `User` `$def` with different structures) are handled by `get_conflicting_type_names` at a later stage. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#returns-2) `FunctionSignature` #### get\_conflicting\_type\_names [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionSignature.get_conflicting_type_names) `@staticmethod` def get_conflicting_type_names(signatures: list[FunctionSignature]) -> frozenset[str] Identify TypedDict name conflicts across multiple tool signatures. Each signature keeps all its referenced types (so it remains self-contained), but identical types (same name and structure) are unified to the same object instance. Returns the set of type names that have conflicts (same name, different structure) and need tool-name prefixes at render time. Pass this set to `FunctionSignature.render(conflicting_type_names=...)`. Use `collect_unique_referenced_types()` when rendering to emit each definition once. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#returns-3) [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] #### render [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionSignature.render) def render( body: str, *, name: str | None = None, description: str | None = None, is_async: bool | None = None, conflicting_type_names: frozenset[str] = frozenset(), ) -> str Render the signature with a specific body. Sets `_type_name_overrides` so that dedup-prefixed types resolve correctly during rendering. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#returns-4) [`str`](https://docs.python.org/3/library/stdtypes.html#str) ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#parameters) **`body`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionSignature.render(body)) The function body (e.g. `'...'` or `'return await tool()'`). **`name`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionSignature.render(name)) The function name (also used for dedup prefix resolution). Falls back to `self.name`. **`description`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionSignature.render(description)) Optional docstring to include. Falls back to `self.description`. **`is_async`** : [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionSignature.render(is_async)) Override async rendering. If `None`, uses `self.is_async`. **`conflicting_type_names`** : [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] _Default:_ `frozenset()` [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionSignature.render(conflicting_type_names)) Set of type names that need tool-name prefixes (from `get_conflicting_type_names`). #### render\_type\_definitions [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionSignature.render_type_definitions) `@staticmethod` def render_type_definitions( signatures: list[FunctionSignature], conflicting_type_names: frozenset[str], ) -> list[str] Render unique TypedDict definitions for a set of function signatures. For types whose names conflict across signatures (as identified by `get_conflicting_type_names`), each definition is rendered with a tool-name prefix (e.g. `get_user_Address`). ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#returns-5) [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] — A list of rendered TypedDict class definitions as strings. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#parameters-1) **`signatures`** : [`list`](https://docs.python.org/3/glossary.html#term-list) \[`FunctionSignature`\] [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionSignature.render_type_definitions(signatures)) The function signatures (after `get_conflicting_type_names`). **`conflicting_type_names`** : [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.FunctionSignature.render_type_definitions(conflicting_type_names)) The set returned by `get_conflicting_type_names`. GenericTypeExpr --------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.GenericTypeExpr) A generic type expression like `list[User]`, `dict[str, User]`, `tuple[int, str]`. LiteralTypeExpr --------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.LiteralTypeExpr) A Literal type expression like `Literal['a', 'b']` or `Literal[42]`. SimpleTypeExpr -------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.SimpleTypeExpr) A simple named type like `str`, `int`, `Any`, `None`. TypeFieldSignature ------------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.TypeFieldSignature) A single field in a TypedDict-style type definition. ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#methods-2) #### \_\_str\_\_ [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.TypeFieldSignature.__str__) def __str__() -> str Render this field as a line in a TypedDict class body. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#returns-6) [`str`](https://docs.python.org/3/library/stdtypes.html#str) TypeSignature ------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.TypeSignature) A TypedDict-style class definition with named fields. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#attributes-1) #### display\_name [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.TypeSignature.display_name) The type name, with tool-name prefix applied if rendering context is set. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#methods-3) #### \_\_str\_\_ [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.TypeSignature.__str__) def __str__() -> str Return the type name (for use in type expressions like `def foo(x: User)`). ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#returns-7) [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### render\_definition [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.TypeSignature.render_definition) def render_definition( *, owner_name: str | None = None, conflicting_type_names: frozenset[str] = frozenset(), ) -> str Render the full TypedDict class definition. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#returns-8) [`str`](https://docs.python.org/3/library/stdtypes.html#str) ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#parameters-2) **`owner_name`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.TypeSignature.render_definition(owner_name)) The owning tool name, used to build prefixed type names for conflicting types (e.g. `get_user_Address`). **`conflicting_type_names`** : [`frozenset`](https://docs.python.org/3/library/stdtypes.html#frozenset) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] _Default:_ `frozenset()` [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.TypeSignature.render_definition(conflicting_type_names)) Set of type names that need tool-name prefixes (from `get_conflicting_type_names`). Only effective when `owner_name` is also provided. #### structurally\_equal [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.TypeSignature.structurally_equal) def structurally_equal(other: TypeSignature) -> bool Compare two TypeSignatures structurally, ignoring descriptions. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#returns-9) [`bool`](https://docs.python.org/3/library/functions.html#bool) UnionTypeExpr ------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.UnionTypeExpr) A union type expression like `User | None`, `str | int`. TypeExpr -------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/function_signature/#pydantic_ai.function_signature.TypeExpr) A type expression node in the signature’s type tree. **Type:** [`TypeAlias`](https://docs.python.org/3/library/typing.html#typing.TypeAlias) **Default:** `'TypeSignature | SimpleTypeExpr | LiteralTypeExpr | GenericTypeExpr | UnionTypeExpr'` Was this page helpful? Thanks for your feedback! --- # Code Mode | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/code-mode/#_top) Code Mode ========= `CodeMode` replaces individual tool calls with a single sandboxed Python execution environment. Instead of the model issuing one tool call per action, it writes a Python program that calls your tools as functions — with loops, conditionals, variables, and `asyncio.gather` — all inside a sandboxed [Monty](https://github.com/pydantic/monty) runtime. [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/code_mode/) The problem ----------- [](https://pydantic.dev/docs/ai/harness/code-mode/#the-problem) Standard tool calling often needs another model turn for each dependent batch of tool calls. An agent that needs to fetch 10 items and then process their results can require many model turns, increasing latency, cost, and context use. Intermediate results also grow the conversation history. The solution ------------ [](https://pydantic.dev/docs/ai/harness/code-mode/#the-solution) `CodeMode` wraps eligible tools into a single `run_code` tool. The model writes sandboxed orchestration code that fans calls out with `asyncio.gather`, filters and transforms results, and returns only what matters. Calls from that code are dispatched through Pydantic AI to the host tools. | Standard tool calling | Code mode | | --- | --- | | Dependent tool batches across model turns | Many dependent calls in one `run_code` | | Parallel only when the model emits a batch | Parallelism expressed in Python | | No local computation | Filter, transform, aggregate in code | | Large conversation history | Compact — fewer messages | Durable execution integrations can record nested calls for deterministic replay. Installation ------------ [](https://pydantic.dev/docs/ai/harness/code-mode/#installation) Code mode requires the Monty sandbox, available via the `codemode` extra (the `code-mode` extra is an equivalent alias): Terminal uv add "pydantic-ai-harness[codemode]" Usage ----- [](https://pydantic.dev/docs/ai/harness/code-mode/#usage) Construct an `Agent` with `CodeMode()` in its `capabilities`, then register tools as usual. Every eligible regular tool becomes callable from inside `run_code`: from pydantic_ai import Agent from pydantic_ai_harness import CodeMode agent = Agent('anthropic:claude-sonnet-4-6', capabilities=[CodeMode()]) @agent.tool_plain def get_weather(city: str) -> dict: """Get current weather for a city.""" return {'city': city, 'temp_f': 72, 'condition': 'sunny'} result = agent.run_sync("What's the weather in Paris and Tokyo, in Celsius?") print(result.output) Inside a single `run_code` call, the model writes code like the following (illustrative — the exact code the model emits will vary): import asyncio paris, tokyo = await asyncio.gather( get_weather(city='Paris'), get_weather(city='Tokyo'), ) paris_c = round((paris['temp_f'] - 32) * 5 / 9, 1) tokyo_c = round((tokyo['temp_f'] - 32) * 5 / 9, 1) {'paris': paris_c, 'tokyo': tokyo_c} Both weather lookups run in parallel and the conversions run inside Monty, all within one `run_code` call. Selective tool sandboxing ------------------------- [](https://pydantic.dev/docs/ai/harness/code-mode/#selective-tool-sandboxing) By default, `CodeMode(tools='all')` sandboxes every eligible regular tool. Framework control tools, undiscovered deferred tools, native fallbacks, and other code-execution tools remain native. The `tools` field is a Pydantic AI `ToolSelector`, so you can control which eligible tools go through the sandbox. Tools that match the selector become callables inside `run_code`; non-matching tools stay visible to the model as regular tool calls. from pydantic_ai_harness import CodeMode # By name -- only these tools are available inside run_code CodeMode(tools=['search', 'fetch']) # By predicate -- (ctx, tool_def) -> bool | Awaitable[bool] CodeMode(tools=lambda ctx, td: td.name != 'dangerous_tool') # By metadata -- combine with SetToolMetadata or a toolset's .with_metadata() CodeMode(tools={'code_mode': True}) ### Metadata-based selection [](https://pydantic.dev/docs/ai/harness/code-mode/#metadata-based-selection) Use metadata when the decision should travel with a tool or toolset, rather than with one `CodeMode` instance. This suits shared toolsets: the toolset author tags the tools that are safe and useful to call from generated code, and each agent opts into that tag with `CodeMode(tools={...})`. `CodeMode(tools={'code_mode': True})` uses the standard Pydantic AI [`ToolSelector`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/) metadata form. A tool is sandboxed when its `ToolDefinition.metadata` contains all of the selector’s key-value pairs. Extra metadata on the tool is fine, and nested dictionaries are matched by deep inclusion. The common pattern is to tag an entire toolset with `.with_metadata(...)`: from pydantic_ai import Agent from pydantic_ai.toolsets import FunctionToolset from pydantic_ai_harness import CodeMode def search(query: str) -> str: """Search the web.""" return f'results for {query}' def fetch(url: str) -> str: """Fetch a URL.""" return f'contents of {url}' search_tools = FunctionToolset(tools=[search, fetch]).with_metadata(code_mode=True) agent = Agent( 'anthropic:claude-sonnet-4-6', toolsets=[search_tools], capabilities=[CodeMode(tools={'code_mode': True})], ) Here `search` and `fetch` are removed from the model-facing tool list and become callable functions inside `run_code`. Tools without `metadata['code_mode'] == True` stay visible as regular tool calls. Tool Search interaction ----------------------- [](https://pydantic.dev/docs/ai/harness/code-mode/#tool-search-interaction) When you mark tools or whole toolsets `defer_loading=True` ([Tool Search](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-search) ), `CodeMode` keeps them out of `run_code` while they’re undiscovered — they pass straight through, so Tool Search drives them as usual (sent on the wire with `defer_loading` on providers with native tool search; otherwise dropped until discovered, with a `search_tools` tool alongside `run_code`). `CodeMode` uses `RunContext.is_tool_available` to follow that reveal state. Once the model discovers a tool — or loads the deferred capability that owns it — `CodeMode` folds it into `run_code` like any other tool from then on, so it’s callable from generated code. (The tool keeps `defer_loading=True`, which records what its author asked for; what changes is its availability for the run.) That fold-in grows `run_code`’s description, which invalidates the prompt-cache prefix once at the moment of discovery (turns with no discovery stay cache-warm). Two ways to avoid the bust: * Pass `dynamic_catalog=True` to keep `run_code`’s description static across discoveries. The catalog of sandboxed-tool signatures moves into the agent instructions (as a dynamic [`InstructionPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.InstructionPart) ) and newly-discovered tools are announced via [`ctx.enqueue`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.enqueue) instead of by rebuilding the description: from pydantic_ai_harness import CodeMode CodeMode(dynamic_catalog=True) This pays off when paired with Tool Search: the tool-definitions block stays byte-stable so the prefix cache survives discoveries, at the cost of a larger (but cache-friendly) system prompt. With a fixed toolset and no Tool Search, the default keeps the system prompt shorter and is the better choice. * To instead keep a Tool Search corpus fully native — never folded into `run_code`, but not callable from inside it — exclude it with a `tools` selector; corpus members carry `with_native` set to the managing native tool: from pydantic_ai_harness import CodeMode CodeMode(tools=lambda ctx, td: td.with_native is None) Return values ------------- [](https://pydantic.dev/docs/ai/harness/code-mode/#return-values) The last expression in the snippet is automatically captured as the return value — the model does not need to `print()`. An assignment stores a value in the REPL but does not return it. A final expression that evaluates to `None` is also treated as no result. Without a non-`None` final expression or print output, `run_code` returns `{}`. Put the assigned name on the final line: result = await get_weather(city='Paris') result Reserve `print()` for supplementary logging: printed text is surfaced separately, wrapped alongside the last-expression result. | Scenario | Return | | --- | --- | | Non-`None` final expression with no print output | Last expression value | | Final assignment or `None` result with no print output | `{}` | | Print output with no final expression or a `None` result | `{'output': ''}` | | Print output with a plain, non-`None` final expression | `{'output': '', 'result': }` | | Multimodal final expression with no print output | Returned natively for model processing | | Print output with a multimodal final expression | List with printed text followed by native multimodal content | Printed output is limited to 10 MiB. Exceeding the limit makes `run_code` return a model retry. REPL state ---------- [](https://pydantic.dev/docs/ai/harness/code-mode/#repl-state) State persists between `run_code` calls within the same agent run — variables, imports, and function definitions carry over. Pass `restart: true` in the tool call to reset state. If a worker crash or host-side execution failure invalidates the session, `run_code` returns a model retry that reports the reset; the next snippet must recreate any required state. Temporal durability ------------------- [](https://pydantic.dev/docs/ai/harness/code-mode/#temporal-durability) Install both integrations: Terminal uv add "pydantic-ai-harness[codemode,temporal]" Construct the named agent and its stable-ID toolsets outside the workflow, then attach `TemporalDurability` alongside `CodeMode`: from pydantic_ai import Agent from pydantic_ai.durable_exec.temporal import TemporalDurability from pydantic_ai_harness import CodeMode agent = Agent( 'openai:gpt-5.6-sol', name='coding-agent', capabilities=[CodeMode(), TemporalDurability()], ) Follow the [Pydantic AI Temporal guide](https://pydantic.dev/docs/ai/capabilities/durable_execution/temporal/) to call the plain agent from a workflow and register its activities with `PydanticAIPlugin` and either `__pydantic_ai_agents__` or `AgentPlugin`. `PydanticAIPlugin` passes `pydantic_monty` through Temporal’s workflow sandbox. This makes Monty runnable there, but `run_code` still executes in workflow code and is re-executed during replay. Model requests and, by default, nested tool calls cross Temporal activity boundaries; `asyncio.gather` can schedule nested tool activities concurrently. The REPL is process-local state for one agent run, not durable storage. Replay reconstructs it by running the recorded snippets again against recorded activity results. Keep workflow-side code deterministic. `mount` reads and writes, `os_access` callbacks, and host-clock calls happen again during replay; changing their results can change which activities the workflow schedules and cause a `NondeterminismError`. Put external reads, writes, clock access, and other side effects in wrapped tools so Temporal records them as activities. Replay may not flag changed arguments when the same activity remains at the same history position, so replay validation is not a substitute for this boundary. Temporal activity timeouts apply to nested tools, not pure computation inside `run_code`, so keep sandbox loops bounded. Observability ------------- [](https://pydantic.dev/docs/ai/harness/code-mode/#observability) Nested tool calls inside `run_code` produce their own spans when instrumented with [Logfire](https://pydantic.dev/logfire) or any OpenTelemetry backend — the easiest way to understand what code mode actually did, since each `run_code` span fans out into the tool calls the model issued from inside the sandbox. See the [Pydantic AI Logfire docs](https://pydantic.dev/docs/ai/integrations/logfire/) for setup. The `run_code` tool return also carries metadata with every nested call, keyed by call id: from pydantic_ai import Agent from pydantic_ai.messages import ToolReturnPart from pydantic_ai_harness import CodeMode agent = Agent('anthropic:claude-sonnet-4-6', capabilities=[CodeMode()]) @agent.tool_plain def get_weather(city: str) -> dict: """Get current weather for a city.""" return {'city': city, 'temp_f': 72} result = agent.run_sync("What's the weather in Paris?") for msg in result.all_messages(): for part in msg.parts: if isinstance(part, ToolReturnPart) and part.tool_name == 'run_code': metadata = part.metadata or {} tool_calls = metadata['tool_calls'] # dict[str, ToolCallPart] tool_returns = metadata['tool_returns'] # dict[str, ToolReturnPart] In practice ----------- [](https://pydantic.dev/docs/ai/harness/code-mode/#in-practice) A representative run wires `CodeMode` up against an MCP server and a web search and asks it to find the most-discussed Hacker News story across three feeds, pull the comment thread and the submitter’s profile, and search the web for follow-up coverage. `CodeMode` collapses that into two `run_code` calls: the first fetches all three feeds in parallel via `asyncio.gather`, dedupes by id, filters by score, and ranks by comment count — in plain Python; the second batches the three follow-up calls (`hn_get_thread`, `hn_get_user`, `duckduckgo_search`) together. [![CodeMode's first run_code: parallel asyncio.gather over three HN feeds, then a dedupe and a score filter](https://pydantic.dev/docs/ai/harness/img/code-mode-trace.png)](https://logfire-us.pydantic.dev/public-trace/84bcf123-2106-49da-9f6f-5c26395339bb?spanId=7650806a0785b946) **[See the full Logfire trace ->](https://logfire-us.pydantic.dev/public-trace/84bcf123-2106-49da-9f6f-5c26395339bb?spanId=7650806a0785b946) ** Each `run_code` span fans out into the tool calls the model issued from inside the sandbox. Filesystem and OS access ------------------------ [](https://pydantic.dev/docs/ai/harness/code-mode/#filesystem-and-os-access) Sandboxed code starts with no access to the host’s files, environment, or clock. Two parameters add controlled filesystem, environment, or clock behavior. Both parameters are fixed when the capability is built, so construct `CodeMode` per request to scope the configured access to that request. ### `mount` — share host directories [](https://pydantic.dev/docs/ai/harness/code-mode/#mount--share-host-directories) Reach for `mount` when the agent works with real files: analyzing a dataset you’ve dropped in a folder and writing a report back, editing a checkout, or processing a batch of documents. Sandboxed `pathlib` code reads and writes under the mounted path. (For environment variables or the clock, use `os_access` instead.) from pydantic_ai import Agent from pydantic_monty import MountDir from pydantic_ai_harness import CodeMode # The agent can read /work/data.csv and write /work/summary.md back to the host: agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[CodeMode(mount=MountDir(virtual_path='/work', host_path='/tmp/agent-workspace', mode='read-write'))], ) A `MountDir` defaults to copy-on-write `mode='overlay'`: the sandbox reads host files and sees writes made during the current `run_code` call, but Monty discards those writes before the next call and they do **not** reach the host. Pass `mode='read-write'` when later calls need to read the writes, or `mode='read-only'` to forbid writes. `mount` also accepts a list of `MountDir` for multiple mount points. ### `os_access` — answer the sandbox’s OS calls yourself [](https://pydantic.dev/docs/ai/harness/code-mode/#os_access--answer-the-sandboxs-os-calls-yourself) Reach for `os_access` when the agent needs environment variables, the current date and time, or filesystem behavior you control. Hand it a ready-made OS implementation (`AbstractOS`), or a callback that decides each call — so you can inject just the secrets it needs, pin “now” for reproducible runs, or route file access to your own store. from pydantic_ai import Agent from pydantic_monty import OSAccess from pydantic_ai_harness import CodeMode # Give the agent a fixed set of environment values: agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[CodeMode(os_access=OSAccess(environ={'API_BASE': 'https://api.example.com'}))], ) A callback receives each OS call and decides its fate: from pydantic_ai import Agent from pydantic_monty import NOT_HANDLED from pydantic_ai_harness import CodeMode allowed_env = {'API_KEY': 'sk-...'} def my_os(fn, args, kwargs): if fn == 'os.getenv': # Answer the call: allow-listed keys resolve, every other key reads back # as None -- absent, exactly like a real unset variable. return allowed_env.get(args[0]) # Refuse everything else: NOT_HANDLED makes the call fail in the sandbox. return NOT_HANDLED agent = Agent('anthropic:claude-sonnet-4-6', capabilities=[CodeMode(os_access=my_os)]) Your callback’s return value decides the call’s fate, and the two outcomes are easy to confuse: * **Return any value** — including `None`, `''`, or `0` — and that becomes the result the sandbox sees. `os.getenv` returning `None` looks exactly like a normal unset variable, so the agent’s code keeps running. This is how you _hide_ something: answer with an empty value. * **Return `NOT_HANDLED`** and the call is treated as unsupported: it raises inside the sandbox and the model gets a retry. This _refuses_ a capability outright — use it to block, not to say “no value”. Returning `NOT_HANDLED` for a key the agent reasonably expects will burn retries. Sandbox restrictions -------------------- [](https://pydantic.dev/docs/ai/harness/code-mode/#sandbox-restrictions) Code runs inside [Monty](https://github.com/pydantic/monty) , a sandboxed Python subset. Key restrictions: * No third-party imports. Allowed stdlib modules: `sys`, `typing`, `asyncio`, `math`, `json`, `re`, `unicodedata`, `datetime`, `os`, `pathlib` (each must be imported before use). * `asyncio.gather(...)` accepts positional awaitables but no keyword arguments. Other task creation and wait APIs are unavailable. * No wall-clock or timing primitives by default: `asyncio.sleep`, `datetime.datetime.now()`, `datetime.date.today()`, and the `time` module. `datetime.datetime.now()` / `datetime.date.today()` become available when an `os_access` handler implements them (the built-in `OSAccess` does); `asyncio.sleep` and `time` never do. * No `import *`. * Filesystem I/O needs an `os_access` handler or a `mount`; `os.getenv` / `os.environ` need an `os_access` handler. * Tools requiring approval or with deferred (`CallDeferred`) execution are sandboxed like any other tool; without a `HandleDeferredToolCalls` (or equivalent) capability on the agent to resolve them inline, calling one from `run_code` raises an error that surfaces to the model as a retry. Agent spec (YAML/JSON) ---------------------- [](https://pydantic.dev/docs/ai/harness/code-mode/#agent-spec-yamljson) `CodeMode` works with Pydantic AI’s [agent spec](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) feature for defining agents in YAML or JSON: # agent.yaml model: anthropic:claude-sonnet-4-6 capabilities: - CodeMode: {} from pydantic_ai import Agent from pydantic_ai_harness import CodeMode agent = Agent.from_file('agent.yaml', custom_capability_types=[CodeMode]) result = agent.run_sync('...') print(result.output) Pass `custom_capability_types` so the spec loader knows how to instantiate `CodeMode`. Arguments can be passed in the YAML too: capabilities: - CodeMode: tools: ['search', 'fetch'] max_retries: 5 Further reading --------------- [](https://pydantic.dev/docs/ai/harness/code-mode/#further-reading) * [Tool use via code](https://www.anthropic.com/engineering/code-execution-with-mcp) (Anthropic) * [Code mode in production](https://blog.cloudflare.com/code-mode/) (Cloudflare) * [Pydantic AI capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) API reference ------------- [](https://pydantic.dev/docs/ai/harness/code-mode/#api-reference) CodeMode -------- [](https://pydantic.dev/docs/ai/harness/code-mode/#pydantic_ai_harness.CodeMode) **Bases:** `AbstractCapability[AgentDepsT]` Capability that exposes selected tools as callables inside a `run_code` sandbox. By default (`tools='all'`) every eligible regular tool the agent has is wrapped behind a single `run_code` tool — the model writes Python that calls them as functions instead of issuing tool calls directly. Framework control tools, undiscovered deferred tools, native fallbacks, and other code-execution tools remain native. Pass a list of tool names or a callable predicate to `tools` to split the toolset: matching tools become callables inside the sandbox, and the rest stay visible to the model as normal tool calls. from pydantic_ai import Agent from pydantic_ai_harness import CodeMode # Sandbox all tools agent = Agent('openai:gpt-5', capabilities=[CodeMode()]) # Sandbox only specific tools agent = Agent('openai:gpt-5', capabilities=[CodeMode(tools=['search', 'fetch'])]) By default, sandboxed code cannot touch the host — no filesystem, environment variables, or clock. Two parameters open it up: * `mount` shares specific host directories: reach for it when the agent reads or writes real files. * `os_access` routes the sandbox’s OS calls to a handler you provide: reach for it when the agent needs environment variables, the clock, or filesystem behavior you control. `mount` exposes selected host directories. The built-in `OSAccess` has an isolated filesystem and environment but uses the host clock by default; custom OS handlers can expose other host resources. from pydantic_monty import MountDir agent = Agent('openai:gpt-5', capabilities=[CodeMode(mount=MountDir(virtual_path='/work', host_path='/tmp/agent-work'))]) ### Attributes [](https://pydantic.dev/docs/ai/harness/code-mode/#attributes) #### dynamic\_catalog [](https://pydantic.dev/docs/ai/harness/code-mode/#pydantic_ai_harness.CodeMode.dynamic_catalog) Keep the `run_code` tool definition cache-stable as the sandboxed toolset grows. By default the signatures of all sandboxed tools are rendered into `run_code`’s description, which lives in the prompt-cache-keyed tool-definitions block. When the toolset changes mid-run — e.g. [`ToolSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ToolSearch) reveals a new tool that then gets folded into `run_code` — the description changes and busts the prefix cache from that point on. Set `dynamic_catalog=True` to instead: * keep only the static base prose (sandbox restrictions, return-value contract) in `run_code.description`, so the tool-definitions block stays byte-stable across discoveries; * move the “available functions” catalog (TypedDict definitions + signatures) into agent instructions as a dynamic [`InstructionPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.InstructionPart) , which providers with static/dynamic instruction splitting (Anthropic, Bedrock) place after the cache breakpoint; * announce newly-discovered tools via a short [`SystemPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SystemPromptPart) enqueued through [`RunContext.enqueue`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.enqueue) , so the model knows the new functions are callable without rewriting the cached description. This pays off when paired with [`ToolSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.ToolSearch) : the tool-definitions cache survives discoveries at the cost of a larger (but cache-friendly) system prompt. With a fixed toolset and no `ToolSearch`, the default keeps the system prompt shorter and is the better choice. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### max\_retries [](https://pydantic.dev/docs/ai/harness/code-mode/#pydantic_ai_harness.CodeMode.max_retries) Maximum number of retries for the `run_code` tool (syntax errors count as retries). **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `3` #### mount [](https://pydantic.dev/docs/ai/harness/code-mode/#pydantic_ai_harness.CodeMode.mount) Host directories to expose to sandboxed `pathlib` code; each mount’s `mode` controls whether writes reach the host. **Type:** `CodeModeMount` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### os\_access [](https://pydantic.dev/docs/ai/harness/code-mode/#pydantic_ai_harness.CodeMode.os_access) Give sandboxed code environment variables, the clock, and file I/O through a handler you provide; unset, they are unavailable. **Type:** `CodeModeOS` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### tools [](https://pydantic.dev/docs/ai/harness/code-mode/#pydantic_ai_harness.CodeMode.tools) Which wrapped tools should be sandboxed inside `run_code`. * `'all'` (default): every eligible regular tool the agent has is sandboxed. * `Sequence[str]`: only tools whose names are listed are sandboxed. * Callable `(ctx, tool_def) -> bool | Awaitable[bool]`: tools where the callable returns `True` are sandboxed; the rest stay as native tool calls. **Type:** `ToolSelector`\[`AgentDepsT`\] **Default:** `field(default='all')` ### Methods [](https://pydantic.dev/docs/ai/harness/code-mode/#methods) #### after\_model\_request [](https://pydantic.dev/docs/ai/harness/code-mode/#pydantic_ai_harness.CodeMode.after_model_request) `@async` def after_model_request( ctx: RunContext[AgentDepsT], *, request_context: ModelRequestContext, response: ModelResponse, ) -> ModelResponse Announce newly-discovered tools from a native (server-side) tool-search return. Only active with `dynamic_catalog=True`. ##### Returns [](https://pydantic.dev/docs/ai/harness/code-mode/#returns) [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) #### after\_tool\_execute [](https://pydantic.dev/docs/ai/harness/code-mode/#pydantic_ai_harness.CodeMode.after_tool_execute) `@async` def after_tool_execute( ctx: RunContext[AgentDepsT], *, call: ToolCallPart, tool_def: ToolDefinition, args: ValidatedToolArgs, result: Any, ) -> Any Announce newly-discovered tools from a local `search_tools` return. Only active with `dynamic_catalog=True`. The native-search path is handled by [`after_model_request`](https://pydantic.dev/docs/ai/harness/code-mode/#pydantic_ai_harness.CodeMode.after_model_request) instead (server-side search emits a `NativeToolSearchReturnPart` rather than a regular tool execute result). ##### Returns [](https://pydantic.dev/docs/ai/harness/code-mode/#returns-1) [`Any`](https://docs.python.org/3/library/typing.html#typing.Any) #### for\_run [](https://pydantic.dev/docs/ai/harness/code-mode/#pydantic_ai_harness.CodeMode.for_run) `@async` def for_run(ctx: RunContext[AgentDepsT]) -> CodeMode[AgentDepsT] Return a fresh instance so concurrent runs don’t share `_announced_tools`. ##### Returns [](https://pydantic.dev/docs/ai/harness/code-mode/#returns-2) `CodeMode`\[`AgentDepsT`\] #### get\_ordering [](https://pydantic.dev/docs/ai/harness/code-mode/#pydantic_ai_harness.CodeMode.get_ordering) def get_ordering() -> CapabilityOrdering CodeMode wraps around ToolSearch so that search\_tools stays native. ##### Returns [](https://pydantic.dev/docs/ai/harness/code-mode/#returns-3) `CapabilityOrdering` #### get\_wrapper\_toolset [](https://pydantic.dev/docs/ai/harness/code-mode/#pydantic_ai_harness.CodeMode.get_wrapper_toolset) def get_wrapper_toolset( toolset: AbstractToolset[AgentDepsT], ) -> AbstractToolset[AgentDepsT] | None Wrap the agent’s assembled toolset, splitting it into native + sandboxed subsets if needed. ##### Returns [](https://pydantic.dev/docs/ai/harness/code-mode/#returns-4) [`AbstractToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AbstractToolset) \[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) Was this page helpful? Thanks for your feedback! --- # HTTP Request Retries | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/models/http-request-retries/#_top) HTTP Request Retries ==================== Pydantic AI provides retry functionality for HTTP requests made by model providers through custom HTTP transports. This is particularly useful for handling transient failures like rate limits, network timeouts, or temporary server errors. This is the lowest of the [several layers that can retry](https://pydantic.dev/docs/ai/core-concepts/retries/) in an agent run, and the only one the model never sees. The retry functionality is built on top of the [tenacity](https://github.com/jd/tenacity) library and integrates seamlessly with httpx clients. You can configure retry behavior for any provider that accepts a custom HTTP client. Installation ------------ [](https://pydantic.dev/docs/ai/models/http-request-retries/#installation) To use the retry transports, you need to install `tenacity`, which you can do via the `retries` dependency group: * [pip](https://pydantic.dev/docs/ai/models/http-request-retries/#tab-panel-116) * [uv](https://pydantic.dev/docs/ai/models/http-request-retries/#tab-panel-117) Terminal pip install 'pydantic-ai-slim[retries]' Terminal uv add 'pydantic-ai-slim[retries]' Usage Example ------------- [](https://pydantic.dev/docs/ai/models/http-request-retries/#usage-example) Here’s an example of adding retry functionality with smart retry handling: smart\_retry\_example.py from httpx import AsyncClient, HTTPStatusError from tenacity import retry_if_exception_type, stop_after_attempt, wait_exponential from pydantic_ai import Agent from pydantic_ai.models.openai import OpenAIChatModel from pydantic_ai.providers.openai import OpenAIProvider from pydantic_ai.retries import AsyncTenacityTransport, RetryConfig, wait_retry_after def create_retrying_client(): """Create a client with smart retry handling for multiple error types.""" def should_retry_status(response): """Raise exceptions for retryable HTTP status codes.""" if response.status_code in (429, 502, 503, 504): response.raise_for_status() # This will raise HTTPStatusError transport = AsyncTenacityTransport( config=RetryConfig( # Retry on HTTP errors and connection issues retry=retry_if_exception_type((HTTPStatusError, ConnectionError)), # Smart waiting: respects Retry-After headers, falls back to exponential backoff wait=wait_retry_after( fallback_strategy=wait_exponential(multiplier=1, max=60), max_wait=300 ), # Stop after 5 attempts stop=stop_after_attempt(5), # Re-raise the last exception if all retries fail reraise=True ), validate_response=should_retry_status ) return AsyncClient(transport=transport) # Use the retrying client with a model client = create_retrying_client() model = OpenAIChatModel('gpt-5.2', provider=OpenAIProvider(http_client=client)) agent = Agent(model) Wait Strategies --------------- [](https://pydantic.dev/docs/ai/models/http-request-retries/#wait-strategies) ### wait\_retry\_after [](https://pydantic.dev/docs/ai/models/http-request-retries/#wait_retry_after) The `wait_retry_after` function is a smart wait strategy that automatically respects HTTP `Retry-After` headers: wait\_strategy\_example.py from tenacity import wait_exponential from pydantic_ai.retries import wait_retry_after # Basic usage - respects Retry-After headers, falls back to exponential backoff wait_strategy_1 = wait_retry_after() # Custom configuration wait_strategy_2 = wait_retry_after( fallback_strategy=wait_exponential(multiplier=2, max=120), max_wait=600 # Never wait more than 10 minutes ) This wait strategy: * Automatically parses `Retry-After` headers from HTTP 429 responses * Supports both seconds format (`"30"`) and HTTP date format (`"Wed, 21 Oct 2015 07:28:00 GMT"`) * Falls back to your chosen strategy when no header is present * Respects the `max_wait` limit to prevent excessive delays Transport Classes ----------------- [](https://pydantic.dev/docs/ai/models/http-request-retries/#transport-classes) ### AsyncTenacityTransport [](https://pydantic.dev/docs/ai/models/http-request-retries/#asynctenacitytransport) For asynchronous HTTP clients (recommended for most use cases): async\_transport\_example.py from httpx import AsyncClient from tenacity import stop_after_attempt from pydantic_ai.retries import AsyncTenacityTransport, RetryConfig def validator(response): """Treat responses with HTTP status 4xx/5xx as failures that need to be retried. Without a response validator, only network errors and timeouts will result in a retry. """ response.raise_for_status() # Create the transport transport = AsyncTenacityTransport( config=RetryConfig(stop=stop_after_attempt(3), reraise=True), validate_response=validator ) # Create a client using the transport: client = AsyncClient(transport=transport) ### TenacityTransport [](https://pydantic.dev/docs/ai/models/http-request-retries/#tenacitytransport) For synchronous HTTP clients: sync\_transport\_example.py from httpx import Client from tenacity import stop_after_attempt from pydantic_ai.retries import RetryConfig, TenacityTransport def validator(response): """Treat responses with HTTP status 4xx/5xx as failures that need to be retried. Without a response validator, only network errors and timeouts will result in a retry. """ response.raise_for_status() # Create the transport transport = TenacityTransport( config=RetryConfig(stop=stop_after_attempt(3), reraise=True), validate_response=validator ) # Create a client using the transport client = Client(transport=transport) Common Retry Patterns --------------------- [](https://pydantic.dev/docs/ai/models/http-request-retries/#common-retry-patterns) ### Rate Limit Handling with Retry-After Support [](https://pydantic.dev/docs/ai/models/http-request-retries/#rate-limit-handling-with-retry-after-support) rate\_limit\_handling.py from httpx import AsyncClient, HTTPStatusError from tenacity import retry_if_exception_type, stop_after_attempt, wait_exponential from pydantic_ai.retries import AsyncTenacityTransport, RetryConfig, wait_retry_after def create_rate_limit_client(): """Create a client that respects Retry-After headers from rate limiting responses.""" transport = AsyncTenacityTransport( config=RetryConfig( retry=retry_if_exception_type(HTTPStatusError), wait=wait_retry_after( fallback_strategy=wait_exponential(multiplier=1, max=60), max_wait=300 # Don't wait more than 5 minutes ), stop=stop_after_attempt(10), reraise=True ), validate_response=lambda r: r.raise_for_status() # Raises HTTPStatusError for 4xx/5xx ) return AsyncClient(transport=transport) # Example usage client = create_rate_limit_client() # Client is now ready to use with any HTTP requests and will respect Retry-After headers The `wait_retry_after` function automatically detects `Retry-After` headers in 429 (rate limit) responses and waits for the specified time. If no header is present, it falls back to exponential backoff. ### Network Error Handling [](https://pydantic.dev/docs/ai/models/http-request-retries/#network-error-handling) network\_error\_handling.py import httpx from tenacity import retry_if_exception_type, stop_after_attempt, wait_exponential from pydantic_ai.retries import AsyncTenacityTransport, RetryConfig def create_network_resilient_client(): """Create a client that handles network errors with retries.""" transport = AsyncTenacityTransport( config=RetryConfig( retry=retry_if_exception_type(( httpx.TimeoutException, httpx.ConnectError, httpx.ReadError )), wait=wait_exponential(multiplier=1, max=10), stop=stop_after_attempt(3), reraise=True ) ) return httpx.AsyncClient(transport=transport) # Example usage client = create_network_resilient_client() # Client will now retry on timeout, connection, and read errors ### Custom Retry Logic [](https://pydantic.dev/docs/ai/models/http-request-retries/#custom-retry-logic) custom\_retry\_logic.py import httpx from tenacity import retry_if_exception, stop_after_attempt, wait_exponential from pydantic_ai.retries import AsyncTenacityTransport, RetryConfig, wait_retry_after def create_custom_retry_client(): """Create a client with custom retry logic.""" def custom_retry_condition(exception): """Custom logic to determine if we should retry.""" if isinstance(exception, httpx.HTTPStatusError): # Retry on server errors but not client errors return 500 <= exception.response.status_code < 600 return isinstance(exception, httpx.TimeoutException | httpx.ConnectError) transport = AsyncTenacityTransport( config=RetryConfig( retry=retry_if_exception(custom_retry_condition), # Use wait_retry_after for smart waiting on rate limits, # with custom exponential backoff as fallback wait=wait_retry_after( fallback_strategy=wait_exponential(multiplier=2, max=30), max_wait=120 ), stop=stop_after_attempt(5), reraise=True ), validate_response=lambda r: r.raise_for_status() ) return httpx.AsyncClient(transport=transport) client = create_custom_retry_client() # Client will retry server errors (5xx) and network errors, but not client errors (4xx) Using with Different Providers ------------------------------ [](https://pydantic.dev/docs/ai/models/http-request-retries/#using-with-different-providers) The retry transports work with any provider that accepts a custom HTTP client: ### OpenAI [](https://pydantic.dev/docs/ai/models/http-request-retries/#openai) openai\_with\_retries.py from pydantic_ai import Agent from pydantic_ai.models.openai import OpenAIChatModel from pydantic_ai.providers.openai import OpenAIProvider from smart_retry_example import create_retrying_client client = create_retrying_client() model = OpenAIChatModel('gpt-5.2', provider=OpenAIProvider(http_client=client)) agent = Agent(model) ### Anthropic [](https://pydantic.dev/docs/ai/models/http-request-retries/#anthropic) anthropic\_with\_retries.py from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModel from pydantic_ai.providers.anthropic import AnthropicProvider from smart_retry_example import create_retrying_client client = create_retrying_client() model = AnthropicModel('claude-sonnet-4-5-20250929', provider=AnthropicProvider(http_client=client)) agent = Agent(model) ### Any OpenAI-Compatible Provider [](https://pydantic.dev/docs/ai/models/http-request-retries/#any-openai-compatible-provider) openai\_compatible\_with\_retries.py from pydantic_ai import Agent from pydantic_ai.models.openai import OpenAIChatModel from pydantic_ai.providers.openai import OpenAIProvider from smart_retry_example import create_retrying_client client = create_retrying_client() model = OpenAIChatModel( 'your-model-name', # Replace with actual model name provider=OpenAIProvider( base_url='https://api.example.com/v1', # Replace with actual API URL api_key='your-api-key', # Replace with actual API key http_client=client ) ) agent = Agent(model) Best Practices -------------- [](https://pydantic.dev/docs/ai/models/http-request-retries/#best-practices) 1. **Start Conservative**: Begin with a small number of retries (3-5) and reasonable wait times. 2. **Use Exponential Backoff**: This helps avoid overwhelming servers during outages. 3. **Set Maximum Wait Times**: Prevent indefinite delays with reasonable maximum wait times. 4. **Handle Rate Limits Properly**: Respect `Retry-After` headers when possible. 5. **Log Retry Attempts**: Add logging to monitor retry behavior in production. (This will be picked up by Logfire automatically if you instrument httpx.) 6. **Consider Circuit Breakers**: For high-traffic applications, consider implementing circuit breaker patterns. Error Handling -------------- [](https://pydantic.dev/docs/ai/models/http-request-retries/#error-handling) The retry transports will re-raise the last exception if all retry attempts fail. Make sure to handle these appropriately in your application: error\_handling\_example.py from pydantic_ai import Agent from pydantic_ai.models.openai import OpenAIChatModel from pydantic_ai.providers.openai import OpenAIProvider from smart_retry_example import create_retrying_client client = create_retrying_client() model = OpenAIChatModel('gpt-5.2', provider=OpenAIProvider(http_client=client)) agent = Agent(model) Performance Considerations -------------------------- [](https://pydantic.dev/docs/ai/models/http-request-retries/#performance-considerations) * Retries add latency to requests, especially with exponential backoff * Consider the total timeout for your application when configuring retry behavior * Monitor retry rates to detect systemic issues * Use async transports for better concurrency when handling multiple requests For more advanced retry configurations, refer to the [tenacity documentation](https://tenacity.readthedocs.io/) . Provider-Specific Retry Behavior -------------------------------- [](https://pydantic.dev/docs/ai/models/http-request-retries/#provider-specific-retry-behavior) ### AWS Bedrock [](https://pydantic.dev/docs/ai/models/http-request-retries/#aws-bedrock) The AWS Bedrock provider uses boto3’s built-in retry mechanisms instead of httpx. To configure retries for Bedrock, use boto3’s `Config`: from botocore.config import Config config = Config(retries={'max_attempts': 5, 'mode': 'adaptive'}) See [Bedrock: Configuring Retries](https://pydantic.dev/docs/ai/models/bedrock/#configuring-retries) for complete examples. Was this page helpful? Thanks for your feedback! --- # Debugging & Monitoring with Pydantic Logfire | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/integrations/logfire/#_top) Debugging & Monitoring with Pydantic Logfire ============================================ Applications that use LLMs have some challenges that are well known and understood: LLMs are **slow**, **unreliable** and **expensive**. These applications also have some challenges that most developers have encountered much less often: LLMs are **fickle** and **non-deterministic**. Subtle changes in a prompt can completely change a model’s performance, and there’s no `EXPLAIN` query you can run to understand why. To build successful applications with LLMs, we need new tools to understand both model performance, and the behavior of applications that rely on them. LLM Observability tools that just let you understand how your model is performing are useless: making API calls to an LLM is easy, it’s building that into an application that’s hard. Pydantic Logfire ---------------- [](https://pydantic.dev/docs/ai/integrations/logfire/#pydantic-logfire) [Pydantic Logfire](https://pydantic.dev/logfire) is an observability platform developed by the team who created and maintain Pydantic Validation and Pydantic AI. Logfire aims to let you understand your entire application: Gen AI, classic predictive AI, HTTP traffic, database queries and everything else a modern application needs, all using OpenTelemetry. Pydantic AI has built-in (but optional) support for Logfire. That means if the `logfire` package is installed and configured and agent instrumentation is enabled then detailed information about agent runs is sent to Logfire. Otherwise there’s virtually no overhead and nothing is sent. Here’s an example showing details of running the [Weather Agent](https://pydantic.dev/docs/ai/examples/getting-started/weather-agent/) in Logfire: ![Weather Agent Logfire](https://pydantic.dev/docs/ai/img/logfire-weather-agent.png) A trace is generated for the agent run, and spans are emitted for each model request and tool call. Using Logfire ------------- [](https://pydantic.dev/docs/ai/integrations/logfire/#using-logfire) To use Logfire, you’ll need a Logfire [account](https://logfire.pydantic.dev/) . The Logfire Python SDK is included with `pydantic-ai`: * [pip](https://pydantic.dev/docs/ai/integrations/logfire/#tab-panel-86) * [uv](https://pydantic.dev/docs/ai/integrations/logfire/#tab-panel-87) Terminal pip install pydantic-ai Terminal uv add pydantic-ai Or if you’re using the slim package, you can install it with the `logfire` optional group: * [pip](https://pydantic.dev/docs/ai/integrations/logfire/#tab-panel-88) * [uv](https://pydantic.dev/docs/ai/integrations/logfire/#tab-panel-89) Terminal pip install "pydantic-ai-slim[logfire]" Terminal uv add "pydantic-ai-slim[logfire]" Then authenticate your local environment with Logfire: * [pip](https://pydantic.dev/docs/ai/integrations/logfire/#tab-panel-90) * [uv](https://pydantic.dev/docs/ai/integrations/logfire/#tab-panel-91) Terminal logfire auth Terminal uv run logfire auth And configure a project to send data to: * [pip](https://pydantic.dev/docs/ai/integrations/logfire/#tab-panel-92) * [uv](https://pydantic.dev/docs/ai/integrations/logfire/#tab-panel-93) Terminal logfire projects new Terminal uv run logfire projects new (Or use an existing project with `logfire projects use`) This will write to a `.logfire` directory in the current working directory, which the Logfire SDK will use for configuration at run time. With that, you can start using Logfire to instrument Pydantic AI code: instrument\_pydantic\_ai.py import logfire from pydantic_ai import Agent logfire.configure() # (1) logfire.instrument_pydantic_ai() # (2) agent = Agent('openai:gpt-5.2', name='hello_world_agent', instructions='Be concise, reply with one sentence.') # (4) result = agent.run_sync('Where does "hello world" come from?') # (3) print(result.output) """ The first known use of "hello, world" was in a 1974 textbook about the C programming language. """ [`logfire.configure()`](https://logfire.pydantic.dev/docs/api/logfire/#logfire.configure) configures the SDK, by default it will find the write token from the `.logfire` directory, but you can also pass a token directly. [`logfire.instrument_pydantic_ai()`](https://logfire.pydantic.dev/docs/api/logfire/#logfire.Logfire.instrument_pydantic_ai) enables instrumentation of Pydantic AI. Since we've enabled instrumentation, a trace will be generated for each run, with spans emitted for models calls and tool function execution Passing `name` is optional but recommended: it labels the agent's run span in Logfire. When omitted, the name is inferred from the variable the agent is assigned to and falls back to `'agent'` when it can't be (e.g. agents kept in a list or dict). This matters most when several agents run in one app and you need to tell their traces apart. _(This example is complete, it can be run “as is”)_ Which will display in Logfire thus: ![Logfire Simple Agent Run](https://pydantic.dev/docs/ai/img/logfire-simple-agent.png) The [Logfire documentation](https://logfire.pydantic.dev/docs/) has more details on how to use Logfire, including how to instrument other libraries like [HTTPX](https://logfire.pydantic.dev/docs/integrations/http-clients/httpx/) and [FastAPI](https://logfire.pydantic.dev/docs/integrations/web-frameworks/fastapi/) . Since Logfire is built on [OpenTelemetry](https://opentelemetry.io/) , you can use the Logfire Python SDK to send data to any OpenTelemetry collector, see [below](https://pydantic.dev/docs/ai/integrations/logfire/#using-opentelemetry) . ### Debugging [](https://pydantic.dev/docs/ai/integrations/logfire/#debugging) To demonstrate how Logfire can let you visualise the flow of a Pydantic AI run, here’s the view you get from Logfire while running the [chat app examples](https://pydantic.dev/docs/ai/examples/conversational-agents/chat-app/) : [Realtime (speech-to-speech) sessions](https://pydantic.dev/docs/ai/realtime/observability/) are instrumented by the same `logfire.instrument_pydantic_ai()` call: a session appears as an agent run whose child spans mark each model response, tool call, and turn boundary as the live conversation unfolds. ### Monitoring Performance [](https://pydantic.dev/docs/ai/integrations/logfire/#monitoring-performance) We can also query data with SQL in Logfire to monitor the performance of an application. Here’s a real world example of using Logfire to monitor Pydantic AI runs inside Logfire itself: ![Logfire monitoring Pydantic AI](https://pydantic.dev/docs/ai/img/logfire-monitoring-pydanticai.png) ### Monitoring HTTP Requests [](https://pydantic.dev/docs/ai/integrations/logfire/#monitoring-http-requests) As per Hamel Husain’s influential 2024 blog post [“Fuck You, Show Me The Prompt.”](https://hamel.dev/blog/posts/prompt/) (bear with the capitalization, the point is valid), it’s often useful to be able to view the raw HTTP requests and responses made to model providers. To observe raw HTTP requests made to model providers, you can use Logfire’s [HTTPX instrumentation](https://logfire.pydantic.dev/docs/integrations/http-clients/httpx/) since all provider SDKs (except for [Bedrock](https://pydantic.dev/docs/ai/models/bedrock/) ) use the [HTTPX](https://www.python-httpx.org/) library internally: with\_logfire\_instrument\_httpx.py import logfire from pydantic_ai import Agent logfire.configure() logfire.instrument_pydantic_ai() logfire.instrument_httpx(capture_all=True) # (1) agent = Agent('openai:gpt-5.2') result = agent.run_sync('What is the capital of France?') print(result.output) #> The capital of France is Paris. See the [`logfire.instrument_httpx` docs](https://logfire.pydantic.dev/docs/api/logfire/#logfire.Logfire.instrument_httpx) more details, `capture_all=True` means both headers and body are captured for both the request and response. ![Logfire with HTTPX instrumentation](https://pydantic.dev/docs/ai/img/logfire-with-httpx.png) Using OpenTelemetry ------------------- [](https://pydantic.dev/docs/ai/integrations/logfire/#using-opentelemetry) Pydantic AI’s instrumentation uses [OpenTelemetry](https://opentelemetry.io/) (OTel), which Logfire is based on. This means you can debug and monitor Pydantic AI with any OpenTelemetry backend. Pydantic AI follows the [OpenTelemetry Semantic Conventions for Generative AI systems](https://opentelemetry.io/docs/specs/semconv/gen-ai/) , so while we think you’ll have the best experience using the Logfire platform 😉, you should be able to use any OTel service with GenAI support. ### Logfire with an alternative OTel backend [](https://pydantic.dev/docs/ai/integrations/logfire/#logfire-with-an-alternative-otel-backend) You can use the Logfire SDK completely freely and send the data to any OpenTelemetry backend. Here’s an example of configuring the Logfire library to send data to the excellent [otel-tui](https://github.com/ymtdzzz/otel-tui) — an open source terminal based OTel backend and viewer (no association with Pydantic Validation). Run `otel-tui` with docker (see [the otel-tui readme](https://github.com/ymtdzzz/otel-tui) for more instructions): Terminal docker run --rm -it -p 4318:4318 --name otel-tui ymtdzzz/otel-tui:latest then run, otel\_tui.py import os import logfire from pydantic_ai import Agent os.environ['OTEL_EXPORTER_OTLP_ENDPOINT'] = 'http://localhost:4318' # (1) logfire.configure(send_to_logfire=False) # (2) logfire.instrument_pydantic_ai() logfire.instrument_httpx(capture_all=True) agent = Agent('openai:gpt-5.2') result = agent.run_sync('What is the capital of France?') print(result.output) #> Paris Set the `OTEL_EXPORTER_OTLP_ENDPOINT` environment variable to the URL of your OpenTelemetry backend. If you're using a backend that requires authentication, you may need to set [other environment variables](https://opentelemetry.io/docs/languages/sdk-configuration/otlp-exporter/) . Of course, these can also be set outside the process, e.g. with `export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318`. We [configure](https://logfire.pydantic.dev/docs/api/logfire/#logfire.configure) Logfire to disable sending data to the Logfire OTel backend itself. If you removed `send_to_logfire=False`, data would be sent to both Logfire and your OpenTelemetry backend. Running the above code will send tracing data to `otel-tui`, which will display like this: ![otel tui simple](https://pydantic.dev/docs/ai/img/otel-tui-simple.png) Running the [weather agent](https://pydantic.dev/docs/ai/examples/getting-started/weather-agent/) example connected to `otel-tui` shows how it can be used to visualise a more complex trace: ![otel tui weather agent](https://pydantic.dev/docs/ai/img/otel-tui-weather.png) For more information on using the Logfire SDK to send data to alternative backends, see [the Logfire documentation](https://logfire.pydantic.dev/docs/how-to-guides/alternative-backends/) . ### OTel without Logfire [](https://pydantic.dev/docs/ai/integrations/logfire/#otel-without-logfire) You can also emit OpenTelemetry data from Pydantic AI without using Logfire at all. To do this, you’ll need to install and configure the OpenTelemetry packages you need. To run the following examples, use Terminal uv run \ --with 'pydantic-ai-slim[openai]' \ --with opentelemetry-sdk \ --with opentelemetry-exporter-otlp \ raw_otel.py raw\_otel.py import os from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter from opentelemetry.sdk.trace import TracerProvider from opentelemetry.sdk.trace.export import BatchSpanProcessor from opentelemetry.trace import set_tracer_provider from pydantic_ai import Agent os.environ['OTEL_EXPORTER_OTLP_ENDPOINT'] = 'http://localhost:4318' exporter = OTLPSpanExporter() span_processor = BatchSpanProcessor(exporter) tracer_provider = TracerProvider() tracer_provider.add_span_processor(span_processor) set_tracer_provider(tracer_provider) Agent.instrument_all() agent = Agent('openai:gpt-5.2') result = agent.run_sync('What is the capital of France?') print(result.output) #> Paris ### Alternative Observability backends [](https://pydantic.dev/docs/ai/integrations/logfire/#alternative-observability-backends) Because Pydantic AI uses OpenTelemetry for observability, you can easily configure it to send data to any OpenTelemetry-compatible backend, not just our observability platform [Pydantic Logfire](https://pydantic.dev/docs/ai/integrations/logfire/#pydantic-logfire) . The following providers have dedicated documentation on Pydantic AI: * [Langfuse](https://langfuse.com/docs/integrations/pydantic-ai) * [W&B Weave](https://weave-docs.wandb.ai/guides/integrations/pydantic_ai/) * [Arize](https://arize.com/docs/ax/observe/tracing-integrations-auto/pydantic-ai) * [Openlayer](https://www.openlayer.com/docs/integrations/pydantic-ai) * [LangWatch](https://docs.langwatch.ai/integration/python/integrations/pydantic-ai) * [Patronus AI](https://docs.patronus.ai/docs/percival/integrations/pydantic) * [Opik](https://www.comet.com/docs/opik/tracing/integrations/pydantic-ai) * [mlflow](https://mlflow.org/docs/latest/genai/tracing/integrations/listing/pydantic_ai) * [Agenta](https://docs.agenta.ai/observability/integrations/pydanticai) * [Braintrust](https://www.braintrust.dev/docs/integrations/sdk-integrations/pydantic-ai) * [SigNoz](https://signoz.io/docs/pydantic-ai-observability/) * [Laminar](https://docs.laminar.sh/tracing/integrations/pydantic-ai) * [Respan](https://respan.ai/docs/integrations/pydantic-ai) * [Raindrop](https://raindrop.ai/docs/integrations/pydantic-ai) * [Sentry](https://docs.sentry.io/platforms/python/integrations/pydantic-ai/) Advanced usage -------------- [](https://pydantic.dev/docs/ai/integrations/logfire/#advanced-usage) ### Emitted metrics [](https://pydantic.dev/docs/ai/integrations/logfire/#emitted-metrics) In addition to spans, the instrumentation records the following [OpenTelemetry metrics](https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-metrics/) , all histograms: | Metric | Unit | Description | | --- | --- | --- | | `gen_ai.client.token.usage` | `{token}` | Number of tokens used per model or embedding request, split by the `gen_ai.token.type` attribute (`input` or `output`). Defined by the GenAI semantic conventions. | | `operation.cost` | `{USD}` | Estimated monetary cost of each model or embedding request, recorded when a price is known for the model. | | `gen_ai.client.operation.time_to_first_chunk` | `s` | Time from issuing a streaming request to the first chunk being surfaced to the consumer. Only recorded for streaming requests; the same value is also set as an attribute of the same name on the model request span. | Each metric point carries the `gen_ai.provider.name` (and legacy `gen_ai.system`), `gen_ai.operation.name`, `gen_ai.request.model`, and `gen_ai.response.model` attributes, so histograms can be broken down by provider and model. ### Aggregated usage attribute names [](https://pydantic.dev/docs/ai/integrations/logfire/#aggregated-usage-attribute-names) By default, model request spans use the standard `gen_ai.usage.input_tokens` and `gen_ai.usage.output_tokens` attributes, while agent run spans use `gen_ai.aggregated_usage.input_tokens`, `gen_ai.aggregated_usage.output_tokens`, and `gen_ai.aggregated_usage.details.*`. This avoids double-counting in observability backends that aggregate usage attributes across parent and child spans, since agent run spans report the sum of their child model request spans’ usage. If you want agent run spans to use the standard `gen_ai.usage.*` attributes and handle double-counting in your backend, disable aggregated usage attribute names: from pydantic_ai import Agent from pydantic_ai.models.instrumented import InstrumentationSettings Agent.instrument_all(InstrumentationSettings(use_aggregated_usage_attribute_names=False)) ### Configuring data format [](https://pydantic.dev/docs/ai/integrations/logfire/#configuring-data-format) Pydantic AI follows the [OpenTelemetry Semantic Conventions for Generative AI systems](https://opentelemetry.io/docs/specs/semconv/gen-ai/) , specifically version 1.37.0 of the conventions. The instrumentation format can be configured using the `version` parameter of [`InstrumentationSettings`](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentationSettings) . **The default is `version=5`**. Versions 2, 3, and 4 are deprecated compatibility formats. Passing one of these versions to [`InstrumentationSettings`](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentationSettings) emits a [`PydanticAIDeprecationWarning`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.PydanticAIDeprecationWarning) ; use version 5 unless you are temporarily preserving an older telemetry pipeline. #### Version 2 (deprecated) [](https://pydantic.dev/docs/ai/integrations/logfire/#version-2-deprecated) Uses the newer OpenTelemetry GenAI spec and stores messages in the following attributes: * `gen_ai.system_instructions` for instructions passed to the agent * `gen_ai.input.messages` and `gen_ai.output.messages` on model request spans * `pydantic_ai.all_messages` on agent run spans Some span and attribute names are not fully spec-compliant for compatibility reasons. Use version 5 for current telemetry. #### Version 3 (deprecated) [](https://pydantic.dev/docs/ai/integrations/logfire/#version-3-deprecated) Builds on version 2 with the following improvements: * **Spec-compliant span names:** * `agent run` becomes `invoke_agent {gen_ai.agent.name}` (with the agent name filled in) * `running tool` becomes `execute_tool {gen_ai.tool.name}` (with the tool name filled in) * **Spec-compliant attribute names:** * `tool_arguments` becomes `gen_ai.tool.call.arguments` * `tool_response` becomes `gen_ai.tool.call.result` * **Thinking tokens support:** Captures thinking/reasoning tokens when available #### Version 4 (deprecated) [](https://pydantic.dev/docs/ai/integrations/logfire/#version-4-deprecated) Builds on version 3 with improved multimodal content handling to better align with the [GenAI semantic conventions for multimodal inputs](https://opentelemetry.io/docs/specs/semconv/gen-ai/non-normative/examples-llm-calls/#multimodal-inputs-example) : **URL-based media (ImageUrl, AudioUrl, VideoUrl):** * Old (v2-3): `{"type": "image-url", "url": "..."}` * New (v4): `{"type": "uri", "modality": "image", "uri": "...", "mime_type": "..."}` **Inline binary content (BinaryContent, FilePart):** * Old (v2-3): `{"type": "binary", "media_type": "...", "content": "..."}` * New (v4): `{"type": "blob", "modality": "image", "mime_type": "...", "content": "..."}` Note: The `modality` field is only included for image, audio, and video content types as specified in the OTel spec. DocumentUrl and unsupported media types omit the `modality` field. #### Version 5 [](https://pydantic.dev/docs/ai/integrations/logfire/#version-5) Builds on version 4 with improved handling of deferred tool calls: * [`CallDeferred`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.CallDeferred) and [`ApprovalRequired`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ApprovalRequired) exceptions no longer record an exception event or set the span status to ERROR — the span is left as UNSET, since deferrals are control flow, not errors. * * * Note that the OpenTelemetry Semantic Conventions are still experimental and are likely to change. ### Setting OpenTelemetry SDK providers [](https://pydantic.dev/docs/ai/integrations/logfire/#setting-opentelemetry-sdk-providers) By default, the global `TracerProvider` is used. This is set automatically by `logfire.configure()`. It can also be set by the `set_tracer_provider` function in the OpenTelemetry Python SDK. You can set custom providers with [`InstrumentationSettings`](https://pydantic.dev/docs/ai/api/models/instrumented/#pydantic_ai.models.instrumented.InstrumentationSettings) . instrumentation\_settings\_providers.py from opentelemetry.sdk.trace import TracerProvider from pydantic_ai import Agent, InstrumentationSettings from pydantic_ai.capabilities import Instrumentation instrumentation_settings = InstrumentationSettings( tracer_provider=TracerProvider(), ) agent = Agent('openai:gpt-5.2', capabilities=[Instrumentation(settings=instrumentation_settings)]) # or to instrument all agents: Agent.instrument_all(instrumentation_settings) ### Excluding binary content [](https://pydantic.dev/docs/ai/integrations/logfire/#excluding-binary-content) When `include_binary_content=False` is set, binary file data (images, audio, documents) is excluded from telemetry: from user prompts and model responses, from tool returns, from the agent’s own output and the arguments its output function receives, and from run and tool deferral metadata. The media type is still recorded everywhere; where the value is recorded as the file itself rather than as a message part, so are its vendor metadata and its identifier, which is derived from the content when you don’t set one. Binary content is found inside dictionaries, lists and [`ToolReturn`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolReturn) s, but not inside your own types: a [`BinaryContent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BinaryContent) held as a field of a model or dataclass you define is still recorded in full. excluding\_binary\_content.py from pydantic_ai import Agent, InstrumentationSettings from pydantic_ai.capabilities import Instrumentation instrumentation_settings = InstrumentationSettings(include_binary_content=False) agent = Agent('openai:gpt-5.2', capabilities=[Instrumentation(settings=instrumentation_settings)]) # or to instrument all agents: Agent.instrument_all(instrumentation_settings) ### Excluding prompts and completions [](https://pydantic.dev/docs/ai/integrations/logfire/#excluding-prompts-and-completions) For privacy and security reasons, you may want to monitor your agent’s behavior and performance without exposing sensitive user data or proprietary prompts in your observability platform. Pydantic AI allows you to exclude the actual content from telemetry while preserving the structural information needed for debugging and monitoring. When `include_content=False` is set, Pydantic AI will exclude sensitive content from telemetry, including user prompts and model completions, tool call arguments and responses, and any other message content. excluding\_sensitive\_content.py from pydantic_ai import Agent from pydantic_ai.capabilities import Instrumentation from pydantic_ai.models.instrumented import InstrumentationSettings instrumentation_settings = InstrumentationSettings(include_content=False) agent = Agent('openai:gpt-5.2', capabilities=[Instrumentation(settings=instrumentation_settings)]) # or to instrument all agents: Agent.instrument_all(instrumentation_settings) This setting is particularly useful in production environments where compliance requirements or data sensitivity concerns make it necessary to limit what content is sent to your observability platform. ### Excluding model request parameters [](https://pydantic.dev/docs/ai/integrations/logfire/#excluding-model-request-parameters) By default, each model request span carries a `model_request_parameters` attribute that serializes the full [`ModelRequestParameters`](https://pydantic.dev/docs/ai/api/models/base/#pydantic_ai.models.ModelRequestParameters) , including the output configuration and every tool definition. Tools that carry large output schemas (some MCP toolsets, for example) can make this attribute big enough to strain span export and inflate memory use. Set `include_model_request_parameters=False` to omit it entirely: excluding\_model\_request\_parameters.py from pydantic_ai import Agent from pydantic_ai.capabilities import Instrumentation from pydantic_ai.models.instrumented import InstrumentationSettings instrumentation_settings = InstrumentationSettings(include_model_request_parameters=False) agent = Agent('openai:gpt-5.2', capabilities=[Instrumentation(settings=instrumentation_settings)]) # or to instrument all agents: Agent.instrument_all(instrumentation_settings) The `gen_ai.tool.definitions` attribute (tool name, description, and parameters) is emitted regardless of this setting, so observability platforms that read the available tools from it are unaffected. ### Adding Custom Metadata [](https://pydantic.dev/docs/ai/integrations/logfire/#adding-custom-metadata) Use the agent’s `metadata` parameter to attach additional data to the agent’s span. When instrumentation is enabled, the computed metadata is recorded on the agent span under the `metadata` attribute. See the [usage and metadata example in the agents guide](https://pydantic.dev/docs/ai/core-concepts/agent/#run-metadata) for details and usage. Was this page helpful? Thanks for your feedback! --- # Google | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/models/google/#_top) Google ====== The `GoogleModel` is a model that uses the [`google-genai`](https://pypi.org/project/google-genai/) package under the hood to access Google’s Gemini models via both the Gemini API and Google Cloud (formerly known as Vertex AI). Two providers wrap those endpoints: * [`GoogleProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.google.GoogleProvider) — the Gemini API (Google AI Studio), surfaced under the `'google:'` prefix. * [`GoogleCloudProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.google_cloud.GoogleCloudProvider) — Google Cloud (formerly known as Vertex AI), surfaced under the `'google-cloud:'` prefix. Install ------- [](https://pydantic.dev/docs/ai/models/google/#install) To use `GoogleModel`, you need to either install `pydantic-ai`, or install `pydantic-ai-slim` with the `google` optional group: * [pip](https://pydantic.dev/docs/ai/models/google/#tab-panel-112) * [uv](https://pydantic.dev/docs/ai/models/google/#tab-panel-113) Terminal pip install "pydantic-ai-slim[google]" Terminal uv add "pydantic-ai-slim[google]" Configuration ------------- [](https://pydantic.dev/docs/ai/models/google/#configuration) `GoogleModel` lets you use Google’s Gemini models through their [Gemini API](https://ai.google.dev/api/all-methods) (`generativelanguage.googleapis.com`) or [Google Cloud](https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models) (`*-aiplatform.googleapis.com`, formerly known as Vertex AI). ### API Key (Gemini API) [](https://pydantic.dev/docs/ai/models/google/#api-key-gemini-api) To use Gemini via the Gemini API, go to [aistudio.google.com](https://aistudio.google.com/apikey) and create an API key. Once you have the API key, set it as an environment variable: Terminal export GOOGLE_API_KEY=your-api-key You can then use `GoogleModel` by name: from pydantic_ai import Agent agent = Agent('google:gemini-3-pro-preview') ... Or you can explicitly create the provider: from pydantic_ai import Agent from pydantic_ai.models.google import GoogleModel from pydantic_ai.providers.google import GoogleProvider provider = GoogleProvider(api_key='your-api-key') model = GoogleModel('gemini-3-pro-preview', provider=provider) agent = Agent(model) ... ### Google Cloud (Enterprise) [](https://pydantic.dev/docs/ai/models/google/#google-cloud-enterprise) If you are an enterprise user, you can also use `GoogleModel` to access Gemini via Google Cloud (formerly known as Vertex AI). This interface has a number of advantages over the Gemini API: 1. The Google Cloud API comes with more enterprise readiness guarantees. 2. You can [purchase provisioned throughput](https://cloud.google.com/vertex-ai/generative-ai/docs/provisioned-throughput#purchase-provisioned-throughput) with Google Cloud to guarantee capacity. 3. If you’re running Pydantic AI inside Google Cloud, you don’t need to set up authentication, it should “just work”. 4. You can decide which region to use, which might be important from a regulatory perspective, and might improve latency. You can authenticate using [application default credentials](https://cloud.google.com/docs/authentication/application-default-credentials) , a service account, or an [API key](https://cloud.google.com/vertex-ai/generative-ai/docs/start/api-keys?usertype=expressmode) . Whichever way you authenticate, you’ll need to have the Vertex AI API (now branded as Google Cloud AI) enabled in your Google Cloud account. #### Application Default Credentials [](https://pydantic.dev/docs/ai/models/google/#application-default-credentials) If you have the [`gcloud` CLI](https://cloud.google.com/sdk/gcloud) installed and configured, you can use the `GoogleCloudProvider` by name: from pydantic_ai import Agent agent = Agent('google-cloud:gemini-3-pro-preview') ... Or you can explicitly create the provider and model: from pydantic_ai import Agent from pydantic_ai.models.google import GoogleModel from pydantic_ai.providers.google_cloud import GoogleCloudProvider provider = GoogleCloudProvider() model = GoogleModel('gemini-3-pro-preview', provider=provider) agent = Agent(model) ... #### Service Account [](https://pydantic.dev/docs/ai/models/google/#service-account) To use a service account JSON file, explicitly create the provider and model: google\_model\_service\_account.py from google.oauth2 import service_account from pydantic_ai import Agent from pydantic_ai.models.google import GoogleModel from pydantic_ai.providers.google_cloud import GoogleCloudProvider credentials = service_account.Credentials.from_service_account_file('path/to/service-account.json') provider = GoogleCloudProvider(credentials=credentials, project='your-project-id') model = GoogleModel('gemini-3-flash-preview', provider=provider) agent = Agent(model) ... #### API Key [](https://pydantic.dev/docs/ai/models/google/#api-key) To use Google Cloud with an API key, [create a key](https://cloud.google.com/vertex-ai/generative-ai/docs/start/api-keys?usertype=expressmode) and set it as an environment variable: Terminal export GOOGLE_API_KEY=your-api-key You can then use `GoogleModel` via [`GoogleCloudProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.google_cloud.GoogleCloudProvider) by name: from pydantic_ai import Agent agent = Agent('google-cloud:gemini-3-pro-preview') ... Or you can explicitly create the provider and model: from pydantic_ai import Agent from pydantic_ai.models.google import GoogleModel from pydantic_ai.providers.google_cloud import GoogleCloudProvider provider = GoogleCloudProvider(api_key='your-api-key') model = GoogleModel('gemini-3-pro-preview', provider=provider) agent = Agent(model) ... #### Customizing Location or Project [](https://pydantic.dev/docs/ai/models/google/#customizing-location-or-project) You can specify the location and/or project when using Google Cloud: google\_model\_location.py from pydantic_ai import Agent from pydantic_ai.models.google import GoogleModel from pydantic_ai.providers.google_cloud import GoogleCloudProvider provider = GoogleCloudProvider(location='asia-east1', project='your-google-cloud-project-id') model = GoogleModel('gemini-3-pro-preview', provider=provider) agent = Agent(model) ... In addition to the single-region values listed in [`GoogleCloudLocation`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.google.GoogleCloudLocation) , `GoogleCloudProvider` accepts the `'global'` location and the `'us'`/`'eu'` multi-regions. The multi-region values are routed to the `aiplatform.{us,eu}.rep.googleapis.com` data-residency endpoints — use them when an org policy blocks the global endpoint for data residency, or when a model is initially available only on `global` and the multi-regions rather than a single region. Model availability differs between single regions, multi-regions, and `global`; see the [Vertex AI locations docs](https://cloud.google.com/vertex-ai/generative-ai/docs/learn/locations#available-regions) . google\_model\_multi\_region.py from pydantic_ai import Agent from pydantic_ai.models.google import GoogleModel from pydantic_ai.providers.google_cloud import GoogleCloudProvider provider = GoogleCloudProvider(location='us', project='your-google-cloud-project-id') model = GoogleModel('gemini-3-pro-preview', provider=provider) agent = Agent(model) ... #### Service tier (`service_tier`, `google_cloud_service_tier`) [](https://pydantic.dev/docs/ai/models/google/#service-tier-service_tier-google_cloud_service_tier) The unified [`service_tier`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.service_tier) field works on both Google subsystems, with [`google_cloud_service_tier`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings.google_cloud_service_tier) available for finer Google Cloud routing control. The provider-specific field wins when both are set. **Gemini API** — sent as the request’s `service_tier` field: | `service_tier` | Sent to Gemini API | | --- | --- | | `'auto'` | _(omitted — server default)_ | | `'default'` | `'standard'` | | `'flex'` | `'flex'` | | `'priority'` | `'priority'` | **Google Cloud** — sent as HTTP routing headers; `'flex'` and `'priority'` always pick the **PT-with-spillover** variant, so customers with [Provisioned Throughput](https://cloud.google.com/vertex-ai/generative-ai/docs/provisioned-throughput/use-provisioned-throughput) (PT) keep using their reserved capacity first: | `service_tier` | Google Cloud routing headers | Effective behavior | | --- | --- | --- | | `'auto'` / `'default'` | _(none)_ | PT first, then standard on-demand spillover | | `'flex'` | `X-Vertex-AI-LLM-Shared-Request-Type: flex` | PT first, then [Flex PayGo](https://cloud.google.com/vertex-ai/generative-ai/docs/flex-paygo)
spillover | | `'priority'` | `X-Vertex-AI-LLM-Shared-Request-Type: priority` | PT first, then [Priority PayGo](https://cloud.google.com/vertex-ai/generative-ai/docs/priority-paygo)
spillover | To bypass PT entirely (or use it exclusively, or any of the other Google Cloud-specific routing combinations) set [`google_cloud_service_tier`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings.google_cloud_service_tier) directly — the unified field is intentionally limited to the safe PT-with-spillover variants. **Google Cloud — full set of routing values** The full [`google_cloud_service_tier`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings.google_cloud_service_tier) values map to these HTTP headers: * `'pt_only'`: PT only (`X-Vertex-AI-LLM-Request-Type: dedicated`). * `'pt_then_flex'`: PT when quota allows, then [Flex PayGo](https://cloud.google.com/vertex-ai/generative-ai/docs/flex-paygo) spillover (`X-Vertex-AI-LLM-Shared-Request-Type: flex`). * `'pt_then_priority'`: PT when quota allows, then [Priority PayGo](https://cloud.google.com/vertex-ai/generative-ai/docs/priority-paygo) spillover (`X-Vertex-AI-LLM-Shared-Request-Type: priority`). * `'on_demand'`: Standard on-demand only (`X-Vertex-AI-LLM-Request-Type: shared`). * `'flex_only'`: [Flex PayGo](https://cloud.google.com/vertex-ai/generative-ai/docs/flex-paygo) only (`X-Vertex-AI-LLM-Request-Type: shared` and `X-Vertex-AI-LLM-Shared-Request-Type: flex`). * `'priority_only'`: [Priority PayGo](https://cloud.google.com/vertex-ai/generative-ai/docs/priority-paygo) only (`X-Vertex-AI-LLM-Request-Type: shared` and `X-Vertex-AI-LLM-Shared-Request-Type: priority`). **Example** from pydantic_ai import Agent from pydantic_ai.models.google import GoogleModel, GoogleModelSettings from pydantic_ai.providers.google_cloud import GoogleCloudProvider provider = GoogleCloudProvider(location='global') model = GoogleModel('gemini-3-flash-preview', provider=provider) agent = Agent(model) result = agent.run_sync( 'Hello!', model_settings=GoogleModelSettings(google_cloud_service_tier='pt_then_flex'), ) Swap `'pt_then_flex'` for any [`GoogleCloudServiceTier`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleCloudServiceTier) value — e.g. `'pt_then_priority'` for [Priority PayGo](https://cloud.google.com/vertex-ai/generative-ai/docs/priority-paygo) spillover, or `'flex_only'` / `'priority_only'` to bypass PT entirely. After the request, inspect [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) `provider_details.get('traffic_type')` (e.g. `ON_DEMAND_FLEX`, `ON_DEMAND_PRIORITY`) to see which tier served it, when the API returns it. #### Model Garden [](https://pydantic.dev/docs/ai/models/google/#model-garden) You can access models from the [Model Garden](https://cloud.google.com/model-garden?hl=en) that support the `generateContent` API and are available under your Google Cloud project, including but not limited to Gemini, using one of the following `model_name` patterns: * `{model_id}` for Gemini models * `{publisher}/{model_id}` * `publishers/{publisher}/models/{model_id}` * `projects/{project}/locations/{location}/publishers/{publisher}/models/{model_id}` from pydantic_ai import Agent from pydantic_ai.models.google import GoogleModel from pydantic_ai.providers.google_cloud import GoogleCloudProvider provider = GoogleCloudProvider( project='your-google-cloud-project-id', location='us-central1', # the region where the model is available ) model = GoogleModel('meta/llama-3.3-70b-instruct-maas', provider=provider) agent = Agent(model) ... Custom HTTP Client ------------------ [](https://pydantic.dev/docs/ai/models/google/#custom-http-client) You can customize the `GoogleProvider` with a custom `httpx.AsyncClient`: from httpx import AsyncClient from pydantic_ai import Agent from pydantic_ai.models.google import GoogleModel from pydantic_ai.providers.google import GoogleProvider custom_http_client = AsyncClient(timeout=30) model = GoogleModel( 'gemini-3-pro-preview', provider=GoogleProvider(api_key='your-api-key', http_client=custom_http_client), ) agent = Agent(model) ... HTTP Retries ------------ [](https://pydantic.dev/docs/ai/models/google/#http-retries) By default, the `google-genai` SDK does not retry requests that fail with a transient HTTP error. You can enable retries by passing a [`HttpRetryOptions`](https://googleapis.github.io/python-genai/genai.html#genai.types.HttpRetryOptions) instance to the `retry_options` argument of `GoogleProvider` or `GoogleCloudProvider`: from google.genai.types import HttpRetryOptions from pydantic_ai import Agent from pydantic_ai.models.google import GoogleModel from pydantic_ai.providers.google import GoogleProvider retry_options = HttpRetryOptions( attempts=4, initial_delay=1.0, max_delay=60.0, http_status_codes=[408, 429, 500, 502, 503, 504], ) model = GoogleModel( 'gemini-3-pro-preview', provider=GoogleProvider(api_key='your-api-key', retry_options=retry_options), ) agent = Agent(model) ... This passes the options through to the SDK’s [`HttpOptions.retry_options`](https://googleapis.github.io/python-genai/genai.html#genai.types.HttpOptions.retry_options) . See the [Vertex AI retry strategy documentation](https://cloud.google.com/vertex-ai/generative-ai/docs/retry-strategy) for guidance on choosing values. Document, Image, Audio, and Video Input --------------------------------------- [](https://pydantic.dev/docs/ai/models/google/#document-image-audio-and-video-input) `GoogleModel` supports multi-modal input, including documents, images, audio, and video. YouTube video URLs can be passed directly to Google models: youtube\_input.py from pydantic_ai import Agent, VideoUrl from pydantic_ai.models.google import GoogleModel agent = Agent(GoogleModel('gemini-3-flash-preview')) result = agent.run_sync( [\ 'What is this video about?',\ VideoUrl(url='https://www.youtube.com/watch?v=dQw4w9WgXcQ'),\ ] ) print(result.output) Files can be uploaded via the [Files API](https://ai.google.dev/gemini-api/docs/files) and passed as URLs: file\_upload.py from pydantic_ai import Agent, DocumentUrl from pydantic_ai.models.google import GoogleModel from pydantic_ai.providers.google import GoogleProvider provider = GoogleProvider() file = provider.client.files.upload(file='pydantic-ai-logo.png') assert file.uri is not None agent = Agent(GoogleModel('gemini-3-flash-preview', provider=provider)) result = agent.run_sync( [\ 'What company is this logo from?',\ DocumentUrl(url=file.uri, media_type=file.mime_type),\ ] ) print(result.output) See the [input documentation](https://pydantic.dev/docs/ai/core-concepts/input/) for more details and examples. Model settings -------------- [](https://pydantic.dev/docs/ai/models/google/#model-settings) You can customize model behavior using [`GoogleModelSettings`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings) : from google.genai.types import HarmBlockThreshold, HarmCategory from pydantic_ai import Agent from pydantic_ai.models.google import GoogleModel, GoogleModelSettings settings = GoogleModelSettings( temperature=0.2, max_tokens=1024, top_k=40, google_safety_settings=[\ {\ 'category': HarmCategory.HARM_CATEGORY_HATE_SPEECH,\ 'threshold': HarmBlockThreshold.BLOCK_LOW_AND_ABOVE,\ }\ ] ) model = GoogleModel('gemini-3-pro-preview') agent = Agent(model, model_settings=settings) ... ### Configure thinking [](https://pydantic.dev/docs/ai/models/google/#configure-thinking) Use the provider-agnostic [`Thinking`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Thinking) capability to enable thinking: from pydantic_ai import Agent from pydantic_ai.capabilities import Thinking agent = Agent('google:gemini-3.5-flash', capabilities=[Thinking(effort='medium')]) ... For advanced usage, you can pass Google’s native thinking config through [`GoogleModelSettings.google_thinking_config`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings.google_thinking_config) : from pydantic_ai import Agent from pydantic_ai.models.google import GoogleModel, GoogleModelSettings model = GoogleModel('gemini-3.5-flash') model_settings = GoogleModelSettings(google_thinking_config={'include_thoughts': True, 'thinking_level': 'MEDIUM'}) agent = Agent(model, model_settings=model_settings) ... See [Thinking](https://pydantic.dev/docs/ai/capabilities/thinking/) for the unified API and [Gemini API docs](https://ai.google.dev/gemini-api/docs/thinking) for Google’s native thinking configuration. ### Safety settings [](https://pydantic.dev/docs/ai/models/google/#safety-settings) You can customize the safety settings by setting the `google_safety_settings` field. from google.genai.types import HarmBlockThreshold, HarmCategory from pydantic_ai import Agent from pydantic_ai.models.google import GoogleModel, GoogleModelSettings model_settings = GoogleModelSettings( google_safety_settings=[\ {\ 'category': HarmCategory.HARM_CATEGORY_HATE_SPEECH,\ 'threshold': HarmBlockThreshold.BLOCK_LOW_AND_ABOVE,\ }\ ] ) model = GoogleModel('gemini-3-flash-preview') agent = Agent(model, model_settings=model_settings) ... See the [Gemini API docs](https://ai.google.dev/gemini-api/docs/safety-settings) for more on safety settings. ### Logprobs [](https://pydantic.dev/docs/ai/models/google/#logprobs) You can return logprobs from the model in your response by setting `google_logprobs` and `google_top_logprobs` in the [`GoogleModelSettings`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings) . This feature is only supported for non-streaming requests and Google Cloud. from pydantic_ai import Agent from pydantic_ai.models.google import GoogleModel, GoogleModelSettings from pydantic_ai.providers.google_cloud import GoogleCloudProvider model_settings = GoogleModelSettings( google_logprobs=True, google_top_logprobs=2, ) model = GoogleModel( model_name='gemini-2.5-flash', provider=GoogleCloudProvider(location='europe-west1'), ) agent = Agent(model, model_settings=model_settings) result = agent.run_sync('Your prompt here') # Access logprobs from provider_details logprobs = result.response.provider_details.get('logprobs') avg_logprobs = result.response.provider_details.get('avg_logprobs') See the [Google Dev Blog](https://developers.googleblog.com/unlock-gemini-reasoning-with-logprobs-on-vertex-ai/) for more information. ### Model Armor (Google Cloud only) [](https://pydantic.dev/docs/ai/models/google/#model-armor-google-cloud-only) [Model Armor](https://docs.cloud.google.com/model-armor/overview) is a Google Cloud security service that screens prompts and responses for risks like prompt injection, jailbreaking, and sensitive data leakage. You can configure it via `google_model_armor_config` in [`GoogleModelSettings`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings) : from pydantic_ai import Agent from pydantic_ai.models.google import GoogleModel, GoogleModelSettings from pydantic_ai.providers.google_cloud import GoogleCloudProvider model_settings = GoogleModelSettings( google_model_armor_config={ 'prompt_template_name': 'projects/my-project/locations/europe-west4/templates/prompt-template', 'response_template_name': 'projects/my-project/locations/europe-west4/templates/response-template', } ) model = GoogleModel( model_name='gemini-2.5-flash', provider=GoogleCloudProvider(location='europe-west4'), ) agent = Agent(model, model_settings=model_settings) ... Templates must be created in advance in the [Google Cloud Console](https://console.cloud.google.com/security/modelarmor) and must reside in the same region as the model endpoint. See the [Model Armor Vertex AI integration docs](https://docs.cloud.google.com/model-armor/model-armor-vertex-integration) for supported locations. When a prompt or response is blocked, a [`ContentFilterError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ContentFilterError) is raised. Note that Model Armor screening — both prompt and response templates — only works with non-streaming requests (`agent.run()`). With streaming (`agent.run_stream()`), Google Cloud does not apply Model Armor: the prompt is not screened and the response text is returned unscreened. If you require streaming and need Model Armor protection, pre-screen prompts using the [`google-cloud-modelarmor` SDK](https://pypi.org/project/google-cloud-modelarmor/) before calling the agent. ### Context caching (`google_cached_content`) [](https://pydantic.dev/docs/ai/models/google/#context-caching-google_cached_content) When you’ve created a Gemini [cached content resource](https://ai.google.dev/gemini-api/docs/caching) , pass its resource name through [`google_cached_content`](https://pydantic.dev/docs/ai/api/models/google/#pydantic_ai.models.google.GoogleModelSettings.google_cached_content) to reuse it across requests: from pydantic_ai import Agent from pydantic_ai.models.google import GoogleModel, GoogleModelSettings model_settings = GoogleModelSettings( google_cached_content='projects/p/locations/global/cachedContents/your-cache-id', ) agent = Agent(GoogleModel('gemini-2.5-pro'), model_settings=model_settings) ... Create a cached content resource Pydantic AI doesn’t wrap the cache-management API — create the resource with the underlying [google-genai](https://googleapis.github.io/python-genai/) SDK, then pass its name through `google_cached_content`: from google.genai.types import Content, CreateCachedContentConfig, Part from pydantic_ai.providers.google import GoogleProvider provider = GoogleProvider(api_key='your-api-key') cache = provider.client.caches.create( model='gemini-2.5-flash', config=CreateCachedContentConfig( system_instruction='You are a geography expert. Be concise.', contents=[Content(role='user', parts=[Part(text='...long context to cache...')])], ttl='3600s', ), ) print(cache.name) #> cachedContents/abc123... Caches have a minimum size (≈1024 tokens for `gemini-2.5-flash`, ≈4096 for `gemini-2.5-pro`) and a TTL — see the [Gemini caching docs](https://ai.google.dev/gemini-api/docs/caching) for the current thresholds, pricing, and `list` / `update` / `delete` operations. Streaming cancellation ---------------------- [](https://pydantic.dev/docs/ai/models/google/#streaming-cancellation) Was this page helpful? Thanks for your feedback! --- # usage | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#_top) usage ===== RequestUsage ------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage) **Bases:** `UsageBase` LLM usage associated with a single request. This is an implementation of `genai_prices.types.AbstractUsage` so it can be used to calculate the price of the request using [genai-prices](https://github.com/pydantic/genai-prices) . ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#methods) #### \_\_add\_\_ [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage.__add__) def __add__(other: RequestUsage) -> RequestUsage Add two RequestUsages together. This is provided so it’s trivial to sum usage information from multiple parts of a response. **WARNING:** this CANNOT be used to sum multiple requests without breaking some pricing calculations. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#returns) [`RequestUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage) #### extract [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage.extract) `@classmethod` def extract( cls, data: Any, *, provider: str, provider_url: str, provider_fallback: str, api_flavor: str = 'default', details: dict[str, Any] | None = None, ) -> RequestUsage Extract usage information from the response data using genai-prices. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#returns-1) [`RequestUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage) ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#parameters) **`data`** : [`Any`](https://docs.python.org/3/library/typing.html#typing.Any) [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage.extract(data)) The response data from the model API. **`provider`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage.extract(provider)) The actual provider ID **`provider_url`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage.extract(provider_url)) The provider base\_url **`provider_fallback`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage.extract(provider_fallback)) The fallback provider ID to use if the actual provider is not found in genai-prices. For example, an OpenAI model should set this to “openai” in case it has an obscure provider ID. **`api_flavor`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) _Default:_ `'default'` [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage.extract(api_flavor)) The API flavor to use when extracting usage information, e.g. ‘chat’ or ‘responses’ for OpenAI. **`details`** : [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage.extract(details)) Becomes the `details` field on the returned `RequestUsage` for convenience. #### incr [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage.incr) def incr(incr_usage: RequestUsage) -> None Increment the usage in place. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#returns-2) [`None`](https://docs.python.org/3/library/constants.html#None) ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#parameters-1) **`incr_usage`** : [`RequestUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage) [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage.incr(incr_usage)) The usage to increment by. RunUsage -------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage) **Bases:** `UsageBase` LLM usage associated with an agent run. Responsibility for calculating request usage is on the model; Pydantic AI simply sums the usage information across requests. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#attributes) #### cache\_audio\_read\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage.cache_audio_read_tokens) Total number of audio tokens read from the cache. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `0` #### cache\_read\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage.cache_read_tokens) Total number of tokens read from the cache. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `0` #### cache\_write\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage.cache_write_tokens) Total number of tokens written to the cache. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `0` #### details [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage.details) Any extra details returned by the model. **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`int`](https://docs.python.org/3/library/functions.html#int)\ \] **Default:** `dataclasses.field(default_factory=(dict[str, int]))` #### input\_audio\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage.input_audio_tokens) Total number of audio input tokens. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `0` #### input\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage.input_tokens) Total number of input/prompt tokens. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `0` #### output\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage.output_tokens) Total number of output/completion tokens. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `0` #### requests [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage.requests) Number of requests made to the LLM API. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `0` #### tool\_calls [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage.tool_calls) Number of successful tool calls executed during the run. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `0` ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#methods-1) #### \_\_add\_\_ [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage.__add__) def __add__(other: RunUsage | RequestUsage) -> RunUsage Add two RunUsages together. This is provided so it’s trivial to sum usage information from multiple runs. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#returns-3) [`RunUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage) #### incr [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage.incr) def incr(incr_usage: RunUsage | RequestUsage) -> None Increment the usage in place. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#returns-4) [`None`](https://docs.python.org/3/library/constants.html#None) ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#parameters-2) **`incr_usage`** : [`RunUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage) | [`RequestUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage) [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage.incr(incr_usage)) The usage to increment by. UsageBase --------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageBase) ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#attributes-1) #### cache\_audio\_read\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageBase.cache_audio_read_tokens) Number of audio tokens read from the cache. Included in `cache_read_tokens` and `input_audio_tokens`. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `0` #### cache\_hit\_ratio [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageBase.cache_hit_ratio) Fraction of input tokens that were read from the provider’s prompt cache. Computed as `cache_read_tokens / input_tokens`. Both counts span all modalities — cached audio tokens are included in `cache_read_tokens` just as audio input tokens are included in `input_tokens` — and `input_tokens` includes cached reads for every provider, so the ratio is comparable across providers: `0.0` means no prompt-cache hits, while values approaching `1.0` mean nearly the entire prompt was served from cache. Returns `0.0` when there are no input tokens. On [`RequestUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage) this is the hit ratio of a single request; on [`RunUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage) it aggregates all requests in the run. **Type:** [`float`](https://docs.python.org/3/library/functions.html#float) #### cache\_read\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageBase.cache_read_tokens) Number of tokens read from the cache, across all modalities (includes `cache_audio_read_tokens`). Included in `input_tokens`. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `0` #### cache\_write\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageBase.cache_write_tokens) Number of tokens written to the cache. Included in `input_tokens`. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `0` #### cost [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageBase.cost) Best-effort cost in USD, or `None` if no cost could be determined. Calculated with [genai-prices](https://github.com/pydantic/genai-prices) . `None` (rather than zero) when the model or provider can’t be priced, so “unknown” stays distinguishable from a genuine zero cost. **Type:** `Decimal` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### details [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageBase.details) Any extra details returned by the model. **Type:** [`Annotated`](https://docs.python.org/3/library/typing.html#typing.Annotated) \[[`dict`](https://docs.python.org/3/reference/expressions.html#dict)\ \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`int`](https://docs.python.org/3/library/functions.html#int)\ \], `BeforeValidator`([`lambda`](https://docs.python.org/3/glossary.html#term-lambda)\ `d`: `d` [`or`](https://docs.python.org/3/reference/expressions.html#or)\ {})\] **Default:** `details or {}` #### input\_audio\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageBase.input_audio_tokens) Number of audio input tokens. Included in `input_tokens`. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `0` #### input\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageBase.input_tokens) Total number of input/prompt tokens, across all modalities. Token counts form inclusive parent/child buckets, not disjoint ones: this total includes cached tokens (`cache_read_tokens`, `cache_write_tokens`) and audio tokens (`input_audio_tokens`). Usage extraction normalizes providers that report these separately (e.g. Anthropic and Bedrock, whose raw `input_tokens` exclude cache reads/writes) so the convention holds everywhere. **Type:** [`Annotated`](https://docs.python.org/3/library/typing.html#typing.Annotated) \[[`int`](https://docs.python.org/3/library/functions.html#int)\ , `Field`(`validation_alias`\=(`AliasChoices`(`input_tokens`, `request_tokens`)))\] **Default:** `0` #### output\_audio\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageBase.output_audio_tokens) Number of audio output tokens. Included in `output_tokens`. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `0` #### output\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageBase.output_tokens) Number of output/completion tokens. **Type:** [`Annotated`](https://docs.python.org/3/library/typing.html#typing.Annotated) \[[`int`](https://docs.python.org/3/library/functions.html#int)\ , `Field`(`validation_alias`\=(`AliasChoices`(`output_tokens`, `response_tokens`)))\] **Default:** `0` #### total\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageBase.total_tokens) Sum of `input_tokens + output_tokens`. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#methods-2) #### \_\_copy\_\_ [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageBase.__copy__) def __copy__() -> UsageBase Shallow copy that also copies mutable fields like `details`. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#returns-5) `UsageBase` #### \_\_get\_pydantic\_core\_schema\_\_ [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageBase.__get_pydantic_core_schema__) `@classmethod` def __get_pydantic_core_schema__( cls, source_type: Any, handler: GetCoreSchemaHandler, ) -> core_schema.CoreSchema Preserve arbitrary usage fields across Pydantic serialization. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#returns-6) `core_schema.CoreSchema` #### has\_values [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageBase.has_values) def has_values() -> bool Whether any values are set and non-zero. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#returns-7) [`bool`](https://docs.python.org/3/library/functions.html#bool) #### opentelemetry\_attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageBase.opentelemetry_attributes) def opentelemetry_attributes() -> dict[str, int] Get the token usage values as OpenTelemetry attributes. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#returns-8) [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`int`](https://docs.python.org/3/library/functions.html#int)\ \] UsageLimits ----------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits) Limits on model usage. The request count is tracked by pydantic\_ai, and the request limit is checked before each request to the model. Token counts are provided in responses from the model, and the token limits are checked after each response. Each of the limits can be set to `None` to disable that limit. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#attributes-2) #### cost\_limit [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits.cost_limit) The maximum cost allowed in USD. **Type:** `Decimal` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### count\_tokens\_before\_request [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits.count_tokens_before_request) If True, perform a token counting pass before sending the request to the model, to enforce `input_tokens_limit` and `per_request_input_tokens_limit` ahead of time. This may incur additional overhead (from calling the model’s `count_tokens` API before making the actual request) and is disabled by default. Supported by: * Anthropic * Google * Bedrock Converse * OpenAI Responses **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### input\_tokens\_limit [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits.input_tokens_limit) The maximum number of input/prompt tokens allowed. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### output\_tokens\_limit [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits.output_tokens_limit) The maximum number of output/response tokens allowed. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### per\_request\_input\_tokens\_limit [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits.per_request_input_tokens_limit) The maximum number of input/prompt tokens allowed per individual request. Unlike `input_tokens_limit` which is cumulative across the entire run, this limit is checked against each request’s input token count independently — ahead of the request when `count_tokens_before_request=True`, otherwise against the provider-reported `input_tokens` of the response. This provides a guard against oversized contexts (which hurt model performance and incur high costs on cache misses), complementing the runaway-loop protection that cumulative limits provide. Note that `input_tokens` (and therefore this limit) includes cached-prefix tokens, normalized consistently across providers: a request served largely from cache still counts its full context size toward this limit. This caps context size, not cache-miss cost. Set `count_tokens_before_request=True` to enforce this preemptively; otherwise the request is sent before the limit is checked, so the oversized request is still billed (matching `input_tokens_limit`). **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### request\_limit [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits.request_limit) The maximum number of requests allowed to the model. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `50` #### tool\_calls\_limit [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits.tool_calls_limit) The maximum number of successful tool calls allowed to be executed. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### total\_tokens\_limit [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits.total_tokens_limit) The maximum number of tokens allowed in requests and responses combined. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#methods-3) #### check\_before\_request [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits.check_before_request) def check_before_request(usage: RunUsage) -> None Raises a `UsageLimitExceeded` exception if the next request would exceed any of the limits. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#returns-9) [`None`](https://docs.python.org/3/library/constants.html#None) #### check\_before\_tool\_call [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits.check_before_tool_call) def check_before_tool_call(projected_usage: RunUsage) -> None Raises a `UsageLimitExceeded` exception if the next tool call(s) would exceed the tool call limit. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#returns-10) [`None`](https://docs.python.org/3/library/constants.html#None) #### check\_cost [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits.check_cost) def check_cost(usage: RunUsage, *, warn_if_cost_unavailable: bool = True) -> None Check whether usage exceeds the cost limit. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#returns-11) [`None`](https://docs.python.org/3/library/constants.html#None) ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#parameters-3) **`usage`** : [`RunUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage) [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits.check_cost(usage)) The accumulated run usage to check. **`warn_if_cost_unavailable`** : [`bool`](https://docs.python.org/3/library/functions.html#bool) _Default:_ `True` [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits.check_cost(warn_if_cost_unavailable)) Whether to warn when a `cost_limit` is set but no cost was calculated. #### check\_per\_request\_input\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits.check_per_request_input_tokens) def check_per_request_input_tokens(request_input_tokens: int) -> None Raises a `UsageLimitExceeded` if the per-request input tokens exceed the limit. This checks a single request’s input token count — not the cumulative `RunUsage.input_tokens` — against `per_request_input_tokens_limit`. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#returns-12) [`None`](https://docs.python.org/3/library/constants.html#None) #### check\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits.check_tokens) def check_tokens(usage: RunUsage) -> None Raises a `UsageLimitExceeded` exception if the usage exceeds any of the token limits. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#returns-13) [`None`](https://docs.python.org/3/library/constants.html#None) #### has\_token\_limits [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits.has_token_limits) def has_token_limits() -> bool Returns `True` if this instance places any limits on token counts. If this returns `False`, the `check_tokens` and `check_per_request_input_tokens` methods will never raise an error. This is useful because if we have token limits, we need to check them after receiving each streamed message. If there are no limits, we can skip that processing in the streaming response iterator. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#returns-14) [`bool`](https://docs.python.org/3/library/functions.html#bool) Was this page helpful? Thanks for your feedback! --- # codec | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/realtime/codec/#_top) codec ===== The lower-level _codec_ vocabulary, for implementing a realtime provider or consuming a [`RealtimeConnection`](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeConnection) directly: the raw events a connection yields, the turn-control verbs and inputs it accepts, and the model-profile merge helpers. Most users only need the session-level API in [`pydantic_ai.realtime`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/) . Low-level _codec_ vocabulary for realtime providers. Most users only need the session-level API in [`pydantic_ai.realtime`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime) ([`AgentRealtime.session`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentRealtime.session) , the events a session yields, and the content passed to [`RealtimeSession.send`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.send) ). This submodule holds the lower-level vocabulary used when _implementing_ a realtime provider or consuming a [`RealtimeConnection`](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeConnection) directly: the raw codec events a connection yields to the session, the turn-control verbs and inputs a connection accepts, and the model-profile merge helpers. AudioDelta ---------- [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.AudioDelta) A chunk of audio output from the model. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/codec/#attributes) #### data [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.AudioDelta.data) Raw PCM audio bytes. The sample rate is provider-specific. **Type:** [`bytes`](https://docs.python.org/3/library/stdtypes.html#bytes) #### item\_id [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.AudioDelta.item_id) Provider item ID for the spoken output this chunk belongs to, when available. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` CancelResponse -------------- [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.CancelResponse) Cancel the model’s in-progress response (maps to the provider’s response-cancel). ClearAudio ---------- [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ClearAudio) Discard any buffered, uncommitted input audio. CommitAudio ----------- [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.CommitAudio) Commit the buffered input audio as a user turn (manual turn-taking / push-to-talk). Only needed when automatic voice activity detection is disabled; with server-side VAD the provider commits audio and triggers a response automatically. ConversationCreated ------------------- [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ConversationCreated) An OpenAI-protocol server assigned a conversation ID. This is a codec-level control event. Providers consume it during their handshake when possible; the session silently consumes any instance that reaches the live stream. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/codec/#attributes-1) #### conversation\_id [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ConversationCreated.conversation_id) Provider-assigned conversation ID. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) ConversationItemCreated ----------------------- [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ConversationItemCreated) An OpenAI-protocol server reported a conversation item. xAI uses `replayed=True` for item events emitted during the resume handshake. The session consumes those items and remembers their newly assigned IDs so any follow-on content or tool events aren’t appended or executed again. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/codec/#attributes-2) #### item\_id [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ConversationItemCreated.item_id) Provider-assigned conversation-item ID, when present. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### replayed [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ConversationItemCreated.replayed) Whether the provider identified this item as part of a resumption replay. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### tool\_call\_id [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ConversationItemCreated.tool_call_id) Provider-assigned tool-call ID, for function call and result items. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` CreateResponse -------------- [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.CreateResponse) Ask the model to generate a response now (manual turn-taking, after `CommitAudio`). InputTranscript --------------- [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.InputTranscript) A transcription of the user’s audio input (partial or final). Providers with per-item IDs use `item_id` to associate interleaved transcripts with the correct user turn. Providers without them retain arrival-order association. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/codec/#attributes-3) #### cumulative [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.InputTranscript.cumulative) Whether `text` is the whole transcript so far rather than an incremental piece. Speech recognition is revisable, and a provider that streams cumulative snapshots may correct what it already transcribed instead of only extending it. Setting this lets the session adopt each snapshot as authoritative rather than guessing from prefixes whether the text appends; it surfaces the difference to callers as a [`SpeechPartDelta.transcript`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SpeechPartDelta.transcript) carrying the corrected whole. Leave `False` for incremental deltas. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### is\_final [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.InputTranscript.is_final) Whether this is the final transcript for the user’s turn. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### item\_id [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.InputTranscript.item_id) Provider item ID for the user’s turn, when available. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### text [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.InputTranscript.text) Transcript text. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) OutputTranscript ---------------- [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.OutputTranscript) The model’s textual output (partial or final): an audio transcript, or plain text output. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/codec/#attributes-4) #### is\_final [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.OutputTranscript.is_final) Whether this is the final transcript for the turn. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### item\_id [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.OutputTranscript.item_id) Provider item ID for the spoken output, when available. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### output\_text [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.OutputTranscript.output_text) Whether this is the model’s plain text output (`output_modalities=('text',)`) rather than a transcription of spoken audio. Text output becomes a [`TextPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.TextPart) ; an audio transcript becomes a [`SpeechPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SpeechPart) . **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### text [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.OutputTranscript.text) Transcript text. A partial event carries the incremental delta; a final event the full turn. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) RealtimeConnection ------------------ [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeConnection) **Bases:** `ABC` A live connection to a realtime model. Providers implement this to handle protocol-specific framing (WebSocket frames, HTTP/2 messages, etc.). Content is fed in via [`send`](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeConnection.send) and events are consumed by iterating the connection. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/codec/#attributes-5) #### input\_transcription\_enabled [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeConnection.input_transcription_enabled) Whether this connection will emit [`InputTranscript`](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.InputTranscript) events for the user’s audio. Providers that transcribe the user’s input (the default) leave this `True`. When it is `False`, no transcript arrives, so [`RealtimeSession`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession) finalizes a user turn from retained input audio instead (see `audio_retention`). Defaults to `True` so a connection that doesn’t override it never triggers the audio-only path (which would risk a duplicate turn if transcripts did arrive). **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### model\_name [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeConnection.model_name) The model id the server reported serving this session, when the provider reports one. Captured from the connect handshake (e.g. the OpenAI protocol’s `session.created`). It can differ from the requested model id: xAI accepts any model slug and silently substitutes its current default, reporting the actually-served model only here. `None` when the provider doesn’t report one (e.g. Gemini Live). The session stamps this on each [`ModelResponse.model_name`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse.model_name) , mirroring how request-response models record the response’s reported model rather than the requested one. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) #### reconnect\_restores\_in\_flight\_state [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeConnection.reconnect_restores_in_flight_state) Whether a reconnect continues the response and tool calls that were in flight when the socket dropped. Otherwise it only brings back the finalized conversation. Native session resumption (xAI Grok Voice) restores the in-flight generation server-side, and Gemini Live settles the cut turn in the connection before its [`RealtimeSessionReconnectEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSessionReconnectEvent) , so in both the [`RealtimeSession`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession) must not settle again and trusts `state_restored`. Local replay (OpenAI, Azure OpenAI) restores only finalized turns, so the session settles the interrupted turn itself and reports `state_restored=False`. Defaults to `True`; the OpenAI connection overrides it. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) #### transport\_errors [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeConnection.transport_errors) The exception types this connection’s transport raises when the link to the provider fails. A [`RealtimeSession`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession) maps these to [`RealtimeError`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeError) so a failed send surfaces as the same typed error as a failed receive, instead of leaking a `websockets` or provider-SDK exception the caller has no reason to expect from a model call. Leave empty if `send` already raises typed errors. The mapping covers the whole of [`send`](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeConnection.send) , so anything it does _besides_ writing to the transport — converting content, downloading media — must not raise these types, or a local failure would be reported as a lost connection. Do that work before the first frame goes out. **Type:** [`tuple`](https://docs.python.org/3/library/stdtypes.html#tuple) \[[`type`](https://docs.python.org/3/glossary.html#term-type)\ \[[`Exception`](https://docs.python.org/3/library/exceptions.html#Exception)\ \], …\] **Default:** `()` ### Methods [](https://pydantic.dev/docs/ai/api/realtime/codec/#methods) #### \_\_aiter\_\_ [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeConnection.__aiter__) `@abstractmethod` def __aiter__() -> AsyncIterator[RealtimeCodecEvent] Iterate over events received from the model. ##### Returns [](https://pydantic.dev/docs/ai/api/realtime/codec/#returns) [`AsyncIterator`](https://docs.python.org/3/library/typing.html#typing.AsyncIterator) \[`RealtimeCodecEvent`\] #### send [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeConnection.send) `@abstractmethod` `@async` def send(content: RealtimeInput) -> None Feed content into the session. Concrete connections accept provider-specific data and control inputs. OpenAI accepts audio, text, images, tool results, manual turn controls, cancellation, and truncation; Gemini accepts audio, text, images, and tool results. A high-level [`RealtimeSession`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession) checks profile-gated operations and raises [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) , as does a connection handed an input it can’t send. ##### Returns [](https://pydantic.dev/docs/ai/api/realtime/codec/#returns-1) [`None`](https://docs.python.org/3/library/constants.html#None) #### set\_message\_history [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeConnection.set_message_history) def set_message_history(message_history: Callable[[], Sequence[ModelMessage]]) -> None Tell the connection how to read the message history as it currently stands. A [`RealtimeSession`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession) calls this when it takes ownership of the connection, so a provider that loses server-side state on [reconnect](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.ReconnectPolicy) can replay the conversation into the new session instead of resuming with total amnesia. The session’s history grows as the call goes on, hence a callable rather than a snapshot. A no-op by default: providers with native session resumption (Gemini Live, xAI) have nothing to replay, and one that can’t seed a session at all has nowhere to put it. ##### Returns [](https://pydantic.dev/docs/ai/api/realtime/codec/#returns-2) [`None`](https://docs.python.org/3/library/constants.html#None) ResponseDone ------------ [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ResponseDone) The provider reported that its current response is done. This codec event is consumed by the session to finalize a [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) . Providers do not necessarily report a terminal for every model response the session records in history. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/codec/#attributes-6) #### event\_kind [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ResponseDone.event_kind) Event type identifier, used as a discriminator. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘response\_done’\] **Default:** `'response_done'` #### finish\_reason [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ResponseDone.finish_reason) Normalized reason the provider finished the response, when available. **Type:** [`FinishReason`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.FinishReason) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### interrupted [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ResponseDone.interrupted) Whether the response ended because it was cancelled (e.g. the user barged in). **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### provider\_details [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ResponseDone.provider_details) Raw provider terminal status details retained on the finalized response, when available. **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### provider\_response\_id [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ResponseDone.provider_response_id) Provider-assigned ID for the completed response, when available. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` SessionUsage ------------ [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.SessionUsage) Usage reported by the provider for a model response or another run-level operation. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/codec/#attributes-7) #### event\_kind [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.SessionUsage.event_kind) Event type identifier, used as a discriminator. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘session\_usage’\] **Default:** `'session_usage'` #### finish\_reason [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.SessionUsage.finish_reason) Normalized completion reason for the response this usage belongs to, when available. **Type:** [`FinishReason`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.FinishReason) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### provider\_response\_id [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.SessionUsage.provider_response_id) Provider-assigned ID for the response this usage belongs to, when available. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### response\_scoped [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.SessionUsage.response_scoped) Whether this usage belongs to a specific model response. `True`, the default, accumulates it both into the run total and the response’s `ModelResponse.usage`. `False` is run-level only, e.g. input audio transcription usage, which is billed on a separate model/meter and is accumulated into the run’s `RunUsage` but attributed to no `ModelResponse`. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `True` #### usage [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.SessionUsage.usage) Normalized usage ready to accumulate into a `RunUsage`. **Type:** [`RequestUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RequestUsage) ToolCall -------- [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ToolCall) The model is requesting a tool call. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/codec/#attributes-8) #### args [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ToolCall.args) Raw JSON-encoded arguments. May be an empty string if the model sent no arguments. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### item\_id [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ToolCall.item_id) Provider conversation-item ID for this call, when available. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### response\_usage\_follows [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ToolCall.response_usage_follows) Whether per-response [`SessionUsage`](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.SessionUsage) will follow this call before the provider’s response is complete. OpenAI-protocol providers report calls before `response.done`, which carries usage; the session uses this signal to keep all calls and their usage on the same `ModelResponse`. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### tool\_call\_id [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ToolCall.tool_call_id) Provider-assigned identifier for this call. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### tool\_name [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ToolCall.tool_name) Name of the tool to invoke. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) ToolCallCancelled ----------------- [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ToolCallCancelled) The model cancelled in-flight tool calls (e.g. the user barged in before they finished). Gemini Live sends this as `toolCallCancellation`; the session cancels the matching running tool tasks so their now-unwanted results are never sent back to the model. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/codec/#attributes-9) #### tool\_call\_ids [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ToolCallCancelled.tool_call_ids) Identifiers of the [`ToolCall`](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ToolCall) s that were cancelled. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] ToolResult ---------- [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ToolResult) The result of a tool call, rendered for the wire and sent back to the model. Built and sent by [`RealtimeSession`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession) after it settles a call: the string-only realtime tool channel and the retry/failure error-key wrapping mean the session renders the [`ToolReturnPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolReturnPart) or [`RetryPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.RetryPromptPart) it records in history down to this flat shape, so every provider sends exactly the same rendering. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/codec/#attributes-10) #### content [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ToolResult.content) Additional user content to send after the tool output when the provider supports it. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`UserContent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.UserContent)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### output [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ToolResult.output) The tool’s output, rendered as a string. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### tool\_call\_id [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ToolResult.tool_call_id) Identifier of the `ToolCall` this result answers. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) TruncateOutput -------------- [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.TruncateOutput) Truncate the model’s current audio output at `audio_end_ms`. After a barge-in the user only heard part of the model’s audio. Truncating tells the provider how much was actually played, so its stored transcript matches and the conversation context stays consistent. The provider resolves which output item to truncate from its own state. ### Attributes [](https://pydantic.dev/docs/ai/api/realtime/codec/#attributes-11) #### audio\_end\_ms [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.TruncateOutput.audio_end_ms) Milliseconds of the current output audio that were actually played before the interruption. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) merge\_realtime\_profile ------------------------ [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.merge_realtime_profile) def merge_realtime_profile( base: RealtimeModelProfile | None, *overrides: RealtimeModelProfile | None, ) -> RealtimeModelProfile Merge realtime profiles, with later layers overriding earlier ones. ### Returns [](https://pydantic.dev/docs/ai/api/realtime/codec/#returns-3) `RealtimeModelProfile` DEFAULT\_AUDIO\_SAMPLE\_RATE ---------------------------- [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.DEFAULT_AUDIO_SAMPLE_RATE) The sample rate, in Hz, assumed for PCM audio when a realtime model profile doesn’t specify one. **Default:** `24000` DEFAULT\_REALTIME\_PROFILE -------------------------- [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.DEFAULT_REALTIME_PROFILE) Default realtime model profile values. **Type:** `RealtimeModelProfile` **Default:** `{'supports_image_input': False, 'supports_manual_turn_control': False, 'supports_interruption': False, 'supports_output_truncation': False, 'supports_text_output': True, 'supports_session_seeding': False, 'supports_webrtc': False, 'supports_seeding_images': False, 'supports_seeding_audio': False, 'supports_async_tool_calls': False, 'supports_tool_return_schema': False, 'supported_native_tools': frozenset(), 'emits_input_speech_events': False, 'audio_input_sample_rate': DEFAULT_AUDIO_SAMPLE_RATE, 'audio_output_sample_rate': DEFAULT_AUDIO_SAMPLE_RATE}` RealtimeCodecEvent ------------------ [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeCodecEvent) Union of the low-level codec events yielded by [`RealtimeConnection`](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeConnection) . This is the provider-facing vocabulary: providers translate their wire protocol into these events, and [`RealtimeSession`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession) translates them again into the shared [`RealtimeEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeEvent) vocabulary while building [`ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage) history. **Default:** `TypeAliasType('RealtimeCodecEvent', AudioDelta | OutputTranscript | InputTranscript | ToolCall | ToolCallCancelled | ResponseDone | RealtimeInputSpeechStartEvent | RealtimeResponseInterruptedEvent | RealtimeInputSpeechEndEvent | RealtimeOutputSpeechStartEvent | RealtimeOutputSpeechEndEvent | RealtimeInputTranscriptionErrorEvent | SessionUsage | RealtimeSessionReconnectEvent | ConversationCreated | ConversationItemCreated | PartStartEvent | PartEndEvent | RealtimeSessionErrorEvent)` RealtimeInput ------------- [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeInput) Union of content types accepted by [`RealtimeConnection.send`](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeConnection.send) . The connection-level counterpart of [`RealtimeSessionInput`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSessionInput) , already normalized: a `str` is a complete text turn, a [`BinaryAudio`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BinaryAudio) carries a raw mono PCM16 chunk at the model’s [`audio_input_sample_rate`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.audio_input_sample_rate) (`media_type='audio/pcm'`), and a [`BinaryImage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BinaryImage) an image frame. The connection additionally accepts the turn-control verbs and [`ToolResult`](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ToolResult) , which [`RealtimeSession`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession) sends on the caller’s behalf. **Default:** `TypeAliasType('RealtimeInput', 'str | BinaryAudio | BinaryImage | CommitAudio | ClearAudio | CreateResponse | CancelResponse | TruncateOutput | ToolResult')` RealtimeSessionInput -------------------- [](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.RealtimeSessionInput) The content types a caller feeds into [`RealtimeSession.send`](https://pydantic.dev/docs/ai/api/pydantic-ai/realtime/#pydantic_ai.realtime.RealtimeSession.send) . Session content only, in the shared message vocabulary: a `str` is a complete text turn, and [`BinaryContent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BinaryContent) carries an image frame, WAV audio (unwrapped to raw PCM before it is streamed, matching the history-seeding path), or a raw PCM chunk (`media_type='audio/pcm'`). The session normalizes these before forwarding them to the connection. Turn-control verbs (`CommitAudio`, `ClearAudio`, `CreateResponse`, `CancelResponse`, `TruncateOutput`) are connection-level vocabulary driven through the dedicated `RealtimeSession` methods (`commit_audio()`, `clear_audio()`, `create_response()`, `interrupt()`), and [`ToolResult`](https://pydantic.dev/docs/ai/api/realtime/codec/#pydantic_ai.realtime.codec.ToolResult) is sent by the session itself when a tool completes — neither is accepted by `send()`. **Default:** `TypeAliasType('RealtimeSessionInput', 'str | BinaryContent')` Was this page helpful? Thanks for your feedback! --- # Anthropic | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/models/anthropic/#_top) Anthropic ========= Install ------- [](https://pydantic.dev/docs/ai/models/anthropic/#install) To use `AnthropicModel` models, you need to either install `pydantic-ai`, or install `pydantic-ai-slim` with the `anthropic` optional group: * [pip](https://pydantic.dev/docs/ai/models/anthropic/#tab-panel-100) * [uv](https://pydantic.dev/docs/ai/models/anthropic/#tab-panel-101) Terminal pip install "pydantic-ai-slim[anthropic]" Terminal uv add "pydantic-ai-slim[anthropic]" Configuration ------------- [](https://pydantic.dev/docs/ai/models/anthropic/#configuration) To use [Anthropic](https://anthropic.com/) through their API, go to [console.anthropic.com/settings/keys](https://console.anthropic.com/settings/keys) to generate an API key. `AnthropicModelName` contains a list of available Anthropic models. Environment variable -------------------- [](https://pydantic.dev/docs/ai/models/anthropic/#environment-variable) Once you have the API key, you can set it as an environment variable: Terminal export ANTHROPIC_API_KEY='your-api-key' You can then use `AnthropicModel` by name: from pydantic_ai import Agent agent = Agent('anthropic:claude-sonnet-4-6') ... Or initialise the model directly with just the model name: from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModel model = AnthropicModel('claude-sonnet-4-5') agent = Agent(model) ... `provider` argument ------------------- [](https://pydantic.dev/docs/ai/models/anthropic/#provider-argument) You can provide a custom `Provider` via the `provider` argument: from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModel from pydantic_ai.providers.anthropic import AnthropicProvider model = AnthropicModel( 'claude-sonnet-4-5', provider=AnthropicProvider(api_key='your-api-key') ) agent = Agent(model) ... Custom HTTP Client ------------------ [](https://pydantic.dev/docs/ai/models/anthropic/#custom-http-client) You can customize the `AnthropicProvider` with a custom `httpx.AsyncClient`: from httpx import AsyncClient from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModel from pydantic_ai.providers.anthropic import AnthropicProvider custom_http_client = AsyncClient(timeout=30) model = AnthropicModel( 'claude-sonnet-4-5', provider=AnthropicProvider(api_key='your-api-key', http_client=custom_http_client), ) agent = Agent(model) ... Model settings -------------- [](https://pydantic.dev/docs/ai/models/anthropic/#model-settings) You can customize model behavior using [`AnthropicModelSettings`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings) : from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModel, AnthropicModelSettings model = AnthropicModel('claude-sonnet-4-5') settings = AnthropicModelSettings( temperature=0.2, top_k=40, service_tier='auto', ) agent = Agent(model, model_settings=settings) ... ### Service tier [](https://pydantic.dev/docs/ai/models/anthropic/#service-tier) Anthropic supports controlling the [service tier](https://platform.claude.com/docs/en/api/service-tiers) to manage latency and throughput. You can use the unified [`service_tier`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.service_tier) field or the provider-specific [`anthropic_service_tier`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_service_tier) field. `anthropic_service_tier` takes precedence over the unified field when both are set, and accepts Anthropic’s native values (`'auto'` or `'standard_only'`). The unified field maps as follows for Anthropic: * `'auto'`: passed through as `'auto'` (Anthropic’s native value — uses priority capacity when available). * `'default'`: maps to `'standard_only'` (forces the standard tier, opting out of priority capacity). * `'flex'` and `'priority'` are not part of Anthropic’s tier model and are silently ignored. Cloud Platform Integrations --------------------------- [](https://pydantic.dev/docs/ai/models/anthropic/#cloud-platform-integrations) You can use Anthropic models through cloud platforms by passing a custom client to [`AnthropicProvider`](https://pydantic.dev/docs/ai/api/pydantic-ai/providers/#pydantic_ai.providers.anthropic.AnthropicProvider) . ### AWS Bedrock [](https://pydantic.dev/docs/ai/models/anthropic/#aws-bedrock) To use Claude models via [AWS Bedrock](https://aws.amazon.com/bedrock/claude/) , follow the [Anthropic documentation](https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock) on how to set up a Bedrock client and then pass it to `AnthropicProvider`. Both the newer `AsyncAnthropicBedrockMantle` client (recommended by Anthropic, using the Messages API) and the legacy `AsyncAnthropicBedrock` client (using the `InvokeModel` API with ARN-versioned model IDs) are supported: from anthropic import AsyncAnthropicBedrockMantle from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModel from pydantic_ai.providers.anthropic import AnthropicProvider bedrock_client = AsyncAnthropicBedrockMantle() # Uses AWS credentials from environment provider = AnthropicProvider(anthropic_client=bedrock_client) model = AnthropicModel('anthropic.claude-haiku-4-5', provider=provider) agent = Agent(model) ... ### Google Cloud [](https://pydantic.dev/docs/ai/models/anthropic/#google-cloud) To use Claude models via [Google Cloud Vertex AI](https://cloud.google.com/vertex-ai/generative-ai/docs/partner-models/use-claude) , follow the [Anthropic documentation](https://docs.anthropic.com/en/api/claude-on-vertex-ai) on how to set up an `AsyncAnthropicVertex` client and then pass it to `AnthropicProvider`: from anthropic import AsyncAnthropicVertex from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModel from pydantic_ai.providers.anthropic import AnthropicProvider vertex_client = AsyncAnthropicVertex(region='us-east5', project_id='your-project-id') provider = AnthropicProvider(anthropic_client=vertex_client) model = AnthropicModel('claude-sonnet-4-5', provider=provider) agent = Agent(model) ... ### Microsoft Foundry [](https://pydantic.dev/docs/ai/models/anthropic/#microsoft-foundry) To use Claude models via [Microsoft Foundry](https://ai.azure.com/) , follow the [Anthropic documentation](https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry) on how to set up an `AsyncAnthropicFoundry` client and then pass it to `AnthropicProvider`: from anthropic import AsyncAnthropicFoundry from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModel from pydantic_ai.providers.anthropic import AnthropicProvider foundry_client = AsyncAnthropicFoundry( api_key='your-foundry-api-key', # Or set ANTHROPIC_FOUNDRY_API_KEY resource='your-resource-name', ) provider = AnthropicProvider(anthropic_client=foundry_client) model = AnthropicModel('claude-sonnet-4-5', provider=provider) agent = Agent(model) ... See [Anthropic’s Microsoft Foundry documentation](https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry) for setup instructions including Entra ID authentication. Task Budgets (Beta) ------------------- [](https://pydantic.dev/docs/ai/models/anthropic/#task-budgets-beta) Anthropic’s [task budgets](https://platform.claude.com/docs/en/build-with-claude/task-budgets) let you give Claude an advisory token budget for a full agentic loop — including thinking, tool calls, tool results, and output — so the model can pace itself and finish gracefully as the budget is consumed. Configure them with [`AnthropicModelSettings.anthropic_task_budget`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_task_budget) , which takes an [`AnthropicTaskBudget`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicTaskBudget) payload and maps to `output_config.task_budget`. Pydantic AI automatically enables Anthropic’s required `task-budgets-2026-03-13` beta when this setting is present. Support is currently limited to native Anthropic `claude-opus-4-7`, `claude-opus-4-8`, `claude-opus-5`, and `claude-sonnet-5` requests, not Bedrock, Vertex, or Microsoft Foundry Anthropic model IDs. anthropic\_task\_budget.py from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModel, AnthropicModelSettings model = AnthropicModel('claude-opus-4-8') settings = AnthropicModelSettings( anthropic_task_budget={'type': 'tokens', 'total': 20_000}, ) agent = Agent(model, model_settings=settings) ... Task budgets compose with [`anthropic_effort`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_effort) : effort tunes per-step reasoning depth, while task budgets cap total work across the loop. Both fields end up under the same `output_config` object. ### Carrying budgets across compaction [](https://pydantic.dev/docs/ai/models/anthropic/#carrying-budgets-across-compaction) If you use [`AnthropicCompaction`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicCompaction) for server-side compaction, you can skip this section: the server tracks the countdown itself, so leave `remaining` unset and let `total` self-regulate. The `remaining` field on `task_budget` is for _client-side_ compaction patterns where you summarize earlier turns yourself between requests, so the server has no memory of how much budget was spent before the rewrite. Pydantic AI does not track `remaining` for you — accumulate token usage across requests yourself (e.g. from [`RunUsage`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.RunUsage) on each run) and pass the updated value on the next request so the countdown continues from where you left off rather than resetting to `total`. Setting `remaining` also invalidates any prompt-cache prefix that contains the budget, so if you want to preserve caching, set `total` once and let the server self-regulate against the running countdown. Prompt Caching -------------- [](https://pydantic.dev/docs/ai/models/anthropic/#prompt-caching) Anthropic supports [prompt caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching) to reduce costs by caching parts of your prompts. Pydantic AI supports automatic caching, per-block message caching, and explicit cache breakpoints: ### Automatic Caching [](https://pydantic.dev/docs/ai/models/anthropic/#automatic-caching) The simplest way to enable prompt caching is with [`AnthropicModelSettings.anthropic_cache`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_cache) . This uses Anthropic’s [automatic caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching#automatic-caching) , passing a top-level `cache_control` parameter so the server automatically applies a cache breakpoint to the last cacheable block in each request: from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModelSettings agent = Agent( 'anthropic:claude-sonnet-4-6', instructions='You are a helpful assistant.', model_settings=AnthropicModelSettings( anthropic_cache=True, ), ) result1 = agent.run_sync('What is the capital of France?') result2 = agent.run_sync( 'What is the capital of Germany?', message_history=result1.all_messages() ) print(f'Cache write: {result1.usage.cache_write_tokens}') print(f'Cache read: {result2.usage.cache_read_tokens}') print(f'Cache hit ratio: {result2.usage.cache_hit_ratio}') This is ideal for multi-turn conversations where the cache breakpoint should move forward as the conversation grows. You can also specify a custom TTL with `anthropic_cache='1h'`. ### Per-block Message Caching [](https://pydantic.dev/docs/ai/models/anthropic/#per-block-message-caching) As an alternative to `anthropic_cache`, [`AnthropicModelSettings.anthropic_cache_messages`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_cache_messages) adds per-block `cache_control` to the last content block of the final message instead of using Anthropic’s top-level automatic caching parameter. Use this with Anthropic-compatible gateways and proxies (such as MiniMax, OpenRouter, or LiteLLM) that accept the Anthropic message format but don’t support top-level automatic caching: from anthropic import AsyncAnthropic from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModel, AnthropicModelSettings from pydantic_ai.providers.anthropic import AnthropicProvider client = AsyncAnthropic( api_key='your-api-key', base_url='https://your-anthropic-compatible-gateway.example.com', ) model = AnthropicModel( 'claude-sonnet-4-6', provider=AnthropicProvider(anthropic_client=client), ) agent = Agent( model, model_settings=AnthropicModelSettings( anthropic_cache_messages=True, ), ) result = agent.run_sync('What is the capital of France?') print(result.output) You can also specify a custom TTL with `anthropic_cache_messages='1h'`. `anthropic_cache_messages` cannot be combined with `anthropic_cache`. ### Explicit Cache Breakpoints [](https://pydantic.dev/docs/ai/models/anthropic/#explicit-cache-breakpoints) In addition to automatic caching, Pydantic AI provides several ways to place cache breakpoints on specific content: 1. **Cache User Messages with [`CachePoint`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.CachePoint) **: Insert a `CachePoint` marker in your user messages to cache everything before it 2. **Cache the Final Message Block**: Set [`AnthropicModelSettings.anthropic_cache_messages`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_cache_messages) to `True` (uses 5m TTL by default) or specify `'5m'` / `'1h'` directly 3. **Cache System Instructions**: Set [`AnthropicModelSettings.anthropic_cache_instructions`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_cache_instructions) to `True` (uses 5m TTL by default) or specify `'5m'` / `'1h'` directly 4. **Cache Tool Definitions**: Set [`AnthropicModelSettings.anthropic_cache_tool_definitions`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_cache_tool_definitions) to `True` (uses 5m TTL by default) or specify `'5m'` / `'1h'` directly #### Example: Comprehensive Caching Strategy [](https://pydantic.dev/docs/ai/models/anthropic/#example-comprehensive-caching-strategy) Combine automatic caching with explicit breakpoints for maximum savings. Automatic caching handles the conversation, while explicit breakpoints pin system instructions and tool definitions: from pydantic_ai import Agent, RunContext from pydantic_ai.models.anthropic import AnthropicModelSettings agent = Agent( 'anthropic:claude-sonnet-4-6', instructions='Detailed instructions...', model_settings=AnthropicModelSettings( anthropic_cache=True, # Server auto-caches last block anthropic_cache_instructions=True, # Explicitly cache system instructions anthropic_cache_tool_definitions='1h', # Explicitly cache tool definitions with 1h TTL ), ) @agent.tool def search_docs(ctx: RunContext, query: str) -> str: """Search documentation.""" return f'Results for {query}' result = agent.run_sync('Search for Python best practices') print(result.output) ### Smart Instruction Caching [](https://pydantic.dev/docs/ai/models/anthropic/#smart-instruction-caching) When you use `anthropic_cache_instructions` with both static and dynamic [instructions](https://pydantic.dev/docs/ai/core-concepts/agent/#instructions) , Pydantic AI automatically places the cache boundary at the optimal point. Static instructions (from `Agent(instructions=...)`) are sorted before dynamic instructions (from `@agent.instructions` functions or [toolsets](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) ), and the cache point is placed after the last static instruction block. This means your stable, static instructions are cached efficiently, while dynamic instructions (which may change between requests) remain outside the cache boundary and don’t cause cache invalidation. from datetime import date from pydantic_ai import Agent, RunContext from pydantic_ai.models.anthropic import AnthropicModelSettings agent = Agent( 'anthropic:claude-sonnet-4-6', deps_type=str, instructions='You are a helpful customer service agent. Follow company policy.', # (1) model_settings=AnthropicModelSettings( anthropic_cache_instructions=True, # (2) ), ) @agent.instructions def dynamic_context(ctx: RunContext[str]) -> str: # (3) return f"Customer name: {ctx.deps}. Today's date: {date.today()}." result = agent.run_sync('What is your return policy?', deps='Alice') print(result.output) Static instructions are cached across requests. Enables smart cache placement at the static/dynamic boundary. Dynamic instructions change per-request and are not cached. ### Fine-Grained Control with CachePoint [](https://pydantic.dev/docs/ai/models/anthropic/#fine-grained-control-with-cachepoint) Use manual `CachePoint` markers to control cache locations precisely: from pydantic_ai import Agent, CachePoint agent = Agent( 'anthropic:claude-sonnet-4-6', instructions='Instructions...', ) # Manually control cache points for specific content blocks result = agent.run_sync([\ 'Long context from documentation...',\ CachePoint(), # Cache everything up to this point\ 'First question'\ ]) print(result.output) ### Accessing Cache Usage Statistics [](https://pydantic.dev/docs/ai/models/anthropic/#accessing-cache-usage-statistics) Access cache usage statistics via `result.usage`: from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModelSettings agent = Agent( 'anthropic:claude-sonnet-4-6', instructions='Instructions...', model_settings=AnthropicModelSettings( anthropic_cache=True, ), ) result = agent.run_sync('Your question') usage = result.usage print(f'Cache write tokens: {usage.cache_write_tokens}') print(f'Cache read tokens: {usage.cache_read_tokens}') ### Cache Point Limits [](https://pydantic.dev/docs/ai/models/anthropic/#cache-point-limits) Anthropic enforces a maximum of 4 cache points per request. Pydantic AI automatically manages this limit to ensure your requests always comply without errors. #### How Cache Points Are Allocated [](https://pydantic.dev/docs/ai/models/anthropic/#how-cache-points-are-allocated) Cache points can come from several sources: 1. **Automatic caching**: Via `anthropic_cache` (the server applies 1 cache point to the last cacheable block) 2. **Final message block**: Via `anthropic_cache_messages` setting (adds cache point to last message content block) 3. **System Prompt**: Via `anthropic_cache_instructions` setting (adds cache point to last system prompt block) 4. **Tool Definitions**: Via `anthropic_cache_tool_definitions` setting (adds cache point to last tool definition) 5. **Messages**: Via `CachePoint` markers (adds cache points to message content) Each setting uses **at most 1 cache point**, but you can combine them — except `anthropic_cache` and `anthropic_cache_messages`, which are mutually exclusive. If the total exceeds 4, Pydantic AI automatically trims excess cache points from older messages. #### Example: Combining Automatic and Explicit Caching [](https://pydantic.dev/docs/ai/models/anthropic/#example-combining-automatic-and-explicit-caching) Define an agent with automatic caching plus explicit breakpoints: from pydantic_ai import Agent, CachePoint from pydantic_ai.models.anthropic import AnthropicModelSettings agent = Agent( 'anthropic:claude-sonnet-4-6', instructions='Detailed instructions...', model_settings=AnthropicModelSettings( anthropic_cache=True, # 1 cache point (server-applied) anthropic_cache_instructions=True, # 1 cache point anthropic_cache_tool_definitions=True, # 1 cache point ), ) @agent.tool_plain def my_tool() -> str: return 'result' # 3 of 4 slots used (1 automatic + 1 instructions + 1 tools) # Room for 1 more explicit CachePoint marker result = agent.run_sync([\ 'Context', CachePoint(), # 4th cache point - OK\ 'Question'\ ]) print(result.output) usage = result.usage print(f'Cache write tokens: {usage.cache_write_tokens}') print(f'Cache read tokens: {usage.cache_read_tokens}') #### Automatic Cache Point Limiting [](https://pydantic.dev/docs/ai/models/anthropic/#automatic-cache-point-limiting) When explicit cache points from all sources (settings + `CachePoint` markers) exceed the available budget, Pydantic AI automatically removes excess cache points from **older message content** (keeping the most recent ones). Define an agent with 2 explicit cache points from settings: from pydantic_ai import Agent, CachePoint from pydantic_ai.models.anthropic import AnthropicModelSettings agent = Agent( 'anthropic:claude-sonnet-4-6', instructions='Instructions...', model_settings=AnthropicModelSettings( anthropic_cache_instructions=True, # 1 cache point anthropic_cache_tool_definitions=True, # 1 cache point ), ) @agent.tool_plain def search() -> str: return 'data' # Already using 2 cache points (instructions + tools) # Can add 2 more CachePoint markers (4 total limit) result = agent.run_sync([\ 'Context 1', CachePoint(), # Oldest - will be removed\ 'Context 2', CachePoint(), # Will be kept (3rd point)\ 'Context 3', CachePoint(), # Will be kept (4th point)\ 'Question'\ ]) # Final cache points: instructions + tools + Context 2 + Context 3 = 4 print(result.output) usage = result.usage print(f'Cache write tokens: {usage.cache_write_tokens}') print(f'Cache read tokens: {usage.cache_read_tokens}') **Key Points**: * System and tool cache points are **always preserved** * `anthropic_cache` counts as 1 cache point, just like `anthropic_cache_instructions` and `anthropic_cache_tool_definitions` * Excess `CachePoint` markers in messages are removed from oldest to newest when the limit is exceeded * This ensures critical caching (instructions/tools) is maintained while still benefiting from message-level caching Mid-conversation system messages -------------------------------- [](https://pydantic.dev/docs/ai/models/anthropic/#mid-conversation-system-messages) Adding an instruction to the agent’s `system_prompt` partway through a long session rewrites the front of the prompt, which invalidates every cached prefix behind it. Anthropic avoids that by accepting a system message _inside_ the conversation, at the instruction’s own position in the history rather than ahead of it, so everything cached up to that point stays cached. Any [`SystemPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SystemPromptPart) outside the first [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) is a mid-conversation instruction — whether it came from a stored `message_history` or from [`RunContext.enqueue`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.enqueue) during a run. There’s nothing extra to turn on: mid\_conversation\_system\_prompt.py from pydantic_ai import Agent, RunContext, SystemPromptPart agent = Agent('anthropic:claude-opus-4-8', system_prompt='You are a code reviewer.') @agent.tool def require_type_annotations(ctx: RunContext[None]) -> str: ctx.enqueue(SystemPromptPart(content='Every suggestion must include explicit type annotations.')) return 'rule added' Keeping the instruction in place leaves the prefix ahead of it reusable, but it doesn’t enable caching on its own — that still comes from [`anthropic_cache`](https://pydantic.dev/docs/ai/models/anthropic/#automatic-caching) , `anthropic_cache_messages`, or an explicit [`CachePoint`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.CachePoint) . A `CachePoint` at the end of an enqueued batch caches everything before it in that batch, the instruction included. One with more content after it caches up to where you put it and leaves the instruction outside: the instruction is sent after the content it accompanies, so it can’t be inside a boundary that content is outside of. Support varies by model and by transport — the [Microsoft Foundry](https://pydantic.dev/docs/ai/models/anthropic/#microsoft-foundry) integration doesn’t serve the role, and some Claude models accept the entry without acting on it. Anthropic’s [mid-conversation system messages docs](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages) have the current list. Pydantic AI picks the rendering that works for the model and transport you’re using, falling back to a ``\-tagged user message at the same position, so the instruction applies where you put it either way. The difference between the two shows up on instructions a model _should_ be wary of taking from its user: given the native entry, Claude will lift a restriction its top-level prompt set, and given the identical text in a `` tag it refuses. For an instruction with nothing to distrust, such as a change of format, both work. See [mid-conversation system prompts](https://pydantic.dev/docs/ai/core-concepts/message-history/#mid-conversation-system-prompts) for how these behave across providers, how to phrase one, and why untrusted content doesn’t belong in one. Fast mode --------- [](https://pydantic.dev/docs/ai/models/anthropic/#fast-mode) Fast mode provides higher output tokens per second and is currently supported on **Claude Opus 4.6**, **Claude Opus 4.7**, **Claude Opus 4.8**, and **Claude Opus 5**. It is a research preview. Set [`anthropic_speed`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_speed) to `'fast'` to enable it; Pydantic AI automatically adds the required `fast-mode-2026-02-01` beta. On unsupported models, `anthropic_speed='fast'` is ignored with a `UserWarning`. For pricing, rate limits, and the latest list of supported models, see the [Anthropic fast mode docs](https://platform.claude.com/docs/en/build-with-claude/fast-mode) . from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModelSettings agent = Agent( 'anthropic:claude-opus-4-8', model_settings=AnthropicModelSettings(anthropic_speed='fast'), ) ... Forced tool choice ------------------ [](https://pydantic.dev/docs/ai/models/anthropic/#forced-tool-choice) Most Anthropic models let you force a tool call via [`tool_choice='required'`](https://pydantic.dev/docs/ai/api/pydantic-ai/settings/#pydantic_ai.settings.ModelSettings.tool_choice) (or a list of tool names), except while [extended thinking](https://pydantic.dev/docs/ai/capabilities/thinking/#anthropic) is enabled — [adaptive thinking](https://pydantic.dev/docs/ai/capabilities/thinking/#adaptive-thinking-effort) is compatible with forcing. **Claude Fable 5** and the **Claude Mythos** models reject a forced tool choice unconditionally — even without thinking — so Pydantic AI marks them with [`anthropic_supports_forced_tool_choice=False`](https://pydantic.dev/docs/ai/api/pydantic-ai/profiles/#pydantic_ai.profiles.anthropic.AnthropicModelProfile.anthropic_supports_forced_tool_choice) . On a model that doesn’t support forcing: * An explicit `tool_choice='required'` (or a list of tool names) raises a [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) ; use `tool_choice='auto'` instead. * A `required` choice that Pydantic AI resolved on your behalf (e.g. from an [output tool](https://pydantic.dev/docs/ai/core-concepts/output/#tool-output) ) falls back softly to `'auto'`. If the resolved choice named a single tool, the available tool list is filtered to that tool while `tool_choice` remains `'auto'`, which invalidates Anthropic’s prompt cache since the cached prefix includes the tool array. The model may therefore answer with text instead of calling it; when an output tool is required, Pydantic AI retries with a prompt to call a tool. Because [Tool Output](https://pydantic.dev/docs/ai/core-concepts/output/#tool-output) resolves to a forced tool choice, extended thinking is also incompatible with it: a bare structured `output_type` switches to [Native Output](https://pydantic.dev/docs/ai/core-concepts/output/#native-output) (or [Prompted Output](https://pydantic.dev/docs/ai/core-concepts/output/#prompted-output) on models without JSON schema support), and an explicit `ToolOutput(...)` raises a [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) . Adaptive thinking keeps Tool Output, except on the models above that reject forcing outright — whenever a thinking setting is configured, those behave as they always have: a bare structured `output_type` switches away from Tool Output, and an explicit `ToolOutput(...)` raises a [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) . Message Compaction ------------------ [](https://pydantic.dev/docs/ai/models/anthropic/#message-compaction) Anthropic supports [automatic context compaction](https://docs.anthropic.com/en/docs/build-with-claude/compaction) to manage long conversations. When input tokens exceed a configured threshold, the API automatically generates a summary that replaces older messages while preserving context. After compaction, subsequent requests send only the compacted window, from the latest compaction block onward, which reduces request size — the API ignores earlier content either way. The standing system prompt is unaffected: it’s sent as the separate `system` parameter, which compaction doesn’t replace. The easiest way to enable compaction is with the [`AnthropicCompaction`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicCompaction) capability: anthropic\_compaction.py from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicCompaction agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[AnthropicCompaction(token_threshold=100_000)], ) The capability accepts: * **`token_threshold`** (default: 150,000, minimum: 50,000): Compaction triggers when input tokens exceed this value. * **`instructions`**: Custom instructions for how the summary should be generated. * **`pause_after_compaction`**: When `True`, the response stops after the compaction block with `stop_reason='compaction'`, allowing explicit handling before continuing. Alternatively, you can configure compaction directly via model settings using [`anthropic_context_management`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_context_management) : anthropic\_compaction\_settings.py from pydantic_ai import Agent from pydantic_ai.models.anthropic import AnthropicModelSettings agent = Agent('anthropic:claude-sonnet-4-6') result = agent.run_sync( 'Hello!', model_settings=AnthropicModelSettings( anthropic_context_management={ 'edits': [{'type': 'compact_20260112', 'trigger': {'type': 'input_tokens', 'value': 100_000}}] } ), ) Code Execution Tool Version --------------------------- [](https://pydantic.dev/docs/ai/models/anthropic/#code-execution-tool-version) By default, Pydantic AI chooses a compatible Anthropic code execution tool version for the selected model. You can override this with [`AnthropicModelSettings.anthropic_code_execution_tool_version`](https://pydantic.dev/docs/ai/api/models/anthropic/#pydantic_ai.models.anthropic.AnthropicModelSettings.anthropic_code_execution_tool_version) when you need a specific supported Anthropic tool version: anthropic\_code\_execution\_tool\_version.py from pydantic_ai import Agent, CodeExecutionTool from pydantic_ai.capabilities import NativeTool from pydantic_ai.models.anthropic import AnthropicModelSettings agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[NativeTool(CodeExecutionTool())], model_settings=AnthropicModelSettings(anthropic_code_execution_tool_version='20260120'), ) Pydantic AI raises a [`UserError`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.UserError) if you explicitly select a tool version that the model does not support. Was this page helpful? Thanks for your feedback! --- # Client | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/mcp/client/#_top) Client ====== Pydantic AI can act as an [MCP client](https://modelcontextprotocol.io/quickstart/client) , connecting to MCP servers to use their tools as part of an agent run. The [`MCPToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset) [toolset](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) wraps the [FastMCP Client](https://gofastmcp.com/clients/) and works with both local (stdio) and remote (Streamable HTTP, SSE) MCP servers. Install ------- [](https://pydantic.dev/docs/ai/mcp/client/#install) You need to either install [`pydantic-ai`](https://pydantic.dev/docs/ai/overview/install/) , or [`pydantic-ai-slim`](https://pydantic.dev/docs/ai/overview/install/#slim-install) with the `mcp` optional group: * [pip](https://pydantic.dev/docs/ai/mcp/client/#tab-panel-98) * [uv](https://pydantic.dev/docs/ai/mcp/client/#tab-panel-99) Terminal pip install "pydantic-ai-slim[mcp]" Terminal uv add "pydantic-ai-slim[mcp]" Usage ----- [](https://pydantic.dev/docs/ai/mcp/client/#usage) An [`MCPToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset) accepts any of the following as its first positional argument: * A URL string (Streamable HTTP, or SSE if the path ends in `/sse`) * A path to a local Python or Node.js script (run via stdio) * A [FastMCP transport](https://gofastmcp.com/clients/transports) like [`StdioTransport`](https://gofastmcp.com/clients/transports) , [`StreamableHttpTransport`](https://gofastmcp.com/clients/transports) , or [`SSETransport`](https://gofastmcp.com/clients/transports) * A pre-built [`fastmcp.Client`](https://gofastmcp.com/clients/client) (for advanced FastMCP-specific configuration like [OAuth](https://gofastmcp.com/clients/auth/oauth) or [tool transformation](https://gofastmcp.com/patterns/tool-transformation) ) * An in-process [FastMCP server](https://gofastmcp.com/servers/) (for testing or single-process deployments — no network round trip) Each `MCPToolset` instance is a [toolset](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) and can be registered with an [`Agent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent) via the `toolsets` argument. You can use [`async with agent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.__aenter__) to open and close connections to all registered MCP toolsets (and in the case of stdio servers, start and stop the subprocesses) around the context where they’ll be used in agent runs. You can also use `async with toolset` to manage the lifecycle of a specific toolset directly, for example if you’d like to share it across multiple agents. If you don’t explicitly enter one of these context managers, the toolset will be opened and closed automatically as needed. Note that a shared `MCPToolset` instance connects to the server as a single identity; if your users have their own credentials for the MCP server, see [per-user authentication](https://pydantic.dev/docs/ai/mcp/client/#per-user-authentication) . ### Streamable HTTP [](https://pydantic.dev/docs/ai/mcp/client/#streamable-http) The [Streamable HTTP](https://modelcontextprotocol.io/introduction#streamable-http) transport is the recommended way to connect to a remote MCP server. Before creating the toolset, we need to run a server that supports the Streamable HTTP transport. streamable\_http\_server.py from mcp.server.fastmcp import FastMCP app = FastMCP() @app.tool() def add(a: int, b: int) -> int: return a + b if __name__ == '__main__': app.run(transport='streamable-http') Then we can create the toolset: mcp\_streamable\_http\_client.py from pydantic_ai import Agent from pydantic_ai.mcp import MCPToolset toolset = MCPToolset('http://localhost:8000/mcp') # (1) agent = Agent('openai:gpt-5.2', toolsets=[toolset]) # (2) async def main(): result = await agent.run('What is 7 plus 5?') print(result.output) #> The answer is 12. Define the MCP toolset with the URL used to connect. Create an agent with the MCP toolset attached. _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main())` to run `main`)_ **What’s happening here?** * The model receives the prompt “What is 7 plus 5?” * The model decides “Oh, I’ve got this `add` tool, that will be a good way to answer this question” * The model returns a tool call * Pydantic AI sends the tool call to the MCP server using the Streamable HTTP transport * The model is called again with the return value of running the `add` tool (12) * The model returns the final answer You can visualise this clearly, and even see the tool call, by adding three lines of code to instrument the example with [logfire](https://logfire.pydantic.dev/docs) : mcp\_streamable\_http\_client\_logfire.py import logfire logfire.configure() logfire.instrument_pydantic_ai() ### SSE [](https://pydantic.dev/docs/ai/mcp/client/#sse) The [HTTP + Server-Sent Events](https://spec.modelcontextprotocol.io/specification/2024-11-05/basic/transports/#http-with-sse) transport is also supported. URLs ending in `/sse` are auto-detected as SSE; for any other path, pass an explicit [`SSETransport`](https://gofastmcp.com/clients/transports) . mcp\_sse\_client.py from pydantic_ai import Agent from pydantic_ai.mcp import MCPToolset toolset = MCPToolset('http://localhost:3001/sse') agent = Agent('openai:gpt-5.2', toolsets=[toolset]) ### Stdio [](https://pydantic.dev/docs/ai/mcp/client/#stdio) MCP also offers the [stdio transport](https://spec.modelcontextprotocol.io/specification/2024-11-05/basic/transports/#stdio) , where the server is run as a subprocess and communicates with the client over `stdin` and `stdout`. Pass a path to a Python or Node.js script, or build a [`StdioTransport`](https://gofastmcp.com/clients/transports) for full control over the command, arguments, and environment. mcp\_stdio\_client.py from fastmcp.client.transports import StdioTransport from pydantic_ai import Agent from pydantic_ai.mcp import MCPToolset toolset = MCPToolset(StdioTransport(command='python', args=['mcp_server.py'])) agent = Agent('openai:gpt-5.2', toolsets=[toolset]) ### In-process FastMCP server [](https://pydantic.dev/docs/ai/mcp/client/#in-process-fastmcp-server) If you already have a [FastMCP server](https://gofastmcp.com/servers/) in the same Python process as your agent, you can hand it directly to `MCPToolset` and save the network round trip: mcp\_in\_process\_server.py from fastmcp import FastMCP from pydantic_ai import Agent from pydantic_ai.mcp import MCPToolset fastmcp_server = FastMCP('my_server') @fastmcp_server.tool() async def add(a: int, b: int) -> int: return a + b toolset = MCPToolset(fastmcp_server) agent = Agent('openai:gpt-5.2', toolsets=[toolset]) async def main(): result = await agent.run('What is 7 plus 5?') print(result.output) #> The answer is 12. _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main())` to run `main`)_ Loading MCP toolsets from configuration --------------------------------------- [](https://pydantic.dev/docs/ai/mcp/client/#loading-mcp-toolsets-from-configuration) Instead of constructing `MCPToolset` instances individually, you can load multiple toolsets from a JSON configuration file using [`load_mcp_toolsets()`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.load_mcp_toolsets) . This is particularly useful when you need to manage multiple MCP servers or want to configure servers externally without modifying code. ### Configuration format [](https://pydantic.dev/docs/ai/mcp/client/#configuration-format) The configuration file should be a JSON file with an `mcpServers` object containing server definitions. Each server is identified by a unique key and contains the configuration for that server: mcp\_config.json { "mcpServers": { "python-runner": { "command": "uv", "args": ["run", "mcp-run-python", "stdio"] }, "weather": { "command": "python", "args": ["mcp_server.py"] }, "weather-api": { "url": "http://localhost:3001/sse" }, "calculator": { "url": "http://localhost:8000/mcp" } } } ### Environment variables [](https://pydantic.dev/docs/ai/mcp/client/#environment-variables) The configuration file supports environment variable expansion using the `${VAR}` and `${VAR:-default}` syntax, [like Claude Code](https://code.claude.com/docs/en/mcp#environment-variable-expansion-in-mcp-json) . This is useful for keeping sensitive information like API keys or host names out of your configuration files: mcp\_config\_with\_env.json { "mcpServers": { "python-runner": { "command": "${PYTHON_CMD:-python3}", "args": ["run", "${MCP_MODULE}", "stdio"], "env": { "API_KEY": "${MY_API_KEY}" } }, "weather-api": { "url": "https://${SERVER_HOST:-localhost}:${SERVER_PORT:-8080}/sse" } } } When loading this configuration with [`load_mcp_toolsets()`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.load_mcp_toolsets) : * `${VAR}` references are replaced with the corresponding environment variable values. * `${VAR:-default}` references use the environment variable value if set, otherwise the default value. ### Usage [](https://pydantic.dev/docs/ai/mcp/client/#usage-1) mcp\_config\_loader.py from pydantic_ai import Agent from pydantic_ai.mcp import load_mcp_toolsets # Load all toolsets from the configuration file toolsets = load_mcp_toolsets('mcp_config.json') # Create an agent with all loaded toolsets agent = Agent('openai:gpt-5.2', toolsets=toolsets) async def main(): result = await agent.run('What is 7 plus 5?') print(result.output) Tool call customization ----------------------- [](https://pydantic.dev/docs/ai/mcp/client/#tool-call-customization) `MCPToolset` accepts a `process_tool_call` callback that lets you customize tool call requests and their responses. A common use case is to inject metadata that the server-side handler needs to read: mcp\_process\_tool\_call.py from typing import Any from fastmcp.client.transports import StdioTransport from pydantic_ai import Agent, RunContext from pydantic_ai.mcp import CallToolFunc, MCPToolset, ToolResult from pydantic_ai.models.test import TestModel async def process_tool_call( ctx: RunContext[int], call_tool: CallToolFunc, name: str, tool_args: dict[str, Any], ) -> ToolResult: """A tool call processor that passes along the deps.""" return await call_tool(name, tool_args, {'deps': ctx.deps}) toolset = MCPToolset( StdioTransport(command='python', args=['mcp_server.py']), process_tool_call=process_tool_call, ) agent = Agent( model=TestModel(call_tools=['echo_deps']), deps_type=int, toolsets=[toolset], ) async def main(): result = await agent.run('Echo with deps set to 42', deps=42) print(result.output) #> {"echo_deps":{"echo":"This is an echo message","deps":42}} How the server reads the injected metadata is MCP server SDK specific. For example, with the [MCP Python SDK](https://github.com/modelcontextprotocol/python-sdk) it’s accessible via the [`ctx: Context`](https://github.com/modelcontextprotocol/python-sdk#context) argument on tool handlers: mcp\_server.py from typing import Any from mcp.server.fastmcp import Context, FastMCP from mcp.server.session import ServerSession mcp = FastMCP('Pydantic AI MCP Server') @mcp.tool() async def echo_deps(ctx: Context[ServerSession, None]) -> dict[str, Any]: """Echo the run context. Args: ctx: Context object containing request and session information. Returns: Dictionary with an echo message and the deps. """ await ctx.info('This is an info message') deps: Any = getattr(ctx.request_context.meta, 'deps') return {'echo': 'This is an echo message', 'deps': deps} if __name__ == '__main__': mcp.run() Tool errors ----------- [](https://pydantic.dev/docs/ai/mcp/client/#tool-errors) When an MCP server reports a tool error, [`MCPToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset) lets you choose whether that error should ask the model to retry, appear as a failed tool result, or escape as an exception: | `tool_error_behavior` | Behavior | | --- | --- | | `'retry'` | Default. Raises [`ModelRetry`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ModelRetry)
, sending the server error back to the model as a retry prompt. Use this when the model may be able to correct the call. | | `'failed'` | Raises [`ToolFailed`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.ToolFailed)
, which is recorded as a tool result with `outcome='failed'`. Use this when the tool call is complete but failed, and the model should decide what to do next. | | `'error'` | Propagates the underlying MCP tool exception and fails the agent run. Use this for errors you want application code to handle outside the model loop. | Structured error content is serialized as JSON in the model-visible message for both `'retry'` and `'failed'`, so retryability hints and other machine-readable details remain available to the model. Protocol and transport errors are not reported as completed failed tool calls. This is the MCP equivalent of the [tool retry](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-retries) vs [failed tool result](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-failed) distinction in local tool code. Tool prefixes to avoid naming conflicts --------------------------------------- [](https://pydantic.dev/docs/ai/mcp/client/#tool-prefixes-to-avoid-naming-conflicts) When connecting to multiple MCP servers that might provide tools with the same name, wrap each `MCPToolset` with [`.prefixed(...)`](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.AbstractToolset.prefixed) to prepend a prefix to its tool names: mcp\_tool\_prefix.py from pydantic_ai import Agent from pydantic_ai.mcp import MCPToolset weather = MCPToolset('http://localhost:3001/sse').prefixed('weather') # `weather_*` calculator = MCPToolset('http://localhost:3002/sse').prefixed('calc') # `calc_*` # Both servers may expose a `get_data` tool, but they're disambiguated as # `weather_get_data` and `calc_get_data`. agent = Agent('openai:gpt-5.2', toolsets=[weather, calculator]) Server instructions ------------------- [](https://pydantic.dev/docs/ai/mcp/client/#server-instructions) MCP servers can provide instructions during initialization that give context about how to best interact with the server’s tools. These are accessible via [`MCPToolset.instructions`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset.instructions) after the connection is established, and can be automatically injected into the agent’s instructions by setting `include_instructions=True`: mcp\_server\_include\_instructions.py from pydantic_ai import Agent from pydantic_ai.mcp import MCPToolset toolset = MCPToolset('http://localhost:8000/mcp', include_instructions=True) agent = Agent('openai:gpt-5.2', toolsets=[toolset]) Tool metadata ------------- [](https://pydantic.dev/docs/ai/mcp/client/#tool-metadata) MCP tools can include metadata that provides additional information about the tool’s characteristics, which can be useful when [filtering tools](https://pydantic.dev/docs/ai/api/pydantic-ai/toolsets/#pydantic_ai.toolsets.FilteredToolset) . The `meta` and `annotations` fields can be found on the `metadata` dict on the [`ToolDefinition`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.ToolDefinition) object that’s passed to filter functions, and the tool’s output schema (if any) is available as the `return_schema` field. [`MCPToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset) additionally exposes a `task: bool` flag indicating whether the toolset will use [task-augmented execution](https://pydantic.dev/docs/ai/mcp/client/#background-tasks) for the tool. For tools where task support is optional, this reflects the `prefer_tasks` setting. Background tasks ---------------- [](https://pydantic.dev/docs/ai/mcp/client/#background-tasks) [`MCPToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset) supports MCP [task-augmented execution](https://modelcontextprotocol.io/specification/2025-11-25/basic/utilities/tasks) (SEP-1686). Servers using SEP-1686, including FastMCP 3 servers, can declare per-tool task support via `execution.taskSupport`, and `MCPToolset` routes calls accordingly: | `execution.taskSupport` | Behavior | | --- | --- | | `"required"` | Always calls with `task=True`. The server creates a task and the client awaits the final result via `tasks/result`. | | `"optional"` | Calls with `task=True` by default. Set [`prefer_tasks=False`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset.prefer_tasks)
to call normally instead. | | `"forbidden"` or absent | Calls normally. | FastMCP 4 uses the newer MCP [Tasks extension](https://tasks.extensions.modelcontextprotocol.io/seps/2663-tasks-extension) (SEP-2663), where the server directs task creation. The `task` metadata and `prefer_tasks` client preference above therefore apply to FastMCP 3, not FastMCP 4. An ordinary call drives a task-only tool to completion with nothing extra installed; explicitly selecting the tasks extension with `use_task=True` requires the separate `fastmcp-tasks` package, available via the `mcp-tasks` optional group: `pip install "pydantic-ai-slim[mcp-tasks]"`. For [FastMCP 3](https://gofastmcp.com/v3/servers/tasks) servers, install the tasks extra with `pip install "fastmcp[tasks]>=3,<4"` and declare task support per tool with `task=TaskConfig(mode=...)`: background\_task\_server.py from fastmcp import FastMCP from fastmcp.server.tasks import TaskConfig mcp = FastMCP('long_running_server') @mcp.tool(task=TaskConfig(mode='optional')) async def deep_research(topic: str) -> str: import asyncio await asyncio.sleep(0) return f'Researched {topic}' if __name__ == '__main__': mcp.run(transport='streamable-http') By default, [`MCPToolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset) uses task-augmented execution when a tool supports it. A client that prefers normal calls for tools where task support is optional can set [`prefer_tasks=False`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset.prefer_tasks) . This setting does not affect tools where task support is required: background\_task\_client.py from pydantic_ai import Agent from pydantic_ai.mcp import MCPToolset toolset = MCPToolset('http://localhost:8000/mcp', prefer_tasks=False) agent = Agent('openai:gpt-5.2', toolsets=[toolset]) Resources --------- [](https://pydantic.dev/docs/ai/mcp/client/#resources) MCP servers can provide [resources](https://modelcontextprotocol.io/docs/concepts/resources) — files, data, or content that can be accessed by the client. Resources in MCP are application-driven, with host applications determining how to incorporate context manually based on their needs. They are _not_ exposed to the LLM automatically (unless a tool returns a `ResourceLink` or `EmbeddedResource`). `MCPToolset` exposes methods to discover and read resources: * [`list_resources()`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset.list_resources) — list all available resources on the server * [`list_resource_templates()`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset.list_resource_templates) — list resource templates with parameter placeholders * [`read_resource(uri)`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset.read_resource) — read the contents of a specific resource by URI Text content is returned as `str`, and binary content as [`BinaryContent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BinaryContent) . Before consuming resources, we need to run a server that exposes some: mcp\_resource\_server.py from mcp.server.fastmcp import FastMCP mcp = FastMCP('Pydantic AI MCP Server') @mcp.resource('resource://user_name.txt', mime_type='text/plain') async def user_name_resource() -> str: return 'Alice' if __name__ == '__main__': mcp.run() Then we can read them from the client: mcp\_resources.py import asyncio from fastmcp.client.transports import StdioTransport from pydantic_ai.mcp import MCPToolset async def main(): toolset = MCPToolset(StdioTransport(command='python', args=['-m', 'mcp_resource_server'])) async with toolset: # List all available resources resources = await toolset.list_resources() for resource in resources: print(f' - {resource.name}: {resource.uri} ({resource.mime_type})') #> - user_name_resource: resource://user_name.txt (text/plain) # Read a text resource user_name = await toolset.read_resource('resource://user_name.txt') print(f'Text content: {user_name}') #> Text content: Alice if __name__ == '__main__': asyncio.run(main()) _(This example is complete, it can be run “as is”)_ HTTP authentication ------------------- [](https://pydantic.dev/docs/ai/mcp/client/#http-authentication) For HTTP transports, `MCPToolset` accepts an `auth` argument: a bearer token string, any [`httpx.Auth`](https://www.python-httpx.org/advanced/authentication/) , or the literal string `'oauth'` to enable [FastMCP’s OAuth flow](https://gofastmcp.com/clients/auth/oauth) . Static headers like API keys can be passed via the `headers` argument instead. ### Per-user authentication [](https://pydantic.dev/docs/ai/mcp/client/#per-user-authentication) In a multi-user or multi-tenant application, each user typically has their own credentials for the MCP server, such as a tenant-scoped bearer token. To make requests with the credentials of the user in question, each concurrent run needs its own `MCPToolset` instance so that it establishes its own authenticated session. The recommended way to do this is to build the toolset [dynamically](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/#dynamically-building-a-toolset) using the [`@agent.toolset`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.toolset) decorator: the decorated function is passed the [run context](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) and can read the user’s credentials from the run’s [dependencies](https://pydantic.dev/docs/ai/core-concepts/dependencies/) : mcp\_per\_user\_auth.py from dataclasses import dataclass from pydantic_ai import Agent, RunContext from pydantic_ai.mcp import MCPToolset @dataclass class UserDeps: mcp_token: str agent = Agent('openai:gpt-5.2', deps_type=UserDeps) @agent.toolset(per_run_step=False) # (1) def user_mcp_server(ctx: RunContext[UserDeps]) -> MCPToolset: return MCPToolset('http://localhost:8000/mcp', auth=ctx.deps.mcp_token) async def main(): result = await agent.run('What is 7 plus 5?', deps=UserDeps(mcp_token='')) print(result.output) #> The answer is 12. `per_run_step=False` builds the toolset once per run instead of ahead of each run step, so the whole run shares a single MCP session. _(This example is complete, it can be run “as is” — you’ll need to add `asyncio.run(main())` to run `main`)_ Because the per-run toolset’s session is established inside the run itself, credentials held in a `ContextVar` also resolve correctly with this pattern — but passing them through deps is more explicit and doesn’t depend on task-local state. As an alternative to a dynamic toolset, you can construct a new `MCPToolset` yourself for each request and pass it to the [`toolsets` argument](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) of the agent run methods. Custom TLS / SSL configuration ------------------------------ [](https://pydantic.dev/docs/ai/mcp/client/#custom-tls--ssl-configuration) In some environments you need to tweak how HTTPS connections are established — for example to trust an internal Certificate Authority, present a client certificate for **mTLS**, or (during local development only!) disable certificate verification altogether. `MCPToolset` exposes an `http_client` parameter so you can pass your own pre-configured [`httpx.AsyncClient`](https://www.python-httpx.org/async/) : mcp\_custom\_tls\_client.py import ssl import httpx from pydantic_ai import Agent from pydantic_ai.mcp import MCPToolset # Trust an internal / self-signed CA ssl_ctx = ssl.create_default_context(cafile='/etc/ssl/private/my_company_ca.pem') # Optional: load a client certificate for mutual TLS ssl_ctx.load_cert_chain(certfile='/etc/ssl/certs/client.crt', keyfile='/etc/ssl/private/client.key') http_client = httpx.AsyncClient(verify=ssl_ctx, timeout=httpx.Timeout(10.0)) toolset = MCPToolset('http://localhost:3001/sse', http_client=http_client) # (1) agent = Agent('openai:gpt-5.2', toolsets=[toolset]) When you supply `http_client`, Pydantic AI reuses this client for every request. Anything supported by **httpx** (`verify`, `cert`, custom proxies, timeouts, etc.) therefore applies to all MCP traffic. Client identification --------------------- [](https://pydantic.dev/docs/ai/mcp/client/#client-identification) When connecting to an MCP server, you can optionally specify an [Implementation](https://modelcontextprotocol.io/specification/2025-11-25/schema#implementation) object as client information that will be sent to the server during initialization. This is useful for: * Identifying your application in server logs * Allowing servers to provide custom behavior based on the client * Debugging and monitoring MCP connections * Version-specific feature negotiation mcp\_client\_with\_name.py from mcp import types as mcp_types from pydantic_ai.mcp import MCPToolset toolset = MCPToolset( 'http://localhost:3001/sse', client_info=mcp_types.Implementation( name='MyApplication', version='2.1.0', ), ) MCP sampling ------------ [](https://pydantic.dev/docs/ai/mcp/client/#mcp-sampling) Sampling diagram Here’s a mermaid diagram that may or may not make the data flow clearer: sequenceDiagram participant LLM participant MCP_Client as MCP client participant MCP_Server as MCP server MCP_Client->>LLM: LLM call LLM->>MCP_Client: LLM tool call response MCP_Client->>MCP_Server: tool call MCP_Server->>MCP_Client: sampling "create message" MCP_Client->>LLM: LLM call LLM->>MCP_Client: LLM text response MCP_Client->>MCP_Server: sampling response MCP_Server->>MCP_Client: tool call response Pydantic AI supports sampling as both a client and server. See the [server](https://pydantic.dev/docs/ai/mcp/server/#mcp-sampling) documentation for details on how to use sampling within a server. To use sampling as a client, an `MCPToolset` needs to have a [`sampling_model`](https://pydantic.dev/docs/ai/api/pydantic-ai/mcp/#pydantic_ai.mcp.MCPToolset.sampling_model) set. This can be done either directly on the toolset using the `sampling_model=` constructor keyword argument, or by using [`agent.set_mcp_sampling_model()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent.set_mcp_sampling_model) to use the agent’s model (or one specified as an argument) as the sampling model on all `MCPToolset`s registered with the agent. Let’s say we have an MCP server that wants to use sampling (in this case to generate an SVG as per the tool arguments): Sampling MCP server generate\_svg.py import re from pathlib import Path from mcp import SamplingMessage from mcp.server.fastmcp import Context, FastMCP from mcp.types import TextContent app = FastMCP() @app.tool() async def image_generator(ctx: Context, subject: str, style: str) -> str: prompt = f'{subject=} {style=}' # `ctx.session.create_message` is the sampling call result = await ctx.session.create_message( [SamplingMessage(role='user', content=TextContent(type='text', text=prompt))], max_tokens=1_024, system_prompt='Generate an SVG image as per the user input', ) assert isinstance(result.content, TextContent) path = Path(f'{subject}_{style}.svg') # remove triple backticks if the svg was returned within markdown if m := re.search(r'^```\w*$(.+?)```$', result.content.text, re.S | re.M): path.write_text(m.group(1), encoding='utf-8') else: path.write_text(result.content.text, encoding='utf-8') return f'See {path}' if __name__ == '__main__': # run the server via stdio app.run() Using this server with an `Agent` will automatically allow sampling: sampling\_mcp\_client.py from fastmcp.client.transports import StdioTransport from pydantic_ai import Agent from pydantic_ai.mcp import MCPToolset toolset = MCPToolset(StdioTransport(command='python', args=['generate_svg.py'])) agent = Agent('openai:gpt-5.2', toolsets=[toolset]) async def main(): agent.set_mcp_sampling_model() result = await agent.run('Create an image of a robot in a punk style.') print(result.output) #> Image file written to robot_punk.svg. _(This example is complete, it can be run “as is”)_ Elicitation ----------- [](https://pydantic.dev/docs/ai/mcp/client/#elicitation) In MCP, [elicitation](https://modelcontextprotocol.io/docs/concepts/elicitation) allows a server to request [structured input](https://modelcontextprotocol.io/specification/2025-06-18/client/elicitation#supported-schema-types) from the client for missing or additional context during a session. Elicitation lets models essentially say “Hold on — I need to know X before I can continue”, rather than requiring everything upfront or taking a shot in the dark. ### How elicitation works [](https://pydantic.dev/docs/ai/mcp/client/#how-elicitation-works) Elicitation introduces a protocol message type called [`ElicitRequest`](https://modelcontextprotocol.io/specification/2025-06-18/schema#elicitrequest) , which is sent from the server to the client when it needs additional information. The client can then respond with an [`ElicitResult`](https://modelcontextprotocol.io/specification/2025-06-18/schema#elicitresult) or an `ErrorData` message. A typical interaction looks like this: * User makes a request to the MCP server (e.g. “Book a table at that Italian place”) * The server identifies that it needs more information (e.g. “Which Italian place?”, “What date and time?”) * The server sends an `ElicitRequest` to the client asking for the missing information. * The client receives the request, presents it to the user (e.g. via a terminal prompt, GUI dialog, or web interface). * User provides the requested information, declines, or cancels. * The client sends an `ElicitResult` back to the server with the user’s response. * With the structured data, the server can continue processing the original request. This allows for a more interactive and user-friendly experience, especially for multi-stage workflows. Instead of requiring all information upfront, the server can ask for it as needed. ### Setting up elicitation [](https://pydantic.dev/docs/ai/mcp/client/#setting-up-elicitation) To enable elicitation, provide an `elicitation_handler` when creating your `MCPToolset`: restaurant\_server.py from mcp.server.fastmcp import Context, FastMCP from pydantic import BaseModel, Field mcp = FastMCP(name='Restaurant Booking') class BookingDetails(BaseModel): """Schema for restaurant booking information.""" restaurant: str = Field(description='Choose a restaurant') party_size: int = Field(description='Number of people', ge=1, le=8) date: str = Field(description='Reservation date (DD-MM-YYYY)') @mcp.tool() async def book_table(ctx: Context) -> str: """Book a restaurant table with user input.""" # Ask user for booking details using Pydantic schema result = await ctx.elicit(message='Please provide your booking details:', schema=BookingDetails) if result.action == 'accept' and result.data: booking = result.data return f'✅ Booked table for {booking.party_size} at {booking.restaurant} on {booking.date}' elif result.action == 'decline': return 'No problem! Maybe another time.' else: # cancel return 'Booking cancelled.' if __name__ == '__main__': mcp.run(transport='stdio') This server demonstrates elicitation by requesting structured booking details from the client when the `book_table` tool is called. Here’s how to wire up the matching client: client\_example.py import asyncio from fastmcp.client.transports import StdioTransport from mcp.types import ElicitRequestParams, ElicitResult from pydantic_ai import Agent from pydantic_ai.mcp import MCPToolset async def handle_elicitation(context, params: ElicitRequestParams) -> ElicitResult: """Handle elicitation requests from MCP server.""" print(f'\n{params.message}') if not params.requestedSchema: response = input('Response: ') return ElicitResult(action='accept', content={'response': response}) # Collect data for each field properties = params.requestedSchema['properties'] data = {} for field, info in properties.items(): description = info.get('description', field) value = input(f'{description}: ') # Convert to proper type based on JSON schema if info.get('type') == 'integer': data[field] = int(value) else: data[field] = value # Confirm confirm = input('\nConfirm booking? (y/n/c): ').lower() if confirm == 'y': print('Booking details:', data) return ElicitResult(action='accept', content=data) elif confirm == 'n': return ElicitResult(action='decline') else: return ElicitResult(action='cancel') toolset = MCPToolset( StdioTransport(command='python', args=['restaurant_server.py']), elicitation_handler=handle_elicitation, ) agent = Agent('openai:gpt-5.2', toolsets=[toolset]) async def main(): """Run the agent to book a restaurant table.""" result = await agent.run('Book me a table') print(f'\nResult: {result.output}') if __name__ == '__main__': asyncio.run(main()) ### Supported schema types [](https://pydantic.dev/docs/ai/mcp/client/#supported-schema-types) MCP elicitation supports string, number, boolean, and enum types with flat object structures only. These limitations ensure reliable cross-client compatibility. See [supported schema types](https://modelcontextprotocol.io/specification/2025-06-18/client/elicitation#supported-schema-types) for details. ### Security [](https://pydantic.dev/docs/ai/mcp/client/#security) MCP elicitation requires careful handling — servers must not request sensitive information, and clients must implement user approval controls with clear explanations. See [security considerations](https://modelcontextprotocol.io/specification/2025-06-18/client/elicitation#security-considerations) for details. Was this page helpful? Thanks for your feedback! --- # native_tools | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#_top) native\_tools ============= AbstractNativeTool ------------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.AbstractNativeTool) **Bases:** `ABC` A native tool that can be used by an agent. This class is abstract and cannot be instantiated directly. The native tools are passed to the model as part of the `ModelRequestParameters`. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#attributes) #### kind [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.AbstractNativeTool.kind) Native tool identifier, this should be available on all native tools as a discriminator. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'unknown_native_tool'` #### label [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.AbstractNativeTool.label) Human-readable label for UI display. Subclasses should override this to provide a meaningful label. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### optional [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.AbstractNativeTool.optional) Whether this instance is a best-effort upgrade rather than a hard requirement. When `True`, the instance is silently dropped from the request on a model that doesn’t support it natively, instead of raising when no local fallback is provided. Use for native tools where a fallback path exists (e.g. a local function tool that takes over when the native one isn’t available). When `False` (the default), the request errors on models that can’t honor the native tool — the user explicitly asked for it, so fail loudly rather than silently substituting different behavior. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### unique\_id [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.AbstractNativeTool.unique_id) A unique identifier for the native tool. If multiple instances of the same native tool can be passed to the model, subclasses should override this property to allow them to be distinguished. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) AdvisorTool ----------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.AdvisorTool) **Bases:** `AbstractNativeTool` A native tool that lets a faster executor model consult a stronger advisor model mid-generation. The fields map 1:1 to the parameters of Anthropic’s advisor tool definition. OpenRouter exposes the advisor as a gateway server tool that honors a subset (`model`, `max_tokens`) and ignores the unsupported fields; see the per-field docstrings for which provider supports each. Supported by: * Anthropic * OpenRouter ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#attributes-1) #### caching [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.AdvisorTool.caching) If provided, caches the advisor context ephemerally with the given TTL. Maps to `caching={'type': 'ephemeral', 'ttl': ...}`. Supported by: * Anthropic OpenRouter’s advisor tool has no equivalent knob and ignores `caching`. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘5m’, ‘1h’\] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### kind [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.AdvisorTool.kind) The kind of tool. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'advisor'` #### max\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.AdvisorTool.max_tokens) If provided, caps the advisor’s output tokens (minimum 1024). Maps to `max_tokens` on Anthropic and `max_completion_tokens` on OpenRouter. When set, the Anthropic advisor result carries a `stop_reason`. Supported by: * Anthropic * OpenRouter **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### max\_uses [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.AdvisorTool.max_uses) If provided, the advisor can be consulted at most this many times per request. Maps to `max_uses`. This is a per-request cap, not a per-run budget: a run that spans multiple requests resets the count each request. Enforce a conversation-wide ceiling yourself if you need one. Supported by: * Anthropic OpenRouter caps advisor consultations per request with a fixed gateway limit and ignores `max_uses`. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### model [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.AdvisorTool.model) The advisor model to consult, i.e. the `model` field of the provider’s advisor tool definition. The executor/advisor pairing is validated by the provider’s API, not here. The accepted namespace depends on the executing provider: Anthropic model IDs (e.g. `claude-opus-4-8`) on Anthropic, OpenRouter catalog slugs (e.g. `anthropic/claude-opus-4.8`) on OpenRouter. Supported by: * Anthropic * OpenRouter **Type:** `AdvisorModelName` CodeExecutionTool ----------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.CodeExecutionTool) **Bases:** `AbstractNativeTool` A native tool that allows your agent to execute code. Supported by: * Anthropic * OpenAI Responses * Google * Bedrock (Nova2.0) * xAI ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#attributes-2) #### files [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.CodeExecutionTool.files) Uploaded files to make available in the code execution environment. Only files matching the model provider are used; files from other providers are ignored. Supported by: * Anthropic * OpenAI Responses **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`UploadedFile`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.UploadedFile)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### kind [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.CodeExecutionTool.kind) The kind of tool. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'code_execution'` FileSearchTool -------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.FileSearchTool) **Bases:** `AbstractNativeTool` A native tool that allows your agent to search through uploaded files using vector search. This tool provides a fully managed Retrieval-Augmented Generation (RAG) system that handles file storage, chunking, embedding generation, and context injection into prompts. Supported by: * OpenAI Responses * Google (Gemini) * xAI (mapped to collections search) ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#attributes-3) #### file\_store\_ids [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.FileSearchTool.file_store_ids) The file store IDs to search through. For OpenAI, these are the IDs of vector stores created via the OpenAI API. For Google, these are file search store names that have been uploaded and processed via the Gemini Files API. For xAI, these are collection IDs for the xAI collections search tool. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] #### instructions [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.FileSearchTool.instructions) Optional instructions that guide how the collections search results are interpreted and ranked. Supported by: * xAI **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### kind [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.FileSearchTool.kind) The kind of tool. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'file_search'` #### max\_num\_results [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.FileSearchTool.max_num_results) The maximum number of results to return. Supported by: * xAI (mapped to collections search `limit`, defaults to 10 server-side) **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### retrieval\_mode [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.FileSearchTool.retrieval_mode) The retrieval strategy for the search. Supported by: * xAI (defaults to `hybrid` server-side) **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘hybrid’, ‘semantic’, ‘keyword’\] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` ImageGenerationTool ------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.ImageGenerationTool) **Bases:** `AbstractNativeTool` A native tool that allows your agent to generate images. Supported by: * OpenAI Responses * Google ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#attributes-4) #### action [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.ImageGenerationTool.action) Whether to generate a new image or edit an existing image. Supported by: * OpenAI Responses. Default: ‘auto’. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘generate’, ‘edit’, ‘auto’\] **Default:** `'auto'` #### aspect\_ratio [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.ImageGenerationTool.aspect_ratio) The aspect ratio to use for generated images. Supported by: * Google image-generation models (Gemini) * OpenAI Responses (maps ‘1:1’, ‘2:3’, and ‘3:2’ to supported sizes) **Type:** `ImageAspectRatio` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### background [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.ImageGenerationTool.background) Background type for the generated image. Supported by: * OpenAI Responses. ‘transparent’ is only supported for ‘png’ and ‘webp’ output formats. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘transparent’, ‘opaque’, ‘auto’\] **Default:** `'auto'` #### input\_fidelity [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.ImageGenerationTool.input_fidelity) Control how much effort the model will exert to match the style and features, especially facial features, of input images. Supported by: * OpenAI Responses. Default: ‘low’. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘high’, ‘low’\] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### kind [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.ImageGenerationTool.kind) The kind of tool. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'image_generation'` #### model [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.ImageGenerationTool.model) The image generation model to use. Supported by: * OpenAI Responses. Defaults to the provider’s image generation model selection. Known image generation models include `gpt-image-2`, `gpt-image-1.5`, `gpt-image-1`, and `gpt-image-1-mini`. This selects the underlying image generation model used by the tool; it does not change the agent’s conversational model. **Type:** `ImageGenerationModelName` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### moderation [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.ImageGenerationTool.moderation) Moderation level for the generated image. Supported by: * OpenAI Responses **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘auto’, ‘low’\] **Default:** `'auto'` #### output\_compression [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.ImageGenerationTool.output_compression) Compression level for the output image. Supported by: * OpenAI Responses. Only supported for ‘jpeg’ and ‘webp’ output formats. Default: 100. * Google (Vertex AI only). Only supported for ‘jpeg’ output format. Default: 75. Setting this will default `output_format` to ‘jpeg’ if not specified. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### output\_format [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.ImageGenerationTool.output_format) The output format of the generated image. Supported by: * OpenAI Responses. Default: ‘png’. * Google (Vertex AI only). Default: ‘png’, or ‘jpeg’ if `output_compression` is set. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘png’, ‘webp’, ‘jpeg’\] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### partial\_images [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.ImageGenerationTool.partial_images) Number of partial images to generate in streaming mode. Supported by: * OpenAI Responses. Supports 0 to 3. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `0` #### quality [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.ImageGenerationTool.quality) The quality of the generated image. Supported by: * OpenAI Responses **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘low’, ‘medium’, ‘high’, ‘auto’\] **Default:** `'auto'` #### size [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.ImageGenerationTool.size) The size of the generated image. * OpenAI Responses: ‘auto’ (default: model selects the size based on the prompt), ‘1024x1024’, ‘1024x1536’, ‘1536x1024’ * Google (Gemini 3 Pro Image and later): ‘512’ (Gemini 3.1 Flash Image only), ‘1K’ (default), ‘2K’, ‘4K’ **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘auto’, ‘1024x1024’, ‘1024x1536’, ‘1536x1024’, ‘512’, ‘1K’, ‘2K’, ‘4K’\] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` MCPServerTool ------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.MCPServerTool) **Bases:** `AbstractNativeTool` A native tool that allows your agent to use MCP servers. Supported by: * OpenAI Responses * Anthropic * xAI ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#attributes-5) #### allowed\_tools [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.MCPServerTool.allowed_tools) A list of tools that the MCP server can use. Supported by: * OpenAI Responses * Anthropic * xAI **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### authorization\_token [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.MCPServerTool.authorization_token) Authorization header to use when making requests to the MCP server. Supported by: * OpenAI Responses * Anthropic * xAI **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### description [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.MCPServerTool.description) A description of the MCP server. Supported by: * OpenAI Responses * xAI **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### headers [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.MCPServerTool.headers) Optional HTTP headers to send to the MCP server. Use for authentication or other purposes. Supported by: * OpenAI Responses * xAI **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### id [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.MCPServerTool.id) A unique identifier for the MCP server. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### url [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.MCPServerTool.url) The URL of the MCP server to use. For OpenAI Responses, it is possible to use `connector_id` by providing it as `x-openai-connector:`. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) MemoryTool ---------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.MemoryTool) **Bases:** `AbstractNativeTool` A native tool that allows your agent to use memory. Supported by: * Anthropic ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#attributes-6) #### kind [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.MemoryTool.kind) The kind of tool. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'memory'` WebFetchTool ------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebFetchTool) **Bases:** `AbstractNativeTool` Allows your agent to access contents from URLs. The parameters that PydanticAI passes depend on the model, as some parameters may not be supported by certain models. Supported by: * Anthropic * Google ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#attributes-7) #### allowed\_domains [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebFetchTool.allowed_domains) If provided, only these domains will be fetched. With Anthropic, you can only use one of `blocked_domains` or `allowed_domains`, not both. Supported by: * Anthropic, see [https://docs.anthropic.com/en/docs/agents-and-tools/tool-use/web-fetch-tool#domain-filtering](https://docs.anthropic.com/en/docs/agents-and-tools/tool-use/web-fetch-tool#domain-filtering) **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### blocked\_domains [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebFetchTool.blocked_domains) If provided, these domains will never be fetched. With Anthropic, you can only use one of `blocked_domains` or `allowed_domains`, not both. Supported by: * Anthropic, see [https://docs.anthropic.com/en/docs/agents-and-tools/tool-use/web-fetch-tool#domain-filtering](https://docs.anthropic.com/en/docs/agents-and-tools/tool-use/web-fetch-tool#domain-filtering) **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### enable\_citations [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebFetchTool.enable_citations) If True, enables citations for fetched content. Supported by: * Anthropic **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### kind [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebFetchTool.kind) The kind of tool. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'web_fetch'` #### max\_content\_tokens [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebFetchTool.max_content_tokens) Maximum content length in tokens for fetched content. Supported by: * Anthropic **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### max\_uses [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebFetchTool.max_uses) If provided, the tool will stop fetching URLs after the given number of uses. Supported by: * Anthropic **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` WebSearchTool ------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebSearchTool) **Bases:** `AbstractNativeTool` A native tool that allows your agent to search the web for information. The parameters that PydanticAI passes depend on the model, as some parameters may not be supported by certain models. OpenRouter uses its Beta web-search server tool, which lets the model decide whether to search and how often. It accepts the portable settings below, though their effect depends on OpenRouter’s selected search engine and downstream provider. Supported by: * Anthropic * OpenAI Responses * Groq * Google * xAI * OpenRouter ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#attributes-8) #### allowed\_domains [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebSearchTool.allowed_domains) If provided, only these domains will be included in results. With Anthropic, you can only use one of `blocked_domains` or `allowed_domains`, not both. Supported by: * Anthropic, see [https://docs.anthropic.com/en/docs/build-with-claude/tool-use/web-search-tool#domain-filtering](https://docs.anthropic.com/en/docs/build-with-claude/tool-use/web-search-tool#domain-filtering) * Groq, see [https://console.groq.com/docs/agentic-tooling#search-settings](https://console.groq.com/docs/agentic-tooling#search-settings) * OpenAI Responses, see [https://platform.openai.com/docs/guides/tools-web-search](https://platform.openai.com/docs/guides/tools-web-search) * xAI, see [https://docs.x.ai/docs/guides/tools/search-tools#web-search-parameters](https://docs.x.ai/docs/guides/tools/search-tools#web-search-parameters) * OpenRouter, see [https://openrouter.ai/docs/guides/features/server-tools/web-search#configuration](https://openrouter.ai/docs/guides/features/server-tools/web-search#configuration) **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### blocked\_domains [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebSearchTool.blocked_domains) If provided, these domains will never appear in results. With Anthropic, you can only use one of `blocked_domains` or `allowed_domains`, not both. Supported by: * Anthropic, see [https://docs.anthropic.com/en/docs/build-with-claude/tool-use/web-search-tool#domain-filtering](https://docs.anthropic.com/en/docs/build-with-claude/tool-use/web-search-tool#domain-filtering) * Groq, see [https://console.groq.com/docs/agentic-tooling#search-settings](https://console.groq.com/docs/agentic-tooling#search-settings) * xAI, see [https://docs.x.ai/docs/guides/tools/search-tools#web-search-parameters](https://docs.x.ai/docs/guides/tools/search-tools#web-search-parameters) * OpenRouter, see [https://openrouter.ai/docs/guides/features/server-tools/web-search#configuration](https://openrouter.ai/docs/guides/features/server-tools/web-search#configuration) **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### external\_web\_access [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebSearchTool.external_web_access) Whether the hosted web search tool may fetch live web content. If `False`, the tool uses only cached or indexed results. If `None`, the parameter is omitted and the provider default is used. OpenAI currently defaults to `True`. Supported by: * OpenAI Responses `web_search` tool, see [https://developers.openai.com/api/docs/guides/tools-web-search#live-internet-access](https://developers.openai.com/api/docs/guides/tools-web-search#live-internet-access) OpenAI’s legacy `web_search_preview` tool ignores this parameter. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### kind [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebSearchTool.kind) The kind of tool. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'web_search'` #### max\_uses [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebSearchTool.max_uses) If provided, the tool will stop searching the web after the given number of uses. For OpenRouter, this limit is enforced with a non-native search engine or Anthropic’s native search. Other native providers ignore it. Supported by: * Anthropic * OpenRouter **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### search\_context\_size [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebSearchTool.search_context_size) The `search_context_size` parameter controls how much context is retrieved from the web to help the tool formulate a response. Supported by: * OpenAI Responses * OpenRouter **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘low’, ‘medium’, ‘high’\] **Default:** `'medium'` #### user\_location [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebSearchTool.user_location) The `user_location` parameter allows you to localize search results based on a user’s location. Supported by: * Anthropic * OpenAI Responses * xAI, see [https://docs.x.ai/docs/guides/tools/search-tools#web-search-parameters](https://docs.x.ai/docs/guides/tools/search-tools#web-search-parameters) * OpenRouter, see [https://openrouter.ai/docs/guides/features/server-tools/web-search#configuration](https://openrouter.ai/docs/guides/features/server-tools/web-search#configuration) **Type:** [`WebSearchUserLocation`](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebSearchUserLocation) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` WebSearchUserLocation --------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebSearchUserLocation) **Bases:** [`TypedDict`](https://docs.python.org/3/library/typing.html#typing.TypedDict) Allows you to localize search results based on a user’s location. Supported by: * Anthropic * OpenAI Responses * xAI * OpenRouter ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#attributes-9) #### city [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebSearchUserLocation.city) The city where the user is located. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### country [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebSearchUserLocation.country) The country where the user is located. For OpenAI and xAI, this must be a 2-letter ISO 3166-1 alpha-2 country code (e.g., ‘US’, ‘GB’). **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### region [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebSearchUserLocation.region) The region or state where the user is located. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### timezone [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.WebSearchUserLocation.timezone) The timezone of the user’s location. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) XSearchTool ----------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.XSearchTool) **Bases:** `AbstractNativeTool` A native tool that allows your agent to search X/Twitter for posts and content. See [https://docs.x.ai/developers/tools/x-search](https://docs.x.ai/developers/tools/x-search) for more details. When used via the [`XSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.XSearch) capability with a `fallback_model` set, this tool also works with non-xAI models by delegating to a subagent running the specified xAI model. Supported by: * xAI ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#attributes-10) #### allowed\_x\_handles [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.XSearchTool.allowed_x_handles) If provided, only posts from these X handles will be included (max 20). Supported by: * xAI, see [https://docs.x.ai/developers/tools/x-search](https://docs.x.ai/developers/tools/x-search) **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### enable\_image\_understanding [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.XSearchTool.enable_image_understanding) Enable image analysis from X posts. Supported by: * xAI, see [https://docs.x.ai/developers/tools/x-search](https://docs.x.ai/developers/tools/x-search) **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### enable\_video\_understanding [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.XSearchTool.enable_video_understanding) Enable video analysis from X content. Supported by: * xAI, see [https://docs.x.ai/developers/tools/x-search](https://docs.x.ai/developers/tools/x-search) **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### excluded\_x\_handles [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.XSearchTool.excluded_x_handles) If provided, posts from these X handles will be excluded (max 20). Supported by: * xAI, see [https://docs.x.ai/developers/tools/x-search](https://docs.x.ai/developers/tools/x-search) **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### from\_date [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.XSearchTool.from_date) If provided, only posts created on or after this datetime will be included. Naive datetimes are interpreted as UTC by the xAI API. Supported by: * xAI, see [https://docs.x.ai/developers/tools/x-search](https://docs.x.ai/developers/tools/x-search) **Type:** [`datetime`](https://docs.python.org/3/library/datetime.html#module-datetime) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### include\_output [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.XSearchTool.include_output) Include raw X search results in the response as [`NativeToolReturnPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.NativeToolReturnPart) . Without this, the model uses the search results internally but only returns its text summary. Enabling it gives programmatic access to searched posts, sources, and metadata. Can also be set via [`XaiModelSettings.xai_include_x_search_output`](https://pydantic.dev/docs/ai/api/models/xai/#pydantic_ai.models.xai.XaiModelSettings.xai_include_x_search_output) . Supported by: * xAI, see [https://docs.x.ai/developers/tools/x-search](https://docs.x.ai/developers/tools/x-search) **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### kind [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.XSearchTool.kind) The kind of tool. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `'x_search'` #### to\_date [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.XSearchTool.to_date) If provided, only posts created on or before this datetime will be included. Naive datetimes are interpreted as UTC by the xAI API. Supported by: * xAI, see [https://docs.x.ai/developers/tools/x-search](https://docs.x.ai/developers/tools/x-search) **Type:** [`datetime`](https://docs.python.org/3/library/datetime.html#module-datetime) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` AdvisorModelName ---------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.AdvisorModelName) Known Anthropic advisor model names, or any other model ID string. These are the models Anthropic currently accepts as the _advisor_ — the stronger model an executor consults mid-generation. The executor/advisor pairing is validated by the API, not here. The literals are Anthropic model IDs; on OpenRouter, pass a catalog slug string instead (e.g. `anthropic/claude-opus-4.8` or the `~anthropic/claude-opus-latest` alias). **Default:** `Literal['claude-fable-5', 'claude-mythos-5', 'claude-opus-5', 'claude-opus-4-8', 'claude-opus-4-7', 'claude-opus-4-6', 'claude-sonnet-4-6'] | str` ImageAspectRatio ---------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.ImageAspectRatio) Supported aspect ratios for image generation tools. **Default:** `Literal['21:9', '16:9', '4:3', '3:2', '1:1', '9:16', '3:4', '2:3', '5:4', '4:5']` ImageGenerationModelName ------------------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.ImageGenerationModelName) Known OpenAI image generation model names, or another OpenAI image model ID. **Default:** `Literal['gpt-image-2', 'gpt-image-1.5', 'gpt-image-1', 'gpt-image-1-mini'] | str` NATIVE\_TOOL\_TYPES ------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.NATIVE_TOOL_TYPES) Registry of all native tool types, keyed by their kind string. This dict is populated automatically via `__init_subclass__` when tool classes are defined. **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`type`](https://docs.python.org/3/glossary.html#term-type)\ \[`AbstractNativeTool`\]\] **Default:** `{}` SUPPORTED\_NATIVE\_TOOLS ------------------------ [](https://pydantic.dev/docs/ai/api/pydantic-ai/native_tools/#pydantic_ai.native_tools.SUPPORTED_NATIVE_TOOLS) Set of all native tool types. **Default:** `frozenset(NATIVE_TOOL_TYPES.values())` Was this page helpful? Thanks for your feedback! --- # run | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#_top) run === AgentRun -------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun) **Bases:** `Generic[AgentDepsT, OutputDataT]` A stateful, async-iterable run of an [`Agent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent) . You generally obtain an `AgentRun` instance by calling `async with my_agent.iter(...) as agent_run:`. Once you have an instance, you can use it to iterate through the run’s nodes as they execute. When an [`End`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.End) is reached, the run finishes and [`result`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.result) becomes available. Example: from pydantic_ai import Agent agent = Agent('openai:gpt-5.2') async def main(): nodes = [] # Iterate through the run, recording each node along the way: async with agent.iter('What is the capital of France?') as agent_run: async for node in agent_run: nodes.append(node) print(nodes) ''' [\ UserPromptNode(\ user_prompt='What is the capital of France?',\ instructions_functions=[],\ system_prompts=(),\ system_prompt_functions=[],\ system_prompt_dynamic_functions={},\ ),\ ModelRequestNode(\ request=ModelRequest(\ parts=[\ UserPromptPart(\ content='What is the capital of France?',\ timestamp=datetime.datetime(...),\ )\ ],\ timestamp=datetime.datetime(...),\ run_id='...',\ conversation_id='...',\ )\ ),\ CallToolsNode(\ model_response=ModelResponse(\ parts=[TextPart(content='The capital of France is Paris.')],\ usage=RequestUsage(\ cost=Decimal('0.000196'), input_tokens=56, output_tokens=7\ ),\ model_name='gpt-5.2',\ timestamp=datetime.datetime(...),\ run_id='...',\ conversation_id='...',\ )\ ),\ End(data=FinalResult(output='The capital of France is Paris.')),\ ] ''' print(agent_run.result.output) #> The capital of France is Paris. You can also manually drive the iteration using the [`next`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.next) method for more granular control. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#attributes) #### conversation\_id [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.conversation_id) The unique identifier for the conversation this run belongs to. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### ctx [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.ctx) The current context of the agent run. **Type:** `GraphRunContext`\[`_agent_graph.GraphAgentState`, `_agent_graph.GraphAgentDeps`\[`AgentDepsT`, [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \]\] #### metadata [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.metadata) Metadata associated with this agent run, if configured. **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) #### next\_node [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.next_node) The next node that will be run in the agent graph. This is the next node that will be used during async iteration, or if a node is not passed to `self.next(...)`. **Type:** `_agent_graph.AgentNode`\[`AgentDepsT`, `OutputDataT`\] | `End`\[`FinalResult`\[`OutputDataT`\]\] #### pending\_messages [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.pending_messages) Internal: live view of the queue mutated by `enqueue` and drained by the internal `PendingMessageDrainCapability`. Exposed for inspection / debugging; use [`enqueue`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.enqueue) to add messages. **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[`PendingMessage`\] #### result [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.result) The final result of the run if it has ended, otherwise `None`. Once the run returns an [`End`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.End) node, `result` is populated with an [`AgentRunResult`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult) . **Type:** [`AgentRunResult`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult) \[`OutputDataT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) #### run\_id [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.run_id) The unique identifier for the agent run. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### usage [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.usage) Get usage statistics for the run so far, including token usage, model requests, and so on. **Type:** `_usage.RunUsage` ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#methods) #### \_\_aiter\_\_ [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.__aiter__) def __aiter__( ) -> AsyncIterator[_agent_graph.AgentNode[AgentDepsT, OutputDataT] | End[FinalResult[OutputDataT]]] Provide async-iteration over the nodes in the agent run. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#returns) [`AsyncIterator`](https://docs.python.org/3/library/typing.html#typing.AsyncIterator) \[`_agent_graph.AgentNode`\[`AgentDepsT`, `OutputDataT`\] | `End`\[`FinalResult`\[`OutputDataT`\]\]\] #### \_\_anext\_\_ [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.__anext__) `@async` def __anext__( ) -> _agent_graph.AgentNode[AgentDepsT, OutputDataT] | End[FinalResult[OutputDataT]] Advance to the next node automatically based on the last returned node. Yields each node before it runs, ending with the [`End`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.End) node. Advancing goes through [`next()`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.next) , so capability hooks fire exactly as they do for [`agent.run()`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AbstractAgent.run) . ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#returns-1) `_agent_graph.AgentNode`\[`AgentDepsT`, `OutputDataT`\] | `End`\[`FinalResult`\[`OutputDataT`\]\] #### all\_messages [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.all_messages) def all_messages() -> list[_messages.ModelMessage] Return all messages for the run so far. Messages from older runs are included. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#returns-2) [`list`](https://docs.python.org/3/glossary.html#term-list) \[`_messages.ModelMessage`\] #### all\_messages\_json [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.all_messages_json) def all_messages_json(*, output_tool_return_content: str | None = None) -> bytes Return all messages from [`all_messages`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.all_messages) as JSON bytes. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#returns-3) [`bytes`](https://docs.python.org/3/library/stdtypes.html#bytes) — JSON bytes representing the messages. #### cancel [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.cancel) def cancel() -> None Cancel the whole agent run. The run stops what it is doing — the in-flight model request is torn down, in-flight tool tasks are cancelled and drained, a suspended server-side job is best-effort cancelled — and the code driving the run sees `asyncio.CancelledError`. When the `agent.iter()` context exits, this becomes [`RunCancelled`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled) (including for `agent.run()`, which wraps `iter()`). Everything that completed before the cancellation took effect is preserved in message history. [`RunCancelled.all_messages()`](https://pydantic.dev/docs/ai/api/pydantic-ai/exceptions/#pydantic_ai.exceptions.RunCancelled.all_messages) returns a complete snapshot that can be passed to a new run as `message_history` to resume the conversation. Cancellation is terminal: capability hooks (`wrap_run`, `wrap_node_run`, `on_run_error`) may observe it and clean up, but cannot recover the run into a successful result. Unlike [`StreamedRunResult.cancel()`](https://pydantic.dev/docs/ai/api/pydantic-ai/result/#pydantic_ai.result.StreamedRunResult.cancel) , which only stops the current model response and lets the run continue, this ends the run itself. Safe to call from another task or thread (e.g. a TUI’s key handler while the run is awaited elsewhere). Idempotent; a no-op once the run has finished — where “finished” means the `agent.iter()`/`agent.run()` context has exited. A `cancel()` issued inside the context after the run has already produced its result (e.g. after iterating to `End`) is still honored and surfaces as `RunCancelled` on exit, so that a hook running at context exit (like `after_run`) can still cancel the run; only after the context has exited is `cancel()` a true no-op. Externally cancelling the task running the agent (`asyncio.Task.cancel()`) remains supported and keeps raising `asyncio.CancelledError` instead; when both happen, the external cancellation wins. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#returns-4) [`None`](https://docs.python.org/3/library/constants.html#None) #### enqueue [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.enqueue) def enqueue( *content: EnqueueContent, priority: PendingMessagePriority = 'asap', ) -> str | None Enqueue content to be injected into the conversation. Designed to be called from the same event loop driving `agent.iter()`. If you’re forwarding events from a different thread (e.g. a webhook handler running on its own loop or thread), marshal the call back onto the agent’s loop first (e.g. `loop.call_soon_threadsafe(agent_run.enqueue, msg)`). The drain’s `queue[:] = remaining` pattern in `_drain_by_priority` isn’t atomic against concurrent appends from a different thread. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#returns-5) [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) — The `enqueue_id` of the queued message, echoed on the [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) — [`EnqueuedMessagesEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.EnqueuedMessagesEvent) emitted when it’s [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) — delivered, or `None` when there was nothing to enqueue (an empty call). ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#parameters) **`*content`** : `EnqueueContent` _Default:_ `()` [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.enqueue(*content)) One or more [`EnqueueContent`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.EnqueueContent) items. Adjacent [`UserContent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.UserContent) (a `str` or multi-modal content like an [`ImageUrl`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ImageUrl) ) is gathered into one [`UserPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.UserPromptPart) , and each [`ModelRequestPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequestPart) (e.g. a [`SystemPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SystemPromptPart) ) is coalesced with adjacent part-style items into one [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) ; a complete [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) or [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) is kept as its own message. The assembled sequence must end in a request. Calling with no positional args is a no-op. **`priority`** : `PendingMessagePriority` _Default:_ `'asap'` [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.enqueue(priority)) When to deliver: `'asap'` (default) — at the earliest opportunity (next model request, or a redirect if the agent would otherwise end). `'when_idle'` — only when the agent would otherwise end, after `'asap'` messages. #### new\_messages [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.new_messages) def new_messages() -> list[_messages.ModelMessage] Return the messages produced during this run so far. Messages provided via `message_history` and messages from older runs are excluded. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#returns-6) [`list`](https://docs.python.org/3/glossary.html#term-list) \[`_messages.ModelMessage`\] #### new\_messages\_json [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.new_messages_json) def new_messages_json() -> bytes Return new messages from [`new_messages`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.new_messages) as JSON bytes. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#returns-7) [`bytes`](https://docs.python.org/3/library/stdtypes.html#bytes) — JSON bytes representing the new messages. #### next [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.next) `@async` def next( node: _agent_graph.AgentNode[AgentDepsT, OutputDataT], ) -> _agent_graph.AgentNode[AgentDepsT, OutputDataT] | End[FinalResult[OutputDataT]] Manually drive the agent run by passing in the node you want to run next. This lets you inspect or mutate the node before continuing execution, or skip certain nodes under dynamic conditions. The agent run should be stopped when you return an [`End`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.End) node. Example: from pydantic_ai import Agent from pydantic_graph import End agent = Agent('openai:gpt-5.2') async def main(): async with agent.iter('What is the capital of France?') as agent_run: next_node = agent_run.next_node # start with the first node nodes = [next_node] while not isinstance(next_node, End): next_node = await agent_run.next(next_node) nodes.append(next_node) # Once `next_node` is an End, we've finished: print(nodes) ''' [\ UserPromptNode(\ user_prompt='What is the capital of France?',\ instructions_functions=[],\ system_prompts=(),\ system_prompt_functions=[],\ system_prompt_dynamic_functions={},\ ),\ ModelRequestNode(\ request=ModelRequest(\ parts=[\ UserPromptPart(\ content='What is the capital of France?',\ timestamp=datetime.datetime(...),\ )\ ],\ timestamp=datetime.datetime(...),\ run_id='...',\ conversation_id='...',\ )\ ),\ CallToolsNode(\ model_response=ModelResponse(\ parts=[TextPart(content='The capital of France is Paris.')],\ usage=RequestUsage(\ cost=Decimal('0.000196'), input_tokens=56, output_tokens=7\ ),\ model_name='gpt-5.2',\ timestamp=datetime.datetime(...),\ run_id='...',\ conversation_id='...',\ )\ ),\ End(data=FinalResult(output='The capital of France is Paris.')),\ ] ''' print('Final result:', agent_run.result.output) #> Final result: The capital of France is Paris. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#returns-8) `_agent_graph.AgentNode`\[`AgentDepsT`, `OutputDataT`\] | `End`\[`FinalResult`\[`OutputDataT`\]\] — The next node returned by the graph logic, or an [`End`](https://pydantic.dev/docs/ai/api/pydantic_graph/basenode/#pydantic_graph.basenode.End) node if `_agent_graph.AgentNode`\[`AgentDepsT`, `OutputDataT`\] | `End`\[`FinalResult`\[`OutputDataT`\]\] — the run has completed. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#parameters-1) **`node`** : `_agent_graph.AgentNode`\[`AgentDepsT`, `OutputDataT`\] [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.next(node)) The node to run next in the graph. AgentRunResult -------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult) **Bases:** `Generic[OutputDataT]` The final result of an agent run. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#attributes-1) #### conversation\_id [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.conversation_id) The unique identifier for the conversation this run belongs to. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### metadata [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.metadata) Metadata associated with this agent run, if configured. **Type:** [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`Any`](https://docs.python.org/3/library/typing.html#typing.Any)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) #### output [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.output) The output data from the agent run. **Type:** `OutputDataT` #### response [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.response) Return the last response from the message history. **Type:** `_messages.ModelResponse` #### run\_id [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.run_id) The unique identifier for the agent run. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) #### timestamp [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.timestamp) Return the timestamp of last response. **Type:** [`datetime`](https://docs.python.org/3/library/datetime.html#module-datetime) #### usage [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.usage) Return the usage of the whole run. **Type:** `_usage.RunUsage` ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#methods-1) #### all\_messages [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.all_messages) def all_messages( *, output_tool_return_content: str | None = None, ) -> list[_messages.ModelMessage] Return the history of \_messages. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#returns-9) [`list`](https://docs.python.org/3/glossary.html#term-list) \[`_messages.ModelMessage`\] — List of messages. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#parameters-2) **`output_tool_return_content`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.all_messages(output_tool_return_content)) The return content of the tool call to set in the last message. This provides a convenient way to modify the content of the output tool call if you want to continue the conversation and want to set the response to the output tool call. If `None`, the last message will not be modified. #### all\_messages\_json [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.all_messages_json) def all_messages_json(*, output_tool_return_content: str | None = None) -> bytes Return all messages from [`all_messages`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.all_messages) as JSON bytes. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#returns-10) [`bytes`](https://docs.python.org/3/library/stdtypes.html#bytes) — JSON bytes representing the messages. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#parameters-3) **`output_tool_return_content`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.all_messages_json(output_tool_return_content)) The return content of the tool call to set in the last message. This provides a convenient way to modify the content of the output tool call if you want to continue the conversation and want to set the response to the output tool call. If `None`, the last message will not be modified. #### new\_messages [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.new_messages) def new_messages( *, output_tool_return_content: str | None = None, ) -> list[_messages.ModelMessage] Return the messages produced during this run. Messages provided via `message_history` and messages from older runs are excluded. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#returns-11) [`list`](https://docs.python.org/3/glossary.html#term-list) \[`_messages.ModelMessage`\] — List of new messages. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#parameters-4) **`output_tool_return_content`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.new_messages(output_tool_return_content)) The return content of the tool call to set in the last message. This provides a convenient way to modify the content of the output tool call if you want to continue the conversation and want to set the response to the output tool call. If `None`, the last message will not be modified. #### new\_messages\_json [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.new_messages_json) def new_messages_json(*, output_tool_return_content: str | None = None) -> bytes Return new messages from [`new_messages`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.new_messages) as JSON bytes. ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#returns-12) [`bytes`](https://docs.python.org/3/library/stdtypes.html#bytes) — JSON bytes representing the new messages. ##### Parameters [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#parameters-5) **`output_tool_return_content`** : [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) _Default:_ `None` [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult.new_messages_json(output_tool_return_content)) The return content of the tool call to set in the last message. This provides a convenient way to modify the content of the output tool call if you want to continue the conversation and want to set the response to the output tool call. If `None`, the last message will not be modified. AgentRunResultEvent ------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResultEvent) **Bases:** `Generic[OutputDataT]` An event indicating the agent run ended and containing the final result of the agent run. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#attributes-2) #### event\_kind [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResultEvent.event_kind) Event type identifier, used as a discriminator. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘agent\_run\_result’\] **Default:** `'agent_run_result'` #### result [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResultEvent.result) The result of the run. **Type:** [`AgentRunResult`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRunResult) \[`OutputDataT`\] EnqueueContent -------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.EnqueueContent) A single item accepted by [`RunContext.enqueue`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.enqueue) and [`AgentRun.enqueue`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.enqueue) . `enqueue` is variadic, so each item is one positional argument: * [`UserContent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.UserContent) (a `str` or a piece of multi-modal content like an [`ImageUrl`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ImageUrl) ): adjacent user content is gathered into a single [`UserPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.UserPromptPart) , so `enqueue('caption', image)` forms one user turn. To pass an existing list, spread it: `enqueue(*items)`. * [`ModelRequestPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequestPart) (e.g. a [`SystemPromptPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.SystemPromptPart) ): included verbatim. * [`ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage) (a complete [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) or [`ModelResponse`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelResponse) ): emitted as its own message. Consecutive part-style items (user content and `ModelRequestPart`s) are coalesced into a single `ModelRequest`; complete `ModelMessage`s stay separate. This lets one `enqueue` call inject an interleaved exchange (e.g. a synthetic tool call + result — a `ModelResponse` followed by a `ModelRequest`). The assembled sequence must end in a `ModelRequest` so the agent has something to respond to. **Type:** [`TypeAlias`](https://docs.python.org/3/library/typing.html#typing.TypeAlias) **Default:** `'UserContent | ModelRequestPart | ModelMessage'` PendingMessage -------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.PendingMessage) One or more [`ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage) s queued for injection into the agent conversation. Enqueued via [`RunContext.enqueue`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.enqueue) or [`AgentRun.enqueue`](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.AgentRun.enqueue) and automatically drained at the appropriate time during the agent run by the internal `PendingMessageDrainCapability`. ### Attributes [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#attributes-3) #### enqueue\_id [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.PendingMessage.enqueue_id) Unique identifier for this enqueue call, surfaced on the [`EnqueuedMessagesEvent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.EnqueuedMessagesEvent) emitted when the messages are delivered, and returned by [`enqueue`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext.enqueue) . **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `field(default_factory=(lambda: str(uuid7())))` #### messages [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.PendingMessage.messages) The message(s) to inject, in order. Always ends in a [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) . **Type:** [`list`](https://docs.python.org/3/glossary.html#term-list) \[[`ModelMessage`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelMessage)\ \] #### priority [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.PendingMessage.priority) When to deliver these messages: * `'asap'`: at the earliest opportunity (next model request, or redirect if the agent would otherwise terminate). * `'when_idle'`: only when the agent would otherwise terminate, after `'asap'` messages. **Type:** `PendingMessagePriority` **Default:** `'asap'` ### Methods [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#methods-2) #### from\_content [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.PendingMessage.from_content) `@classmethod` def from_content( cls, *content: EnqueueContent, priority: PendingMessagePriority = 'asap', ) -> PendingMessage | None Build a `PendingMessage` from `enqueue` arguments, or `None` when there’s nothing to send. Returns `None` for an empty call (enqueueing nothing is a no-op rather than an error). ##### Returns [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#returns-13) `PendingMessage` | [`None`](https://docs.python.org/3/library/constants.html#None) ##### Raises [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#raises) * `UserError` — If the assembled messages don’t end in a [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) — e.g. a lone `ModelResponse` — since the agent needs a request to respond to. PendingMessagePriority ---------------------- [](https://pydantic.dev/docs/ai/api/pydantic-ai/run/#pydantic_ai.run.PendingMessagePriority) When to deliver a pending message. * `'asap'`: Delivered at the earliest opportunity — either prepended to the next [`ModelRequest`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ModelRequest) , or, if the agent would otherwise terminate before another request, used to redirect the run into one more request. * `'when_idle'`: Delivered only when the agent would otherwise terminate, after any `'asap'` messages. Doesn’t interrupt in-flight work. **Type:** [`TypeAlias`](https://docs.python.org/3/library/typing.html#typing.TypeAlias) **Default:** `Literal['asap', 'when_idle']` Was this page helpful? Thanks for your feedback! --- # Exa Search | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/harness/exa-search/#_top) Exa Search ========== `ExaSearch` gives an agent web research tools backed by the [Exa](https://exa.ai/) search API: search that returns the most relevant excerpts from each hit (with an optional synthesized text summary), full-page retrieval for digging into a specific URL, and opt-in deep search that synthesizes a cited answer in one call. The separate `ExaAgent` capability delegates long-running research to the Exa Agent API as deferred tool calls. [Source](https://github.com/pydantic/pydantic-ai-harness/tree/main/pydantic_ai_harness/exa/) The problem ----------- [](https://pydantic.dev/docs/ai/harness/exa-search/#the-problem) Search tools that return only titles and snippets force a second round of fetching before the agent can judge a source, while search tools that return full page text flood the context with pages the agent will discard. Wiring a search API together with a page fetcher, budgeting what each tool returns, and prompting the agent to research methodically is boilerplate every research agent reinvents. `ExaSearch` bundles that plumbing into a single [capability](https://pydantic.dev/docs/ai/capabilities/overview/) : the research tools, per-tool output budgets, and short research guidance in the system prompt. Usage ----- [](https://pydantic.dev/docs/ai/harness/exa-search/#usage) Install the `exa` extra and set the `EXA_API_KEY` environment variable (create a key at [https://dashboard.exa.ai](https://dashboard.exa.ai/) ): Terminal uv add "pydantic-ai-harness[exa]" Then pass `ExaSearch` to an `Agent` via the `capabilities` parameter: from pydantic_ai import Agent from pydantic_ai_harness.exa import ExaSearch agent = Agent('anthropic:claude-sonnet-4-6', capabilities=[ExaSearch()]) result = agent.run_sync('What changed in the latest stable Python release?') print(result.output) Tools ----- [](https://pydantic.dev/docs/ai/harness/exa-search/#tools) `ExaSearch` contributes these tools to the agent: | Tool | Purpose | | --- | --- | | `web_search` | Search the web and return the top `num_results` pages, each with title, URL, and its most relevant excerpts. | | `get_page` | Retrieve the full text of one specific URL — a promising `web_search` hit, or a URL the user provided. | | `deep_search` | Run Exa’s multi-step deep search and return a synthesized, cited answer. Opt-in via `include_deep_search=True`. | | `exa_agent` | Delegate a research task to an asynchronous Exa agent run. Provided by the separate `ExaAgent` capability. | `web_search` returns short excerpts (Exa highlights) rather than full page text, following [Exa’s own guidance for agents](https://exa.ai/docs/reference/search-api-guide-for-coding-agents) , so surveying several sources stays cheap; the agent reads a chosen page with `get_page`. `get_page` text is capped at `max_text_chars` characters, keeping the **head** (a page’s lead carries the substance). One character of headroom above the cap is requested from Exa, so when a page exceeds the cap the output ends with a `[... page text truncated at N characters]` marker; at the API ceiling of 10,000 characters no headroom exists, so the marker cannot appear there. The result count is bounded the same way: `num_results` is requested from Exa and re-applied to the response. A URL or question that returns no content, a rate limit, or a transient API or network failure surfaces to the model as a [`ModelRetry`](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-retries) rather than a hard error: the run continues and the model can correct the URL, rephrase, or try again. Authentication failures (401/403) are configuration errors and propagate. Deep search ----------- [](https://pydantic.dev/docs/ai/harness/exa-search/#deep-search) `deep_search` calls Exa search with `type='deep'` and a plain-text output schema: Exa expands the question into multiple queries, searches, and returns an answer grounded in citations — all in **one tool call**, with the cited sources listed under the answer. Each call invests more time and search depth than `web_search` (Exa’s research-grade mode), and the model decides when to invoke tools, so the tool is off by default — enable it explicitly: from pydantic_ai_harness.exa import ExaSearch ExaSearch(include_deep_search=True) When enabled, the capability’s instructions tell the model to treat it as an escalation from `web_search`, not a replacement. The synthesized answer is returned in full (it is Exa-generated and inherently bounded); `max_text_chars` only applies to `get_page`. Text summary ------------ [](https://pydantic.dev/docs/ai/harness/exa-search/#text-summary) Set `text_summary` to have every `web_search` call also request Exa’s plain-text output schema, so the response carries a short summary synthesized from the results for question-style queries. Pass `True` for an unconstrained summary, or a string describing the desired format (sent as the schema’s `description`): from pydantic_ai_harness.exa import ExaSearch ExaSearch(text_summary='One concise sentence with the requested facts.') The tool’s return shape is unchanged and backward compatible: the result list is returned as before, and when Exa returns a summary it is prepended as a `Summary:` line. Structured citations -------------------- [](https://pydantic.dev/docs/ai/harness/exa-search/#structured-citations) Every tool returns a [`ToolReturn`](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#advanced-tool-returns) : `return_value` carries the readable text the model sees (unchanged from previous releases, including the `Sources:` blocks), and `metadata` carries the sources as structured `ExaSource` records (`{'url': ..., 'title': ...}`) under the `'sources'` key. Metadata is never sent to the model; the application reads it from the `ToolReturnPart` in the message history, so rendering citations needs no text parsing: from pydantic_ai.messages import ModelRequest, ToolReturnPart for message in result.all_messages(): if isinstance(message, ModelRequest): for part in message.parts: if isinstance(part, ToolReturnPart) and part.metadata is not None: for source in part.metadata.get('sources', []): print(source['url'], source['title']) `exa_agent` results additionally carry the Exa run ID in metadata under `RUN_ID_METADATA_KEY`. Instructions ------------ [](https://pydantic.dev/docs/ai/harness/exa-search/#instructions) `ExaSearch` contributes short research guidance to the system prompt: search wide with `web_search` first, read the most promising pages in full with `get_page` before drawing conclusions, prefer primary sources, and cite the URLs relied on. With `include_deep_search=True`, the guidance also covers when to escalate to `deep_search`. Set `guidance` to replace the default text, or to `''` to contribute no instructions at all. Configuration ------------- [](https://pydantic.dev/docs/ai/harness/exa-search/#configuration) Every field of `ExaSearch` with its default: from pydantic_ai_harness.exa import ExaSearch ExaSearch( num_results=5, # results per web_search call (1 to 100) max_text_chars=10_000, # get_page text cap, in characters (1 to 10,000) text_summary=False, # web_search also returns a synthesized text summary include_deep_search=False, # also expose the deep_search tool include_domains=[], # only search these domains (allowlist) exclude_domains=[], # never search these domains (denylist) guidance=None, # None = default instructions, '' = none, str = custom client=None, # ExaClient -- None builds exa_py.AsyncExa from EXA_API_KEY ) `include_domains` and `exclude_domains` apply to `web_search` and `deep_search`, and are mutually exclusive — set one, not both. Out-of-range limits and setting both domain lists raise at construction. Exa agent runs -------------- [](https://pydantic.dev/docs/ai/harness/exa-search/#exa-agent-runs) The Exa [Agent API](https://exa.ai/docs/reference/agent-api-guide) runs open-ended research tasks asynchronously: a run is created, moves through `queued -> running`, and reaches a terminal status (`completed`, `failed`, or `cancelled`) after up to an hour. The separate `ExaAgent` capability maps that lifecycle onto Pydantic AI’s [deferred tool calls](https://pydantic.dev/docs/ai/tools-toolsets/deferred-tools/) : its `exa_agent` tool creates the run and defers, carrying the Exa run ID in the deferred call’s metadata. from pydantic_ai import Agent from pydantic_ai_harness.exa import ExaAgent agent = Agent('anthropic:claude-sonnet-4-6', capabilities=[ExaAgent()]) By default (`execution='inline'`) the capability resolves its own deferred calls within the agent run by polling the Exa run to completion, so the tool behaves like a regular (if slow) tool. With `execution='external'` the calls bubble up as `DeferredToolRequests` output for the host application to resolve out of band — including from a different process, since the Exa run ID survives in the request metadata under `RUN_ID_METADATA_KEY`. The agent’s `output_type` must include `DeferredToolRequests`, otherwise the run raises instead of returning the deferred requests: from pydantic_ai import Agent from pydantic_ai.tools import DeferredToolRequests from pydantic_ai_harness.exa import ExaAgent agent = Agent( 'anthropic:claude-sonnet-4-6', output_type=[str, DeferredToolRequests], capabilities=[ExaAgent(execution='external')], ) Render a finished run with the `agent_run_result` helper, passing the same `output_schema` the capability was constructed with so external resolution applies the same validation and produces the same tool result shape as inline execution, then feed the results and the original message history back into the agent to resume the deferred run: from pydantic_ai.tools import DeferredToolResults from pydantic_ai_harness.exa import RUN_ID_METADATA_KEY, agent_run_result async def resolve(requests, runs, output_schema=None): # e.g. in a worker process results = DeferredToolResults() for call in requests.calls: run_id = requests.metadata[call.tool_call_id][RUN_ID_METADATA_KEY] run = await runs.poll_until_finished(run_id) results.calls[call.tool_call_id] = agent_run_result(run, output_schema=output_schema) return results async def resume(agent, messages, results): return await agent.run(message_history=messages, deferred_tool_results=results) Every field of `ExaAgent` with its default: from pydantic_ai_harness.exa import ExaAgent ExaAgent( effort=None, # 'low' | 'medium' | 'high' | 'xhigh' | 'auto' -- None = API default execution='inline', # 'inline' polls to completion; 'external' bubbles DeferredToolRequests output_schema=None, # BaseModel class or dict schema for structured output system_prompt=None, # forwarded to the Exa agent run poll_interval=1000, # ms between polls when resolving inline timeout_ms=3_600_000, # ms to wait for a run when resolving inline guidance=None, # None = default instructions, '' = none, str = custom runs=None, # ExaAgentRuns -- None builds AsyncExa().agent.runs from EXA_API_KEY ) An `output_schema` model class is validated against the completed run’s structured output (mismatches surface as `ModelRetry`), while a dict schema is forwarded without client-side validation and is the agent-spec form. Terminal failures (`failed`, `cancelled`) are returned to the model as a structured message rather than raised, so the agent can decide how to proceed. Each result includes the run ID, which the model can pass back as `previous_run_id` to ask follow-up questions in the context of a previous run. Multiple instances ------------------ [](https://pydantic.dev/docs/ai/harness/exa-search/#multiple-instances) Two instances of the same capability register the same tool names, which is an error. To run several differently configured instances in one agent (for example one open-web `ExaSearch` and one pinned to specific domains), wrap the extra instances in core’s `PrefixTools` capability, which prefixes their tool names: from pydantic_ai import Agent from pydantic_ai.capabilities import PrefixTools from pydantic_ai_harness.exa import ExaSearch agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[\ ExaSearch(), # web_search, get_page\ PrefixTools(\ wrapped=ExaSearch(include_domains=['crunchbase.com'], guidance=''),\ prefix='cb',\ ), # cb_web_search, cb_get_page\ ], ) Set `guidance=''` on the wrapped instance (or replace it with text that tells the model when to use the prefixed tools), since each instance otherwise contributes the same default research guidance. This also works for `ExaAgent`: it identifies its deferred calls by metadata it wrote when deferring, not by tool name, so a prefixed `exa_agent` still resolves inline, and multiple `ExaAgent` instances never claim each other’s calls. Custom client ------------- [](https://pydantic.dev/docs/ai/harness/exa-search/#custom-client) The default client is `exa_py.AsyncExa`, configured from the `EXA_API_KEY` environment variable; when the variable is missing, construction fails with a setup hint. Pass any object satisfying the `ExaClient` protocol — the subset of `AsyncExa` the toolset calls — to configure authentication or the base URL explicitly, or to substitute a fake in tests: from exa_py import AsyncExa from pydantic_ai_harness.exa import ExaSearch ExaSearch(client=AsyncExa(api_key='...')) The API may change between releases while the capability settles; breaking changes ship deprecation warnings where practical. ExaSearch vs core WebSearch --------------------------- [](https://pydantic.dev/docs/ai/harness/exa-search/#exasearch-vs-core-websearch) Pydantic AI core ships a provider-adaptive [`WebSearch`](https://pydantic.dev/docs/ai/capabilities/overview/#provider-adaptive-tools) capability: on models with a native search tool it uses the provider’s own search, executed server-side; elsewhere it falls back to a local DuckDuckGo tool. Reach for it when you want search that follows the model. Reach for `ExaSearch` when you want the same search behavior on every model: one vendor, excerpts with every hit, explicit page retrieval, domain filters, and opt-in deep search. One caveat when combining them: on Anthropic models the provider-native search tool is also named `web_search` on the wire, so `capabilities=[WebSearch(), ExaSearch()]` puts two tools with the same name in the request. Use one search capability per agent on native-search models, or force the local fallback with `WebSearch(native=False)` (its DuckDuckGo tool is named `duckduckgo_search`, which does not collide). ExaSearch vs Exa’s MCP server ----------------------------- [](https://pydantic.dev/docs/ai/harness/exa-search/#exasearch-vs-exas-mcp-server) Exa also ships an official hosted MCP server at `https://mcp.exa.ai/mcp` ([exa-labs/exa-mcp-server](https://github.com/exa-labs/exa-mcp-server) ). By default it exposes `web_search_exa` and `web_fetch_exa`; the full catalog adds `web_search_advanced_exa` and an agent-run set (`agent_create_run`, `agent_wait_for_run`, `agent_get_run_output`, `agent_cancel_run`). `ExaSearch` is the curated, typed path: bounded output, a retry-on-empty contract, bundled research instructions, and a client seam that makes it testable offline. The MCP server is how you get Exa’s full catalog with zero wrapper code, via Pydantic AI core’s MCP capability. Their agent runs are create-then-poll (`agent_create_run` returns an ID immediately; `agent_wait_for_run` polls it), where `deep_search` returns the answer in a single call. The two compose in one `capabilities` list, and none of the MCP tool names collide with `web_search`, `get_page`, or `deep_search`: from pydantic_ai import Agent from pydantic_ai.capabilities import MCP from pydantic_ai_harness.exa import ExaSearch agent = Agent('anthropic:claude-sonnet-4-6', capabilities=[ExaSearch(), MCP('https://mcp.exa.ai/mcp')]) Agent spec (YAML/JSON) ---------------------- [](https://pydantic.dev/docs/ai/harness/exa-search/#agent-spec-yamljson) `ExaSearch` works with Pydantic AI’s [agent spec](https://pydantic.dev/docs/ai/core-concepts/agent-spec/) , so you can declare it in a config file instead of Python: # agent.yaml model: anthropic:claude-sonnet-4-6 capabilities: - ExaSearch: num_results: 3 include_deep_search: true - ExaAgent: effort: low from pydantic_ai import Agent from pydantic_ai_harness.exa import ExaAgent, ExaSearch agent = Agent.from_file('agent.yaml', custom_capability_types=[ExaSearch, ExaAgent]) Pass `custom_capability_types` so the spec loader knows how to instantiate the capabilities. The `client` and `runs` fields are not spec-serializable; spec-loaded instances always build the default client from `EXA_API_KEY`. In specs, `output_schema` takes the JSON-schema dict form; Pydantic model classes are only available when constructing the capability in Python. Further reading --------------- [](https://pydantic.dev/docs/ai/harness/exa-search/#further-reading) * [Pydantic AI capabilities](https://pydantic.dev/docs/ai/capabilities/overview/) * [Toolsets](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) * [Exa API documentation](https://docs.exa.ai/) API reference ------------- [](https://pydantic.dev/docs/ai/harness/exa-search/#api-reference) ExaSearch --------- [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaSearch) **Bases:** `AbstractCapability[AgentDepsT]` Web research for agents, backed by the [Exa](https://exa.ai/) search API. Adds two tools: `web_search`, which returns search results with their most relevant excerpts, and `get_page`, which retrieves the full text of a specific URL. Set `text_summary` to have `web_search` also return a synthesized text summary with the results. Set `include_deep_search=True` to also expose `deep_search`, which runs Exa’s multi-step deep search and returns a synthesized, cited answer in one tool call. from pydantic_ai import Agent from pydantic_ai_harness.exa import ExaSearch agent = Agent('anthropic:claude-sonnet-4-6', capabilities=[ExaSearch()]) Authentication comes from the `EXA_API_KEY` environment variable by default; pass `client` to configure it explicitly. ### Attributes [](https://pydantic.dev/docs/ai/harness/exa-search/#attributes) #### client [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaSearch.client) Exa client to use; when `None`, an `exa_py.AsyncExa` is built from `EXA_API_KEY`. Any object satisfying the `ExaClient` protocol works: use it to pass an API key explicitly, point at a different base URL, or substitute a fake in tests. **Type:** `ExaClient` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### exclude\_domains [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaSearch.exclude_domains) Search results never come from these domains (denylist). Applies to `web_search` and `deep_search`. Mutually exclusive with `include_domains`. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `field(default_factory=(list[str]))` #### guidance [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaSearch.guidance) Custom research guidance for the system prompt. Leave as `None` for the default guidance (which adapts to `include_deep_search`), or set `''` to contribute no instructions at all. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### include\_deep\_search [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaSearch.include_deep_search) Also expose the `deep_search` tool. Off by default. Deep search (Exa search `type='deep'`) runs a multi-step agentic search and synthesizes a cited answer in one call. Each call invests more time and search depth than `web_search` (Exa’s research-grade mode), and the model decides when to invoke tools, so that investment is opt-in rather than the default. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) **Default:** `False` #### include\_domains [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaSearch.include_domains) If non-empty, search results only come from these domains (allowlist). Applies to `web_search` and `deep_search`. Mutually exclusive with `exclude_domains`. **Type:** [`Sequence`](https://docs.python.org/3/library/typing.html#typing.Sequence) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] **Default:** `field(default_factory=(list[str]))` #### max\_text\_chars [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaSearch.max_text_chars) Maximum characters of page text `get_page` returns (1 to 10,000, the Exa API range). One character of headroom above the cap is requested from Exa so local truncation can detect a longer page and append a truncation marker. At the API ceiling of 10,000 no headroom exists, so the marker cannot fire there. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `10000` #### num\_results [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaSearch.num_results) Number of results `web_search` returns per query (1 to 100, the Exa API range). **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `5` #### text\_summary [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaSearch.text_summary) Have `web_search` also return a synthesized text summary above the results. Off by default. When enabled, each `web_search` call requests Exa’s plain-text output schema, so the response carries a short summary synthesized from the results in addition to the result list. Pass a string to describe the desired summary format (it is sent as the schema’s `description`), or `True` for an unconstrained summary. The tool’s return shape is unchanged: the summary is prepended as a `Summary:` line when Exa returns one. **Type:** [`bool`](https://docs.python.org/3/library/functions.html#bool) | [`str`](https://docs.python.org/3/library/stdtypes.html#str) **Default:** `False` ### Methods [](https://pydantic.dev/docs/ai/harness/exa-search/#methods) #### \_\_post\_init\_\_ [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaSearch.__post_init__) def __post_init__() -> None Validate configuration against the Exa API’s documented bounds. ##### Returns [](https://pydantic.dev/docs/ai/harness/exa-search/#returns) [`None`](https://docs.python.org/3/library/constants.html#None) #### from\_spec [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaSearch.from_spec) `@classmethod` def from_spec( cls, *, num_results: int = 5, max_text_chars: int = 10000, text_summary: bool | str = False, include_deep_search: bool = False, include_domains: Sequence[str] = (), exclude_domains: Sequence[str] = (), guidance: str | None = None, ) -> ExaSearch[AgentDepsT] Construct the capability from serializable spec options. The `client` field is not spec-serializable, so spec-loaded instances always build the default `exa_py.AsyncExa` from `EXA_API_KEY`. ##### Returns [](https://pydantic.dev/docs/ai/harness/exa-search/#returns-1) `ExaSearch`\[`AgentDepsT`\] #### get\_instructions [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaSearch.get_instructions) def get_instructions() -> AgentInstructions[AgentDepsT] | None Static research guidance: search wide, read the promising pages in full, cite URLs. When `include_deep_search` is set, the default guidance also covers when to escalate to `deep_search`. A non-`None` `guidance` replaces the default; `''` disables instructions entirely. ##### Returns [](https://pydantic.dev/docs/ai/harness/exa-search/#returns-2) `AgentInstructions`\[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) #### get\_toolset [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaSearch.get_toolset) def get_toolset() -> ExaSearchToolset[AgentDepsT] Build the toolset providing `web_search`, `get_page`, and the optional `deep_search` tool. ##### Returns [](https://pydantic.dev/docs/ai/harness/exa-search/#returns-3) `ExaSearchToolset`\[`AgentDepsT`\] ExaAgent -------- [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaAgent) **Bases:** `AbstractCapability[AgentDepsT]` Delegation of deep research tasks to the [Exa](https://exa.ai/) Agent API. Adds one tool, `exa_agent`, which creates an asynchronous Exa agent run (queued -> running -> terminal, up to an hour) and defers the tool call. By default the capability resolves its own deferred calls inline by polling the run to completion. With `execution='external'` the calls surface as `DeferredToolRequests` output instead, for the host application to resolve out of band (the run ID is in the request metadata under `RUN_ID_METADATA_KEY`), including across process restarts. from pydantic_ai import Agent from pydantic_ai_harness.exa import ExaAgent agent = Agent('anthropic:claude-sonnet-4-6', capabilities=[ExaAgent()]) Authentication comes from the `EXA_API_KEY` environment variable by default; pass `runs` to configure it explicitly. ### Attributes [](https://pydantic.dev/docs/ai/harness/exa-search/#attributes-1) #### effort [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaAgent.effort) How much work the Exa agent invests per run; `None` uses the API default. **Type:** `AgentEffort` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### execution [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaAgent.execution) How deferred `exa_agent` calls are resolved. With `'inline'`, the capability polls the run to completion during the agent run. With `'external'`, calls bubble up as `DeferredToolRequests` output for the host application to resolve (see `agent_run_result`), which suits durable workers that outlive a single process. **Type:** [`Literal`](https://docs.python.org/3/library/typing.html#typing.Literal) \[‘inline’, ‘external’\] **Default:** `'inline'` #### guidance [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaAgent.guidance) Custom delegation guidance for the system prompt. Leave as `None` for the default guidance, or set `''` to contribute no instructions at all. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### output\_schema [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaAgent.output_schema) Structured output schema for the Exa agent’s result. `None` returns prose. Accepts a Pydantic model class or a JSON-schema-style dict. A model class is forwarded to the API and a completed run’s structured output is validated against it (a mismatch surfaces as a retry). The dict form skips client-side validation and is the serializable shape used by agent specs. **Type:** [`type`](https://docs.python.org/3/glossary.html#term-type) \[`BaseModel`\] | [`dict`](https://docs.python.org/3/reference/expressions.html#dict) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ , [`object`](https://docs.python.org/3/glossary.html#term-object)\ \] | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### poll\_interval [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaAgent.poll_interval) Milliseconds between polls while resolving a run inline. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `1000` #### runs [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaAgent.runs) Exa Agent runs client; when `None`, `exa_py.AsyncExa().agent.runs` is built from `EXA_API_KEY`. Any object satisfying the `ExaAgentRuns` protocol works: use it to pass an API key explicitly or substitute a fake in tests. **Type:** `ExaAgentRuns` | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### system\_prompt [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaAgent.system_prompt) System prompt forwarded to the Exa agent run; `None` uses the API default. **Type:** [`str`](https://docs.python.org/3/library/stdtypes.html#str) | [`None`](https://docs.python.org/3/library/constants.html#None) **Default:** `None` #### timeout\_ms [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaAgent.timeout_ms) Milliseconds to wait for a run to finish when resolving inline. **Type:** [`int`](https://docs.python.org/3/library/functions.html#int) **Default:** `3600000` ### Methods [](https://pydantic.dev/docs/ai/harness/exa-search/#methods-1) #### from\_spec [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaAgent.from_spec) `@classmethod` def from_spec( cls, *, effort: AgentEffort | None = None, execution: Literal['inline', 'external'] = 'inline', output_schema: dict[str, object] | None = None, system_prompt: str | None = None, poll_interval: int = 1000, timeout_ms: int = 3600000, guidance: str | None = None, ) -> ExaAgent[AgentDepsT] Construct the capability from serializable spec options. The `runs` field is not spec-serializable, so spec-loaded instances always build the default client from `EXA_API_KEY`. `output_schema` takes the JSON-schema dict form here; Pydantic model classes are only available when constructing the capability in Python. ##### Returns [](https://pydantic.dev/docs/ai/harness/exa-search/#returns-4) `ExaAgent`\[`AgentDepsT`\] #### get\_instructions [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaAgent.get_instructions) def get_instructions() -> AgentInstructions[AgentDepsT] | None Static delegation guidance: when to hand a task to `exa_agent`, and run ID continuation. A non-`None` `guidance` replaces the default; `''` disables instructions entirely. ##### Returns [](https://pydantic.dev/docs/ai/harness/exa-search/#returns-5) `AgentInstructions`\[`AgentDepsT`\] | [`None`](https://docs.python.org/3/library/constants.html#None) #### get\_toolset [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaAgent.get_toolset) def get_toolset() -> ExaAgentToolset[AgentDepsT] Build the toolset providing the `exa_agent` tool. ##### Returns [](https://pydantic.dev/docs/ai/harness/exa-search/#returns-6) `ExaAgentToolset`\[`AgentDepsT`\] #### handle\_deferred\_tool\_calls [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaAgent.handle_deferred_tool_calls) `@async` def handle_deferred_tool_calls( ctx: RunContext[AgentDepsT], *, requests: DeferredToolRequests, ) -> DeferredToolResults | None Resolve deferred `exa_agent` calls inline by polling the Exa run to completion. With `execution='external'` all calls are left unresolved, so they bubble up as `DeferredToolRequests` output for the host application to resolve; the Exa run ID is available in `requests.metadatatool_call_id`. Calls are claimed by the instance token in the deferred-call metadata rather than by tool name, so tool renaming or prefixing wrappers (e.g. `PrefixTools`) do not break inline resolution. ##### Returns [](https://pydantic.dev/docs/ai/harness/exa-search/#returns-7) [`DeferredToolResults`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.DeferredToolResults) | [`None`](https://docs.python.org/3/library/constants.html#None) agent\_run\_result ------------------ [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.agent_run_result) def agent_run_result( run: AgentRun, *, output_schema: type[BaseModel] | dict[str, object] | None = None, ) -> ToolReturn[str] Render a terminal Exa agent run as the `exa_agent` tool result. Use it when resolving externally executed `exa_agent` calls in a host application (building `DeferredToolResults`), so external and inline execution produce the same result shape. When `output_schema` is a Pydantic model class, a completed run’s structured output is validated against it; a mismatch raises `ModelRetry`. The `ToolReturn.return_value` text is what the model sees; its metadata carries the run ID under `RUN_ID_METADATA_KEY` and the citation `sources` (`ExaSource` dicts) for the application to use directly. ### Returns [](https://pydantic.dev/docs/ai/harness/exa-search/#returns-8) [`ToolReturn`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolReturn) \[[`str`](https://docs.python.org/3/library/stdtypes.html#str)\ \] ExaSource --------- [](https://pydantic.dev/docs/ai/harness/exa-search/#pydantic_ai_harness.exa.ExaSource) **Bases:** [`TypedDict`](https://docs.python.org/3/library/typing.html#typing.TypedDict) One source behind a tool result, carried in `ToolReturn.metadata['sources']`. Was this page helpful? Thanks for your feedback! --- # On-Demand Capabilities | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/capabilities/on-demand/#_top) On-Demand Capabilities ====================== A capability is a bundle of instructions and/or tools, optionally with settings and hooks. A multi-workflow agent normally sends every workflow’s instructions and tool schemas on every turn, and applies every workflow’s settings and hooks for the whole run — even though most requests need just one workflow. That cost grows with each workflow you add: more input tokens, and worse tool selection once the visible tool set passes the ~30–50-tool mark where models start picking the wrong one (the same pressure behind [tool search](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-search) ). Mark a [capability](https://pydantic.dev/docs/ai/capabilities/overview/) with `defer_loading=True` and give it a stable `id`, and it collapses to a one-line catalog entry — its `id` plus an optional `description` — that the model pulls in on demand. Here’s the minimal shape: on\_demand\_capability.py from pydantic_ai import Agent from pydantic_ai.capabilities import Capability refunds = Capability( id='refunds', description='Use for refund eligibility, refund status, or processing a refund.', instructions='Always confirm the order ID before issuing a refund.', defer_loading=True, ) @refunds.tool_plain def refund_status(order_id: str) -> str: """Look up the refund status for an order.""" return f'Order {order_id}: refund issued on 2026-05-01.' agent = Agent( 'openai-responses:gpt-5.4', instructions='You are a customer support assistant.', capabilities=[refunds], ) On the first turn, the refund workflow is collapsed to a catalog entry. The model sees its base instructions, the framework-managed `load_capability` tool, and the catalog appended to the instructions: The following capabilities are deferred and can be loaded using the `load_capability` tool. A capability may have tools; they stay hidden until it is loaded: - refunds: Use for refund eligibility, refund status, or processing a refund. The model does not receive the refund instructions or the `refund_status` tool definition yet, so it has no reason to call the tool. Depending on the active model, Pydantic AI may also send provider/tool-search plumbing to preserve the hidden state; that plumbing does not expose the refund tool definition until the capability is loaded. The exchange unfolds across model requests within a single `agent.run_sync` call: 1. **Request 1.** The model sees the catalog above and the user’s prompt. It calls the `load_capability` tool with `id='refunds'`. 2. **Load.** Pydantic AI returns the capability’s instructions — _“Always confirm the order ID before issuing a refund.”_ — as the tool result and exposes the `refund_status` definition on the next request. 3. **Request 2.** The model now sees those instructions in history and `refund_status` in its tool list. It calls `refund_status(order_id='ABC-123')` and answers the user from the result. Already-loaded capabilities stay loaded for the rest of the run — the model never needs to re-open one. Searching cannot reveal a capability-owned tool: it stays hidden until its capability loads. In runs that also have searchable deferred tools, the catalog explicitly steers the model to load the capability rather than search for its tools; in capability-only runs — where no search surface exists — the catalog omits any mention of searching. Loading activates the whole bundle, not just instructions: the capability’s function tools, model settings, and lifecycle hooks come live together (see [What you can defer](https://pydantic.dev/docs/ai/capabilities/on-demand/#what-you-can-defer) ). It’s a one-line change to a capability you already register, it works on [every provider](https://pydantic.dev/docs/ai/capabilities/on-demand/#cross-provider-behavior) , and it [survives history replay](https://pydantic.dev/docs/ai/capabilities/on-demand/#resumable-across-runs) . What you can defer ------------------ [](https://pydantic.dev/docs/ai/capabilities/on-demand/#what-you-can-defer) Every part of a capability bundle activates together as a single unit: | Part | Before load | After load | | --- | --- | --- | | Instructions (static or dynamic) | Not sent | Returned as the `load_capability` tool result; included in subsequent requests | | Function tools | Not exposed | Exposed on the next request | | Model settings (static or per-step) | Not applied | Merged into the run’s settings for subsequent requests | | Lifecycle [hooks](https://pydantic.dev/docs/ai/capabilities/custom/#hooking-into-the-lifecycle) | Do not fire | Fire after the capability is loaded | | [Native tools](https://pydantic.dev/docs/ai/tools-toolsets/native-tools/) | Not exposed | Exposed on the next request — see [Cache implications](https://pydantic.dev/docs/ai/capabilities/on-demand/#cache-implications) | When to use it -------------- [](https://pydantic.dev/docs/ai/capabilities/on-demand/#when-to-use-it) **Reach for on-demand capabilities when:** * the agent serves multiple distinct workflows (refunds, returns, fraud review, account security…) where most turns need one * a workflow needs _more than instructions_ — its own tools, raised reasoning effort, an approval hook — and those should travel together as a unit * you want skills-style progressive disclosure but also want the loaded bundle to bring tools and settings, not just a runbook **Skip it when:** * the capability is used on most turns — the discovery round-trip costs more than the tokens it saves * you have a flat catalog of individually-discoverable tools with no shared instructions — use [tool search](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-search) instead, which discovers individual tools by name rather than loading bundles If you’ve used [Anthropic’s Agent Skills](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) , this is the same idea generalised: a skill is a markdown file the model can pull in on demand. An on-demand capability does that _plus_ typed function tools, per-step model settings, and lifecycle hooks. Retrofitting an existing capability ----------------------------------- [](https://pydantic.dev/docs/ai/capabilities/on-demand/#retrofitting-an-existing-capability) `defer_loading=True` is not specific to the [`Capability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Capability) convenience class. The shared fields live on [`AbstractCapability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability) , and built-in capabilities expose `id`, `description`, and `defer_loading` on construction. For custom capabilities, set those attributes on the instance. defer\_existing\_capability.py from pydantic_ai import Agent from pydantic_ai.capabilities import MCP agent = Agent( 'openai-responses:gpt-5.4', capabilities=[\ MCP(\ url='https://mcp.example.com/analytics',\ native=True,\ id='analytics-mcp',\ description='Use for analytics queries, dashboards, and metric lookups.',\ defer_loading=True,\ ),\ ], ) Until the model loads `analytics-mcp`, none of the MCP server’s tool definitions enter the prompt. The same flag works on [`WebSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebSearch) , [`WebFetch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebFetch) , [`Hooks`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Hooks) , and any custom [`AbstractCapability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability) subclass — see [Building custom capabilities](https://pydantic.dev/docs/ai/capabilities/custom/) for adding `defer_loading` to your own subclass. Resumable across runs --------------------- [](https://pydantic.dev/docs/ai/capabilities/on-demand/#resumable-across-runs) Loaded-capability and tool-availability state live in message history, not in the agent. When a conversation is persisted to a database and resumed later — possibly on a different process, machine, or model — Pydantic AI reconstructs the loaded capability IDs from `load_capability` call/return pairs and the revealed tool names from [`ToolAvailabilityDeltaPart`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolAvailabilityDeltaPart) . Capabilities the model loaded earlier stay loaded; capabilities it never loaded stay collapsed in the catalog. No re-discovery round-trip on resume. This is why deferred capabilities require a stable explicit `id`: history replay matches calls to capabilities by id, so a class-derived id would silently break the moment a class is renamed. The same property makes cross-provider replay work — a run that loaded `refunds` on Anthropic and continued on OpenAI Responses keeps `refunds` loaded after the switch. History carries _which_ capability ids were loaded, not the capabilities themselves: the resuming agent must be constructed with the same capabilities (matching `id`s), just as it must be constructed with the same tools. State lives in history; definitions live in code. Runtime state in `RunContext` ----------------------------- [](https://pydantic.dev/docs/ai/capabilities/on-demand/#runtime-state-in-runcontext) Several [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) fields expose progressive-disclosure state to tools, hooks, and capability-owned callbacks: * `ctx.loaded_capability_ids` — deferred capability IDs explicitly loaded through the `load_capability` tool, reconstructed from message history before each model request. A capability loaded during a step appears from the _next_ step onwards, which is also the first step on which its instructions and tools reach the model. * `ctx.available_capability_ids` — the currently-live capability IDs: always-available capabilities plus `ctx.loaded_capability_ids`. * `ctx.capability_loaded` — only meaningful while Pydantic AI is running a capability-owned hook or callback. It is scoped to that capability; deferred hooks and callbacks are skipped until this value would be true. * `ctx.discovered_tool_names` — deferred function tools revealed by durable history, whether through tool search, [`ToolReturn.tools`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolReturn) , or a capability load. * `ctx.available_tool_names` — function tool names currently known as available: always-visible tools from the current step’s assembled tool manager plus names revealed in history. Early hooks such as `before_run` may see only the history-derived names, or an empty set if none exist yet, before tool definitions have been prepared. See [Hook ordering](https://pydantic.dev/docs/ai/core-concepts/hooks/#hook-ordering) for how hook timing affects what is populated. * `ctx.is_tool_available(tool)` — whether a function tool is currently visible. Wrapping toolsets should pass the [`ToolDefinition`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.ToolDefinition) they hold; model-request hooks and tool execution can pass a name from the current `ctx.tools` snapshot. * `ctx.usage_limits` — the [`UsageLimits`](https://pydantic.dev/docs/ai/api/pydantic-ai/usage/#pydantic_ai.usage.UsageLimits) the run is enforcing (defaulting to `UsageLimits()` when none were passed, so it’s only `None` outside of a run), alongside `ctx.usage` for the usage so far. A capability can read the run’s limits to disclose or adapt to the remaining budget (e.g. budget disclosure) without being configured with a duplicate copy. Treat it as read-only: it’s the live object the run enforces against, so mutating a field would change what the run enforces on subsequent requests. Loading a capability updates the capability state immediately, but the loaded bundle’s function tools, native tools, and model settings take effect on the next model request. Cross-provider behavior ----------------------- [](https://pydantic.dev/docs/ai/capabilities/on-demand/#cross-provider-behavior) On-demand capabilities work on every model, and where the provider can express an availability change natively, loading one leaves the prompt prefix intact. A capability-owned tool is hidden until its capability loads, and it is never searchable — the model reaches it by loading the capability, not by asking for it. The unified rule is that an unrevealed deferred tool stays outside the model’s usable context; each provider’s reveal mechanism determines its wire representation. * **Anthropic `tool_addition_mode='by_reference'`** references the revealed name in a `tool_addition` block. A capability-only run pre-advertises the definition with `defer_loading=True`; a mixed run with a search surface withholds it, then appends the deferred definition in the same request as the reveal. * **OpenAI Responses `tool_addition_mode='with_definitions'`** carries the full revealed definition in an appended `additional_tools` input item and leaves it out of `tools`. * **No provider-native reveal-item support (`tool_addition_mode=None`)** announces `The following tool(s) are now available: {names}` when the schema is visible. It synthesizes a `search_tools` exchange only when a result must reveal a schema that is still withheld. Add a standalone `defer_loading=True` tool to the same run and tool search comes back for it, since that one genuinely is searchable. Capability-owned tools stay off the wire entirely while a search surface is present, so search remains fully native — server-executed where the model supports it — and no query can surface a tool whose capability has not loaded. ### Cache implications [](https://pydantic.dev/docs/ai/capabilities/on-demand/#cache-implications) Calling the `load_capability` tool reveals capability behavior between requests. Whether that breaks the provider’s prompt-cache prefix depends on what’s revealed: `load_capability` returns the loaded capability’s function-tool names through [`ToolReturn.tools`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.ToolReturn) , and the executor records the availability delta beside the tool result. Any user tool can use the same source. Histories that contain a complete capability-load exchange without its delta are translated before the next model request. | What loads | Cache prefix | | --- | --- | | Instructions only | **Stable** — instructions land in the message history, not the request prefix. | | Function tools with provider-native reveal-item support (`tool_addition_mode='by_reference'` or `'with_definitions'`) | **Stable on Anthropic and OpenAI Responses** — deferred Anthropic entries are outside its cache key, and OpenAI Responses appends `additional_tools` without changing `tools[]`. | | Function tools without provider-native reveal-item support (`tool_addition_mode=None`) | **May break between turns** — function-tool visibility can change as capabilities load. | | Native tools | **Always breaks the prefix on load** — native tool definitions are part of the request prefix on every provider. | When preserving the cache prefix matters, prefer instruction-only or function-tool-only on-demand capabilities on a model that can express an availability change natively. The provider-specific mechanics that keep the prefix stable live in [Tool search and prompt caching](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-search-caching) . The `Capability` convenience class ---------------------------------- [](https://pydantic.dev/docs/ai/capabilities/on-demand/#the-capability-convenience-class) [`Capability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Capability) bundles instructions, function tools, and toolsets without subclassing. Register tools with the decorator that mirrors [`@agent.tool`](https://pydantic.dev/docs/ai/tools-toolsets/tools/#registering-function-tools-via-decorator) : capability\_decorator.py from pydantic_ai import RunContext from pydantic_ai.capabilities import Capability refunds = Capability( id='refunds', description='Use for refund eligibility and refund status.', instructions='Always confirm the order ID before issuing a refund.', defer_loading=True, ) @refunds.tool def refund_status(ctx: RunContext[None], order_id: str) -> str: """Look up the refund status for an order.""" return f'Order {order_id}: refund issued on 2026-05-01.' In addition to `@capability.tool` and `@capability.tool_plain`, you can pass existing functions or [`Tool`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.Tool) instances via `tools=`, or hand in one or more [toolsets](https://pydantic.dev/docs/ai/tools-toolsets/toolsets/) via `toolsets=`. For dynamic instructions, use the [`@capability.instructions`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Capability.instructions) decorator. For a dynamic catalog entry, pass a callable as `description=`. `@capability.tool` and `@capability.tool_plain` mirror [`@agent.tool`](https://pydantic.dev/docs/ai/tools-toolsets/tools/#registering-function-tools-via-decorator) exactly, including the `defer_loading` argument. On a deferred capability that per-tool flag is a no-op — the capability gates all its tools as a unit — so it only has an effect on a non-deferred `Capability`, where it opts an individual tool into [tool search](https://pydantic.dev/docs/ai/tools-toolsets/tools-advanced/#tool-search) discovery. For anything beyond instructions, function tools, toolsets, and descriptions — model settings, hooks, native tools, wrapper toolsets, or custom per-run logic — subclass [`AbstractCapability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability) directly. When subclassing, override [`get_description`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability.get_description) if the catalog entry needs to vary by run. Beyond instructions: tools, settings, hooks, native tools --------------------------------------------------------- [](https://pydantic.dev/docs/ai/capabilities/on-demand/#beyond-instructions) The [`Capability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Capability) example above deferred instructions and a function tool, but the same flag gates the whole bundle — what the model knows, what it can do, and how it does it (see [What you can defer](https://pydantic.dev/docs/ai/capabilities/on-demand/#what-you-can-defer) ). The snippets below show the remaining pieces in turn: model settings, hooks, and native tools. ### Deferred model settings [](https://pydantic.dev/docs/ai/capabilities/on-demand/#deferred-model-settings) [`get_model_settings`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability.get_model_settings) is collected during capability assembly, but its settings are only applied after the deferred capability is loaded. That means per-step settings like raised reasoning effort only apply for workflows the model opts into: deferred\_model\_settings.py from dataclasses import dataclass from typing import Any from pydantic_ai import Agent, ModelSettings from pydantic_ai.capabilities import AbstractCapability @dataclass class DeepReasoning(AbstractCapability[Any]): def get_model_settings(self) -> ModelSettings: return ModelSettings(extra_body={'reasoning_effort': 'high'}) agent = Agent( 'openai-responses:gpt-5.4', capabilities=[\ DeepReasoning(\ id='deep-reasoning',\ description='Use for multi-step planning or hard analytical problems.',\ defer_loading=True,\ ),\ ], ) ### Lifecycle hooks with deferred workflows [](https://pydantic.dev/docs/ai/capabilities/on-demand/#lifecycle-hooks-with-deferred-workflows) Hooks can live on deferred capabilities too. They do not run until the model loads the capability that owns them: deferred\_hooks.py from dataclasses import dataclass from pydantic_ai import Agent from pydantic_ai.capabilities import AbstractCapability @dataclass class AccountSecurityWorkflow(AbstractCapability[None]): id: str = 'account-security' description: str = 'Use when the next action may be destructive.' defer_loading: bool = True def get_instructions(self) -> str: return 'Confirm the customer identity before taking destructive action.' async def before_tool_execute(self, ctx, *, call, tool_def, args): # Inspect the call, prompt the operator, raise to block. return args agent = Agent('openai-responses:gpt-5.4', capabilities=[AccountSecurityWorkflow()]) ### Deferred native tools [](https://pydantic.dev/docs/ai/capabilities/on-demand/#deferred-native-tools) Any [native capability](https://pydantic.dev/docs/ai/capabilities/overview/#built-in-capabilities) (`WebSearch`, `WebFetch`, `MCP`, …) can be deferred the same way. The native tool definition only enters the request after the `load_capability` tool loads the capability — see [Cache implications](https://pydantic.dev/docs/ai/capabilities/on-demand/#cache-implications) for the trade-off: deferred\_native\_tool.py from pydantic_ai import Agent from pydantic_ai.capabilities import WebSearch agent = Agent( 'anthropic:claude-sonnet-4-6', capabilities=[\ WebSearch(\ local='duckduckgo',\ id='web-research',\ description='Use when the question requires up-to-date information.',\ defer_loading=True,\ ),\ ], ) Putting it together: a multi-workflow support agent --------------------------------------------------- [](https://pydantic.dev/docs/ai/capabilities/on-demand/#putting-it-together-a-multi-workflow-support-agent) A realistic on-demand capability rarely consists of just one piece. The example below defines a customer-support agent with two deferred workflows that exercise different parts of the bundle: * `orders` — instructions plus a function tool, defined inline with [`Capability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Capability) . * `account-security` — instructions, a function tool, raised reasoning effort, _and_ an approval hook, all bundled as one [`AbstractCapability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability) subclass. For those workflows, turn 1 exposes only the two-line catalog. Base instructions, always-on tools, the framework-managed `load_capability` tool, and any provider/tool-search plumbing still appear as usual. Loading `account-security` activates the runbook, the destructive tool, the higher reasoning effort, _and_ the approval gate together — that’s what we mean by bundle-level disclosure. support\_agent.py from dataclasses import dataclass from pydantic_ai import Agent, ModelSettings, RunContext from pydantic_ai.capabilities import AbstractCapability, Capability from pydantic_ai.toolsets import AgentToolset, FunctionToolset @dataclass class Store: orders: dict[str, str] # Workflow 1: instructions + function tool, defined inline. orders = Capability[Store]( id='orders', description='Use for order tracking, delivery status, or questions involving an order ID.', instructions='Quote the order ID and item name when discussing an order.', defer_loading=True, ) @orders.tool def order_status(ctx: RunContext[Store], order_id: str) -> str: """Look up shipping or delivery status for an order.""" return ctx.deps.orders.get(order_id, f'No order found with id {order_id}.') # Workflow 2: instructions + tool + per-step model settings + approval hook, # all hidden until the model loads `account-security`. security_tools = FunctionToolset[Store]() @security_tools.tool def revoke_sessions(ctx: RunContext[Store], account_id: str) -> str: """Revoke all active sessions for an account.""" return f'Revoked sessions for {account_id}.' @dataclass class AccountSecurity(AbstractCapability[Store]): id: str = 'account-security' description: str = 'Use for suspicious logins, account takeover, or session revocation.' defer_loading: bool = True def get_instructions(self) -> str: return 'Confirm the customer identity before revoking sessions.' def get_toolset(self) -> AgentToolset[Store]: return security_tools def get_model_settings(self) -> ModelSettings: # Raise reasoning effort just for sensitive workflows. return ModelSettings(extra_body={'reasoning_effort': 'high'}) async def before_tool_execute(self, ctx, *, call, tool_def, args): # Approval gate: inspect the call and raise to block, active once the model has loaded `account-security`. return args support_agent = Agent( 'openai-responses:gpt-5.4', deps_type=Store, instructions='You are a customer-support agent for an e-commerce store.', capabilities=[orders, AccountSecurity()], ) A “where is my order?” request loads only `orders`. A “someone is logging into my account” request loads only `account-security` — and from that point on, every tool call in the run passes through the approval hook _and_ benefits from the raised reasoning effort, without either being visible to the model on requests that never touched the workflow. Enforcing read-before-act ------------------------- [](https://pydantic.dev/docs/ai/capabilities/on-demand/#enforcing-read-before-act) Want the model to actually _read the runbook_ before taking a destructive action? Make the runbook a deferred capability, then check `ctx.loaded_capability_ids` in a one-method hook: runbook\_required.py from dataclasses import dataclass, field from pydantic_ai import Agent, ModelRetry from pydantic_ai.capabilities import AbstractCapability, Capability @dataclass class RunbookRequired(AbstractCapability[None]): """Bounces a tool call back until the matching runbook has been loaded.""" requirements: dict[str, str] = field(default_factory=dict) async def before_tool_execute(self, ctx, *, call, tool_def, args): required = self.requirements.get(tool_def.name) if required and required not in ctx.loaded_capability_ids: raise ModelRetry( f'Call the `load_capability` tool with `id={required!r}` and follow its ' f'guidance before calling `{tool_def.name}`.' ) return args refund_policy = Capability( id='refund-policy', description='Read before issuing refunds. Eligibility rules and approval limits.', instructions=( 'Refunds over $500 require manager approval. ' 'Refunds outside the 30-day window require a documented exception.' ), defer_loading=True, ) agent = Agent( 'openai-responses:gpt-5.4', capabilities=[\ refund_policy,\ RunbookRequired(requirements={'issue_refund': 'refund-policy'}),\ ], ) @agent.tool_plain def issue_refund(order_id: str, amount: float) -> str: """Issue a refund for an order.""" return f'Refund of ${amount} issued for {order_id}.' The model sees `issue_refund` from turn 1. If it tries to call it before opening `refund-policy`, the hook bounces the call back with a message pointing at the exact `load_capability` tool call to make. The model loads the policy, the policy text lands in its recent context, and the refund runs _within_ the rules — and only then. Same shape for any tool-and-runbook pair. Because the loaded set is just runtime data on [`RunContext`](https://pydantic.dev/docs/ai/api/pydantic-ai/tools/#pydantic_ai.tools.RunContext) , the pattern generalises: dynamic instructions can warn when a risky pair of workflows is open, audit hooks can tag traces with the loaded set, escalation hooks can require an extra confirmation when both `payments` and `account-security` are active. Loading skills from Markdown files ---------------------------------- [](https://pydantic.dev/docs/ai/capabilities/on-demand/#loading-skills-from-markdown-files) If you already keep your skills as Markdown files with YAML frontmatter — the format used by [Anthropic Agent Skills](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) — you can wrap each one in a [`Capability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.Capability) with a few lines of glue. Given a skill file `skills/refunds.md`: skills/refunds.md --- id: refunds description: Use for refund eligibility, refund status, or processing a refund. --- Always confirm the order ID before issuing a refund. Never issue refunds over $500 without manager approval. Load it into an agent as an on-demand capability: skill\_from\_markdown.py from pathlib import Path import yaml from pydantic_ai import Agent from pydantic_ai.capabilities import Capability def load_skill(path: Path) -> Capability: _, frontmatter, body = path.read_text().split('---', 2) meta = yaml.safe_load(frontmatter) return Capability( id=meta['id'], description=meta['description'], instructions=body.strip(), defer_loading=True, ) agent = Agent( 'openai-responses:gpt-5.4', instructions='You are a customer support assistant.', capabilities=[load_skill(p) for p in Path('skills').glob('*.md')], ) Each file shows up in the model’s catalog as its `id` plus `description`; the body is only sent once the model calls the `load_capability` tool. To go beyond instructions — add function tools, model settings, or hooks for a particular skill — subclass [`AbstractCapability`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.AbstractCapability) as in the examples above. Was this page helpful? Thanks for your feedback! --- # Camera Agent | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/examples/realtime/realtime-camera/#_top) Camera Agent ============ This camera agent streams microphone audio and one camera frame per second into a [realtime session](https://pydantic.dev/docs/ai/realtime/overview/) , then plays and captions the spoken response. Point it at objects to ask about them, enable _Watch_ for proactive narration, or show it a sketch to redraw. The example demonstrates: * provider-agnostic [realtime sessions](https://pydantic.dev/docs/ai/realtime/overview/) with profile-derived PCM sample rates * [image input](https://pydantic.dev/docs/ai/realtime/audio/#images) using [`BinaryContent`](https://pydantic.dev/docs/ai/api/pydantic-ai/messages/#pydantic_ai.messages.BinaryContent) * live vision with `turn_coverage='all_input'` and a _Watch_ toggle * a regular function tool that delegates diagram rendering to a second [`Agent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent) * [web search](https://pydantic.dev/docs/ai/realtime/tools/#native-tools) with [`WebSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebSearch) and clickable citations * a model picker and provider-aware voice, modality, VAD, and Gemini settings Running the Example ------------------- [](https://pydantic.dev/docs/ai/examples/realtime/realtime-camera/#running-the-example) Add credentials for a picker model to the repository-root `.env`, for example: GOOGLE_API_KEY=your-google-api-key The sketch-redraw tool delegates to a separate drawing agent — `google:gemini-3.5-flash` by default, which reuses the same `GOOGLE_API_KEY`. Set `CAMERA_DRAW_MODEL` to any other `provider:model` your credentials cover, or `CAMERA_DRAW=false` to disable drawing; the rest of the assistant works either way. With [dependencies installed and environment variables set](https://pydantic.dev/docs/ai/examples/setup/#usage) , start the local server: * [pip](https://pydantic.dev/docs/ai/examples/realtime/realtime-camera/#tab-panel-48) * [uv](https://pydantic.dev/docs/ai/examples/realtime/realtime-camera/#tab-panel-49) Terminal python -m pydantic_ai_examples.realtime_camera.app Terminal uv run -m pydantic_ai_examples.realtime_camera.app Open [http://localhost:8000](http://localhost:8000/) , select **Start**, and allow camera and microphone access. The model defaults to `google:gemini-3.1-flash-live-preview`; set `CAMERA_REALTIME_MODEL` to change it, or use the picker to switch to any Google, OpenAI, or Azure OpenAI `provider:model` per session (xAI realtime doesn’t support camera image input). The selected model’s realtime profile supplies the browser’s PCM input and output sample rates: Gemini input uses 16 kHz, while OpenAI and Azure input uses 24 kHz. Watch mode ---------- [](https://pydantic.dev/docs/ai/examples/realtime/realtime-camera/#watch-mode) Camera frames add visual context but do not start a model turn. _Watch_ periodically sends a short text turn while the model is idle, prompting it to report a visual change without interrupting speech already in progress. Set `CAMERA_WATCH_PROMPT` to customize that instruction. Gemini native-audio models can decide that nothing needs saying: Terminal export CAMERA_PROACTIVE=true export CAMERA_AFFECTIVE=true `CAMERA_TURN_COVERAGE` defaults to `all_input`, which works with both the Gemini Developer API and Vertex AI. Watch mode consumes tokens while enabled. Search and citations -------------------- [](https://pydantic.dev/docs/ai/examples/realtime/realtime-camera/#search-and-citations) With `CAMERA_WEB_SEARCH=true` (the default), the example adds [`WebSearch`](https://pydantic.dev/docs/ai/api/pydantic-ai/capabilities/#pydantic_ai.capabilities.WebSearch) when the selected model profile supports native search. Native-tool return events are converted into citation chips; the browser accepts only HTTP(S) source URLs. Redraw a diagram ---------------- [](https://pydantic.dev/docs/ai/examples/realtime/realtime-camera/#redraw-a-diagram) With `CAMERA_DRAW=true` (the default), the realtime agent can call `redraw_diagram`. It gives a detailed textual description of the visible sketch to a separate [`Agent`](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.Agent) , which produces self-contained HTML. The browser displays that HTML in an opaque-origin iframe that blocks scripts and network access, and retains the PNG export action. The default is a fast small model because the user is waiting on a live call: the redraw’s latency is dominated by HTML output tokens, so a larger model mostly adds thinking time, not quality. Configure the drawing model independently: Terminal export CAMERA_DRAW_MODEL=anthropic:claude-haiku-4-5 Drawing and web search remain enabled together when the selected realtime model supports both. Tools [run concurrently](https://pydantic.dev/docs/ai/api/pydantic-ai/agent/#pydantic_ai.agent.AgentRealtime.session) , so drawing does not replace the voice conversation. Vertex AI --------- [](https://pydantic.dev/docs/ai/examples/realtime/realtime-camera/#vertex-ai) Use Application Default Credentials when your organization does not allow Gemini API keys: Terminal gcloud auth application-default login export GOOGLE_GENAI_USE_VERTEXAI=true export GOOGLE_CLOUD_PROJECT=your-project export GOOGLE_CLOUD_LOCATION=us-central1 How the bridge works -------------------- [](https://pydantic.dev/docs/ai/examples/realtime/realtime-camera/#how-the-bridge-works) The browser and provider are connected by two small concurrent pumps in `_run_session`: browser ── PCM16 + JPEG/text ──▶ FastAPI /ws ──▶ RealtimeSession browser ◀── PCM16 + JSON events ──────────────── RealtimeSession Before microphone capture begins, the server sends `session_config` over the JSON channel with the profile-derived audio rates. The inbound pump then forwards PCM, image, text, and Watch messages. The event pump returns audio, transcripts, barge-in notifications, grounding citations, drawing updates, and turn completion. Either side ending cancels the other pump and closes the session cleanly. Example Code ------------ [](https://pydantic.dev/docs/ai/examples/realtime/realtime-camera/#example-code) The server contains the realtime bridge and the subordinate Watch, grounding, and drawing helpers: app.py from __future__ import annotations import base64 import json import os import re from collections.abc import Awaitable, Callable, Mapping from contextlib import suppress from dataclasses import dataclass from functools import lru_cache from pathlib import Path from typing import cast from urllib.parse import urlsplit import anyio import logfire from dotenv import load_dotenv from fastapi import FastAPI, WebSocket, WebSocketDisconnect from fastapi.responses import HTMLResponse from pydantic_ai import ( Agent, BinaryContent, PartDeltaEvent, PartEndEvent, RunContext, SpeechPartDelta, ) from pydantic_ai.capabilities import WebSearch from pydantic_ai.exceptions import ModelAPIError, UserError from pydantic_ai.messages import NativeToolReturnPart, TextPartDelta from pydantic_ai.native_tools import WebSearchTool from pydantic_ai.realtime import ( RealtimeError, RealtimeEvent, RealtimeInputSpeechStartEvent, RealtimeModel, RealtimeModelSettings, RealtimeResponseInterruptedEvent, RealtimeSession, RealtimeTurnCompleteEvent, ReconnectPolicy, TurnDetection, infer_realtime_model, ) from pydantic_ai.realtime.google import ( AutomaticVAD, GoogleRealtimeModel, GoogleRealtimeModelSettings, ) from pydantic_ai.realtime.openai import ( OpenAIRealtimeModel, OpenAIRealtimeModelSettings, ) load_dotenv() # 'if-token-present' means nothing will be sent (and the example will work) if you don't have logfire configured. # Configure after `load_dotenv()` so a `LOGFIRE_TOKEN` in `.env` is picked up. logfire.configure(send_to_logfire='if-token-present') logfire.instrument_pydantic_ai() def _truthy(value: str | None) -> bool: """Parse an env/query flag: `'1'`, `'true'`, `'yes'`, or `'on'` (any case) mean enabled.""" return (value or '').lower() in ('1', 'true', 'yes', 'on') MODEL = os.environ.get('CAMERA_REALTIME_MODEL', 'google:gemini-3.1-flash-live-preview') # Empty by default so each provider picks its own default voice — no need to change it when switching # between Gemini and OpenAI, whose voice names differ (Gemini rejects `alloy`, OpenAI rejects `Puck`). VOICE = os.environ.get('CAMERA_REALTIME_VOICE', '') # Use Vertex AI (Application Default Credentials) instead of a Gemini API key — handy where org # policy disallows API keys. Needs `gcloud auth application-default login` + `GOOGLE_CLOUD_PROJECT`. USE_VERTEX = _truthy(os.environ.get('GOOGLE_GENAI_USE_VERTEXAI')) # `all_input` keeps every camera frame in the model's context — the live scene the assistant reasons # about — and works on both the Gemini Developer API and Vertex AI (the newer `all_video` doesn't yet). TURN_COVERAGE = os.environ.get('CAMERA_TURN_COVERAGE', 'all_input') # Gemini native-audio-only knobs, off by default so the default model still connects: proactive audio # lets the model stay silent when a Watch nudge finds nothing new; affective dialog adapts delivery # to emotion in the conversation. PROACTIVE = _truthy(os.environ.get('CAMERA_PROACTIVE')) AFFECTIVE = _truthy(os.environ.get('CAMERA_AFFECTIVE')) # Sketch-to-diagram: the `redraw_diagram` tool passes the realtime model's text description of a # sketch to a separate drawing agent that renders it as clean HTML. The default drawing model reuses # the `GOOGLE_API_KEY` the default realtime model already needs, and is a fast small model because # the user is waiting on a live call: output tokens dominate the redraw's latency, and a larger # model mostly adds thinking time. `CAMERA_DRAW_MODEL` takes any `provider:model` string. DRAW = _truthy(os.environ.get('CAMERA_DRAW', 'true')) DRAW_MODEL = os.environ.get('CAMERA_DRAW_MODEL', 'google:gemini-3.5-flash') # Web search (the `WebSearch` capability) — on by default, but only enabled for a session when the # selected model supports web search natively (see `_web_search_supported`), so switching models # drops the capability instead of failing the session. WEB_SEARCH = _truthy(os.environ.get('CAMERA_WEB_SEARCH', 'true')) WATCH_PROMPT = os.environ.get( 'CAMERA_WATCH_PROMPT', "Look at the current camera view. In a few words, say what's changed since you last spoke; " 'if nothing notable changed, stay silent.', ) _INDEX_PATH = Path(__file__).parent / 'index.html' def _same_origin(socket: WebSocket) -> bool: """Accept browser WebSockets only from the origin serving this development example. Any web page can open a WebSocket to this server (which spends your API credits), so the browser-reported `Origin` must match the host the request was addressed to. Three ways in: a loopback origin matching `Host` (direct local use); an origin matching `X-Forwarded-Host` (a reverse proxy such as Codespaces or a dev tunnel — trustworthy because the browser WebSocket API cannot send custom headers, so its presence proves a real proxy hop); or an origin listed in `CAMERA_ALLOWED_ORIGINS` (comma-separated `scheme://host[:port]`, for proxies that forward neither). """ origin = socket.headers.get('origin') if not origin: return False allowed = os.environ.get('CAMERA_ALLOWED_ORIGINS', '') if origin in {value.strip() for value in allowed.split(',') if value.strip()}: return True parsed = urlsplit(origin) if parsed.scheme not in ('http', 'https'): return False if parsed.netloc == socket.headers.get('x-forwarded-host'): return True return parsed.hostname in ( 'localhost', '127.0.0.1', '::1', ) and parsed.netloc == socket.headers.get('host') def _instructions(*, web_search: bool) -> str: """The assistant's instructions, built per connection. The web-search guidance is included only when web search is actually enabled for the selected model (see `_web_search_supported`), so the model isn't told about a tool it doesn't have. """ return ( 'You are a friendly, concise voice assistant. The user is talking to you and may show you things ' 'through their camera — when relevant, describe and reason about what you can see. Keep replies ' 'short and natural, like a conversation.' + ( ' Search the web when a question needs current or external facts.' if web_search else '' ) + ( ' You can redraw a hand-drawn sketch the user shows you — a diagram, system design, flow ' 'chart, or wireframe — into a clean version with the `redraw_diagram` tool. Do NOT call it ' 'the moment you see a drawing. First make sure you understand what they actually want: if ' "they haven't said, ask one short question — keep it faithful but tidier, turn it into a " 'flowchart, restructure it, add or label something? Once their intent is clear, FIRST tell ' "them out loud that you're about to redraw it and that it takes a few moments (around ten" "seconds) — don't leave them waiting in silence — THEN call the tool. The drawing tool " 'cannot see the camera, so pass it a thorough text description as `instructions`: every box ' 'and its label, every arrow and what it connects, groupings, and the overall layout, plus ' 'what the user asked you to change. Be specific — it can only draw what you describe. ' 'After calling the tool, stop talking until its result arrives — never say the redraw is ' 'done in the same breath as calling it, because the drawing takes several seconds. Once ' 'the result arrives, briefly describe what you drew.' if DRAW else '' ) ) @dataclass class CameraDeps: """Per-connection hooks the `redraw_diagram` tool needs. `emit` pushes a JSON message back to this connection's browser — the tool uses it to show and then clear the drawing overlay while the diagram is being generated. """ emit: Callable[[dict[str, object]], Awaitable[None]] app = FastAPI() logfire.instrument_fastapi(app) DRAW_INSTRUCTIONS = ( 'You turn a text description of a hand-drawn sketch — a diagram, system design, flow chart, or ' 'wireframe — into a clean, modern, self-contained HTML page that recreates and tidies up the ' 'drawing. Faithfully render every box, label, arrow, and connection the description mentions, ' 'and lay everything out neatly with clear typography, generous spacing, and restrained color on ' 'a light background. ' 'Design it to fit comfortably on a phone screen in portrait: prefer a vertical flow over very ' 'wide horizontal layouts, let content wrap, and use relative widths so nothing is cut off. ' # The user is waiting on a live call while this generates, so latency is part of the spec: # output tokens dominate the wall-clock time, and a compact page halves it. 'Keep the page LEAN so it generates fast: one short `
Lens Camera Assistant
tap start
Tap Start, then talk and point your camera at things to ask about them. Show it a hand-drawn sketch and ask it to redraw it as a clean diagram. The model sees your camera the whole time; toggle Watch to have it speak up on its own when the scene changes, instead of only when you ask.
Was this page helpful? Thanks for your feedback! --- # Online Evaluation | Pydantic Docs [Skip to content](https://pydantic.dev/docs/ai/evals/online-evaluation/#_top) Online Evaluation ================= Online evaluation lets you attach evaluators to production (or staging) functions so that every call (or a sampled subset) is automatically evaluated in the background. The same [`Evaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Evaluator) classes used with [`Dataset.evaluate()`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset.evaluate) work here; the difference is just in how they’re wired up. When to Use Online Evaluation ----------------------------- [](https://pydantic.dev/docs/ai/evals/online-evaluation/#when-to-use-online-evaluation) Online evaluation is useful when you want to: * **Monitor production quality:** continuously score LLM outputs against rubrics * **Catch regressions:** detect degradation in agent behavior across deploys * **Collect evaluation data:** build datasets from real traffic for offline analysis * **Control costs:** sample expensive LLM judges on a fraction of traffic while running cheap checks on everything For testing against curated datasets before deployment, use [offline evaluation](https://pydantic.dev/docs/ai/evals/getting-started/quick-start/) with [`Dataset.evaluate()`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset.evaluate) instead. Quick Start ----------- [](https://pydantic.dev/docs/ai/evals/online-evaluation/#quick-start) The [`evaluate()`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.evaluate) decorator attaches evaluators to any function. Evaluators run in the background without blocking the caller, and results are emitted as [OpenTelemetry events](https://pydantic.dev/docs/ai/evals/online-evaluation/#default-otel-event-emission) : from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext from pydantic_evals.online import evaluate @dataclass class OutputNotEmpty(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return bool(ctx.output) @evaluate(OutputNotEmpty()) async def summarize(text: str) -> str: return f'Summary of: {text}' Wire up OTel export (e.g. [`logfire.configure()`](https://pydantic.dev/docs/ai/integrations/logfire/#using-logfire) ) elsewhere in your application startup so that the emitted `gen_ai.evaluation.result` events reach your backend. When using [Pydantic Logfire](https://pydantic.dev/docs/logfire/evaluate/live-evals/) , these events surface in the **Live Evaluations** view, where you can browse results by target, drill into the originating trace, and watch scores over a time window. Each decorated call emits one `gen_ai.evaluation.result` OTel event per evaluator result, following the [OTel GenAI evaluation semconv](https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-events/#event-gen_aievaluationresult) . This mirrors how offline evaluation emits OTel spans via `logfire.span`: if any OTel SDK is configured in the process (via [`logfire.configure()`](https://pydantic.dev/docs/ai/integrations/logfire/#using-logfire) , the OTel SDK directly, or a vendor instrumentation), events flow to your backend; if not, emission is a cheap no-op. To additionally handle results in Python code — for alerting, bespoke aggregation, in-memory test capture, or non-OTel destinations — register a [sink](https://pydantic.dev/docs/ai/evals/online-evaluation/#sinks) . Sinks run _in addition to_ OTel event emission. The module-level [`configure()`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.configure) and [`evaluate()`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.evaluate) functions delegate to a global [`OnlineEvalConfig`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.OnlineEvalConfig) . For multiple configurations or isolated setups, create your own config instances (see [OnlineEvalConfig](https://pydantic.dev/docs/ai/evals/online-evaluation/#onlineevalconfig) below). Target ------ [](https://pydantic.dev/docs/ai/evals/online-evaluation/#target) Each decorated function (or agent) emits results tagged with a **target** — a name that groups results in downstream sinks and dashboards. By default the target is the decorated function’s `__name__`, but you can override it with `target=...`: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext from pydantic_evals.online import evaluate @dataclass class OutputNotEmpty(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return bool(ctx.output) # Default: target='summarize' (function name) @evaluate(OutputNotEmpty()) async def summarize(text: str) -> str: ... # Override: use a friendly name @evaluate(OutputNotEmpty(), target='customer_support') async def run_agent(prompt: str) -> str: ... The target name is supplied to sinks on every `submit()` call as a plain `str` — a single sink instance handles any number of decorated functions or agents. For agent capabilities, the target name is taken from the agent’s own `name` attribute (see [Agent Integration](https://pydantic.dev/docs/ai/evals/online-evaluation/#agent-integration) ); to categorize or route on agent-ness, add metadata on the config (e.g., `metadata={'kind': 'agent'}`). Core Concepts ------------- [](https://pydantic.dev/docs/ai/evals/online-evaluation/#core-concepts) ### OnlineEvaluator [](https://pydantic.dev/docs/ai/evals/online-evaluation/#onlineevaluator) Different evaluators need different settings. A cheap heuristic could run on 100% of traffic; an expensive LLM judge might run on 1%. [`OnlineEvaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.OnlineEvaluator) wraps an [`Evaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Evaluator) with per-evaluator configuration: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext, LLMJudge from pydantic_evals.online import OnlineEvaluator @dataclass class IsHelpful(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return len(str(ctx.output)) > 10 # Cheap evaluator: run on every request always_check = OnlineEvaluator(evaluator=IsHelpful(), sample_rate=1.0) # Expensive evaluator: run on 1% of requests, limit concurrency rare_check = OnlineEvaluator( evaluator=LLMJudge(rubric='Is the response helpful?'), sample_rate=0.01, max_concurrency=5, ) When you pass a bare [`Evaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Evaluator) to the [`evaluate()`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.evaluate) decorator, it’s automatically wrapped in an [`OnlineEvaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.OnlineEvaluator) with the config’s default sample rate. ### OnlineEvalConfig [](https://pydantic.dev/docs/ai/evals/online-evaluation/#onlineevalconfig) [`OnlineEvalConfig`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.OnlineEvalConfig) holds cross-evaluator defaults (sample rate, metadata, optional additional sinks, OTel-emission toggle). There’s a global default instance, plus you can create custom instances for different configurations: import asyncio from collections.abc import Sequence from dataclasses import dataclass from pydantic_evals.evaluators import ( EvaluationResult, Evaluator, EvaluatorContext, EvaluatorFailure, ) from pydantic_evals.online import OnlineEvalConfig, wait_for_evaluations @dataclass class IsNonEmpty(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return bool(ctx.output) results_log: list[str] = [] async def log_sink( results: Sequence[EvaluationResult], failures: Sequence[EvaluatorFailure], context: EvaluatorContext, ) -> None: for r in results: results_log.append(f'{r.name}={r.value}') my_eval = OnlineEvalConfig( default_sink=log_sink, default_sample_rate=1.0, metadata={'service': 'my-app'}, ) @my_eval.evaluate(IsNonEmpty()) async def my_function(query: str) -> str: return f'Answer to: {query}' async def main(): result = await my_function('What is 2+2?') print(result) #> Answer to: What is 2+2? await wait_for_evaluations() print(results_log) #> ['IsNonEmpty=True'] asyncio.run(main()) ### Sinks [](https://pydantic.dev/docs/ai/evals/online-evaluation/#sinks) OTel event emission is the default observability surface for online evaluation (see [Default OTel event emission](https://pydantic.dev/docs/ai/evals/online-evaluation/#default-otel-event-emission) ). Sinks are for _additional_ handling in Python code — in-memory test capture, alerting, fan-out to non-OTel destinations, or bespoke aggregation. [`EvaluationSink`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.EvaluationSink) is the protocol; multiple sinks can be registered on a single config. The built-in [`CallbackSink`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.CallbackSink) wraps any callable (sync or async) that accepts results, failures, and context. You can also pass a bare callable wherever a sink is expected — it’s auto-wrapped in a [`CallbackSink`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.CallbackSink) . For custom sinks, implement the [`EvaluationSink`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.EvaluationSink) protocol. Each `submit()` call receives a [`SinkPayload`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.SinkPayload) bundling the results, failures, context, span reference, and target from one or more evaluators that ran for a given function call: from pydantic_evals.online import SinkPayload class PrintSink: """Prints evaluation results to stdout.""" async def submit(self, payload: SinkPayload) -> None: for r in payload.results: version = f' ({r.evaluator_version})' if r.evaluator_version else '' print(f' [{payload.target}] {r.name}{version}: {r.value}') for f in payload.failures: version = f' ({f.evaluator_version})' if f.evaluator_version else '' print(f' [{payload.target}] FAILED {f.name}{version}: {f.error_message}') `payload.results` and `payload.failures` may cover one or more evaluators from a single function call — when multiple evaluators share a sink, their results are batched into a single `submit()` call. Each result carries its own attribution (name, `evaluator_version` on [`EvaluationResult`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluationResult) and [`EvaluatorFailure`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorFailure) , and source spec), so sinks can separate them downstream; see [Evaluator Versioning](https://pydantic.dev/docs/ai/evals/online-evaluation/#evaluator-versioning) . The `payload.target` identifies the function or agent being evaluated (see [Target](https://pydantic.dev/docs/ai/evals/online-evaluation/#target) ). ### Default OTel event emission [](https://pydantic.dev/docs/ai/evals/online-evaluation/#default-otel-event-emission) Every dispatched evaluator emits one `gen_ai.evaluation.result` OTel log event per [`EvaluationResult`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluationResult) or [`EvaluatorFailure`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorFailure) , unconditionally — no sink registration required. Events are parented to the span that produced them, so they appear nested under the original function call in the trace. If no OTel SDK is configured in the process, emission is a cheap no-op. Each event has `event.name = 'gen_ai.evaluation.result'` and a short human-readable body (e.g. `evaluation: accuracy=0.87`, or `evaluation: accuracy failed: `). Emission follows the [OpenTelemetry GenAI evaluation semconv](https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-events/#event-gen_aievaluationresult) , with these attributes: * `gen_ai.evaluation.name` — the evaluator class name when `evaluate()` returns a scalar, or the mapping key when it returns `{'accuracy': ..., 'score': ...}`. Source: [`EvaluationResult.name`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluationResult) / [`EvaluatorFailure.name`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorFailure) . * `gen_ai.evaluation.score.value` — populated for `bool` (`True`→`1.0`, `False`→`0.0`) and numeric returns. Omitted for `str` returns. * `gen_ai.evaluation.score.label` — populated for `bool` (`True`→`'pass'`, `False`→`'fail'`) and `str` returns (used directly as the label). Omitted for numeric returns. * `gen_ai.evaluation.explanation` — `EvaluationResult.reason` on success or `EvaluatorFailure.error_message` on failure. Omitted when absent. Set via `reason=...` when constructing an `EvaluationResult` inside a custom evaluator. * `error.type` (failure events only) — the exception class name (e.g. `'ValueError'`) when the failure was built from a caught exception; falls back to `'pydantic_evals.EvaluatorFailure'` for `EvaluatorFailure` instances constructed without it. Absent on successful evaluations. Source: `EvaluatorFailure.error_type`. * `gen_ai.evaluation.target` — `@evaluate(target=...)` or agent `name`. See [Target](https://pydantic.dev/docs/ai/evals/online-evaluation/#target) . * `gen_ai.evaluation.evaluator.version` — `Evaluator.evaluator_version` class attribute; omitted when the class doesn’t set it. See [Evaluator Versioning](https://pydantic.dev/docs/ai/evals/online-evaluation/#evaluator-versioning) . * `gen_ai.evaluation.evaluator.source` — JSON-serialized [`EvaluatorSpec`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorSpec) identifying the evaluator class and its constructor arguments, so downstream queries can group by evaluator identity without relying on `name` alone (two different `LLMJudge(rubric=...)` instances share a name but have different sources). [OTel baggage](https://pydantic.dev/docs/logfire/reference/baggage/) entries (if any) are also attached to each event as attributes — configurable via `include_baggage` on the config. The `gen_ai.*` and `error.type` attributes above always win on conflict with baggage. For example, the `OutputNotEmpty` evaluator above, decorated as `@evaluate(OutputNotEmpty(), target='customer_support')` and returning `True` for a given call, emits one event with: * `gen_ai.evaluation.name = 'OutputNotEmpty'` * `gen_ai.evaluation.score.value = 1.0` * `gen_ai.evaluation.score.label = 'pass'` * `gen_ai.evaluation.target = 'customer_support'` * `gen_ai.evaluation.evaluator.source = '{"name":"OutputNotEmpty","arguments":null}'` An evaluator with constructor arguments gets those rendered into `source` — e.g. [`LLMJudge(rubric='Is the response helpful?')`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.LLMJudge) emits `gen_ai.evaluation.evaluator.source = '{"name":"LLMJudge","arguments":["Is the response helpful?"]}'`, so two `LLMJudge` instances with different rubrics remain distinguishable downstream. Attributes under `gen_ai.evaluation.evaluator.*` are pydantic-evals extensions — they aren’t in the current OTel GenAI semconv, and their names may change to align with future semconv additions. To disable the default emission (e.g. in a test harness that only wants to assert on a custom sink), set `emit_otel_events=False` on the config: from pydantic_evals.online import OnlineEvalConfig config = OnlineEvalConfig(emit_otel_events=False) #### Evaluator Versioning [](https://pydantic.dev/docs/ai/evals/online-evaluation/#evaluator-versioning) Override [`get_evaluator_version`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Evaluator.get_evaluator_version) on an [`Evaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Evaluator) subclass to stamp every result it emits with a version string — surfaced as `gen_ai.evaluation.evaluator.version` on emitted events and as `evaluator_version` on each [`EvaluationResult`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluationResult) and [`EvaluatorFailure`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorFailure) . This lets trend lines and dashboards filter out results produced by retired evaluator versions without deleting historical rows — useful when you change an LLM judge’s prompt or rework a heuristic in a way that invalidates prior scores: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext @dataclass class ToneCheck(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> str: return 'neutral' def get_evaluator_version(self) -> str | None: return 'v2' # bumped after prompt rewrite The version applies to all results the evaluator produces (so one evaluator class maps to one version, even when the evaluator returns a mapping of named results). Sampling -------- [](https://pydantic.dev/docs/ai/evals/online-evaluation/#sampling) Control evaluation frequency with per-evaluator sample rates to balance quality monitoring against cost. ### Static Sample Rates [](https://pydantic.dev/docs/ai/evals/online-evaluation/#static-sample-rates) A `sample_rate` between 0.0 and 1.0 sets the probability of evaluating each call: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext from pydantic_evals.online import OnlineEvaluator @dataclass class QuickCheck(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return bool(ctx.output) # Run on every request always = OnlineEvaluator(evaluator=QuickCheck(), sample_rate=1.0) # Run on 10% of requests sometimes = OnlineEvaluator(evaluator=QuickCheck(), sample_rate=0.1) # Never run (effectively disabled) never = OnlineEvaluator(evaluator=QuickCheck(), sample_rate=0.0) ### Dynamic Sample Rates [](https://pydantic.dev/docs/ai/evals/online-evaluation/#dynamic-sample-rates) Pass a callable to enable runtime-configurable or input-dependent sampling. The callable receives a [`SamplingContext`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.SamplingContext) with the evaluator instance, function inputs, config metadata, and a per-call random seed, and returns a `float` (probability) or `bool` (always/never): from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext from pydantic_evals.online import OnlineEvaluator, SamplingContext def get_current_rate(ctx: SamplingContext) -> float: return 0.5 @dataclass class QuickCheck(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return bool(ctx.output) dynamic = OnlineEvaluator(evaluator=QuickCheck(), sample_rate=get_current_rate) This enables integration with feature flags, managed variables, or configuration systems — for example, you could replace `get_current_rate` with a function that reads from a remote config service (such as [Logfire managed variables](https://logfire.pydantic.dev/docs/reference/advanced/managed-variables/) ) at runtime, allowing you to change the probability without redeploying the application. You can also use the [`SamplingContext`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.SamplingContext) to make sampling decisions based on the function inputs: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext from pydantic_evals.online import OnlineEvaluator, SamplingContext def sample_long_inputs(ctx: SamplingContext) -> bool: """Only evaluate calls with long input text.""" return len(str(ctx.inputs.get('text', ''))) > 100 @dataclass class QualityCheck(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return len(str(ctx.output)) > 10 expensive = OnlineEvaluator(evaluator=QualityCheck(), sample_rate=sample_long_inputs) ### Correlated Sampling [](https://pydantic.dev/docs/ai/evals/online-evaluation/#correlated-sampling) By default, each evaluator samples independently. With three evaluators each at 10%, roughly 27% of calls incur evaluation overhead (`1 − 0.9³`). If you’d prefer that the _same_ 10% of calls run _all_ evaluators, set `sampling_mode='correlated'`: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext from pydantic_evals.online import OnlineEvalConfig, OnlineEvaluator @dataclass class CheckA(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return True @dataclass class CheckB(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return True config = OnlineEvalConfig( default_sink=lambda results, failures, ctx: None, sampling_mode='correlated', ) # Both run on the same ~10% of calls check_a = OnlineEvaluator(evaluator=CheckA(), sample_rate=0.1) check_b = OnlineEvaluator(evaluator=CheckB(), sample_rate=0.1) In correlated mode, a single random `call_seed` (uniformly distributed between 0.0 and 1.0) is generated per function call and shared across all evaluators. An evaluator runs when `call_seed < sample_rate`, so lower-rate evaluators’ calls are always a subset of higher-rate ones, and the total overhead probability equals the maximum rate rather than accumulating. The `call_seed` is also available on [`SamplingContext`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.SamplingContext) for custom `sample_rate` callables that want to implement their own correlated logic regardless of mode. ### Disabling Evaluation [](https://pydantic.dev/docs/ai/evals/online-evaluation/#disabling-evaluation) Use [`disable_evaluation()`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.disable_evaluation) to suppress all online evaluation in a scope. This may be useful in tests: import asyncio from collections.abc import Sequence from dataclasses import dataclass from pydantic_evals.evaluators import ( EvaluationResult, Evaluator, EvaluatorContext, EvaluatorFailure, ) from pydantic_evals.online import ( OnlineEvalConfig, disable_evaluation, wait_for_evaluations, ) @dataclass class OutputCheck(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return bool(ctx.output) results_log: list[str] = [] async def log_sink( results: Sequence[EvaluationResult], failures: Sequence[EvaluatorFailure], context: EvaluatorContext, ) -> None: for r in results: results_log.append(f'{r.name}={r.value}') config = OnlineEvalConfig(default_sink=log_sink) @config.evaluate(OutputCheck()) async def my_function(x: int) -> int: return x * 2 async def main(): # Evaluators suppressed inside this block with disable_evaluation(): result = await my_function(21) print(result) #> 42 await wait_for_evaluations() print(f'evaluations run: {len(results_log)}') #> evaluations run: 0 # Evaluators resume outside the block await my_function(21) await wait_for_evaluations() print(f'evaluations run: {len(results_log)}') #> evaluations run: 1 asyncio.run(main()) Conditional Evaluation ---------------------- [](https://pydantic.dev/docs/ai/evals/online-evaluation/#conditional-evaluation) For cost control, you can run expensive evaluation logic conditionally within a single custom evaluator. Return a mapping where you only include keys for checks that have run — checks you don’t want to perform can simply be omitted from the results: import asyncio from collections.abc import Sequence from dataclasses import dataclass from pydantic_evals.evaluators import ( EvaluationResult, Evaluator, EvaluatorContext, EvaluatorFailure, ) from pydantic_evals.online import ( OnlineEvalConfig, wait_for_evaluations, ) results_log: list[str] = [] async def log_sink( results: Sequence[EvaluationResult], failures: Sequence[EvaluatorFailure], context: EvaluatorContext, ) -> None: for r in results: results_log.append(f'{r.name}={r.value}') @dataclass class ConditionalAnalysis(Evaluator): """Runs a cheap check on every call, and an expensive check only on long outputs.""" def evaluate(self, ctx: EvaluatorContext) -> dict[str, float | bool]: output = str(ctx.output) results: dict[str, float | bool] = { 'has_content': len(output) > 0, } # Only run the expensive analysis on long outputs if len(output) > 20: # pretend the following line is expensive.. results['detail_score'] = len(output) / 100.0 return results config = OnlineEvalConfig(default_sink=log_sink) @config.evaluate(ConditionalAnalysis()) async def generate(prompt: str) -> str: return f'Response to: {prompt}' async def main(): await generate('hi') # short output — only cheap check runs await wait_for_evaluations() print(results_log) #> ['has_content=True'] results_log.clear() await generate('tell me a long story about dragons') # long output — both checks run await wait_for_evaluations() print(sorted(results_log)) #> ['detail_score=0.47', 'has_content=True'] asyncio.run(main()) This pattern lets you combine cheap and expensive checks in one evaluator, avoiding unnecessary work when conditions aren’t met. Sync Function Support --------------------- [](https://pydantic.dev/docs/ai/evals/online-evaluation/#sync-function-support) The [`evaluate()`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.evaluate) decorator works with both async and sync functions: import asyncio from collections.abc import Sequence from dataclasses import dataclass from pydantic_evals.evaluators import ( EvaluationResult, Evaluator, EvaluatorContext, EvaluatorFailure, ) from pydantic_evals.online import OnlineEvalConfig, wait_for_evaluations results_log: list[str] = [] async def log_sink( results: Sequence[EvaluationResult], failures: Sequence[EvaluatorFailure], context: EvaluatorContext, ) -> None: for r in results: results_log.append(f'{r.name}={r.value}') @dataclass class OutputCheck(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return bool(ctx.output) config = OnlineEvalConfig(default_sink=log_sink) @config.evaluate(OutputCheck()) def process(text: str) -> str: return text.upper() async def main(): # Sync decorated functions work from async contexts too result = process('hello') print(result) #> HELLO await wait_for_evaluations() print(results_log) #> ['OutputCheck=True'] asyncio.run(main()) Sync decorated functions work from both sync and async contexts. When a running event loop is available, evaluators are dispatched as background tasks on that loop. Otherwise, a background thread with its own event loop is spawned. Per-Evaluator Sink Overrides ---------------------------- [](https://pydantic.dev/docs/ai/evals/online-evaluation/#per-evaluator-sink-overrides) Individual evaluators can override the config’s default sink. This is useful if different evaluators need to send results to different destinations: import asyncio from collections.abc import Sequence from dataclasses import dataclass from pydantic_evals.evaluators import ( EvaluationResult, Evaluator, EvaluatorContext, EvaluatorFailure, ) from pydantic_evals.online import ( OnlineEvalConfig, OnlineEvaluator, wait_for_evaluations, ) default_log: list[str] = [] special_log: list[str] = [] async def default_sink( results: Sequence[EvaluationResult], failures: Sequence[EvaluatorFailure], context: EvaluatorContext, ) -> None: for r in results: default_log.append(r.name) async def special_sink( results: Sequence[EvaluationResult], failures: Sequence[EvaluatorFailure], context: EvaluatorContext, ) -> None: for r in results: special_log.append(r.name) @dataclass class FastCheck(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return True @dataclass class ImportantCheck(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return True config = OnlineEvalConfig(default_sink=default_sink) @config.evaluate( FastCheck(), # uses default sink OnlineEvaluator(evaluator=ImportantCheck(), sink=special_sink), # uses special sink ) async def my_function(x: int) -> int: return x async def main(): await my_function(42) await wait_for_evaluations() print(f'default: {default_log}') #> default: ['FastCheck'] print(f'special: {special_log}') #> special: ['ImportantCheck'] asyncio.run(main()) Re-running Evaluators from Stored Data -------------------------------------- [](https://pydantic.dev/docs/ai/evals/online-evaluation/#re-running-evaluators-from-stored-data) A key capability of online evaluation is re-running evaluators without re-executing the original function. This is useful when you want to evaluate historical data with updated rubrics, or run additional evaluators on existing traces. ### run\_evaluators [](https://pydantic.dev/docs/ai/evals/online-evaluation/#run_evaluators) [`run_evaluators()`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.run_evaluators) runs a list of evaluators against an [`EvaluatorContext`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorContext) and returns the results: import asyncio from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext from pydantic_evals.online import run_evaluators from pydantic_evals.otel.span_tree import SpanTree @dataclass class LengthCheck(Evaluator): min_length: int = 10 def evaluate(self, ctx: EvaluatorContext) -> bool: return len(str(ctx.output)) >= self.min_length @dataclass class HasKeyword(Evaluator): keyword: str = 'hello' def evaluate(self, ctx: EvaluatorContext) -> bool: return self.keyword in str(ctx.output).lower() async def main(): # Build a context manually (in practice, you'd get this from stored data) # Normally EvaluatorContext would not be manually constructed — # it is built automatically by the @evaluate decorator or OnlineEvaluation capability, # or from an EvaluatorContextSource (see below). ctx = EvaluatorContext( name='example', inputs={'query': 'greet the user'}, output='Hello! How can I help you today?', expected_output=None, metadata=None, duration=0.5, _span_tree=SpanTree(), attributes={}, metrics={}, ) results, failures = await run_evaluators( [LengthCheck(min_length=10), HasKeyword(keyword='hello')], ctx, ) for r in results: print(f'{r.name}: {r.value}') #> LengthCheck: True #> HasKeyword: True print(f'failures: {len(failures)}') #> failures: 0 asyncio.run(main()) ### EvaluatorContextSource Protocol [](https://pydantic.dev/docs/ai/evals/online-evaluation/#evaluatorcontextsource-protocol) For fetching context data from external storage (like Pydantic Logfire), implement the [`EvaluatorContextSource`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.EvaluatorContextSource) protocol. It defines `fetch()` and `fetch_many()` methods that return [`EvaluatorContext`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorContext) objects from stored data: import asyncio from collections.abc import Sequence from pydantic_evals.evaluators import EvaluatorContext from pydantic_evals.online import SpanReference from pydantic_evals.otel.span_tree import SpanTree class MyContextSource: """Example source that fetches context from a hypothetical store.""" def __init__(self, store: dict[str, EvaluatorContext]) -> None: self._store = store async def fetch(self, span: SpanReference) -> EvaluatorContext: return self._store[span.span_id] async def fetch_many(self, spans: Sequence[SpanReference]) -> list[EvaluatorContext]: return [self._store[s.span_id] for s in spans] def _make_context( *, inputs: object = None, output: object = None, metadata: object = None, duration: float = 0.0, ) -> EvaluatorContext: # Normally EvaluatorContext would not be manually constructed — # it is built automatically by the @evaluate decorator or OnlineEvaluation capability. return EvaluatorContext( name=None, inputs=inputs, output=output, expected_output=None, metadata=metadata, duration=duration, _span_tree=SpanTree(), attributes={}, metrics={}, ) async def main(): source = MyContextSource({ 'span_abc': _make_context( inputs={'query': 'What is AI?'}, output='AI is artificial intelligence.', metadata={'model': 'gpt-4o'}, duration=1.2, ), 'span_def': _make_context( inputs={'query': 'What is ML?'}, output='ML is machine learning.', metadata={'model': 'gpt-4o'}, duration=0.8, ), }) # Fetch a single context ctx = await source.fetch(SpanReference(trace_id='t1', span_id='span_abc')) print(f'inputs: {ctx.inputs}') #> inputs: {'query': 'What is AI?'} print(f'output: {ctx.output}') #> output: AI is artificial intelligence. # Fetch multiple contexts in a batch spans = [\ SpanReference(trace_id='t1', span_id='span_abc'),\ SpanReference(trace_id='t1', span_id='span_def'),\ ] contexts = await source.fetch_many(spans) print(f'batch size: {len(contexts)}') #> batch size: 2 asyncio.run(main()) #### Serializing an `EvaluatorContext` [](https://pydantic.dev/docs/ai/evals/online-evaluation/#serializing-an-evaluatorcontext) To populate a store like the one above, you need to serialize an [`EvaluatorContext`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorContext) to JSON (and read it back). `EvaluatorContext` is a Pydantic-serializable dataclass, so a [`TypeAdapter`](https://docs.pydantic.dev/latest/api/pydantic/type_adapter/#pydantic.type_adapter.TypeAdapter) handles both directions. Bind it to the concrete `inputs`, `output`, and `metadata` types your contexts carry so those fields are reconstructed faithfully: from pydantic import TypeAdapter from pydantic_evals.evaluators import EvaluatorContext from pydantic_evals.otel.span_tree import SpanTree context_adapter = TypeAdapter(EvaluatorContext[dict[str, str], str, dict[str, str]]) ctx = EvaluatorContext[dict[str, str], str, dict[str, str]]( name='span_abc', inputs={'query': 'What is AI?'}, output='AI is artificial intelligence.', expected_output=None, metadata={'model': 'gpt-4o'}, duration=1.2, _span_tree=SpanTree(), attributes={}, metrics={}, ) json_bytes = context_adapter.dump_json(ctx) restored = context_adapter.validate_json(json_bytes) print(restored.output) #> AI is artificial intelligence. Concurrency Control ------------------- [](https://pydantic.dev/docs/ai/evals/online-evaluation/#concurrency-control) Each [`OnlineEvaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.OnlineEvaluator) has a `max_concurrency` limit (default: 10). When the limit is reached, new evaluation requests for that evaluator are **dropped** (not queued). This prevents expensive evaluators from consuming unbounded resources: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext from pydantic_evals.online import OnlineEvaluator @dataclass class ExpensiveCheck(Evaluator): async def evaluate(self, ctx: EvaluatorContext) -> bool: # Imagine this makes a slow call to an LLM return True # Allow at most 3 concurrent evaluations limited = OnlineEvaluator( evaluator=ExpensiveCheck(), sample_rate=0.1, max_concurrency=3, ) To react to dropped evaluations, set `on_max_concurrency` on the [`OnlineEvaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.OnlineEvaluator) or as a default on [`OnlineEvalConfig`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.OnlineEvalConfig) . The callback receives the [`EvaluatorContext`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorContext) that would have been evaluated, and can be sync or async: import warnings from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext from pydantic_evals.online import OnlineEvalConfig, OnlineEvaluator @dataclass class ExpensiveCheck(Evaluator): async def evaluate(self, ctx: EvaluatorContext) -> bool: return True def warn_on_drop(ctx: EvaluatorContext) -> None: warnings.warn('Evaluation dropped due to max concurrency', stacklevel=1) # Per-evaluator handler limited = OnlineEvaluator( evaluator=ExpensiveCheck(), max_concurrency=3, on_max_concurrency=warn_on_drop, ) # Or set a global default for all evaluators in a config config = OnlineEvalConfig(on_max_concurrency=warn_on_drop) Error Handling -------------- [](https://pydantic.dev/docs/ai/evals/online-evaluation/#error-handling) There are two types of error handling: * **`on_sampling_error`**: Called synchronously when a `sample_rate` callable raises. Receives the exception and the [`Evaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Evaluator) . Must be sync (not async). If set, the evaluator is skipped. If not set, the exception **propagates to the caller**. * **`on_error`**: Called when an exception occurs in a `sink` or `on_max_concurrency` callback. Receives the exception, [`EvaluatorContext`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorContext) , [`Evaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.Evaluator) , and a [`OnErrorLocation`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.OnErrorLocation) string. Can be sync or async. If not set, exceptions are **silently suppressed**. The `'sink'` location is broad — it covers both custom sink failures and the rarer default OTel event emission failures, so handlers that branch on location should treat `'sink'` as “result delivery went wrong”. Set these on [`OnlineEvalConfig`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.OnlineEvalConfig) for global defaults, or on [`OnlineEvaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.OnlineEvaluator) to override per-evaluator: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext from pydantic_evals.online import OnErrorLocation, OnlineEvalConfig, OnlineEvaluator def log_errors( exc: Exception, ctx: EvaluatorContext, evaluator: Evaluator, location: OnErrorLocation, ) -> None: print(f'[{location}] {type(exc).__name__}: {exc}') @dataclass class MyCheck(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return True # Global default — applies to all evaluators in this config config = OnlineEvalConfig( default_sink=lambda results, failures, context: None, on_error=log_errors, ) # Per-evaluator override custom = OnlineEvaluator(evaluator=MyCheck(), on_error=log_errors) Key behaviors: * **Evaluator exceptions** are handled by converting them to [`EvaluatorFailure`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorFailure) objects passed to sinks — they do not go through `on_error`. * **One evaluator’s error doesn’t affect siblings** — each evaluator runs in its own task with isolated error handling. * **One sink’s error doesn’t affect other sinks** — each sink submission is wrapped individually. * **If `on_error` itself raises**, the exception is silently suppressed to protect sibling evaluators. * **If no `on_error` is set**, exceptions are silently suppressed — this is the safe default. ### Evaluating Failed Calls [](https://pydantic.dev/docs/ai/evals/online-evaluation/#evaluating-failed-calls) By default, when the decorated function or wrapped agent run raises, **no evaluators are dispatched** — only successful results reach evaluators. The exception propagates to the caller as usual. To score failure modes (e.g. classify exception types, count tool errors, alert on regressions), opt an evaluator in by setting `run_on_errors=True` on its [`OnlineEvaluator`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.OnlineEvaluator) . When the call raises, those evaluators are dispatched with the exception as `EvaluatorContext.output`; the exception still propagates after dispatch: from dataclasses import dataclass from pydantic_evals.evaluators import Evaluator, EvaluatorContext from pydantic_evals.online import OnlineEvaluator, evaluate @dataclass class CategorizeError(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> str: # On failed calls, ctx.output is the raised exception. if isinstance(ctx.output, Exception): return type(ctx.output).__name__ return 'ok' @evaluate(OnlineEvaluator(evaluator=CategorizeError(), run_on_errors=True)) async def my_function(x: int) -> int: if x < 0: raise ValueError('negative input') return x * 2 Evaluators sampled for the call but without `run_on_errors=True` are skipped on the error path, so a cheap success-only check can sit alongside a dedicated error categorizer in the same decorator. The flag is also honored by the [`OnlineEvaluation`](https://pydantic.dev/docs/ai/api/pydantic_evals/online_capability/#pydantic_evals.online_capability.OnlineEvaluation) agent capability. Agent Integration ----------------- [](https://pydantic.dev/docs/ai/evals/online-evaluation/#agent-integration) The [`OnlineEvaluation`](https://pydantic.dev/docs/ai/api/pydantic_evals/online_capability/#pydantic_evals.online_capability.OnlineEvaluation) capability brings online evaluation to Pydantic AI agents. Instead of decorating a function, you add the capability to your agent. As with the `@evaluate` decorator, evaluators dispatch in the background and results are emitted as OTel events by default — no sink registration required: from dataclasses import dataclass from pydantic_ai import Agent from pydantic_evals.evaluators import Evaluator, EvaluatorContext from pydantic_evals.online_capability import OnlineEvaluation @dataclass class OutputNotEmpty(Evaluator): def evaluate(self, ctx: EvaluatorContext) -> bool: return bool(ctx.output) agent = Agent( 'openai:gpt-5.2', name='assistant', capabilities=[OnlineEvaluation(evaluators=[OutputNotEmpty()])], ) The target name written to each emitted event is the agent’s own `name` attribute, so events from `agent = Agent(..., name='assistant')` land under `gen_ai.evaluation.target = 'assistant'`. If the agent has no name, the target falls back to the literal string `'agent'`. After each completed agent run, the capability: 1. Samples evaluators based on their `sample_rate` configuration 2. Builds an [`EvaluatorContext`](https://pydantic.dev/docs/ai/api/pydantic_evals/evaluators/#pydantic_evals.evaluators.EvaluatorContext) from the run result (output, prompt, token usage, duration, span tree) — `context.name` is populated with the agent run’s `run_id` 3. Dispatches evaluators asynchronously in the background 4. Returns control to the caller without waiting for evaluators to finish To attach additional sinks or override sampling defaults, pass an [`OnlineEvalConfig`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.OnlineEvalConfig) — same as with the `@evaluate` decorator: `OnlineEvaluation(evaluators=[...], config=OnlineEvalConfig(default_sample_rate=0.1))`. The capability supports all the same features as the [`@evaluate()`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.evaluate) decorator: sampling, per-evaluator sinks, concurrency control, and error handling. The `config` parameter is optional and defaults to the global [`DEFAULT_CONFIG`](https://pydantic.dev/docs/ai/api/pydantic_evals/online/#pydantic_evals.online.DEFAULT_CONFIG) . API Reference ------------- [](https://pydantic.dev/docs/ai/evals/online-evaluation/#api-reference) The complete API for the `pydantic_evals.online` module is documented in the [API reference](https://pydantic.dev/docs/ai/api/pydantic_evals/online/) . Next Steps ---------- [](https://pydantic.dev/docs/ai/evals/online-evaluation/#next-steps) * **[Custom Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/custom/) ** — Write evaluators for your domain * **[Native Evaluators](https://pydantic.dev/docs/ai/evals/evaluators/built-in/) ** — Use ready-made evaluators * **[Live Evaluations in Logfire](https://pydantic.dev/docs/logfire/evaluate/live-evals/) ** — Browse, filter, and trend online evaluation results in the Logfire web UI * **[Logfire Integration](https://pydantic.dev/docs/ai/evals/how-to/logfire-integration/) ** — Visualize evaluation results in Logfire * **[Quick Start](https://pydantic.dev/docs/ai/evals/getting-started/quick-start/) ** — Offline evaluation with [`Dataset.evaluate()`](https://pydantic.dev/docs/ai/api/pydantic_evals/dataset/#pydantic_evals.dataset.Dataset.evaluate) Was this page helpful? Thanks for your feedback! ---