Anthropic
This provider supports the Anthropic Claude series of models.
Note: Anthropic models can also be accessed through Azure AI Foundry, AWS Bedrock, and Google Vertex.
For agentic evaluations that need built-in file access and skill plugins on top of the Messages API, see the Claude Agent SDK provider. The Messages provider documented here speaks directly to MCP servers via the mcp config so you can plug in your own tools without changing providers.
Setup
To use Anthropic, you need to set the ANTHROPIC_API_KEY environment variable or specify the apiKey in the provider configuration.
Create Anthropic API keys here.
Example of setting the environment variable:
export ANTHROPIC_API_KEY=your_api_key_here
Authenticating via a Claude Code session
If you already have an active Claude Code session (for example as a Claude Pro or Max subscriber), you can reuse its OAuth credential instead of creating a separate Anthropic Console API key. Set apiKeyRequired: false on the provider config:
providers:
- id: anthropic:messages:claude-sonnet-5
config:
apiKeyRequired: false
When apiKeyRequired is false and no ANTHROPIC_API_KEY is available, Promptfoo loads the Claude Code OAuth credential from:
- The macOS keychain entry
Claude Code-credentials(darwin only), then $HOME/.claude/.credentials.jsonon Linux and macOS, or%USERPROFILE%\.claude\.credentials.jsonon Windows.
Set CLAUDE_CONFIG_DIR to read the credential from a different Claude Code profile — the same environment variable the Claude Code CLI itself uses to relocate ~/.claude. It can be set in your shell, in the config's top-level env: block, or in a provider's env: block (the provider-scoped value wins). On macOS, where Claude Code stores credentials in the system keychain, Promptfoo mirrors the CLI's profile-specific keychain entry: when CLAUDE_CONFIG_DIR is set, the credential is looked up under that profile's keychain service (derived from the configured directory) rather than the default one, so evals authenticate as the profile you selected.
Promptfoo authenticates requests with a Bearer token, sends the claude-code-20250219,oauth-2025-04-20 beta headers, and prepends the required Claude Code identity system block ("You are Claude Code, Anthropic's official CLI for Claude.") to every Messages request. Your own system prompt is still forwarded as the next system block.
If you haven't logged in yet, run claude /login to create a credential. Re-run it if Promptfoo warns that the credential has expired. Requests made this way are expected to count against your Claude subscription the same way calls from the Claude Code CLI do — check Anthropic's documentation for current billing behavior.
This also enables model-graded assertions such as llm-rubric to run without a separate Anthropic Console key — see the example below.
Models
These models currently resolve on the Anthropic Messages API:
| Model ID | Description |
|---|---|
anthropic:messages:claude-fable-5-1 | Claude Fable 5.1 |
anthropic:messages:claude-mythos-5-1 | Claude Mythos 5.1 |
anthropic:messages:claude-fable-5 | Claude Fable 5 |
anthropic:messages:claude-mythos-5 | Claude Mythos 5 |
anthropic:messages:claude-opus-5-5 | Claude Opus 5.5 |
anthropic:messages:claude-sonnet-5-5 | Claude Sonnet 5.5 |
anthropic:messages:claude-opus-5 | Claude Opus 5 |
anthropic:messages:claude-opus-4-8 | Claude 4.8 Opus |
anthropic:messages:claude-opus-4-7 | Claude 4.7 Opus |
anthropic:messages:claude-sonnet-5 | Claude Sonnet 5 |
anthropic:messages:claude-sonnet-4-6 | Claude 4.6 Sonnet |
anthropic:messages:claude-opus-4-6 | Claude 4.6 Opus |
anthropic:messages:claude-opus-4-5-20251101 (claude-opus-4-5) | Claude 4.5 Opus |
anthropic:messages:claude-sonnet-4-5-20250929 (claude-sonnet-4-5) | Claude 4.5 Sonnet |
anthropic:messages:claude-haiku-4-5-20251001 (claude-haiku-4-5) | Claude 4.5 Haiku |
The Mythos rows are limited-access models: the ID is correct, but an org without access gets
the same not_found_error a retired model returns.
Retired on the Anthropic API
These IDs return 404 not_found_error from Anthropic, so a direct anthropic:messages:
call will fail. Promptfoo still keeps their pricing, because cost attribution on historical
evals needs it and because partner platforms set their own lifecycle dates.
Availability elsewhere is per-model, not a blanket rule — check the row in
Cross-Platform Model Availability before assuming a
retired ID still works somewhere. Three of these are withdrawn from Bedrock as well
(claude-3-opus-20240229, claude-opus-4-20250514, claude-3-5-haiku-20241022) and are
rejected locally with Unknown Amazon Bedrock model.
| Model ID | Description | Suggested replacement |
|---|---|---|
claude-opus-4-1-20250805 | Claude 4.1 Opus | claude-opus-5-5 |
claude-opus-4-20250514 | Claude 4 Opus | claude-opus-5-5 |
claude-sonnet-4-20250514 | Claude 4 Sonnet | claude-sonnet-5 |
claude-3-7-sonnet-20250219 | Claude 3.7 Sonnet | claude-sonnet-5 |
claude-3-5-sonnet-20241022 | Claude 3.5 Sonnet (v2) | claude-sonnet-5 |
claude-3-5-sonnet-20240620 | Claude 3.5 Sonnet (v1) | claude-sonnet-5 |
claude-3-5-haiku-20241022 | Claude 3.5 Haiku | claude-haiku-4-5 |
claude-3-opus-20240229 | Claude 3 Opus | claude-opus-5-5 |
claude-3-haiku-20240307 | Claude 3 Haiku | claude-haiku-4-5 |
Anthropic does not publish -latest aliases — claude-sonnet-4-5-latest returns a
not_found_error. Where a shorter alias exists it is the bare family ID shown in
parentheses above: claude-sonnet-4-5 resolves to claude-sonnet-4-5-20250929.
Claude 4.6 and newer are already unversioned IDs, so there is nothing to shorten.
Rows without a parenthetical have no alias. Retired on the Anthropic API lists retirements from the direct Anthropic API. Check the platform guides below for availability through other providers.
Cross-Platform Model Availability
Claude models are available across multiple platforms. Here's how the model names map across different providers:
| Model | Anthropic API | Azure AI Foundry (docs) | AWS Bedrock (docs) | GCP Vertex AI (docs) |
|---|---|---|---|---|
| Claude Fable 5.1 | claude-fable-5-1 | claude-fable-5-1 | anthropic.claude-fable-5-1 | claude-fable-5-1 |
| Claude Mythos 5.1 | claude-mythos-5-1 | claude-mythos-5-1 (limited) | anthropic.claude-mythos-5-1 (limited) | claude-mythos-5-1 (limited) |
| Claude Fable 5 | claude-fable-5 | claude-fable-5 | anthropic.claude-fable-5 | claude-fable-5 |
| Claude Mythos 5 | claude-mythos-5 | Not available | anthropic.claude-mythos-5 (limited) | Limited availability; ID not public |
| Claude Opus 5.5 | claude-opus-5-5 | claude-opus-5-5 | anthropic.claude-opus-5-5 | claude-opus-5-5 |
| Claude Sonnet 5.5 | claude-sonnet-5-5 | claude-sonnet-5-5 | global.anthropic.claude-sonnet-5-5 | claude-sonnet-5-5 |
| Claude Opus 5 | claude-opus-5 | claude-opus-5 | anthropic.claude-opus-5 | claude-opus-5 |
| Claude 4.8 Opus | claude-opus-4-8 | claude-opus-4-8 | anthropic.claude-opus-4-8 | claude-opus-4-8 |
| Claude 4.7 Opus | claude-opus-4-7 | claude-opus-4-7 | anthropic.claude-opus-4-7 | claude-opus-4-7 |
| Claude Sonnet 5 | claude-sonnet-5 | claude-sonnet-5 | anthropic.claude-sonnet-5 | claude-sonnet-5 |
| Claude 4.6 Sonnet | claude-sonnet-4-6 | claude-sonnet-4-6 | anthropic.claude-sonnet-4-6 | claude-sonnet-4-6 |
| Claude 4.6 Opus | claude-opus-4-6 | claude-opus-4-6-20260205 | anthropic.claude-opus-4-6-v1 | claude-opus-4-6 |
| Claude 4.5 Opus | claude-opus-4-5-20251101 (claude-opus-4-5) | claude-opus-4-5-20251101 | anthropic.claude-opus-4-5-20251101-v1:0 | claude-opus-4-5@20251101 |
| Claude 4.5 Sonnet | claude-sonnet-4-5-20250929 (claude-sonnet-4-5) | claude-sonnet-4-5-20250929 | anthropic.claude-sonnet-4-5-20250929-v1:0 | claude-sonnet-4-5@20250929 |
| Claude 4.5 Haiku | claude-haiku-4-5-20251001 (claude-haiku-4-5) | claude-haiku-4-5-20251001 | anthropic.claude-haiku-4-5-20251001-v1:0 | claude-haiku-4-5@20251001 |
| Claude 4.1 Opus | Retired on the direct API | claude-opus-4-1-20250805 | anthropic.claude-opus-4-1-20250805-v1:0 | claude-opus-4-1@20250805 |
| Claude 4 Opus | Retired on the direct API | claude-opus-4-20250514 | Withdrawn from Bedrock | claude-opus-4@20250514 |
| Claude 4 Sonnet | Retired on the direct API | claude-sonnet-4-20250514 | anthropic.claude-sonnet-4-20250514-v1:0 | claude-sonnet-4@20250514 |
| Claude 3.7 Sonnet | Retired on the direct API | claude-3-7-sonnet-20250219 | anthropic.claude-3-7-sonnet-20250219-v1:0 | claude-3-7-sonnet@20250219 |
| Claude 3.5 Sonnet | Retired on the direct API | claude-3-5-sonnet-20241022 | anthropic.claude-3-5-sonnet-20241022-v2:0 | claude-3-5-sonnet-v2@20241022 |
| Claude 3.5 Haiku | Retired on the direct API | claude-3-5-haiku-20241022 | Withdrawn from Bedrock | claude-3-5-haiku@20241022 |
| Claude 3 Opus | Retired on the direct API | claude-3-opus-20240229 | Withdrawn from Bedrock | claude-3-opus@20240229 |
| Claude 3 Haiku | Retired on the direct API | claude-3-haiku-20240307 | anthropic.claude-3-haiku-20240307-v1:0 | claude-3-haiku@20240307 |
Supported Parameters
| Config Property | Environment Variable | Description |
|---|---|---|
| apiKey | ANTHROPIC_API_KEY | Your API key from Anthropic |
| apiKeyRequired | - | Skip the API key preflight and authenticate via a local Claude Code session |
| apiBaseUrl | ANTHROPIC_BASE_URL | The base URL for requests to the Anthropic API |
| temperature | ANTHROPIC_TEMPERATURE | Controls the randomness of the output (default: 0). Omitted when top_p is set. |
| max_tokens | ANTHROPIC_MAX_TOKENS | The maximum length of the generated text (default: 1024, or 2048 whenever thinking will consume output tokens — either because you enabled it or because the model thinks by default) |
| cost | - | Legacy per-token override applied to both input and output pricing |
| inputCost | - | Override input token pricing in promptfoo cost estimates |
| outputCost | - | Override output token pricing in promptfoo cost estimates |
| top_p | - | Controls nucleus sampling. Mutually exclusive with temperature. |
| top_k | - | Only sample from the top K options for each subsequent token |
| stop_sequences | - | Array of strings that will stop generation when encountered |
| stream | - | Enable streaming (required when max_tokens > 21,333) |
| tools | - | An array of tool or function definitions for the model to call |
| tool_choice | - | An object specifying the tool to call |
| effort | - | Output effort level: low, medium, high, xhigh, or max |
| output_format | - | JSON schema configuration for structured outputs |
| thinking | - | Configuration for Claude's extended thinking (enabled, adaptive, disabled, or between_tools) |
| showThinking | - | Whether to include thinking content in the output (default: true) |
| cache_control | - | Auto-apply cache_control to the last cacheable block in the request |
| metadata | - | Request metadata such as user_id for tracking purposes |
| service_tier | - | Priority tier: auto (default) or standard_only |
| beta | - | Array of anthropic-beta feature flags to send with the request |
| headers | - | Additional headers to be sent with the API request |
| extra_body | - | Additional parameters to be included in the API request body |
For MCP tools, set mcp.enabled: true. The
max_tool_calls option caps MCP tool executions per request (default: 8).
Sampling support varies by model. Promptfoo omits unsupported temperature,
top_p, and top_k settings; see the model notes below.
Prompt Template
To allow for compatibility with the OpenAI prompt template, the following format is supported:
[
{
"role": "system",
"content": "{{ system_message }}"
},
{
"role": "user",
"content": "{{ question }}"
}
]
Promptfoo extracts system messages into the API's system prompt and forwards user and
assistant messages. Set system_message and question in your test's vars.
Claude 4.6 and later models reject prompts ending with an assistant message. End
with a user message instead. To constrain the response format, use
structured outputs or a system instruction. The 4.5 models
still accept assistant prefill.
Options
Set supported parameters under config:
providers:
- id: anthropic:messages:claude-sonnet-5
config:
effort: medium
max_tokens: 2048 # Includes thinking and the final answer
prompts:
- file://prompt.json
Stop Sequences
Use stop_sequences to halt generation when Claude encounters specific strings:
providers:
- id: anthropic:messages:claude-sonnet-5
config:
stop_sequences:
- "\n\nHuman:"
- 'STOP'
Classifier Refusals
An Anthropic safety-classifier refusal can arrive as a successful Messages response with stop_reason: refusal. When the response also includes stop_details, Promptfoo returns the provider content as output and normalizes:
{
"finishReason": "content_filter",
"guardrails": {
"flagged": true,
"reason": "Content refused by Anthropic safety filters — category: ..."
}
}
Promptfoo currently requires stop_details to create the top-level guardrails signal. A model-written refusal or stop_reason: refusal without details is therefore not the same assertion result. Treat the optional stop_details.category and stop_details.explanation as diagnostic evidence. API validation errors remain provider errors and skip assertions.
In a stream, the terminal refusal reason can arrive after partial text in the final message delta. Promptfoo merges those details before returning the provider response, and cached structured refusals preserve the signal. Use not-guardrails to require the structured classifier signal. Use is-refusal for model-written refusal text.
Cost estimates follow Anthropic's refusal billing rules: refusals before any output cost zero unless the category is bio, frontier_llm, or reasoning_extraction. Those categories and refusals after output begins use normal token pricing.
Metadata
Pass request metadata to the API for tracking or auditing purposes:
providers:
- id: anthropic:messages:claude-sonnet-5
config:
metadata:
user_id: 'user-123'
Tool Calling
The Anthropic provider supports tool calling (function calling). Here's an example configuration for defining tools.
providers:
- id: anthropic:messages:claude-sonnet-5
config:
tools:
- name: get_weather
description: Get the current weather in a given location
input_schema:
type: object
properties:
location:
type: string
description: The city and state, e.g., San Francisco, CA
unit:
type: string
enum:
- celsius
- fahrenheit
required:
- location
Web Search and Web Fetch Tools
Anthropic provides specialized tools for web search and web fetching capabilities:
Web Fetch Tool
The web fetch tool allows Claude to retrieve full content from web pages and PDF documents. This is useful when you want Claude to access and analyze specific web content.
providers:
- id: anthropic:messages:claude-sonnet-5
config:
tools:
- type: web_fetch_20250910
name: web_fetch
max_uses: 5
allowed_domains:
- docs.example.com
- help.example.com
citations:
enabled: true
max_content_tokens: 50000
Use one fetch version per request. web_fetch_20260209 adds dynamic filtering,
web_fetch_20260309 adds cache control, and web_fetch_20260318 adds response inclusion control:
providers:
- id: anthropic:messages:claude-sonnet-5
config:
tools:
- type: web_fetch_20260318
name: web_fetch
max_uses: 3
use_cache: false
response_inclusion: excluded
response_inclusion: excluded omits nested tool-use/result pairs consumed by a
completed code-execution call. Direct fetches and paused calls still return their
full results. See Anthropic's web fetch documentation.
Web Fetch Tool Configuration Options:
| Parameter | Type | Description |
|---|---|---|
type | string | web_fetch_20250910 (beta), web_fetch_20260209, web_fetch_20260309 (adds use_cache), or web_fetch_20260318 (adds response_inclusion) |
name | string | Must be web_fetch |
max_uses | number | Maximum number of web fetches per request (optional) |
allowed_callers | string[] | Restrict which tool callers may invoke the server tool (optional) |
allowed_domains | string[] | List of domains to allow fetching from (optional, mutually exclusive with blocked_domains) |
blocked_domains | string[] | List of domains to block fetching from (optional, mutually exclusive with allowed_domains) |
defer_loading | boolean | Load the tool lazily instead of including it in the initial system prompt (optional) |
citations | object | Enable citations with { enabled: true } (optional) |
max_content_tokens | number | Maximum tokens for web content (optional) |
cache_control | object | Apply Anthropic cache control to the tool definition (optional) |
strict | boolean | Enable strict schema validation for tool names and inputs (optional) |
url_sources | object | Limit fetchable URLs by source: user_input, client_tool_results, or server_tool_results, each with a tagged filter such as { type: none } (optional) |
use_cache | boolean | Whether to use cached content (web_fetch_20260309 and web_fetch_20260318, optional) |
response_inclusion | string | full (default) or excluded; applies to completed nested code-execution calls (web_fetch_20260318 only) |
Web Search Tool
The web search tool allows Claude to search the internet for information:
providers:
- id: anthropic:messages:claude-sonnet-5
config:
tools:
- type: web_search_20260209
name: web_search
max_uses: 3
Web Search Tool Configuration Options:
| Parameter | Type | Description |
|---|---|---|
type | string | web_search_20250305, web_search_20260209, or web_search_20260318 (adds response_inclusion) |
name | string | Must be web_search |
max_uses | number | Maximum number of searches per request (optional) |
allowed_callers | string[] | Restrict which tool callers may invoke the server tool (optional) |
allowed_domains | string[] | Restrict results to specific domains (optional, mutually exclusive with blocked_domains) |
blocked_domains | string[] | Exclude domains from results (optional, mutually exclusive with allowed_domains) |
cache_control | object | Apply Anthropic cache control to the tool definition (optional) |
defer_loading | boolean | Load the tool lazily instead of including it in the initial system prompt (optional) |
strict | boolean | Enable strict schema validation for tool names and inputs (optional) |
response_inclusion | string | full (default) or excluded — see the web fetch table above (web_search_20260318 only, optional) |
user_location | object | Approximate user location to improve search relevance (optional) |
Combined Web Search and Web Fetch
Use web search to find URLs and web fetch to read their contents:
providers:
- id: anthropic:messages:claude-sonnet-5
config:
tools:
- type: web_search_20260209
name: web_search
max_uses: 3
- type: web_fetch_20260309
name: web_fetch
max_uses: 5
citations:
enabled: true
This configuration allows the model to first search for relevant information, then fetch full content from the most promising results.
Paused Turns
A long server-tool run can stop with stop_reason: pause_turn before Claude finishes. Promptfoo resumes it up to 5 times, preserving the code execution container and summing each request's usage and cost. Structured output uses the final response's JSON; generated file references remain in metadata.fileReferences. If resuming fails or reaches the limit, the result keeps its partial output with finishReason: pause_turn, logs a warning, and is not cached. Cancellation stops further requests.
Memory Tool
Anthropic's memory_20250818 tool can be included in tools. Promptfoo passes this native tool definition through unchanged, which is useful for evaluating whether a model requests memory operations. Promptfoo does not manage Anthropic memory stores or run local memory handlers for you.
providers:
- id: anthropic:messages:claude-sonnet-5
config:
tools:
- type: memory_20250818
name: memory
allowed_callers:
- direct
Memory Tool Configuration Options:
| Parameter | Type | Description |
|---|---|---|
type | string | Must be memory_20250818 |
name | string | Must be memory |
allowed_callers | string[] | Restrict which tool callers may invoke the memory tool (optional) |
cache_control | object | Apply Anthropic cache control to the tool definition (optional) |
defer_loading | boolean | Load the tool lazily instead of including it in the initial prompt |
input_examples | object[] | Example memory commands to include in the tool definition (optional) |
strict | boolean | Enable strict schema validation for tool names and inputs (optional) |
Important Security Notes:
- The web fetch tool requires trusted environments due to potential data exfiltration risks
- The model cannot dynamically construct URLs - only URLs provided by users or from search results can be fetched
- Use domain filtering to restrict access to specific sites:
- Use
allowed_domainsto whitelist trusted domains (recommended) - Use
blocked_domainsto blacklist specific domains - Note: Only one of
allowed_domainsorblocked_domainscan be specified, not both
- Use
Model Context Protocol (MCP)
The Anthropic Messages provider can connect to any MCP server — stdio, SSE, or streamable HTTP — and execute the model's tool_use blocks against that server, feeding the tool_result back into the conversation until Claude produces a final reply.
providers:
- id: anthropic:messages:claude-sonnet-5
config:
mcp:
enabled: true
# Inline command-based stdio server, or `path` to a local script
server:
command: npx
args: ['-y', '@modelcontextprotocol/server-filesystem', '/tmp/workspace']
# Or use a remote SSE / streamable HTTP server
# servers:
# - name: deepwiki
# url: https://mcp.deepwiki.com/mcp
# Optional cap on MCP rounds per request (default 8). Enforced locally;
# not sent to Anthropic.
max_tool_calls: 5
How it works:
- Tools discovered on the MCP server are passed to Claude alongside any inline
tools. - When Claude returns a
tool_useblock whose name matches an MCP tool, promptfoo calls the tool with the model's arguments and appends a matchingtool_resultblock on the user turn. - The loop repeats until Claude returns text (no more
tool_use) ormax_tool_callsis hit. Tool errors are forwarded astool_resultblocks withis_error: trueso the model can recover. - Non-MCP
tool_useblocks (regular function tools, or built-ins likeweb_search) are passed through to the existing output and not auto-executed.
The disk response cache is skipped while mcp.enabled is true, because tool results can be non-deterministic between runs. Use max_tool_calls to bound spend.
See the MCP integration guide for full server configuration options (auth, timeouts, multiple servers, etc.) and the Anthropic MCP example.
See the Anthropic Tool Use Guide for more information on how to define tools and the tool use example here.
Images / Vision
All current Claude models accept images in the prompt.
See the Claude vision example.
Claude accepts base64 images and image URLs. Use Anthropic image content blocks;
their shape differs from OpenAI's image_url blocks. See the
vision API guide for request examples.
Prompt Caching
Prompt caching reuses unchanged prompt prefixes and is supported by current Claude
models. Mark the prefix to cache with cache_control:
providers:
- id: anthropic:messages:claude-sonnet-5
prompts:
- file://prompts.yaml
- role: system
content:
- type: text
text: 'System message'
cache_control:
type: ephemeral
- type: text
text: '{{context}}'
cache_control:
type: ephemeral
- role: user
content: '{{question}}'
As a simpler alternative, use the top-level cache_control parameter to automatically apply a cache marker to the last cacheable block in the request, without annotating each block individually:
providers:
- id: anthropic:messages:claude-sonnet-5
config:
cache_control:
type: ephemeral
Common use cases for caching:
- System messages and instructions
- Tool/function definitions
- Large context documents
- Frequently used images
Cache read and creation token counts are tracked in the response's token usage details. Cost estimates use the reported cache duration: five-minute writes cost 1.25 times the base input rate, and one-hour writes cost twice the base input rate.
See Anthropic's Prompt Caching Guide for more details on requirements, pricing, and best practices.
Citations
Claude can provide detailed citations when answering questions about documents. Basic example:
providers:
- id: anthropic:messages:claude-sonnet-5
prompts:
- file://prompts.yaml
- role: user
content:
- type: document
source:
type: text
media_type: text/plain
data: 'Your document text here'
citations:
enabled: true
- type: text
text: 'Your question here'
See Anthropic's Citations Guide for more details.
PDF Documents
Claude can process PDF files using document content blocks. Pass the PDF as base64-encoded data:
- role: user
content:
- type: document
source:
type: base64
media_type: application/pdf
data: '{{pdf_base64}}'
- type: text
text: 'Summarize this document'
Use a test var to supply the base64-encoded PDF content:
tests:
- vars:
pdf_base64: file://document.pdf
Claude Fable 5.1 and Mythos 5.1
Use the pinned model IDs below. Mythos 5.1 requires Project Glasswing access and provider approval.
providers:
- id: anthropic:messages:claude-fable-5-1
config:
max_tokens: 4096
effort: high
- id: anthropic:messages:claude-mythos-5-1
config:
max_tokens: 4096
effort: high
Both models have a 1M-token context window and support up to 128K output tokens. Input and output cost $10 and $50 per million tokens, respectively. Cache reads cost $0.25 per million tokens, down from $1 on Fable 5 and Mythos 5; promptfoo includes this discount in its cost estimates. See Anthropic's pricing.
Thinking is always on, and the sampling and thinking normalization described below
also applies to 5.1. Unlike Fable 5, 5.1 rejects forced tool use: promptfoo omits
tool_choice with type any or tool and warns. Use auto or none instead.
When replaying Fable 5.1 thinking blocks, keep earlier messages, system prompts,
and tools unchanged; edited prefixes can cause API errors. See
Anthropic's migration notes.
Claude Fable 5 and Mythos 5 notes
Fable 5 and Mythos 5 use always-on adaptive thinking. Promptfoo omits unsupported
temperature, top_p, and top_k values, converts legacy
thinking: { type: 'enabled', budget_tokens: N } configs to adaptive thinking, and
omits thinking: { type: 'disabled' } because thinking cannot be disabled.
Set thinking: { type: 'adaptive', display: 'summarized' } to include a readable
thinking summary; the default display: 'omitted' returns an empty thinking block,
which Promptfoo excludes from the output.
Both models use a 1M-token context window, support up to 128K output tokens, and are priced at $10 per million input tokens and $50 per million output tokens. Mythos 5 access is limited through Project Glasswing and may require provider approval. Both model IDs are pinned.
Claude Opus 5.5 notes
Opus 5.5 is priced below Opus 5 and has the same context window, maximum output, and tokenizer. Its request rules match Fable 5.1's, and promptfoo adjusts requests to follow them:
- Thinking is always on. Opus 5.5 rejects
thinking: { type: 'disabled' }at every effort level, so promptfoo removes it and logs a warning once. Manualthinking: { type: 'enabled', budget_tokens: N }configs becomethinking: { type: 'adaptive' }. effortdefaults tomedium, one level below Opus 5'shigh. Seteffortexplicitly when you compare the two. It is the only way to control how much the model thinks.- Forced tool use and sampling controls are rejected. Promptfoo omits
tool_choicevalues of typeanyortool(useautoornone) and all oftemperature,top_p, andtop_k.
Opus 5.5 costs a flat $4 per million input tokens and $20 per million output tokens across
its 1M-token context window. Cache reads cost $0.20 per million tokens, and promptfoo's cost
estimates include that rate. To track Anthropic's fast mode, set inputCost: 8 / 1e6 and
outputCost: 40 / 1e6.
providers:
- id: anthropic:messages:claude-opus-5-5
config:
effort: high
max_tokens: 16000
Claude Opus 5 notes
Opus 5 is the Opus-tier Claude 5 model, aimed at complex agentic coding and long-horizon work. It keeps Opus 4.8's request surface and pricing, with two behavior changes promptfoo handles for you:
- Thinking is on by default. Unlike Opus 4.7/4.8 — where omitting
thinkingmeant no extended thinking — an omittedthinkingblock on Opus 5 runs adaptive thinking. Becausemax_tokenscaps thinking plus response text, promptfoo sizes its defaultmax_tokenswith thinking headroom (2048 instead of 1024) so responses aren't truncated mid-answer. Setmax_tokensexplicitly for anything longer. - Disabling thinking is effort-gated.
thinking: { type: 'disabled' }is only accepted atefforthighor below; pairing it withxhighormaxreturns a 400. Promptfoo drops the rejectedthinking: { type: 'disabled' }(keeping youreffort) and logs a one-time warning. Lowerefforttohighif you actually need thinking off. - Sampling controls are managed for you. Like Opus 4.7/4.8, Opus 5 rejects
temperature,top_p, andtop_kwith a 400; promptfoo omits all three from every request, including its built-intemperature: 0default. A legacythinking: { type: 'enabled', budget_tokens: N }config is converted tothinking: { type: 'adaptive' }. - The full
low→maxeffort ladder is available. Start atxhighfor coding and agentic work, then sweep downward —lowandmediumare unusually strong on this model and are the main cost and latency lever. See the Effort Level section.
Opus 5 uses a 1M-token context window (both the default and the maximum) billed at a flat
$5 per million input / $25 per million output — the same list rates as Opus 4.8, with no
long-context surcharge above 200K tokens. Anthropic's fast mode ($10 / $50, Claude API only)
is a separate research-preview rate that promptfoo does not encode. To track it, set
inputCost: 10 / 1e6 and outputCost: 50 / 1e6 — a single cost cannot express asymmetric
rates, because it is applied as both the input and the output per-token price.
providers:
- id: anthropic:messages:claude-opus-5
config:
effort: xhigh
max_tokens: 8192
Claude Sonnet 5.5 notes
Sonnet 5.5 has a 1M-token context window and a 128K-token output limit. Promptfoo adjusts requests to its API requirements:
- Thinking is on by default, and
disabledis rejected. Sonnet 5.5's lowest thinking setting isthinking: { type: 'between_tools' }, which turns off up-front thinking. Promptfoo sendsbetween_toolsin place ofdisabledand logs a warning once.between_toolsis only accepted atefforthighor below, so withxhighormaxpromptfoo omitsdisabledorbetween_tools, warns, and the model thinks adaptively. effortdefaults tohigh. Compare effort levels for your workload when migrating.- Forced tool use is rejected. Promptfoo omits
tool_choicevalues of typeanyortool. Useautoornone, and say in the prompt when a tool applies. - Sampling controls and manual budgets are rejected, as on Sonnet 5. Promptfoo omits
temperature,top_p, andtop_k, and convertsthinking: { type: 'enabled', budget_tokens: N }tothinking: { type: 'adaptive' }. - Text between tool calls comes back in
thinkingblocks. With the defaultdisplay: 'omitted'those blocks are empty. Setthinking: { type: 'adaptive', display: 'summarized' }or usebetween_toolsto keep that text in the output.
Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens across its 1M-token context window, with cache reads at $0.20 per million tokens.
providers:
- id: anthropic:messages:claude-sonnet-5-5
config:
effort: medium
thinking:
type: between_tools # no up-front thinking; accepted at effort high or below
max_tokens: 4096
Claude Sonnet 5 notes
Sonnet 5 has a 1M-token context window and supports effort levels
from low through max. It does not support manual sampling controls:
-
Sampling controls: Promptfoo omits
temperature,top_p, andtop_k, including its default temperature. Explicit settings trigger a one-time warning. The parameters are also omitted through AWS Bedrock, GCP Vertex, and Azure AI Foundry. -
Manual thinking budgets convert to adaptive. A legacy
thinking: { type: 'enabled', budget_tokens: N }config is converted tothinking: { type: 'adaptive' }; useeffortto control reasoning depth. -
Adaptive thinking is on by default. Requests without a
thinkingfield still use adaptive thinking. Passthinking: { type: 'disabled' }to turn it off, and leave enoughmax_tokensheadroom for thinking plus the visible response.
Sonnet 5 uses a 1M-token context window billed at $2 per million input / $10 per million output, with no long-context surcharge above 200K tokens. Anthropic made these rates permanent on August 10, 2026, canceling the planned September increase. See Anthropic's pricing documentation. The newer tokenizer can produce more tokens for the same text, so compare total request costs when migrating from Sonnet 4.6.
Claude Opus 4.8 notes
Opus 4.8 uses the same sampling and thinking settings as Opus 4.7:
- Sampling controls are managed for you. Like Opus 4.7, Opus 4.8 samples adaptively and rejects
temperature,top_p, andtop_k(any of them returns a 400); promptfoo omits all three from every request. Setting any of them in config orANTHROPIC_TEMPERATURElogs a one-time heads-up so you can clean the values out of your eval. - Adaptive thinking is opt-in. Set
thinking: { type: 'adaptive' }to let the model decide how much to reason per request. Without an explicitthinkingblock the model runs without extended thinking, even at high effort. Manual budget-based thinking (thinking: { type: 'enabled', budget_tokens: N }) is rejected with a 400. effortdefaults tohighandxhighis available. Settingeffort: highbehaves the same as omitting it. Start withxhighfor coding and agentic work. See the Effort Level section.
The same suppression applies when you reach Opus 4.8 through AWS Bedrock, GCP Vertex, or Azure AI Foundry — promptfoo omits the unsupported sampling parameters on each of those paths too (silently; the one-time warning above is specific to the Anthropic Messages provider).
Claude Opus 4.7 notes
Opus 4.7 supports adaptive thinking when explicitly enabled:
- Temperature is managed for you. Opus 4.7 samples adaptively and does not accept
temperature; promptfoo omits the field from every request. Passingtemperaturein config orANTHROPIC_TEMPERATURElogs a one-time heads-up so you can clean the value out of your eval. - Adaptive thinking is opt-in. Set
thinking: { type: 'adaptive' }to let the model choose how much to reason per request; leavingthinkingunset runs Opus 4.7 without extended thinking, even at high effort. (Opus 5 is the model where an omitted block means adaptive.) Budget-based modes from older models aren't used on 4.7. xhigheffort level is available. It sits betweenhighandmaxand is a good starting point for coding and agentic tasks. See the Effort Level section.- Updated tokenizer. The same input can map to 1.0–1.35× more tokens than Opus 4.6, so measure real traffic if you're comparing costs.
The same guidance applies when you reach Opus 4.7 through AWS Bedrock, GCP Vertex, or Azure AI Foundry — promptfoo suppresses temperature on each of those paths as well.
Extended Thinking
Use thinking to configure reasoning before the final answer. When available,
display: summarized returns a summary of that reasoning:
providers:
# Adaptive thinking
- id: anthropic:messages:claude-opus-5
config:
max_tokens: 20000
thinking:
type: 'adaptive'
display: 'summarized' # Opt in to readable reasoning; default is 'omitted'
effort: xhigh # Controls reasoning depth instead of a token budget
# Manual thinking budget — Opus 4.6 / Sonnet 4.6 and older (deprecated on 4.6)
- id: anthropic:messages:claude-haiku-4-5
config:
max_tokens: 20000
thinking:
type: 'enabled'
budget_tokens: 16000 # Must be ≥1024 and less than max_tokens
The thinking configuration has three possible values:
- Adaptive thinking (supported on Claude 4.6 and later):
thinking:
type: 'adaptive'
In adaptive mode, Claude decides when and how much to think based on the complexity of the request. Control depth with effort rather than a token budget.
- Manual thinking budgets (Claude 4.5 and 4.6):
thinking:
type: 'enabled'
budget_tokens: 16000 # Must be at least 1024 and less than max_tokens
Claude 4.5 and 4.6 accept manual budgets; this mode is deprecated on 4.6. On
adaptive-only models, Promptfoo converts enabled to adaptive, removes the
budget, and logs a warning.
- Disabled thinking:
thinking:
type: 'disabled'
Not accepted on Opus 5.5 or Fable 5 / Mythos 5 / Fable 5.1 / Mythos 5.1, where thinking is always on,
nor on Opus 5 above effort: high. Promptfoo omits it in both cases and warns.
On Sonnet 5.5, it becomes between_tools at high effort or below and is omitted at higher effort.
- Between tools (Claude Sonnet 5.5 only):
thinking:
type: between_tools
This mode turns off up-front thinking and requires effort: high or below. Progress updates
between tool calls still return as thinking blocks. It takes no other fields, including display
or budget_tokens. At xhigh or max, Promptfoo omits it so the model uses adaptive thinking.
Thinking display
display controls the reasoning text returned by the API. Omitting that text does
not disable thinking or remove its token cost:
| Value | Behavior |
|---|---|
'omitted' | Default on Fable 5/5.1, Mythos 5/5.1, Opus 5.5, Opus 5, Opus 4.7/4.8, and Sonnet 5. The thinking block is returned with empty text plus a signature for multi-turn continuity. |
'summarized' | Returns a readable summary of the reasoning. Default on Claude 4.6 and earlier. |
Set display: summarized to request a readable summary:
thinking:
type: adaptive
display: summarized
Thinking tokens count toward max_tokens in both modes. Only manual thinking
requires budget_tokens, which must be at least 1,024 and below max_tokens.
Promptfoo omits temperature and top_k when thinking is enabled, and clamps
top_p to [0.95, 1.0] on models that support it. On models without sampling
controls, all three parameters are omitted.
Forced tool use (tool_choice type any or tool) is incompatible with manual
thinking and with Opus 5.5, Sonnet 5.5, Fable 5.1, and Mythos 5.1. Promptfoo omits it with a
warning in those cases; use auto or none. Other adaptive models accept forced
tool use.
Example response with thinking enabled:
{
"content": [
{
"type": "thinking",
"thinking": "Let me analyze this step by step...",
"signature": "WaUjzkypQ2mUEVM36O2TxuC06KN8xyfbJwyem2dw3URve/op91XWHOEBLLqIOMfFG/UvLEczmEsUjavL...."
},
{
"type": "text",
"text": "Based on my analysis, here is the answer..."
}
]
}
Controlling Thinking Output
thinking.display controls what the API returns; showThinking controls whether
Promptfoo includes that text in the output (default: true). It cannot reveal
reasoning omitted by the API. For example, request a summary but exclude it from
the output being graded:
providers:
- id: anthropic:messages:claude-sonnet-5
config:
thinking:
type: 'adaptive'
display: 'summarized'
showThinking: false # Exclude thinking content from the output
Redacted Thinking
Sometimes Claude's internal reasoning may be flagged by safety systems. When this occurs, the thinking block will be encrypted and returned as a redacted_thinking block:
{
"content": [
{
"type": "redacted_thinking",
"data": "EmwKAhgBEgy3va3pzix/LafPsn4aDFIT2Xlxh0L5L8rLVyIwxtE3rAFBa8cr3qpP..."
},
{
"type": "text",
"text": "Based on my analysis..."
}
]
}
Redacted thinking blocks are automatically decrypted when passed back to the API, allowing Claude to maintain context without compromising safety guardrails.
Extended Output with Thinking
For longer responses, raise max_tokens and enable streaming:
providers:
- id: anthropic:messages:claude-sonnet-5
config:
max_tokens: 64000 # Sonnet 5 supports up to 128K output tokens
stream: true # Required above 21,333 output tokens
thinking:
type: 'adaptive'
effort: high
Fable 5/5.1, Opus 5.5, Opus 5, Opus 4.6–4.8, Sonnet 5, and Sonnet 4.6 all support up to 128K output
tokens; the 4.5 generation caps at 64K. The output-128k-2025-02-19 beta feature is specific to
Claude 3.7 Sonnet and is not needed on any current model.
When using extended output:
- Streaming is required when
max_tokensis greater than 21,333 - Thinking shares the
max_tokensbudget with the answer, so leave headroom for both - The model may not use the entire allocated budget
See Anthropic's Extended Thinking Guide for more details on requirements and best practices.
Effort Level
effort controls how many tokens Claude spends on a response, including thinking
and tool calls. Higher settings can increase quality, cost, and latency. It is not
a sampling-temperature control or a hard token limit.
providers:
- id: anthropic:messages:claude-opus-5
config:
effort: xhigh # Options: low, medium, high, xhigh, max
Support varies by model — sending an unsupported level returns a 400:
| Model | Supported levels |
|---|---|
| Fable/Mythos 5 and 5.1, Opus 5.5, Opus 5, Sonnet 5, Opus 4.7/4.8 | low, medium, high, xhigh, max |
| Opus 4.6, Sonnet 4.6 | low, medium, high, max (no xhigh) |
| Opus 4.5 | low, medium, high |
| Sonnet 4.5, Haiku 4.5 | Not supported — omit effort |
The API defaults to medium on Opus 5.5 and high on the other models that support
effort. Set it explicitly when comparing models. See Anthropic's
effort guide for model-specific guidance.
This can be combined with other features like structured outputs:
providers:
- id: anthropic:messages:claude-opus-5
config:
effort: high
output_format:
type: json_schema
schema:
type: object
properties:
analysis:
type: string
required:
- analysis
additionalProperties: false
Structured Outputs
Structured outputs constrain responses to a JSON schema and are supported by
current Claude models. Promptfoo maps config.output_format to the API's
output_config.format and adds the structured-outputs-2025-11-13 beta flag to
the anthropic-beta header.
JSON Outputs
Add output_format to get structured responses:
providers:
- id: anthropic:messages:claude-sonnet-5
config:
output_format:
type: json_schema
schema:
type: object
properties:
name:
type: string
email:
type: string
required:
- name
- email
additionalProperties: false
You can also load the entire output_format from an external file:
config:
output_format: file://./schemas/analysis-format.json
Nested file references are supported for the schema:
{
"type": "json_schema",
"schema": "file://./schemas/analysis-schema.json"
}
Variable rendering is supported in file paths:
config:
output_format: file://./schemas/{{ schema_name }}.json
Strict Tool Use
Add strict: true to tool definitions for schema-validated parameters:
providers:
- id: anthropic:messages:claude-sonnet-5
config:
tools:
- name: get_weather
strict: true
input_schema:
type: object
properties:
location:
type: string
required:
- location
additionalProperties: false
Limitations
Supported: object, array, string, integer, number, boolean, null, enum, required, additionalProperties: false
Not supported: recursive schemas, minimum/maximum, minLength/maxLength
Incompatible with: citations, message prefilling
See Anthropic's guide and the structured outputs example.
Model-Graded Tests
Model-graded assertions such as factuality or llm-rubric will automatically use Anthropic as the grading provider if ANTHROPIC_API_KEY is set and OPENAI_API_KEY is not set.
If both API keys are present, OpenAI will be used by default. You can explicitly override the grading provider in your configuration.
Claude Pro/Max subscribers without a separate Anthropic Console key can wire up llm-rubric through a local Claude Code session by pointing the grader at anthropic:messages:<model> with apiKeyRequired: false:
defaultTest:
options:
provider:
id: anthropic:messages:claude-sonnet-5
config:
apiKeyRequired: false
See Authenticating via a Claude Code session above for how the credential is loaded and what beta headers Promptfoo sets.
Because of how model-graded evals are implemented, the model must support chat-formatted prompts (except for embedding or classification models).
You can override the grading provider in several ways:
- For all test cases using
defaultTest:
defaultTest:
options:
provider: anthropic:messages:claude-sonnet-5
- For individual assertions:
assert:
- type: llm-rubric
value: Do not mention that you are an AI or chat assistant
provider:
id: anthropic:messages:claude-sonnet-5
config:
effort: low
- For specific tests:
tests:
- vars:
question: What is the capital of France?
options:
provider:
id: anthropic:messages:claude-sonnet-5
assert:
- type: llm-rubric
value: Answer should mention Paris
Additional Capabilities
- Caching: Promptfoo caches previous LLM requests by default.
- Token Usage Tracking: Provides detailed information on the number of tokens used in each request, aiding in usage monitoring and optimization.
- Cost Calculation: Estimates each request from its model and reported token usage.
For Claude 4.6 and later, U.S. inference adds 10% to token costs, including cache
reads and writes. The response's
usage.inference_geotakes precedence overconfig.extra_body.inference_geo. Explicitcost,inputCost, oroutputCostoverrides take precedence over this surcharge. See Anthropic pricing.
When using the Anthropic disk response cache across runs, assign each provider a distinct, non-secret label. Promptfoo uses that stable label to isolate cached responses without persisting API-key or OAuth-token fingerprints. Unlabeled providers receive an ephemeral per-instance namespace, which safely preserves repeated calls within a run without reusing responses across tenants or processes. Requests with custom headers bypass the disk response cache because those headers may contain tenant credentials.
providers:
- id: anthropic:messages:claude-sonnet-5
label: tenant-a
config:
apiKey: '{{env.ANTHROPIC_API_KEY_A}}'
- id: anthropic:messages:claude-sonnet-5
label: tenant-b
config:
apiKey: '{{env.ANTHROPIC_API_KEY_B}}'
See Also
Examples
We provide several example implementations demonstrating Claude's capabilities:
Core Features
- Tool Use Example - Shows how to use Claude's tool calling capabilities
- MCP Example - Connect Claude to a Model Context Protocol server and let it execute the discovered tools
- Structured Outputs Example - Demonstrates JSON outputs and strict tool use for guaranteed schema compliance
- Vision Example - Demonstrates using Claude's vision capabilities
Model Comparisons & Evaluations
- Claude vs GPT - Compares Claude with GPT-5.4 on various tasks
- Claude vs GPT Image Analysis - Compares Claude's and GPT's image analysis capabilities
Cloud Platform Integrations
- Azure AI Foundry - Using Claude through Azure AI Foundry
- AWS Bedrock - Using Claude through AWS Bedrock
- Google Vertex AI - Using Claude through Google Vertex AI
Agentic Evaluations
- Claude Agent SDK - For agentic evals with file access, tool use, and MCP servers
For more examples and general usage patterns, visit our examples directory on GitHub.