DeepSeek
DeepSeek provides an OpenAI-compatible chat API. The provider accepts OpenAI chat options; DeepSeek determines which options each model supports.
Setup
- Get an API key from the DeepSeek Platform
- Set
DEEPSEEK_API_KEYenvironment variable or specifyapiKeyin your config
Configuration
Basic configuration example:
providers:
- id: deepseek:deepseek-flash
config:
max_tokens: 4000
passthrough:
thinking: { type: disabled }
- id: deepseek:deepseek-v4-pro
config:
max_tokens: 8192
showThinking: true
passthrough:
thinking:
type: enabled
reasoning_effort: high
Configuration Options
temperaturemax_tokensreasoning_effort-nonedisables thinking;low,high, andmaxenable it. DeepSeek defaults tohigh. When neithermax_tokensnorOPENAI_MAX_TOKENSis set, Promptfoo uses DeepSeek's output budget.cost,inputCost,outputCost,cacheReadCost- Set cost estimates in USD per token.inputCostandoutputCosttake precedence overcost;cacheReadCostsets a separate cached-input rate.top_p,presence_penalty,frequency_penaltyshowThinking- Control whether reasoning content is included in the output (default:true, applies to thinking-capable models)
Promptfoo requests complete responses; this provider does not support streaming.
Available Models
DeepSeek lists deepseek-flash and deepseek-v4-pro in its model catalog. The older deepseek-chat and deepseek-reasoner IDs are retired. The shorthand deepseek: uses deepseek-flash with thinking disabled; use the full ID for DeepSeek's default thinking mode.
deepseek-flash
Use deepseek:deepseek-flash for V4.1 Flash, which supports text and image inputs. The older deepseek-v4-flash and deepseek-v4-flash-vision-exp IDs route to the same model. It supports a 1M-token context window and up to 384K output tokens.
deepseek-v4-pro
V4 Pro supports text input, thinking and non-thinking modes, a 1M-token context window, and up to 384K output tokens.
DeepSeek charges different peak and off-peak rates. Promptfoo estimates current models at peak rates: Flash costs $0.30 input, $0.006 cached input, and $1.20 output per million tokens; Pro costs $1.32, $0.044, and $3.96 respectively. Off-peak rates are half these amounts. Set inputCost, outputCost, and cacheReadCost in USD per token to override the estimate using the current rates. Promptfoo does not infer the billing period or Chinese public holidays.
Sampling support differs by mode. temperature has no effect in thinking mode. top_p only affects thinking mode and values below 0.95 are treated as 0.95. See the Chat Completions reference for parameter limits.
Example Usage
Compare DeepSeek with OpenAI on a reasoning task:
# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
providers:
- id: deepseek:deepseek-v4-pro
config:
max_tokens: 8000
showThinking: true # Include reasoning content in output (default)
- id: openai:gpt-5.4-mini
config:
reasoning_effort: medium
prompts:
- 'Solve this step by step: {{math_problem}}'
tests:
- vars:
math_problem: 'What is the derivative of x^3 + 2x with respect to x?'
Controlling Reasoning Output
Set showThinking: false to exclude reasoning content from the output:
providers:
- id: deepseek:deepseek-v4-pro
config:
showThinking: false # Hide reasoning content from output
passthrough:
thinking:
type: enabled
With showThinking: true (the default), the output includes reasoning when DeepSeek returns it:
Thinking: <reasoning content>
<final answer>
With showThinking: false, assertions see only the final answer. This option does not turn off thinking at the API; use config.passthrough.thinking: { type: disabled } for that.
API Details
- Base URL:
https://api.deepseek.com/v1 - OpenAI-compatible API format
- DeepSeek API documentation
See Also
- OpenAI Provider - Compatible configuration options
- Historical MMLU comparison - Replace its retired provider IDs with the current IDs shown above before running it.