Skip to main content

DeepSeek

DeepSeek provides an OpenAI-compatible chat API. The provider accepts OpenAI chat options; DeepSeek determines which options each model supports.

Setup​

  1. Get an API key from the DeepSeek Platform
  2. Set DEEPSEEK_API_KEY environment variable or specify apiKey in your config

Configuration​

Basic configuration example:

providers:
- id: deepseek:deepseek-flash
config:
max_tokens: 4000
passthrough:
thinking: { type: disabled }

- id: deepseek:deepseek-v4-pro
config:
max_tokens: 8192
showThinking: true
passthrough:
thinking:
type: enabled
reasoning_effort: high

Configuration Options​

  • temperature
  • max_tokens
  • reasoning_effort - none disables thinking; low, high, and max enable it. DeepSeek defaults to high. When neither max_tokens nor OPENAI_MAX_TOKENS is set, Promptfoo uses DeepSeek's output budget.
  • cost, inputCost, outputCost, cacheReadCost - Set cost estimates in USD per token. inputCost and outputCost take precedence over cost; cacheReadCost sets a separate cached-input rate.
  • top_p, presence_penalty, frequency_penalty
  • showThinking - Control whether reasoning content is included in the output (default: true, applies to thinking-capable models)

Promptfoo requests complete responses; this provider does not support streaming.

Available Models​

DeepSeek lists deepseek-flash and deepseek-v4-pro in its model catalog. The older deepseek-chat and deepseek-reasoner IDs are retired. The shorthand deepseek: uses deepseek-flash with thinking disabled; use the full ID for DeepSeek's default thinking mode.

deepseek-flash​

Use deepseek:deepseek-flash for V4.1 Flash, which supports text and image inputs. The older deepseek-v4-flash and deepseek-v4-flash-vision-exp IDs route to the same model. It supports a 1M-token context window and up to 384K output tokens.

deepseek-v4-pro​

V4 Pro supports text input, thinking and non-thinking modes, a 1M-token context window, and up to 384K output tokens.

DeepSeek charges different peak and off-peak rates. Promptfoo estimates current models at peak rates: Flash costs $0.30 input, $0.006 cached input, and $1.20 output per million tokens; Pro costs $1.32, $0.044, and $3.96 respectively. Off-peak rates are half these amounts. Set inputCost, outputCost, and cacheReadCost in USD per token to override the estimate using the current rates. Promptfoo does not infer the billing period or Chinese public holidays.

warning

Sampling support differs by mode. temperature has no effect in thinking mode. top_p only affects thinking mode and values below 0.95 are treated as 0.95. See the Chat Completions reference for parameter limits.

Example Usage​

Compare DeepSeek with OpenAI on a reasoning task:

promptfooconfig.yaml
# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
providers:
- id: deepseek:deepseek-v4-pro
config:
max_tokens: 8000
showThinking: true # Include reasoning content in output (default)
- id: openai:gpt-5.4-mini
config:
reasoning_effort: medium

prompts:
- 'Solve this step by step: {{math_problem}}'

tests:
- vars:
math_problem: 'What is the derivative of x^3 + 2x with respect to x?'

Controlling Reasoning Output​

Set showThinking: false to exclude reasoning content from the output:

providers:
- id: deepseek:deepseek-v4-pro
config:
showThinking: false # Hide reasoning content from output
passthrough:
thinking:
type: enabled

With showThinking: true (the default), the output includes reasoning when DeepSeek returns it:

Thinking: <reasoning content>

<final answer>

With showThinking: false, assertions see only the final answer. This option does not turn off thinking at the API; use config.passthrough.thinking: { type: disabled } for that.

API Details​

See Also​