Nscale
Use Nscale's Serverless Inference API for OpenAI-compatible chat, completion, embedding, and image requests.
Setup
Set your Nscale service token as an environment variable:
export NSCALE_SERVICE_TOKEN=your_service_token_here
Alternatively, you can add it to your .env file:
NSCALE_SERVICE_TOKEN=your_service_token_here
You can also supply config.apiKey or select a credential variable with config.apiKeyEnvar. Nscale does not fall back to OPENAI_API_KEY by default; set apiKeyEnvar: OPENAI_API_KEY to use that variable explicitly. The same credential rules apply when you set a custom apiBaseUrl.
Obtaining Credentials
You can obtain service tokens by:
- Signing up at Nscale
- Navigating to your account settings
- Going to "Service Tokens" section
Configuration
To use Nscale models in your promptfoo configuration, use the nscale: prefix followed by the model name:
providers:
- nscale:openai/gpt-oss-120b
- nscale:meta-llama/Llama-3.3-70B-Instruct
- nscale:Qwen/Qwen3-235B-A22B-Instruct-2507
Model IDs are the upstream Hugging Face repository IDs and are case-sensitive.
Model Types
Nscale supports different types of models through specific endpoint formats:
Chat Completion Models (Default)
For chat completion models, you can use either format:
providers:
- nscale:chat:<model-id>
- nscale:<model-id> # Defaults to chat
Completion Models
For text completion models:
providers:
- nscale:completion:<model-id>
Embedding Models
For embedding models:
providers:
- nscale:embedding:Qwen3-Embedding-8B
- nscale:embeddings:Qwen3-Embedding-8B # Alternative format
Text-to-Image Models
For image generation models:
providers:
- nscale:image:black-forest-labs/FLUX.1-schnell
Popular Models
The authoritative list for your account is GET https://inference.api.nscale.com/v1/models,
which also returns pricing and context length:
curl -fsS https://inference.api.nscale.com/v1/models \
-H "Authorization: Bearer $NSCALE_SERVICE_TOKEN"
Text Generation Models
Use a returned id after the nscale: or nscale:chat: prefix.
| Model | Provider Format | Use Case |
|---|---|---|
| GPT OSS 120B | nscale:openai/gpt-oss-120b | General-purpose reasoning and tasks |
| GPT OSS 20B | nscale:openai/gpt-oss-20b | Lightweight general-purpose model |
| Kimi K2.5 | nscale:moonshotai/Kimi-K2.5 | Large-scale agentic reasoning |
| Qwen 3 235B A22B | nscale:Qwen/Qwen3-235B-A22B | Large-scale language understanding |
| Qwen 3 235B A22B Instruct 2507 | nscale:Qwen/Qwen3-235B-A22B-Instruct-2507 | Qwen 3 235B 2507 variant |
| Qwen 3 4B Instruct 2507 | nscale:Qwen/Qwen3-4B-Instruct-2507 | Lightweight instruction following |
| Qwen 3 4B Thinking 2507 | nscale:Qwen/Qwen3-4B-Thinking-2507 | Reasoning and thinking tasks |
| Qwen 3 8B | nscale:Qwen/Qwen3-8B | Mid-size general-purpose model |
| Qwen 3 14B | nscale:Qwen/Qwen3-14B | Enhanced reasoning capabilities |
| Qwen 3 32B | nscale:Qwen/Qwen3-32B | Large-scale reasoning and analysis |
| Qwen 2.5 Coder 3B Instruct | nscale:Qwen/Qwen2.5-Coder-3B-Instruct | Lightweight code generation |
| Qwen 2.5 Coder 7B Instruct | nscale:Qwen/Qwen2.5-Coder-7B-Instruct | Code generation and programming |
| Qwen 2.5 Coder 32B Instruct | nscale:Qwen/Qwen2.5-Coder-32B-Instruct | Advanced code generation |
| Qwen QwQ 32B | nscale:Qwen/QwQ-32B | Specialized reasoning model |
| Llama 3.3 70B Instruct | nscale:meta-llama/Llama-3.3-70B-Instruct | High-quality instruction following |
| Llama 3.1 8B Instruct | nscale:meta-llama/Llama-3.1-8B-Instruct | Efficient instruction following |
| Llama 3.2 11B Vision Instruct | nscale:meta-llama/Llama-3.2-11B-Vision-Instruct | Vision-language tasks |
| Llama 4 Scout 17B | nscale:meta-llama/Llama-4-Scout-17B-16E-Instruct | Image-Text-to-Text capabilities |
| DeepSeek R1 Distill Llama 70B | nscale:deepseek-ai/DeepSeek-R1-Distill-Llama-70B | Efficient reasoning model |
| DeepSeek R1 Distill Llama 8B | nscale:deepseek-ai/DeepSeek-R1-Distill-Llama-8B | Lightweight reasoning model |
| DeepSeek R1 Distill Qwen 1.5B | nscale:deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B | Ultra-lightweight reasoning |
| DeepSeek R1 Distill Qwen 7B | nscale:deepseek-ai/DeepSeek-R1-Distill-Qwen-7B | Compact reasoning model |
| DeepSeek R1 Distill Qwen 14B | nscale:deepseek-ai/DeepSeek-R1-Distill-Qwen-14B | Mid-size reasoning model |
| DeepSeek R1 Distill Qwen 32B | nscale:deepseek-ai/DeepSeek-R1-Distill-Qwen-32B | Large reasoning model |
| Devstral Small 2505 | nscale:mistralai/Devstral-Small-2505 | Code generation and development |
| Mixtral 8x22B Instruct | nscale:mistralai/Mixtral-8x22B-Instruct-v0.1 | Large mixture-of-experts model |
Embedding Models
Nscale's embedding API reference demonstrates Qwen3-Embedding-8B. Confirm it in
your organization's /v1/models response before running an eval.
Text-to-Image Models
Nscale's image API reference demonstrates
nscale:image:black-forest-labs/FLUX.1-schnell. Confirm it in your organization's catalog before
running an eval.
| Model | Provider Format | Use Case |
|---|---|---|
| Flux.1 Schnell | nscale:image:black-forest-labs/FLUX.1-schnell | Fast image generation |
| Stable Diffusion XL | nscale:image:stabilityai/stable-diffusion-xl-base-1.0 | High-quality image generation |
| SDXL Lightning | nscale:image:ByteDance/SDXL-Lightning | Ultra-fast image generation |
Configuration Options
Nscale supports standard OpenAI-compatible parameters:
providers:
- id: nscale:meta-llama/Llama-4-Scout-17B-16E-Instruct
config:
temperature: 0.7
max_tokens: 1024
top_p: 0.9
frequency_penalty: 0.1
presence_penalty: 0.2
stop: ['END', 'STOP']
seed: 42
Supported Parameters
temperature: Controls randomness (0.0 to 2.0). Defaults to0unless set.max_tokens: Maximum number of tokens to generate. Defaults to1024unless set.top_p: Nucleus sampling parameterfrequency_penalty: Reduces repetition based on frequencypresence_penalty: Reduces repetition based on presencestop: Stop sequences to halt generationseed: Deterministic sampling seed
Any other parameter is forwarded to the Nscale API unchanged.
Streaming is not supported. Promptfoo reads each response as a single JSON body, so
setting stream: true produces a response it cannot parse.
Example Configuration
Here's a complete example configuration:
# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
providers:
- id: nscale:openai/gpt-oss-120b
label: nscale-gpt-oss
config:
temperature: 0.7
max_tokens: 512
- id: nscale:meta-llama/Llama-3.3-70B-Instruct
label: nscale-llama
config:
temperature: 0.5
max_tokens: 1024
prompts:
- 'Explain {{concept}} in simple terms'
- 'What are the key benefits of {{concept}}?'
tests:
- vars:
concept: quantum computing
assert:
- type: contains
value: 'quantum'
Pricing
Nscale prices text generation and embeddings per token, and images per megapixel. Check Nscale's model endpoint API for current rates available to your organization.
Promptfoo has no Nscale-specific token price table. For chat, completion, or embedding estimates, set cost or inputCost/outputCost in USD per token. Image estimates use fixed model rates multiplied by the number of images; they do not adjust for resolution or use those token-cost settings.
Key Features
Nscale hosts the models and exposes an OpenAI-compatible API. See Nscale's documentation for throughput, rate limits, and available regions.
Error Handling
The Nscale provider includes built-in error handling for common issues:
- Network timeouts and retries
- Rate limiting
- Invalid API key errors
- Model availability issues
Support
For support with the Nscale provider: