Llama 3.1 405B Instruct Turbo API
Pricing & Specs
Meta · $3.5 input / $3.5 output per 1M tokens · 130.815K context window
No longer offered on OpenMark
Llama 3.1 405B Instruct Turbo API is no longer offered on OpenMark (deprecated or superseded upstream). Its specs are kept here for reference — if you're still running it in production, it's worth benchmarking current alternatives: newer models are often cheaper AND better on the same task.
Llama 3.1 405B Instruct Turbo API API Pricing
| Tokens | Price |
|---|---|
| Input | $3.5 / 1M tokens |
| Output | $3.5 / 1M tokens |
50% discount available for batch API requests. Batch API pricing is available via Together at 50%; cache pricing is not published.
What that means in practice
Llama 3.1 405B Instruct Turbo API Specs
| Spec | Value |
|---|---|
| Context window | 130.815K tokens |
| Input modalities | text |
| Output modalities | text |
| Latency (measured by OpenMark) | ~1.2s median response |
| Reasoning model | No |
| Tool / function calling | Yes |
| JSON mode | Yes |
| Streaming | Yes |
| Prompt caching | No |
| Batch API | Yes |
Capabilities
FAQ
How much does Llama 3.1 405B Instruct Turbo API cost?
Llama 3.1 405B Instruct Turbo API costs $3.5 per 1M input tokens and $3.5 per 1M output tokens.
What is Llama 3.1 405B Instruct Turbo API's context window?
Llama 3.1 405B Instruct Turbo API has a 130.815K-token context window.
Can I test Llama 3.1 405B Instruct Turbo API on my own task?
Llama 3.1 405B Instruct Turbo API is no longer offered on OpenMark, but you can benchmark its current alternatives from Meta and other providers on your own task — free tier available.
Still using Llama 3.1 405B Instruct Turbo API?
It's been superseded. Benchmark its current alternatives on your actual
workload — you'll likely find something cheaper AND better. Free tier available.