Llama 4 Scout (17B-16E Instruct)

GA

Open-weight, natively multimodal MoE: 17B active / 109B total params, 16 experts; fits on a single H100. License: Llama 4 Community License Agreement. Open weights support up to 10M-token context; the official Llama API serves it at 128k (model ID 'Llama-4-Scout-17B-16E-Instruct-FP8'). Hosted price is OpenRouter slug 'meta-llama/llama-4-scout' = $0.10 in / $0.30 out per 1M (page accessed 2026-06-28); Together AI reported ~$0.08 in / $0.30 out per 1M. No cached-input discount published. Same 12 supported languages as Maverick.

Llama 4 Scout (17B-16E Instruct) by Meta costs $0.1 per 1M input tokens and $0.3 per 1M output tokens ($0.15/1M blended), with a 10M (10.000.000-token) context window. It is generally available (GA), with a knowledge cutoff of Aug 2024. Its weights are publicly downloadable, so it can be self-hosted on your own infrastructure.

Last verified: 13 Aug 2026 · sourced from official provider documentation

Provider
Meta
Status
GA
Input price
$0.1 / 1M tokens
Output price
$0.3 / 1M tokens
Cached input
Blended price
$0.15 / 1M tokens
Context window
10.000.000 tokens (10M)
Max output
Modality
text, image
Open weights
Yes — self-hostable
Knowledge cutoff
2024-08
Released
5 Apr 2025
API string
meta-llama/Llama-4-Scout-17B-16E-Instruct-FP8

Source: Meta official documentation ↗

FAQ

Llama 4 Scout (17B-16E Instruct) — questions & answers

How much does Llama 4 Scout (17B-16E Instruct) cost?

Llama 4 Scout (17B-16E Instruct) is priced at $0.1 per 1M input tokens and $0.3 per 1M output tokens ($0.15/1M blended at a 3:1 input-to-output ratio).

What is the context window of Llama 4 Scout (17B-16E Instruct)?

Llama 4 Scout (17B-16E Instruct) has a 10.000.000-token context window (10M).

Is Llama 4 Scout (17B-16E Instruct) deprecated?

No — Llama 4 Scout (17B-16E Instruct) is generally available (GA) and not currently scheduled for deprecation or retirement in our tracker.

Does Llama 4 Scout (17B-16E Instruct) support image or vision input?

Yes — Llama 4 Scout (17B-16E Instruct) accepts image input. Listed modalities: text, image.

Can I self-host Llama 4 Scout (17B-16E Instruct)?

Yes — Llama 4 Scout (17B-16E Instruct) is an open-weight model: its weights are publicly downloadable, so you can run it on your own infrastructure (subject to its license) instead of only calling Meta's API.

What is the API model string for Llama 4 Scout (17B-16E Instruct)?

The API model identifier for Llama 4 Scout (17B-16E Instruct) is "meta-llama/Llama-4-Scout-17B-16E-Instruct-FP8" when calling Meta's API.

What is Llama 4 Scout (17B-16E Instruct)'s knowledge cutoff?

Llama 4 Scout (17B-16E Instruct)'s training knowledge cutoff is Aug 2024.