Download canonical Markdown for Study
Price facts
Prices checked 2026-08-29.
| Model | Input / 1M | Output / 1M | Cached input / 1M |
|---|---|---|---|
| GPT-5.6 Luna | $0.20 | $1.20 | $0.02 |
| GPT-5.6 Terra | $2.00 | $12.00 | $0.20 |
| GPT-5.6 Sol | $4.00 | $20.00 | $0.40 |
| Claude Sonnet 5 | $2.00 | $10.00 | — |
| Claude Opus 5 | $5.00 | $25.00 | — |
| DeepSeek V4 Flash off-peak | $0.22 | $0.66 | $0.007 |
What is measured, estimated, and unknown
| Class | Claim |
|---|---|
| Measurement | ML.ENERGY records roughly 0.39 J per output token for one highly optimized 70B FP8 H100 configuration; it is an aggregate benchmark measure. |
| Estimate | Joule query energy figures are bottom-up estimates, not a production electricity meter for a provider or closed model. |
| Unknown | The true production J/token and unit economics of closed models are not public facts. |
The short version
Electricity becomes computation; computation reads input tokens and produces output tokens; those tokens can become code, answers, products and action. A kWh can correspond to hundreds of thousands through millions of output tokens, depending on model, context length, batching and reasoning. That is a useful scale, not a promise about any particular model.
Cache is not magic
Provider prompt caching and client context minimisation are different. A provider cache needs a byte-stable prefix. Anthropic documents cache reads at 0.1x input pricing; that arithmetic is 90% lower than an uncached read, but it is not a blanket savings promise. Client-side context packs should be addressed by hash, send only deltas, invalidate on change, and never share secrets across agents.
From tokens to products
Token counts are engineering estimates, not quotations. The useful question is not only dollars per token, but dollars per accepted task, rework, and solved problem. Keep uncertainty visible, review volatile prices regularly, and prefer current primary sources.