Token cost calculation
Info
The GPU rates and per-token prices below are the figures used to derive the the given examples here. The authoritative and up-to-date prices for each model are published in our current pricing list.
What is our cost?
Everything you spend is billed token-exact: the input and output tokens of each request are counted and priced according to the model that served it.
| External Models | Internal Models | |
|---|---|---|
| Where the model runs | on the cloud provider’s own servers, for example OpenAI and Anthropic models reached through Microsoft Azure | on our own HPC clusters in Göttingen |
| What you are charged | the standard token rates the provider publishes for that model at the time of each request | a rate we calculate ourselves from what it costs to run that model on our clusters, as broken down below |
What goes into internal models’ price
| Cost factor | What it includes |
|---|---|
| Infrastructure | hardware investment, maintenance, energy costs including cooling |
| Staff | developers, system administrators and support; roughly one full-time employee per 400,000 users, estimated at about 15 % of the hardware investment |
| Overheads | GWDG’s general overheads |
| Idle time | GPUs are not busy every hour of the day, and the hours they sit idle have to be paid for by the hours they are working |
The base unit: one GPU-second
All of the factors above are rolled into a single figure: what one second of one GPU costs. Costs are calculated in one-second intervals.
| GPU | Cost per second |
|---|---|
| A100 | 0.0005991 € |
| H100 | 0.0007408 € |
| H200 | 0.0008847 € |
| B200 | 0.0013898 € |
Newer and faster GPUs cost more per second (a B200 is roughly 2.3 times an A100), and a model served on one is priced accordingly.
From GPU-seconds to token prices
For every model we define which GPU type it runs on, how many GPUs it uses, and how many requests it serves in parallel. Performance tests then measure, for that model, the time per input token and the time per output token. The price of a token is those two numbers multiplied:
Worked through for one model:
| Input parameter | Value |
|---|---|
| Model | qwen3.6-35b-a3b |
| GPU | 1 × H100 95 GB, at 0.0007408 € per second |
| Concurrency | 16 requests served at once |
| Time per input token | 0.000162604 s |
| Time per output token | 0.001648210 s |
| Result | Calculation | Price |
|---|---|---|
| 1 million input tokens | 0.0007408 × 1,000,000 × 0.000162604 | 0.11 € |
| 1 million output tokens | 0.0007408 × 1,000,000 × 0.001648210 | 1.11 € |
Concurrency is what makes this affordable. The GPU serves 16 requests at the same time, so the cost of each second is effectively divided between them. That is already contained in the measured per-token times.
Why output tokens cost about ten times more than input tokens
In the example above, output tokens cost roughly ten times what input tokens cost: 1.11 € against 0.11 € per million. That is not a pricing decision but a property of how the models work: input tokens are processed in parallel in a single pass, while output tokens have to be generated one after another, each one requiring a full pass through the model. The same ratio appears in the commercial providers’ price lists.
The practical consequence for you: a long prompt is cheap, a long answer is not. Capping the output length of your requests is the single most effective way to reduce what you spend.
External models
Commercial models (OpenAI, Anthropic and others) are hosted by Microsoft on Azure rather than in Göttingen, so their cost is the provider’s token price rather than a GWDG calculation. They are billed on token usage and draw on the external models budgets. Because they usualy are more expensive per token compared to open-weight-models, they are the models a cost cap most often applies to.