Apps · Calculator

Will it run?

Four choices. One answer — whether an open model fits on your machine, and how fast it will talk back.

Context
Fine-tune the numbers
B
B
GB
GB
GB/s

Will it run?

Yes.

It fits with room to spare.

Linux default lets the GPU use 96 GB; it can be raised.

66.1 GB needed96 GB available
  • Weights 62.2
  • Context (KV cache) 2.4
  • Runtime 1.5
Expected speed
33–71tokens/s
Ceiling 94.5 tokens/s if memory were the only limit.
Longest context that fits
426Ktokens
Before memory runs out. The model’s own limit may be lower.

How it estimates

Three sums, nothing hidden.

Good enough to choose a download. Not a guarantee — real backends differ by a few gigabytes either way.

Weights

Total parameters × bits per weight ÷ 8. A mixture-of-experts model still has to hold every expert in memory, so total — not active — parameters count here.

Context

The KV cache grows with every token of context. Sizes are class defaults; check the model card and adjust under “Fine-tune” for an exact figure. Runtime adds about 1.5 GB.

Speed

Each token reads the active weights once, so bandwidth ÷ active size is the ceiling. Real runs reach 35–75% of it — the band is calibrated on measured results from the test bench.