Apps · Calculator
Will it run?
Four choices. One answer — whether an open model fits on your machine, and how fast it will talk back.
Will it run?
Yes.
It fits with room to spare.
Linux default lets the GPU use 96 GB; it can be raised.
- Weights 62.2
- Context (KV cache) 2.4
- Runtime 1.5
- Expected speed
- 33–71tokens/s
- Ceiling 94.5 tokens/s if memory were the only limit.
- Longest context that fits
- 426Ktokens
- Before memory runs out. The model’s own limit may be lower.
How it estimates
Three sums, nothing hidden.
Good enough to choose a download. Not a guarantee — real backends differ by a few gigabytes either way.
Weights
Total parameters × bits per weight ÷ 8. A mixture-of-experts model still has to hold every expert in memory, so total — not active — parameters count here.
Context
The KV cache grows with every token of context. Sizes are class defaults; check the model card and adjust under “Fine-tune” for an exact figure. Runtime adds about 1.5 GB.
Speed
Each token reads the active weights once, so bandwidth ÷ active size is the ceiling. Real runs reach 35–75% of it — the band is calibrated on measured results from the test bench.