VVRAM Lens

LOCAL INFERENCE / MEMORY PLANNING

Will your model fit?

A transparent VRAM estimate for quantized weights, KV cache, runtime overhead and breathing room. No benchmarks. No uploads.

Everything runs in your browser
01

Your setup

Hardware
Model

For MoE, enter total parameters—not active parameters. All experts occupy weight memory.

Runtime

Block costs are lower bounds, not mixed Q_K_M files. For Q4_K_M / Q5_K_M use measured custom bpw. No universal 15% margin guarantees coverage.

Context includes generated tokens. INT8 KV excludes scales and padding.

02

Estimate

Block costs are lower bounds, not mixed Q_K_M files. Use measured custom bpw for mixed files.

Estimated required VRAM range

Planning estimate, not a guarantee. Actual usage varies by runtime, kernels, allocator and GPU.

Planning band, not a confidence interval.

Where it goes
Quantized weights
KV cache
Runtime overheadYour planning allowance
With headroom
KV cache = layers × 2 (K+V) × KV heads × head dimension × context × batch × bytes per KV value.
03

Method, in plain English

Weights use total parameters and the effective bits-per-weight of the selected format, using single-block costs, not Q_K_M mixed recipes. Use measured effective bpw for mixed files. KV cache grows linearly with context and concurrent sequences. GQA uses the model's KV-head count.

MLA and other architectures whose KV layout is not represented here are not supported. This tool does not benchmark speed or predict quality.

05

Privacy

Calculations stay in this page. No analytics or model uploads. Share query settings are sent to the static host when opened and may appear in logs. Do not put secrets in URLs.