Method, in plain English
Weights use total parameters and the effective bits-per-weight of the selected format, using single-block costs, not Q_K_M mixed recipes. Use measured effective bpw for mixed files. KV cache grows linearly with context and concurrent sequences. GQA uses the model's KV-head count.
MLA and other architectures whose KV layout is not represented here are not supported. This tool does not benchmark speed or predict quality.