GPU Memory Calculator
Estimate VRAM per GPU before you launch the job.
LLMem · arXiv:2404.10933
0
Inference
Fine-tuning
Model
Preset
Paste a HuggingFace config.json
or upload the file
Parallelism
TP
PP
DP
EP (expert-parallel)
Inference
Weight dtype
KV cache dtype
Sequence length
Batch size
Logits mode
Last token (decode)
Full sequence (prefill scoring)
Fine-tuning
Method
Full fine-tuning
LoRA
QLoRA
Fidelity
Simple (fast, analytic)
Granular (LLMem-faithful)
ZeRO stage
0 (none)
1 (optimizer)
2 (+ gradients)
3 (+ params, ADP)
Optimizer
AdamW
AdamW 8-bit
SGD (momentum)
Sequence length
Batch size
Gradient checkpointing
Mixed precision
LoRA rank
Target modules (comma-separated)
GPU
Device
Usable memory (gpu_memory_utilization)
%
Breakdown
Component
Total
Per-GPU (peak)