The Mathematical Foundation of 70B Model Weights
To understand how much VRAM a 70B model requires, we must first look at the raw parameter count. A 70 billion parameter model consists of 70 billion individual weights. In standard half-precision (FP16 or BF16), each parameter occupies 2 bytes of memory. The baseline calculation is straightforward: 70,000,000,000 parameters multiplied by 2 bytes equals 140,000,000,000 bytes, or approximately 130.39 GiB. This is the absolute minimum VRAM required just to load the model weights into memory without accounting for any activations, KV cache, or system overhead.
In a production environment, you cannot provision 140GB of VRAM and expect the model to run. CUDA kernels, library overheads, and the operating system itself consume a portion of the available memory. Furthermore, the model architecture dictates how these weights are distributed. For a 70B model, this usually necessitates a multi-GPU setup: two A100 or H100 80GB cards hold 160GB, which leaves almost no headroom once the safety margin below is applied, so production deployments typically move to four 80GB cards. When using multiple GPUs, communication buffers for technologies like NCCL (NVIDIA Collective Communications Library) also add to the memory footprint. Engineers must account for a safety margin of at least 10 to 15 percent above the raw weight size to ensure stability during inference. Without this buffer, even a small increase in input sequence length can trigger an immediate OOM error, crashing the entire inference service.
Quantization Strategies and Memory Reduction
Quantization is the most effective technique for reducing the VRAM footprint of 70B models. By reducing the precision of the weights from 16-bit to 8-bit or 4-bit, you can significantly lower the entry barrier for hardware. In 8-bit quantization (INT8), each parameter uses 1 byte, bringing the model weight size down to approximately 70GB. This allows a 70B model to fit onto a single 80GB GPU, though with very limited room for context. The real breakthrough for many teams comes with 4-bit quantization techniques like GPTQ, AWQ, or GGUF, all of which are supported in current inference stacks.
At 4-bit precision, each parameter occupies roughly 0.5 bytes. However, due to the need for quantization constants and metadata, the actual footprint is closer to 0.55 to 0.7 bytes per parameter. For a 70B model, 4-bit quantization results in a weight size of approximately 40GB to 48GB, with a common Q4_K_M build of Llama 3.1 70B landing near 42GB. This makes it possible to run a 70B model on a single NVIDIA L40S (48GB) or even a dual RTX 4090 (24GB x 2) setup. While there is a slight degradation in perplexity when moving from FP16 to 4-bit, the trade-off is often worth it for the massive reduction in infrastructure costs. Lyceum offers the L40S (48GB) on its On-demand GPU VM product with per-second billing, so you can match the card to the quantization level of your deployment instead of paying for VRAM you never load.
VRAM Requirements for Full Fine-Tuning
Fine-tuning a 70B model is an entirely different challenge compared to inference. During training, the GPU must store not only the model weights but also the gradients, optimizer states, and forward activations. In mixed-precision training with the Adam optimizer, each parameter requires 6 bytes for the weights (an FP16 copy for the forward and backward pass plus an FP32 master copy), 4 bytes for the FP32 gradient, and 8 bytes for the optimizer states (momentum and variance in FP32). This totals 18 bytes per parameter. For a 70B model, full fine-tuning requires approximately 1.26 terabytes of VRAM for model states alone, before activations.
This level of memory requirement necessitates a massive cluster of GPUs, typically at least 16 to 20 A100 80GB or H100 80GB cards, connected via high-speed interconnects like NVLink. Even with DeepSpeed ZeRO-3 redundancy reduction, which partitions the optimizer states, gradients, and parameters across all data-parallel processes, the aggregate VRAM needed remains the same. For most scaleups and mid-market companies, full fine-tuning is economically unfeasible without specialized orchestration. Lyceum addresses this with GPU compute in European data centres in Paris and Finland, providing the high-performance compute needed for these workloads.
PEFT and QLoRA: Fine-Tuning on a Budget
Parameter-Efficient Fine-Tuning (PEFT) techniques, specifically LoRA (Low-Rank Adaptation) and QLoRA, have made 70B models accessible to fine-tune on a single GPU. LoRA works by freezing the original model weights and only training a small number of adapter weights. This drastically reduces the memory needed for gradients and optimizer states. QLoRA takes this a step further by quantizing the base model to 4-bit (using a special NormalFloat4 data type) and using paged optimizers to handle memory spikes.
With QLoRA, the VRAM requirement for the base 70B model drops to roughly 44GB, scaled from the 41GB checkpoint size the QLoRA paper reports for a 65B model, and the paper demonstrates a 65B model finetuning in under 48GB on a single GPU, down from more than 780GB for 16-bit finetuning. The additional memory needed for the adapters and activations depends on the rank (r) and alpha settings, but it typically adds only a few gigabytes. This allows a 70B model to be fine-tuned on a single 80GB GPU or a small cluster of 48GB GPUs. This is a practical shift for AI teams that need to specialize a large model on proprietary data while keeping costs under control. When running these jobs, Lyceum's Serverless Training is priced per GPU-hour and, like all its GPU compute, billed per second, so you can size the instance to the job instead of provisioning headroom for an activation spike that may never come.
Now that you know the VRAM, see what it costs. Use the GPU Pricing Calculator to compare costs across RunPod, Lambda, AWS, GCP, CoreWeave, and Lyceum.
Multi-GPU Orchestration: Sharding and Parallelism
When a model exceeds the VRAM of a single GPU, you must employ parallelism strategies. The two most common are Pipeline Parallelism (PP) and Tensor Parallelism (TP). Pipeline Parallelism splits the model layers across different GPUs. For example, in a 2-GPU setup, layers 1-40 might live on GPU 0, and layers 41-80 on GPU 1. While this is simple to implement, it can lead to GPU idle time (bubbles) as one GPU waits for the other to finish its computation. Tensor Parallelism is more complex but more efficient within a single node, as it splits individual weight matrices across GPUs, allowing them to work on the same layer simultaneously.
For a 70B model in FP16, you would typically use a 2-way or 4-way Tensor Parallelism setup. This requires high-bandwidth interconnects like NVLink to minimize the communication overhead between GPUs. If you are using a cloud provider without high-speed interconnects, the latency of moving data over PCIe can negate the benefits of multi-GPU scaling. Lyceum's On-demand GPU VMs cover these high-performance workloads, with H100 and A100 nodes billed per second and no base fee. This lets ML engineers size the tensor-parallel deployment to hardware they control, rather than fighting CUDA stream synchronization and NCCL collective operations on an interconnect that cannot keep up.
Predicting Memory Footprint Before Deployment
One of the most common frustrations for AI teams is the 'trial and error' approach to provisioning GPUs. You launch a job, wait for it to initialize, and then watch it crash with a 'CUDA Out of Memory' error five minutes later. This waste of time and resources is exactly what Lyceum's per-second GPU billing is built to reduce: you can validate a workload on a single card for a few minutes instead of committing to an oversized cluster for the day. This is particularly important for 70B models where the difference between a 40GB and 80GB GPU is significant in terms of cost.
Predicting memory usage involves analyzing the model's computational graph and accounting for the specific framework overheads (PyTorch vs. JAX). For instance, PyTorch's caching allocator might hold onto memory that is technically free, leading to reported usage that is higher than actual requirements. Lyceum's Pythia tool does this for PyTorch workloads: it predicts the memory requirement and runtime on each GPU option and recommends a card, currently choosing between T4, A100 and H100. This workload-aware approach means that if your 70B model only needs part of an A100's VRAM due to efficient quantization, you can select smaller or cheaper hardware rather than paying for capacity you do not use.
Data Residency and Jurisdiction for Large Models
For European companies, the decision of where to host a 70B model is not just a technical one, but a legal one. Large language models are often fine-tuned on sensitive internal data, including PII or proprietary intellectual property. Using US-based hyperscalers can introduce compliance risks under GDPR, especially when data is processed in regions without equivalent privacy protections. Lyceum runs its GPU compute in European data centres in Paris and Finland. Choosing one of those regions lets you name the jurisdiction your hardware sits in, which is the first question a data-protection review asks.
Data transfer is a cost that is easy to miss with large models. When working with 70B models, moving model checkpoints (which can be 140GB each) or large datasets between regions can result in thousands of euros in unexpected charges on platforms like AWS or GCP. Read each provider's data-transfer terms before planning a workflow that moves checkpoints between regions.