NVIDIA L4. 24 GB Inference GPU at 72 Watts
Rent the NVIDIA L4: 24 GB GDDR6, 4th-gen Tensor cores with FP8 and dedicated AV1 video engines in a 72 W data center GPU. Built for LLM serving, image generation, computer vision and video transcoding. Deploy in minutes from German data centers.
Why the NVIDIA L4
24 GB VRAM
More memory than the RTX 4000 Ada. Room for 7B-8B LLMs in FP16, 13B-14B models in FP8 and larger KV caches for more concurrent users.
Built for Inference
4th-gen Tensor cores with FP8, BF16 and sparsity. Runs vLLM, TensorRT-LLM, Triton, ComfyUI and all common AI frameworks.
AV1 Video Engines
2 NVENC and 4 NVDEC units with AV1 support. Transcode, analyze and stream many video streams in parallel.
72 W Efficiency
Data center GPU designed for continuous 24/7 inference at a fraction of the power draw of large training GPUs.
Dedicated, Not Shared
The full GPU is yours. No time-slicing, no noisy neighbors, predictable latency.
Hosted in Germany
GDPR-compliant data centers, low EU latency, no US-cloud lock-in.
What's Included
- 1× NVIDIA L4 (24 GB GDDR6)
- AMD EPYC vCPU cores and RAM
- Fast NVMe SSD storage
- 1× IPv4 + /64 IPv6
- Generous traffic, no overage
- Full root access, bring your own AI stack
- Snapshots and backups available
NVIDIA L4 Specs
Rent NVIDIA L4
NVIDIA L4 plans from €179.99/month. Multi-month contracts unlock additional discounts.
Frequently Asked Questions
What is the NVIDIA L4 best for?+
Serving LLMs and embedding models, image generation with Stable Diffusion, computer vision, speech recognition and video workloads such as transcoding and video analytics. It is designed for inference that runs around the clock.
Which LLMs fit into 24 GB?+
7B-8B models such as Llama 3 8B or Mistral 7B run in FP16 with room for the KV cache. 13B-14B models fit in FP8 or INT8, and with 4-bit quantization even models around 30B parameters are possible.
How does it compare to the RTX 4000 Ada?+
Both use the Ada Lovelace architecture. The L4 has 24 GB instead of 20 GB VRAM, dedicated AV1 video engines and a 72 W data center design. The RTX 4000 Ada has higher memory bandwidth and is the better fit for rendering. For inference with larger models or video, choose the L4.
Can I use it for training?+
For fine-tuning smaller models with LoRA or QLoRA, yes. For training larger models, the RTX 6000 Ada with 48 GB VRAM is the better choice.