What LLM Can I Run? Free GPU & VRAM Model Checker

LLM Model Checker

What LLM can I run on my hardware?

GPU Your graphics card. Auto-detected from your browser, or select manually. RTX 4070
VRAM Video memory available for model weights. Determines which models can run. On Apple Silicon, VRAM is shared with system RAM (unified memory). 12 GB
BW Memory bandwidth in GB/s. Determines inference speed (tokens per second). 504 GB/s
RAM System memory. On Apple Silicon, RAM is unified with VRAM. Model and OS share this pool. 32 GB

Estimates based on browser APIs. Actual specs may vary.

Models you can run

13 models fit your RTX 4070

Model Quality Speed VRAM Context
Qwen3.5-9B Excellent ~71 tok/s 5GB 262K ctx
Qwen3.5-4B Excellent ~176 tok/s 2GB 262K ctx
Gemma 4 12B Good ~59 tok/s 6GB 262K ctx
GPT-oss 20B Weak ~32 tok/s 11GB 128K ctx

See full benchmarks and quality scores for each model.

What LLM Can I Run? Free GPU & VRAM Model Checker | Onyx AI