LOCAL INFERENCE PLANNING SYSTEM

VRAM Modeler

GB
Accelerate Inference with the Plugable TBT5-AI Learn more →
DEPLOYMENT INPUTS
Context
tokens

DEPLOYMENT TIER 1

High-precision dense syntax

FITS
96 GB VRAM BUDGET
Weights 32.5 GB FP8 estimate
KV cache 16.8 GB 131K tokens, 1 sequence
Runtime reserve 2.0 GB vLLM heuristic
Free headroom 44.7 GB 46.6% of VRAM
SYSTEM PARAMETERS
Target model
Quantization
Architecture
Context configured
Total allocated
DEPLOYMENT METRICS
Maximum context that fits
KV cost per 1K tokens
Model utilization
Recommended use
Confidence
CONTEXT COST CURVE

VRAM required by context length

Weights and runtime reserve stay fixed. KV cache grows linearly for this model profile.

Required Budget
RESEARCH NOTES
Verified recently

Memory values are planning estimates, not measured allocations. Runtime behavior varies by engine, GPU, kernels, tensor parallelism, and model implementation.