Menu
Total Experiments
1
Sweep runs + 1 baseline
Est. Total Runtime
10 mins
Sustained soak & cycling
EST. Total Runtime Formula
Total Experiments × 10 mins
Calculates cumulative runtime across all sweep points assuming standard 10-minute soak per run.
Current: 1 runs × 10m = 10 mins
Est. GPU Hours
1.5 hrs
Characterization load
EST. GPU Hours Formula
Total Experiments × 1.5 GPU Hours
Assuming parallel profiling overhead on an 8-GPU node structure.
Current: 1 runs × 1.5h = 1.5 hrs
Est. KV Cache
640.00 GB
Serving stream requirements
EST. KV Cache Size Formula
2 × Layers × KV Heads × Head Dim × Seq Len × Batch × Concurrency × Precision Bytes
Evaluates active key-value state requirement for model serving streams in Gigabytes.
Current Size: 640.00 GB
Est. Memory Occupancy
100.0%
780.0 GB / 288 GB
EST. Memory Occupancy Formula
(Model Size × Precision Bytes) + KV Cache GB
Predicts total Memory usage against standard 8x 36GB (288GB total) hardware capacity.
Current HBM: 780.0 GB / 288 GB
Est. Bandwidth
19.50 TB/s
Memory read bandwidth
Bandwidth Demand Formula
(Weights GB + KV Cache GB) × Concurrency × 25 tokens/s / 1000
Estimates total concurrent read bandwidth requirement for weight and state updates in TB/s.
Current: 19.50 TB/s

General Information & Templates

Factor Sweep Table

Define parameter sweeps. Values in sweep columns override the baseline reference for that run.

ParameterBaseline (Immutable)Sweep 1Sweep 2Sweep 3Sweep 4Sweep 5RunsActions
Max Active Sequences2560
Context Length81920
Request Rate / Client Concurrency10
Precisionbfloat160
Input tokens10240
Output tokens1280
Number of Requests1000
GPU Memory Utilization0.90
KV Block Size160
Model Weight CPU Offload (GB)00
KV Cache Host Memory (GB)00
KV Cache Dtypeauto0
Memory TechnologyHBM20
Power Cap3000
Tensor Parallelism10
Datasetsharegpt0
Benchmark modethroughput0

LLM-Specific Factors (Built-in Editors)

Ergonomic quick-editors for model-serving parameters. Click a pill to set the Baseline reference, or click the + badge to add to the Sweep Table.

Precision
Baseline: bfloat16
Dataset
Baseline: sharegpt
Benchmark mode
Baseline: throughput
Input tokens
Baseline: 1024
Output tokens
Baseline: 128
Context Length
Baseline: 8192
Number of Requests
Baseline: 100
Request Rate / Client Concurrency
Baseline: 1
Max Active Sequences
Baseline: 256
GPU Memory Utilization
Baseline: 0.9
Model Weight CPU Offload (GB)
Baseline: 0
KV Cache Host Memory (GB)
Baseline: 0
KV Cache Dtype
Baseline: auto
KV Block Size
Baseline: 16
Power Cap
Baseline: 300
Tensor Parallelism
Baseline: 1

Baseline Configuration

Reference Point

Define the baseline parameter settings. These remain immutable during sweep generation, and act as the single reference point.

DOE Statistics

Total Sweep Factors17
Total Sweep Runs0
Baseline Reference Runs1
Est. KV Cache Size 640.00 GB
Est. Memory Occupancy 780.0 GB / 288 GB (100.0%)
Est. Memory Bandwidth 19.50 TB/s
Total Experiments1