DOE Designer
Design factor sweep experiments for memory bandwidth & reliability characterization.
Total Experiments
1
Sweep runs + 1 baseline
Est. Total Runtime
10 mins
Sustained soak & cycling
EST. Total Runtime Formula
Total Experiments × 10 mins
Calculates cumulative runtime across all sweep points assuming standard 10-minute soak per run.
Current: 1 runs × 10m = 10 mins
Est. GPU Hours
1.5 hrs
Characterization load
EST. GPU Hours Formula
Total Experiments × 1.5 GPU Hours
Assuming parallel profiling overhead on an 8-GPU node structure.
Current: 1 runs × 1.5h = 1.5 hrs
Est. KV Cache
640.00 GB
Serving stream requirements
EST. KV Cache Size Formula
2 × Layers × KV Heads × Head Dim × Seq Len × Batch × Concurrency × Precision Bytes
Evaluates active key-value state requirement for model serving streams in Gigabytes.
Current Size: 640.00 GB
Est. Memory Occupancy
100.0%
780.0 GB / 288 GB
EST. Memory Occupancy Formula
(Model Size × Precision Bytes) + KV Cache GB
Predicts total Memory usage against standard 8x 36GB (288GB total) hardware capacity.
Current HBM: 780.0 GB / 288 GB
Est. Bandwidth
19.50 TB/s
Memory read bandwidth
Bandwidth Demand Formula
(Weights GB + KV Cache GB) × Concurrency × 25 tokens/s / 1000
Estimates total concurrent read bandwidth requirement for weight and state updates in TB/s.
Current: 19.50 TB/s
General Information & Templates
Factor Sweep Table
Define parameter sweeps. Values in sweep columns override the baseline reference for that run.
| Parameter | Baseline (Immutable) | Sweep 1 | Sweep 2 | Sweep 3 | Sweep 4 | Sweep 5 | Runs | Actions |
|---|---|---|---|---|---|---|---|---|
| Max Active Sequences | 256 | 0 | ||||||
| Context Length | 8192 | 0 | ||||||
| Request Rate / Client Concurrency | 1 | 0 | ||||||
| Precision | bfloat16 | 0 | ||||||
| Input tokens | 1024 | 0 | ||||||
| Output tokens | 128 | 0 | ||||||
| Number of Requests | 100 | 0 | ||||||
| GPU Memory Utilization | 0.9 | 0 | ||||||
| KV Block Size | 16 | 0 | ||||||
| Model Weight CPU Offload (GB) | 0 | 0 | ||||||
| KV Cache Host Memory (GB) | 0 | 0 | ||||||
| KV Cache Dtype | auto | 0 | ||||||
| Memory Technology | HBM2 | 0 | ||||||
| Power Cap | 300 | 0 | ||||||
| Tensor Parallelism | 1 | 0 | ||||||
| Dataset | sharegpt | 0 | ||||||
| Benchmark mode | throughput | 0 |
LLM-Specific Factors (Built-in Editors)
Ergonomic quick-editors for model-serving parameters. Click a pill to set the Baseline reference, or click the + badge to add to the Sweep Table.
Precision
Baseline: bfloat16
Dataset
Baseline: sharegpt
Benchmark mode
Baseline: throughput
Input tokens
Baseline: 1024
Output tokens
Baseline: 128
Context Length
Baseline: 8192
Number of Requests
Baseline: 100
Request Rate / Client Concurrency
Baseline: 1
Max Active Sequences
Baseline: 256
GPU Memory Utilization
Baseline: 0.9
Model Weight CPU Offload (GB)
Baseline: 0
KV Cache Host Memory (GB)
Baseline: 0
KV Cache Dtype
Baseline: auto
KV Block Size
Baseline: 16
Power Cap
Baseline: 300
Tensor Parallelism
Baseline: 1
Baseline Configuration
Reference PointDefine the baseline parameter settings. These remain immutable during sweep generation, and act as the single reference point.
DOE Statistics
Total Sweep Factors17
Total Sweep Runs0
Baseline Reference Runs1
Est. KV Cache Size 640.00 GB
Est. Memory Occupancy 780.0 GB / 288 GB (100.0%)
Est. Memory Bandwidth 19.50 TB/s
Total Experiments1
