| Host | Dense BF16 Peak | HBM Capacity | HBM Bandwidth | Storage | Hourly price |
|---|---|---|---|---|---|
| 8 × v6e | 7.344 PFLOP/s | 256 GB | 13.104 TB/s | 2000 GB Hyperdisk Balanced | $21.60 |
| 8 × H100 | 7.916 PFLOP/s | 640 GB | 26.80 TB/s | 1280 GB local SSD | $31.20 |
Your prompt goes to both inference hosts over persistent HTTPS connections. Reasoning effort is Low for all models. TTFT and token rate are measured on EC2 as reasoning and answer text arrive; the browser displays those measurements and streams both replies. Your browser’s network and rendering time are excluded. Different storage configurations can contribute to startup-time differences; preparation also includes engine initialization and warmups.
Your description goes to both inference hosts to generate one 1024 × 1024 image per host. Generation time is measured by the web server through receipt of the complete image response, including network transfer. Images appear in the conversation panes when ready. Image size is fixed.
RTT means round-trip network latency. Labels are approximate measurements from this web server on 24 Sep 2026 (median of five TCP connection probes), not live readings or one-way delays. Network transit is included in TTFT, so these results differ from host-local benchmarks.
Ending demo
Model selection will return when both hosts confirm shutdown.
QYJOHN · 8 × v6e (TPU)
- Starting
- Loading
- Warming up
8 × H100 (GPU)
- Starting
- Loading
- Warming up