At a glance
RTX Pro 6000 Server Edition
- GPU
- RTX Pro 6000 (Blackwell)
- VRAM
- 96 GB
- Configurations
- 1 / 2 / 4 / 8 GPUs
- Hourly rate
- $1.89 / GPU-hour
Current hourly rate: $1.89/GPU-hour, billed per second. See pricing for the live rate.
Configurations
Sizes, specs and live availability
Every configuration runs on a single machine, so all GPUs share one host. Billed per second at the same GPU-hour rate.
| Configuration | vCPU | RAM | VRAM | Price / hour | Availability |
|---|---|---|---|---|---|
| 1× RTX Pro 6000 | 16 | 64 GB | 96 GB | $1.89 | Available now |
| 2× RTX Pro 6000 | 32 | 128 GB | 192 GB | $3.78 | Available now |
| 4× RTX Pro 6000 | 64 | 256 GB | 384 GB | $7.56 | Available now |
| 8× RTX Pro 6000 | 128 | 512 GB | 768 GB | $15.12 | Limited |
Availability is live and moves with demand. Last updated . Need a size that shows as on request? Talk to us — we can usually free capacity.
What it's good for
Workloads that fit 96 GB
RTX Pro 6000 is EcoHash's hero GPU for cost-effective open-model inference and GPU workspaces.
20B–35B open models
96 GB of VRAM fits popular 20B–35B models with room for larger batch sizes and KV-cache headroom.
LoRA & QLoRA fine-tuning
Adapter fine-tuning and experimentation on a single Blackwell GPU, with root access and a browser terminal.
Quantized larger models
Validate selected quantized larger models where 96 GB is enough for single-GPU inference.
Batch & data jobs
Batch inference, data preprocessing, and offline evaluation on workstation-class GPUs.
Best-fit models
Popular models on RTX Pro 6000
Start with these open models through the API, then move to a dedicated endpoint or GPU workspace as usage grows.
Voice and lightweight models (such as Kokoro TTS and Whisper STT) are available through the same API and can scale onto EcoHash GPU capacity for high-throughput workloads — they don't require a dedicated RTX Pro 6000.
Product options
Three ways to use it
GPU workspace
Launch an hourly RTX Pro 6000 instance with root access, Jupyter, and a web terminal.
Best for tuning, rendering, and experimentation.
Dedicated endpoint
Reserve RTX Pro 6000 capacity for predictable latency and throughput on your chosen model.
Best for steady production traffic.
Hosted model API
Call open models through one OpenAI-compatible API — no GPU to manage until you need one.
Best for getting started fast.
GPU instances
A GPU workspace with root access
Provision exactly what a job needs and tear it down when you're done — root access, your choice of image, and per-second billing at the GPU-hour rate.
1, 2, 4 or 8 GPUs per instance
Every GPU brings 16 vCPU, 64 GB RAM and 400 GiB of local scratch. Bring any container image, or start from the CUDA, JupyterLab or vLLM images we maintain.
SSH, browser terminal, JupyterLab
Native SSH for scp, rsync and VS Code Remote; a shell in any browser with no key required; JupyterLab wired up automatically for notebook images.
File upload & public service URL
Upload files straight into the pod from the console, and expose any HTTP port inside the container at a public URL.
Auto-expiry protection
Set a duration when you launch and the instance stops on its own — extend it any time. Billed per second, so you never get a surprise bill.
3D rendering & ray tracing
Render in Blender, Octane and more on Blackwell RT cores with 96 GB — workstation-class GPUs that clouds renting H100/A100 don't offer.
Persistent storage
Attach a Cloud Drive or Shared Filesystem in the launch form; it mounts under /workspace and outlives the instance.
GPU clusters
The same instance, replicated
A cluster is N identical replicas of one instance configuration in one region, sharing a filesystem and a single endpoint — for data-parallel training, multi-replica serving or any worker pool that wants horizontal scale.
2–8 identical replicas
One image, one startup command, one GPU configuration — repeated across every replica and managed from a single panel.
One load-balanced endpoint
Every replica sits behind the same HTTPS URL, so a multi-replica model server or worker pool looks like one service to callers.
Shared Filesystem on every replica
Attach a Shared Filesystem at the cluster level and it mounts at the same path on all replicas — one copy of the weights, one copy of the data.
Scale replicas in place
Change the replica count between 1 and 8 while the cluster runs. Per-replica terminals and uploads are available from the console.
Storage
Persistent storage that travels with your work
Keep datasets and model weights close to the GPUs. Storage outlives any single instance and reattaches on demand.
Cloud Drive
Block storage (read-write-once) for a single instance — a durable workspace that survives restarts.
Shared Filesystem
CephFS-backed shared storage (read-write-many) mounted across instances and cluster replicas — ideal for datasets and model weights.
Sizes, mount paths, export and the live rate table are on the Storage page.
Boundaries
What it's not for
RTX Pro 6000 is best for 20B–35B open models, selected quantized larger models, and production inference workloads. Very large models (235B+ / 671B) that need multi-GPU tensor sharding are better served by H200 or a partner path. RTX Pro 6000 is offered as PCIe capacity for single-GPU and multi-instance inference, not as an NVLink training cluster.