Skip to main content

AI Infrastructure

RTX Pro 6000 GPU cloud for open-model inference

Rent NVIDIA RTX Pro 6000 Blackwell GPUs with 96 GB of VRAM. Run hourly workloads — LoRA fine-tuning, 3D rendering & ray tracing — or reserve a long-term dedicated endpoint for any open model, or your own fine-tuned weights.

At a glance

RTX Pro 6000 Server Edition

GPU
RTX Pro 6000 (Blackwell)
VRAM
96 GB
Configurations
1 / 2 / 4 / 8 GPUs
Hourly rate
$1.89 / GPU-hour

Current hourly rate: $1.89/GPU-hour, billed per second. See pricing for the live rate.

Configurations

Sizes, specs and live availability

Every configuration runs on a single machine, so all GPUs share one host. Billed per second at the same GPU-hour rate.

ConfigurationvCPURAMVRAMPrice / hourAvailability
1× RTX Pro 60001664 GB96 GB$1.89Available now
2× RTX Pro 600032128 GB192 GB$3.78Available now
4× RTX Pro 600064256 GB384 GB$7.56Available now
8× RTX Pro 6000128512 GB768 GB$15.12Limited

Availability is live and moves with demand. Last updated . Need a size that shows as on request? Talk to us — we can usually free capacity.

What it's good for

Workloads that fit 96 GB

RTX Pro 6000 is EcoHash's hero GPU for cost-effective open-model inference and GPU workspaces.

20B–35B open models

96 GB of VRAM fits popular 20B–35B models with room for larger batch sizes and KV-cache headroom.

LoRA & QLoRA fine-tuning

Adapter fine-tuning and experimentation on a single Blackwell GPU, with root access and a browser terminal.

Quantized larger models

Validate selected quantized larger models where 96 GB is enough for single-GPU inference.

Batch & data jobs

Batch inference, data preprocessing, and offline evaluation on workstation-class GPUs.

Best-fit models

Popular models on RTX Pro 6000

Start with these open models through the API, then move to a dedicated endpoint or GPU workspace as usage grows.

Voice and lightweight models (such as Kokoro TTS and Whisper STT) are available through the same API and can scale onto EcoHash GPU capacity for high-throughput workloads — they don't require a dedicated RTX Pro 6000.

Product options

Three ways to use it

GPU workspace

Launch an hourly RTX Pro 6000 instance with root access, Jupyter, and a web terminal.

Best for tuning, rendering, and experimentation.

Dedicated endpoint

Reserve RTX Pro 6000 capacity for predictable latency and throughput on your chosen model.

Best for steady production traffic.

Hosted model API

Call open models through one OpenAI-compatible API — no GPU to manage until you need one.

Best for getting started fast.

GPU instances

A GPU workspace with root access

Provision exactly what a job needs and tear it down when you're done — root access, your choice of image, and per-second billing at the GPU-hour rate.

1, 2, 4 or 8 GPUs per instance

Every GPU brings 16 vCPU, 64 GB RAM and 400 GiB of local scratch. Bring any container image, or start from the CUDA, JupyterLab or vLLM images we maintain.

SSH, browser terminal, JupyterLab

Native SSH for scp, rsync and VS Code Remote; a shell in any browser with no key required; JupyterLab wired up automatically for notebook images.

File upload & public service URL

Upload files straight into the pod from the console, and expose any HTTP port inside the container at a public URL.

Auto-expiry protection

Set a duration when you launch and the instance stops on its own — extend it any time. Billed per second, so you never get a surprise bill.

3D rendering & ray tracing

Render in Blender, Octane and more on Blackwell RT cores with 96 GB — workstation-class GPUs that clouds renting H100/A100 don't offer.

Persistent storage

Attach a Cloud Drive or Shared Filesystem in the launch form; it mounts under /workspace and outlives the instance.

GPU clusters

The same instance, replicated

A cluster is N identical replicas of one instance configuration in one region, sharing a filesystem and a single endpoint — for data-parallel training, multi-replica serving or any worker pool that wants horizontal scale.

2–8 identical replicas

One image, one startup command, one GPU configuration — repeated across every replica and managed from a single panel.

One load-balanced endpoint

Every replica sits behind the same HTTPS URL, so a multi-replica model server or worker pool looks like one service to callers.

Shared Filesystem on every replica

Attach a Shared Filesystem at the cluster level and it mounts at the same path on all replicas — one copy of the weights, one copy of the data.

Scale replicas in place

Change the replica count between 1 and 8 while the cluster runs. Per-replica terminals and uploads are available from the console.

Storage

Persistent storage that travels with your work

Keep datasets and model weights close to the GPUs. Storage outlives any single instance and reattaches on demand.

Cloud Drive

Block storage (read-write-once) for a single instance — a durable workspace that survives restarts.

Shared Filesystem

CephFS-backed shared storage (read-write-many) mounted across instances and cluster replicas — ideal for datasets and model weights.

Sizes, mount paths, export and the live rate table are on the Storage page.

Boundaries

What it's not for

RTX Pro 6000 is best for 20B–35B open models, selected quantized larger models, and production inference workloads. Very large models (235B+ / 671B) that need multi-GPU tensor sharding are better served by H200 or a partner path. RTX Pro 6000 is offered as PCIe capacity for single-GPU and multi-instance inference, not as an NVLink training cluster.

Launch an RTX Pro 6000 GPU

Start an hourly workspace, call a model through the API, or reserve a dedicated endpoint.