EcoHash
Blog
News, engineering deep-dives, and product updates from the EcoHash team.
How to deploy and tune Qwen3.8-27B on one RTX Pro 6000
Qwen3.8-27B is 27B dense in FP8, so one RTX Pro 6000 holds it with room for a large KV cache. The serving config EcoHash runs in production, the five settings whose obvious alternative fails, and the one flag that doubled what a single card serves.
Run your Retell AI agent on an open model: the custom LLM bridge
How to connect a Retell AI voice agent to EcoHash's OpenAI-compatible API with a small WebSocket bridge, what it does to the per-minute LLM bill, and what stays on Retell.
Give your Vapi voice agent a cheaper voice: Kokoro TTS via custom-voice
How to plug EcoHash's Kokoro TTS into a Vapi voice agent as a custom voice: a small adapter server, the assistant config, a local test, and what it costs per minute compared to ElevenLabs.
Build a voice agent with Whisper, Kokoro, and an OpenAI-compatible API
How to build a speech-to-text to LLM to text-to-speech voice agent on EcoHash using Whisper, a chat model, and Kokoro, all through one OpenAI-compatible API key.
Qwen3 Coder 30B on RTX Pro 6000: API, dedicated endpoint, or GPU workspace?
Three ways to run Qwen3 Coder 30B on EcoHash — the shared OpenAI-compatible API, a dedicated endpoint, or an RTX Pro 6000 GPU workspace — and how to choose between them.
Best models to run on RTX Pro 6000 96GB
A task-by-task guide to choosing open models for a 96GB RTX Pro 6000 on EcoHash: coding, general chat, embeddings and rerankers for RAG, and speech models, all through an OpenAI-compatible API.
What makes RTX Pro 6000 96GB a good GPU for AI inference?
How the NVIDIA RTX Pro 6000 Blackwell Server Edition with 96GB VRAM fits open-model inference on EcoHash, which model sizes it suits, how it compares to other GPUs, and how to access it.