Newsletter
Join the Community
Subscribe to our newsletter for the latest news and updates
Low-cost, scalable inference for high-volume AI workloads
Doubleword is a low-cost, scalable inference platform for high-volume AI workloads. It runs open-weight models on its own GPUs, offering async and batch inference up to 90% cheaper than real-time APIs. It is OpenAI-compatible and designed for agents, evals, and pipelines.
https://api.doubleword.ai/v1 and send a request.service_tier: "flex"./batch endpoint with a JSONL file, upload via dashboard, CLI, or Autobatcher.Doubleword uses simple pay-per-token pricing with prepaid credits and no minimum. Prices drop with latency flexibility: batch is cheaper than async, and async is cheaper than realtime. For example, Kimi-K3 costs $11.25 per 1M output tokens async and $7.50 batch. GLM-5.3-Flash costs $0.38 async and $0.25 batch. Prompt caching can reduce input costs further.
How do I get started? Sign up, create an API key, choose a model, and point the OpenAI SDK at the API endpoint.
Is Doubleword a router? No. Doubleword owns and runs the GPUs, ensuring consistent model builds and pricing.
Which models are supported? Leading open-source models including DeepSeek, GLM, Kimi, Qwen, and more.
Are open-source models reliable? Yes, they show task-level parity with closed-source APIs on intelligence indices.
What are the rate limits? New accounts have initial limits; adding a payment method lifts them. Async tier offers higher limits.
Can I download partial batch results? Yes, you can download completed results while a batch is running.
What happens if a batch fails? Doubleword retries and recovers automatically, saving partial results.
Is data secure? Data is encrypted in transit and at rest, with opt-in Zero Data Retention and no training on data by default.