Product Calculator Compare Developers Sign In Request Access
1-Click Model Deployment

Same models. Same GPUs. Up to 85% less per hour.

Pick a model from our curated catalog and deploy it in one click on spot GPUs arbitraged across AWS, GCP and RunPod — the guaranteed <30s preemption recovery SLA keeps endpoints serving even on spot capacity.

1-click
Deploys from a curated model catalog
<30s
Recovery SLA — benchmarked in CI
3
Clouds arbitraged — AWS · GCP · RunPod
730h
Per-month always-on pricing basis
GPU Rates

The same silicon, a fraction of the hourly rate.

Per-GPU hourly list rates, July 2026. AetherCompute rates are live spot-arbitrage estimates refreshed hourly.

On AetherCompute What you'd pay elsewhere
GPU VRAM AetherCompute HF Inference Endpoints Replicate You save vs HF
RTX 4090 24 GB $0.22
L40S 48 GB $0.35 $1.80 $3.51 −81%
A100 SXM 80 GB $0.65 $2.50 $5.04 −74%
H100 SXM 80 GB $0.95 $4.50 $5.49 −79%
B200 192 GB $1.85 $9.25 −80%

RunPod and Modal H100 list rates (~$2.69 and $3.95/hr respectively) sit between HF and AetherCompute. Per-second serverless platforms only win at low utilization — see methodology.

Model Hosting Costs

What 1-click deployment saves on the models you already use.

The curated launch catalog on always-on (730 hrs/mo) serving configs. GPU count × type shown per deployment.

On AetherCompute What you'd pay elsewhere
Model Type AetherCompute config AetherCompute $/mo HF Endpoints $/mo Replicate $/mo Savings
Llama 3.1 8B Instructmeta-llama/Llama-3.1-8B-Instruct LLM 1× RTX 4090 · FP16 $161 $584 $2,562 −72%
Mistral Small 3 24Bmistralai/Mistral-Small-3-24B-Instruct LLM 1× L40S · INT8 $256 $1,314 $2,562 −81%
Gemma 3 27B ITgoogle/gemma-3-27b-it LLM 1× L40S · INT8 $256 $1,314 $2,562 −81%
DeepSeek-R1 Distill 32Bdeepseek-ai/DeepSeek-R1-Distill-Qwen-32B LLM 1× A100 · FP16 $475 $1,825 $3,679 −74%
Mixtral 8×7B Instructmistralai/Mixtral-8x7B-Instruct-v0.1 LLM 2× L40S · FP16 $511 $2,628 $5,124* −81%
Llama 3.3 70B Instructmeta-llama/Llama-3.3-70B-Instruct LLM 2× A100 · FP8 $949 $3,650 $7,358 −74%
Qwen2.5 72B InstructQwen/Qwen2.5-72B-Instruct LLM 2× A100 · FP8 $949 $3,650 $7,358 −74%
Llama 3.1 405B Instructmeta-llama/Llama-3.1-405B-Instruct LLM 4× H100 · FP8 $2,774 $13,140 contract-only −79%
FLUX.1-devblack-forest-labs/FLUX.1-dev Vision 1× RTX 4090 · FP8 $161 $1,314 $2,562 −88%
SDXL 1.0stabilityai/stable-diffusion-xl-base-1.0 Vision 1× RTX 4090 $161 $584 $2,562 −72%
Whisper Large v3openai/whisper-large-v3 Audio 1× RTX 4090 $161 $584 $2,562 −72%

*Mixtral Replicate: 2× L40S × $3.51 × 730 = $5,124. Savings % are vs HF Inference Endpoints. Monthly = hourly × 730, 24/7 always-on.

How It Works

From model pick to live endpoint in three steps.

No Dockerfiles, no GPU hunting, no spot anxiety.

Pick a model from the catalog

A curated set of production-ready open models. Weights hydrate instantly over the FUSE layer — no multi-GB stage-in wait.

We pick the cheapest GPU config

Hourly arbitrage across AWS Spot, GCP preemptible and RunPod places the model on the cheapest compliant $/GPU-hr, with GPU bin-packing for multi-tenant efficiency.

Spot-proof serving

If a node is preempted, checkpoint + failover migration restores the endpoint in under 30 seconds — guaranteed by a CI-benchmarked SLA.

Methodology & assumptions

  • Competitor figures are public list prices as of July 2026 (sources: huggingface.co/pricing, replicate.com/pricing); HF Inference Endpoints AWS region rates; multi-GPU HF configs scale linearly from published 1×/4×/8× rates.
  • Replicate per-second rates converted to hourly equivalents assuming 100% utilization — bursty low-traffic workloads can cost less on per-second billing.
  • AetherCompute rates are illustrative spot-arbitrage rates that vary by region and time.
  • Monthly basis: 730 hours always-on, single replica, vLLM serving with the quantization shown. Models fit the stated GPU config in the stated precision; configs are illustrative minimums.
Illustrative estimates — verify current list prices before purchase decisions.

Deploy your first model in one click.

Pick a model, leave with an OpenAI-compatible endpoint.

Request Access