GitHub ↗
Inference

Inference & Model Serving

Run open models as an API — throughput, latency, and cost per token.

6 products · last updated 2026-08-30

Baseten hero section
Baseten Inference
Inference is everything
The fastest model runtimes, cross-cloud high availability, and seamless developer workflows. Powered by the Baseten Inference Stack.
Updated 2026-08-29 Visit ↗
DeepInfra hero section
DeepInfra Inference
LOW-COST AI Inference
Accelerate your AI with developer-friendly APIs designed for performance and cost-efficiency.
Updated 2026-08-29 Visit ↗
Fireworks AI hero section
Fireworks AI Inference
Own your model. Own your future.
Fireworks' SoTA training and inference take you beyond the frontier, transforming open models into your specialized intelligence.
Updated 2026-08-30 Visit ↗
Modal hero section
Modal Inference
AI infrastructure that developers love
Run inference, training, batch processing, and sandboxes with sub-second cold starts, instant autoscaling, and a developer experience that feels local.
Updated 2026-08-29 Visit ↗
vLLM by Inferact hero section
vLLM by Inferact Inference
The High-Throughput and Memory-Efficient inference and serving engine for LLMs
Easy, fast, and cost-efficient LLM serving for everyone.
Updated 2026-08-30 Visit ↗