Skip to main content

For ML APIs

FastAPI Hosting for ML APIs

A FastAPI ML API has a specific shape: long-lived container, model weights cached in memory, persistent volume for the weights so they do not redownload, async streaming for token-by-token responses. Hostim runs this shape natively. No cold starts, no per-request billing, predictable monthly cost.

# docker-compose.yml
services:
  api:
    image: my-fastapi-app
    command: uvicorn main:app --host 0.0.0.0 --port 8000
  db:
    image: postgres:16
850+apps deployed
320+developers building
200+services running now
Oct 2025in production since

Why FastAPI on Hostim for ML APIs

Serverless platforms are a bad fit for ML inference: cold starts reload the model, per-request billing punishes long generation, and the function timeout caps long completions. Hostim runs your FastAPI container as a normal long-lived process. Mount a persistent volume at /models and load the weights once at startup; they survive every redeploy. Async streaming responses (SSE, token-by-token) work because the container stays connected for the duration of the request. CPU and RAM scale per app — GPU support is roadmap.

What ML APIs need from a host

Teams running model inference behind a FastAPI or Flask endpoint, typically with model weights cached on disk.

  • Long-lived containers (no cold starts that reload models)
  • Persistent volumes for model weights, several GB each
  • GPU access (if the model needs it) — or fast CPUs with enough RAM
  • Streaming responses for token-by-token output
  • Predictable cost per inference, not a per-request bill

Hostim runs containers as long-lived processes. Model weights stay in memory. Persistent volumes hold the weights so a redeploy does not redownload 5 GB. CPU and RAM scale per app — GPU support is roadmap.

How Hostim runs FastAPI

FastAPI hosting means running a Uvicorn or Gunicorn-Uvicorn worker process and exposing it over HTTPS. The framework is async by default, so the host has to support long-lived connections — websockets, server-sent events, streaming responses.

Deploy model

Hostim runs your FastAPI Docker image as a normal container. Long-lived connections work. Managed PostgreSQL is attached at runtime. If you serve an ML model, mount a persistent volume for the weights so they do not redownload on every deploy.

Common pitfalls

Cold starts from serverless platforms are not a fit for ML inference workloads. Hostim runs a permanent container, so model weights stay in memory across requests.

Typical env vars

DATABASE_URL, OPENAI_API_KEY, MODEL_PATH, LOG_LEVEL

Questions

FAQ

How do I keep model weights across deploys?

Mount a persistent volume at /models. Download or build weights into the volume once; subsequent deploys mount the same volume — no redownload.

Are streaming token responses supported?

Yes. FastAPI on Hostim runs as a long-lived process behind HTTPS. SSE and token-by-token streaming work without extra config.

Is there GPU support?

Not yet. CPU inference with enough RAM works for many smaller models (embeddings, classical ML). GPU is on the roadmap; ask if you need it for evaluation.

How is per-request billing avoided?

Hostim charges by reserved CPU, RAM and storage — flat per month. Inference cost does not scale with the number of requests; the bill is predictable.

Ready to deploy FastAPI?

Spin up an app in minutes. Managed database on the free tier, custom domain included.

The numbers behind this

Platform figures from the published price list, docs and our own benchmark.

€2.50/month

Hostim entry app plan

Plan sa-1-1: 1 vCPU, 1 GB RAM. Flat price, billed hourly, no usage meter.

€0

Managed database free tier

PostgreSQL 256 MB, MySQL 256 MB, Redis 128 MB, volume 1 GB. No cap on how many free instances you create.

Included on every plan

Database high availability

PostgreSQL and MySQL run as a primary plus hot standby with automatic failover, shared plans included.

2,708 TPS

Managed PostgreSQL write throughput

pgbench, 4 clients, 300 s, plan drp-50 (2 vCPU / 4 GB). AWS RDS db.t4g.medium scored 1,080 TPS on the same test.

€0/GB

Traffic charges

No ingress fees and no per-GB egress line on app plans.

1 today (Falkenstein, Germany)

Regions

Bare metal, EU only, with more regions planned. No AWS, GCP or Azure underneath.

Managed PostgreSQL runs as a replicated cluster with automatic failover on every plan, shared and dedicated alike.

Hostim.dev docs, managed PostgreSQL

Sources

  1. Hostim.dev price listHostim.dev (checked 2026-08-06)
  2. Pricing model — plan-based billing, no meteringHostim.dev docs (checked 2026-08-06)
  3. Managed PostgreSQL — high availability and failoverHostim.dev docs (checked 2026-08-06)
  4. PostgreSQL benchmark: AWS RDS vs Hostim vs self-hosted on HetznerHostim.dev blog, July 2026 (checked 2026-08-06)