FastAPI has revolutionized Python API development with its incredible performance, automatic documentation, and modern Python features. Used by Microsoft, Uber, and Netflix, FastAPI combines the simplicity of Flask with the performance of NodeJS. Deploying FastAPI on Ubuntu servers provides complete infrastructure control while leveraging async capabilities. This comprehensive guide shows you how to deploy, scale, and optimize FastAPI applications on Ubuntu servers for production in 2025.

Note: This guide focuses on deploying FastAPI on Ubuntu servers. CloudPloy currently supports Laravel applications, with FastAPI support coming soon. Stay tuned for updates!

Why FastAPI Changes the Deployment Game

FastAPI isn’t just another Python framework - it’s a paradigm shift in how we build and deploy APIs:

  • 3x faster than Flask: Comparable to NodeJS and Go
  • Automatic API documentation: OpenAPI and JSON Schema generation
  • Native async support: True asynchronous request handling
  • Type hints everywhere: Catch errors before runtime
  • WebSocket support: Real-time communication built-in
  • Standards-based: OpenAPI, JSON Schema, OAuth2

FastAPI Production Architecture

Modern FastAPI Stack

# production_app.py
from fastapi import FastAPI, Depends, HTTPException
from fastapi.middleware.cors import CORSMiddleware
from fastapi.middleware.trustedhost import TrustedHostMiddleware
from fastapi.middleware.gzip import GZipMiddleware
from contextlib import asynccontextmanager
import asyncpg
import redis.asyncio as redis
from prometheus_fastapi_instrumentator import Instrumentator

# Lifecycle management
@asynccontextmanager
async def lifespan(app: FastAPI):
    # Startup
    app.state.db = await asyncpg.create_pool(
        "postgresql://user:pass@localhost/db",
        min_size=10,
        max_size=20,
        command_timeout=60
    )
    app.state.redis = await redis.from_url(
        "redis://localhost",
        encoding="utf-8",
        decode_responses=True
    )
    yield
    # Shutdown
    await app.state.db.close()
    await app.state.redis.close()

# Create app with lifespan
app = FastAPI(
    title="Production API",
    version="2.0.0",
    lifespan=lifespan,
    docs_url="/docs",
    redoc_url="/redoc"
)

# Middleware stack
app.add_middleware(
    CORSMiddleware,
    allow_origins=["https://example.com"],
    allow_credentials=True,
    allow_methods=["*"],
    allow_headers=["*"],
)

app.add_middleware(TrustedHostMiddleware, allowed_hosts=["example.com", "*.example.com"])
app.add_middleware(GZipMiddleware, minimum_size=1000)

# Prometheus metrics
Instrumentator().instrument(app).expose(app)

# Health checks
@app.get("/health")
async def health_check():
    try:
        # Check database
        async with app.state.db.acquire() as conn:
            await conn.fetchval("SELECT 1")
        
        # Check Redis
        await app.state.redis.ping()
        
        return {"status": "healthy", "database": "up", "cache": "up"}
    except Exception as e:
        raise HTTPException(status_code=503, detail=str(e))

High-Performance ASGI Server Configuration

# gunicorn.conf.py
import multiprocessing
import os

# Gunicorn with Uvicorn workers
bind = "0.0.0.0:8000"
workers = multiprocessing.cpu_count() * 2 + 1
worker_class = "uvicorn.workers.UvicornWorker"
worker_connections = 1000
keepalive = 5
max_requests = 1000
max_requests_jitter = 50
preload_app = True

# Performance tuning
worker_tmp_dir = "/dev/shm"
threads = 4
timeout = 120
graceful_timeout = 30

# Logging
accesslog = "-"
errorlog = "-"
loglevel = "info"
access_log_format = '%(h)s %(l)s %(u)s %(t)s "%(r)s" %(s)s %(b)s "%(f)s" "%(a)s" %(D)s'

# StatsD integration
statsd_host = "localhost:8125"
statsd_prefix = "fastapi"

Dockerizing FastAPI Applications

Production-Ready Multi-Stage Dockerfile

# Build stage
FROM python:3.11-slim as builder

WORKDIR /app

# Install build dependencies
RUN apt-get update && \
    apt-get install -y --no-install-recommends \
    gcc \
    g++ \
    python3-dev \
    && rm -rf /var/lib/apt/lists/*

# Install Python dependencies
COPY requirements.txt .
RUN pip install --no-cache-dir --user -r requirements.txt

# Production stage
FROM python:3.11-slim

WORKDIR /app

# Install runtime dependencies
RUN apt-get update && \
    apt-get install -y --no-install-recommends \
    curl \
    && rm -rf /var/lib/apt/lists/*

# Copy Python packages from builder
COPY --from=builder /root/.local /root/.local

# Copy application
COPY . .

# Create non-root user
RUN groupadd -r fastapi && useradd -r -g fastapi fastapi
RUN chown -R fastapi:fastapi /app
USER fastapi

# Environment
ENV PATH=/root/.local/bin:$PATH
ENV PYTHONPATH=/app
ENV PYTHONUNBUFFERED=1

# Health check
HEALTHCHECK --interval=30s --timeout=3s --start-period=40s --retries=3 \
    CMD curl -f http://localhost:8000/health || exit 1

# Start server
CMD ["gunicorn", "main:app", "-c", "gunicorn.conf.py"]

Docker Compose for Development

version: '3.8'

services:
  api:
    build: .
    ports:
      - "8000:8000"
    environment:
      - DATABASE_URL=postgresql://fastapi:password@postgres:5432/fastapi_db
      - REDIS_URL=redis://redis:6379
      - ENV=development
    depends_on:
      - postgres
      - redis
    volumes:
      - .:/app
    command: uvicorn main:app --host 0.0.0.0 --port 8000 --reload

  postgres:
    image: postgres:15-alpine
    environment:
      - POSTGRES_USER=fastapi
      - POSTGRES_PASSWORD=password
      - POSTGRES_DB=fastapi_db
    volumes:
      - postgres_data:/var/lib/postgresql/data
    ports:
      - "5432:5432"

  redis:
    image: redis:7-alpine
    ports:
      - "6379:6379"
    command: redis-server --appendonly yes

  nginx:
    image: nginx:alpine
    ports:
      - "80:80"
      - "443:443"
    volumes:
      - ./nginx.conf:/etc/nginx/nginx.conf
      - ./ssl:/etc/nginx/ssl
    depends_on:
      - api

volumes:
  postgres_data:

Async Database Operations

Async SQLAlchemy with FastAPI

# database.py
from sqlalchemy.ext.asyncio import create_async_engine, AsyncSession
from sqlalchemy.orm import sessionmaker, declarative_base
from sqlalchemy import Column, Integer, String, DateTime, Boolean, Float
from datetime import datetime
import os

DATABASE_URL = os.getenv("DATABASE_URL", "postgresql+asyncpg://user:pass@localhost/db")

engine = create_async_engine(
    DATABASE_URL,
    echo=False,
    pool_size=20,
    max_overflow=40,
    pool_pre_ping=True,
    pool_recycle=3600
)

AsyncSessionLocal = sessionmaker(
    engine,
    class_=AsyncSession,
    expire_on_commit=False
)

Base = declarative_base()

# Dependency
async def get_db():
    async with AsyncSessionLocal() as session:
        try:
            yield session
            await session.commit()
        except Exception:
            await session.rollback()
            raise
        finally:
            await session.close()

# Models
class User(Base):
    __tablename__ = "users"
    
    id = Column(Integer, primary_key=True, index=True)
    email = Column(String, unique=True, index=True)
    username = Column(String, unique=True, index=True)
    hashed_password = Column(String)
    is_active = Column(Boolean, default=True)
    created_at = Column(DateTime, default=datetime.utcnow)

# Async CRUD operations
from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession

class UserCRUD:
    @staticmethod
    async def get_user(db: AsyncSession, user_id: int):
        result = await db.execute(
            select(User).where(User.id == user_id)
        )
        return result.scalar_one_or_none()
    
    @staticmethod
    async def get_users(db: AsyncSession, skip: int = 0, limit: int = 100):
        result = await db.execute(
            select(User).offset(skip).limit(limit)
        )
        return result.scalars().all()
    
    @staticmethod
    async def create_user(db: AsyncSession, user_data: dict):
        db_user = User(**user_data)
        db.add(db_user)
        await db.commit()
        await db.refresh(db_user)
        return db_user

High-Performance Async Endpoints

# routes/users.py
from fastapi import APIRouter, Depends, HTTPException, Query
from sqlalchemy.ext.asyncio import AsyncSession
from typing import List, Optional
import asyncio
from cachetools import TTLCache
from functools import wraps

router = APIRouter(prefix="/api/v1/users", tags=["users"])

# In-memory cache
cache = TTLCache(maxsize=1000, ttl=300)

def async_cache(key_prefix: str):
    def decorator(func):
        @wraps(func)
        async def wrapper(*args, **kwargs):
            cache_key = f"{key_prefix}:{args}:{kwargs}"
            if cache_key in cache:
                return cache[cache_key]
            result = await func(*args, **kwargs)
            cache[cache_key] = result
            return result
        return wrapper
    return decorator

@router.get("/", response_model=List[UserResponse])
@async_cache("users_list")
async def list_users(
    skip: int = Query(0, ge=0),
    limit: int = Query(100, le=1000),
    db: AsyncSession = Depends(get_db)
):
    """
    List users with pagination and caching
    """
    users = await UserCRUD.get_users(db, skip=skip, limit=limit)
    return users

@router.get("/{user_id}", response_model=UserResponse)
async def get_user(
    user_id: int,
    db: AsyncSession = Depends(get_db)
):
    """
    Get user by ID with automatic 404 handling
    """
    user = await UserCRUD.get_user(db, user_id)
    if not user:
        raise HTTPException(status_code=404, detail="User not found")
    return user

@router.post("/bulk", response_model=List[UserResponse])
async def create_users_bulk(
    users: List[UserCreate],
    db: AsyncSession = Depends(get_db)
):
    """
    Bulk create users with async concurrency
    """
    tasks = [UserCRUD.create_user(db, user.dict()) for user in users]
    created_users = await asyncio.gather(*tasks)
    return created_users

WebSocket Implementation

Real-Time Features with FastAPI

# websocket_manager.py
from fastapi import WebSocket, WebSocketDisconnect
from typing import Dict, Set
import json
import asyncio

class ConnectionManager:
    def __init__(self):
        self.active_connections: Dict[str, Set[WebSocket]] = {}
        self.user_connections: Dict[str, WebSocket] = {}
    
    async def connect(self, websocket: WebSocket, room: str, user_id: str):
        await websocket.accept()
        if room not in self.active_connections:
            self.active_connections[room] = set()
        self.active_connections[room].add(websocket)
        self.user_connections[user_id] = websocket
    
    def disconnect(self, websocket: WebSocket, room: str, user_id: str):
        self.active_connections[room].discard(websocket)
        if not self.active_connections[room]:
            del self.active_connections[room]
        if user_id in self.user_connections:
            del self.user_connections[user_id]
    
    async def send_personal_message(self, message: str, websocket: WebSocket):
        await websocket.send_text(message)
    
    async def broadcast_to_room(self, message: str, room: str):
        if room in self.active_connections:
            tasks = []
            for connection in self.active_connections[room]:
                tasks.append(connection.send_text(message))
            await asyncio.gather(*tasks, return_exceptions=True)

manager = ConnectionManager()

@app.websocket("/ws/{room}/{user_id}")
async def websocket_endpoint(
    websocket: WebSocket,
    room: str,
    user_id: str
):
    await manager.connect(websocket, room, user_id)
    try:
        while True:
            data = await websocket.receive_text()
            message = json.loads(data)
            
            # Process message
            response = {
                "user_id": user_id,
                "room": room,
                "message": message["content"],
                "timestamp": datetime.utcnow().isoformat()
            }
            
            # Broadcast to room
            await manager.broadcast_to_room(
                json.dumps(response),
                room
            )
    except WebSocketDisconnect:
        manager.disconnect(websocket, room, user_id)
        await manager.broadcast_to_room(
            json.dumps({"user_id": user_id, "status": "disconnected"}),
            room
        )

Background Tasks and Job Queues

Celery Integration with FastAPI

# celery_app.py
from celery import Celery
from celery.result import AsyncResult
import os

celery_app = Celery(
    "fastapi_tasks",
    broker=os.getenv("REDIS_URL", "redis://localhost:6379"),
    backend=os.getenv("REDIS_URL", "redis://localhost:6379"),
    include=["app.tasks"]
)

celery_app.conf.update(
    task_serializer="json",
    accept_content=["json"],
    result_serializer="json",
    timezone="UTC",
    enable_utc=True,
    result_expires=3600,
    task_track_started=True,
    task_time_limit=300,
    task_soft_time_limit=240,
    worker_prefetch_multiplier=4,
    worker_max_tasks_per_child=100,
)

# tasks.py
from celery import Task
from .celery_app import celery_app
import asyncio
from typing import Any

class CallbackTask(Task):
    """Task with callback support"""
    def on_success(self, retval, task_id, args, kwargs):
        """Success callback"""
        print(f"Task {task_id} succeeded with result: {retval}")
    
    def on_failure(self, exc, task_id, args, kwargs, einfo):
        """Failure callback"""
        print(f"Task {task_id} failed with exception: {exc}")

@celery_app.task(base=CallbackTask, bind=True, max_retries=3)
def process_heavy_computation(self, data: dict) -> dict:
    """
    Heavy computation task with retry logic
    """
    try:
        # Simulate heavy computation
        result = perform_computation(data)
        return {"status": "completed", "result": result}
    except Exception as exc:
        # Exponential backoff retry
        raise self.retry(exc=exc, countdown=2 ** self.request.retries)

# FastAPI endpoint
@app.post("/api/v1/tasks/compute")
async def create_computation_task(
    data: ComputationRequest,
    background_tasks: BackgroundTasks
):
    """
    Create async computation task
    """
    task = process_heavy_computation.delay(data.dict())
    
    # Also run a fast background task
    background_tasks.add_task(
        send_notification,
        user_id=data.user_id,
        message="Computation started"
    )
    
    return {
        "task_id": task.id,
        "status": "processing",
        "status_url": f"/api/v1/tasks/{task.id}"
    }

@app.get("/api/v1/tasks/{task_id}")
async def get_task_status(task_id: str):
    """
    Get task status and result
    """
    result = AsyncResult(task_id, app=celery_app)
    
    if result.ready():
        return {
            "task_id": task_id,
            "status": "completed" if result.successful() else "failed",
            "result": result.get() if result.successful() else str(result.info)
        }
    else:
        return {
            "task_id": task_id,
            "status": "processing",
            "current": result.info.get("current", 0) if result.info else 0,
            "total": result.info.get("total", 100) if result.info else 100
        }

Authentication and Security

JWT Authentication with OAuth2

# auth.py
from fastapi import Depends, HTTPException, status
from fastapi.security import OAuth2PasswordBearer, OAuth2PasswordRequestForm
from jose import JWTError, jwt
from passlib.context import CryptContext
from datetime import datetime, timedelta
from typing import Optional
import os

# Configuration
SECRET_KEY = os.getenv("SECRET_KEY", "your-secret-key")
ALGORITHM = "HS256"
ACCESS_TOKEN_EXPIRE_MINUTES = 30
REFRESH_TOKEN_EXPIRE_DAYS = 7

pwd_context = CryptContext(schemes=["bcrypt"], deprecated="auto")
oauth2_scheme = OAuth2PasswordBearer(tokenUrl="/api/v1/auth/token")

class AuthManager:
    @staticmethod
    def verify_password(plain_password: str, hashed_password: str) -> bool:
        return pwd_context.verify(plain_password, hashed_password)
    
    @staticmethod
    def get_password_hash(password: str) -> str:
        return pwd_context.hash(password)
    
    @staticmethod
    def create_access_token(data: dict, expires_delta: Optional[timedelta] = None):
        to_encode = data.copy()
        if expires_delta:
            expire = datetime.utcnow() + expires_delta
        else:
            expire = datetime.utcnow() + timedelta(minutes=ACCESS_TOKEN_EXPIRE_MINUTES)
        
        to_encode.update({"exp": expire, "type": "access"})
        return jwt.encode(to_encode, SECRET_KEY, algorithm=ALGORITHM)
    
    @staticmethod
    def create_refresh_token(data: dict):
        to_encode = data.copy()
        expire = datetime.utcnow() + timedelta(days=REFRESH_TOKEN_EXPIRE_DAYS)
        to_encode.update({"exp": expire, "type": "refresh"})
        return jwt.encode(to_encode, SECRET_KEY, algorithm=ALGORITHM)

# Dependency
async def get_current_user(
    token: str = Depends(oauth2_scheme),
    db: AsyncSession = Depends(get_db)
):
    credentials_exception = HTTPException(
        status_code=status.HTTP_401_UNAUTHORIZED,
        detail="Could not validate credentials",
        headers={"WWW-Authenticate": "Bearer"},
    )
    
    try:
        payload = jwt.decode(token, SECRET_KEY, algorithms=[ALGORITHM])
        username: str = payload.get("sub")
        if username is None:
            raise credentials_exception
    except JWTError:
        raise credentials_exception
    
    user = await UserCRUD.get_user_by_username(db, username=username)
    if user is None:
        raise credentials_exception
    return user

# Rate limiting
from slowapi import Limiter, _rate_limit_exceeded_handler
from slowapi.util import get_remote_address
from slowapi.errors import RateLimitExceeded

limiter = Limiter(key_func=get_remote_address)
app.state.limiter = limiter
app.add_exception_handler(RateLimitExceeded, _rate_limit_exceeded_handler)

@app.post("/api/v1/auth/token")
@limiter.limit("5/minute")
async def login(
    request: Request,
    form_data: OAuth2PasswordRequestForm = Depends(),
    db: AsyncSession = Depends(get_db)
):
    """
    OAuth2 compatible token endpoint
    """
    user = await authenticate_user(db, form_data.username, form_data.password)
    if not user:
        raise HTTPException(
            status_code=status.HTTP_401_UNAUTHORIZED,
            detail="Incorrect username or password",
            headers={"WWW-Authenticate": "Bearer"},
        )
    
    access_token = AuthManager.create_access_token(data={"sub": user.username})
    refresh_token = AuthManager.create_refresh_token(data={"sub": user.username})
    
    return {
        "access_token": access_token,
        "refresh_token": refresh_token,
        "token_type": "bearer"
    }

Performance Optimization

Response Caching with Redis

# caching.py
import redis.asyncio as redis
import json
from functools import wraps
from fastapi import Request
import hashlib

redis_client = redis.from_url("redis://localhost", decode_responses=True)

def cache_response(expire: int = 300):
    """
    Decorator to cache FastAPI responses
    """
    def decorator(func):
        @wraps(func)
        async def wrapper(request: Request, *args, **kwargs):
            # Generate cache key
            cache_key = f"api:{request.url.path}:{hashlib.md5(str(kwargs).encode()).hexdigest()}"
            
            # Try to get from cache
            cached = await redis_client.get(cache_key)
            if cached:
                return json.loads(cached)
            
            # Call function and cache result
            result = await func(request, *args, **kwargs)
            await redis_client.setex(
                cache_key,
                expire,
                json.dumps(result, default=str)
            )
            
            return result
        return wrapper
    return decorator

@app.get("/api/v1/products")
@cache_response(expire=600)
async def get_products(
    request: Request,
    category: Optional[str] = None,
    limit: int = Query(100, le=1000)
):
    """
    Cached product endpoint
    """
    # This will be cached for 10 minutes
    products = await fetch_products(category, limit)
    return products

# Cache invalidation
async def invalidate_cache(pattern: str):
    """
    Invalidate cache by pattern
    """
    cursor = 0
    while True:
        cursor, keys = await redis_client.scan(
            cursor, match=pattern, count=100
        )
        if keys:
            await redis_client.delete(*keys)
        if cursor == 0:
            break

Database Connection Pooling

# connection_pool.py
import asyncpg
from contextlib import asynccontextmanager
from typing import AsyncGenerator

class DatabasePool:
    def __init__(self, database_url: str):
        self.database_url = database_url
        self.pool = None
    
    async def create_pool(self):
        self.pool = await asyncpg.create_pool(
            self.database_url,
            min_size=10,
            max_size=20,
            max_queries=50000,
            max_inactive_connection_lifetime=300,
            command_timeout=60,
            statement_cache_size=0,  # Disable for prepared statements
            server_settings={
                'application_name': 'fastapi',
                'jit': 'off'
            }
        )
    
    async def close_pool(self):
        if self.pool:
            await self.pool.close()
    
    @asynccontextmanager
    async def connection(self) -> AsyncGenerator[asyncpg.Connection, None]:
        async with self.pool.acquire() as conn:
            async with conn.transaction():
                yield conn
    
    async def execute_query(self, query: str, *args):
        async with self.connection() as conn:
            return await conn.fetch(query, *args)
    
    async def execute_many(self, query: str, args_list):
        async with self.connection() as conn:
            return await conn.executemany(query, args_list)

# Usage in FastAPI
db_pool = DatabasePool(DATABASE_URL)

@app.on_event("startup")
async def startup():
    await db_pool.create_pool()

@app.on_event("shutdown")
async def shutdown():
    await db_pool.close_pool()

@app.get("/api/v1/users/search")
async def search_users(q: str):
    query = """
        SELECT id, username, email 
        FROM users 
        WHERE username ILIKE $1 OR email ILIKE $1
        LIMIT 10
    """
    results = await db_pool.execute_query(query, f"%{q}%")
    return [dict(r) for r in results]

Monitoring and Observability

Comprehensive Monitoring Setup

# monitoring.py
from prometheus_client import Counter, Histogram, Gauge, generate_latest
from opentelemetry import trace
from opentelemetry.exporter.jaeger import JaegerExporter
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
import time

# Metrics
request_count = Counter(
    'fastapi_requests_total',
    'Total requests',
    ['method', 'endpoint', 'status']
)

request_duration = Histogram(
    'fastapi_request_duration_seconds',
    'Request duration',
    ['method', 'endpoint']
)

active_requests = Gauge(
    'fastapi_active_requests',
    'Active requests'
)

# Tracing
trace.set_tracer_provider(TracerProvider())
tracer = trace.get_tracer(__name__)

jaeger_exporter = JaegerExporter(
    agent_host_name="localhost",
    agent_port=6831,
)

span_processor = BatchSpanProcessor(jaeger_exporter)
trace.get_tracer_provider().add_span_processor(span_processor)

# Middleware
@app.middleware("http")
async def monitoring_middleware(request: Request, call_next):
    # Metrics
    start_time = time.time()
    active_requests.inc()
    
    # Tracing
    with tracer.start_as_current_span(f"{request.method} {request.url.path}") as span:
        span.set_attribute("http.method", request.method)
        span.set_attribute("http.url", str(request.url))
        
        try:
            response = await call_next(request)
            
            # Record metrics
            duration = time.time() - start_time
            request_count.labels(
                method=request.method,
                endpoint=request.url.path,
                status=response.status_code
            ).inc()
            
            request_duration.labels(
                method=request.method,
                endpoint=request.url.path
            ).observe(duration)
            
            span.set_attribute("http.status_code", response.status_code)
            
            return response
            
        except Exception as e:
            span.record_exception(e)
            span.set_status(trace.Status(trace.StatusCode.ERROR))
            raise
        finally:
            active_requests.dec()

@app.get("/metrics")
async def metrics():
    """
    Prometheus metrics endpoint
    """
    return Response(generate_latest(), media_type="text/plain")

Deployment Strategies

Kubernetes Deployment

# kubernetes/fastapi-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: fastapi-app
spec:
  replicas: 3
  selector:
    matchLabels:
      app: fastapi
  template:
    metadata:
      labels:
        app: fastapi
    spec:
      containers:
      - name: fastapi
        image: myregistry/fastapi:latest
        ports:
        - containerPort: 8000
        env:
        - name: DATABASE_URL
          valueFrom:
            secretKeyRef:
              name: fastapi-secrets
              key: database-url
        - name: REDIS_URL
          valueFrom:
            secretKeyRef:
              name: fastapi-secrets
              key: redis-url
        resources:
          requests:
            memory: "256Mi"
            cpu: "250m"
          limits:
            memory: "512Mi"
            cpu: "500m"
        livenessProbe:
          httpGet:
            path: /health
            port: 8000
          initialDelaySeconds: 30
          periodSeconds: 10
        readinessProbe:
          httpGet:
            path: /health
            port: 8000
          initialDelaySeconds: 5
          periodSeconds: 5
---
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: fastapi-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: fastapi-app
  minReplicas: 3
  maxReplicas: 10
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 70
  - type: Resource
    resource:
      name: memory
      target:
        type: Utilization
        averageUtilization: 80

NGINX Configuration

# nginx.conf
upstream fastapi {
    least_conn;
    server api1:8000 max_fails=3 fail_timeout=30s;
    server api2:8000 max_fails=3 fail_timeout=30s;
    server api3:8000 max_fails=3 fail_timeout=30s;
    keepalive 32;
}

server {
    listen 80;
    server_name api.example.com;
    return 301 https://$server_name$request_uri;
}

server {
    listen 443 ssl http2;
    server_name api.example.com;
    
    ssl_certificate /etc/nginx/ssl/cert.pem;
    ssl_certificate_key /etc/nginx/ssl/key.pem;
    ssl_protocols TLSv1.2 TLSv1.3;
    ssl_ciphers HIGH:!aNULL:!MD5;
    
    # Security headers
    add_header X-Content-Type-Options nosniff;
    add_header X-Frame-Options DENY;
    add_header X-XSS-Protection "1; mode=block";
    add_header Strict-Transport-Security "max-age=31536000; includeSubDomains" always;
    
    # API routes
    location /api/ {
        proxy_pass http://fastapi;
        proxy_http_version 1.1;
        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection "upgrade";
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
        
        # Timeouts
        proxy_connect_timeout 60s;
        proxy_send_timeout 60s;
        proxy_read_timeout 60s;
        
        # Buffering
        proxy_buffering off;
        proxy_request_buffering off;
    }
    
    # WebSocket
    location /ws/ {
        proxy_pass http://fastapi;
        proxy_http_version 1.1;
        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection "upgrade";
        proxy_set_header Host $host;
        proxy_read_timeout 86400;
    }
    
    # Health check
    location /health {
        proxy_pass http://fastapi/health;
        access_log off;
    }
}

Deploying FastAPI with CloudPloy

CloudPloy simplifies FastAPI deployment with automatic configuration:

# cloudploy.yml
name: fastapi-app
framework: fastapi
python: "3.11"

build:
  command: pip install -r requirements.txt

run:
  command: gunicorn main:app -c gunicorn.conf.py
  workers: auto

services:
  - postgres:15
  - redis:7

environment:
  - ENV=production
  - WORKERS=4

scaling:
  min: 2
  max: 10
  target_cpu: 70

health_check:
  path: /health
  interval: 30

domains:
  - api.example.com

CloudPloy Quick Deploy

# Deploy FastAPI app
cloudploy init --framework fastapi
cloudploy deploy --env production

# Enable WebSocket support
cloudploy config set websocket.enabled true

# Scale dynamically
cloudploy scale --min 2 --max 10

Performance Benchmarks

FastAPI vs Other Frameworks

FrameworkRequests/secLatency (p99)MemoryStartup
FastAPI15,00015ms85MB0.8s
Flask3,50095ms120MB1.2s
Django2,100180ms250MB2.5s
Express.js12,00025ms95MB0.5s
Go Gin25,0008ms25MB0.2s

Common Deployment Issues and Solutions

Issue 1: Async Context Errors

# Problem: Synchronous code in async context
@app.get("/bad")
async def bad_endpoint():
    time.sleep(5)  # Blocks event loop!
    return {"status": "done"}

# Solution: Use async alternatives
import asyncio

@app.get("/good")
async def good_endpoint():
    await asyncio.sleep(5)  # Non-blocking
    return {"status": "done"}

Issue 2: Connection Pool Exhaustion

# Problem: Creating new connections per request
@app.get("/users")
async def get_users():
    conn = await asyncpg.connect(DATABASE_URL)  # Bad!
    users = await conn.fetch("SELECT * FROM users")
    await conn.close()
    return users

# Solution: Use connection pool
@app.get("/users")
async def get_users(db=Depends(get_db_pool)):
    users = await db.fetch("SELECT * FROM users")
    return users

Conclusion

FastAPI represents the future of Python API development, combining incredible performance with developer-friendly features. Its native async support, automatic documentation, and type safety make it ideal for modern microservices and API-first architectures.

Whether you’re building real-time applications with WebSockets, high-throughput APIs, or microservices, FastAPI provides the performance and features you need. With CloudPloy, deploying FastAPI becomes even simpler, with automatic configuration, scaling, and monitoring built-in.

Start building lightning-fast APIs with FastAPI and deploy them in minutes with CloudPloy.

Deploy FastAPI with CloudPloy →