FDE Bootcamp β€” Detailed Notes: Modules 1–4

Foundations: Python & Linux Β· APIs Β· Cloud Β· Containers

Part 1 of 5. All code in this document was executed and verified before publication.


Module 1 Β· Python & Linux Foundations

Bank coverage: S26 (25 questions). Mostly algorithmic β€” the async material below is genuinely new.

1.1 The event loop, properly understood

The single most important thing to internalise: async gives you concurrency, not parallelism. One thread, one event loop, cooperatively switching between tasks whenever one awaits.

That means it wins on I/O-bound work (network, disk, database) and does nothing for CPU-bound work.

import asyncio, time

async def io_task(n):
    await asyncio.sleep(0.1)          # simulates network wait
    return n

async def sequential():
    return [await io_task(i) for i in range(10)]

async def concurrent():
    return await asyncio.gather(*(io_task(i) for i in range(10)))

Measured output:

sequential : 1.003s   (10 Γ— 0.1s awaited one at a time)
concurrent : 0.101s   (10 Γ— 0.1s overlapped via gather)
speedup    : 10.0x

The list comprehension in sequential() looks concurrent but is not β€” await inside a comprehension runs them one after another. gather() schedules them all onto the loop together.

The trap that breaks production systems

Calling a blocking function inside async code stalls the entire event loop β€” every other task freezes, including unrelated requests.

def blocking_cpu(n):
    time.sleep(0.1)                   # time.sleep BLOCKS; asyncio.sleep does not
    return n

async def wrong():
    return [blocking_cpu(i) for i in range(5)]        # loop stalled

async def right():
    loop = asyncio.get_running_loop()
    return await asyncio.gather(
        *(loop.run_in_executor(None, blocking_cpu, i) for i in range(5))
    )

Measured output:

blocking in async  : 0.501s   ← event loop stalled
run_in_executor    : 0.104s   ← offloaded to a thread pool

Interview-grade statement: β€œAny synchronous library β€” requests, psycopg2, pyodbc, most SDK clients β€” blocks the loop. Either use the async variant (httpx, asyncpg) or push it to run_in_executor. In a FastAPI service this is the difference between 1,000 rps and 20.”

Concurrency vs parallelism vs the GIL

Β  Mechanism Good for Limited by
asyncio One thread, cooperative I/O-bound Blocking calls
Threads OS threads, preemptive I/O-bound, blocking libs GIL for CPU work
Processes Separate interpreters CPU-bound Memory, IPC cost

The GIL permits only one thread to execute Python bytecode at a time. Threads still help for I/O because the GIL is released during I/O waits. For CPU work you need multiprocessing or a native extension that releases the GIL (NumPy does).

Note: free-threaded CPython (PEP 703) is progressively removing this constraint, but assume the GIL applies unless you’ve explicitly verified otherwise on your runtime.

1.2 Memory management

Python uses reference counting plus a generational cycle collector. Reference counting frees objects immediately at zero refs; the cycle collector exists solely for reference cycles that counting cannot reclaim.

Practical consequences for AI workloads:

import tracemalloc
tracemalloc.start()
# ... run the suspect workload ...
snap = tracemalloc.take_snapshot()
for stat in snap.statistics('lineno')[:5]:
    print(stat)

1.3 Linux for engineers

The subset that actually matters when you are on a client’s box at 2am:

# What is eating the machine
top -o %CPU              # or htop
free -h                  # memory, including cache
df -h                    # disk full is the #1 silent cause of weird failures
du -sh /var/log/*        # find the offender

# What is this process doing
ps aux | grep python
lsof -p <PID>            # open files and sockets
strace -p <PID>          # syscalls (invasive; use briefly)

# Networking
ss -tulpn                # listening ports (replaces netstat)
curl -v https://host     # verbose TLS + headers
dig +short host          # DNS resolution

# Permissions
ls -la                   # note: symlinks show as l-------
chmod 600 key.pem        # SSH keys must not be group/world readable
id / groups              # who am I, effectively

Permissions detail that catches people: a file’s permissions are checked against the effective user of the process, not the login user. In a container running as USER appuser, a volume mounted with root-owned files is unreadable β€” a very common Docker debugging session.


Module 2 Β· Modern API Development

Bank coverage: S26 + S20 (60 questions). Framework specifics are new.

2.1 FastAPI: the parts that matter

FastAPI’s value is that type hints become validation, serialisation and documentation simultaneously.

from fastapi import FastAPI, Depends, HTTPException, status
from pydantic import BaseModel, Field
from typing import Literal, Annotated

app = FastAPI(title="OmniGuard API", version="1.0.0")

class QueryRequest(BaseModel):
    question: str = Field(min_length=1, max_length=2000)
    top_k: int = Field(default=5, ge=1, le=20)
    mode: Literal["hybrid", "dense", "sparse"] = "hybrid"

class Citation(BaseModel):
    doc_id: str
    span: tuple[int, int]

class QueryResponse(BaseModel):
    answer: str
    citations: list[Citation]
    model_version: str

async def get_current_user(token: Annotated[str, Depends(oauth2_scheme)]):
    user = await verify_jwt(token)
    if not user:
        raise HTTPException(status.HTTP_401_UNAUTHORIZED, "Invalid token")
    return user

@app.post("/v1/query", response_model=QueryResponse)
async def query(
    req: QueryRequest,
    user: Annotated[User, Depends(get_current_user)],
):
    # user.groups flows into retrieval as a PRE-filter (see Module 12)
    return await pipeline.run(req, entitlements=user.groups)

What you get free: request validation with precise 422 errors, response serialisation, OpenAPI schema at /docs, and dependency injection.

Pydantic validation, verified

from pydantic import BaseModel, Field, ValidationError
from typing import Literal

class Ticket(BaseModel):
    id: int
    priority: Literal["low","medium","high"]   # enum beats free string
    score: float = Field(ge=0.0, le=1.0)       # range enforced

Measured output:

valid  : id=1 priority='high' score=0.9
caught : 2 errors -> ['priority', 'score']

Both the invalid enum value and the out-of-range float were caught in one pass. This same pattern is how you constrain LLM structured output in Module 5 β€” the model’s JSON is parsed into a Pydantic model and validation errors are fed back for a retry.

Dependency injection is the security seam

Depends() is not just convenience β€” it is where authentication, entitlement resolution and rate limiting attach, once, rather than being repeated in every handler. An FDE reviewing a codebase looks for exactly this: is authorisation a dependency, or is it copy-pasted into 40 endpoints (and missing from three)?

async def vs def in FastAPI

So a handler using a synchronous DB driver should be def, not async def. Getting this backwards is the most common FastAPI performance bug.

2.2 GraphQL and the N+1 problem

REST over-fetches or under-fetches; GraphQL lets the client specify the shape. The cost is a specific and severe performance failure:

query {
  tickets(limit: 100) {      # 1 query
    id
    assignee { name }        # naive resolver: 100 more queries
  }
}

101 queries instead of 2. The fix is a DataLoader β€” batch and cache resolutions within a single request:

from strawberry.dataloader import DataLoader

async def load_users(keys: list[int]) -> list[User]:
    rows = await db.fetch("SELECT * FROM users WHERE id = ANY($1)", keys)
    by_id = {r["id"]: r for r in rows}
    return [by_id.get(k) for k in keys]   # MUST preserve input order

user_loader = DataLoader(load_fn=load_users)

The ordering requirement is non-negotiable β€” DataLoader maps results back to keys positionally.

Other GraphQL production concerns: query depth limiting and cost analysis (a malicious nested query is a DoS), and persisted queries to prevent arbitrary query execution.

API versioning

Approach Pro Con
URL path /v1/, /v2/ Explicit, cacheable, obvious URL proliferation
Header Accept-Version Clean URLs Invisible in logs/browsers, easy to forget

Recommendation: URL versioning for public APIs β€” the version appears in every log line and trace, which matters enormously during an incident.

2.3 Testing with pytest

import pytest
from httpx import AsyncClient, ASGITransport

@pytest.fixture
async def client():
    async with AsyncClient(
        transport=ASGITransport(app=app), base_url="http://test"
    ) as c:
        yield c

@pytest.fixture
def mock_llm(monkeypatch):
    """Never call a real LLM in a unit test β€” cost, latency, non-determinism."""
    async def fake(*args, **kwargs):
        return {"answer": "stub", "citations": []}
    monkeypatch.setattr("app.pipeline.generate", fake)

@pytest.mark.parametrize("top_k,expected", [(0, 422), (1, 200), (21, 422)])
async def test_top_k_bounds(client, mock_llm, top_k, expected):
    r = await client.post("/v1/query", json={"question": "hi", "top_k": top_k})
    assert r.status_code == expected

The AI-specific testing point: LLM output is non-deterministic, so do not assert exact strings. Assert on structure (schema valid), on properties (contains a citation, refuses when it should), and keep genuine quality measurement in a separate evaluation suite (Module 14) β€” not in unit tests.


Module 3 Β· Cloud Fundamentals & Networking

Bank coverage: S19 + S20 (65 questions). Conceptual coverage is strong; hands-on VPC/IAM is new.

3.1 VPC: the mental model

VPC  10.0.0.0/16
β”œβ”€β”€ Public subnet   10.0.1.0/24   β†’ route 0.0.0.0/0 β†’ Internet Gateway
β”‚     └── NAT Gateway, Load Balancer
└── Private subnet  10.0.2.0/24   β†’ route 0.0.0.0/0 β†’ NAT Gateway
      └── ECS tasks, RDS, your model service

The rule that defines everything: a subnet is β€œpublic” if and only if its route table has a route to an Internet Gateway. Nothing else makes it public.

Security Groups vs NACLs

Β  Security Group NACL
Level Instance/ENI Subnet
State Stateful β€” return traffic auto-allowed Stateless β€” must allow both directions
Rules Allow only Allow and deny

The stateless nature of NACLs is the classic gotcha: you allow inbound 443 and the response fails because you forgot the outbound ephemeral port range (1024–65535).

Practical default: use Security Groups for almost everything; reach for NACLs only for coarse subnet-level deny rules.

Security group referencing β€” the pattern to know

ALB-SG:  inbound 443 from 0.0.0.0/0
APP-SG:  inbound 8000 from ALB-SG          ← references the SG, not a CIDR
DB-SG:   inbound 5432 from APP-SG

Referencing security groups rather than IP ranges means the rules survive autoscaling and IP churn. This is what a reviewer expects to see.

3.2 IAM: least privilege in practice

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Action": ["s3:GetObject"],
    "Resource": "arn:aws:s3:::corpus-bucket/tenant-a/*",
    "Condition": {
      "StringEquals": {"aws:PrincipalTag/tenant": "a"},
      "Bool": {"aws:SecureTransport": "true"}
    }
  }]
}

Note what makes this least-privilege: a single action, a path-scoped resource (not *), and conditions binding it to a tagged principal over TLS.

Evaluation order: explicit Deny > explicit Allow > implicit deny. An SCP or permission boundary denying something cannot be overridden by any role policy β€” which is how you enforce guardrails across an org.

For AI workloads specifically:

3.3 Cost control

# Budget with an action, not just an alert
aws budgets create-budget --account-id <id> --budget file://budget.json

The three levers that matter most for AI workloads, in order of impact:

  1. Tag everything, enforced in IaC. Without attribution, β€œreduce cloud cost” is guesswork. Enforce with SCPs or Config rules so untagged resources cannot be created.
  2. Shut down non-production GPU and endpoint resources on a schedule. Consistently the largest and easiest saving.
  3. VPC Endpoints instead of NAT for AWS-service traffic at volume.

Module 4 Β· Containerization & CI/CD

Bank coverage: S20 + S16 (90 questions). Strong on principle; Dockerfile craft and Fargate specifics are new.

4.1 A production Dockerfile

# ---- build stage ----
FROM python:3.12-slim AS builder
WORKDIR /build
RUN pip install --no-cache-dir uv
COPY requirements.txt .
RUN uv pip install --system --no-cache -r requirements.txt

# ---- runtime stage ----
FROM python:3.12-slim
RUN useradd -m -u 1000 appuser
WORKDIR /app

COPY --from=builder /usr/local/lib/python3.12/site-packages \
                    /usr/local/lib/python3.12/site-packages
COPY --from=builder /usr/local/bin /usr/local/bin
COPY --chown=appuser:appuser ./app ./app

USER appuser
EXPOSE 8000
HEALTHCHECK --interval=30s --timeout=3s --start-period=40s \
  CMD python -c "import urllib.request;urllib.request.urlopen('http://localhost:8000/health')"

CMD ["uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8000"]

Every line above earns its place:

Model weights: do not bake them in

# ❌ Adds gigabytes to every image, invalidated on every model change
COPY ./models /app/models

# βœ… Mount at runtime or pull from S3 on start
ENV MODEL_PATH=/mnt/models

Multi-gigabyte layers make pulls slow, which directly damages autoscaling responsiveness β€” a new task cannot serve until the image is pulled.

4.2 ECS Fargate

{
  "family": "omniguard-api",
  "networkMode": "awsvpc",
  "requiresCompatibilities": ["FARGATE"],
  "cpu": "1024", "memory": "2048",
  "executionRoleArn": "arn:aws:iam::...:role/ecsTaskExecutionRole",
  "taskRoleArn":      "arn:aws:iam::...:role/omniguardTaskRole",
  "containerDefinitions": [{
    "name": "api",
    "image": "<acct>.dkr.ecr.<region>.amazonaws.com/omniguard:sha-abc123",
    "portMappings": [{"containerPort": 8000}],
    "secrets": [
      {"name": "OPENAI_API_KEY",
       "valueFrom": "arn:aws:secretsmanager:...:secret:llm/openai"}
    ],
    "logConfiguration": {"logDriver": "awslogs", "options": {...}}
  }]
}

Two roles, and the distinction is examinable:

Role Used by For
executionRoleArn The ECS agent Pulling the image, writing logs, fetching secrets
taskRoleArn Your application Calling S3, Bedrock, RDS at runtime

secrets not environment β€” values are injected from Secrets Manager at start and never appear in the task definition, in docker inspect, or in the console.

Image tagged by commit SHA, never latest β€” latest makes rollback ambiguous and breaks the audit trail.

Fargate sizing

CPU and memory come in fixed valid combinations (e.g. 1 vCPU pairs with 2–8 GB). For AI services, memory is usually the binding constraint β€” model weights plus request concurrency. Size from measured peak RSS plus headroom, not from a guess.

4.3 GitHub Actions

name: deploy
on:
  push: { branches: [main] }

permissions:
  id-token: write        # OIDC β€” no long-lived AWS keys
  contents: read

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: '3.12', cache: pip }
      - run: pip install -r requirements-dev.txt
      - run: pytest --cov=app --cov-fail-under=80
      - run: python scripts/run_evals.py --threshold 0.85   # ← the AI gate

  deploy:
    needs: test
    runs-on: ubuntu-latest
    steps:
      - uses: aws-actions/configure-aws-credentials@v4
        with:
          role-to-assume: arn:aws:iam::<acct>:role/github-deploy
          aws-region: eu-west-1
      - name: Build & push
        run: |
          docker build -t $ECR/omniguard:${{ github.sha }} .
          docker push  $ECR/omniguard:${{ github.sha }}
      - name: Deploy
        run: aws ecs update-service --cluster prod --service api --force-new-deployment

The two things to highlight in an interview:

  1. OIDC (id-token: write) β€” GitHub assumes an AWS role via short-lived federated tokens. No stored AWS keys anywhere. This is the modern correct answer and it directly addresses β€œhow do you manage CI credentials”.

  2. The evaluation gate. Line run_evals.py --threshold 0.85 is what distinguishes AI CI/CD from ordinary CI/CD. It is statistical rather than binary, so:

    • The suite must be large enough to detect the regression size you claim
    • You must know its noise floor β€” run it twice unchanged; a gate tighter than the noise fails randomly and teaches people to re-run until green
    • Report the specific failing cases, not an aggregate score

Checkpoint: what to build after Modules 1–4

Self-test β€” can you answer these without notes?

  1. Why is a def handler sometimes faster than async def in FastAPI?
  2. What exactly makes a subnet public?
  3. Why does --start-period matter more for an AI container than a web app?
  4. What is the difference between the ECS execution role and task role?
  5. Why should an eval gate’s threshold be set relative to its noise floor?

Part 1 of 5 Β· Modules 1–4 Β· Next: Modules 5–8 (LLM Fundamentals, RAG, Knowledge Graphs, Multimodal)