FDE Bootcamp β€” Detailed Notes: Modules 12–14

Production: Identity & Access Β· Security & Guardrails Β· Observability & Gateways

Part 4 of 5. All code executed and verified before publication.


Module 12 Β· Identity & Access Management

Bank coverage: S22 + S29 (90 questions). Principles are strong; OAuth/SAML mechanics and data-level RBAC implementation are new.

12.1 Authentication vs authorisation

They fail differently, and conflating them is the root of most enterprise AI leaks. A system can authenticate perfectly and still show a user another department’s documents, because tool-level access was granted without data-level entitlement.

12.2 JWT: what verification actually does

import jwt, datetime
tok = jwt.encode({"sub":"krunal","groups":["finance","eu"],"exp":...},
                 SECRET, algorithm="HS256")

Measured output:

-- decode WITHOUT verification (what attackers exploit) --
   ['finance', 'eu']

-- verified decode --
   ['finance', 'eu']

  tampered -> InvalidSignatureError
  expired  -> ExpiredSignatureError

The first block is the point. A JWT is base64 β€” anyone can read the claims without the key. Two consequences:

  1. Never put secrets in a JWT. The payload is public.
  2. Never trust claims without verifying the signature. verify_signature: False returned the same groups list β€” an attacker forging a token with groups: ["admin"] is trivially successful if you skip verification.

Always pin algorithms=["RS256"] explicitly. Accepting the token’s own alg header enables the classic alg: none and RS256β†’HS256 confusion attacks, where the attacker signs with the public key as an HMAC secret.

12.3 OAuth 2.0 grant types

Grant Use for Notes
Authorization Code + PKCE Users in web/mobile apps The default. PKCE now recommended for all clients, not just public ones
Client Credentials Service-to-service, no user Your agent calling an internal API as itself
On-Behalf-Of Service acting for a user The one that matters for FDE work
Implicit / Password β€” Deprecated. Do not use.

On-Behalf-Of is the control that bounds blast radius. The agent exchanges the user’s token for a downstream token carrying that user’s identity, so it cannot exceed what the user could do. Without it, the agent acts with a broad service identity and every retrieval is a potential leak.

SAML vs OIDC

Β  SAML 2.0 OIDC
Format XML assertions JSON / JWT
Age Enterprise incumbent Modern, built on OAuth 2.0
Where Legacy enterprise SSO, AD FS New applications, mobile, APIs

You will meet SAML in enterprises whether you like it or not. And SAML assertions are XML β€” so everything from Module 11 about XXE applies: use a hardened parser, and validate the assertion signature and the audience, issuer and time conditions.

12.4 Data-level RBAC: pre-filter vs post-filter β€” demonstrated

This is the single most important implementation detail in the whole module. Simulation: 1,000 documents, user entitled to ~3%.

# POST-FILTER: search everything, discard after
topk = sorted(docs, key=lambda d: -d["score"])[:10]
kept = [d for d in topk if d["owner"] in user_groups]

# PRE-FILTER: restrict the search space first
allowed = [d for d in docs if d["owner"] in user_groups]
pre = sorted(allowed, key=lambda d: -d["score"])[:10]

Measured output:

post-filter: searched all 1000, top-10 -> 0 survive
             user asked for 10 results, got 0
pre-filter : searched 31 permitted, top-10 -> 10 results

quality gap: 0 vs 10 results for the same request
security   : post-filter TRAVERSED 970 docs the user cannot see

Post-filtering returned nothing at all. That is the quality failure β€” and it degrades silently, so it presents as β€œthe AI can’t find anything” rather than as a bug.

The security failure is worse: the search already read 970 documents the user is not entitled to see. Any bug, refactor, or code path that skips the filter leaks immediately, because the filter is application-level rather than enforced by the store.

Correct implementation:

results = index.query(
    vector=qvec,
    top_k=10,
    filter={"allowed_groups": {"$in": user.groups}},   # PRE-filter, in the DB
    namespace=f"tenant-{user.tenant_id}",              # native isolation
)

Three requirements:

  1. Derive entitlements from the authenticated session, never from client- or model-supplied input
  2. Capture ACLs at ingestion as chunk metadata, flattening inherited folder permissions and group memberships
  3. Refresh permissions independently of content β€” ACLs change far more often than documents, and re-embedding for a permission change is wasteful

The most common real-world leak is a permission change at source that never propagated to the index. Monitor propagation lag as a first-class metric, and test adversarially: query as user A for user B’s content, assert nothing returns β€” as a standing suite, not a one-off.


Module 13 Β· Production AI Security & Guardrails

Bank coverage: S22 + S45 (65 questions, median 146w). Strong on principle; NeMo Colang and Presidio implementation are new.

13.1 OWASP LLM Top 10 β€” the four that compose

The taxonomy’s value is as a coverage checklist. For an agentic system, four items compose into the characteristic failure:

Prompt Injection  β†’  Excessive Agency  β†’  Insecure Output Handling  β†’  real-world consequence
   (entry point)      (the amplifier)      (unvalidated action)

Injection alone is not the incident. Injection + over-permissioned tool + unvalidated action = incident.

Insecure Output Handling is the one most often missed, because teams think about what goes into the model. Model output is untrusted input to whatever consumes it β€” rendered as HTML it is XSS, passed to a shell it is command injection, executed as SQL it is you-know-what.

13.2 Direct vs indirect injection

Β  Direct Indirect
Where User’s own input Retrieved document, web page, email, tool result, image
Adversary The user A third party
Victim Your policy Your user
Input filtering sees it? Yes No

Indirect is the serious class in agentic systems β€” the attack arrives through the data path, so no amount of input filtering helps, and it reaches a model holding live credentials.

13.3 The defence stack, honestly ranked

PROBABILISTIC β€” reduce frequency          DETERMINISTIC β€” bound consequence
─────────────────────────────────         ──────────────────────────────────
input classifiers                          least-privilege credentials
system prompt hardening                    server-side value/rate limits
content delimiting                         permission-aware retrieval
output filtering                           egress allowlisting
                                           human approval on irreversible actions

Safety comes from the right-hand column. An action the agent cannot perform cannot be induced. A refund cap in a prompt is a request; a refund cap in the tool implementation is a control β€” and it leaves an audit trail.

Do not answer an incident with β€œadd more guardrails.” If an agent took a damaging action, the fix is scoped permissions, not another classifier that will sometimes miss.

13.4 Presidio: PII detection and its false-positive problem

from presidio_analyzer import AnalyzerEngine
from presidio_anonymizer import AnonymizerEngine
from presidio_anonymizer.entities import OperatorConfig

analyzer, anonymizer = AnalyzerEngine(), AnonymizerEngine()

results = analyzer.analyze(
    text=prompt,
    entities=["PERSON","EMAIL_ADDRESS","CREDIT_CARD","UK_NHS","IBAN_CODE"],
    language="en",
    score_threshold=0.6,
)
masked = anonymizer.anonymize(
    text=prompt, analyzer_results=results,
    operators={"DEFAULT": OperatorConfig("replace", {"new_value":"<REDACTED>"})},
)

Why β€œevaluating false-positive redaction rates” is in the syllabus

TEXT                            regex   +Luhn     note
payment 4539578763621486        True    True      real test card
order ref 1234567890123456      True    False     16-digit ORDER NUMBER, not a card
invoice 2024011500123456        True    False     date-prefixed invoice id

-> regex alone flags all three; Luhn removes 2 false positives

Two of three were false positives until a checksum validator was applied. In production this means order numbers and invoice IDs get redacted out of prompts, and the model then answers β€œI don’t have that reference” β€” a quality failure attributed to the model rather than to the redactor.

Detection combines regex + validators (Luhn, modulus checks) for structured identifiers, and NER for names and addresses, which have no reliable pattern. Free-text fields are where PII actually hides β€” not the columns labelled as sensitive.

Reversible masking

# Tokenise so the response can be restored, rather than dropped
vault = {}                                     # keep OUT of model context
def mask(text, results):
    for r in sorted(results, key=lambda x:-x.start):
        token = f"<PII_{uuid4().hex[:8]}>"
        vault[token] = text[r.start:r.end]
        text = text[:r.start] + token + text[r.end:]
    return text

The vault becomes the most sensitive asset in the system. State that when you propose it.

Placement matters more than technique: redact before the prompt leaves your boundary, and before logging β€” logs are retained longer, replicated more widely and accessed by more people than production data, so an unscrubbed log pipeline is frequently the worse exposure.

13.5 NeMo Guardrails and Colang

Rails apply at four points: input, dialog, retrieval, output, plus execution rails around tool calls.

# config.yml
models:
  - type: main
    engine: openai
    model: gpt-4o
rails:
  input:
    flows: [self check input, check sensitive data]
  output:
    flows: [self check output, check hallucination]
  config:
    sensitive_data_detection:
      input:
        entities: [PERSON, EMAIL_ADDRESS, CREDIT_CARD]
# rails.co - topical boundary
define user ask off topic
  "what do you think about politics"
  "recommend a restaurant"

define bot refuse off topic
  "I can only help with questions about your policy documents."

define flow
  user ask off topic
  bot refuse off topic
  stop

Colang is a dialogue-flow DSL, not a classifier. It defines canonical user intents by example, matches semantically, and executes a flow. Two versions exist (1.0 and 2.0) with meaningfully different syntax β€” check which your install defaults to before writing rails.

Test rails against a jailbreak library, and track the success rate per attack category over time. A rising rate after a model update is exactly the silent regression worth catching.

13.6 Calibrating aggressiveness

Over-blocking is a real cost, not a safe default:

Fix: tier by severity, fail closed only where a miss is serious, prefer redirect-with-reason over flat refusal, sample blocked requests for human review, and provide an appeal path.


Module 14 Β· AI Observability & Gateway Management

Bank coverage: S16 + S21 + S15 (135 questions) β€” your largest overlap. Tool APIs are the only gap.

14.1 Reliability primitives

Retry with full jitter β€” verified

def fixed(attempt):   return 0.5 * (2 ** attempt)
def jitter(attempt):  return random.uniform(0, 0.5 * (2 ** attempt))

Measured output (20 clients, attempt 3):

fixed       : spread=4.00-4.00s  stdev=0.00
full jitter : spread=0.01-3.78s  stdev=1.27
-> fixed backoff SYNCHRONISES clients into a retry storm; jitter spreads them

Every client waiting exactly 4.00s means they all retry simultaneously β€” a thundering herd against a service that is already struggling. Full jitter spreads them across the window.

def call_with_retry(fn, max_attempts=4, budget_s=30):
    start = time.monotonic()
    for attempt in range(max_attempts):
        try:
            return fn()
        except PermanentError:
            raise                                   # 400/401/404 - never retry
        except TransientError as e:
            if attempt == max_attempts - 1: raise
            wait = e.retry_after or random.uniform(0, 0.5 * 2**attempt)
            if time.monotonic() - start + wait > budget_s: raise
            time.sleep(wait)

Four things this gets right: classify permanent vs transient, honour Retry-After, cap total elapsed time (not just attempts), and use full jitter.

Add a retry budget so retries can never exceed a fraction of total traffic β€” otherwise retries amplify an outage into a self-inflicted DoS.

Idempotency

key = hashlib.sha256(f"{trace_id}:{tool}:{canonical_args}".encode()).hexdigest()
if (prior := store.get(key)): return prior          # already executed
result = execute(); store.set(key, result, ttl=86400)

This is what prevents an agent retrying a completed payment. Required for any tool with side effects, and it is the mitigation for the β€œpartial failure” case β€” action succeeded, response lost.

14.2 The LLM gateway

One control point between all internal agents and all providers:

Function Why it belongs at the gateway
Provider abstraction Switching becomes config, not code
Auth + per-team quotas Applied once, covers everything
Cost attribution The only place with complete visibility
PII redaction Before egress, for every consumer
Fallback + circuit breaking Health-aware routing
Audit logging Complete by construction
# LiteLLM: one interface, provider-agnostic
import litellm
resp = litellm.completion(
    model="bedrock/anthropic.claude-3-5-sonnet-20241022-v2:0",
    messages=msgs,
    fallbacks=["azure/gpt-4o", "ollama/llama3"],
    metadata={"team":"risk","feature":"omniguard","trace_id":trace_id},
)

The organisational value exceeds the technical. It is the only place you can enforce a spend cap, see aggregate usage, apply a security control once, and negotiate provider contracts from a single consumption picture.

The risk to name: the gateway is now on the critical path for everything, so it needs its own reliability engineering and is a single point of failure if not made redundant.

14.3 Evaluation with RAGAS

from ragas import evaluate
from ragas.metrics import faithfulness, answer_relevancy, context_precision, context_recall

result = evaluate(dataset, metrics=[faithfulness, answer_relevancy,
                                     context_precision, context_recall])
Metric Measures Ground truth?
faithfulness Claims entailed by retrieved context No
answer_relevancy Answer addresses the question No
context_precision Retrieved chunks are relevant No
context_recall Ground truth covered by context Yes

The first two run on live production traffic, because they need only the context and the output. That is the practical difference between an offline eval and a production monitor β€” instrument faithfulness first.

LLM-as-judge, and its biases

Measured, documented, and all working against you: position bias (favours first/second in pairwise), verbosity bias (longer scores higher), self-preference (own model family), style over substance (confident, well-formatted, wrong).

Calibrate against human ratings and compare judge-human agreement against human-human agreement. If two experts agree at ΞΊ=0.6, a judge at 0.55 is near the ceiling β€” further tuning is chasing noise.

Re-calibrate whenever the judge model version changes, since judges drift silently when the provider updates.

14.4 Tracing

from langfuse.decorators import observe, langfuse_context

@observe()
def rag_pipeline(question: str, user):
    langfuse_context.update_current_trace(user_id=user.id, session_id=sess,
                                          tags=["omniguard","prod"])
    chunks = retrieve(question, user.groups)     # child span
    return generate(question, chunks)            # child span

One user request = one trace. Every model call, retrieval, tool call and agent handoff is a span carrying inputs, outputs, tokens, latency, cost and resolved model version.

Why tracing is the decisive pillar for AI (not logs or metrics): a multi-step agent failure is many steps removed from its symptom. Only a trace shows which retrieval returned what, and which tool failed.

Requirements teams miss: capture full prompts and responses (the reasoning lives in text that metrics cannot represent), scrub PII at capture rather than downstream, sample intelligently (keep all errors and slow traces plus a fraction of successes), and surface the trace ID to support so a user complaint maps to the actual execution.

14.5 The metrics that actually drive decisions

metrics = {
  "ttft_p95": ...,             # perceived latency when streaming
  "tpot_p95": ...,             # decode speed
  "tokens_in / tokens_out": ..., # priced differently; input usually dominates
  "cost_per_successful_task": ...,   # ← THE headline metric
  "calls_per_user_request": ...,     # ← early warning of agent amplification
  "schema_validity_rate": ...,
  "groundedness": ...,
  "refusal_rate": ...,               # spike = abuse OR over-tightened guardrail
}

Two ratios matter more than absolute spend:

Alert on trend, not absolute thresholds, and run synthetic monitoring β€” a fixed probe set on a schedule, compared to stored baselines. That is the only direct detector of a silent provider-side model update, where output changes on unchanged input with no error and no latency change.


Checkpoint: what to build after Modules 12–14

Self-test:

  1. Why can an attacker read JWT claims without the signing key, and what follows?
  2. Post-filtering returned 0 of 10 results. Name both failures it caused.
  3. Why does a checksum validator matter for PII redaction, and what breaks without it?
  4. Which column of the defence stack provides safety, and why?
  5. Why does fixed exponential backoff make an outage worse?
  6. Which two RAGAS metrics run on production traffic, and why can they?
  7. What is the only reliable detector of a silent provider-side model change?

Part 4 of 5 Β· Modules 12–14 Β· Next: Capstones 1 & 2 β€” including the consulting lifecycle, your largest gap (0% bank coverage)