FDE Bootcamp β€” Detailed Notes: Capstones 1 & 2

OmniGuard Β· AuditMesh Β· and the Consulting Lifecycle

Part 5 of 5. The consulting material below is your only 0%-coverage area β€” and the actual difference between an AI Engineer and a Forward Deployed Engineer.


Why Pillar 1 is the whole point

Both capstones split into Pillar 1 (Consulting Lifecycle) and Pillar 2 (Technical Build). Your instinct will be to spend 90% of the time on Pillar 2, because that is what you are good at.

Do the opposite. You can already architect these systems β€” 890 questions in your bank say so. What you cannot yet do is run a discovery workshop, write a SOW that survives legal review, or present to a CISO who does not care about your architecture.

The FDE role, defined in one line: you are judged on integration and delivery under constraint, not on modelling. The constraint is usually organisational, not technical.


The Consulting Lifecycle

1. Technical discovery workshop

Purpose: establish what is actually true, not what the sponsor believes is true. Most projects fail here β€” the requirement is misunderstood, not the technology.

Agenda (half day, 6–10 people)

Time Item Who
0:00 Business outcome and how it will be measured Sponsor
0:30 Current process, walked through end to end Practitioners
1:15 Data: what exists, where, who owns it, quality Data owners
2:00 Constraints: regulatory, security, residency Legal, security
2:45 Systems: what must it integrate with Architecture
3:15 Success criteria and out-of-scope All

The questions that actually surface risk

On the process:

On data:

On constraints:

On success:

What to listen for

Signal What it means
Nobody knows the current baseline You cannot prove value later. Baseline first.
β€œThe data is all in the warehouse” It is not. Ask to see it.
Sponsor and practitioners describe different processes The documented process is not the real one
No named data owner Access will take weeks, not days
β€œWe just need a chatbot on our documents” Underspecified. Find the actual decision being made.

Output: a discovery memo

One page, circulated within 48 hours: what we heard, what we verified, what we assumed, what is unresolved, and recommended scope. The assumptions list is the most valuable section β€” it is what protects you when something turns out to be false.

2. Data classification

Public       β†’ no restriction
Internal     β†’ all employees
Confidential β†’ need-to-know; most customer and commercial data
Restricted   β†’ regulated; special-category PII, health, payment, MNPI

This is not a paperwork exercise β€” it drives the architecture. Each tier maps to:

Implementation: classify at ingestion, carry the label as metadata through retrieval and into the request, and let the gateway enforce routing policy automatically. Enforcement in the gateway is what makes classification operational rather than aspirational.

3. The architecture SOW

The section that matters most is the one people write last.

## 1. Background and objective
   One paragraph. The business outcome, not the technology.

## 2. Scope
### 2.1 In scope
   - Specific, enumerated deliverables
### 2.2 OUT OF SCOPE          ← the most important section in the document
   - Explicitly named. Everything not listed in 2.1 that a
     reasonable reader might assume is included.
   e.g. "Migration of historical tickets prior to 2024"
        "Integration with the Oracle HR system"
        "Training of end users beyond the two sessions in 5.3"

## 3. Assumptions
   Numbered. Each with an owner and a date by which it must hold.
   "A3. Client provides read access to the policy corpus by 14 March.
        Owner: J. Smith. Slippage impacts timeline 1:1."

## 4. Architecture
   Diagram + component responsibilities + data flows + trust boundaries.

## 5. Deliverables
   Artefacts, not activities. "Deployed service" not "development work".

## 6. Acceptance criteria
   Measurable and testable. "Retrieval recall@5 β‰₯ 0.85 on the agreed
   200-question evaluation set" β€” NOT "system performs well".

## 7. Dependencies and RACI
## 8. Timeline and milestones
## 9. Change control
   How scope changes are requested, priced and approved.

## 10. Commercials

Three rules learned expensively:

  1. Out-of-scope is longer than in-scope. Every assumption a client might reasonably make and you are not delivering goes here.
  2. Acceptance criteria must be measurable and the measurement method agreed upfront. β€œAccurate” is not a criterion. β€œRecall@5 β‰₯ 0.85 on the evaluation set attached as Appendix B” is.
  3. Assumptions carry an owner and a date. An assumption without an owner is a risk you have silently accepted.

For AI specifically, always include: a statement that outputs are probabilistic and that acceptance is against a statistical threshold, not per-case correctness.

4. ROI presentation to CISO / executives

The cost model, built bottom-up

tokens in : 6,010  (system 450, tools 800, history 1200, retrieved 3500)
tokens out: 400
cost/request        : $0.02595
monthly requests    : 100,800
monthly LLM cost    : $2,616

Note the composition. Input tokens are 94% of the volume, and retrieved context is over half of that. This is why β€œreduce retrieved chunks” is usually the largest cost lever β€” and why estimating from the user’s question length is wrong by an order of magnitude.

--- if it becomes agentic (4 model calls per user action) ---
monthly LLM cost    : $10,464   <- 4x, same user traffic

Agent amplification is the multiplier people forget. Same users, same requests, four times the cost. Present this as a scenario, not a footnote.

The value side β€” and how ROI decks lose credibility

My first pass produced this:

CLAIM: 14 min saved x 12 req/day x 21 days
  -> 58.8 hrs/user/month saved
  -> working month ~168 hrs, so claim = 35% of ALL their time
  -> IMPLAUSIBLE. A CFO will kill the whole case on this line.

Always run this sanity check. Convert your claimed saving into percentage of the user’s working time and say it out loud. If it exceeds ~10%, you will not be believed β€” and once one line is disbelieved, the whole deck is discounted.

The defensible version:

REVISED: 6 min x 4 req/day
  hrs/user/month    : 8.4  (5% of working time - defensible)
  monthly requests  : 33,600
  LLM cost          : $872
  gross value       : $174,720
  x realisation 50% : $87,360
  adoption 60%      : net $51,893/mo
  adoption 35%      : net $30,271/mo

Two multipliers that separate a credible deck from a fantasy:

Structure for the room

Slide Content
1 The problem, in their words, with the current cost
2 What we propose β€” one diagram, no jargon
3 Unit economics: cost per request, cost per outcome
4 Value, as a range with adoption and realisation visible
5 Risks and controls β€” what could go wrong and what stops it
6 The decision you need, and by when

For a CISO specifically, lead with slide 5. Their question is not β€œwhat is the ROI” β€” it is β€œwhat is the worst thing this can do, and what stops it?” Answer that first and the rest of the meeting is easier.

Have ready: data flows and where data crosses a boundary; what the AI cannot do (scoped credentials, approval gates); the audit trail; the incident and rollback plan; and the residual risk, named honestly with a proposed accepter.

5. UAT runbook

# UAT: OmniGuard v1.0
Tester: ____  Date: ____  Environment: UAT  Build: sha-________

## Prerequisites
- [ ] Test accounts provisioned: analyst_a (finance), analyst_b (hr), admin
- [ ] Test corpus loaded (247 documents, manifest in Appendix A)

## TC-01  Retrieval accuracy
Steps  : Submit each of the 20 questions in Appendix B
Expect : β‰₯17/20 return the correct source document in the top 3
Result : ____ / 20        Pass / Fail

## TC-07  Entitlement enforcement  [CRITICAL]
Steps  : As analyst_a, ask "What is the HR grievance procedure?"
Expect : No HR content returned. Response states info is unavailable.
         Audit log records the query and the applied entitlement filter.
Result : Pass / Fail       ← any failure BLOCKS go-live

## TC-11  Guardrail - PII in prompt
Steps  : Submit a prompt containing a test card number
Expect : Card masked before egress; masked form visible in the trace
Result : Pass / Fail

## TC-14  Graceful degradation
Steps  : Disable the primary model provider
Expect : Failover within 30s; response quality degraded but valid;
         degraded state visible in telemetry
Result : Pass / Fail

## Sign-off
Business owner: ______  Security: ______  Date: ______
Defects raised: ______  Blocking: ______

Mark the security test cases CRITICAL and make them blocking. A failed entitlement test is not a defect to triage β€” it stops the release.


Capstone 1 Β· OmniGuard: Secure AI Integration

Architecture

User β†’ SSO (OAuth2/OIDC)
   ↓ JWT with groups
FastAPI  ── PII redaction (Presidio) ──┐
   β”‚                                   β”‚
   β”œβ”€β”€ Hybrid RAG                      β”‚  ← entitlement PRE-filter
   β”‚     dense + BM25 β†’ RRF β†’ rerank   β”‚     from JWT groups
   β”‚                                   β”‚
   β”œβ”€β”€ Text-to-SQL (MS SQL)            β”‚  ← read-only role, sqlglot
   β”‚                                   β”‚     validation, row limit
   ↓                                   β”‚
NeMo Guardrails (in/out rails) β†β”€β”€β”€β”€β”€β”€β”€β”˜
   ↓
Response + citations
   ↓
Langfuse trace (PII scrubbed at capture)

Build order

Do not build it in architecture order. Build the thin end-to-end path first, then harden.

  1. FastAPI + one endpoint + Pydantic models β€” deployable skeleton
  2. Hybrid RAG β€” dense + BM25 + RRF. Measure recall@5 before the re-ranker.
  3. Re-ranker β€” measure again. This is your biggest single quality jump; have the number.
  4. Entitlement pre-filter β€” then immediately write the adversarial test (analyst_a for HR content)
  5. Text-to-SQL β€” read-only role first, then sqlglot validation, then the prompt
  6. Presidio β€” and measure the false-positive rate on your identifier formats
  7. NeMo rails β€” topical boundary + output check
  8. Dockerise, deploy to Fargate, wire Langfuse

The five things to get demonstrably right

  1. Pre-filter, never post-filter. Show the adversarial test passing.
  2. Read-only DB role as the primary Text-to-SQL control; parser as backstop.
  3. Secrets from Secrets Manager, injected via the ECS secrets block β€” never environment.
  4. PII redacted before egress and before logging. The logging path is the one people miss.
  5. Citations validated, not just requested β€” check the cited passage actually supports the claim.

Capstone 2 Β· AuditMesh: Multi-Agent Compliance

Pillar 1: mapping the manual workflow

Before any code, map the five-step process as it actually runs:

Step Who Time Bottleneck? Automatable?
1. Retrieve transactions Analyst 20 min No Yes β€” deterministic query
2. Cross-check against policy Analyst 45 min Yes Partly β€” retrieval + judgement
3. Flag exceptions Analyst 15 min No Yes β€” rules
4. Approve escalation Manager 2–3 days YES No β€” must stay human
5. Raise Jira ticket Analyst 10 min No Yes β€” tool call

The finding that shapes the whole design: step 4 is 95% of elapsed time and is a human approval, not work. Automating steps 1–3 and 5 saves 90 minutes; it does not change the 2–3 day cycle time.

This is the most valuable consulting insight in either capstone. The sponsor asked for automation of the analysis. The actual constraint is approval latency. Say so β€” and propose reducing approval friction (better context, mobile approval, tiered thresholds) alongside the automation.

MCP trust boundaries

β”Œβ”€ TRUSTED: your control ──────────────────────┐
β”‚  Orchestrator (LangGraph)                     β”‚
β”‚  MCP client, credential vault                 β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
              β”‚ scoped, per-tool credentials
              ↓
β”Œβ”€ SEMI-TRUSTED: your code, external system ───┐
β”‚  MCP server (Jira)                            β”‚
β”‚  ← enforces authorisation, validates args     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
              ↓
β”Œβ”€ UNTRUSTED: everything returned ─────────────┐
β”‚  Tool results, retrieved documents            β”‚
β”‚  ← may contain injected instructions          β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

State this boundary explicitly in the design document. Three rules follow:

  1. Credentials live in the execution layer, never in model context β€” context is attacker-reachable via retrieved content
  2. Authorisation is enforced at the server, against the underlying system’s own permissions
  3. Every tool result is untrusted input to the next reasoning step β€” delimit it, and never let it drive a consequential action without validation

Latency and cost SLAs

## Service levels

### Latency
- Simple query        p95 ≀ 4s
- Full audit workflow p95 ≀ 90s (excluding human approval time)
- Time to first token p95 ≀ 1.2s

  Human approval time is EXCLUDED and measured separately.
  Reported as approval-wait, not system latency.

### Cost
- Cost per completed audit  ≀ $0.85
- Monthly ceiling           $4,000 (hard cap at gateway)
- Alert at 80% of monthly ceiling

### Quality
- Groundedness β‰₯ 0.90 on sampled production traffic
- Zero tolerance: out-of-scope actions attempted

Excluding human time from the latency SLA is not a trick β€” it is the only honest way to measure a system with a human in the loop, and it keeps the approval bottleneck visible as its own metric rather than hidden inside a blended number.

The technical build

class AuditState(TypedDict):
    transactions: list[dict]
    findings:     Annotated[list[dict], operator.add]   # parallel-safe
    risk_score:   float
    approval:     dict | None
    step_count:   int

def route(s: AuditState):
    if s["step_count"] >= 12:            return "escalate"   # hard limit
    if s["risk_score"] >= 0.7:           return "approve"    # human gate
    if not s["findings"]:                return "analyse"
    return "ticket"

g.add_conditional_edges("assess", route,
    {"analyse":"analyse", "approve":"human_approval",
     "ticket":"create_ticket", "escalate":"escalate_to_human"})

app = g.compile(checkpointer=PostgresSaver(conn))   # NOT InMemorySaver

Non-negotiables: an append reducer on findings (parallel branches), a step limit in the router, a durable checkpointer, and interrupt() showing the exact parsed action with its arguments.

Operations and training handoff

The deliverable most often skipped, and the one that determines whether the system is still running in six months:

Include β€œhow to turn it off” prominently. A client who knows they can stop it is far more willing to run it.


Consolidated build checklist

Capstone 1

Capstone 2


Final self-test

  1. In discovery, why is β€œcan you show me ten real examples” the highest-value question?
  2. What is the most important section of a SOW, and why?
  3. Your ROI claims 14 minutes saved per request. What check do you run before presenting?
  4. Why lead with risks rather than value when presenting to a CISO?
  5. AuditMesh automates 90 minutes of a process with a 3-day cycle time. What do you tell the sponsor?
  6. Why is human approval time excluded from the latency SLA?
  7. Which UAT test cases should block go-live rather than raise a defect?

Part 5 of 5 Β· Series complete: Modules 1–14 + both capstones