FDE Bootcamp β Detailed Notes: Capstones 1 & 2
OmniGuard Β· AuditMesh Β· and the Consulting Lifecycle
Part 5 of 5. The consulting material below is your only 0%-coverage area β and the actual difference between an AI Engineer and a Forward Deployed Engineer.
Why Pillar 1 is the whole point
Both capstones split into Pillar 1 (Consulting Lifecycle) and Pillar 2 (Technical Build). Your instinct will be to spend 90% of the time on Pillar 2, because that is what you are good at.
Do the opposite. You can already architect these systems β 890 questions in your bank say so. What you cannot yet do is run a discovery workshop, write a SOW that survives legal review, or present to a CISO who does not care about your architecture.
The FDE role, defined in one line: you are judged on integration and delivery under constraint, not on modelling. The constraint is usually organisational, not technical.
The Consulting Lifecycle
1. Technical discovery workshop
Purpose: establish what is actually true, not what the sponsor believes is true. Most projects fail here β the requirement is misunderstood, not the technology.
Agenda (half day, 6β10 people)
| Time | Item | Who |
|---|---|---|
| 0:00 | Business outcome and how it will be measured | Sponsor |
| 0:30 | Current process, walked through end to end | Practitioners |
| 1:15 | Data: what exists, where, who owns it, quality | Data owners |
| 2:00 | Constraints: regulatory, security, residency | Legal, security |
| 2:45 | Systems: what must it integrate with | Architecture |
| 3:15 | Success criteria and out-of-scope | All |
The questions that actually surface risk
On the process:
- βWalk me through the last time this went wrong. What happened?β
- βWho does this today, and what do they do when they are unsure?β β this defines your escalation path
- βWhat percentage of cases are straightforward?β β scope to those first
On data:
- βCan you show me ten real examples?β β always ask. It routinely changes the plan in week one.
- βWho owns this data, and what is the process to get access?β
- βHow current is it? What is the lag from source to here?β
- βAre there records in here that some of your own staff cannot see?β β entitlement requirements
On constraints:
- βWhat would make legal say no to this?β
- βWhere must this data physically reside?β
- βIs there an existing model risk or AI governance process?β
On success:
- βWhat number changes if this works?β
- βWhat is that number today?β β if nobody knows, baselining is your first deliverable
- βWho has to approve this going live?β
What to listen for
| Signal | What it means |
|---|---|
| Nobody knows the current baseline | You cannot prove value later. Baseline first. |
| βThe data is all in the warehouseβ | It is not. Ask to see it. |
| Sponsor and practitioners describe different processes | The documented process is not the real one |
| No named data owner | Access will take weeks, not days |
| βWe just need a chatbot on our documentsβ | Underspecified. Find the actual decision being made. |
Output: a discovery memo
One page, circulated within 48 hours: what we heard, what we verified, what we assumed, what is unresolved, and recommended scope. The assumptions list is the most valuable section β it is what protects you when something turns out to be false.
2. Data classification
Public β no restriction
Internal β all employees
Confidential β need-to-know; most customer and commercial data
Restricted β regulated; special-category PII, health, payment, MNPI
This is not a paperwork exercise β it drives the architecture. Each tier maps to:
- Permitted processing location β Restricted may require self-hosted or a specific in-region endpoint
- Permitted model endpoint β which provider, under what contract
- Logging policy β whether prompt content may be retained at all
- Approval requirement β which tier triggers review
Implementation: classify at ingestion, carry the label as metadata through retrieval and into the request, and let the gateway enforce routing policy automatically. Enforcement in the gateway is what makes classification operational rather than aspirational.
3. The architecture SOW
The section that matters most is the one people write last.
## 1. Background and objective
One paragraph. The business outcome, not the technology.
## 2. Scope
### 2.1 In scope
- Specific, enumerated deliverables
### 2.2 OUT OF SCOPE β the most important section in the document
- Explicitly named. Everything not listed in 2.1 that a
reasonable reader might assume is included.
e.g. "Migration of historical tickets prior to 2024"
"Integration with the Oracle HR system"
"Training of end users beyond the two sessions in 5.3"
## 3. Assumptions
Numbered. Each with an owner and a date by which it must hold.
"A3. Client provides read access to the policy corpus by 14 March.
Owner: J. Smith. Slippage impacts timeline 1:1."
## 4. Architecture
Diagram + component responsibilities + data flows + trust boundaries.
## 5. Deliverables
Artefacts, not activities. "Deployed service" not "development work".
## 6. Acceptance criteria
Measurable and testable. "Retrieval recall@5 β₯ 0.85 on the agreed
200-question evaluation set" β NOT "system performs well".
## 7. Dependencies and RACI
## 8. Timeline and milestones
## 9. Change control
How scope changes are requested, priced and approved.
## 10. Commercials
Three rules learned expensively:
- Out-of-scope is longer than in-scope. Every assumption a client might reasonably make and you are not delivering goes here.
- Acceptance criteria must be measurable and the measurement method agreed upfront. βAccurateβ is not a criterion. βRecall@5 β₯ 0.85 on the evaluation set attached as Appendix Bβ is.
- Assumptions carry an owner and a date. An assumption without an owner is a risk you have silently accepted.
For AI specifically, always include: a statement that outputs are probabilistic and that acceptance is against a statistical threshold, not per-case correctness.
4. ROI presentation to CISO / executives
The cost model, built bottom-up
tokens in : 6,010 (system 450, tools 800, history 1200, retrieved 3500)
tokens out: 400
cost/request : $0.02595
monthly requests : 100,800
monthly LLM cost : $2,616
Note the composition. Input tokens are 94% of the volume, and retrieved context is over half of that. This is why βreduce retrieved chunksβ is usually the largest cost lever β and why estimating from the userβs question length is wrong by an order of magnitude.
--- if it becomes agentic (4 model calls per user action) ---
monthly LLM cost : $10,464 <- 4x, same user traffic
Agent amplification is the multiplier people forget. Same users, same requests, four times the cost. Present this as a scenario, not a footnote.
The value side β and how ROI decks lose credibility
My first pass produced this:
CLAIM: 14 min saved x 12 req/day x 21 days
-> 58.8 hrs/user/month saved
-> working month ~168 hrs, so claim = 35% of ALL their time
-> IMPLAUSIBLE. A CFO will kill the whole case on this line.
Always run this sanity check. Convert your claimed saving into percentage of the userβs working time and say it out loud. If it exceeds ~10%, you will not be believed β and once one line is disbelieved, the whole deck is discounted.
The defensible version:
REVISED: 6 min x 4 req/day
hrs/user/month : 8.4 (5% of working time - defensible)
monthly requests : 33,600
LLM cost : $872
gross value : $174,720
x realisation 50% : $87,360
adoption 60% : net $51,893/mo
adoption 35% : net $30,271/mo
Two multipliers that separate a credible deck from a fantasy:
- Realisation β saved time only becomes value if it is redeployed. Two hours saved across a team does not equal two hours of output unless something absorbs it. 50% is a defensible default; state it explicitly.
- Adoption β a tool used by 35% of intended users delivers 35% of the benefit. This is where most business cases are fiction. Present a range, not a point.
Structure for the room
| Slide | Content |
|---|---|
| 1 | The problem, in their words, with the current cost |
| 2 | What we propose β one diagram, no jargon |
| 3 | Unit economics: cost per request, cost per outcome |
| 4 | Value, as a range with adoption and realisation visible |
| 5 | Risks and controls β what could go wrong and what stops it |
| 6 | The decision you need, and by when |
For a CISO specifically, lead with slide 5. Their question is not βwhat is the ROIβ β it is βwhat is the worst thing this can do, and what stops it?β Answer that first and the rest of the meeting is easier.
Have ready: data flows and where data crosses a boundary; what the AI cannot do (scoped credentials, approval gates); the audit trail; the incident and rollback plan; and the residual risk, named honestly with a proposed accepter.
5. UAT runbook
# UAT: OmniGuard v1.0
Tester: ____ Date: ____ Environment: UAT Build: sha-________
## Prerequisites
- [ ] Test accounts provisioned: analyst_a (finance), analyst_b (hr), admin
- [ ] Test corpus loaded (247 documents, manifest in Appendix A)
## TC-01 Retrieval accuracy
Steps : Submit each of the 20 questions in Appendix B
Expect : β₯17/20 return the correct source document in the top 3
Result : ____ / 20 Pass / Fail
## TC-07 Entitlement enforcement [CRITICAL]
Steps : As analyst_a, ask "What is the HR grievance procedure?"
Expect : No HR content returned. Response states info is unavailable.
Audit log records the query and the applied entitlement filter.
Result : Pass / Fail β any failure BLOCKS go-live
## TC-11 Guardrail - PII in prompt
Steps : Submit a prompt containing a test card number
Expect : Card masked before egress; masked form visible in the trace
Result : Pass / Fail
## TC-14 Graceful degradation
Steps : Disable the primary model provider
Expect : Failover within 30s; response quality degraded but valid;
degraded state visible in telemetry
Result : Pass / Fail
## Sign-off
Business owner: ______ Security: ______ Date: ______
Defects raised: ______ Blocking: ______
Mark the security test cases CRITICAL and make them blocking. A failed entitlement test is not a defect to triage β it stops the release.
Capstone 1 Β· OmniGuard: Secure AI Integration
Architecture
User β SSO (OAuth2/OIDC)
β JWT with groups
FastAPI ββ PII redaction (Presidio) βββ
β β
βββ Hybrid RAG β β entitlement PRE-filter
β dense + BM25 β RRF β rerank β from JWT groups
β β
βββ Text-to-SQL (MS SQL) β β read-only role, sqlglot
β β validation, row limit
β β
NeMo Guardrails (in/out rails) βββββββββ
β
Response + citations
β
Langfuse trace (PII scrubbed at capture)
Build order
Do not build it in architecture order. Build the thin end-to-end path first, then harden.
- FastAPI + one endpoint + Pydantic models β deployable skeleton
- Hybrid RAG β dense + BM25 + RRF. Measure recall@5 before the re-ranker.
- Re-ranker β measure again. This is your biggest single quality jump; have the number.
- Entitlement pre-filter β then immediately write the adversarial test (analyst_a for HR content)
- Text-to-SQL β read-only role first, then sqlglot validation, then the prompt
- Presidio β and measure the false-positive rate on your identifier formats
- NeMo rails β topical boundary + output check
- Dockerise, deploy to Fargate, wire Langfuse
The five things to get demonstrably right
- Pre-filter, never post-filter. Show the adversarial test passing.
- Read-only DB role as the primary Text-to-SQL control; parser as backstop.
- Secrets from Secrets Manager, injected via the ECS
secretsblock β neverenvironment. - PII redacted before egress and before logging. The logging path is the one people miss.
- Citations validated, not just requested β check the cited passage actually supports the claim.
Capstone 2 Β· AuditMesh: Multi-Agent Compliance
Pillar 1: mapping the manual workflow
Before any code, map the five-step process as it actually runs:
| Step | Who | Time | Bottleneck? | Automatable? |
|---|---|---|---|---|
| 1. Retrieve transactions | Analyst | 20 min | No | Yes β deterministic query |
| 2. Cross-check against policy | Analyst | 45 min | Yes | Partly β retrieval + judgement |
| 3. Flag exceptions | Analyst | 15 min | No | Yes β rules |
| 4. Approve escalation | Manager | 2β3 days | YES | No β must stay human |
| 5. Raise Jira ticket | Analyst | 10 min | No | Yes β tool call |
The finding that shapes the whole design: step 4 is 95% of elapsed time and is a human approval, not work. Automating steps 1β3 and 5 saves 90 minutes; it does not change the 2β3 day cycle time.
This is the most valuable consulting insight in either capstone. The sponsor asked for automation of the analysis. The actual constraint is approval latency. Say so β and propose reducing approval friction (better context, mobile approval, tiered thresholds) alongside the automation.
MCP trust boundaries
ββ TRUSTED: your control βββββββββββββββββββββββ
β Orchestrator (LangGraph) β
β MCP client, credential vault β
βββββββββββββββββββββββββββββββββββββββββββββββββ
β scoped, per-tool credentials
β
ββ SEMI-TRUSTED: your code, external system ββββ
β MCP server (Jira) β
β β enforces authorisation, validates args β
βββββββββββββββββββββββββββββββββββββββββββββββββ
β
ββ UNTRUSTED: everything returned ββββββββββββββ
β Tool results, retrieved documents β
β β may contain injected instructions β
βββββββββββββββββββββββββββββββββββββββββββββββββ
State this boundary explicitly in the design document. Three rules follow:
- Credentials live in the execution layer, never in model context β context is attacker-reachable via retrieved content
- Authorisation is enforced at the server, against the underlying systemβs own permissions
- Every tool result is untrusted input to the next reasoning step β delimit it, and never let it drive a consequential action without validation
Latency and cost SLAs
## Service levels
### Latency
- Simple query p95 β€ 4s
- Full audit workflow p95 β€ 90s (excluding human approval time)
- Time to first token p95 β€ 1.2s
Human approval time is EXCLUDED and measured separately.
Reported as approval-wait, not system latency.
### Cost
- Cost per completed audit β€ $0.85
- Monthly ceiling $4,000 (hard cap at gateway)
- Alert at 80% of monthly ceiling
### Quality
- Groundedness β₯ 0.90 on sampled production traffic
- Zero tolerance: out-of-scope actions attempted
Excluding human time from the latency SLA is not a trick β it is the only honest way to measure a system with a human in the loop, and it keeps the approval bottleneck visible as its own metric rather than hidden inside a blended number.
The technical build
class AuditState(TypedDict):
transactions: list[dict]
findings: Annotated[list[dict], operator.add] # parallel-safe
risk_score: float
approval: dict | None
step_count: int
def route(s: AuditState):
if s["step_count"] >= 12: return "escalate" # hard limit
if s["risk_score"] >= 0.7: return "approve" # human gate
if not s["findings"]: return "analyse"
return "ticket"
g.add_conditional_edges("assess", route,
{"analyse":"analyse", "approve":"human_approval",
"ticket":"create_ticket", "escalate":"escalate_to_human"})
app = g.compile(checkpointer=PostgresSaver(conn)) # NOT InMemorySaver
Non-negotiables: an append reducer on findings (parallel branches), a step limit in the router, a durable checkpointer, and interrupt() showing the exact parsed action with its arguments.
Operations and training handoff
The deliverable most often skipped, and the one that determines whether the system is still running in six months:
- Runbook β the top five failure modes, each with symptom, diagnosis and fix
- Kill switch β documented, tested, and accessible to their on-call, not just you
- Named owner on the client side, with a review date
- Monitoring dashboard they can read without you
- What to do when it is wrong β the escalation path, and the feedback route into the eval set
- Two training sessions: operators (how to use, when to distrust) and support (how to triage)
Include βhow to turn it offβ prominently. A client who knows they can stop it is far more willing to run it.
Consolidated build checklist
Capstone 1
- SSO with JWT verification, pinned algorithm
- Hybrid RAG with recall@5 measured before and after re-ranking
- Entitlement pre-filter + passing adversarial cross-user test
- Text-to-SQL: read-only role, sqlglot validation, row limit, SQL shown to user
- Presidio with measured false-positive rate
- NeMo topical + output rails, tested against a jailbreak set
- Fargate deploy, secrets from Secrets Manager, Langfuse tracing
- Discovery memo, SOW, ROI deck, UAT runbook
Capstone 2
- Workflow map identifying the real bottleneck
- LangGraph supervisor: append reducers, step limit, Postgres checkpointer
interrupt()approval showing exact action; approve/edit/reject- Custom MCP server for Jira with scoped credentials and arg validation
- Documented trust boundaries
- Cost-per-successful-audit dashboard
- Streamlit/Gradio approval UI
- Latency/cost SLA doc + ops runbook + training materials
Final self-test
- In discovery, why is βcan you show me ten real examplesβ the highest-value question?
- What is the most important section of a SOW, and why?
- Your ROI claims 14 minutes saved per request. What check do you run before presenting?
- Why lead with risks rather than value when presenting to a CISO?
- AuditMesh automates 90 minutes of a process with a 3-day cycle time. What do you tell the sponsor?
- Why is human approval time excluded from the latency SLA?
- Which UAT test cases should block go-live rather than raise a defect?
Part 5 of 5 Β· Series complete: Modules 1β14 + both capstones