Forward Deployed Engineer Bootcamp β Study Guide
Companion notes for the Krish Naik Academy FDE Bootcamp v1.0 syllabus (14 modules + 2 capstones), mapped against the AI Prep Buddy question bank.
Read this first: you already own half of it
I mapped all 16 syllabus units against your 1,761-question bank:
| Β | Β |
|---|---|
| Questions covering this syllabus | 890 |
| Share of your bank | 51% |
| Median answer depth | 146 words |
But the overlap is not even, and the shape of the gap is the useful finding:
Your bank is conceptual and architectural. The bootcamp is tool-specific and hands-on, plus consulting.
You can already explain why RRF fuses ranks rather than scores, why post-filtering is a security failure, and why permission inheritance bounds blast radius. What you cannot yet demonstrate is a working
zeepSOAP client, a Colang rail file, or an ECS Fargate task definition.
So do not study this syllabus from scratch. Use your bank as the theory layer and spend your time on the three things it does not contain:
- Tool fluency β LangGraph, Presidio, NeMo, LiteLLM, Neo4j, Zeep, Docker/ECS
- The legacy-integration muscle β SOAP/WSDL/XML, Oracle & MS SQL, on-prem realities
- The consulting lifecycle β discovery workshops, SOWs, UAT runbooks, ROI decks (0% coverage in your bank)
Module-by-module map
| # | Module | Your sections | Qs | Genuinely new to you |
|---|---|---|---|---|
| 1 | Python & Linux Foundations | S26 | 25 | asyncio internals, Linux ops |
| 2 | Modern API Development | S26, S20 | 60 | FastAPI, GraphQL, Pytest |
| 3 | Cloud Fundamentals & Networking | S19, S20 | 65 | VPC/NAT/SG hands-on, IAM policy authoring |
| 4 | Containerization & CI/CD | S20, S16 | 90 | Dockerfile craft, ECS Fargate, Actions |
| 5 | LLM Fundamentals & Prompting | S8, S9 | 90 | Little β you are strong here |
| 6 | Vector Search & Core RAG | S10, S11, S40 | 95 | Pinecone/Qdrant APIs |
| 7 | Enterprise Graph Architecture | S42 | 15 | Cypher + Neo4j hands-on (thin in bank) |
| 8 | Multimodal RAG & Vision AI | S32, S6 | 60 | ColPali / byaldi implementation |
| 9 | Agentic Frameworks & LangGraph | S12 | 45 | LangGraph API specifics |
| 10 | Advanced Agent Orchestration | S12, S56, S30 | 105 | Checkpointers, Bedrock AgentCore |
| 11 | Legacy Systems & Integrations | S18, S30 | 80 | SOAP/WSDL/Zeep, pyodbc/oracledb |
| 12 | Identity & Access Management | S22, S29 | 90 | OAuth/SAML flows hands-on |
| 13 | Production AI Security & Guardrails | S22, S45 | 65 | NeMo Colang, Presidio |
| 14 | AI Observability & Gateway Mgmt | S16, S21, S15 | 135 | LiteLLM, Langfuse, RAGAS APIs |
| C1 | OmniGuard | S13, S29, S23 | 145 | Consulting lifecycle |
| C2 | AuditMesh | S12, S30, S43, S48 | 135 | Consulting lifecycle |
Weakest coverage: Module 7 (Graph) has only 15 questions behind it, and Modules 11/13 lean on sections that discuss the principles without the tooling. Those three are where I would spend disproportionate time.
The five concepts to nail cold
If an interviewer probes depth, these are the ones where your bank already gives you a defensible mechanism-level answer. Rehearse them β they carry the most signal per minute.
1. Why hybrid retrieval, and why RRF specifically.
Dense and sparse have near-complementary failure modes. RRF fuses ranks, not scores, because cosine similarity in ~[0,1] and unbounded BM25 are incomparable β naive weighted addition is dominated by whichever scale is larger. score = Ξ£ 1/(k + rank_i), kβ60.
2. Pre-filter vs post-filter in permission-aware RAG. Post-filtering is both a quality bug and a security failure: you request 10 results, discard 9, relevance collapses β and the search already traversed data the user cannot see. Pre-filter, ideally via native namespace/partition isolation the database enforces. The most common real-world leak is a permission change at source that never propagated to the index β so monitor propagation lag.
3. Guardrails are probabilistic; permissions are deterministic. Probabilistic layers reduce frequency; deterministic layers bound consequence. If an agent took a damaging action, the fix is scoped credentials and a server-side limit β not another classifier. A refund cap in a prompt is a request, not a control, and leaves no audit trail.
4. Excessive agency is the amplifier. Injection alone is not the incident. Injection + over-permissioned tool + unvalidated action = incident. Test it by enumerating what the agentβs credentials permit, not what its prompt says.
5. Why the KV cache decides serving economics.
Decode is memory-bandwidth-bound. Cache per token = 2 Γ layers Γ kv_heads Γ head_dim Γ bytes. That number β not the weights β sets max concurrency, which is why GQA and paged attention exist.
Reference links
Links marked β I verified by search during this session. Unmarked ones are stable official roots I am confident about but have not individually confirmed β check before relying on a deep path.
Krish Naikβs own free material (most relevant β same instructors)
- β Agentic AI roadmap repo β https://github.com/krishnaik06/Roadmap-To-Learn-Agentic-AI
- β RAG tutorials repo β https://github.com/krishnaik06/RAG-Tutorials
- β YouTube playlists β https://www.youtube.com/@krishnaik06/playlists
- β Agentic AI playlist β https://www.youtube.com/playlist?list=PLZoTAELRMXVMBr14UQ30AFlnlQ7eL5wjl
- β Building Agentic AI with LangGraph (live series) β https://www.youtube.com/watch?v=4ETMQkmm41g
- β Getting Started with LangGraph for AI Agents β https://www.youtube.com/watch?v=UltwJqpNA04
- β LangGraph + MCP crash course (2h27m, with timestamps) β https://www.classcentral.com/course/youtube-agentic-ai-with-langgraph-and-mcp-crash-course-part-1-460868
- β Past live classes incl. Neo4j/Cypher crash course β https://www.krishnaik.in/liveclasses
M1βM2 Β· Python, async, APIs
- FastAPI β https://fastapi.tiangolo.com/
- Pydantic β https://docs.pydantic.dev/
- asyncio β https://docs.python.org/3/library/asyncio.html
- Strawberry GraphQL β https://strawberry.rocks/
- pytest β https://docs.pytest.org/
M3βM4 Β· AWS, Docker, CI/CD
- AWS VPC β https://docs.aws.amazon.com/vpc/
- AWS IAM β https://docs.aws.amazon.com/iam/
- ECS Fargate β https://docs.aws.amazon.com/AmazonECS/latest/developerguide/AWS_Fargate.html
- Docker β https://docs.docker.com/
- GitHub Actions β https://docs.github.com/en/actions
M5βM6 Β· LLMs, RAG, vector search
- Pinecone β https://docs.pinecone.io/
- Qdrant β https://qdrant.tech/documentation/
- sentence-transformers β https://sbert.net/
- MTEB leaderboard β https://huggingface.co/spaces/mteb/leaderboard
β οΈ Treat MTEB as a shortlist, not a decision. Evaluate on your own query-document pairs β what counts as βsimilarβ is domain-specific.
M7 Β· Knowledge graphs (your thinnest section)
- Neo4j Cypher manual β https://neo4j.com/docs/cypher-manual/current/
- Neo4j GraphAcademy (free courses) β https://graphacademy.neo4j.com/
- Amazon Neptune β https://docs.aws.amazon.com/neptune/
M8 Β· Multimodal & vision
- β ColPali official repo β https://github.com/illuin-tech/colpali
- β ColPali paper β https://arxiv.org/abs/2407.01449
- ViDoRe benchmark β https://huggingface.co/vidore
M9βM10 Β· Agents, LangGraph, MCP
- β LangGraph supervisor tutorial β https://langchain-ai.github.io/langgraph/tutorials/multi_agent/agent_supervisor/
- β langgraph-supervisor-py β https://github.com/langchain-ai/langgraph-supervisor-py
- β LangGraph supervisor API ref β https://reference.langchain.com/python/langgraph-supervisor
- LangGraph docs β https://langchain-ai.github.io/langgraph/
- MCP spec β https://modelcontextprotocol.io/
β οΈ LangChainβs own docs now note the supervisor library is being superseded by the direct tool-calling pattern for most cases. Learn the pattern, not just the helper.
M11 Β· Legacy integration
- Zeep (SOAP) β https://docs.python-zeep.org/
- SQLAlchemy β https://docs.sqlalchemy.org/
- pyodbc β https://github.com/mkleehammer/pyodbc
M12βM13 Β· IAM & security
- β NeMo Guardrails docs β https://docs.nvidia.com/nemo/guardrails/home
- β NeMo Guardrails repo β https://github.com/NVIDIA-NeMo/Guardrails
- β Microsoft Presidio β https://github.com/Microsoft/presidio
- β Presidio supported entities β https://microsoft.github.io/presidio/supported_entities/
- OWASP Top 10 for LLM Apps β https://owasp.org/www-project-top-10-for-large-language-model-applications/
- OAuth 2.0 β https://oauth.net/2/
M14 Β· Observability & evaluation
- RAGAS β https://docs.ragas.io/
- Langfuse β https://langfuse.com/docs
- LangSmith β https://docs.smith.langchain.com/
- LiteLLM β https://docs.litellm.ai/
- DeepEval β https://github.com/confident-ai/deepeval
Suggested study order
The syllabus order is pedagogical (foundations first). Given what you already know, that order wastes your time. Reorder by gap size:
Phase 1 β close the tooling gaps (highest value)
- Module 7 (Neo4j/Cypher) β thinnest bank coverage, and GraphAcademy is free and fast
- Module 9 + 10 (LangGraph) β you know agent theory cold; you need the API
- Module 13 (NeMo/Presidio) β you know the principles; write actual Colang and Presidio configs
Phase 2 β the unglamorous differentiator
- Module 11 (SOAP/legacy) β few people can do this, and it is most of what an FDE actually hits in a bank or insurer
- Module 3 + 4 (AWS/Docker) β if not already fluent
Phase 3 β skim, do not study
- Modules 5, 6, 14 β your bank covers these at interview depth already. Do the labs, skip the theory.
Phase 4 β the real gap
- Both capstonesβ Pillar 1 (consulting lifecycle). Your bank has zero coverage of discovery workshops, SOWs, UAT runbooks and ROI decks. This is the actual differentiator between an AI Engineer and an FDE, and it is the part you cannot learn from documentation.
Build-first checklist
Reading these docs will not make you an FDE. Ship these instead β each proves a claim on your CV:
- FastAPI service, Dockerised, deployed to ECS Fargate via GitHub Actions
- Hybrid RAG (dense + BM25 + RRF) with a cross-encoder re-ranker β measure recall@k before and after
- Same pipeline with pre-filtered permission enforcement, then adversarially test: query as user A for user Bβs content, assert nothing returns
- Neo4j graph from a relational dataset; answer a 2-hop question vector search cannot
- ColPali/byaldi retrieval over scanned PDFs; compare against an OCR-first baseline on the same queries
- LangGraph supervisor with a checkpointer, a human-in-the-loop interrupt, and a step limit
- MCP server exposing one internal tool with scoped credentials and argument validation
- NeMo Colang rails + Presidio redaction; measure the false-positive rate, not just false negatives
- LiteLLM gateway with fallback routing; kill the primary provider and prove failover works
- Langfuse tracing end-to-end; produce cost-per-successful-task, not cost-per-request
Notes on the syllabus itself
Three observations worth having, since interviewers may probe them:
βBeginnerβ is optimistic. The prerequisites say Python/SQL/APIs/Git with no GenAI/cloud/DevOps needed β but Modules 3, 4 and 11 (VPCs, Fargate, SOAP/WSDL, Oracle drivers) are solidly mid-level infrastructure work. Budget more than 8 hrs/week for those.
The role matrix is directionally right. The FDE profile β high on rapid prototyping, client interaction and enterprise integration; lower on model development and scalability than the other two roles β matches what the job actually is. Worth internalising: an FDE is judged on integration and delivery under constraint, not on modelling.
Verify the tool list before interviews. Bedrock AgentCore, MCP registry conventions and the LangGraph supervisor API have all moved recently. Quoting a stale feature comparison is worse than reasoning from principles.
Generated from FDE_BootCamp_V1_0.pdf against the AI Prep Buddy bank at commit af7374f (1,761 answers, median 146 words).