FDE Bootcamp — Detailed Notes: Modules 9–11

Agents: Frameworks & LangGraph · Advanced Orchestration · Legacy Systems

Part 3 of 5. All code executed and verified before publication — LangGraph examples ran against a real install.


Module 9 · Agentic Frameworks & LangGraph

Bank coverage: S12 (45 questions, median 144w). You know agent theory cold. What follows is the API.

9.1 The design patterns, ranked by when to use them

Pattern Shape Use when
ReAct Thought → Action → Observation, loop Path genuinely unknown in advance
Plan & Execute Plan upfront, then run steps Multi-step, high-stakes, want approval before acting
Router Classify, dispatch to one handler Distinct intents, no iteration needed
Supervisor Orchestrator delegates to workers Sub-tasks need different tools/permissions

The judgement your bank already contains: plan-and-execute is more efficient and auditable but is formed with less information than the executor will have — so re-planning must be designed in, not treated as an exception. ReAct adapts but its cost is unbounded until a step limit fires.

9.2 LangGraph: state, channels and reducers

This is the concept that everything else depends on, and it is where people’s mental model breaks.

State is not a dict you mutate. It is a set of channels, each with a reducer that says how updates merge.

from typing import Annotated, TypedDict
import operator
from langgraph.graph import StateGraph, START, END
from langgraph.checkpoint.memory import InMemorySaver

class S(TypedDict):
    question: str
    steps: Annotated[list[str], operator.add]   # APPEND  (parallel-safe)
    count: int                                  # OVERWRITE (last write wins)

def plan(s: S):   return {"steps": ["plan"],   "count": s["count"] + 1}
def search(s: S): return {"steps": ["search"], "count": s["count"] + 1}
def answer(s: S): return {"steps": ["answer"], "count": s["count"] + 1}

Measured output:

final state : {'question': 'q', 'steps': ['plan','search','answer'], 'count': 3}
steps merged: ['plan','search','answer'] <- operator.add APPENDED, did not overwrite
checkpoint  : 3 steps recorded, next = ()

Note what happened: each node returned {"steps": ["..."]} — a one-element list — and the reducer concatenated them. A node returns a partial update, never the whole state.

Why the reducer matters more than it looks: with parallel branches, two nodes may write the same channel in one superstep. operator.add merges both. The default (overwrite) means one silently wins — this is the single most common source of “my parallel LangGraph nodes lose data”.

9.3 Nodes, edges and conditional routing

def route(s: S):
    if s["count"] >= 4:                      # hard step limit -> loop guard
        return "answer"
    return "search" if "search" not in s["steps"] else "answer"

g = StateGraph(S)
g.add_node("plan", plan); g.add_node("search", search); g.add_node("answer", answer)
g.add_edge(START, "plan")
g.add_conditional_edges("plan", route, {"search": "search", "answer": "answer"})
g.add_edge("search", "answer")
g.add_edge("answer", END)

app = g.compile(checkpointer=InMemorySaver())

The architectural claim to make in an interview: LangGraph models the agent as an explicit state machine rather than an implicit loop. That means the control flow is inspectable, branching is defined, and resumption is possible — which is why it suits production work where a role-based framework does not.

9.4 Checkpointers: what actually makes this production-grade

cfg = {"configurable": {"thread_id": "t1"}}
out  = app.invoke({"question":"q","steps":[],"count":0}, cfg)
snap = app.get_state(cfg)          # -> full state + what runs next

The checkpointer persists state after every superstep, keyed by thread_id. That single mechanism gives you four things:

  1. Durability — a crash or deploy resumes rather than restarts
  2. Human-in-the-loop — the graph genuinely halts with state persisted, and can resume days later from a different process
  3. Time-travel debugging — inspect or fork from any prior checkpoint
  4. Multi-turn memory — same thread_id, state carries across invocations

Use InMemorySaver for tests; PostgresSaver (or equivalent) in production. A checkpointer backed by memory in production means every restart loses in-flight work.

9.5 Parallel execution and subgraphs

g.add_edge("dispatch", "search_web")     # fan-out: same source,
g.add_edge("dispatch", "search_docs")    # multiple targets -> parallel
g.add_edge("search_web", "merge")        # fan-in
g.add_edge("search_docs", "merge")

Parallel nodes run in the same superstep, and their writes to a shared channel must have an append-style reducer or you lose one. That is the state-conflict issue the syllabus lists, and the fix is the Annotated[..., operator.add] above.

Subgraphs — a compiled graph is itself a node. Use them for bounded sub-tasks with their own fresh context, which is the cleanest way to stop scratchpad growth (your Q464).

9.6 When not to reach for multi-agent

Your bank’s Q470 and Q1054 are the strong answers here, and they are contrarian in a good way:

Each handoff is a lossy boundary. Sequential per-step reliability compounds multiplicatively — ten steps at 95% is 60%. Empirically, more agents often reduces success rate.

Justify each split by a concrete constraint — different tools, different permissions, genuine parallelism. If the problem is tool-catalogue size, use tool retrieval (inject only relevant tools per query), not more agents.

⚠️ Verify before quoting: LangChain’s own docs now note the langgraph-supervisor library is being superseded by the direct tool-calling supervisor pattern for most cases. Learn the pattern; treat the helper library as optional.


Module 10 · Advanced Agent Orchestration

Bank coverage: S12 + S56 + S30 (105 questions). Memory and MCP are well covered; Bedrock AgentCore is new.

10.1 Human-in-the-loop with interrupts

from langgraph.types import interrupt, Command

def approve_refund(state):
    decision = interrupt({                    # graph HALTS here, state persisted
        "action": "issue_refund",
        "amount": state["amount"],
        "customer": state["customer_id"],
        "reason": state["reason"],
    })
    if decision["approved"]:
        return {"status": "approved", "approver": decision["user"]}
    return {"status": "rejected", "reason": decision.get("note")}

# Later - possibly days later, from a different process:
app.invoke(Command(resume={"approved": True, "user": "krunal"}), cfg)

What makes an approval useful rather than theatre (your Q461):

Gate by risk tier, not uniformly. Gating everything destroys the usability that justified the agent; gating nothing is how production data gets deleted.

10.2 Memory: the distinction that matters

The model has no long-term memory at all. It has retrieval. Everything remembered must be explicitly written and explicitly fetched back — so every memory system is a write policy plus a retrieval policy, not a model property.

  Short-term Long-term
Where Context window / checkpointer Vector store, DB, graph
Scope Current task Across sessions
Content Goal, plan, recent observations Preferences, decisions, what was already tried
from langgraph.store.memory import InMemoryStore
store = InMemoryStore()
app = g.compile(checkpointer=checkpointer, store=store)

def remember(state, *, store):
    ns = ("memories", state["user_id"])       # NAMESPACE = tenant isolation
    store.put(ns, key=str(uuid4()), value={
        "fact": state["learned"],
        "source": state["trace_id"],          # provenance, for conflict resolution
        "ts": datetime.utcnow().isoformat(),
    })

Store what is about the interaction — preferences, decisions, approaches already tried. Re-fetch what is about the world — a cached balance or order status will be stale exactly when it matters.

The most valuable thing to retain is failure. “This approach did not work” prevents the agent re-attempting broken paths — and it is the first thing summarisation destroys, which is why it needs a pinned, never-summarised region.

10.3 Error handling and loop detection

Classify before responding — the three cases need opposite handling:

Failure Response
Transient (timeout, 5xx, 429) Retry, exponential backoff + jitter, capped
Permanent (400, auth, not-found) Never retry — same call fails identically, budget wasted
Partial (action succeeded, response lost) Idempotency key — or the agent retries a completed payment
def detect_loop(history: list[dict], window: int = 3) -> bool:
    if len(history) < window:
        return False
    recent = [(h["tool"], json.dumps(h["args"], sort_keys=True)) for h in history[-window:]]
    return len(set(recent)) == 1        # identical action+args repeated = bug

Distinguish an error from an empty-but-valid result in the observation text. An agent that cannot tell “no data” from “call failed” retries forever — this is the most common cause of runaway agent cost.

10.4 MCP: what it actually buys you

The problem is combinatorial. Without a standard, N agent frameworks × M tools = N×M bespoke integrations, each re-implementing auth, error handling and schema description. MCP reduces this to N + M.

The secondary benefit is organisational and matters more at enterprise scale: it creates a boundary where governance, authentication and logging apply once at the server, rather than being re-implemented per agent.

from mcp.server.fastmcp import FastMCP
mcp = FastMCP("jira")

@mcp.tool()
def create_ticket(project: str, summary: str, priority: str = "medium") -> dict:
    """Create a Jira ticket.

    Use for logging a NEW issue. Do NOT use to comment on an existing
    ticket — use add_comment for that.

    Args:
        project:  Jira project key, e.g. 'OPS'
        summary:  One-line issue summary
        priority: One of low|medium|high
    """
    ...

Note the docstring. Tool descriptions are API surface the model reads — they belong in code review. Stating when not to use a tool is what disambiguates confusable tools, and it is the fastest fix for a selection problem.

Credentials never enter model context. An agent’s context is attacker-reachable through retrieved content, so a token in the prompt is effectively published. Credentials belong to the execution layer — the server holds them, injected at invocation.


Module 11 · Legacy Systems & Integrations

Bank coverage: S18 + S30 (80 questions), but on data engineering principles, not SOAP/Oracle. This is the biggest genuine skills gap — and the strongest differentiator, because few AI engineers can do it.

11.1 SOAP: the mental model

SOAP is XML-RPC with a contract. The WSDL is a machine-readable description of every operation, its types and its endpoint — which is genuinely better than most REST APIs, since there is no ambiguity about the schema.

from zeep import Client
from zeep.transports import Transport
from requests import Session
from requests.auth import HTTPBasicAuth

session = Session()
session.auth = HTTPBasicAuth(user, password)
session.verify = "/etc/ssl/corp-ca.pem"        # enterprise CA, not disabled TLS

client = Client(
    "https://legacy.corp.internal/PolicyService?wsdl",
    transport=Transport(session=session, timeout=30, operation_timeout=30),
)

print(client.wsdl.dump())                       # ← ALWAYS do this first

result = client.service.GetPolicy(policyNumber="POL-123", asOfDate="2026-01-01")

client.wsdl.dump() is the first thing you run. It prints every operation and type. Legacy WSDLs are frequently undocumented, and this is faster than asking the team that owns it — who may no longer exist.

Handling SOAP faults

from zeep.exceptions import Fault, TransportError

try:
    result = client.service.GetPolicy(policyNumber=num)
except Fault as e:
    # SOAP faults are APPLICATION errors returned with HTTP 500
    code, msg = e.code, e.message
    detail = getattr(e, "detail", None)          # legacy error codes live here
    if code == "soap:Client":
        raise ValueError(f"bad request: {msg}")  # permanent — do NOT retry
    raise TransientError(msg)                    # server fault — retry
except TransportError as e:
    raise TransientError(f"transport: {e}")      # network — retry

The trap: a SOAP fault arrives as HTTP 500 with a valid body. Naive retry logic sees 500 and retries a permanently-invalid request forever. Classify on the fault code, not the HTTP status.

XML security — verified

EVIL = """<?xml version="1.0"?>
<!DOCTYPE r [ <!ENTITY xxe SYSTEM "file:///etc/hostname"> ]>
<r><data>&xxe;</data></r>"""

Measured output:

stdlib ET  : raised ParseError
defusedxml : BLOCKED -> EntitiesForbidden

XXE (XML External Entity) injection can read local files, reach internal endpoints (SSRF), or cause a billion-laughs DoS. Stdlib behaviour has varied across Python versions — do not rely on it.

from defusedxml.ElementTree import fromstring     # not xml.etree

defusedxml refuses entity declarations outright. Use it for any XML you did not generate.

XML → JSON: the ambiguity

from zeep.helpers import serialize_object
data = serialize_object(result)                   # OrderedDict -> then json.dumps

The genuine difficulty: XML does not distinguish a single element from a one-element list. <items><item>A</item></items> may deserialise to a scalar where two items give a list. Normalise explicitly:

def as_list(x):
    if x is None: return []
    return x if isinstance(x, list) else [x]

This single helper prevents a large class of intermittent production bugs.

11.2 Oracle & MS SQL

from sqlalchemy import create_engine, text

# MS SQL via pyodbc
engine = create_engine(
    "mssql+pyodbc://@dsn_name",
    pool_size=5, max_overflow=10, pool_pre_ping=True,   # pre_ping: survives idle drops
    connect_args={"timeout": 30},
)

with engine.connect() as conn:
    conn.execute(text("SET LOCK_TIMEOUT 5000"))          # never block prod indefinitely
    rows = conn.execute(
        text("SELECT TOP 100 name, salary FROM employees WHERE dept = :dept"),
        {"dept": dept},                                   # PARAMETERISED
    ).fetchall()

pool_pre_ping=True is not optional against enterprise databases — firewalls silently drop idle connections, and without it your first query after a quiet period fails.

SQL injection — verified

unsafe query : SELECT name FROM users WHERE name = 'x' OR '1'='1'
unsafe result: [('alice',), ('bob',)]   <- ALL ROWS LEAKED
safe   result: []

The parameterised version treated the payload as a literal string and correctly returned nothing.

This is not theoretical for an FDE. An LLM asked to write SQL will happily produce string-formatted queries. Never interpolate — including values that “came from the model”, which are the least trustworthy of all.

Text-to-SQL, safely

import sqlglot

def validate_sql(generated: str, allowed_tables: set[str]) -> str:
    parsed = sqlglot.parse_one(generated, dialect="tsql")   # real parser, not regex
    if parsed.key != "select":
        raise ValueError("only SELECT permitted")
    for tbl in parsed.find_all(sqlglot.exp.Table):
        if tbl.name.lower() not in allowed_tables:
            raise ValueError(f"table not allowed: {tbl.name}")
    return generated

Layered controls, in order of importance:

  1. Read-only database role — the primary control. Everything else is a backstop.
  2. Parse and validate with a real SQL parser, not regex or substring checks
  3. Row limit and statement timeout enforced server-side
  4. Query a replica, not production
  5. Show the generated SQL to the user — a plausible-but-wrong query returning a confident number is the dangerous failure

Schema in context: do not dump the whole schema. Retrieve the relevant tables and columns with descriptions and example values, since a large schema exceeds context and degrades accuracy. Few-shot with verified query examples from your own schema helps far more than prompt tuning.

11.3 Slack and Jira

# Slack: ACKNOWLEDGE within 3 seconds or Slack retries the event
@app.event("app_mention")
def handle(body, say, ack):
    ack()                                        # immediately
    threading.Thread(target=slow_agent_work, args=(body, say)).start()

Slack’s 3-second rule causes duplicate work if you run the agent inline — Slack retries, and the agent runs twice. Acknowledge first, process asynchronously.

Also: verify the request signature (X-Slack-Signature) on every inbound webhook, or anyone can post to your endpoint.

# Jira: issue keys are stable identifiers - use them for idempotency
existing = jira.search_issues(f'project=OPS AND labels="agent-{trace_id}"')
if existing:
    return existing[0]                            # don't create a duplicate

Checkpoint: what to build after Modules 9–11

Self-test:

  1. Why does a LangGraph node return a partial update rather than the full state?
  2. What breaks if two parallel nodes write the same channel without an append reducer?
  3. Why does a checkpointer enable human-in-the-loop, not just crash recovery?
  4. A SOAP fault returns HTTP 500. Why is retrying on status code wrong?
  5. Why is defusedxml required rather than xml.etree for a legacy integration?
  6. In Text-to-SQL, which control is primary — the SQL parser or the database role, and why?

Part 3 of 5 · Modules 9–11 · Next: Modules 12–14 (IAM, Security & Guardrails, Observability)