FDE Bootcamp — Detailed Notes: Modules 5–8

Applied AI: LLM Fundamentals · Vector Search & RAG · Knowledge Graphs · Multimodal

Part 2 of 5. All code executed and verified before publication.


Module 5 · LLM Fundamentals & Prompting

Bank coverage: S8 + S9 (90 questions), median 152w and 147w. This is your strongest area. Skim the theory; do the labs. Notes below cover only what the bank does not.

5.1 The three things worth re-reading in your own bank

Before anything new, these are the answers that carry the most interview signal and are already written:

5.2 Structured output: the enforcement ladder

The bootcamp treats “forcing LLMs to output structured data” as one topic. It is actually four, in strict order of reliability:

Level Mechanism Guarantee
1 Ask nicely in the prompt None
2 JSON mode Syntactically valid JSON — not your schema
3 Tool/function calling with a schema Provider enforces during decoding
4 Grammar-constrained decoding Structurally impossible to violate

Level 2 is where teams stop and get burned. JSON mode gives you parseable JSON with the wrong fields. Always validate afterwards.

from pydantic import BaseModel, Field, ValidationError
from typing import Literal

class Extraction(BaseModel):
    invoice_id: str
    total: float = Field(ge=0)
    currency: Literal["GBP","USD","EUR"]     # enum >> free string

def extract_with_retry(client, doc: str, max_retries: int = 1):
    messages = [{"role":"user","content": PROMPT.format(doc=doc)}]
    for attempt in range(max_retries + 1):
        raw = client.chat(messages, response_format={"type":"json_object"})
        try:
            return Extraction.model_validate_json(raw)
        except ValidationError as e:
            if attempt == max_retries:
                raise                      # fail explicitly, don't half-parse
            # Feed the SPECIFIC error back — far better than a blind retry
            messages += [
                {"role":"assistant","content": raw},
                {"role":"user","content": f"Validation failed:\n{e}\nReturn corrected JSON only."},
            ]

Why the error feedback matters: a blind retry re-rolls the same dice. Feeding the specific validation error recovers most residual failures, because the model can see what it got wrong.

Before reaching for retries, check max_tokens. Output truncated mid-object is a very common cause of “malformed JSON” that no amount of prompt tuning fixes.

5.3 Tool calling and hallucinated calls

The syllabus explicitly lists “Managing hallucinated tool calls” — worth being precise about what that means:

  1. Non-existent tool — the model invents a function name. Fix: validate against the registered tool list; return a specific error naming the available tools.
  2. Fabricated arguments — the request lacked the information, but the field was marked required, so the model invented a value. Fix: mark only genuinely required fields required. This is the single highest-leverage change.
  3. Wrong tool selected — semantic overlap between tool descriptions. Fix: rewrite descriptions to say when NOT to use each tool.
def dispatch(call, registry: dict, args_model_by_name: dict):
    if call.name not in registry:
        return {"error": f"unknown tool '{call.name}'",
                "available": list(registry)}          # actionable feedback
    try:
        args = args_model_by_name[call.name].model_validate(call.arguments)
    except ValidationError as e:
        return {"error": "invalid_arguments", "detail": str(e)}
    return registry[call.name](**args.model_dump())   # validated before execution

The principle: the model’s output is a proposal. Validation happens before execution, in code. Same principle as Module 4’s task role and Module 12’s RBAC.

5.4 Token accounting

Costs diverge from estimates for reasons that are all avoidable — count with the actual tokeniser, and remember these are billed and usually invisible in application code:

Put stable content first. Prefix caching is prefix-exact: one differing token near the top invalidates everything after it.


Module 6 · Vector Search & Core RAG

Bank coverage: S10 + S11 + S40 (95 questions), all at 145w+. Theory is covered; the code below is the implementation layer.

6.1 Chunking, with the bug that bites everyone

def chunk(tokens, size, overlap):
    if overlap >= size:
        raise ValueError(f"overlap({overlap}) >= size({size}) -> infinite loop")
    out, step, i = [], size - overlap, 0
    while i < len(tokens):
        out.append(tokens[i:i+size])
        i += step
    return out

Measured output:

size=8 overlap=2: 4 chunks, starts=[0, 6, 12, 18]
size=8 overlap=0: 3 chunks, starts=[0, 8, 16]
guard fired: overlap(8) >= size(8) -> infinite loop

Three non-negotiables:

  1. Chunk by tokens, using the embedding model’s own tokeniser. Character or whitespace approximations silently truncate content past the model’s sequence limit.
  2. Guard overlap >= size. Without it the loop never advances.
  3. Overlap is a mitigation for arbitrary boundaries. If you need large overlap, the chunking strategy is the real problem — go structure-aware.

Chunk size: sweep it against end-to-end answer accuracy, not retrieval recall. Typical range 200–800 tokens, but it is corpus-dependent.

6.2 Embeddings: the efficiency point

import numpy as np
cos = (M @ q) / (np.linalg.norm(M, axis=1) * np.linalg.norm(q))

Mn = M / np.linalg.norm(M, axis=1, keepdims=True)   # normalise ONCE, at index time
qn = q / np.linalg.norm(q)
dot = Mn @ qn                                        # one BLAS matmul

Measured output:

max abs diff cosine vs dot(normalised): 1.56e-17
cosine rank order: [3 4 2 1 0]
euclid rank order: [3 4 2 1 0]  <- identical for normalised vectors

Two facts fall out, both worth stating in an interview:

Use float32, not float64 — half the memory, better cache behaviour, no meaningful precision loss for embeddings.

6.3 Hybrid search and RRF — verified

def rrf(rankings: dict[str, list[str]], k: int = 60):
    scores = {}
    for _, ranked in rankings.items():
        for rank, doc in enumerate(ranked, start=1):
            scores[doc] = scores.get(doc, 0.0) + 1.0 / (k + rank)
    return sorted(scores.items(), key=lambda x: -x[1])

Measured output:

dense  : ['d1', 'd2', 'd3', 'd4']
sparse : ['d3', 'd9', 'd1', 'd5']

RRF fused:
  d1  0.03227
  d3  0.03227     <- both retrievers ranked these highly; correctly co-promoted
  d2  0.01613
  d9  0.01613

Why RRF and not weighted score addition — demonstrated

--- naive weighted sum with real-world scales ---
  dense scores : 0.92, 0.88, 0.85, 0.80    (cosine, bounded ~[0,1])
  sparse scores: 41.2, 38.7, 22.1, 19.5    (BM25, unbounded)
  naive top-3  : ['d3', 'd9', 'd1']
  -> BM25 magnitudes dominate; dense contributes essentially nothing

The mechanism: RRF fuses ranks, not scores, so it needs no normalisation and is immune to incomparable scales. d9 — ranked 2nd by sparse and absent from dense — placed above d4 which dense ranked 4th. That is the intended behaviour.

The constant k≈60 damps top-rank influence, so a document ranked 1st by one retriever and 50th by another does not automatically win.

6.4 The evaluation split that makes debugging possible

Never report one “RAG accuracy” number. Report four, because they have different fixes:

Metric Question Needs ground truth?
recall@k Was the answer-bearing chunk retrieved at all? Yes — but see below
groundedness Is each claim entailed by the retrieved context? No
citation accuracy Does the cited passage support that specific claim? No — checkable
answer correctness Is it actually true? Yes

recall@k is the ceiling. What is not retrieved cannot be used, however good the generator.

Groundedness needs no ground truth, which means it runs on live production traffic as a monitor rather than only offline. That is the single most useful RAG metric to instrument first.

Bootstrapping an eval set in an afternoon

# For each chunk, generate a question it answers.
# The chunk is then the known-relevant document -> free retrieval labels.
prompt = f"Write one specific question that this passage answers:\n\n{chunk}"

Questions are more literal than real user queries, but this gives a usable recall@k metric immediately. Replace with mined production queries as they accumulate.

6.5 Re-ranking

Two-stage exists because the stages have incompatible constraints:

The accuracy gain comes from what the model is allowed to look at, not from model size. A cross-encoder cannot search a corpus; a bi-encoder cannot see the pair together.

Retrieve 50–100, re-rank, pass 3–8. Re-ranking is typically the single highest-value addition to a naive RAG system.


Module 7 · Enterprise Graph Architecture

Bank coverage: S42 — only 15 questions. Your thinnest section relative to syllabus weight. Spend disproportionate time here.

7.1 When a graph earns its cost

Vector RAG is fundamentally local: it returns the k most similar passages. So a question whose answer is distributed across many documents, or that depends on a connection rather than a passage, is unanswerable by similarity search.

Graphs win on exactly three query classes:

  1. Multi-hop — the connecting fact never co-occurs with either endpoint in a single chunk
  2. Relational — “how are X and Y connected”
  3. Global/aggregative — “what are the main themes across the corpus”

The honest cost: graph construction needs an LLM pass over the entire corpus, extraction is lossy and error-prone, schema design is real work, and updates are harder than re-embedding a chunk. Adopt it when you can name the failing query class — not on principle.

7.2 Cypher: the working subset

// Nodes have labels + properties; relationships are typed and directed
CREATE (a:Customer {id: 'C1', name: 'Acme Ltd', tier: 'enterprise'})
CREATE (p:Product  {sku: 'P9', name: 'Widget'})
CREATE (a)-[:PURCHASED {date: date('2026-01-15'), qty: 40}]->(p)
// MATCH is pattern matching, not a join. Read it as a picture.
MATCH (c:Customer {tier:'enterprise'})-[r:PURCHASED]->(p:Product)
WHERE r.date > date('2026-01-01')
RETURN c.name, p.name, r.qty
ORDER BY r.qty DESC
LIMIT 10

The queries that justify the graph

// 2-hop: customers who bought what THIS customer bought (collaborative signal)
MATCH (me:Customer {id:'C1'})-[:PURCHASED]->(:Product)<-[:PURCHASED]-(peer:Customer)
WHERE peer <> me
RETURN peer.name, count(*) AS shared
ORDER BY shared DESC
// Variable-length path - THE thing SQL cannot express cleanly
MATCH path = (a:Company {name:'Acme'})-[:SUPPLIES*1..4]->(b:Company {name:'Zenith'})
RETURN path, length(path) AS hops
ORDER BY hops
LIMIT 1
// Shortest path - supply-chain risk, fraud rings, org reachability
MATCH (a:Account {id:'A1'}), (b:Account {id:'A9'})
MATCH p = shortestPath((a)-[:TRANSFERRED*..6]-(b))
RETURN p

*1..4 is the whole argument for graph databases. In SQL, a variable-depth traversal is a recursive CTE that becomes unreadable and slow. In Cypher it is three characters.

Performance rules

CREATE INDEX customer_id FOR (c:Customer) ON (c.id);           // always
CREATE CONSTRAINT customer_unique FOR (c:Customer) REQUIRE c.id IS UNIQUE;
PROFILE MATCH (c:Customer {id:'C1'})-[:PURCHASED*1..3]->(p) RETURN p;

7.3 Relational → graph modelling

The translation rule is short:

Relational Graph
Row in an entity table Node with a label
Column Property
Foreign key Relationship
Join table (many-to-many) Relationship, with the extra columns as relationship properties

The judgement call: something is a node if you will traverse from it or attach properties to it; otherwise it is a property. city as a property is fine until you need “other customers in the same city” — then it becomes a node.

7.4 GraphRAG in practice

# Text2Cypher: LLM generates the query, but you NEVER execute it unvalidated
def safe_cypher(generated: str) -> str:
    banned = ("CREATE","DELETE","SET","MERGE","REMOVE","DROP","CALL db.")
    upper = generated.upper()
    if any(b in upper for b in banned):
        raise ValueError("write operation rejected")
    if "LIMIT" not in upper:
        generated += "\nLIMIT 100"
    return generated

Same principle as Module 11’s Text-to-SQL: parse and validate before execution, and connect with a read-only role. The allowlist is a backstop, not the primary control — the read-only credential is.

Hybrid is the production answer: graph for entity and thematic questions, vectors for everything else, routed by query classification.


Module 8 · Multimodal RAG & Vision AI

Bank coverage: S32 + S6 (60 questions), median 152w/142w. Strong on architecture; ColPali implementation is new.

8.1 Why document parsing is the hardest part

A PDF stores positioned glyphs, not a document model. There is no reliable notion of paragraphs, columns or reading order. Naive extraction:

The consequence that matters: chunk quality is bounded by parse quality. A table mangled at parse time is unrecoverable no matter how good your retriever is — and it produces confidently wrong numbers, which is worse than a miss because it looks authoritative.

8.2 The two architectures

  Text-mediated (OCR-first) Vision-native (ColPali-style)
Method OCR + layout → index text Embed page images directly
Strengths Cheap, uses existing infra, exact-identifier matching, explainable Preserves layout/tables/figures; no OCR error propagation
Weaknesses Inherits every OCR and layout error; charts lost Large multi-vector index, higher compute, weak on exact identifiers

The production answer is hybrid: OCR text + BM25 for lexical and identifier matching, page embeddings for layout- and figure-dependent queries — and pass the original page image to the vision model at generation time, so detail is not lost twice.

8.3 ColPali: the mechanism

ColPali applies ColBERT-style late interaction to document images. Rather than one vector per page, it produces a multi-vector representation from ViT patch embeddings (PaliGemma-3B, projected linearly), and scores query-to-page with MaxSim: for each query token, take the maximum similarity across all page patches, then sum.

# Simplified: byaldi wraps ColPali for practical use
from byaldi import RAGMultiModalModel

model = RAGMultiModalModel.from_pretrained("vidore/colpali-v1.2")
model.index(input_path="./contracts", index_name="contracts", overwrite=True)

results = model.search("What is the termination notice period?", k=3)
# Returns PAGE IMAGES, not text -> feed directly to a VLM

Why it removes OCR: the model sees layout, fonts, figures and tables as visual signal. There is no brittle layout-recognition stage to fail.

The costs, stated honestly:

References: paper · repo

8.4 The token-budget problem

Image tokens dominate cost and latency. A high-resolution page can consume thousands of tokens. Size this first, before anything else in a multimodal design.

Levers, in order of impact:

  1. Resolution / pixel budget — most tasks do not need full resolution; cap it and downscale adaptively by detected content density
  2. Tiling — split into standard-resolution crops plus a downscaled global thumbnail. Compute then scales linearly with area instead of quadratically, and the thumbnail restores the context tiling destroys
  3. Cache image embeddings by content hash — highly effective for repeated documents
  4. Route — does this request genuinely need vision?

8.5 Tables: the specific failure

A table’s meaning is two-dimensional — a cell is interpretable only with its row and column headers. Naive extraction discards that, so numbers get attributed to the wrong row or period.

Handling:

8.6 Images are an injection vector

Text embedded inside an image can carry instructions the model reads and follows — typographic prompt injection. No text filter on user input sees it.

Treat image content as untrusted input, exactly like a retrieved document: delimit it, disable instruction-following where possible, and gate consequential actions behind approval regardless of what the content said.


Checkpoint: what to build after Modules 5–8

Self-test:

  1. Why does RRF use ranks rather than scores — and what specifically breaks if you use scores?
  2. For normalised vectors, why do cosine and Euclidean give the same ranking?
  3. Which RAG metric can run on production traffic without ground truth, and why?
  4. What can MATCH ... *1..4 express that a SQL join cannot?
  5. Why does tiling a 4K image plus a thumbnail beat monolithic encoding?
  6. Where does a mangled table produce a worse outcome than a retrieval miss?

Part 2 of 5 · Modules 5–8 · Next: Modules 9–11 (Agentic Frameworks, Advanced Orchestration, Legacy Systems)