AI systems
Designs where a model sits in the request path: retrieve the right evidence, serve tokens within a budget, evaluate the answers, and contain what the model can do when input is hostile.
For: Engineers prepping for RAG, AI assistant or LLM platform prompts
After this path
Design a grounded AI feature end to end: permission-aware retrieval, vector-search choices, a serving and cost story, eval gates, and prompt-injection containment.
- 1Deep dive
RAG and retrieval
Grounding answers in your own data: chunking, embeddings, hybrid retrieval and reranking, the context budget, permissions, freshness and citations.
Why this, here: A common core of AI system prompts: chunk, embed, filter by permission, retrieve and cite.
Checkpoint
A user asks about a document they can't open. Say where the permission filter runs and why filtering the final answer is too late.
- 2Deep dive
Vector search
Approximate nearest-neighbour search: HNSW and IVF, filtering, the recall and latency trade-off, and re-embedding without downtime.
Why this, here: What the vector index is doing, and how recall, latency and filters trade off.
- 3Deep dive
LLM serving and cost
Tokens as the unit of latency and cost: batching, streaming, safe caching, routing and fallbacks, quotas and the usage ledger.
Why this, here: Tokens drive the bill; TTFT, output length and decode speed shape latency. Stream, cache safely, route and meter.
Checkpoint
Name what drives cost (input and output tokens), what drives visible latency (TTFT, output length and decode speed), and one safe way to cut each.
- 4Deep dive
AI evaluation and guardrails
Knowing a model feature works: offline eval sets, retrieval and answer metrics, regression gates, online signals, prompt injection and guardrails.
Why this, here: How you know the answers are good before and after a change, and where guardrails sit in the request path.
- 5Deep dive
Prompt injection is a security problem
A card on Security, authorization and multi-tenancy
Why this, here: Retrieved documents and tool results are untrusted input. Contain tools instead of trusting the prompt.