LLM Cost & Latency Interview Guide
Production interview guide for LLM cost and latency — caching, routing, streaming, rate limits, and semantic caches.
Interview Guides
Curated study guides with deep-dive question lists for RAG, Agents, Models, and System Design
Browse by topic
19 guides · 49 curated questions · 250 total in library
Production interview guide for LLM cost and latency — caching, routing, streaming, rate limits, and semantic caches.
Interview prep for LLM evaluation — classic metrics limits, LLM-as-judge design, offline vs online eval, and red-teaming.
Observability and evaluation for coding agents — traces, cost/quality metrics, alerts, and private-repo bakeoffs.
Threat model MCP coding agents — prompt injection, malicious servers, .mcp.json attacks, and runtime governance.
Enterprise AI governance for coding agents — approval workflows, least privilege, sandboxes, and AI platform engineering ownership.
What belongs in AGENTS.md, why concise files win, company-wide templates, and securing agent config against repo attacks.
From prompt engineering to context engineering — budgets, retrieval, compression, and company-wide repository context.
Model Context Protocol interview guide: hosts, clients, servers, agent gateways, and enterprise MCP deployments.
Handbook-style interview prep: why LLMs hallucinate, hallucination types, and production defenses (RAG, guardrails, citations, eval) — asked at OpenAI, Google, Meta, Anthropic, Microsoft, and Amazon.
System design for AI coding agents — planner, tools, memory, verifiers, MCP, sandboxes, and a platform for 1,000 engineers.
OpenAI Codex CLI interview guide — when to use a CLI coding agent vs Copilot, plus cost optimization for high-volume agent jobs.
Cursor vs GitHub Copilot interview prep: codebase indexing, agent edits, .cursor/rules, enterprise fit, and review gates.
Interview guide for Claude Code — CLI coding agents, AGENTS.md workflows, sandboxes, and how Claude Code compares to Cursor and Copilot.

ANN search, hybrid retrieval, and scaling to 500M vectors — system design for vector DB interviews.
Chain-of-thought, few-shot, and evaluation strategies — the craft behind reliable LLM outputs.
State machines, cyclic workflows, and the Model Context Protocol — essential for agent infrastructure interviews.
Model-based interview questions for every major LLM family — know the trade-offs before you walk in.
Master multi-agent systems, tool use, and agent orchestration — the hottest topic in GenAI hiring loops.
Everything you need to ace RAG interviews — from fundamentals and hallucination mitigation to enterprise pipeline design.