Exam guide v1.0 · effective July 2026 · code CCAR-F

Claude Certified Architect Foundations

Reading notes built from the official exam guide. Every task statement in the blueprint is covered with the concept, the exam traps, a scenario to reason through, and a reference design where one helps. The exam rewards judgment about tradeoffs, so the notes lean on why one option beats another.

60items, single & multi-response
120 minabout 2 min per item
720pass, on a 100–1,000 scale
4 of 6scenarios drawn at random
D1 · 27%Agentic architecture & orchestration
D2 · 18%Tool design & MCP integration
D3 · 20%Claude Code configuration & workflows
D4 · 20%Prompt engineering & structured output
D5 · 15%Context management & reliability

Fee $125 · credential valid 12 months (free renewal assessment if on time) · retake waits 14 → 30 → 90 days, max 4 attempts per rolling 12 months · score report shows percent-correct per domain, but pass/fail uses the total scaled score only.

★

How to pick the right answer

Ten rules that resolve most questions. The sample questions in the guide follow them consistently.
  1. Guarantees come from code, not prompts.If a rule must hold every time (identity check before refund, refund cap), choose hooks or programmatic gates. Prompt instructions and few-shot examples have a non-zero failure rate.
  2. Take the proportionate first step that fixes the root cause.Better tool descriptions, explicit criteria, or few-shot examples come before routing layers, classifiers, or new ML infrastructure. "Over-engineered" is a common reason a distractor is wrong.
  3. Fix the component the evidence points at.If logs show the coordinator decomposed a topic narrowly, the bug is in the coordinator, not in the subagents that did what they were told.
  4. Least privilege for tools.Give each agent the 4–5 tools its role needs. Add a narrow cross-role tool for the frequent case; route the rare complex case through the coordinator.
  5. Structured beats generic.Errors carry category, retryability, what was attempted and partial results. Handoffs carry IDs, root cause, amounts and a recommendation. Findings carry source, date and location.
  6. Never hide a failure, never let one failure kill everything.Returning empty-as-success and terminating the whole workflow are both wrong. Recover locally, then propagate what you could not fix.
  7. More context is not more attention.A bigger context window does not fix diluted attention. Split into focused passes, trim tool output, put key facts at the start.
  8. Self-reported confidence and sentiment are weak proxies.Do not route escalation on "confidence 1–10" or frustration level. Field-level confidence is acceptable only after calibration against a labeled validation set.
  9. A fresh instance reviews better than the author.The generating session keeps its own reasoning and rarely challenges it. Use an independent reviewer instance without that context.
  10. Distrust options that name features you have never seen.The guide's distractors include CLAUDE_HEADLESS=true, a --batch flag and .claude/config.json with a commands array. None exist.
D1

Agentic Architecture & Orchestration

27% of scored items, the largest domain · primary in scenarios 1, 3 and 4

The domain tests whether you can build the loop correctly, split work across agents without losing information, and decide when the model may choose freely versus when code must force the order.

TASK 1.1

Agentic loops for autonomous task execution

An agentic loop is plain control flow around the Messages API. You send the conversation plus tool definitions. Claude either asks for tools or finishes. Your code decides what to do by reading stop_reason, never by reading Claude's prose.

THE LOOP 1 · Call Messages APIhistory + tool definitions 2 · Inspect stop_reasonnever parse the text Done · return answerstop_reason = "end_turn" 3 · Run every tool_useblock in the response 4 · Append to historyassistant turn + tool_result end_turn tool_use next iteration
MINIMAL LOOP (PYTHON, CLAUDE API)
messages = [{"role": "user", "content": task}]
while True:
    resp = client.messages.create(model=MODEL, max_tokens=4096,
                                  system=SYSTEM, tools=TOOLS, messages=messages)
    messages.append({"role": "assistant", "content": resp.content})  # keep the full turn

    if resp.stop_reason == "end_turn":
        break                                                    # the model decided it is finished
    if resp.stop_reason == "tool_use":
        results = []
        for block in resp.content:
            if block.type == "tool_use":
                out = run_tool(block.name, block.input)
                results.append({"type": "tool_result", "tool_use_id": block.id,
                                "content": out.text, "is_error": out.failed})
        messages.append({"role": "user", "content": results})    # results go back as a user turn
        continue
    handle_other(resp)   # e.g. max_tokens = output truncated; decide to continue or fail

Key ideas

  • Tool results are appended to history. That is how the model "sees" what happened and reasons about the next action. Each tool_result must reference the matching tool_use_id.
  • Model-driven vs pre-configured. In an agent, Claude decides which tool to call next from context. A hard-coded decision tree or fixed tool sequence is a workflow, not an agent. Use the agent when the path depends on what is discovered; use fixed code when the path is known.
  • A single response can contain text and several tool_use blocks. Execute all of them before the next call.
Anti-patterns the exam lists
  • Stopping when the text says "I'm done" or "Task complete" (parsing natural language for termination).
  • Using an iteration cap as the primary stop condition. A cap is fine as a safety net, not as the design.
  • Treating "the response contains text" as completion. Claude often writes text alongside a tool call.
Scenario

Your support agent sometimes ends the conversation mid-investigation. The loop exits whenever resp.content[0].type == "text". Claude had written "Let me check your order" and then requested lookup_order. Fix: branch on stop_reason only.

TASK 1.2

Multi-agent systems with coordinator–subagent patterns

Hub-and-spoke: one coordinator owns decomposition, delegation, routing between subagents, error handling, and aggregation. Subagents never talk to each other directly. Every message passes through the hub, which gives you one place for observability, consistent error handling and controlled information flow.

REFERENCE DESIGN · RESEARCH SYSTEM (SCENARIO 3)
Coordinatoranalyse query → pick subagents → partition scope → aggregate → check gaps → re-delegate
Task calls ↓  ·  structured results ↑  ·  no spoke-to-spoke links
Web searchtools: web_search, load_document
Document analysistools: extract_data_points, summarize
Synthesistools: verify_fact (scoped)
Reporttools: write_report

Key ideas

  • Isolated context. A subagent starts with only what its prompt gives it. It does not inherit the coordinator's history and keeps no memory between invocations.
  • Dynamic selection. A good coordinator reads the query and invokes only the subagents it needs. A simple factual question should not run the full five-stage pipeline.
  • Partition scope to avoid duplicate work: give each subagent distinct subtopics or source types.
  • Iterative refinement. The coordinator evaluates the synthesis for gaps, sends targeted follow-up queries to search and analysis, and re-runs synthesis until coverage is sufficient.
Trap: narrow decomposition

"Impact of AI on creative industries" split into digital art, graphic design and photography. Every subagent succeeds, yet music, writing and film are missing. The root cause is the coordinator's decomposition, not search queries or synthesis instructions. Blame the component whose output was wrong, which here is the task list itself.

TASK 1.3

Subagent invocation, context passing and spawning

  • The Task tool spawns subagents in the Agent SDK. The coordinator's allowedTools must include "Task" or it cannot delegate at all.
  • An AgentDefinition configures each subagent type: a description (used to decide when to delegate), a system prompt, and a restricted tools list.
  • Context must be explicit. Pass the complete findings of earlier agents inside the subagent's prompt, for example web results and document analysis handed to synthesis.
  • Keep metadata separate from content using structured formats (claim, excerpt, URL, document name, page) so attribution survives the hand-off.
  • Parallelism happens when the coordinator emits several Task calls in one response. Calls spread across separate turns run sequentially.
  • Prompt the coordinator with goals and quality criteria ("cover all major sectors, two independent sources per claim") rather than step-by-step procedures, so subagents can adapt.
  • Fork-based sessions let you branch from a shared analysis baseline to explore divergent approaches (see 1.7).
AGENT SDK SHAPE (PYTHON, ILLUSTRATIVE)
options = ClaudeAgentOptions(
    system_prompt=COORDINATOR_GOALS,              # goals + quality bar, not a procedure
    allowed_tools=["Task"],                        # required to spawn subagents
    agents={
        "web-searcher": AgentDefinition(
            description="Finds and fetches sources for one assigned subtopic.",
            prompt="Return findings as JSON: claim, excerpt, url, published_at.",
            tools=["web_search", "load_document"]),
        "synthesizer": AgentDefinition(
            description="Merges structured findings into a cited draft.",
            prompt="Preserve every claim→source mapping. Flag conflicts.",
            tools=["verify_fact"]),
    })
Scenario

The synthesis agent writes vague, uncited text. Its prompt says "synthesize the research so far." It never received the research, because subagents do not share memory. Fix: put the structured findings, with their source metadata, into the synthesis prompt.

TASK 1.4

Multi-step workflows with enforcement and handoff

Two ways to make steps happen in order: prompt guidance (probabilistic, usually works) and programmatic enforcement (hooks, prerequisite gates; deterministic). When a mistake costs money, breaks compliance or touches identity, prompts alone are not enough.

PREREQUISITE GATE · SUPPORT AGENT
get_customerreturns verified customer_id
→unlocks
lookup_orderscoped to that customer
→unlocks
process_refund≤ $500 else escalate

A gate in code checks session state before executing the tool. If verified_customer_id is missing, it returns an error that tells the model to verify first. The model cannot skip the step, even when the customer volunteers an order number.

Multi-concern requests

"My order arrived damaged, I was double-charged, and I want to change my email." Decompose into distinct items, investigate each in parallel against shared context (same customer, same session), then synthesize one unified reply.

Structured handoff to a human

The human agent does not see the transcript. The escalation payload must stand alone:

{
  "customer_id": "C-88213",
  "issue_summary": "Charged twice for order #12345 on 2026-09-28",
  "root_cause": "Payment retry created a duplicate capture",
  "evidence": ["capture_id pi_91a", "capture_id pi_91b"],
  "refund_amount": 742.10,
  "why_escalated": "Exceeds $500 autonomous refund limit",
  "recommended_action": "Void duplicate capture pi_91b",
  "customer_expectation": "Refund before Friday"
}
Trap

Sample Q1: the agent skips get_customer in 12% of cases. Stronger system-prompt wording and more few-shot examples are both still probabilistic. A tool-routing classifier changes which tools exist, not their order. Only a programmatic prerequisite fixes ordering.

TASK 1.5

Agent SDK hooks for interception and normalization

Hook pointInterceptsTypical use
Before a tool runs (PreToolUse)The outgoing tool call and its argumentsBlock policy violations (refund > $500), require prerequisites, redirect to escalate_to_human
After a tool runs (PostToolUse)The tool result, before the model reads itNormalize formats (Unix epoch, ISO 8601, numeric status codes → one schema), trim fields, redact
# Illustrative hook logic
async def pre_tool(call, ctx):
    if call.name == "process_refund" and call.input["amount"] > 500:
        return deny("Refunds over $500 need a human. Call escalate_to_human "
                    "with a structured summary.")          # redirect, not just block

async def post_tool(call, result, ctx):
    for k in ("created_at", "shipped_at"):
        result[k] = to_iso8601(result.get(k))               # 1727049600 → 2024-09-23T00:00:00Z
    result["status"] = STATUS_NAMES.get(result["status"], result["status"])  # 3 → "shipped"
    return result
The one-line rule

Business rule must always hold → hook. Preference or style → prompt. Hooks give deterministic guarantees; prompts give probabilistic compliance.

Scenario

Three MCP tools return dates as 1727049600, "2024-09-23T00:00:00Z" and "09/23/24". The agent miscomputes return windows. Adding "be careful with date formats" to the prompt is weak. A PostToolUse hook that normalizes every timestamp to ISO 8601 removes the problem at the source.

TASK 1.6

Task decomposition strategies

Prompt chaining (fixed pipeline)Dynamic adaptive decomposition
ShapeKnown sequence of focused stepsSubtasks generated from what each step discovers
Use forPredictable, multi-aspect work: code review, document extractionOpen-ended investigation: "add comprehensive tests to a legacy codebase", root-cause hunts
ExamplePer-file local review → cross-file integration passMap structure → find high-impact areas → prioritized plan → revise as dependencies appear
StrengthConsistent, testable, cheapAdapts to surprises
CHAINED REVIEW FOR A 14-FILE PR
file 1 · local pass
file 2 · local pass
… file 14
→
Integration passcross-file data flow, contracts, call sites
→
Merge & dedupeone consistent finding list

A single pass over many files causes attention dilution: deep comments on some files, shallow ones on others, missed bugs, and contradictory verdicts on identical code.

TASK 1.7

Session state, resumption and forking

  • --resume <session-name> continues a specific named conversation, for example an investigation spread across days.
  • fork_session creates an independent branch from a shared baseline. Analyse the codebase once, then fork to compare two refactoring or testing strategies without them contaminating each other.
  • When resuming after code changed, tell the agent which files changed so it re-reads those specifically instead of re-exploring everything or trusting stale results.
  • When most earlier tool results are stale, start a new session with a structured summary. That is more reliable than resuming a history full of outdated file contents.
SituationChoose
Prior context is still mostly validResume (--resume name)
A few files changed sinceResume and state exactly which files changed
Large refactor landed; old reads are wrongFresh session + injected summary of key findings
Compare two approaches from the same analysisFork the session
D2

Tool Design & MCP Integration

18% · primary in scenarios 1, 3 and 4

The model chooses tools by reading their names and descriptions. Most tool problems on the exam are fixed by better interfaces: clearer descriptions, narrower tools, structured errors and scoped access.

TASK 2.1

Tool interfaces with clear descriptions and boundaries

The description is the primary mechanism the model uses to select a tool. Minimal or overlapping descriptions cause misrouting between similar tools.

BEFORE → AFTER
// before: the model cannot tell these apart
get_customer:  "Retrieves customer information"
lookup_order:  "Retrieves order details"

// after
lookup_order:
  "Look up a single order by order number (format: #12345 or ORD-12345).
   Use when the user mentions an order, shipment, delivery, return or refund
   for a specific purchase. Returns status, items, amounts, ship/delivery dates.
   Do NOT use to find who a customer is; use get_customer for identity,
   contact details or account status. If only a name is given, call
   get_customer first."

A good description states

  • What it does and what it returns
  • Input formats, with an example query
  • Edge cases and limits
  • When to use it versus its nearest neighbour

Fixes when tools overlap

  • Rename + rescope: analyze_content → extract_web_results with a web-specific description.
  • Split generic tools into purpose-specific ones with defined contracts: analyze_document → extract_data_points, summarize_content, verify_claim_against_source.
  • Audit the system prompt for keyword-sensitive wording. "Always analyze the content first" can pull the model toward a tool named analyze_content regardless of its description.
Trap

Sample Q2 asks for the most effective first step. Better descriptions win. Few-shot examples add tokens without fixing the root cause. A keyword routing layer is over-engineered. Merging into one lookup_entity tool is a valid design but too large a first step.

TASK 2.2

Structured error responses for MCP tools

MCP tools signal failure with isError: true on the result. The content should tell the agent what kind of failure it is, because the right recovery differs.

CategoryExampleRetry?Agent should
TransientTimeout, 503YesRetry with backoff, locally
ValidationBad order-number formatAfter fixing inputCorrect the input or ask the user
BusinessOutside 30-day return windowNoExplain the policy kindly; offer alternatives or escalate
PermissionNot authorized for this accountNoStop; escalate or request access
{
  "isError": true,
  "content": [{ "type": "text", "text": JSON.stringify({
    "errorCategory": "business",
    "isRetryable": false,
    "code": "RETURN_WINDOW_EXPIRED",
    "message": "Order #12345 was delivered 41 days ago; returns close at 30.",
    "customerMessage": "This order is past our 30-day return window, but I can check warranty options."
  })}]
}
  • A uniform "Operation failed" leaves the agent guessing, so it retries things that will never succeed or gives up on things that would.
  • Empty is not failure. "Query succeeded, zero matches" and "could not reach the database" must look different.
  • Subagents retry transient errors themselves and only propagate what they cannot resolve, with partial results and a list of what was tried.
TASK 2.3

Distributing tools across agents and configuring tool_choice

  • Fewer tools, better choices. 18 tools on one agent degrades selection compared with 4–5 focused ones.
  • Agents given tools outside their role misuse them. A synthesis agent with web search starts searching instead of synthesizing.
  • Constrained alternatives: replace a generic fetch_url with load_document that validates document URLs.
  • Scoped cross-role tools: give synthesis a narrow verify_fact for the common simple check; send deep verification through the coordinator.
tool_choiceBehaviourUse when
{"type":"auto"}Model may call a tool or reply in textNormal conversational agents
{"type":"any"}Must call some tool, model picks whichYou need structured output and several schemas could apply (document type unknown)
{"type":"tool","name":"extract_metadata"}Must call this exact toolA specific step must run first; do later steps in follow-up turns
Scenario (sample Q9)

Synthesis bounces to the coordinator for every fact check, adding 2–3 round trips and 40% latency. 85% are simple lookups. Give synthesis a scoped verify_fact tool and keep the 15% complex checks on the coordinator path. Full web-search access over-provisions; batching the checks creates blocking dependencies; speculative caching cannot predict needs.

TASK 2.4

Integrating MCP servers into Claude Code and agents

ScopeFileShared via git?Use for
Project.mcp.json in repo rootYesTeam tooling: GitHub, Jira, internal APIs
User~/.claude.jsonNoPersonal or experimental servers
.mcp.json WITH ENV-VAR EXPANSION
{
  "mcpServers": {
    "github": {
      "command": "npx",
      "args": ["-y", "@modelcontextprotocol/server-github"],
      "env": { "GITHUB_TOKEN": "${GITHUB_TOKEN}" }   // secret stays out of git
    },
    "orders-api": {
      "type": "http",
      "url": "https://mcp.internal.example.com/orders",
      "headers": { "Authorization": "Bearer ${ORDERS_API_TOKEN}" }
    }
  }
}
  • Tools from all configured servers are discovered at connection time and available at once. A project server and a personal server work side by side.
  • Tools vs resources. Tools perform actions. Resources expose readable content catalogs (issue summaries, doc hierarchies, database schemas) so the agent can see what exists without exploratory tool calls.
  • If the agent keeps using built-in Grep instead of your more capable MCP search, improve the MCP tool's description of what it can do and return.
  • Prefer an existing community server for standard systems (Jira, GitHub). Build custom servers only for team-specific workflows.
TASK 2.5

Choosing built-in tools: Read, Write, Edit, Bash, Grep, Glob

I need to…ToolExample
Find text inside filesGrepAll callers of processRefund, an error string, an import
Find files by name or pathGlob**/*.test.tsx, src/**/migrations/*.sql
See a whole fileReadFollow an import to understand a flow
Change a specific snippetEditRequires the old text to match uniquely
Create or fully replace a fileWriteNew file, or fallback when Edit's anchor is not unique
Run commandsBashTests, builds, git
  • Edit fails on non-unique matches. Fallback: Read the file, then Write the full new content.
  • Explore incrementally: Grep for entry points, Read to follow imports, trace the flow. Do not read every file up front.
  • Wrapper modules: list everything a module exports first, then Grep for each exported name to find all usages, including re-exports under different names.
D3

Claude Code Configuration & Workflows

20% · primary in scenarios 2, 4 and 5

Most questions here are "where does this file go" and "which mechanism fits." Learn the file layout and the reason each mechanism exists: always-loaded context, conditional context, on-demand workflows, or headless automation.

FILE MAP · WHAT IS SHARED AND WHAT IS PERSONAL
~/.claude/                       personal · never in git
  CLAUDE.md                      your instructions, every project
  commands/*.md                  personal slash commands
  skills/<name>/SKILL.md          personal skills (use distinct names)
~/.claude.json                   user-scoped MCP servers

repo/                            shared · committed
  CLAUDE.md  (or .claude/CLAUDE.md)  project standards, always loaded
  .mcp.json                      team MCP servers, ${ENV} expansion
  .claude/
    rules/*.md                   topic rules; optional paths: globs
    commands/*.md                team slash commands  → /review
    skills/<name>/SKILL.md        team skills with frontmatter
  packages/api/CLAUDE.md         directory-level, loads for that subtree
TASK 3.1

CLAUDE.md hierarchy, scoping and modular organization

  • Three levels: user (~/.claude/CLAUDE.md), project (CLAUDE.md at root or .claude/CLAUDE.md), directory (a CLAUDE.md in a subfolder).
  • User-level is private. Instructions there are not in version control, so teammates never get them.
  • @import pulls other files into a CLAUDE.md, e.g. @docs/standards/testing.md. Each package can import only the standards that apply to it.
  • .claude/rules/ replaces one huge CLAUDE.md with focused files: testing.md, api-conventions.md, deployment.md.
  • /memory shows which memory files are loaded. Use it first when behaviour differs between people or sessions.
# packages/billing/CLAUDE.md
Billing service. Money is integer cents; never floats.
@../../docs/standards/api-errors.md
@../../docs/standards/pci-logging.md
Scenario

A new hire's Claude ignores the team's naming convention that works for everyone else. /memory shows the convention lives only in a senior engineer's ~/.claude/CLAUDE.md. Move it to the project CLAUDE.md (or a rules file) and commit it.

TASK 3.2

Custom slash commands and skills

Project (shared)User (personal)
Commands.claude/commands/~/.claude/commands/
Skills.claude/skills/<name>/SKILL.md~/.claude/skills/<name>/SKILL.md
SKILL.md FRONTMATTER
---
name: analyze-module
description: Deep-dive one module and report structure, hotspots and risks.
context: fork                     # run in an isolated sub-agent; only the summary returns
allowed-tools: Read, Grep, Glob   # restrict what the skill may do
argument-hint: <module-path>      # prompts the dev when invoked with no argument
---
Map the module at $ARGUMENTS. Return: entry points, dependencies,
untested paths, and the three riskiest functions with reasons.
  • context: fork keeps verbose output (codebase analysis) or exploratory noise (brainstorming) out of the main conversation.
  • allowed-tools limits blast radius, for example a skill that may only write files and never run shell commands.
  • To customise a team skill for yourself, create a personal copy with a different name in ~/.claude/skills/ so teammates are not affected.
Skill or CLAUDE.md?

CLAUDE.md = always loaded, universal standards. Skill = on-demand, task-specific workflow. A rule every edit should follow does not belong in a skill; a 40-line release checklist does not belong in CLAUDE.md.

TASK 3.3

Path-specific rules for conditional loading

# .claude/rules/testing.md
---
paths: ["**/*.test.tsx", "**/*.test.ts"]
---
Use React Testing Library. Query by role, not test id.
One behaviour per test. Use fixtures from test/fixtures/.

# .claude/rules/terraform.md
---
paths: ["terraform/**/*"]
---
Every resource gets owner and cost-center tags. No inline IAM policies.
  • Rules with a paths glob load only when Claude edits matching files, saving tokens and avoiding irrelevant guidance.
  • Glob rules beat subdirectory CLAUDE.md files for conventions that span directories, such as test files living next to their components everywhere.
Trap (sample Q6)

Tests sit beside components across the tree. Subdirectory CLAUDE.md files are bound to directories; one big root CLAUDE.md relies on inference; skills need invocation. Path-scoped rules are the deterministic, maintainable answer.

TASK 3.4

Plan mode vs direct execution

Signal in the taskMode
Architectural decision, service boundaries, monolith → microservicesPlan mode
Library migration touching 45+ filesPlan mode
Several valid approaches with different infra needsPlan mode
Single-file bug with a clear stack traceDirect execution
Add one date-validation conditionalDirect execution
  • Plan mode allows safe, read-only exploration and design before any change, which prevents expensive rework.
  • Combine them: plan the migration, approve the plan, then execute it directly.
  • The Explore subagent isolates verbose discovery and returns a summary, keeping the main context from filling up in multi-phase work.
Trap (sample Q5)

"Start direct and switch to plan mode if it gets complex" ignores that the complexity is stated in the requirements. "Let the implementation reveal the boundaries" risks late rework.

TASK 3.5

Iterative refinement techniques

When prose is read inconsistently

Input/output examples

Give 2–3 concrete before/after pairs. Examples communicate a transformation better than descriptions.

When correctness is checkable

Test-driven iteration

Write tests first (behaviour, edge cases, performance), then iterate by pasting the failing output.

In unfamiliar domains

Interview pattern

Ask Claude to question you first: cache invalidation, failure modes, scale. It surfaces decisions you missed.

Batching feedback

Interacting vs independent

Fixes that affect each other go in one detailed message. Independent issues can be fixed one by one.

Scenario

A migration script mishandles nulls. Instead of "handle nulls properly," give a test case: input {"discount": null} → expected {"discount_cents": 0}.

TASK 3.6

Claude Code in CI/CD pipelines

REFERENCE DESIGN · PR REVIEW JOB (SCENARIO 5)
PR opened / pusheddiff + changed files
→
claude -pindependent instance · CLAUDE.md gives criteria
→
JSON findings--output-format json + --json-schema
→
Post inline commentsdedupe vs prior findings
claude -p "Review this PR diff. Report only bugs and security issues per CLAUDE.md.
Previously reported findings are in prior.json; report only new or still-unaddressed ones." \
  --output-format json \
  --json-schema '{"type":"object","properties":{"findings":{"type":"array","items":{
     "type":"object","required":["file","line","severity","issue","fix"],
     "properties":{"file":{"type":"string"},"line":{"type":"integer"},
       "severity":{"enum":["critical","major","minor"]},
       "issue":{"type":"string"},"fix":{"type":"string"}}}}}}'
  • -p / --print = non-interactive: run, print, exit. Without it the job hangs waiting for input.
  • CLAUDE.md is the CI context: review criteria, testing standards, fixture conventions.
  • Session isolation: the session that wrote the code is a weak reviewer of it. Review in a separate instance.
  • On re-runs, include prior findings and ask for new or unresolved issues only, to avoid duplicate comments.
  • For test generation, include existing test files so Claude does not propose scenarios already covered, and document valuable-test criteria and fixtures in CLAUDE.md.
D4

Prompt Engineering & Structured Output

20% · primary in scenarios 5 and 6

Three themes: say exactly what counts (criteria and examples), make the shape guaranteed (tool use with schemas), and check the meaning separately (validation, retries, independent review).

TASK 4.1

Explicit criteria to raise precision and cut false positives

Vague (does not work)Explicit (works)
"Check that comments are accurate.""Flag a comment only when the behaviour it claims contradicts what the code does."
"Be conservative." / "Only high-confidence findings.""Report: null derefs, injection, auth bypass, race conditions. Skip: naming, formatting, patterns used consistently in this repo."
"Rate severity appropriately.""Critical = data loss or security exposure, e.g. db.query(`…${req.query.id}`). Minor = …"
  • False positives cost more than missed nits. One noisy category makes developers distrust the accurate ones too.
  • Tactic: temporarily disable the high-false-positive category to restore trust while you improve its prompt.
  • Define severity levels with a concrete code example for each, so classification is consistent.
TASK 4.2

Few-shot prompting for consistency and quality

When detailed instructions still give inconsistent results, few-shot examples are the most effective lever. They teach format and judgment, and they help the model generalise to cases you did not list.

  • Use 2–4 targeted examples aimed at the ambiguous cases, each showing why one action beat the plausible alternative.
  • Show the exact output format: location, issue, severity, suggested fix.
  • Include an example of acceptable code that looks suspicious, to cut false positives.
  • For extraction, show varied document structures (inline citations vs bibliography, tables vs narrative). This reduces hallucination and empty fields.
<example>
Request: "I want to return the blender but I also think I was charged twice."
Decision: lookup_order first, then check payments.
Why: both concerns hinge on order #; get_customer is unnecessary because the
session already holds a verified customer_id. Escalation not needed: both are
within policy and tools.
</example>
TASK 4.3

Structured output with tool use and JSON schemas

Define an "extraction tool" whose input_schema is the shape you want. When Claude calls it, the tool_use.input is your structured data. This is the most reliable way to get schema-compliant JSON and removes syntax errors.

{
  "name": "extract_invoice",
  "description": "Record fields from one invoice. Use null when a field is absent.",
  "input_schema": {
    "type": "object",
    "required": ["vendor_name", "line_items", "stated_total", "document_type"],
    "properties": {
      "vendor_name":   { "type": "string" },
      "invoice_date":  { "type": ["string", "null"], "description": "ISO 8601; null if absent" },
      "po_number":     { "type": ["string", "null"] },
      "document_type": { "enum": ["invoice", "credit_note", "receipt", "unclear", "other"] },
      "document_type_detail": { "type": ["string", "null"], "description": "Fill when type is other" },
      "line_items": { "type": "array", "items": { "type": "object",
          "properties": { "desc": {"type":"string"}, "amount_cents": {"type":"integer"} } } },
      "stated_total":     { "type": "integer" },
      "calculated_total": { "type": "integer", "description": "Sum of line_items" },
      "conflict_detected":{ "type": "boolean" }
    }
  }
}
  • Nullable fields stop the model inventing values to satisfy a required field.
  • "unclear" handles ambiguity; "other" + detail string keeps categories extensible.
  • Put normalization rules in the prompt (dates to ISO, amounts to cents) alongside the strict schema.
  • tool_choice: "any" when several extraction schemas exist and the document type is unknown; forced {"type":"tool","name":"extract_metadata"} to run one step before enrichment.
Syntax vs semantics

Schemas guarantee valid JSON of the right shape. They do not guarantee that line items sum to the total or that a value landed in the right field. That needs validation (4.4).

TASK 4.4

Validation, retry and feedback loops

VALIDATION–RETRY LOOP
Extracttool_use + schema
→
ValidatePydantic / business rules
→fail
Retry with feedbackoriginal doc + failed output + exact errors
→still fails
Human reviewinfo likely absent
FailureWill a retry help?
Date in wrong format, wrong field placement, structural mistakesYes
Line items do not sum to stated total because the model misread a rowUsually
Value exists only in an external document that was not providedNo; route elsewhere
  • Self-checking fields: extract calculated_total next to stated_total; add conflict_detected when the source contradicts itself.
  • Feedback loops for reviews: add a detected_pattern field to each finding so you can analyse which patterns developers dismiss and tune the prompt.
TASK 4.5

Batch processing strategies

Message Batches API

What you get

  • 50% lower cost
  • Up to 24 h processing, no latency SLA
  • custom_id correlates request ↔ result
  • Poll for completion
Limits

What you do not get

  • No multi-turn tool calling inside a request
  • No guaranteed order or speed
  • Unfit for anything a person is waiting on
WorkloadAPI
Blocking pre-merge checkSynchronous
Overnight tech-debt report, weekly audit, nightly test generationBatch

SLA arithmetic

Worst case = waiting for the next submission + up to 24 h processing + a buffer for validation and resubmits. For a 30 h SLA: 4 h windows → 4 + 24 = 28 h, leaving 2 h of slack. A 6 h window would hit 30 h with zero slack.

Failure handling

  • Resubmit only failed items, identified by custom_id, with a fix (e.g. chunk documents that exceeded the context limit).
  • Refine the prompt on a sample set first to maximise first-pass success before sending 10,000 documents.
TASK 4.6

Multi-instance and multi-pass review

  • Self-review is weak. The generating session keeps its reasoning and tends to defend it. Instructions to "double-check" and extended thinking do not fix that.
  • Independent instance without the generator's context catches subtler issues.
  • Multi-pass: per-file local passes plus an integration pass for cross-file data flow (see 1.6).
  • Confidence per finding in a verification pass can route review attention, once calibrated.
Trap (sample Q12)

A bigger context window does not fix attention quality. Requiring 2-of-3 agreement suppresses real bugs that only appear intermittently. Making developers split PRs shifts the burden without improving the system.

D5

Context Management & Reliability

15% · appears in scenarios 1, 2, 3 and 6

Long sessions lose detail, tool output crowds out what matters, and multi-agent systems drop sources and errors at every hand-off. This domain is about keeping facts, failures and provenance intact.

TASK 5.1

Preserving critical information across long interactions

Three ways context goes wrong

  • Progressive summarization blurs numbers, dates, percentages and what the customer was promised ("refund $742.10 by Friday" becomes "customer wants a refund").
  • Lost in the middle: models handle the start and end of long inputs well and can miss findings in the middle.
  • Tool-result bloat: an order lookup returns 40+ fields when 5 matter, and every call adds more.
PROMPT LAYOUT FOR A LONG SUPPORT SESSION
System promptrole, policies, escalation criteria with examples
Case facts block · never summarizedcustomer_id C-88213 · order #12345 · $742.10 · delivered 2026-09-12 · promised: refund by Fri
Summary of older turnsnarrative only; the facts live above
Recent turns, full fidelitywith trimmed tool results (only return-relevant fields)
  • Extract transactional facts into a persistent case facts block sent with every request, outside the summarized history. For multi-issue sessions keep one structured entry per issue.
  • Trim tool output to relevant fields before it enters context (a PostToolUse hook is a good place).
  • For aggregated inputs, put a key-findings summary at the top and organise details under explicit section headers.
  • Ask upstream agents for structured data (key facts, citations, relevance scores, dates) instead of long prose and reasoning chains when downstream budgets are tight.
  • The API is stateless: send the complete conversation history each request to stay coherent.
TASK 5.2

Escalation and ambiguity resolution

SituationCorrect behaviour
Customer explicitly asks for a humanEscalate immediately, no investigation first
Customer is frustrated, issue is simple and in scopeAcknowledge, offer to resolve; escalate if they repeat the request
Policy is silent or ambiguous (competitor price match when policy covers only own-site adjustments)Escalate: a policy gap, not just a hard case
Needs a policy exceptionEscalate
No meaningful progress possibleEscalate with a structured handoff
Lookup returns several matching customersAsk for another identifier (email, order #, postcode); never guess
Standard damage replacement with photo evidenceResolve autonomously
Unreliable escalation signals
  • Sentiment measures mood, not case complexity.
  • Self-reported confidence (1–10) is poorly calibrated; the agent is often confidently wrong on the hard cases.
  • A separate classifier is over-engineered before prompt criteria have been tried.
The fix pattern (sample Q3)

Explicit escalation criteria in the system prompt plus few-shot examples of escalate vs resolve. It targets the real cause, unclear decision boundaries, at proportionate cost.

TASK 5.3

Error propagation across multi-agent systems

Subagent returns…VerdictWhy
Structured error: type, attempted query, partial results, alternativesRightCoordinator can retry differently, try another source, or proceed and annotate
"search unavailable" after silent retriesWrongGeneric status hides what was tried and what partially worked
Empty result marked successWrongSuppresses the failure; report looks complete but is not
Exception that ends the whole workflowWrongOne failure should not discard all the other work
{
  "status": "partial_failure",
  "failure_type": "timeout",
  "attempted": { "query": "AI music generation royalties 2025", "source": "web_search", "retries": 2 },
  "partial_results": [ { "claim": "...", "url": "...", "published_at": "2025-11-03" } ],
  "alternatives": ["query trade-press archive", "narrow to EU market"],
  "coverage_gap": "Music licensing economics"
}

The final report should carry coverage annotations: which findings are well supported and which areas have gaps because sources were unavailable.

TASK 5.4

Context in large codebase exploration

Symptom of degradation: late in a long session the model answers inconsistently and talks about "typical patterns" instead of the specific classes it found earlier.

  • Scratchpad files: the agent writes key findings to a file (e.g. notes/refund-flow.md) and re-reads it for later questions.
  • Delegate to subagents for focused questions ("find all test files", "trace refund flow dependencies"); the main agent keeps the high-level picture.
  • Summarize between phases and inject that summary into the next phase's subagents.
  • /compact shrinks context when it fills with verbose discovery output.
  • Crash recovery: each agent exports its state to a known location; on resume the coordinator loads a manifest and injects the state into agent prompts.
// state/manifest.json
{ "run_id": "audit-2026-10-04", "phase": "analysis",
  "agents": {
    "mapper":   { "status": "done",    "state": "state/mapper.json" },
    "security": { "status": "running", "state": "state/security.json", "last_file": "src/auth/session.ts" },
    "tests":    { "status": "pending" } } }
TASK 5.5

Human review workflows and confidence calibration

REVIEW ROUTING · EXTRACTION PIPELINE (SCENARIO 6)
Extraction+ field-level confidence
→
Routerthresholds calibrated on a labeled validation set
→
Human queuelow confidence, contradictory or ambiguous source
Auto-accept+ stratified random sample audited
  • Aggregate accuracy hides weak segments. 97% overall can mean 99% on invoices and 70% on handwritten receipts.
  • Before reducing human review, measure accuracy by document type and by field.
  • Stratified random sampling of high-confidence output keeps measuring the real error rate and catches new error patterns.
  • Field-level confidence is only useful after calibration against labeled data. Spend scarce reviewer time on low-confidence or contradictory cases.
TASK 5.6

Provenance and uncertainty in multi-source synthesis

{ "claim": "Generative tools used by 38% of indie musicians",
  "excerpt": "...38 percent of surveyed independent artists...",
  "source": { "url": "https://example.org/survey", "title": "Indie Artist Survey" },
  "published_at": "2025-06-10", "collected": "2025-Q1",
  "method": "online survey, n=1,204",
  "conflicts_with": [{ "value": "21%", "source": "Label Assoc. report", "published_at": "2024-02-01" }] }
  • Summaries lose attribution unless each claim keeps its claim → source mapping; synthesis must preserve and merge those mappings.
  • Conflicting credible numbers: report both with attribution. Do not pick one arbitrarily. Let the coordinator decide how to reconcile.
  • Dates prevent false conflicts: 21% in 2024 and 38% in 2025 may be a trend, not a contradiction.
  • Structure the report into well-established vs contested findings, keeping methodological context.
  • Render content by type: financial data as tables, news as prose, technical findings as lists.
S

The six scenarios, as system designs

You will get four. For each, know the reference architecture and the decisions the questions are likely to probe.
Scenario 1

Customer support resolution agent

D1D2D5

Agent SDK loop + MCP tools get_customer, lookup_order, process_refund, escalate_to_human. Target ≥ 80% first-contact resolution.

  • Gate: refund tools blocked until a verified customer ID exists
  • PreToolUse hook: refund > $500 → escalate
  • PostToolUse: normalize dates, trim order fields
  • Case-facts block; structured handoff
  • Explicit escalation criteria + few-shot
  • Multiple matches → ask for identifiers
Scenario 2

Code generation with Claude Code

D3D5

Team use for generation, refactoring, debugging, docs.

  • Project CLAUDE.md + @import; rules with paths globs
  • Team commands in .claude/commands/
  • Skills with context: fork, allowed-tools
  • Plan mode for architecture; direct for scoped fixes
  • Explore subagent, /compact, scratchpads
  • I/O examples, TDD, interview pattern
Scenario 3

Multi-agent research system

D1D2D5

Coordinator → search, analysis, synthesis, report subagents; cited output.

  • allowedTools includes Task; parallel Task calls in one turn
  • Broad decomposition, partitioned scope
  • Findings passed explicitly, as claim/source/date records
  • Scoped verify_fact for synthesis
  • Structured partial-failure errors; coverage gaps in report
  • Conflicts kept with attribution
Scenario 4

Developer productivity agent

D2D3D1

Agent SDK with Read, Write, Bash, Grep, Glob + MCP servers; explores legacy code, generates boilerplate.

  • Grep vs Glob; Edit then Read+Write fallback
  • Incremental exploration from entry points
  • MCP descriptions strong enough to beat Grep
  • Resources for catalogs (schemas, docs)
  • Dynamic decomposition for "add tests to legacy code"
  • Resume / fork sessions; manifests for recovery
Scenario 5

Claude Code in CI/CD

D3D4

Automated reviews, test generation, PR feedback with few false positives.

  • claude -p + --output-format json + --json-schema
  • Independent reviewer instance
  • Explicit report/skip criteria; severity examples
  • Per-file + integration passes
  • Prior findings in context to avoid duplicates
  • Sync for pre-merge; batch for nightly
Scenario 6

Structured data extraction

D4D5

Unstructured docs → schema-validated JSON → downstream systems.

  • Extraction tool schema; nullable fields; enum + unclear/other
  • tool_choice any / forced
  • Validation-retry with specific errors; know when retry is futile
  • calculated vs stated totals; conflict flags
  • Batches with custom_id; SLA windows
  • Calibrated confidence, stratified sampling, per-segment accuracy
⇢

Signal in the question → likely answer

Read the stem for one of these signals, then check the option that matches.
If the question says…Lean toward…
"must always", "financial", "compliance", "X% of the time it skips"Programmatic gate or hook, not prompt text
"most effective first step"Smallest change at the root cause (descriptions, criteria, examples)
Loop ends early / runs foreverBranch on stop_reason; append tool_result blocks
Subagent output vague, missing earlier infoPass findings explicitly in its prompt
Coordinator cannot spawn subagents"Task" missing from allowedTools
Subagents run one after anotherEmit multiple Task calls in one response
Report misses whole subtopics, subagents fineCoordinator decomposition too narrow
Wrong tool picked between two similar toolsRicher, differentiated descriptions; rename/split
Agent retries hopeless calls / gives up on recoverable onesStructured errors: category + isRetryable
Agent has 15+ tools and misuses someScope tools per role; constrained alternatives
Need a tool call guaranteed, type unknowntool_choice: "any"
Need one specific step to run firstForced {"type":"tool","name":…}
Shared MCP server, secret token.mcp.json + ${ENV_VAR}
Agent prefers Grep over MCP toolImprove MCP tool description
Teammate does not get instructionsThey are in user-level config; move to project
Conventions for files spread across dirs.claude/rules/ with paths globs
Command for everyone who clones.claude/commands/
Skill output floods the conversationcontext: fork
Architecture / many files / multiple approachesPlan mode
CI job hangs waiting for input-p
Reviewer approves its own bugsIndependent instance
Big PR, inconsistent reviewPer-file passes + integration pass
Too many false positivesExplicit report/skip categories; disable noisy category for now
Inconsistent format despite instructionsFew-shot examples
Model invents missing valuesNullable fields + "use null if absent"
Retry still fails, data not in docStop retrying; route to human / fetch source
Cut cost, latency-tolerantMessage Batches API; sync for blocking
Numbers lost in long chatsPersistent case-facts block
"Customer asks for a human"Escalate immediately
Policy does not cover the requestEscalate (policy gap)
Two customers matchAsk for another identifier
Subagent timeoutStructured error with partial results; continue and annotate gaps
Model forgets earlier classes in long explorationScratchpad file, subagents, /compact
"97% accurate, can we drop review?"Segment by doc type/field; stratified sampling
Two sources disagreeKeep both with attribution and dates
?

Practice set

Original questions written in the exam's style for these notes; not taken from the exam. Try each before opening it.
D1 · Support agent · choose oneYour refund policy says amounts over $500 must go to a human. In testing, the agent follows this 98% of the time. What should you change?
  1. Move the rule to the top of the system prompt in capitals.
  2. Add a PreToolUse hook that denies process_refund above $500 and tells the agent to call escalate_to_human.
  3. Add five few-shot examples of large refunds being escalated.
  4. Lower the model temperature to zero.
Show answer ▾
B. A financial rule needs a deterministic guarantee. A, C and D all improve probabilistic compliance; 98% is still a failure in 1 of 50 cases. The hook also redirects, so the customer still gets help.
D1 · Research system · choose oneThe coordinator delegates to four subagents, but logs show each runs only after the previous one finishes, even though they are independent. What is the cause?
  1. Subagents share one context window and must queue.
  2. The coordinator issues one Task call per turn instead of several in a single response.
  3. allowedTools is missing "Task".
  4. The Agent SDK cannot run subagents in parallel.
Show answer ▾
B. Parallel execution happens when one coordinator response contains multiple Task tool calls. C would stop delegation entirely, not serialize it. A and D are false.
D1 · Sessions · choose oneYesterday you analysed a service in a named session. Overnight a teammate refactored 3 of its 40 files. You want to continue the analysis. Best approach?
  1. Resume the named session and tell Claude which 3 files changed so it re-reads them.
  2. Start a brand-new session and re-explore all 40 files.
  3. Resume and continue without mentioning the change.
  4. Fork the session twice and compare outputs.
Show answer ▾
A. Most prior context is still valid, so resume, and target re-analysis at the changed files. B wastes work; C trusts stale reads; D solves a different problem (comparing approaches).
D2 · MCP config · choose twoYour team wants everyone to use the same Jira integration in Claude Code. The API token must not be committed. Which two actions are correct?
  1. Configure the server in the project's .mcp.json using ${JIRA_TOKEN}.
  2. Ask each developer to add the server to their ~/.claude.json.
  3. Use an existing community Jira MCP server rather than writing a custom one.
  4. Put the token in .mcp.json and add the file to .gitignore.
  5. Describe the Jira API in CLAUDE.md so Claude can call it with Bash.
Show answer ▾
A and C. Project scope shares the config through git, env-var expansion keeps the secret out, and Jira is a standard integration with community servers. B is not shared; D stops sharing the config at all; E is fragile.
D2 · Errors · choose oneThe agent calls lookup_order five times in a row for an order that is past the return window, because the tool returns "Error: request failed". Best fix?
  1. Cap retries at two in the agent loop.
  2. Return a structured error with errorCategory: "business", isRetryable: false and a customer-friendly explanation.
  3. Return an empty result instead of an error.
  4. Tell the agent in the system prompt not to retry failed tools.
Show answer ▾
B. The agent cannot tell a policy decision from an outage. Structured metadata lets it stop retrying and explain the policy. A is a blunt cap; C hides the failure; D would also stop useful retries of transient errors.
D2 · Built-in tools · choose oneClaude must update a config value, but Edit fails because the line "timeout": 30 appears in four places in the file. What should it do?
  1. Use Bash with sed -i on every match.
  2. Read the full file, then Write the corrected full content.
  3. Use Glob to find a different file to edit.
  4. Retry Edit until it succeeds.
Show answer ▾
B. The documented fallback when Edit lacks a unique anchor is Read + Write. A changes all four, which may be wrong. (Including more surrounding text to make the anchor unique is also fine in practice.)
D3 · Skills · choose oneA team skill that maps a module's dependencies dumps thousands of lines into the conversation, and later answers get worse. Best fix?
  1. Add context: fork to the skill's frontmatter.
  2. Move the skill's instructions into CLAUDE.md.
  3. Run /compact after each use.
  4. Restrict the skill with allowed-tools: Read.
Show answer ▾
A. Forked context runs the skill in an isolated sub-agent and returns only its result. C treats the symptom every time; B loads it always; D limits actions, not output volume.
D3 · CI · choose oneYour nightly review job re-runs on every new commit and posts the same comments again and again. What fixes it?
  1. Switch to the Message Batches API.
  2. Include the previous findings in context and ask for only new or still-unaddressed issues.
  3. Run the review in the same session that generated the code.
  4. Raise the severity threshold so fewer comments are posted.
Show answer ▾
B. Prior findings in context let Claude de-duplicate. D hides real issues; C makes reviews worse; A changes cost, not duplication.
D4 · Extraction · choose oneInvoices without a PO number come back with plausible but fake PO numbers. The schema marks po_number as a required string. Best fix?
  1. Make po_number nullable and instruct the model to return null when absent.
  2. Add a retry when the PO number fails a regex.
  3. Lower temperature.
  4. Remove po_number from the schema.
Show answer ▾
A. A required non-null field pressures the model to fabricate. B fails because the information is absent; retries cannot add it. D loses data you need when it exists.
D4 · Batch · choose oneDocuments arrive continuously; results are due within 30 hours. Batch processing may take up to 24 hours. Which submission cadence safely meets the SLA with room for resubmitting failures?
  1. Every 12 hours.
  2. Every 8 hours.
  3. Every 4 hours.
  4. Once a day.
Show answer ▾
C. Worst case is window + 24 h. 4 + 24 = 28 h leaves 2 h of buffer. 8 + 24 = 32 h already breaks the SLA.
D5 · Escalation · choose oneA customer writes: "I just want to talk to a person." Their issue is a simple address change the agent could finish in one step. What should the agent do?
  1. Complete the address change first, then offer a human.
  2. Escalate to a human immediately, with a structured summary.
  3. Explain that it can handle this faster than a human.
  4. Check sentiment and escalate only if negative.
Show answer ▾
B. An explicit request for a human is honoured immediately, without investigating first. Offering to resolve applies when the customer is frustrated but has not asked for a person.
D5 · Human review · choose oneAn extraction pipeline is 97% accurate on 10,000 documents. Leadership wants to stop human review for high-confidence results. What should you do first?
  1. Approve it; 97% exceeds the target.
  2. Measure accuracy by document type and field, and set up stratified sampling of high-confidence output.
  3. Ask the model to self-rate confidence 1–10 and auto-accept above 8.
  4. Run every document twice and accept only matches.
Show answer ▾
B. Aggregates can hide a weak segment. Validate per segment and keep sampling to detect new error patterns. C uses uncalibrated self-reports.
✓

Hands-on labs and exam scope

The guide's four preparation exercises, condensed into checklists, plus what will not be tested.
Lab 1 · D1 D2 D5

Multi-tool agent with escalation

  • 3–4 MCP tools, two deliberately similar
  • Loop on stop_reason
  • Errors with category + isRetryable
  • Hook blocking amounts over a threshold → escalate
  • Test multi-concern messages
Lab 2 · D3 D2

Team Claude Code setup

  • Project CLAUDE.md
  • Rules with paths for API and tests
  • Skill with context: fork + allowed-tools
  • .mcp.json with env vars + a personal server
  • Compare plan mode vs direct on 3 task sizes
Lab 3 · D4 D5

Extraction pipeline

  • Schema: required, nullable, enum + other
  • Validation-retry; log retryable vs not
  • Few-shot for varied layouts
  • Batch of 100 docs; resubmit failures by custom_id
  • Field confidence → human routing; per-segment accuracy
Lab 4 · D1 D2 D5

Multi-agent research pipeline

  • Coordinator with Task; explicit context passing
  • Parallel Task calls; measure latency gain
  • Findings: claim, excerpt, source, date
  • Simulate a timeout; continue with partial results
  • Conflicting stats preserved with attribution
Out of scope (not tested)
Fine-tuning or training models · Claude internals, training, weights · Constitutional AI, RLHF
API auth, billing, account management · OAuth, key rotation · rate limits, quotas, pricing maths
Hosting MCP servers (infra, networking, containers) · cloud-provider-specific setup
Embeddings and vector databases · computer use · vision · streaming / SSE
Prompt caching internals (only know it exists) · tokenization details · benchmarks and model comparisons