Learn AI from scratch, the developer's way

Concepts explained simply, with real-world examples — so you actually understand what's happening under the hood.

⚡
AI Fundamentals — Core 10
click any term to expand
🗺️
How It All Connects

The flow below shows how raw text travels through an AI system — from tokenisation to embeddings to RAG to grounded output.

INPUT ENCODE RETRIEVE GENERATE Raw Text"Hello world" Tokeniserwords → tokens Context Windowprompt + history+ docs (tokens) KnowledgeCutoffAug 2025 Temperature0 = precise1 = creative Embeddingstext → vectors Vector DBPinecone /pgvector Semantic Searchmeaning ≠ keyword encode RAGRetrieval-Augmented Generationretrieve → inject into context → generate top-k docs Groundingcite sourcesverify facts Fine-Tuningstyle / formatnot new facts LLMTransformer + Attentionpredicts next token inject shape Hallucinationconfident butwrong output Grounded Outputanswer + citation→ user withoutRAG/ground AI AgentTool Use /Function Callplan → act → loop uses LLM RLHF → safe
📖
Complete AI Glossary
The big picture
TermSimple meaning
AIMachines that simulate human intelligence
ML (Machine Learning)AI that learns from data instead of explicit rules
Deep LearningML using neural networks with many layers
GenAIAI that generates content — text, images, code
LLMLarge Language Model — the brain behind ChatGPT, Claude
Foundation ModelA massive pre-trained model others build on top of
How LLMs work
TermSimple meaning
TokenA chunk of text (roughly 1 word ≈ 1.3 tokens)
Context WindowHow much text the model can "see" at once
PromptThe input you give the model
InferenceRunning the model to get an output
TemperatureControls randomness — 0 = predictable, 1 = creative
EmbeddingConverting text into numbers (vectors) for comparison
VectorA list of numbers representing meaning
Semantic SearchRetrieval that understands intent, not just keyword matches
Building with AI
TermSimple meaning
RAGRetrieval-Augmented Generation — give the model your own docs
Fine-tuningRe-training a model on your specific data
Prompt EngineeringCrafting inputs to get better outputs
AgentAn AI that can take actions, use tools, make decisions
Tool Use / Function CallingLetting the LLM call your code or APIs
ChainConnecting multiple AI steps together
HallucinationWhen the model confidently says something wrong
Models & training
TermSimple meaning
ParametersThe "weights" inside a model (GPT-4 ≈ 1 trillion)
TrainingTeaching a model on massive datasets
Pre-trainingInitial training on general internet data
RLHFHuman feedback used to make models more helpful and safe
TransformerThe architecture behind almost all modern LLMs
AttentionHow the model decides what parts of text matter most
💰
Anthropic API Pricing
June 2026
Pay-as-you-go — no monthly subscription. You pay only for tokens used, billed per million. 1M tokens ≈ 750,000 words.
ModelInput / 1M tokensOutput / 1M tokensBest for
Haiku 4.5$1.00$5.00Learning & cheap tasks
Sonnet 4.6$3.00$15.00Best balance — use this for Week 1
Opus 4.8$5.00$25.00Most powerful tasks
⚙️
Cost Optimization
🗂️
Batch Processing
50% cheaper on input + output. Submit thousands of requests asynchronously, processed within 24 hrs.
💾
Prompt Caching
Cache your static context (docs, instructions) and pay 90% less when it's reused across calls.
🐣
Use Haiku While Learning
Haiku 4.5 is $1/1M input — cheapest and still great for all Week 1 practice tasks.
Batch processing — when to use it
Good for batch
  • Processing thousands of documents
  • Bulk summarization
  • Data extraction / classification
  • Nightly jobs
  • Offline evaluations
  • Content generation queues
Not suitable for batch
  • Chatbots
  • Real-time APIs
  • Interactive applications
  • Anything needing a response in seconds
Prompt caching — how it works
Cost breakdown per request
Without cache
full prompt cost
100%
With caching
cached 10%
dynamic 100%
~42%
Static context (large docs, instructions) gets cached — you pay only 10% on reuse.
Dynamic context (user's question) is always charged at full rate.

Claude Code — Project Structure Guide

How the .claude/ config folder works, explained with a real example project.

runbook-bot/ ├── CLAUDE.md ├── CLAUDE.local.md ├── .mcp.json └── .claude/ ├── settings.json ├── settings.local.json ├── rules/ │ └── code-style.md ├── commands/ │ └── fix-issue.md ├── skills/ │ └── deploy/SKILL.md ├── agents/ │ └── security-auditor.md └── hooks/ └── validate-bash.sh
CLAUDE.md Loaded every session
  • The main brief Claude Code reads at the start of every session
  • Defines project overview, tech stack, build/run commands
  • Documents architecture and conventions so Claude doesn't need re-explaining each time
# Runbook Bot AI assistant that answers questions from AWS runbooks using RAG. ## Tech Stack - Python 3.11, LangChain, ChromaDB, Anthropic API - Embeddings: HuggingFace sentence-transformers (local, free) ## Commands - Run bot: `python runbook_bot.py` - Re-index docs: `python ingest.py --docs ./docs` ## Architecture - `ingest.py` — loads PDFs, chunks, embeds, stores in ./chroma_db - `runbook_bot.py` — CLI chat loop with ConversationalRetrievalChain
CLAUDE.local.md Personal, gitignored
  • Same idea as CLAUDE.md, but personal overrides not shared with the team
  • Use it for machine-specific notes — local AWS profile, test file paths, Python version
# Local notes (not shared with team) My AWS profile: `personal-sandbox` Test PDFs are in ~/Downloads/test-runbooks/ Running Python 3.11 via pyenv, not system Python
.mcp.json Shared via git
  • Stores MCP (Model Context Protocol) integration configs
  • Connects Claude Code to GitHub, Jira, Slack, databases
  • Checked into git, so the whole team gets the same integrations automatically
{ "mcpServers": { "github": { "command": "npx", "args": ["-y", "@modelcontextprotocol/server-github"] } } }
.claude/settings.json Permissions & model
  • Controls what Claude Code is allowed to do — like an IAM policy for the agent
  • Defines model selection and hooks
  • settings.local.json works the same way for personal, untracked overrides
{ "permissions": { "allow": ["Bash(pytest:*)", "Bash(python:*)"], "deny": ["Bash(rm -rf:*)"] }, "model": "claude-sonnet-4-6" }
.claude/rules/ Modular conventions
  • Topic-specific markdown files instead of one giant CLAUDE.md
  • Covers style, testing, API design — can target specific files/paths
# code-style.md - Type hints required on all functions - Use dataclasses for config objects, not raw dicts - f-strings only, never .format() or % - Max function length: 40 lines
.claude/commands/ Slash commands
  • Custom reusable workflows, triggered like /project:fix-issue
  • Good for repeatable tasks: bug fixes, releases, audits
# fix-issue.md Given a GitHub issue number, reproduce the bug, write a failing test, fix the code, and confirm the test passes. Show the diff before committing.
.claude/skills/ Auto-triggered
  • Loaded automatically only when the task context matches
  • Keeps everyday context lightweight — deploy instructions don't load unless you're deploying
# skills/deploy/SKILL.md --- name: deploy-runbook-bot description: Use when deploying the runbook bot to AWS Lambda or ECS --- 1. Run `pip freeze > requirements.txt` 2. Build Docker image, push to ECR 3. Update Lambda via `aws lambda update-function-code`
.claude/agents/ Specialized sub-agents
  • A narrow-focus agent with its own isolated context and tool preferences
  • Useful for a dedicated security or code-review pass
# agents/security-auditor.md --- name: security-auditor description: Reviews code for security issues before merging --- Check for: hardcoded credentials, unsafe boto3 IAM scopes, injection risks in user-input-to-prompt paths.
.claude/hooks/ Event-driven safety
  • Scripts that run automatically before/after a tool executes
  • Blocks unsafe operations, automates linting and validation
# hooks/validate-bash.sh #!/bin/bash if [[ "$1" == *"rm -rf"* ]]; then echo "BLOCKED: dangerous command" exit 1 fi
None of this is required to start — Claude Code works fine with zero config. Add these pieces gradually as a project grows, the same way you'd add IAM policies and CI rules over time.
Reference structure based on a Claude Code project layout diagram. Adapted with an AWS RAG-bot example.

Retrieval-Augmented Generation

RAG connects a language model to an external knowledge source, so it answers using your documents instead of guessing from what it memorised during training.

🎯
What problem it solves
01

A language model only knows what was in its training data, frozen at a cutoff date. It cannot see your company's docs, last week's changelog, or anything private. Ask it a question outside that training data and it either says it doesn't know, or worse, makes something up with total confidence.

RAG fixes this by retrieving relevant text at the moment of the question and handing it to the model as context, so the answer is grounded in real, current, verifiable source material rather than memory.

🔁
The two phases
02

Indexing happens once, up front. Query happens every time someone asks a question.

indexing — offline, run once or on doc updates
Documents
PDFs, wikis, tickets
→
Chunks
split with overlap
→
Embeddings
text → vectors
→
Vector DB
stored for search
query — online, run per question
User question
natural language
→
Query vector
same embed model
→
Retrieval
top-k similar chunks
→
LLM
question + context
→
Answer
grounded, citable

The vector DB is the only thing the two phases share. This decoupling is why RAG scales well — you can re-index documents without touching the model, and swap the LLM without re-embedding anything.

🧩
Core pieces
03
Chunking
Splitting documents into pieces small enough to embed and retrieve precisely, usually 200–1000 tokens with some overlap so meaning isn't cut mid-sentence.
Embeddings
A model that converts text into a vector of numbers, positioned so semantically similar text ends up close together in that space.
Vector database
Stores those vectors and does fast similarity search — Chroma, Pinecone, Qdrant, Weaviate, or pgvector are common choices.
Reranking
An optional second pass that reorders retrieved chunks by relevance, since raw embedding similarity is a decent but imperfect signal.
📖
Glossary
04
TermSimple meaning
ChunkA piece of a document small enough to embed and retrieve on its own
EmbeddingA vector representing the meaning of a piece of text
Cosine similarityHow the DB measures closeness between two vectors
Top-kThe number of most-similar chunks pulled back per query
GroundingBasing an answer on retrieved source text rather than model memory
HallucinationA confident but incorrect answer, the failure mode RAG reduces
⌨️
Minimal working example
05

Chroma runs embedded with a built-in embedding model, so this needs no API key to try.

terminal
pip install chromadb
rag_with_chroma.py
import chromadb client = chromadb.Client() collection = client.create_collection("docs") collection.add( documents=[ "Chunking splits documents into smaller pieces with overlap.", "Embeddings turn text into vectors that capture meaning.", "Vector databases support fast similarity search.", ], ids=["d0", "d1", "d2"], ) results = collection.query( query_texts=["How does chunking work?"], n_results=2, ) print(results["documents"][0])
⚖️
RAG vs fine-tuning
06
RAGFine-tuning
Updates knowledgeInstantly, re-index docsRequires retraining
Best forFacts, current data, citationsTone, format, behaviour
Cost to updatelowhigher, needs training runs
TraceabilityCan cite exact sourceOpaque, baked into weights
⚠️
Common pitfalls
07
  • Chunks too large or too small — either buries the answer in noise or splits it across boundaries.
  • No reranking — top embedding matches aren't always the most relevant ones.
  • No fallback for "not found" — the model should say it doesn't know rather than force an answer from weak matches.
  • Stale embeddings — re-index when source documents change, or answers drift out of date silently.
Concept guide — retrieval-augmented generation. Part of the AI learning resource series.

ChromaDB, from scratch

How text becomes vectors, how similarity search actually works, and the full Python API — with runnable examples at every step.

🧠
What Chroma actually is
01

Not a replacement for your normal database — a companion for anything that needs "find similar," not "find exact."

A normal database matches on exact values: WHERE user_id = 42. Chroma matches on meaning. It stores each piece of text as an embedding — a list of numbers capturing what that text is about — and lets you ask "what's most similar to this?" even when no words overlap.

01
Text in
"Dogs are loyal companions"
02
Embed
model → [0.12, -0.4, …]
03
Store
vector + id + metadata
04
Query
nearest vectors by distance

This is the backbone of most retrieval-augmented generation (RAG) systems: embed a knowledge base once, then at question time retrieve the nearest chunks and hand them to an LLM as context.

📐
Vectors, angles & dimensions
02

The geometry that makes "find similar" a maths problem instead of a keyword problem.

"The cat sat on the mat" "Dogs are loyal companions" "Python is a programming language" θ small θ large

Each sentence is a vector pointing away from the origin. The angle between two vectors — not the distance between their tips — is what cosine similarity measures. The cat and dog sentences point in nearly the same direction (small θ, high similarity); the Python sentence points somewhere else entirely (large θ, low similarity).

Why angle instead of raw distance

A long sentence and a short sentence about the exact same topic can produce vectors of different lengths. If you measured plain distance, the longer one would look artificially "further away" just because its vector is bigger — not because the meaning differs. Measuring the angle instead cancels out length and isolates direction, which is what actually encodes meaning.

cosine_similarity(A, B) = (A · B) / (||A|| × ||B||)

The dot product on top captures how much the two vectors point the same way; dividing by both magnitudes normalises them to unit length first, so only the angle survives. Result ranges from -1 (opposite meaning) to 1 (identical direction) — most related-but-different sentences land somewhere around 0.3–0.9. Chroma stores this as cosine distance (1 − similarity), so lower always means closer, matching the convention of other distance metrics.

What "dimensions" really means here

The diagram above is 2D so it's drawable, but real embeddings from models like all-MiniLM-L6-v2 or OpenAI's text-embedding-3-small have 384 to 1536 dimensions. Each dimension is one axis in that space — not something you can point to and label ("dimension 42 = animal-ness"), but collectively the axes encode aspects like topic, tone, formality, and tense, learned automatically during training.

ModelTypical dimensions
all-MiniLM-L6-v2 Chroma default384
text-embedding-3-small1536
all-mpnet-base-v2768
text-embedding-3-large3072
more dimensions ≠ always better

Higher-dimensional embeddings can capture finer distinctions, but they cost more to store and compare, and beyond a point the extra dimensions mostly encode noise rather than useful meaning. Model choice is a trade-off between quality, storage, and query speed — not "bigger is strictly better."

📦
Install & first run
03
terminal
pip install chromadb
test_chroma.py
import chromadb client = chromadb.Client() collection = client.create_collection(name="my_docs") collection.add( documents=["The cat sat on the mat", "Dogs are loyal companions", "Python is a programming language"], ids=["id1", "id2", "id3"] ) results = collection.query(query_texts=["tell me about pets"], n_results=2) print(results["documents"])
first run

Chroma silently downloads a small default embedding model (~80 MB, all-MiniLM-L6-v2) the first time you call .add(). That's expected — subsequent runs are fast.

🔌
Client types
04

Pick based on whether data needs to survive a restart, and whether multiple processes share it.

ephemeralchromadb.Client()

In-memory only. Fast to try things, gone the moment the process ends. Good for scratch scripts and this guide's examples.

persistentchromadb.PersistentClient(path)

Saves to disk at the given path. This is what you use for anything real — a script you'll re-run, a local app.

client = chromadb.PersistentClient(path="./chroma_db")
serverchromadb.HttpClient(host, port)

Connects to a Chroma server running elsewhere (start one with chroma run). Use this when multiple apps or processes need to share the same database over the network.

client = chromadb.HttpClient(host="localhost", port=8000)
🗂️
Collections
05

A collection is roughly a table — a named group of embedded documents.

collection = client.create_collection( name="my_docs", metadata={"hnsw:space": "cosine"} # or "l2", "ip" ) # avoid errors on re-run collection = client.get_or_create_collection(name="my_docs") client.list_collections() client.delete_collection(name="my_docs") collection = client.get_collection(name="my_docs")
note

The distance metric (hnsw:space) is set at creation time and can't be changed after — plan ahead if you know you'll need something other than the cosine default.

🔢
Embedding functions
06

Whatever turns your text into vectors. Swappable — but not mid-collection.

from chromadb.utils import embedding_functions # OpenAI openai_ef = embedding_functions.OpenAIEmbeddingFunction( api_key="sk-...", model_name="text-embedding-3-small" ) # local, open-source st_ef = embedding_functions.SentenceTransformerEmbeddingFunction( model_name="all-mpnet-base-v2" ) collection = client.create_collection(name="my_docs", embedding_function=openai_ef)
important constraint

Every document in a collection must use the same embedding function. Mixing models produces vectors from different, non-aligned spaces — distances between them mean nothing.

✏️
Add, query, update, delete
07
Add
collection.add( documents=["text one", "text two"], metadatas=[{"source": "manual", "page": 1}, {"source": "manual", "page": 2}], ids=["doc1", "doc2"] )
Query
results = collection.query( query_texts=["find something like this"], n_results=5, where={"source": "manual"}, include=["documents", "metadatas", "distances"] )
Update, upsert, delete
collection.update(ids=["doc1"], documents=["updated text"]) collection.upsert(ids=["doc1"], documents=["text"]) # insert or update collection.delete(ids=["doc1"]) collection.delete(where={"source": "manual"})

upsert is the one you'll reach for most — it lets you re-run an ingestion script without hitting duplicate-id errors.

Inspect
collection.count() collection.peek(limit=5) collection.get(ids=["doc1"])
🔍
Metadata filters
08

Narrow the search space before similarity even runs.

OperatorMeaning
$eq / $neequals / not equals
$gt / $gtegreater than / or equal
$lt / $lteless than / or equal
$and / $orcombine multiple conditions
{"page": {"$gt": 5}} {"$and": [{"source": "manual"}, {"page": {"$gt": 5}}]}

There's also where_document for full-text filtering on the raw content, e.g. {"$contains": "python"}.

⚡
How the search is actually fast
09

Comparing against every vector doesn't scale — HNSW makes it approximate but fast.

With a few hundred documents, brute-force comparison is instant. With millions, it isn't. Chroma uses HNSW (Hierarchical Navigable Small World graphs): a layered graph where the top layer has few nodes and long jumps, and the bottom layer holds every vector densely connected to its neighbours.

A query starts at the top, greedily hops toward whatever's closest, drops a layer once it can't get closer, and repeats — funnelling down to a precise local search. Roughly binary search, but for high-dimensional space.

ParameterEffect
MNeighbour connections per node — higher = more accurate, more memory
ef_searchCandidates explored per query — higher = more accurate, slower
cosine, not euclidean

Chroma defaults to cosine similarity — it compares vector direction, ignoring length. A short and a long sentence about the same topic still score as similar, which raw distance would unfairly penalise.

🔗
A minimal RAG pipeline
10

Embed once, retrieve at question time, hand the result to an LLM.

client = chromadb.PersistentClient(path="./chroma_db") collection = client.get_or_create_collection("knowledge_base") # 1. ingest (run once, or whenever docs change) chunks = ["...", "..."] # split long docs into ~500-token pieces collection.upsert(documents=chunks, ids=[f"chunk_{i}" for i in range(len(chunks))]) # 2. retrieve (run per question) question = "How do I reset a user's password?" hits = collection.query(query_texts=[question], n_results=3) context = "\n\n".join(hits["documents"][0]) # 3. generate — hand `context` + `question` to your LLM of choice

The one detail that matters most for quality: chunking. Embedding a whole document as one vector averages away detail. Splitting into ~500-token pieces with a little overlap keeps each chunk sharply about one thing.

New to the concept itself? The RAG guide covers the two phases, chunking strategy, and common pitfalls in more depth.
📖
Glossary
11
TermMeaning
EmbeddingA vector of numbers representing the meaning of a piece of text
CollectionA named group of documents + their embeddings, like a table
Cosine similaritySimilarity measured by angle between vectors, ignoring length
HNSWThe graph structure Chroma uses for fast approximate nearest-neighbour search
ChunkingSplitting long documents into smaller pieces before embedding
RAGRetrieval-Augmented Generation — retrieve relevant text, then generate an answer grounded in it
UpsertUpdate if the id exists, insert if it doesn't
Personal learning notes on ChromaDB — built for reference, not production.