Learn AI from scratch, the developer's way
Concepts explained simply, with real-world examples — so you actually understand what's happening under the hood.
The flow below shows how raw text travels through an AI system — from tokenisation to embeddings to RAG to grounded output.
| Term | Simple meaning |
|---|---|
| AI | Machines that simulate human intelligence |
| ML (Machine Learning) | AI that learns from data instead of explicit rules |
| Deep Learning | ML using neural networks with many layers |
| GenAI | AI that generates content — text, images, code |
| LLM | Large Language Model — the brain behind ChatGPT, Claude |
| Foundation Model | A massive pre-trained model others build on top of |
| Term | Simple meaning |
|---|---|
| Token | A chunk of text (roughly 1 word ≈ 1.3 tokens) |
| Context Window | How much text the model can "see" at once |
| Prompt | The input you give the model |
| Inference | Running the model to get an output |
| Temperature | Controls randomness — 0 = predictable, 1 = creative |
| Embedding | Converting text into numbers (vectors) for comparison |
| Vector | A list of numbers representing meaning |
| Semantic Search | Retrieval that understands intent, not just keyword matches |
| Term | Simple meaning |
|---|---|
| RAG | Retrieval-Augmented Generation — give the model your own docs |
| Fine-tuning | Re-training a model on your specific data |
| Prompt Engineering | Crafting inputs to get better outputs |
| Agent | An AI that can take actions, use tools, make decisions |
| Tool Use / Function Calling | Letting the LLM call your code or APIs |
| Chain | Connecting multiple AI steps together |
| Hallucination | When the model confidently says something wrong |
| Term | Simple meaning |
|---|---|
| Parameters | The "weights" inside a model (GPT-4 ≈ 1 trillion) |
| Training | Teaching a model on massive datasets |
| Pre-training | Initial training on general internet data |
| RLHF | Human feedback used to make models more helpful and safe |
| Transformer | The architecture behind almost all modern LLMs |
| Attention | How the model decides what parts of text matter most |
| Model | Input / 1M tokens | Output / 1M tokens | Best for |
|---|---|---|---|
| Haiku 4.5 | $1.00 | $5.00 | Learning & cheap tasks |
| Sonnet 4.6 | $3.00 | $15.00 | Best balance — use this for Week 1 |
| Opus 4.8 | $5.00 | $25.00 | Most powerful tasks |
- Processing thousands of documents
- Bulk summarization
- Data extraction / classification
- Nightly jobs
- Offline evaluations
- Content generation queues
- Chatbots
- Real-time APIs
- Interactive applications
- Anything needing a response in seconds
Dynamic context (user's question) is always charged at full rate.
How the .claude/ config folder works, explained with a real example project.
- The main brief Claude Code reads at the start of every session
- Defines project overview, tech stack, build/run commands
- Documents architecture and conventions so Claude doesn't need re-explaining each time
- Same idea as CLAUDE.md, but personal overrides not shared with the team
- Use it for machine-specific notes — local AWS profile, test file paths, Python version
- Stores MCP (Model Context Protocol) integration configs
- Connects Claude Code to GitHub, Jira, Slack, databases
- Checked into git, so the whole team gets the same integrations automatically
- Controls what Claude Code is allowed to do — like an IAM policy for the agent
- Defines model selection and hooks
settings.local.jsonworks the same way for personal, untracked overrides
- Topic-specific markdown files instead of one giant CLAUDE.md
- Covers style, testing, API design — can target specific files/paths
- Custom reusable workflows, triggered like
/project:fix-issue - Good for repeatable tasks: bug fixes, releases, audits
- Loaded automatically only when the task context matches
- Keeps everyday context lightweight — deploy instructions don't load unless you're deploying
- A narrow-focus agent with its own isolated context and tool preferences
- Useful for a dedicated security or code-review pass
- Scripts that run automatically before/after a tool executes
- Blocks unsafe operations, automates linting and validation
Retrieval-Augmented Generation
RAG connects a language model to an external knowledge source, so it answers using your documents instead of guessing from what it memorised during training.
A language model only knows what was in its training data, frozen at a cutoff date. It cannot see your company's docs, last week's changelog, or anything private. Ask it a question outside that training data and it either says it doesn't know, or worse, makes something up with total confidence.
RAG fixes this by retrieving relevant text at the moment of the question and handing it to the model as context, so the answer is grounded in real, current, verifiable source material rather than memory.
Indexing happens once, up front. Query happens every time someone asks a question.
The vector DB is the only thing the two phases share. This decoupling is why RAG scales well — you can re-index documents without touching the model, and swap the LLM without re-embedding anything.
| Term | Simple meaning |
|---|---|
| Chunk | A piece of a document small enough to embed and retrieve on its own |
| Embedding | A vector representing the meaning of a piece of text |
| Cosine similarity | How the DB measures closeness between two vectors |
| Top-k | The number of most-similar chunks pulled back per query |
| Grounding | Basing an answer on retrieved source text rather than model memory |
| Hallucination | A confident but incorrect answer, the failure mode RAG reduces |
Chroma runs embedded with a built-in embedding model, so this needs no API key to try.
terminal| RAG | Fine-tuning | |
|---|---|---|
| Updates knowledge | Instantly, re-index docs | Requires retraining |
| Best for | Facts, current data, citations | Tone, format, behaviour |
| Cost to update | low | higher, needs training runs |
| Traceability | Can cite exact source | Opaque, baked into weights |
- Chunks too large or too small — either buries the answer in noise or splits it across boundaries.
- No reranking — top embedding matches aren't always the most relevant ones.
- No fallback for "not found" — the model should say it doesn't know rather than force an answer from weak matches.
- Stale embeddings — re-index when source documents change, or answers drift out of date silently.
ChromaDB, from scratch
How text becomes vectors, how similarity search actually works, and the full Python API — with runnable examples at every step.
Not a replacement for your normal database — a companion for anything that needs "find similar," not "find exact."
A normal database matches on exact values: WHERE user_id = 42. Chroma matches on meaning. It stores each piece of text as an embedding — a list of numbers capturing what that text is about — and lets you ask "what's most similar to this?" even when no words overlap.
This is the backbone of most retrieval-augmented generation (RAG) systems: embed a knowledge base once, then at question time retrieve the nearest chunks and hand them to an LLM as context.
The geometry that makes "find similar" a maths problem instead of a keyword problem.
Each sentence is a vector pointing away from the origin. The angle between two vectors — not the distance between their tips — is what cosine similarity measures. The cat and dog sentences point in nearly the same direction (small θ, high similarity); the Python sentence points somewhere else entirely (large θ, low similarity).
A long sentence and a short sentence about the exact same topic can produce vectors of different lengths. If you measured plain distance, the longer one would look artificially "further away" just because its vector is bigger — not because the meaning differs. Measuring the angle instead cancels out length and isolates direction, which is what actually encodes meaning.
The dot product on top captures how much the two vectors point the same way; dividing by both magnitudes normalises them to unit length first, so only the angle survives. Result ranges from -1 (opposite meaning) to 1 (identical direction) — most related-but-different sentences land somewhere around 0.3–0.9. Chroma stores this as cosine distance (1 − similarity), so lower always means closer, matching the convention of other distance metrics.
The diagram above is 2D so it's drawable, but real embeddings from models like all-MiniLM-L6-v2 or OpenAI's text-embedding-3-small have 384 to 1536 dimensions. Each dimension is one axis in that space — not something you can point to and label ("dimension 42 = animal-ness"), but collectively the axes encode aspects like topic, tone, formality, and tense, learned automatically during training.
| Model | Typical dimensions |
|---|---|
| all-MiniLM-L6-v2 Chroma default | 384 |
| text-embedding-3-small | 1536 |
| all-mpnet-base-v2 | 768 |
| text-embedding-3-large | 3072 |
Higher-dimensional embeddings can capture finer distinctions, but they cost more to store and compare, and beyond a point the extra dimensions mostly encode noise rather than useful meaning. Model choice is a trade-off between quality, storage, and query speed — not "bigger is strictly better."
Chroma silently downloads a small default embedding model (~80 MB, all-MiniLM-L6-v2) the first time you call .add(). That's expected — subsequent runs are fast.
Pick based on whether data needs to survive a restart, and whether multiple processes share it.
ephemeralchromadb.Client()
In-memory only. Fast to try things, gone the moment the process ends. Good for scratch scripts and this guide's examples.
persistentchromadb.PersistentClient(path)
Saves to disk at the given path. This is what you use for anything real — a script you'll re-run, a local app.
serverchromadb.HttpClient(host, port)
Connects to a Chroma server running elsewhere (start one with chroma run). Use this when multiple apps or processes need to share the same database over the network.
A collection is roughly a table — a named group of embedded documents.
The distance metric (hnsw:space) is set at creation time and can't be changed after — plan ahead if you know you'll need something other than the cosine default.
Whatever turns your text into vectors. Swappable — but not mid-collection.
Every document in a collection must use the same embedding function. Mixing models produces vectors from different, non-aligned spaces — distances between them mean nothing.
upsert is the one you'll reach for most — it lets you re-run an ingestion script without hitting duplicate-id errors.
Narrow the search space before similarity even runs.
| Operator | Meaning |
|---|---|
| $eq / $ne | equals / not equals |
| $gt / $gte | greater than / or equal |
| $lt / $lte | less than / or equal |
| $and / $or | combine multiple conditions |
There's also where_document for full-text filtering on the raw content, e.g. {"$contains": "python"}.
Comparing against every vector doesn't scale — HNSW makes it approximate but fast.
With a few hundred documents, brute-force comparison is instant. With millions, it isn't. Chroma uses HNSW (Hierarchical Navigable Small World graphs): a layered graph where the top layer has few nodes and long jumps, and the bottom layer holds every vector densely connected to its neighbours.
A query starts at the top, greedily hops toward whatever's closest, drops a layer once it can't get closer, and repeats — funnelling down to a precise local search. Roughly binary search, but for high-dimensional space.
| Parameter | Effect |
|---|---|
| M | Neighbour connections per node — higher = more accurate, more memory |
| ef_search | Candidates explored per query — higher = more accurate, slower |
Chroma defaults to cosine similarity — it compares vector direction, ignoring length. A short and a long sentence about the same topic still score as similar, which raw distance would unfairly penalise.
Embed once, retrieve at question time, hand the result to an LLM.
The one detail that matters most for quality: chunking. Embedding a whole document as one vector averages away detail. Splitting into ~500-token pieces with a little overlap keeps each chunk sharply about one thing.
| Term | Meaning |
|---|---|
| Embedding | A vector of numbers representing the meaning of a piece of text |
| Collection | A named group of documents + their embeddings, like a table |
| Cosine similarity | Similarity measured by angle between vectors, ignoring length |
| HNSW | The graph structure Chroma uses for fast approximate nearest-neighbour search |
| Chunking | Splitting long documents into smaller pieces before embedding |
| RAG | Retrieval-Augmented Generation — retrieve relevant text, then generate an answer grounded in it |
| Upsert | Update if the id exists, insert if it doesn't |