Tag: Learning

AI Foundations 1 - How AI Models Work

Calling an AI model feels like calling a function: you pass in some input, you get output back. But it breaks one rule you’d normally rely on — call it twice with the same input, and you can get two different answers.

Deterministic vs. probabilistic

A sorting function is deterministic. Feed it [3, 1, 2] a hundred times, get [1, 2, 3] a hundred times. The logic is fixed, written line by line by a person.

[…445 words]

AI Foundations 2 - Hallucinations and Limitations

Sometimes a model states something false with total confidence. No hedging, no “I’m not sure” — just a wrong answer delivered like a fact. That’s a hallucination.

It’s not lying, exactly. The model isn’t tracking truth at all. It’s producing text that looks like a correct answer would look.

Why hallucinations happen

A model is trained to predict plausible next words, not to check facts against a database. When it doesn’t know something, it doesn’t stop and say so — it keeps generating the most likely-sounding continuation, true or not.

[…326 words]

AI Foundations 3 - Tokens and Pricing

Models don’t read words the way you do. They break text into chunks called tokens, and a token isn’t the same thing as a word. “Cat” might be one token. “Unbelievable” might split into two or three. Code is chunked the same way — brackets, keywords, and indentation all count.

Why tokens matter

Every token costs something, and every token takes time to produce. A short prompt on a small file runs fast and cheap. Paste in a ten-thousand-line log file and ask for a summary, and you’ll feel both the wait and the bill.

Input and output tokens

Input tokens are everything you send in: your prompt, any files, the conversation so far. Output tokens are what the model generates back. They’re counted and priced separately, and they behave differently — one you control directly, the other you only shape indirectly by asking for shorter or longer answers.

Pricing

Providers charge per token, usually in price-per-million units, and output tokens almost always cost more per token than input tokens. Makes sense — generating text is the harder, slower half of the job.

Streaming

Rather than making you wait for the entire response, most models can stream tokens out one at a time as they’re generated. That’s why chat interfaces show text appearing word by word instead of all at once. It doesn’t make the model faster overall, but it makes the wait feel shorter.

Token optimization

A few habits keep both cost and latency down:

  • Don’t paste in more context than the task actually needs.
  • Reuse a stable prompt prefix where the provider supports caching it — repeated setup shouldn’t be repriced every call.
  • Ask explicitly for a short answer when a short answer is all you need. Models default to being thorough, which is usually not what you’re paying for.

Previous: AI Foundations 2 - Hallucinations and Limitations Next: AI Foundations 4 - Context

AI Foundations 4 - Context

Everything a model can see for the current task lives in its context — think of it as short-term memory that gets rebuilt fresh each time you ask something. Nothing outside that window exists to the model, no matter how obvious it seems to you.

System and user prompts

Most setups split instructions into two layers. A system prompt sets the general rules — tone, role, boundaries — and usually stays fixed across a whole session. A user prompt is the specific ask for this particular turn. Same system prompt, different user prompts, very different outcomes each time.

[…253 words]

AI Foundations 5 - Tool Calling

On its own, a model can only produce text. It can’t check today’s date, read a file, or run your test suite — it can only describe what those things might look like. Tools close that gap: they let the model reach out and actually do something, then bring the result back into the conversation.

Basic tool-call flow

The pattern is always roughly the same. The model decides a tool would help, picks which one, and fills in the arguments it needs. Something outside the model — the application — actually runs it. Whatever comes back gets fed into the model’s context, and it decides what to do next based on that.

[…280 words]

AI Foundations 6 - Prompting and Evaluation

The exact same model can look brilliant or mediocre depending entirely on how you ask it something. Swap a vague instruction for a specific one, and the quality, the format, even the correctness of the answer can shift noticeably — without touching the model at all.

Prompt structure

A prompt that works tends to spell out the same handful of things: what the task actually is, what context matters, and what shape the answer should take. Keeping the role, the task, any constraints, and any examples visibly separate makes a long prompt easier for the model to parse — and easier for you to debug later.

[…317 words]

AI Foundations 7 - Agents and Autonomous Loops

One tool call answers one question. An agent chains many of them together: decide what to do, do it, look at what happened, decide again. It keeps cycling through that loop until the task looks done, or until something stops it.

Planning

Before diving into a large task, an agent can lay out a rough sequence of steps first rather than acting on the very first idea that comes to mind. That plan isn’t fixed — new information from an earlier step can send it back to revise later steps, sometimes more than once.

[…267 words]

micromlp: A From-Scratch Neural Net That Predicts Housing Prices

I built micromlp: a single file of Python, no dependencies, no PyTorch. It downloads a real dataset, builds a 2-layer MLP, implements automatic differentiation from scratch, trains with gradient descent, and makes predictions. I use the California housing dataset from chapter 2 of Hands-On Machine Learning. The task: predict a district’s median house value from its census stats.

This is inspired by Karpathy’s microgpt. There’s a difference between knowing .backward() exists and knowing what it does when you call it. I found it worth closing that gap once, on a toy.

[…1482 words]

Partial Derivatives, Reverse-Mode Autodiff

A partial derivative measures how much a function changes when you nudge one input, holding every other input fixed. For f(x, y) = x²y + y + 2, the partial derivative with respect to x asks: if y stays put, how fast does f move as x moves?

Manual differentiation

Mathematically, we know that ∂f/∂x = 2xy and ∂f/∂y = x² + 1, using a handful of rules:

  • the derivative of a constant is 0
  • the derivative of ax is a
  • the derivative of x^a is a·x^(a-1)
  • the derivative of a sum is the sum of the derivatives: (u + v)' = u' + v'
  • the derivative of a product follows the product rule: (u·v)' = u'v + uv'

But how can a program know that?

[…913 words]

Vanilla Neural Networks

A neural network is just a math function: y = FNN(x).

FNN has a nested form. Think of it as a stack of layers. A 3-layer neural network that returns a scalar value looks like this:

y = FNN(x) = f₃(f₂(f₁(x)))

x flows through three layers, f one, f two, f three, each computing g of W x plus b, to produce y.

Each f — f₁, f₂, … fₙ — has the same form:

f(x) = g(Wx + b)

W (the weight matrix) and b (a bias vector) are the learned parameters, usually trained via gradient descent. g is the activation function, and it can be chosen differently for each layer.

Wx + b is linear — wrapping it in g is what makes each layer non-linear. Without g (or with g chosen to be linear), the whole FNN collapses into a single linear function: stack 100 such layers and the composition of linear maps is still just one linear map, no matter how deep the network looks. g is what lets a stack of layers approximate anything more than a straight line. Popular choices for g are sigmoid and ReLU.

There are many variants of neural networks — CNNs, RNNs, transformers, and more — each shaped by assumptions about the data they process. The example above, where every neuron in one layer connects to every neuron in the next, is the plainest of them: a multilayer perceptron (MLP), also called a vanilla neural network.

ELI5

Thariq at Anthropic posted that it’s a skill people there have been using a lot recently: /eli5 <what you want to explain>.

I tried it on neural networks. Here’s what it made.

The first pass had bad colors, low contrast, hard to read. One follow-up prompt asking for a color fix and it was done.

No formulas, no sigmoid, no softmax, none of the math that actually makes a neural network work. It’s missing for good. What’s left is the dataflow — inputs go in, get combined, come out the other end as a decision — drawn simply enough to follow at a glance.

That’s the trade the skill makes, and it’s the right one for someone who isn’t about to read a textbook. eli5 is good exactly because it throws out the complexity most explanations lead with. I imagine this isn’t just for self-learning — it’s a genuinely useful tool for collaboration, communication, meetings. Conveying an idea isn’t easy. ;P


Updated at 2026-08-23. The neural network example is ported to julin.ai now.