Components of A Coding Agent
by Sebastian Raschka (Ahead of AI)
The six building blocks behind Claude Code-style coding agents, with a pure-Python reference implementation you can run locally
Overview
Components of A Coding Agent, subtitled How Coding Agents Use Tools, Memory, and Repo Context to Make LLMs Work Better in Practice, was published free on April 4, 2026 in Ahead of AI, the newsletter of Sebastian Raschka, author of Build a Large Language Model (From Scratch). It opens by separating LLMs, reasoning models and agents, then defines the coding harness as the software scaffold around the model. The core sections cover live repo context (git status, branch, docs and file structure gathered up front), prompt shape and cache reuse (stable instructions and tool descriptions kept in a cacheable prefix, separate from the changing transcript), tool access through predefined, validated actions with approval gates, minimizing context bloat by clipping outputs, deduplicating file reads and compressing older turns, structured session memory that splits working state from the full transcript, and delegation to bounded subagents. A components summary and a comparison with OpenClaw follow. The article ships with rasbt/mini-coding-agent on GitHub (Apache-2.0, about 1.2k stars), a single-file, stdlib-only Python agent that implements all six components, runs against local models through Ollama (default qwen3.5:4b), and includes approval modes, session resumption and a REPL with slash commands.
At a Glance
- Topic
- Agentic
- Level
- Intermediate
- Format
- Tutorial
- Cost
- Free
- Duration
- ~45 min read; a few more hours to run and read the companion mini-coding-agent code
- Provider
- Sebastian Raschka (Ahead of AI)
- Hands-on
- Yes — code/exercises
- Certificate
- None
What You’ll Learn
- ✓How LLMs, reasoning models and agents differ, and where the coding harness sits between them
- ✓How to collect live repo context such as git status, branch and file structure before the first turn
- ✓How to order the prompt so instructions and tool descriptions stay in a reusable cached prefix
- ✓How to expose tools as validated, predefined actions with approval gates instead of free-form shell commands
- ✓How to keep context small by clipping tool output, deduplicating file reads and compressing old turns
- ✓How to separate working memory from the full transcript so a session can be resumed later
- ✓How to delegate subtasks to bounded subagents without duplicated work or runaway recursion
- ✓How a minimal Python agent loop on Ollama implements all six components in one file
Highlights
- •Companion repo rasbt/mini-coding-agent has about 1.2k stars and 218 forks, uses only the Python standard library and includes tests and CI
- •Runs fully locally on small open-weight models through Ollama, so you can experiment without an API bill
- •Explains the mechanics behind Claude Code and Codex CLI (caching, permissions, memory) without tying you to either vendor
- •Free to read, from the author of Build a Large Language Model (From Scratch)
- •Short enough to finish in an evening, unlike survey papers on agent harnesses
Who It’s For
Best For
- ✓Developers who use Claude Code or Codex CLI daily and want to understand what the harness is doing
- ✓Engineers building an internal coding agent or tool-using assistant from scratch
- ✓Learners who prefer one readable Python file over a large agent framework
- ✓Teams tuning prompt caching and context size to cut agent token costs
Prerequisites
- •Intermediate Python, enough to read a single long script
- •Basic LLM concepts: prompts, tool calling and context windows
- •Ollama installed and Python 3.10+ if you want to run the companion code
FAQ
What is Components of A Coding Agent?
A free article by Sebastian Raschka that breaks a coding agent into six components: live repo context, prompt caching, validated tools, context reduction, session memory and bounded subagents. It is for developers who use tools like Claude Code or Codex CLI and want to know how they work. Afterwards you can read or build a minimal coding agent harness yourself.
Is Components of A Coding Agent free?
Components of A Coding Agent is free to access.
What level is Components of A Coding Agent for?
Components of A Coding Agent is aimed at a intermediate audience. Recommended background: Intermediate Python, enough to read a single long script, Basic LLM concepts: prompts, tool calling and context windows, Ollama installed and Python 3.10+ if you want to run the companion code.
How long does Components of A Coding Agent take?
Expect roughly ~45 min read; a few more hours to run and read the companion mini-coding-agent code. Most learners work through it at their own pace.
What will I learn from Components of A Coding Agent?
You'll learn: How LLMs, reasoning models and agents differ, and where the coding harness sits between them; How to collect live repo context such as git status, branch and file structure before the first turn; How to order the prompt so instructions and tool descriptions stay in a reusable cached prefix; How to expose tools as validated, predefined actions with approval gates instead of free-form shell commands; How to keep context small by clipping tool output, deduplicating file reads and compressing old turns; How to separate working memory from the full transcript so a session can be resumed later; How to delegate subtasks to bounded subagents without duplicated work or runaway recursion; How a minimal Python agent loop on Ollama implements all six components in one file.
Topics
Sources
This page was written from 2 sources, 1 on domains other than magazine.sebastianraschka.com.