AI Agents Explained: Perception, Planning, Memory, and Tool Use
A chatbot and an AI agent can be built from the exact same underlying language model, and yet they behave completely differently. Ask a chatbot to fix a bug in your code, and it writes you a suggested fix in the chat window — you copy it, paste it, run it, and tell the chatbot what happened. Ask an agent to fix the same bug, and it opens the file itself, edits it, runs your test suite, reads the failure output, adjusts its fix, and runs the tests again — on its own, without you doing any of the copying and pasting. That difference isn’t about the model being smarter. It’s about architecture: an agent is a language model wrapped in a system that gives it perception, planning, memory, and the ability to use tools, wired together in a loop.
Perception: knowing what’s actually going on
Before an agent can do anything useful, it needs an accurate picture of its current situation — not just the user’s instruction, but the actual state of whatever it’s working in. For a coding agent, that means reading the relevant files in a codebase, not just the one line the user mentioned. For a research agent, it means reading the actual content of a search result, not just its title. This sounds obvious, but it’s a real engineering problem: an agent working with stale or incomplete information about its environment will confidently act on a wrong picture of reality, which is one of the most common sources of agent failure in practice.
Planning: breaking a goal into steps — and revising when needed
A single goal like “fix the failing test” or “write a report comparing these three vendors” almost never maps to one action. Planning is the process of decomposing that goal into an ordered sequence of smaller steps: read the test file, understand what it expects, read the function it’s testing, identify the mismatch, write a fix, run the test again.
The important part isn’t just making the plan — it’s that a capable agent revises the plan as new information comes in. If running the test reveals the actual bug is somewhere else entirely, a well-built agent updates its plan around that discovery instead of stubbornly executing its original steps. This is what separates agentic planning from a fixed script: the plan is a working hypothesis, not a rigid checklist.
Memory: short-term context versus long-term persistence
Memory in an agent comes in two meaningfully different flavors. Short-term memory is the context the agent keeps within a single task — what file it just opened, what the last tool call returned, what step it’s currently on. This lives only as long as the task does, and every agent has some version of it by necessity, because you can’t execute a multi-step plan without remembering what you already did.
Long-term memory is different: it’s information that persists across sessions, deliberately stored somewhere the agent can search or read later — a notes file, a database, a record of past user preferences. This isn’t automatic. A language model itself doesn’t retain anything between separate conversations; if an agent “remembers” something you told it last week, that’s because the surrounding system explicitly saved it and fed it back in, not because the model has some biological-style memory of its own. This distinction matters because it’s the difference between an agent that starts fresh every time and one that actually improves or personalizes over repeated use.
Tool use: doing things instead of just describing them
This is arguably the component that turns a language model into an agent at all. On its own, a language model can only generate text — it can describe what code should do, but it can’t run that code and see if it works. Tool use is the mechanism that lets an agent call something outside itself: executing code, searching the web, querying a database, calling an API, reading or writing a file. The model decides when to call a tool and what to call it with, based on its plan, and then reads the tool’s actual output back into its context before deciding what to do next.
This is the piece that makes agents genuinely useful for real work rather than just conversation. A model that can only talk about running your test suite is a chatbot. A model that can actually invoke the test runner, read the pass/fail output, and act on it is an agent.
The loop: observe, plan, act, observe again
Put those four pieces together and you get the core agent loop that runs underneath basically every real agentic system: observe the current state (perception), decide what to do (planning, using memory of what’s happened so far), take an action (tool use), then observe the result of that action, and repeat — replanning whenever the new observation doesn’t match what was expected. This loop is what lets an agent handle a task where the correct sequence of steps genuinely can’t be known in advance, because the right next step depends on what actually happens after each action.
A concrete example: a coding agent given the task “make the test suite pass” reads the codebase (perception), plans an initial fix (planning), edits a file and runs the tests (action), reads the resulting failure message (observe), realizes the bug was somewhere else (replan), edits the correct file, and runs the tests again — repeating until they pass or it runs out of reasonable options. A research agent works the same loop with different tools: it searches the web (action), reads what comes back (observe), decides its search terms were too narrow (replan), searches again with better terms, and eventually synthesizes what it found into an answer.
Where agents still break down
It’s worth being honest about the current limitations, because they’re a direct consequence of this architecture, not incidental bugs. Every step in the loop depends on the accuracy of the step before it — if a tool call returns something the model misreads, or if the model simply hallucinates what a tool returned instead of reading it carefully, that mistaken belief carries forward into every subsequent step instead of getting caught immediately. On a short, simple task this rarely causes visible problems. On a long task with many steps, small errors compound: a slightly wrong assumption in step three can quietly corrupt steps four through twenty. Agents can also get stuck in unproductive loops — retrying a failing approach with minor variations instead of stepping back to reconsider the plan — especially when the failure signal itself is ambiguous. This is exactly why, for anything with real consequences, current agentic systems are best used with a human reviewing the outcome rather than left to run fully unsupervised.
Why this matters beyond text
Everything described above assumes an agent working mostly with text: files, search results, code, API responses. But a growing number of real agentic tasks involve information that isn’t text at all — a screenshot the agent needs to interpret, a photo it needs to analyze, a diagram it needs to reason about. That requires a model that can perceive more than one type of input in the first place, reasoning over images and text together rather than needing a separate specialized system bolted on for each. That capability — a single model understanding multiple types of data at once — is called multimodal AI, and it’s the subject of the next article in this series.
What's the actual difference between a chatbot and an AI agent?
A chatbot takes a message and generates a text reply — the interaction is one input, one output, and the conversation is the whole product. An agent takes a goal and works toward completing it, potentially across many internal steps the user never sees: searching, writing code, running it, checking the result, and trying again. The output of an agent is a completed task, not just a message.
Do AI agents actually remember things between sessions?
It depends on the system. Within a single task, an agent keeps short-term context — what it just did, what it just observed. Whether it retains anything after that task ends depends on whether it's been given a separate long-term memory system, like a database it can write notes to and search later. Without that, an agent forgets everything the moment the session ends, no matter how sophisticated its reasoning was during the task.
Why do AI agents sometimes get stuck in loops or fail on long tasks?
Because every step in an agent's process depends on the ones before it — if the model misreads a tool's output, hallucinates a result, or plans around a wrong assumption, that error carries forward into every later step instead of being caught and corrected immediately. On short tasks this rarely matters; on long, multi-step tasks small errors compound, which is why current agents still need human review for anything with real consequences.