Volver a las appsPortalDemo
Agent Loops · Beginner Guide

Agent loops, made simple

Everyone online means something a little different by "agent loop." This is the version that's simple, correct, and something you can build tonight.

An agent loop is just an AI that reasons what to do, acts, and observes the result, over and over, until the goal is met.
Why you're confused (it's not you)

Nobody draws it the same way

Watch five videos and you'll see five different pictures. They're all the same idea from different angles. Here are the four you'll bump into most:

New here? An LLM is just an AI model, the kind of thing behind ChatGPT. "Tools" are things it can do for itself: search the web, run code, edit a file. That's all those words mean below.
ReasonActObserve

"Think → Act → See"

The same three beats, older words.

the ReAct paper
AI modelTools

"Model uses tools, on repeat"

The simplest version. A model calling tools, over and over.

Anthropic
GoalSelf-promptruns unattended

"Runs on its own"

The same loop, just left to run unattended toward a goal.

AutoGPT
Managerhelperhelperhelper

"One boss, many helpers"

A lead agent hands tasks to sub-agents.

multi-agent
Four pictures.One skeleton underneath. That's the part you actually learn 👇
The one picture that matters

Underneath, it's always this

Strip the jargon and every agent loop is the same three beats, reason → act → observe, repeating toward a goal until a "done?" check says stop.

🧑‍💼

Think of a smart intern you don't micromanage. You hand them the goal. They figure out the next step, do it, check their own work, and go again, only coming back when it's actually done, or stuck. An agent loop is that, in software.

Your goal+ what "done" meansReasonthink the next step✓ DonestopActuse a toolObserveread what just happeneddone?yes ✓no → act

Each lap is one step. The agent keeps circling (act, observe, act, observe) until the "done?" check passes, or a guardrail (a safety limit, like a max number of tries) stops it. The Observe beat is the mechanism that makes a loop work: it reads its own output instead of assuming it worked. That's exactly why the "done?" check (coming up next) is everything.

Where loops actually fit

Most tasks don't need a loop

A loop is worth the setup only when a couple of things are true. Run your task down this ladder before you build anything.

Your taskDoes it repeat, or takemany unknown steps?noJust prompt itone shot is fasteryesCan the AI check"done" by itself?e.g. run the tests, count the wordsnoDo it yourselfit can't tell whenit's done yetyes✓ Build a loop
Pick a shape

Three loops you'll actually use

Start with the first one. Reach for the others only when a single agent genuinely can't keep up.

↑ These are the same shapes from up top, just grouped by when you'd actually reach for one. ("Runs on its own" isn't a separate shape; it's any of these left unattended, which is exactly why those need the strongest stop-limits.)

Solo loop

Start here, covers most work
1 agentreason → act → observe → repeat

One agent runs the loop on one task. Easiest to build and debug.

= "Think → Act → See" / "model + tools"

Maker → Checker

When quality matters
MakerCheckerfix until it passes

One agent does the work, a second grades it: a fresh AI given only the grading job, so it can't rubber-stamp its own work.

= the "Observe / check" beat, given its own agent

Manager → Helpers

When the job is big
Managerhelperhelperhelper

A lead splits the goal and hands pieces to sub-agents working in parallel.

= "one boss, many helpers"
Where beginners actually go wrong

A loop is only as good as its "done" check

The loop will happily produce confident, polished, wrong work and call it finished, unless you tell it exactly what "done" means and give it a way to check.

Before you build

Lock your finish line: answer these two first

These two answers are the "done?" check from the loop diagram.

1

What does "done" mean?

Write it so a machine could check it. "The tests pass." "Under 50 words and mentions the price." Not "make it good."

2

How will it check?

Different jobs need different checks. Pick the one that fits 👇

↓ that's question 2: pick one of these four ↓

Functional

A machine can answer yes/no with zero opinion.
e.g. tests pass, the app runs, the build compileseasiest: start here

Visual

It has to be seen to be judged.
e.g. UI, a thumbnail, a layoutmost agents can look at images for you

Judgment

Needs taste, but you can write a checklist.
score it against your rubrica second AI does the scoring for you

You decide

Irreversible, or pure taste no rubric captures.
loop pauses, you approve, then continuesfor risky or one-way-door steps
Your turn

Build your first loop tonight

You don't need a framework. If you have Claude Code (a free AI assistant you run on your computer, and it already works as a loop), just tell it the goal, what "done" means, and how to check. The same pattern works in any agent tool. Copy this:

your first agent loop · paste into Claude Code
# the goal
Here's the task: fix the one failing test in this project

# what "done" means (make it checkable)
You're done when: that test passes when you run the test suite

# how to check it
Verify by: running the tests and reading the output

# the guardrail
Keep going until it passes, or stop after 5 tries and tell me.
📝 Not a coder? Same exact shape, different job:  goal = rewrite this paragraph · done = under 50 words AND mentions the price · check = count the words, look for "price" · guardrail = stop after 5 tries.That's a Functional check: it checks the rules, not whether it reads well. Judging the writing would be a Judgment check.
🔒 Risky task? Add a human gate to the prompt: "Pause and ask me before deleting, sending, or paying for anything." That's the "You decide" check from above.
1

Name the goal

One clear sentence. "Fix the failing test." Not "improve the code."

↳ that's the GOAL box
2

Define + check "done"

Give it a finish line it can verify itself, plus a hard stop so it can't run forever.

↳ that's the DONE? check
3

Watch the first run

See where it trips. Fix the instructions, not the output. Then let it run on its own.

↳ that's reason → act → observe
Once it works, make it solid

8 things that make a loop actually work

You've already met a few of these (tagged recap). Here they are as one checklist to run before you trust a loop to go on its own. Most loops fail for boring reasons: they run forever, burn money, or ship junk. The good ones get these right.

A checkable goalrecap

Define "done" so a machine could verify it, not "make it good."

tests pass / under 50 words

A hard stoprecap

Max tries, a budget, or a time limit. Always. So it can't run forever.

stop after 5 tries

Good tools

The actions it can take must be reliable and clearly described.

search · run code · edit a file

Memory

Keep the history, but summarize it so the context doesn't bloat.

a running notes file

A separate checkerrecap

Generate → judge → fix → repeat. The maker never grades itself.

a 2nd agent scores it

Plan first?

Big, multi-step job? Have it write a plan before acting. Small job? Skip it.

outline before building

Logging

Save every thought, action, and result so you can debug at 3am.

log each step

Cost sense

Loops burn tokens fast. Start small and bounded, then scale.

watch the first runs

Do this

  • Start with one small, repeatable task
  • Make "done" something a machine can check
  • Always set a max number of tries
  • Use a separate agent to grade important work

Skip this (for now)

  • 24/7 swarms of 10 agents prompting 10 agents
  • Letting it run with no stop limit
  • Trusting "looks done" with no real check
  • Looping a one-off task you could just prompt

Six named here (Anthropic, OpenAI, Google, LangChain, the ReAct paper, Simon Willison), plus 39 more, each independently verified.