Everyone online means something a little different by "agent loop." This is the version that's simple, correct, and something you can build tonight.
Watch five videos and you'll see five different pictures. They're all the same idea from different angles. Here are the four you'll bump into most:
The same three beats, older words.
The simplest version. A model calling tools, over and over.
The same loop, just left to run unattended toward a goal.
A lead agent hands tasks to sub-agents.
Strip the jargon and every agent loop is the same three beats, reason → act → observe, repeating toward a goal until a "done?" check says stop.
Think of a smart intern you don't micromanage. You hand them the goal. They figure out the next step, do it, check their own work, and go again, only coming back when it's actually done, or stuck. An agent loop is that, in software.
Each lap is one step. The agent keeps circling (act, observe, act, observe) until the "done?" check passes, or a guardrail (a safety limit, like a max number of tries) stops it. The Observe beat is the mechanism that makes a loop work: it reads its own output instead of assuming it worked. That's exactly why the "done?" check (coming up next) is everything.
A loop is worth the setup only when a couple of things are true. Run your task down this ladder before you build anything.
Start with the first one. Reach for the others only when a single agent genuinely can't keep up.
One agent runs the loop on one task. Easiest to build and debug.
One agent does the work, a second grades it: a fresh AI given only the grading job, so it can't rubber-stamp its own work.
A lead splits the goal and hands pieces to sub-agents working in parallel.
The loop will happily produce confident, polished, wrong work and call it finished, unless you tell it exactly what "done" means and give it a way to check.
Lock your finish line: answer these two first
These two answers are the "done?" check from the loop diagram.
Write it so a machine could check it. "The tests pass." "Under 50 words and mentions the price." Not "make it good."
Different jobs need different checks. Pick the one that fits 👇
You don't need a framework. If you have Claude Code (a free AI assistant you run on your computer, and it already works as a loop), just tell it the goal, what "done" means, and how to check. The same pattern works in any agent tool. Copy this:
One clear sentence. "Fix the failing test." Not "improve the code."
↳ that's the GOAL boxGive it a finish line it can verify itself, plus a hard stop so it can't run forever.
↳ that's the DONE? checkSee where it trips. Fix the instructions, not the output. Then let it run on its own.
↳ that's reason → act → observeYou've already met a few of these (tagged recap). Here they are as one checklist to run before you trust a loop to go on its own. Most loops fail for boring reasons: they run forever, burn money, or ship junk. The good ones get these right.
Define "done" so a machine could verify it, not "make it good."
tests pass / under 50 wordsMax tries, a budget, or a time limit. Always. So it can't run forever.
stop after 5 triesThe actions it can take must be reliable and clearly described.
search · run code · edit a fileKeep the history, but summarize it so the context doesn't bloat.
a running notes fileGenerate → judge → fix → repeat. The maker never grades itself.
a 2nd agent scores itBig, multi-step job? Have it write a plan before acting. Small job? Skip it.
outline before buildingSave every thought, action, and result so you can debug at 3am.
log each stepLoops burn tokens fast. Start small and bounded, then scale.
watch the first runsSix named here (Anthropic, OpenAI, Google, LangChain, the ReAct paper, Simon Willison), plus 39 more, each independently verified.