MarsDawn On the Mac App Store

Four agent design patterns and the documents each one hands you

In March 2024, Andrew Ng used his newsletter, The Batch, to describe four design patterns for AI agents: reflection, tool use, planning and multi-agent collaboration. They're usually discussed from the builder's side, as ways to get better results out of a model. This post looks from the other side. If you use an agent built on one of these patterns, what lands in your folder, and what should you read first?

The four patterns are Andrew Ng's. Which documents each one tends to hand you, and what to check in them, is our own inference. He doesn't write about either, and he doesn't argue for human review in this series.

The four patterns, briefly

Ng describes them in “Agentic Design Patterns Part 1.” In short: with reflection, the model looks over its own work and improves it. With tool use, it can call tools such as web search or code execution. With planning, it comes up with a multistep plan and carries it out. With multi-agent collaboration, several agents split up the work and discuss it.

In Part 1 he shows the payoff with a coding benchmark, HumanEval, using results his team gathered from several research groups: “GPT-3.5 (zero shot) was 48.1% correct. GPT-4 (zero shot) does better at 67.0%. However, the improvement from GPT-3.5 to GPT-4 is dwarfed by incorporating an iterative agent workflow. Indeed, wrapped in an agent loop, GPT-3.5 achieves up to 95.1%.” Those numbers are about one coding benchmark, and 95.1% is a best case (“up to”). They show agent workflows can improve output. They don't say anything about who checks it.

From here on, the documents and the checks are our reading, not Ng's. Real agents mix patterns, too. A coding agent may plan, run tools and review its own work in one session, so you'll often get all four kinds of file.

1. Reflection: a draft that has already reviewed itself

Ng's post on reflection frames it as automating the feedback a person would otherwise give: “What if you automate the step of delivering critical feedback, so the model automatically criticizes its own output and improves its response?”

What it tends to hand you: a revised document, sometimes with a self-review section or lines like “double-checked the edge cases.”

What to check: the result against your request, not against the agent's own critique. Self-review can go wrong in its own way. Chip Huyen: “An interesting mode of planning failure is caused by errors in reflection. The agent is convinced that it’s accomplished a task when it hasn’t.” Lilian Weng, writing on her blog Lil’Log in June 2023 while at OpenAI, about models of that time: “The lack of expertise may cause LLMs not knowing its flaws and thus cannot well judge the correctness of task results.” (In the study she was describing, an LLM's evaluation of the results and human experts' evaluation didn't agree.) If it says “verified,” check one thing yourself.

2. Tool use: a report of what ran

What it tends to hand you: a summary of what the agent ran or searched and what came back. “Ran the test suite: all passing.” A results table. Links it found.

Anthropic's guide describes tool results as the agent's own check on itself: “During execution, it's crucial for the agents to gain “ground truth” from the environment at each step (such as tool call results or code execution) to assess its progress.” That check happens inside the agent. What reaches you is the agent's retelling of it.

What to check: that each claim traces back to output you can see. Match one number in the summary to the real output. Open one of the links.

3. Planning: plan.md

What it tends to hand you: a plan, a spec, a task list with checkboxes the agent ticks as it goes.

Ng is candid about this pattern in Part 4:

“On one hand, Planning is a very powerful capability; on the other, it leads to less predictable results. In my experience, while I can get the agentic design patterns of Reflection and Tool Use to work reliably and improve my applications’ performance, Planning is a less mature technology, and I find it hard to predict in advance what it will do.”

He's optimistic, too: “But the field continues to evolve rapidly, and I'm confident that Planning abilities will improve quickly.”

What to check: the plan before it runs, using the five-minute review: shape, one claim, steps that can't be undone, diagrams, scope. If the agent rewrites the plan midway, compare it with the version you approved; if it's in git, git diff plan.md shows what changed. In MarsDawn, the Outline tab shows a long plan's shape, and a rewritten plan reloads without losing your place, as long as you have no unsaved edits of your own.

4. Multi-agent collaboration: several files, several authors

What it tends to hand you: a spec from one agent, implementation notes from another, a review from a third, and summaries passed between them. Sometimes each works in its own branch or worktree.

What to check: the handoffs. Where one agent summarizes another's work, look for a requirement that didn't make it across. Look for two files that disagree, and decide which one is the source of truth before anyone builds on the other. In MarsDawn, open the shared folder with File ▸ Open Folder… (⇧⌘O): new files show up in the Files tab within about a second as the agents write them, and for a git checkout the header names the branch or worktree, so two windows on the same file name from different branches don't look alike. When the result has to go to people who don't read Markdown, Sharing exported PDFs covers that step.

At a glance

Pattern (Ng)What it tends to hand you (our inference)Read first (our suggestion)
ReflectionA revised draft, maybe with a self-reviewThe result against your own request; check one “verified”
Tool useA report of what ran and what came backOne claim traced to real output
Planningplan.md, a spec, a task listThe five-minute review, before it runs
Multi-agent collaborationSeveral files from several agents, maybe on several branchesThe handoffs, and which file is the source of truth

None of the authors quoted here mention MarsDawn, and none of them endorse it or any other Markdown tool. MarsDawn has no AI model inside: it doesn't know which pattern produced a file, and it won't do these checks for you. It keeps the files readable while you do.

Try it

MarsDawn is on the Mac App Store. There's also the free marsdawn command-line tool:

brew install redtear1115/tap/marsdawn

It exports Markdown to PDF without the app: see Markdown to PDF.

Command Line · Know before you buy: What MarsDawn doesn't do

Next

Sources