The AI Agent Breakthrough That Finally Teaches Machines to Plan Before They Click
Plan-and-Act splits thinking from doing, and the results change what long-horizon agents can actually finish.

Plan-and-Act: separating planning from execution for long-horizon agents
Long-horizon tasks require an agent to hold a goal across 5 to 50+ actions. The agent must track intermediate state and adapt when the environment changes. Examples include buying the highest-rated Bluetooth earphones under $50 on an e-commerce site, adding user authentication to a Flask app, analyzing Q3 sales data and writing a report, researching Rust stream-processing frameworks, automating employee onboarding across email, permissions, Slack, and HR systems, and building a house in Minecraft.
ReAct-style agents fail on these tasks. Zero-shot LLaMA-3.3-70B scores 9.85% on WebArena-Lite. GPT-4o scores 13.9%. Planning is the bottleneck.
Why single-model agents break
Three problems explain the failure.
The first is cognitive overload. ReAct asks one model to track current state, final goal, completed steps, next action, and consequences at the same time. Context grows. Attention dilutes. Early strategy gets buried. The model forgets what it decided ten steps ago.
The second is the absence of an explicit global plan. CoT reasons but cannot act. ReAct acts but decides step by step with no global view. Tree of Thoughts searches but costs too much and cannot adapt to dynamic environments. Planning and execution are fused. For short tasks this is fine. For long tasks it collapses.
The third is error propagation. One wrong click early poisons every later step. Reflexion tries to fix this with self-reflection, but that requires multiple full runs. For long-horizon tasks, that is too expensive, and reflection quality depends on the model's self-evaluation.
Plan-and-Act separates planning from execution. The Planner decides what to do. The Executor decides how to do it.
Architecture
The architecture has two modules. The Planner takes the user query and initial HTML and returns a structured step list. The Executor takes the plan, current HTML, and history and returns one environment action.
Planner outputs steps with two fields: Reasoning and Step. Reasoning explains why this step matters. Step describes what to do. For example: "Reasoning: need to log in before accessing user settings; Step: log in with saved credentials."
This format gives the Executor enough guidance without forcing it to re-derive the global goal. The Executor does not need the whole plan. It needs the next action.
Dynamic replanning
A perfect initial plan cannot handle real environments. The plan might say "follow top contributors," but who they are only becomes clear during execution. Pages change. Options differ. Search results surprise.
Plan-and-Act calls the Planner again after each action. The plan absorbs new information. This mechanism adds 10.31 percentage points on WebArena-Lite, moving from 43.63% to 53.94%. It is the largest single improvement in the paper.
Dynamic replanning also acts as a memory system. Key information survives through continuous plan updates. No separate memory module is required.
Synthetic data pipeline
Planner training data is scarce. One trajectory yields about one plan sample but eight Executor samples. The paper builds a four-stage pipeline.
First, action trajectory collection. GPT-4o generates diverse queries. WebRL-Llama-3.1-70B executes them. An ORM filters successful trajectories. The result is 923 successful trajectories.
Second, grounded plan generation. The pipeline reverse-engineers high-level plans from successful trajectories. Direct LLM plan generation often mismatches execution. Reverse-engineered plans align with what actually worked.
Third, plan expansion. GPT-4o generates 10,000 additional query-plan pairs from existing ones. This solves Planner data scarcity.
Fourth, targeted enhancement. The authors analyze validation failures, classify failure modes, and generate 5,000 more samples for hard cases. This is curriculum learning for planning.
Results
WebArena-Lite success rate reaches 57.58%. Previous SOTA WebRL reaches 49.1%.
The ablation gains matter more than the headline number. Dynamic replanning adds 10.31 percentage points. Plan expansion with 10,000 samples adds 9.10. Adding an untrained base Planner adds 4.36. Targeted enhancement with 5,000 samples adds 4.23. CoT reasoning adds 3.64. More executor trajectory data alone adds 0.61.
Planning is the bottleneck. Adding executor data barely moves the number. Improving the Planner moves it a lot.
Planner-Executor alignment
Attaching an untrained Planner to a trained Executor drops performance from 36.36% to 17.16%. The Executor expects a specific plan format. Untrained Planner output does not match. Both modules must be trained together or share a strict format standard.
Small model plus CoT
An 8B model with CoT matches a 70B model without CoT. On WebVoyager, QWQ-32B reaches 81.36% text-mode SOTA. This makes cost-sensitive deployment viable. Use a large model for the Planner and a small model for the Executor. The Planner runs less often. The Executor runs every step.
Comparison with other frameworks
ReAct is simple and general but decides locally. ToT explores but costs too much and cannot adapt. Reflexion improves through multiple full runs. AutoGPT is conceptually close but lacks dedicated training. WebRL improves execution but does not separate planning. Plan-and-Act chooses a pragmatic path: one plan path, dynamic replanning, and dedicated Planner training.
Practical guidance for agent developers
Use plan-execution separation for tasks over 5 steps. You do not need two separate models. Prompt one model to generate a plan first, then execute. The key is forcing global thinking before action.
Dynamic replanning is mandatory. Do not assume the initial plan survives contact. The simplest approach is to replan every N steps, feeding current state and completed steps as input.
Generate data by reverse-engineering plans from real trajectories. Do not let an LLM invent plans from scratch. Collect successful trajectories, extract plans, then expand with a strong model. This keeps plans aligned with execution.
Unify plan format between Planner and Executor. Mismatched formats destroy performance. Standardize step granularity, state description, and output schema.
Use a large model for the Planner and a small model for the Executor. The Planner runs less often. The Executor runs every step. CoT lets small models approach large-model quality.
Use failure cases for curriculum learning. Classify failure modes. Generate targeted data. Uniform expansion is less efficient.
Test small model plus CoT. 8B plus CoT can match 70B without CoT. This is a direct cost reduction path.
Limitations
The method is validated mainly on web navigation. Robot control, game AI, and other structured decision environments are untested. The synthetic data pipeline depends on strong models like GPT-4o and seed successful trajectories. 70B fine-tuning requires a lot of compute. Dynamic replanning uses LoRA, not full fine-tuning. Multimodal integration with vision is not fully explored. Combining synthetic data with RL methods like WebRL is a future direction.
Conclusion
Plan-and-Act's contribution is a validated architecture pattern: explicit planning, execution separation, dynamic replanning, and a trainable data pipeline.
The core insight is simple. Do not make one model carry all cognitive load. The Planner decides what to do. The Executor decides how to do it. Think globally first, then execute locally.
Long-horizon agents fail because they lose the plot. Clicking is rarely the problem. Plan-and-Act makes planning a first-class citizen. That is the step from demo to production tool.
About the Creator
Jin
Writer of reamstories
https://reamstories.com/jin
Enjoyed the story? Support the Creator.
Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.
Comments
There are no comments for this story
Be the first to respond and start the conversation.