Asking a model to explain an error is different from asking it to open a project, locate files, change code, and run tests. The latter requires observing the environment, choosing the next action, and judging whether the job is done. Many agent research questions live in that process.
An agent places a model inside an ongoing process: receive a task, inspect available information, choose an action, call a tool, and use the result to decide what happens next. A search followed by an answer and a coding loop ending in passing tests can both be studied this way. The key question is whether observations actually change later behavior.
Research covers more than the model: tool interfaces, memory, task decomposition, recovery, and evaluation all matter. Keep execution traces and distinguish a bad model decision from a tool failure or an incorrect evaluator.
- Connect retrieval, calculation, and code execution to information and capabilities outside model parameters.
- Complete tasks requiring several operations, such as locating a bug, editing code, and verifying the change.
- Use feedback to revise actions and study reliability and cost on longer tasks.
- 01Task input
- 02Planning
- 03Act and observe
- 04Response
Research milestones
- 2023
ReAct: connect reasoning, actions, and observations
Choose actions from current information and continue reasoning after observing their results. Start here to understand a trajectory.
- 2023
Toolformer: learn when to call tools
Tool use can be part of training data, rather than only instructions in a prompt.
- 2023
Reflexion: reuse feedback in another attempt
Turn feedback into written memory for subsequent attempts. This improvement does not require weight updates.
- 2024
From SWE-bench to SWE-agent: working in real repositories
SWE-bench evaluates changes against real issues and tests. In 2024, SWE-agent and OpenHands explored how editing interfaces, command execution, browsers, and sandboxes support complete tasks.
- 2025
Training models to use tools
Search-R1 brings search into multi-turn reinforcement learning; Agent Lightning connects execution traces to training. Learning actions from task feedback brings trajectory data, rewards, and credit assignment into focus.
- 2025
Coding agents enter everyday development
Tools such as Claude Code and Codex connect reading repositories, editing files, and running tests into an ongoing process. Models and harnesses improve together, bringing context management, tool interfaces, and verification into everyday development.
- 2026
Ongoing collaboration, personal memory, and research workflows
Muse and Dots extend tasks to cloud computers and everyday apps. Hermes develops cross-session memory and reusable skills, while ARIS connects execution and review in research. These projects explore how to sustain long tasks and reuse past experience.
Key concepts
- Tool call
- The model emits a tool name and arguments; an external program executes them and returns a result.
- Trajectory
- The sequence of observations, actions, and results in a task. It helps locate the first error instead of judging only the final answer.
- Context and memory
- Context is what the model currently sees. Memory persists across steps or attempts and must be retrieved to affect the current decision.
- Scaffold
- The surrounding program manages tools, prompts, loops, stopping conditions, and feedback. The same model can behave differently in different scaffolds.
ReAct: reasoning, action and feedback
On the ReAct project page, follow an example alternating reasoning, action, and environmental feedback. Notice whether new information changes the next action.
Then choose a basic exercise from Hello Agents or the Hugging Face Agents Course. Build a small system whose tool calls you can inspect and keep its full trace. For Python or API difficulties, return to foundations.
What can you study within agent research?
The same agent task can be studied through decisions, coordination, data, training, evaluation, or system design. These directions often overlap: recovering from failure on a long task might involve changing memory, adding training examples, or adjusting a tool interface.
Planning, tools, and memory
As tasks get longer, do early errors propagate, does important information disappear, and can the system replan after failure? When a model misuses an available tool, the description, visible information, call format, or action-selection ability may be responsible. Inspect what it actually saw at each step, then study which memories to retain, when to call a tool, and when to change the plan. These questions also lead to context management, recovery, and long-task evaluation.
- Reflexion: follow a failure as it becomes written memory and influences another attempt, making the role of feedback in decisions concrete.
- Voyager: inspect task selection, code revision, and the skill library in Minecraft to see how later tasks reuse successful behavior.
Multi-agent systems (MAS): division of work, communication, and organization
Do multiple agents help? Identify work that benefits from division, what communication adds, and its time and cost. Research questions include who messages whom, how much context to share, who checks results, and how disagreements are resolved. Compare a simpler system at the same budget to understand the source of gains. Another line studies collective behavior: how individual memories and interactions shape information spreading and social relationships.
- AutoGen: examine configurable roles and conversation patterns to understand how agents complete tasks together and explore ways to divide work and communicate.
- Generative Agents: trace a party invitation through a town's social network to connect individual behavior with collective outcomes.
Trajectory data: which experiences teach an agent?
Agent training examples can contain observations, actions, tool responses, and final results. Which tasks to collect, how to select successful trajectories, whether failures and retries help, and how to mix data from different tasks are all research questions. Follow a complete trajectory to check what the model could see, what it was trained to predict, and how tool responses enter training.
- Toolformer: inspect how candidate tool calls are generated, executed, and filtered to understand how a few demonstrations can support tool-use data generation.
- AgentTuning: look at how its AgentInstruct interaction trajectories are mixed with general instructions, and how transfer to unseen agent tasks is evaluated.
Agent RL: improving behavior through feedback
How can task feedback teach a model to choose better actions? This connects training data, rewards, RL, and automated evaluation. Success on a multi-step task may depend on a much earlier search or tool choice. Questions include how to assign rewards across steps, explore new actions, and handle environment responses during training. One concrete question is whether a model learns to search more effectively or simply makes more calls.
- Search-R1: study how search enters multi-turn RL training, focusing on outcome rewards and the treatment of retrieved text, then compare actual search trajectories.
- Agent Lightning: see how agent execution records become training data and how credit assignment connects task feedback to particular decisions.
Environments and evaluation: did the task actually get done?
Agent research includes building environments for repeated interaction and methods for judging completion. Can a task be reset, are initial conditions consistent, and does the evaluator recognize a partially finished result? These choices affect whether comparisons are reliable. Alongside overall success, study where failures occur, how many calls a task takes, and what changes on a different type of task.
- AgentBench: explore interactive settings such as operating systems and databases, comparing their task definitions, evaluation methods, and typical failures.
- WebArena: examine reproducible website environments and evaluation based on task completion to understand how browser actions become comparable experiments.
Harness and infrastructure
The program surrounding a model determines which tools it can use, how results return, how failures are retried, and when a task ends. Research can focus on tool interfaces and context organization, or on sandboxes, parallel execution, logs, and state recovery. Hold the model fixed and compare success, time, and cost under different interfaces or execution methods to understand the effect of system design.
- SWE-agent: inspect interfaces for viewing files, editing, and execution to understand how interface design affects the process of completing coding tasks.
- OpenHands: see how code execution, a command line, a browser, and sandboxes form a platform, and how the runtime and evaluation support agent experiments.
What agents are doing: from code to everyday life
Real agents combine tools, memory, planning, and feedback. Coding assistants, personal assistants, and research tools put these capabilities to work on different tasks.
Coding agents: working in a repository
Codex ↗
Codex reads projects, edits code, runs tests, and reviews changes, with support for parallel tasks and ongoing background work. Follow a bug fix to see how it locates files, checks an edit, and responds to test results.
Claude Code ↗
Claude Code reads repositories, edits files, and runs commands from the terminal, connecting code exploration, implementation, and testing. Its tools, Skills, and project instructions show how a coding agent can adapt to a team’s workflow.
Kimi Code ↗
Kimi Code works through the terminal and editor to search code, change a project, run commands, and adjust its next steps from feedback. A small feature and its execution trace show how code understanding, tools, and verification work together.
ZCode ↗
ZCode is a coding-agent workspace for GLM that brings projects, conversations, and task execution together. It uses AGENTS.md for project conventions, making it useful for studying how that context shapes edits and checks.
Pi ↗
Pi is a lightweight, extensible agent harness built around reading and writing files and executing commands. Extensions, Skills, and prompt templates let you adapt its workflow; the code is a useful entry into agent loops and tool interfaces.
Personal agents: remembering and following through
Muse ↗
Meta’s Muse has its own cloud computer and browser and connects to everyday apps for email, travel, and longer-term goals. It brings personal preferences and ongoing tasks together, linking memory, cross-app actions, and task state.
Dots ↗
OpenAI’s Dots are always-on agents that use cloud computers and connected apps to work on tasks and learn user preferences from conversations and feedback. They connect ongoing collaboration with long-term context, background execution, and learning from feedback.
Hermes ↗
Hermes is an open-source personal agent from Nous Research, accessible through the terminal and messaging apps. It keeps memory across sessions and turns methods developed during tasks into reusable skills; its implementation offers a concrete way to study both.
Research agents: connecting literature and experiments
ARIS ↗
ARIS uses Skills to organize machine-learning research, including literature review, idea development, experiments, analysis, and paper revision. An executor advances the work while another model reviews it, with feedback feeding the next round—a concrete multi-agent research workflow.
Health agents: supporting everyday health needs
蚂蚁阿福 / Ant A-Fu ↗
Ant A-Fu brings together health questions, report interpretation, health records, and access to healthcare services. The application connects longitudinal personal information, specialist knowledge, and services, making information organization and task workflows concrete in a specific domain.
Working with AI · Building a research workflow
Evaluate an agent on retrieval tasks
Give the system a set of similar tasks, such as retrieving specified facts from public documents with sources. Count successes, incorrect citations, missing information, and tool failures. Keep tasks and conditions fixed across changes and check whether failures decrease.
For coding agents, see SWE-bench; AgentBench above also provides environment and evaluation code. Read their task and evaluation definitions separately.
Agent overviews and paper lists
Lilian Weng's LLM Powered Autonomous Agents introduces an early component-based view. The LLM Agent Paper List helps once you have a question to follow.
For using agents in your own research, see working with AI. For studying why they fail, see experiments and evaluation.
Evaluation analysis and multi-agent organization
Chai's research notes include AutomationBench scoring analysis and a simulation study of multi-agent organization. More tutorials and paper lists are in the agent catalog.
Representative papers
Start with ReAct, tool use, feedback, and evaluation. Generative Agents and Voyager explore memory and accumulating skills in a town and in Minecraft.