Guide home/Research directions
中文
Research directions

Agents: getting models to do things

Asking a model to explain an error is different from asking it to open a project, locate files, change code, and run tests. The latter requires observing the environment, choosing the next action, and judging whether the job is done. Many agent research questions live in that process.

An agent places a model inside an ongoing process: receive a task, inspect available information, choose an action, call a tool, and use the result to decide what happens next. A search followed by an answer and a coding loop ending in passing tests can both be studied this way. The key question is whether observations actually change later behavior.

Research covers more than the model: tool interfaces, memory, task decomposition, recovery, and evaluation all matter. Keep execution traces and distinguish a bad model decision from a tool failure or an incorrect evaluator.

Research tasks
  • Connect retrieval, calculation, and code execution to information and capabilities outside model parameters.
  • Complete tasks requiring several operations, such as locating a bug, editing code, and verifying the change.
  • Use feedback to revise actions and study reliability and cost on longer tasks.
Tools inside an agent task
  1. 01Task input
  2. 02Planning
  3. 03Act and observe
  4. 04Response

Research milestones

  1. 2023

    ReAct: connect reasoning, actions, and observations

    Choose actions from current information and continue reasoning after observing their results. Start here to understand a trajectory.

  2. 2023

    Toolformer: learn when to call tools

    Tool use can be part of training data, rather than only instructions in a prompt.

  3. 2023

    Reflexion: reuse feedback in another attempt

    Turn feedback into written memory for subsequent attempts. This improvement does not require weight updates.

  4. 2024

    From SWE-bench to SWE-agent: working in real repositories

    SWE-bench evaluates changes against real issues and tests. In 2024, SWE-agent and OpenHands explored how editing interfaces, command execution, browsers, and sandboxes support complete tasks.

  5. 2025

    Training models to use tools

    Search-R1 brings search into multi-turn reinforcement learning; Agent Lightning connects execution traces to training. Learning actions from task feedback brings trajectory data, rewards, and credit assignment into focus.

  6. 2025

    Coding agents enter everyday development

    Tools such as Claude Code and Codex connect reading repositories, editing files, and running tests into an ongoing process. Models and harnesses improve together, bringing context management, tool interfaces, and verification into everyday development.

  7. 2026

    Ongoing collaboration, personal memory, and research workflows

    Muse and Dots extend tasks to cloud computers and everyday apps. Hermes develops cross-session memory and reusable skills, while ARIS connects execution and review in research. These projects explore how to sustain long tasks and reuse past experience.

Key concepts

Tool call
The model emits a tool name and arguments; an external program executes them and returns a result.
Trajectory
The sequence of observations, actions, and results in a task. It helps locate the first error instead of judging only the final answer.
Context and memory
Context is what the model currently sees. Memory persists across steps or attempts and must be retrieved to affect the current decision.
Scaffold
The surrounding program manages tools, prompts, loops, stopping conditions, and feedback. The same model can behave differently in different scaffolds.

ReAct: reasoning, action and feedback

On the ReAct project page, follow an example alternating reasoning, action, and environmental feedback. Notice whether new information changes the next action.

Then choose a basic exercise from Hello Agents or the Hugging Face Agents Course. Build a small system whose tool calls you can inspect and keep its full trace. For Python or API difficulties, return to foundations.

What can you study within agent research?

The same agent task can be studied through decisions, coordination, data, training, evaluation, or system design. These directions often overlap: recovering from failure on a long task might involve changing memory, adding training examples, or adjusting a tool interface.

Planning, tools, and memory

As tasks get longer, do early errors propagate, does important information disappear, and can the system replan after failure? When a model misuses an available tool, the description, visible information, call format, or action-selection ability may be responsible. Inspect what it actually saw at each step, then study which memories to retain, when to call a tool, and when to change the plan. These questions also lead to context management, recovery, and long-task evaluation.

  • Reflexion: follow a failure as it becomes written memory and influences another attempt, making the role of feedback in decisions concrete.
  • Voyager: inspect task selection, code revision, and the skill library in Minecraft to see how later tasks reuse successful behavior.

Multi-agent systems (MAS): division of work, communication, and organization

Do multiple agents help? Identify work that benefits from division, what communication adds, and its time and cost. Research questions include who messages whom, how much context to share, who checks results, and how disagreements are resolved. Compare a simpler system at the same budget to understand the source of gains. Another line studies collective behavior: how individual memories and interactions shape information spreading and social relationships.

  • AutoGen: examine configurable roles and conversation patterns to understand how agents complete tasks together and explore ways to divide work and communicate.
  • Generative Agents: trace a party invitation through a town's social network to connect individual behavior with collective outcomes.

Trajectory data: which experiences teach an agent?

Agent training examples can contain observations, actions, tool responses, and final results. Which tasks to collect, how to select successful trajectories, whether failures and retries help, and how to mix data from different tasks are all research questions. Follow a complete trajectory to check what the model could see, what it was trained to predict, and how tool responses enter training.

  • Toolformer: inspect how candidate tool calls are generated, executed, and filtered to understand how a few demonstrations can support tool-use data generation.
  • AgentTuning: look at how its AgentInstruct interaction trajectories are mixed with general instructions, and how transfer to unseen agent tasks is evaluated.

Agent RL: improving behavior through feedback

How can task feedback teach a model to choose better actions? This connects training data, rewards, RL, and automated evaluation. Success on a multi-step task may depend on a much earlier search or tool choice. Questions include how to assign rewards across steps, explore new actions, and handle environment responses during training. One concrete question is whether a model learns to search more effectively or simply makes more calls.

  • Search-R1: study how search enters multi-turn RL training, focusing on outcome rewards and the treatment of retrieved text, then compare actual search trajectories.
  • Agent Lightning: see how agent execution records become training data and how credit assignment connects task feedback to particular decisions.

Environments and evaluation: did the task actually get done?

Agent research includes building environments for repeated interaction and methods for judging completion. Can a task be reset, are initial conditions consistent, and does the evaluator recognize a partially finished result? These choices affect whether comparisons are reliable. Alongside overall success, study where failures occur, how many calls a task takes, and what changes on a different type of task.

  • AgentBench: explore interactive settings such as operating systems and databases, comparing their task definitions, evaluation methods, and typical failures.
  • WebArena: examine reproducible website environments and evaluation based on task completion to understand how browser actions become comparable experiments.

Harness and infrastructure

The program surrounding a model determines which tools it can use, how results return, how failures are retried, and when a task ends. Research can focus on tool interfaces and context organization, or on sandboxes, parallel execution, logs, and state recovery. Hold the model fixed and compare success, time, and cost under different interfaces or execution methods to understand the effect of system design.

  • SWE-agent: inspect interfaces for viewing files, editing, and execution to understand how interface design affects the process of completing coding tasks.
  • OpenHands: see how code execution, a command line, a browser, and sandboxes form a platform, and how the runtime and evaluation support agent experiments.

What agents are doing: from code to everyday life

Real agents combine tools, memory, planning, and feedback. Coding assistants, personal assistants, and research tools put these capabilities to work on different tasks.

Coding agents: working in a repository

Codex ↗

Codex reads projects, edits code, runs tests, and reviews changes, with support for parallel tasks and ongoing background work. Follow a bug fix to see how it locates files, checks an edit, and responds to test results.

Claude Code ↗

Claude Code reads repositories, edits files, and runs commands from the terminal, connecting code exploration, implementation, and testing. Its tools, Skills, and project instructions show how a coding agent can adapt to a team’s workflow.

Kimi Code ↗

Kimi Code works through the terminal and editor to search code, change a project, run commands, and adjust its next steps from feedback. A small feature and its execution trace show how code understanding, tools, and verification work together.

ZCode ↗

ZCode is a coding-agent workspace for GLM that brings projects, conversations, and task execution together. It uses AGENTS.md for project conventions, making it useful for studying how that context shapes edits and checks.

Pi ↗

Pi is a lightweight, extensible agent harness built around reading and writing files and executing commands. Extensions, Skills, and prompt templates let you adapt its workflow; the code is a useful entry into agent loops and tool interfaces.

Personal agents: remembering and following through

Muse ↗

Meta’s Muse has its own cloud computer and browser and connects to everyday apps for email, travel, and longer-term goals. It brings personal preferences and ongoing tasks together, linking memory, cross-app actions, and task state.

Dots ↗

OpenAI’s Dots are always-on agents that use cloud computers and connected apps to work on tasks and learn user preferences from conversations and feedback. They connect ongoing collaboration with long-term context, background execution, and learning from feedback.

Hermes ↗

Hermes is an open-source personal agent from Nous Research, accessible through the terminal and messaging apps. It keeps memory across sessions and turns methods developed during tasks into reusable skills; its implementation offers a concrete way to study both.

Research agents: connecting literature and experiments

ARIS ↗

ARIS uses Skills to organize machine-learning research, including literature review, idea development, experiments, analysis, and paper revision. An executor advances the work while another model reviews it, with feedback feeding the next round—a concrete multi-agent research workflow.

Health agents: supporting everyday health needs

蚂蚁阿福 / Ant A-Fu ↗

Ant A-Fu brings together health questions, report interpretation, health records, and access to healthcare services. The application connects longitudinal personal information, specialist knowledge, and services, making information organization and task workflows concrete in a specific domain.

Working with AI · Building a research workflow

Evaluate an agent on retrieval tasks

Give the system a set of similar tasks, such as retrieving specified facts from public documents with sources. Count successes, incorrect citations, missing information, and tool failures. Keep tasks and conditions fixed across changes and check whether failures decrease.

For coding agents, see SWE-bench; AgentBench above also provides environment and evaluation code. Read their task and evaluation definitions separately.

Agent overviews and paper lists

Lilian Weng's LLM Powered Autonomous Agents introduces an early component-based view. The LLM Agent Paper List helps once you have a question to follow.

For using agents in your own research, see working with AI. For studying why they fail, see experiments and evaluation.

Evaluation analysis and multi-agent organization

Chai's research notes include AutomationBench scoring analysis and a simulation study of multi-agent organization. More tutorials and paper lists are in the agent catalog.

Representative papers

Start with ReAct, tool use, feedback, and evaluation. Generative Agents and Voyager explore memory and accumulating skills in a town and in Minecraft.

2023 · ICLR 2023

ReAct: Synergizing Reasoning and Acting in Language Models

ReAct: Synergizing Reasoning and Acting in Language Models · Yao et al. · Fig. 1
Figure excerpt from the paper · Yao et al. · Fig. 1 · Original paper

The problem

Can a model gather external information while adjusting its approach?

The key idea

Interleave reasoning text and actions, then add environmental observations to context.

Why this paper

It makes task execution visible as a sequence you can inspect.

Where to start

Start with the project's HotpotQA example and then its ALFWorld failure.

A question to keep asking

Models can ignore feedback or repeat mistakes. Written reasoning is not a complete account of internal computation.

2023 · NeurIPS 2023

Toolformer: Language Models Can Teach Themselves to Use Tools

Toolformer: Language Models Can Teach Themselves to Use Tools · Schick et al., 2023 · Figure 2
Figure excerpt from the paper · Schick et al., 2023 · Figure 2 · Original paper

The problem

How can tool-use training data be obtained without labeling every call?

The key idea

Generate candidate calls, execute them, retain those that improve subsequent text prediction, and fine-tune on the resulting data.

Why this paper

Contrast it with ReAct to separate execution design from trained capability.

Where to start

Figure 2 shows sampling, execution, and filtering; then inspect tool examples.

A question to keep asking

The original method has limitations with chained calls and interactive tools. Better text prediction is not the same as long-task success.

2023 · NeurIPS 2023

Reflexion: Language Agents with Verbal Reinforcement Learning

Reflexion: Language Agents with Verbal Reinforcement Learning · Shinn et al., 2023 · Figure 2(a)
Figure excerpt from the paper · Shinn et al., 2023 · Figure 2(a) · Original paper

The problem

How can another attempt use lessons from a failure?

The key idea

Turn task feedback into written reflections, store them, and supply them to subsequent attempts.

Why this paper

It provides a concrete contrast between changing context and changing parameters.

Where to start

Follow execution, evaluation, reflection, and the next attempt through the method.

A question to keep asking

Reflections need useful feedback. This is not traditional reinforcement-learning training through gradient updates.

2024 · ICLR 2024

SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

SWE-bench: Can Language Models Resolve Real-World GitHub Issues? · Jimenez et al., 2024 · Figure 1
Figure excerpt from the paper · Jimenez et al., 2024 · Figure 1 · Original paper

The problem

Can a model turn a repository and issue into a working fix?

The key idea

Collect real issues and fixes, provide the repository and problem, and evaluate proposed patches with tests.

Why this paper

It joins reading, localization, editing, and verification in one task.

Where to start

Start with task construction and evaluation, then a failure case.

A question to keep asking

The original data comes from Python repositories. Passing tests does not cover all software quality or establish transfer to other languages.

2023 · UIST 2023

Generative Agents: Interactive Simulacra of Human Behavior

Generative Agents: Interactive Simulacra of Human Behavior · Park et al. · Smallville
Figure excerpt from the paper · Park et al. · Smallville · Original paper

The problem

How can 25 town residents remember experiences, plan their days, and interact?

The key idea

Record observations in language, retrieve memories by relevance, importance, and recency, then reflect and plan.

Why this paper

Follow how a party invitation spreads through individual memories, conversations, and actions.

Where to start

Start with Smallville and the Valentine’s Day party, then study memory retrieval, reflection, and planning.

2023 · arXiv 2023

Voyager: An Open-Ended Embodied Agent with Large Language Models

Voyager: An Open-Ended Embodied Agent with Large Language Models · Wang et al. · Fig. 1
Figure excerpt from the paper · Wang et al. · Fig. 1 · Original paper

The problem

How can an agent build on what it knows, from chopping wood to mining diamonds?

The key idea

An automatic curriculum proposes tasks. GPT-4 writes and revises executable code using feedback; successful skills are stored for retrieval and reuse.

Why this paper

The tech tree makes cumulative progress tangible: gathering, crafting, and tool use support later skills.

Where to start

Start with the tech-tree exploration figure, then study the curriculum, skill library, and code-refinement loop.

Getting started

Stuck on a step? Bring your attempt to the AMA ↗