Guide home/Start learning
中文
Start learning

Choose a paper to read

Representative papers

Choose a field, then a paper: what question did it ask, and how did it test its idea? Each card includes an explanation, a reading starting point, and the original.

49 papers

2012 · NIPS 2012

ImageNet Classification with Deep Convolutional Neural Networks

ImageNet Classification with Deep Convolutional Neural Networks · Krizhevsky et al. · Fig. 2
Figure excerpt from the paper · Krizhevsky et al. · Fig. 2 · Original paper

The problem

Can a network learn recognition features from a large image collection?

The key idea

Stack convolutional layers and combine GPU training with ReLU activations and regularization.

Why this paper

It connects advances in data, compute, and training practice.

Where to start

Start with the architecture and training choices, then inspect classification results and errors.

A question to keep asking

It classifies predefined categories; it does not directly locate objects.

2014 · NIPS 2014

Generative Adversarial Nets

Generative Adversarial Nets · Goodfellow et al., 2014 · Figure 1
Figure excerpt from the paper · Goodfellow et al., 2014 · Figure 1 · Original paper

The problem

How can a model learn to generate without directly specifying a complex image probability?

The key idea

A generator turns random inputs into samples; a discriminator learns to distinguish them from training examples and supplies feedback.

Why this paper

The two-model game gives a clear alternative training principle.

Where to start

Start with both objectives and alternating updates, then the samples.

A question to keep asking

Training can be unstable or cover only some patterns in the data.

2015 · Nature

Human-level control through deep reinforcement learning

Human-level control through deep reinforcement learning · Mnih et al. · Fig. 1
Figure excerpt from the paper · Mnih et al. · Fig. 1 · Original paper

The problem

How can pixels guide discrete actions?

The key idea

Predict action values; replay experience and use a delayed target network to stabilize learning.

Why this paper

A clear entry to observations, values, and deep RL training targets.

Where to start

Trace the interaction loop, Q updates, and Atari evaluation.

A question to keep asking

The original method enumerates discrete actions; continuous control needs other machinery.

2016 · CVPR 2016

Deep Residual Learning for Image Recognition

Deep Residual Learning for Image Recognition · He et al. · Fig. 5
Figure excerpt from the paper · He et al. · Fig. 5 · Original paper

The problem

Why can a deeper network fit even its training data worse?

The key idea

Learn a change F(x) and add the input back, producing x + F(x).

Why this paper

A clear architectural idea is tested against plain networks at different depths.

Where to start

Read the degradation problem and residual block, then compare training errors.

A question to keep asking

Easier optimization does not make arbitrary depth beneficial; data and compute still matter.

2016 · CVPR 2016

You Only Look Once: Unified, Real-Time Object Detection

You Only Look Once: Unified, Real-Time Object Detection · Redmon et al., YOLO, Figure 2
Figure excerpt from the paper · Redmon et al., YOLO, Figure 2 · Original paper

The problem

Can one network pass predict both object categories and locations?

The key idea

Predict boxes and scores from the full image, training recognition and localization together.

Why this paper

It connects directly to a detection demo and its misses, false positives, and runtime.

Where to start

Start with the detection pipeline and grid formulation, then localization errors.

A question to keep asking

The original struggles with small or crowded objects and precise localization. Modern YOLO tools have changed substantially.

2017 · ICML

Neural Message Passing for Quantum Chemistry

Neural Message Passing for Quantum Chemistry · Gilmer et al. · Fig. 1
Figure excerpt from the paper · Gilmer et al. · Fig. 1 · Original paper

The problem

How can atoms and their relationships predict molecular properties?

The key idea

Update node representations through messages, then aggregate them for a molecular prediction.

Why this paper

It grounds graph learning in a concrete chemistry task.

Where to start

Start with message passing, readout, targets, and data.

2017 · NIPS 2017

Attention Is All You Need

Attention Is All You Need · Vaswani et al. · Fig. 1
Figure excerpt from the paper · Vaswani et al. · Fig. 1 · Original paper

The problem

Can a sequence model handle dependencies without recurrent computation?

The key idea

Each position reads other positions through attention, combined with feed-forward layers and positional information.

Why this paper

Later models reuse these components. It is a useful first map of information flow.

Where to start

Start with the architecture and attention explanation, then the translation experiments.

A question to keep asking

The original is an encoder–decoder translation model.

2017 · arXiv

Proximal Policy Optimization Algorithms

Proximal Policy Optimization Algorithms · Schulman et al. (2017), Fig. 1
Figure excerpt from the paper · Schulman et al. (2017), Fig. 1 · Original paper

The problem

How can simple policy updates avoid overly large changes?

The key idea

The clipped objective removes incentives for excessive changes in action-probability ratios.

Why this paper

A practical entry to policy gradients and many RL implementations.

Where to start

Inspect the clipped objective and the sampling/minibatch-update loop.

A question to keep asking

Clipping does not guarantee monotonic improvement; collection and implementation choices still matter.

2018 · Physical Review Letters

Deep Potential Molecular Dynamics: A Scalable Model with the Accuracy of Quantum Mechanics

Deep Potential Molecular Dynamics: A Scalable Model with the Accuracy of Quantum Mechanics · Zhang et al., Deep Potential Molecular Dynamics, Figure 1
Figure excerpt from the paper · Zhang et al., Deep Potential Molecular Dynamics, Figure 1 · Original paper

The problem

How can accurate energy and force calculations become cheaper?

The key idea

Learn local atomic contributions from first-principles data while preserving problem symmetries.

Why this paper

An example of replacing an expensive step in a scientific computation.

Where to start

Trace energies and forces into simulation, then inspect physical comparisons.

A question to keep asking

Accuracy depends on covered configurations and reference calculations; new conditions need testing.

2018 · OpenAI technical report · 2018

Improving Language Understanding by Generative Pre-Training

Improving Language Understanding by Generative Pre-Training · Radford et al., 2018 · Figure 1
Figure excerpt from the paper · Radford et al., 2018 · Figure 1 · Original paper

The problem

How can unlabeled text help supervised language-understanding tasks?

The key idea

Pretrain by predicting text, then fine-tune on task data arranged as input sequences.

Why this paper

It makes the transfer from general representations to particular tasks concrete.

Where to start

Figure 1 shows the training stages and task-specific input transformations.

A question to keep asking

Downstream task fine-tuning is still required.

2018 · ICML

Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor · Haarnoja et al. (2018), Fig. 1
Figure excerpt from the paper · Haarnoja et al. (2018), Fig. 1 · Original paper

The problem

How can experience be reused without committing too early to one action?

The key idea

Optimize reward and policy entropy, using a critic and replayed experience.

Why this paper

Compare with PPO to understand data reuse and exploration objectives.

Where to start

Connect the maximum-entropy objective, actor, critic, and replay buffer.

A question to keep asking

This paper focuses on continuous control; reward and entropy scales affect behavior.

2018 · arXiv

World Models

World Models · Ha & Schmidhuber · Fig. 4
Figure excerpt from the paper · Ha & Schmidhuber · Fig. 4 · Original paper

The problem

Can a learned representation support a small controller?

The key idea

Compress images, learn dynamics with a recurrent model, and feed their representations to a controller.

Why this paper

The interactive article makes the components tangible.

Where to start

Read Agent Model and Car Racing, then the imagined VizDoom experiment.

A question to keep asking

Controllers can exploit model errors; test model-trained behavior in the original environment.

2019 · KDD

Optuna: A Next-generation Hyperparameter Optimization Framework

Optuna: A Next-generation Hyperparameter Optimization Framework · Akiba et al. · Figure 6
Figure excerpt from the paper · Akiba et al. · Figure 6 · Original paper

The problem

How can a limited training budget find a good configuration?

The key idea

Define the search space in code, then combine sampling and pruning to allocate the budget across trials.

Where to start

Start with the search space and trial loop, then examine how early stopping saves computation.

2020 · arXiv

Qlib: An AI-oriented Quantitative Investment Platform

Qlib: An AI-oriented Quantitative Investment Platform · Figure 1
Figure excerpt from the paper · Figure 1 · Original paper

The problem

How do data, models, portfolios, and backtesting form a research workflow?

The key idea

Connect financial data processing and machine learning pipelines to strategy evaluation so that different models can be compared within a common workflow.

Where to start

Trace the data flow through the framework diagram, then identify the time split, labels, and backtesting configuration in an example.

2020 · NeurIPS 2020

Denoising Diffusion Probabilistic Models

Denoising Diffusion Probabilistic Models · Ho et al. · Fig. 1
Figure excerpt from the paper · Ho et al. · Fig. 1 · Original paper

The problem

Can learning denoising tasks produce clear images?

The key idea

Train on known noise added to real images; repeatedly use the learned predictions to sample from an initial random state.

Why this paper

Its training and sampling algorithms provide a useful foundation for later diffusion work.

Where to start

Compare the two algorithms, tracing model inputs and predictions.

A question to keep asking

Original sampling uses many sequential model calls.

2020 · NeurIPS 2020

Language Models are Few-Shot Learners

Language Models are Few-Shot Learners · Brown et al., 2020 · Figure 2.1
Figure excerpt from the paper · Brown et al., 2020 · Figure 2.1 · Original paper

The problem

Can a few examples specify a new task without weight updates?

The key idea

Scale an autoregressive model and evaluate instructions and demonstrations under zero-, one-, and few-shot settings.

Why this paper

It distinguishes in-context learning from fine-tuning and connects to questions about scale.

Where to start

Start with the three evaluation settings, then inspect a familiar task.

A question to keep asking

Results depend on the task and prompt; overlap between web training data and tests complicates evaluation.

2020 · SC 2020

ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

ZeRO: Memory Optimizations Toward Training Trillion Parameter Models · Rajbhandari et al. (2020), Fig. 1
Figure excerpt from the paper · Rajbhandari et al. (2020), Fig. 1 · Original paper

The problem

How much memory is wasted by replicating training state on every GPU?

The key idea

Progressively partition optimizer states, gradients, and parameters, communicating required pieces when needed.

Why this paper

It explains training memory beyond parameter count.

Where to start

Identify what each of the three stages partitions, then compare memory and communication costs.

A question to keep asking

Sharding changes communication. Benefits depend on bandwidth and device count.

2020 · Nature

Mastering Atari, Go, chess and shogi by planning with a learned model

Mastering Atari, Go, chess and shogi by planning with a learned model · Schrittwieser et al. (2020), Fig. 1
Figure excerpt from the paper · Schrittwieser et al. (2020), Fig. 1 · Original paper

The problem

Can a learned model support search without supplied dynamics?

The key idea

Predict reward, value, and policy in internal states, then compare choices with tree search.

Why this paper

Useful models need not reconstruct every pixel.

Where to start

Follow planning, acting, and training before game results.

A question to keep asking

The paper evaluates specific games and action interfaces.

2021 · Nature

Highly accurate protein structure prediction with AlphaFold

Highly accurate protein structure prediction with AlphaFold · Jumper et al., AlphaFold, Figure 1(e) · CC BY 4.0
Figure excerpt from the paper · Jumper et al., AlphaFold, Figure 1(e) · CC BY 4.0 · Original paper

The problem

How can sequence information support three-dimensional structure prediction?

The key idea

Combine evolutionary sequence information and residue relationships, iteratively refining representations and structure.

Why this paper

It connects data, geometry, and system design around a disciplinary problem.

Where to start

Start with the architecture, CASP14 evaluation, and confidence estimates.

2021 · ICLR

Fourier Neural Operator for Parametric Partial Differential Equations

Fourier Neural Operator for Parametric Partial Differential Equations · Figure 1
Figure excerpt from the paper · Figure 1 · Original paper

The problem

How can a set of initial conditions directly predict continuous physical fields?

The key idea

Parameterize integral kernels in the frequency domain to learn mappings from functions to functions for a family of partial differential equation problems.

Where to start

Start with the operator layer and the inputs and outputs, then compare errors across resolutions and equation settings.

2021 · ICML 2021

Learning Transferable Visual Models From Natural Language Supervision

Learning Transferable Visual Models From Natural Language Supervision · Radford et al. · Fig. 1
Figure excerpt from the paper · Radford et al. · Fig. 1 · Original paper

The problem

Can captions teach visual concepts beyond fixed category labels?

The key idea

Train image and text encoders so matched pairs are more similar than mismatches. Compare an image with candidate descriptions at inference.

Why this paper

Retrieval and zero-shot classification become an observable matching problem.

Where to start

Follow pretraining and zero-shot prediction, then change the official example's candidate text.

A question to keep asking

CLIP produces representations or scores, not conversation. Correct matching need not imply relational understanding.

2021 · ICLR 2021

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale · Dosovitskiy et al. · Fig. 1
Figure excerpt from the paper · Dosovitskiy et al. · Fig. 1 · Original paper

The problem

Does image classification require convolutional structure?

The key idea

Turn image patches into vectors and exchange information using a Transformer.

Why this paper

It connects vision to sequence processing and highlights pretraining scale.

Where to start

Follow patches into the model, then compare pretraining dataset sizes.

A question to keep asking

The original strong results rely on large-scale pretraining; they do not imply Transformers always beat CNNs on small datasets.

2022 · arXiv

RT-1: Robotics Transformer for Real-World Control at Scale

RT-1: Robotics Transformer for Real-World Control at Scale · Brohan et al., RT-1, Figure 1(a)
Figure excerpt from the paper · Brohan et al., RT-1, Figure 1(a) · Original paper

The problem

Can one model learn instruction-following from multitask robot data?

The key idea

Condition a Transformer on images and language to predict discretized actions.

Why this paper

It makes data scale and diversity concrete in robotics.

Where to start

Inspect inputs, outputs, data collection, and generalization evaluations.

A question to keep asking

Its results depend on a particular robot body and action interface.

2022 · CVPR 2022

High-Resolution Image Synthesis with Latent Diffusion Models

High-Resolution Image Synthesis with Latent Diffusion Models · Rombach et al. · Fig. 3
Figure excerpt from the paper · Rombach et al. · Fig. 3 · Original paper

The problem

Can diffusion process less information while retaining useful image detail?

The key idea

Compress images with an autoencoder, diffuse in latent space, and decode. Cross-attention introduces conditions such as text.

Why this paper

It links compression, quality, and compute, and underpins Stable Diffusion.

Where to start

Trace encoding, diffusion, and decoding; then compare compression levels.

A question to keep asking

Compression loses information. Text conditioning does not guarantee accurate counts or spatial relations.

2022 · NeurIPS 2022

Training language models to follow instructions with human feedback

Training language models to follow instructions with human feedback · Ouyang et al. · Fig. 2
Figure excerpt from the paper · Ouyang et al. · Fig. 2 · Original paper

The problem

How can a text-completion model better follow a user's request?

The key idea

Fine-tune on demonstrations, learn a reward model from ranked answers, and use reinforcement learning to adjust behavior.

Why this paper

It separates pretrained capability from the way answers are produced.

Where to start

Follow the training pipeline and identify the data needed for demonstrations, ranking, and policy updates.

A question to keep asking

Preferences depend on labelers and prompt distributions; preferred answers can still be factually wrong.

2022 · NeurIPS 2022

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness · Dao et al. · Fig. 1
Figure excerpt from the paper · Dao et al. · Fig. 1 · Original paper

The problem

Can attention be limited by data movement rather than multiplication count?

The key idea

Tile work into fast on-chip memory instead of repeatedly materializing the full attention matrix in device memory.

Why this paper

It shows how the same mathematical operation can have very different execution costs.

Where to start

Start with GPU memory levels and tiling, then IO analysis and runtime.

A question to keep asking

Exact dense attention still has quadratic arithmetic complexity. Speedup depends on shapes, hardware, and implementation.

2023 · ICLR 2023

ReAct: Synergizing Reasoning and Acting in Language Models

ReAct: Synergizing Reasoning and Acting in Language Models · Yao et al. · Fig. 1
Figure excerpt from the paper · Yao et al. · Fig. 1 · Original paper

The problem

Can a model gather external information while adjusting its approach?

The key idea

Interleave reasoning text and actions, then add environmental observations to context.

Why this paper

It makes task execution visible as a sequence you can inspect.

Where to start

Start with the project's HotpotQA example and then its ALFWorld failure.

A question to keep asking

Models can ignore feedback or repeat mistakes. Written reasoning is not a complete account of internal computation.

2023 · NeurIPS 2023

Toolformer: Language Models Can Teach Themselves to Use Tools

Toolformer: Language Models Can Teach Themselves to Use Tools · Schick et al., 2023 · Figure 2
Figure excerpt from the paper · Schick et al., 2023 · Figure 2 · Original paper

The problem

How can tool-use training data be obtained without labeling every call?

The key idea

Generate candidate calls, execute them, retain those that improve subsequent text prediction, and fine-tune on the resulting data.

Why this paper

Contrast it with ReAct to separate execution design from trained capability.

Where to start

Figure 2 shows sampling, execution, and filtering; then inspect tool examples.

A question to keep asking

The original method has limitations with chained calls and interactive tools. Better text prediction is not the same as long-task success.

2023 · NeurIPS 2023

Reflexion: Language Agents with Verbal Reinforcement Learning

Reflexion: Language Agents with Verbal Reinforcement Learning · Shinn et al., 2023 · Figure 2(a)
Figure excerpt from the paper · Shinn et al., 2023 · Figure 2(a) · Original paper

The problem

How can another attempt use lessons from a failure?

The key idea

Turn task feedback into written reflections, store them, and supply them to subsequent attempts.

Why this paper

It provides a concrete contrast between changing context and changing parameters.

Where to start

Follow execution, evaluation, reflection, and the next attempt through the method.

A question to keep asking

Reflections need useful feedback. This is not traditional reinforcement-learning training through gradient updates.

2023 · UIST 2023

Generative Agents: Interactive Simulacra of Human Behavior

Generative Agents: Interactive Simulacra of Human Behavior · Park et al. · Smallville
Figure excerpt from the paper · Park et al. · Smallville · Original paper

The problem

How can 25 town residents remember experiences, plan their days, and interact?

The key idea

Record observations in language, retrieve memories by relevance, importance, and recency, then reflect and plan.

Why this paper

Follow how a party invitation spreads through individual memories, conversations, and actions.

Where to start

Start with Smallville and the Valentine’s Day party, then study memory retrieval, reflection, and planning.

2023 · arXiv 2023

Voyager: An Open-Ended Embodied Agent with Large Language Models

Voyager: An Open-Ended Embodied Agent with Large Language Models · Wang et al. · Fig. 1
Figure excerpt from the paper · Wang et al. · Fig. 1 · Original paper

The problem

How can an agent build on what it knows, from chopping wood to mining diamonds?

The key idea

An automatic curriculum proposes tasks. GPT-4 writes and revises executable code using feedback; successful skills are stored for retrieval and reuse.

Why this paper

The tech tree makes cumulative progress tangible: gathering, crafting, and tool use support later skills.

Where to start

Start with the tech-tree exploration figure, then study the curriculum, skill library, and code-refinement loop.

2023 · arXiv

DreamDiffusion: Generating High-Quality Images from Brain EEG Signals

DreamDiffusion: Generating High-Quality Images from Brain EEG Signals · Figure 1
Figure excerpt from the paper · Figure 1 · Original paper

The problem

How can EEG representations connect to an image generation model?

The key idea

Learn EEG representations through temporally masked signal modeling, combine them with CLIP visual supervision, and use the signals for conditional generation.

Where to start

Trace the role of each data source and supervision signal through the training stages, then examine participant splits and controls.

2023 · Robotics: Science and Systems

Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Diffusion Policy: Visuomotor Policy Learning via Action Diffusion · Chi et al. · Fig. 3
Figure excerpt from the paper · Chi et al. · Fig. 3 · Original paper

The problem

How can a policy represent several valid, coherent action sequences?

The key idea

Denoise an observation-conditioned action sequence, execute part, then observe and predict again.

Why this paper

A tangible bridge from generative models to control.

Where to start

Follow denoising and receding-horizon execution, then Push-T videos.

A question to keep asking

Iterative denoising costs time; demonstration coverage and control frequency constrain deployment.

2023 · CoRL

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control · Zitkovich et al. · Fig. 1
Figure excerpt from the paper · Zitkovich et al. · Fig. 1 · Original paper

The problem

Can web-trained knowledge help robots interpret new instructions?

The key idea

Encode actions as tokens and co-fine-tune with vision-language data.

Why this paper

It connects semantic knowledge to existing action capabilities.

Where to start

Read action encoding, then semantic-generalization tasks.

A question to keep asking

Semantic transfer does not automatically create new low-level motor skills; robot demonstrations remain important.

2023 · ICCV 2023

Scalable Diffusion Models with Transformers

Scalable Diffusion Models with Transformers · Peebles & Xie · Fig. 3
Figure excerpt from the paper · Peebles & Xie · Fig. 3 · Original paper

The problem

Can a scalable Transformer serve as the diffusion network?

The key idea

Process latent patches with a Transformer and compare depth, width, and patch size.

Why this paper

It separates the generation objective from architecture and examines scaling through compute.

Where to start

Read the architecture, then compare compute and generation metrics across sizes.

A question to keep asking

The original studies class-conditioned ImageNet images, not text-to-video generation.

2023 · ICML 2023

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models · Li et al., BLIP-2, Figure 1
Figure excerpt from the paper · Li et al., BLIP-2, Figure 1 · Original paper

The problem

Can existing vision and language models be reused with less retraining?

The key idea

Freeze both models and train a Q-Former to extract a compact set of visual features and connect them to the language model.

Why this paper

It turns the idea of giving a language model vision into modules and training stages.

Where to start

Mark frozen and trainable parts in the two-stage framework, then trace visual features into language inputs.

A question to keep asking

Frozen components retain their limitations, and visual errors can accompany fluent but inaccurate answers.

2023 · NeurIPS 2023

Visual Instruction Tuning

Visual Instruction Tuning · Liu et al., Visual Instruction Tuning, Figure 1
Figure excerpt from the paper · Liu et al., Visual Instruction Tuning, Figure 1 · Original paper

The problem

How can a model answer natural-language instructions about an image?

The key idea

Connect an image encoder to a language model with a projection and train on visual instructions. The original uses captions and box information with text-only GPT-4 to create part of its dialogue data.

Why this paper

Data organization can change interaction as much as the connection itself.

Where to start

Start with instruction-data construction, then the connector and staged training.

A question to keep asking

Answers can invent visual details.

2023 · ICML 2023

SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models · Xiao et al. (2023), Fig. 2
Figure excerpt from the paper · Xiao et al. (2023), Fig. 2 · Original paper

The problem

Why do a few very large activations make INT8 quantization difficult?

The key idea

Redistribute scale between weights and activations with an equivalent transformation, then quantize.

Why this paper

It connects outliers, numerical error, and hardware execution in one example.

Where to start

Start with activation outliers and scale migration, then accuracy and speed comparisons.

A question to keep asking

The pre-quantization transformation is equivalent; quantization still introduces error. Representative calibration data and suitable INT8 kernels matter.

2023 · SOSP 2023

Efficient Memory Management for Large Language Model Serving with PagedAttention

Efficient Memory Management for Large Language Model Serving with PagedAttention · Kwon et al. (2023), Fig. 6
Figure excerpt from the paper · Kwon et al. (2023), Fig. 6 · Original paper

The problem

How can changing request lengths avoid wasting KV-cache memory?

The key idea

Split the cache into blocks, map logical sequences to noncontiguous storage, and support cache sharing.

Why this paper

It connects operating-system paging to the capacity of a model-serving system.

Where to start

Start with fragmentation and block mapping, then throughput under matched latency conditions.

A question to keep asking

It optimizes cache management. Concurrency gains must be assessed with latency targets and request lengths.

2024 · ICLR 2024

SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

SWE-bench: Can Language Models Resolve Real-World GitHub Issues? · Jimenez et al., 2024 · Figure 1
Figure excerpt from the paper · Jimenez et al., 2024 · Figure 1 · Original paper

The problem

Can a model turn a repository and issue into a working fix?

The key idea

Collect real issues and fixes, provide the repository and problem, and evaluate proposed patches with tests.

Why this paper

It joins reading, localization, editing, and verification in one task.

Where to start

Start with task construction and evaluation, then a failure case.

A question to keep asking

The original data comes from Python repositories. Passing tests does not cover all software quality or establish transfer to other languages.

2024 · Nature

Solving olympiad geometry without human demonstrations

Solving olympiad geometry without human demonstrations · Google DeepMind · AlphaGeometry 官方方法示意图
Figure excerpt from the paper · Google DeepMind · AlphaGeometry 官方方法示意图 · Original paper

The problem

How can auxiliary constructions and symbolic reasoning form a proof process?

The key idea

A symbolic engine first derives consequences from the given conditions. When needed, a language model proposes auxiliary constructions, and proof search continues.

Where to start

Trace the added constructions and the deductions they enable in a simple geometry problem, then examine how the training data are synthesized.

2024 · arXiv

π0: A Vision-Language-Action Flow Model for General Robot Control

π0: A Vision-Language-Action Flow Model for General Robot Control · Physical Intelligence · Fig. 1
Figure excerpt from the paper · Physical Intelligence · Fig. 1 · Original paper

The problem

How can a vision-language model produce continuous actions across robot tasks and platforms?

The key idea

An action expert uses flow matching to generate continuous action sequences alongside a pretrained VLM. The policy is pretrained on demonstrations from multiple robots, then fine-tuned on task data.

Why this paper

It brings VLM pretraining, action generation, and cross-robot data together, connecting RT-2 and Diffusion Policy.

Where to start

Start with the first-page overview, then the action-expert architecture. Compare direct use after pretraining with task-specific fine-tuning.

2025 · arXiv

Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents

Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents · Figure 1
Figure excerpt from the paper · Figure 1 · Original paper

The problem

Can a coding agent improve its abilities by modifying its own code?

The key idea

Maintain multiple agent versions, generate self-modifications, and select candidates using coding tasks, creating an expanding version tree.

Where to start

Start with one specific code change, then examine version selection, evaluation tasks, and the computation budget.

2025 · arXiv

The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search · Yamada et al. · Figure 1
Figure excerpt from the paper · Yamada et al. · Figure 1 · Original paper

The problem

How can experimental results guide the next round of research?

The key idea

Use agentic tree search to organize experiments, connecting plans, execution, analysis, and paper writing.

Where to start

Start with one complete experimental branch and check whether the final claims correspond to actual results.

2025 · arXiv

AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model

AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model · Figure 1
Figure excerpt from the paper · Figure 1 · Original paper

The problem

How can family information and verification feedback support protein design?

The key idea

Model proteins with Bayesian flow networks, use MSA context, and explore design workflows with larger inference-time search budgets.

Where to start

Trace generation conditions, verifiers, and experimental measurements separately, comparing their roles in the iterative process.

2025 · arXiv

CMRINet: Joint Groupwise Registration and Segmentation for Cardiac Function Quantification from Cine-MRI

CMRINet: Joint Groupwise Registration and Segmentation for Cardiac Function Quantification from Cine-MRI · Figure 1
Figure excerpt from the paper · Figure 1 · Original paper

The problem

How can dynamic cardiac MRI provide both structure and motion information?

The key idea

Jointly handle groupwise registration and segmentation, connecting temporal image analysis to cardiac function quantification.

Where to start

Start with the temporal dimension of cine MRI, then examine how the two tasks work together and how functional measures are calculated.

2025 · arXiv

π0.5: a Vision-Language-Action Model with Open-World Generalization

π0.5: a Vision-Language-Action Model with Open-World Generalization · Physical Intelligence · Fig. 1
Figure excerpt from the paper · Physical Intelligence · Fig. 1 · Original paper

The problem

Can a robot carry out long tasks in homes it did not visit during training?

The key idea

Building on π0, the model co-trains on robot actions, subtask prediction, object detection, and web vision-language data, connecting high-level subtasks to continuous actions.

Why this paper

It studies multi-stage manipulation in new homes and how the mix of training data affects generalization.

Where to start

Start with the first-page data and deployment overview, then inspect the new-home evaluation and data-source ablations.

2025 · Nature

Mastering diverse control tasks through world models

Mastering diverse control tasks through world models · Hafner et al. (2025), Fig. 1 (CC BY 4.0)
Figure excerpt from the paper · Hafner et al. (2025), Fig. 1 (CC BY 4.0) · Original paper

The problem

How can one algorithm handle different task signal scales?

The key idea

Learn a world model from experience, train actor and critic on imagined trajectories, and stabilize losses and signal scales.

Why this paper

It connects predictive learning with behavior learning and cross-task evaluation.

Where to start

Start with the training diagram, learning algorithm, and ablations.

A question to keep asking

A shared configuration does not mean one pretrained agent solves every task without training.

2026 · ICML

FLAG: Foundation model representation with Latent diffusion Alignment via Graph for spatial gene expression prediction

FLAG: Foundation model representation with Latent diffusion Alignment via Graph for spatial gene expression prediction · Figure 1
Figure excerpt from the paper · Figure 1 · Original paper

The problem

How can tissue images predict spatial gene expression while preserving structural relationships?

The key idea

Combine spatial graph encoding with gene foundation model representations, and use a diffusion process to model expression distributions.

Where to start

Start with input images, spatial adjacency, and expression outputs, then compare errors at individual locations with structural evaluations.

Here are papers selected by question, each with an angle to begin from. For the reading process, see finding and reading papers.

A first encounter with a paper

ResNet
What training problem did the authors face, what structure changed, and which comparisons supported it? Pair with Mu Li's readings, then return to vision.

Attention Is All You Need
Trace inputs and outputs in the architecture diagram, then ask which computation attention handles. The illustrated explanation helps a first pass. Continue to foundations and architectures.

ReAct
Read one complete trace: when does the model act, and what changes after feedback? Then inspect comparisons between reasoning, acting, and their combination. Continue to agents.

Follow your interests

  • BERT: how objectives shape representations and how those representations serve tasks. See LLMs.
  • LoRA: what reducing trainable parameters saves and how effectiveness is evaluated. See LLMs and systems.
  • CLIP: image–text matching as training and how the trained model is used. See multimodal models.
  • DDPM: connections between noisy data, objectives, and generation. Pair with generation.
  • World Models: combining compressed observations, dynamics prediction, and action selection. See world models.
  • DreamerV3: learned models for behavior learning and the tasks covered by evaluation. See RL.
  • ACT / ALOHA: connecting demonstrations and action prediction to real robot tasks. See embodied AI.
  • Diffusion Policy: generative modeling of actions and the reason for that representation. See generation and embodied AI.

Longer lists once you have a question

Agents · Multimodal models · World Models · Embodied guide · AI for Science

For new work, see Hugging Face Daily Papers.

Read with someone experienced

Compare your own reading notes with Mu Li's close readings. For more projects and paper lists, use the direction catalog.