When a demo makes you curious, ask what goes in, what should come out, how success is judged, and what is still difficult. These pages start with concrete tasks.
Understanding and generating
- Large language models: how models learn from text, and how data, training, and computation shape capability.
- Computer vision: recognizing objects, locating them, and understanding scenes and 3D structure.
- Multimodal models: connecting images, language, and other information to answer questions.
- Generative models: producing images from noise or simple distributions and controlling the process.
Acting
- Agents: giving models tools and environments, then observing how they complete tasks.
- Reinforcement learning: choosing actions from feedback when actions change future situations.
- World models: predicting what happens after an action and using that prediction.
- Embodied AI: moving from observation to action and doing things in real environments.
These pages connect. Robots can use vision, generative models, and RL; agents can learn with RL too. Follow your question across the boundaries.
Putting models to use
- Efficiency and systems: understanding slow programs, limited GPU memory, and quality–latency–cost trade-offs.
- Interdisciplinary AI: understanding molecular, medical, or relational-data problems beyond the model.
Each page offers a starting point, foundations to revisit, and further reading. For individual papers, see paper entry points; for what researchers are thinking about, see Lookout.
Courses, papers, and people
The direction catalog includes Lumina, awesome lists, courses, papers, and projects. Continue to the lab directory to meet the people working on these questions.