Blogs often restore the process papers omit: why a method was tried, what disappointed, and how the next step was chosen.
For experience, start with Zou's Cognitive Compound Interest or Chai's FAQ. For experiments, read LoopDiT and AutomationBench. For mechanisms, consult Su and Purshow.
For more authors and topics, explore OpenEnvision's BlogrXiv, which organizes research blogs, lab essays, and technical notes by field.
Yuwei Niu / Purshow
Why remove the encoder? An infrastructure perspective
The article explains encoder-free design through differences in workload, parallel configuration, and scheduling between vision encoders and language models. It also notes that removing the encoder does not eliminate workload variation from variable-length visual tokens.
Purshow Notes: consult by question
Repository · WAM · Encoder-Free
Topics include attention, MoE, on-policy distillation, world-action models, residual connections, sparse vocabulary embeddings, and YOCO. Several have Chinese and English versions.
Two short perspectives
- When Vision Is Pushed to the Roadside
- The (Possible) Future of Multimodal Understanding: From Describing the World to Entering the World
Wenhao Chai
FAQ for Juniors
Updated February 25, 2026. Covers PhD choices, entering from outside CS, lacking a suitable local lab, contacting researchers, selecting directions, and the value of benchmark work.
LoopDiT: Loop Transformers for Diffusion Models
Chinese and English, September 10, 2026. Compares weight-sharing looped Transformers and ordinary models under fixed parameter and fixed compute conditions, including more loops versus more denoising steps.
What Happened to AutomationBench?
Chinese and English, September 19, 2026. Examines scoring definitions and string-matching rules in agent-workflow evaluation and how they affect completion rates and model rankings.
Predictable Swarm Scaling
Chinese and English, September 27, 2026. Uses task-dependency graphs and simulations to compare single agents and standard or recursive swarms. Distinguishes coverage from best results while discussing division of work and cost limits.
Jiaxuan Zou
Cognitive Compound Interest: reflections on two undergraduate years
Chinese, July 4, 2026. Connects seminars, public technical writing, collaboration, and internship opportunities through undergraduate experience, including changing plans as new information arrived.
Pretraining and scaling as methodology and scientific perspective
Chinese, September 7, 2026. Uses pretraining and embodied-AI examples to connect general learning frameworks, training/inference compute scaling, efficiency, stability, and extrapolation in research.
How to build a scientific Scaling Ladder
Chinese, September 20, 2026. Organizes a workflow around small experiments, consistent definitions, configuration search, data and model scale, fitting, and validation of extrapolations. Discusses dense, MoE, and data-limited settings.
Xiaohan Ding
Writing AI Conference Papers: A Handbook for Beginners
By Zhewei Huang and Xiaohan Ding, 2024. Covers extracting contributions, overall structure, introductions and related work, logic, defensibility, reader effort, and information density. Includes common negative reviews and a pre-submission check.
Rebuttals and doctoral experience
- Article on rebuttals
- An Anxious Beginning, a Calmer Later Stage: the First Two PhD Years at Tsinghua, reprinted by BAAI
Xiaoguang Han
University profile. The materials below include academic discussion, mentoring interviews, and research experience.
Highlights from the first GAMES academic salon
A 2022 discussion of neural implicit representations and 3D research. Han's contribution concerns the relationship between understanding and reconstruction. The full discussion shows researchers questioning one another about concrete problems.
Good Mentors Online, episode one
Official GAMES-Webinar video on Bilibili
Guests Sida Peng and Guanbin Li, hosted by Xiaoguang Han; published January 16, 2025.
Talking Shape and AI: academia and industry, two sides of research
Official GAMES-Webinar video on Bilibili
Guests Xiaojuan Qi and Yingqing He, hosted by Xiaoguang Han; page dated August 2, 2026.
For What Makes Good Research, see the original post and research conversations. More talks are in the GAMES archive.
Jianlin Su / BoJone: Scientific Spaces
Scientific Spaces · Article archive
Upgrading the Transformer, part 1: tracing sinusoidal positional encodings
Original, 2021.
Why do we need positional representations, and how can desired properties guide the analysis of sinusoidal encodings? Useful for LLM mechanisms and applied mathematics; assumes basic linear algebra, attention, and Taylor expansions. Ask where the sine and cosine terms come from.
Upgrading the Transformer, part 2: rotary positional embeddings
Original, 2021.
Introduces RoPE and RoFormer, constructing a positional transformation from the goal of making inner products reflect relative position. It is both a mechanism explanation and a firsthand account of design emerging from a question and constraints. Read after part one.
A discussion of diffusion models, part 1: DDPM as demolition and reconstruction
Original, 2022.
Starts with a gradual destruction-and-rebuilding analogy, then connects noising and generation. Use the analogy for intuition before following the probability-based derivation.
More on Muon: why did we choose to try it?
Original, 2025.
Explains the team's choice of Muon through optimizer practice and discusses spectral norms, learning rates, weight decay, and related theory and experiments. Start with the question, motivation, and experimental adjustments before the equations.
The thousandth article
Original, 2020.
A reflection on long-term blogging and the site as a place for personal notes. Useful when beginning to record and share your own learning.