A reading track that runs alongside the levels. Every entry was checked on 2026-10-03 against its arXiv record or its Crossref DOI record: title, first author and year below are as those records give them. Nothing here is cited from memory.
[!recommendation] How to read a paper in this field
Read the evaluation before the method. For each paper, write down three things before you summarize what it claims: the benchmark and its version, how success is defined and aggregated (per episode or per task, how many seeds and evaluation episodes), and which generalization axis the test actually varies. Then ask what the evaluation cannot support. A method tested only on its training distribution says nothing about generalization; a method tested on one seed says little about anything. Research mode has the tools for this check.
Each section lists papers in suggested reading order with the course level where they fit, and one question to answer while reading.
Primary sources: MuJoCo itself
The documentation is the reference for everything this course states about MuJoCo’s behaviour. Read the computation chapter after Level 1 and again after Level 20.
- MuJoCo documentation: Computation, Modeling, XML reference, API reference.
- Changelog. Read it before upgrading. Several lessons cite it for when a feature appeared.
- Source code and MuJoCo Menagerie, the curated collection of robot models.
- E. Todorov, T. Erez, Y. Tassa. MuJoCo: A physics engine for model-based control. IROS 2012, pp. 5026-5033. doi:10.1109/IROS.2012.6386109. Level 0. The original design goals. Question: which of the 2012 design decisions does the current computation chapter still describe, and which changed?
- E. Todorov. Convex and analytically-invertible dynamics with contacts and constraints: Theory and implementation in MuJoCo. ICRA 2014, pp. 6054-6061. doi:10.1109/ICRA.2014.6907751. Levels 9 and 20. The convex soft-contact formulation the solver still uses. Question: what does convexity buy, and what physical behaviour does it give up compared with a complementarity formulation?
Mathematics and robotics foundations
- K. M. Lynch, F. C. Park. Modern Robotics: Mechanics, Planning, and Control. Cambridge University Press, 2017. doi:10.1017/9781316661239. Levels 4 to 8. Screw theory, kinematics, dynamics and control in one notation. Question: map its spatial and body twists onto MuJoCo’s
cvel and free-joint qvel.
- R. Featherstone. Rigid Body Dynamics Algorithms. Springer, 2008. doi:10.1007/978-1-4899-7560-7. Levels 7 and 20. The recursive Newton-Euler and composite rigid body algorithms behind
mj_rne and mj_crb. Question: which quantities does MuJoCo cache in mjData that the book’s algorithms recompute?
- J. Solà. Quaternion kinematics for the error-state Kalman filter. arXiv:1711.02508, 2017. arXiv. Level 4. Read the section on conventions: quaternion libraries disagree on component order and on local versus global perturbations, and MuJoCo picks one of each.
- J. Solà, J. Deray, D. Atchuthan. A micro Lie theory for state estimation in robotics. arXiv:1812.01537, 2018. arXiv. Levels 4 and 6. The cleanest short treatment of exponential maps and Jacobians on SO(3) and SE(3). Question: why does
mj_differentiatePos return a vector of size nv rather than nq?
- O. Khatib. A unified approach for motion and force control of robot manipulators: The operational space formulation. IEEE Journal on Robotics and Automation 3(1), 1987, pp. 43-53. doi:10.1109/JRA.1987.1087068. Level 8. The source of operational-space control. Question: what does the controller assume about the model, and what happens when that model is wrong (Project 6 measures it)?
- D. E. Stewart, J. C. Trinkle. An implicit time-stepping scheme for rigid body dynamics with inelastic collisions and Coulomb friction. International Journal for Numerical Methods in Engineering 39(15), 1996, pp. 2673-2691. doi:10.1002/(SICI)1097-0207(19960815)39:15<2673::AID-NME972>3.0.CO;2-I. Level 9. The velocity-level time-stepping idea that most robotics simulators share.
- M. Anitescu, F. A. Potra. Formulating dynamic multi-rigid-body contact problems with friction as solvable linear complementarity problems. Nonlinear Dynamics 14(3), 1997, pp. 231-247. doi:10.1023/A:1008292328909. Level 9. Why polyhedral friction cones make the problem solvable, and what they cost (compare the pyramidal cone measurements in Lesson 0.3).
- E. Todorov. A convex, smooth and invertible contact model for trajectory optimization. ICRA 2011. doi:10.1109/ICRA.2011.5979814. Levels 9 and 21. The precursor of MuJoCo’s contact model, written for optimization rather than simulation. Question: why does trajectory optimization want contact forces that are smooth in the state?
- J. Horak, J. C. Trinkle. On the similarities and differences among contact models in robot simulation. IEEE Robotics and Automation Letters 4(2), 2019, pp. 493-499. doi:10.1109/LRA.2019.2891085. Level 9. A side-by-side reading of the formulations used by MuJoCo and other engines.
- T. Erez, Y. Tassa, E. Todorov. Simulation tools for model-based robotics: Comparison of Bullet, Havok, MuJoCo, ODE and PhysX. ICRA 2015, pp. 4397-4404. doi:10.1109/ICRA.2015.7139807. Level 0. Read for the method of comparison (speed against accuracy as the timestep varies), not for the numbers: every engine compared has changed since 2015.
- T. A. Howell et al. Dojo: A Differentiable Physics Engine for Robotics. arXiv:2203.00806, 2022. arXiv. Level 20. A hard-contact alternative designed for useful gradients. Question: what does it give up to get them?
Reinforcement learning and its evaluation
- J. Schulman et al. High-Dimensional Continuous Control Using Generalized Advantage Estimation. arXiv:1506.02438, 2015. arXiv. Level 13.
- J. Schulman et al. Proximal Policy Optimization Algorithms. arXiv:1707.06347, 2017. arXiv. Level 13. Question: which of the details that make PPO work in practice are in the paper, and which are only in the code?
- L. Engstrom et al. Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO. arXiv:2005.12729, 2020. arXiv. Level 13. The answer to the previous question, measured.
- T. Haarnoja et al. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. arXiv:1801.01290, 2018. arXiv. Level 13.
- S. Fujimoto, H. van Hoof, D. Meger. Addressing Function Approximation Error in Actor-Critic Methods. arXiv:1802.09477, 2018. arXiv. Level 13. TD3.
- P. Henderson et al. Deep Reinforcement Learning that Matters. arXiv:1709.06560, 2017. arXiv. Levels 13 and 21. Seeds, implementations and hyperparameters change conclusions; the experiments use MuJoCo locomotion tasks.
- M. Andrychowicz et al. What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study. arXiv:2006.05990, 2020. arXiv. Level 13. What it cannot support: its tasks are locomotion, so its recommendations are hypotheses, not results, for manipulation.
- R. Agarwal et al. Deep Reinforcement Learning at the Edge of the Statistical Precipice. arXiv:2108.13264, 2021. arXiv. Level 21. Interquartile means and bootstrap intervals over runs; Research mode uses both.
- Y. Tassa et al. DeepMind Control Suite. arXiv:1801.00690, 2018. arXiv. Level 12. A benchmark built on MuJoCo; read how it fixes the reward scale and episode length across tasks.
- M. Towers et al. Gymnasium: A Standard Interface for Reinforcement Learning Environments. arXiv:2407.17032, 2024. arXiv. Level 12. Why
terminated and truncated are separate.
Imitation learning
- S. Ross, G. J. Gordon, J. A. Bagnell. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. arXiv:1011.0686, 2010 (AISTATS 2011). arXiv. Level 14. DAgger, and the compounding-error argument that motivates it.
- A. Mandlekar et al. What Matters in Learning from Offline Human Demonstrations for Robot Manipulation. arXiv:2108.03298, 2021. arXiv. Level 14. Demonstration quality, observation choice and evaluation for behaviour cloning.
- C. Chi et al. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. arXiv:2303.04137, 2023. arXiv. Level 14. Multimodal action distributions. Question: on which tasks does a unimodal Gaussian policy fail, and why?
- T. Z. Zhao et al. Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv:2304.13705, 2023. arXiv. Level 14. Action chunking (ACT).
Vision, language and vision-language-action models
- M. Laskin et al. Reinforcement Learning with Augmented Data. arXiv:2004.14990, 2020. arXiv. Level 15.
- I. Kostrikov, D. Yarats, R. Fergus. Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels. arXiv:2004.13649, 2020. arXiv. Level 15. Both papers evaluate on the DeepMind Control Suite; ask what their augmentations would do to a task where object position is the signal.
- A. Brohan et al. RT-1: Robotics Transformer for Real-World Control at Scale. arXiv:2212.06817, 2022. arXiv. Level 16.
- A. Brohan et al. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. arXiv:2307.15818, 2023. arXiv. Level 16. Where the term vision-language-action model comes from.
- Open X-Embodiment Collaboration. Open X-Embodiment: Robotic Learning Datasets and RT-X Models. arXiv:2310.08864, 2023. arXiv. Level 16.
- Octo Model Team. Octo: An Open-Source Generalist Robot Policy. arXiv:2405.12213, 2024. arXiv. Level 16.
- M. J. Kim et al. OpenVLA: An Open-Source Vision-Language-Action Model. arXiv:2406.09246, 2024. arXiv. Level 16.
- K. Black et al. π0: A Vision-Language-Action Flow Model for General Robot Control. arXiv:2410.24164, 2024. arXiv. Level 16. For each of these four, find which generalization axis (object, position, instruction, embodiment) each reported number tests.
World models and planning
- Y. Tassa, T. Erez, E. Todorov. Synthesis and stabilization of complex behaviors through online trajectory optimization. IROS 2012, pp. 4906-4913. doi:10.1109/IROS.2012.6386025. Level 17. Model predictive control with MuJoCo as the model.
- G. Williams et al. Information theoretic MPC for model-based reinforcement learning. ICRA 2017, pp. 1714-1721. doi:10.1109/ICRA.2017.7989202. Level 17. MPPI.
- T. Howell et al. Predictive Sampling: Real-time Behaviour Synthesis with MuJoCo. arXiv:2212.00541, 2022. arXiv. Level 17. MuJoCo MPC, and the case for a very simple sampling planner.
- D. Ha, J. Schmidhuber. World Models. arXiv:1803.10122, 2018. arXiv. Level 17.
- D. Hafner et al. Learning Latent Dynamics for Planning from Pixels. arXiv:1811.04551, 2018. arXiv. Level 17. PlaNet.
- D. Hafner et al. Mastering Diverse Domains through World Models. arXiv:2301.04104, 2023. arXiv. Level 17. DreamerV3.
- N. Hansen, X. Wang, H. Su. Temporal Difference Learning for Model Predictive Control. arXiv:2203.04955, 2022. arXiv. Level 17. TD-MPC.
- N. Hansen, H. Su, X. Wang. TD-MPC2: Scalable, Robust World Models for Continuous Control. arXiv:2310.16828, 2023. arXiv. Level 17.
- M. Assran et al. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning. arXiv:2506.09985, 2025. arXiv. Level 17. For each world-model paper, separate three claims the course keeps apart: prediction accuracy, controllability, and the success of the policy or planner that uses the model.
Domain randomization, system identification and sim-to-real
- J. Tobin et al. Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World. arXiv:1703.06907, 2017. arXiv. Level 18. Visual randomization.
- X. B. Peng et al. Sim-to-Real Transfer of Robotic Control with Dynamics Randomization. arXiv:1710.06537, 2017. arXiv. Level 18. Dynamics randomization.
- J. Tan et al. Sim-to-Real: Learning Agile Locomotion For Quadruped Robots. arXiv:1804.10332, 2018. arXiv. Levels 18 and 19. System identification and randomization used together, with an actuator model.
- J. Hwangbo et al. Learning agile and dynamic motor skills for legged robots. Science Robotics 4(26), 2019, eaau5872. arXiv:1901.08652. Level 19. A learned actuator model closes the gap that rigid-body identification left.
- OpenAI et al. Solving Rubik’s Cube with a Robot Hand. arXiv:1910.07113, 2019. arXiv. Level 18. Automatic domain randomization at large compute cost.
- A. Kumar et al. RMA: Rapid Motor Adaptation for Legged Robots. arXiv:2107.04034, 2021. arXiv. Level 19. Adapting to the dynamics online instead of being robust to all of them.
- W. Zhao, J. P. Queralta, T. Westerlund. Sim-to-Real Transfer in Deep Reinforcement Learning for Robotics: a Survey. IEEE SSCI 2020, pp. 737-744. arXiv:2009.13303. Level 19.
- F. Muratore et al. Robot Learning from Randomized Simulations: A Review. arXiv:2111.00956, 2021. arXiv. Level 18.
Benchmarks and batched simulation
Read these for their protocols (Level 21): task set, splits, initial states, success definition, aggregation, and simulator version.
- T. Yu et al. Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning. arXiv:1910.10897, 2019. arXiv. Built on MuJoCo. Its task definitions changed between releases: always report the version.
- Y. Zhu et al. robosuite: A Modular Simulation Framework and Benchmark for Robot Learning. arXiv:2009.12293, 2020. arXiv. Built on MuJoCo.
- B. Liu et al. LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning. arXiv:2306.03310, 2023. arXiv. Built on robosuite.
- X. Zhou et al. LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization. arXiv:2510.03827, 2025. arXiv, and S. Fei et al. LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models. arXiv:2510.13626, 2025. arXiv. Both perturb LIBERO’s evaluation conditions; read them together with LIBERO to see how much of a reported success rate survives a change of initial state or instruction.
- O. Mees et al. CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks. arXiv:2112.03227, 2021. arXiv. Built on PyBullet. Its splits (for example ABC to D versus ABCD to D) are not interchangeable.
- S. James et al. RLBench: The Robot Learning Benchmark & Learning Environment. arXiv:1909.12271, 2019. arXiv. Built on CoppeliaSim.
- X. Li et al. Evaluating Real-World Robot Manipulation Policies in Simulation. arXiv:2405.05941, 2024. arXiv. SimplerEnv: how well simulated success predicts real success, and how to measure that.
- S. Tao et al. ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI. arXiv:2410.00425, 2024. arXiv. Built on SAPIEN.
- C. D. Freeman et al. Brax: A Differentiable Physics Engine for Large Scale Rigid Body Simulation. arXiv:2106.13281, 2021. arXiv. Accelerator-batched simulation in JAX, the setting MJX later brought to MuJoCo’s own pipeline.
- K. Zakka et al. MuJoCo Playground. arXiv:2502.08844, 2025. arXiv. Environments and training on MJX. Read with the Performance lab: which of its throughput numbers depend on the accelerator, and which on the batch size?