MuJoCo gives Python two ways to see a simulation, built for two different jobs.
The interactive viewer (mujoco.viewer) is for you: a window with mouse navigation, perturbation by dragging, toggles for contacts, forces, frames and inertias, and a profiler. You use it to debug a model and to watch a controller. launch_passive hands the stepping loop to your code; launch runs its own loop.
The renderer (mujoco.Renderer) is for your program: it renders a camera’s view into a NumPy array, offscreen, as RGB, depth or segmentation. You use it to generate datasets, to feed vision policies, to make videos for papers.
Both draw the same thing: MuJoCo turns mjModel and mjData into an abstract scene (mjvScene, a list of geoms, lights and cameras) with mjv_updateScene, and an OpenGL backend draws the scene. The renderer needs an OpenGL context even though it never opens a window, which is why headless machines need MUJOCO_GL (Lesson 0.2).
The passive viewer is the one you will use most, because your code stays in charge of stepping and control:
INPUT: `cartpole.xml`; a desktop session with a display (on macOS, run with `mjpython`)
PROCESS: balance the pole with state feedback, step at real-time speed, toggle contact drawing, sync the viewer every step
OUTPUT: an interactive window, and the final pole angle
```python file=examples/l3_2_viewer.py “"”Lesson 3.2: the passive viewer with your own control loop.
INPUT cartpole.xml; a desktop session with a display (on macOS run with mjpython)
PROCESS open the passive viewer, balance the pole with a linear state-feedback law
at real-time speed,
toggle contact-point drawing every 2 s, close after 20 s or when the window closes
OUTPUT an interactive window; the final pole angle printed at exit
Run: python examples/l3_2_viewer.py (mjpython on macOS) “””
import time
import mujoco import mujoco.viewer
from mjcourse import model_path
def main() -> None: model = mujoco.MjModel.from_xml_path(str(model_path(“cartpole”))) data = mujoco.MjData(model) mujoco.mj_resetDataKeyframe(model, data, 0) # pole tilted 0.15 rad with mujoco.viewer.launch_passive(model, data) as viewer: start = time.time() while viewer.is_running() and time.time() - start < 20: step_start = time.time() # Balance by pushing the cart. The gains are an LQR design on MuJoCo’s own # linearization (mjd_transitionFD), derived in Level 21.4. Feedback on the cart’s # position and velocity is needed too: feedback on the pole alone keeps it up for # a while, then the cart runs into the end of its rail. x, angle = data.qpos xdot, rate = data.qvel data.ctrl[0] = 9.81 * x + 70.0 * angle + 11.3 * xdot + 12.7 * rate mujoco.mj_step(model, data) with viewer.lock(): # edit viewer state safely viewer.opt.flags[mujoco.mjtVisFlag.mjVIS_CONTACTPOINT] = int(data.time % 4 < 2) viewer.sync() # push the new state to the window remaining = model.opt.timestep - (time.time() - step_start) if remaining > 0: time.sleep(remaining) print(f”final pole angle {data.joint(‘hinge’).qpos[0]:+.4f} rad at t = {data.time:.2f} s”)
if name == “main”: main()
> [!unverified] What this course could and could not run
> The build machine for this course has no display, so the window itself was not opened while writing the lesson; the test suite only byte-compiles this script. The control law inside it *was* run headlessly: from the 0.15 rad start it keeps the pole up for 20 s with a peak force of 10.5 N. The viewer calls (`launch_passive`, `lock`, `sync`, `is_running`, `opt.flags`) follow the 3.14.0 Python documentation's own passive-viewer example. If the window misbehaves on your machine, that is the part to suspect first.
> [!warning] An earlier draft of this script was wrong
> The first version balanced the pole with feedback on the pole angle and rate only. Run headlessly it held the pole for a while, then the cart drifted into the end of its rail and the pole fell. Stabilizing a cart-pole needs feedback on all four state variables; the gains above come from an LQR design and the positive gain on the cart's position is not a typo (it is the non-minimum-phase character of the system: to move the cart right, the pole must first be tipped right). Testing control code without a viewer caught a bug the viewer would have made look like a tuning problem.
The parts of the passive viewer you will use:
| Call | Does |
|---|---|
| `mujoco.viewer.launch_passive(model, data)` | opens the window and returns a handle (use it as a context manager) |
| `viewer.sync()` | copies the current `mjData` to the window and applies mouse perturbations to `data`; call it every step or every frame |
| `with viewer.lock():` | holds the viewer's lock while you change its options or camera |
| `viewer.is_running()` | false once the user closes the window |
| `viewer.opt`, `viewer.cam`, `viewer.user_scn` | visualization flags, the camera, and a scene you can add your own geoms to (markers, targets, trajectories) |
## Rendering to arrays
The renderer takes a model, an image size and, for each frame, a state and a camera:
```io
INPUT: `pick_place.xml` at its home pose; an OpenGL backend
PROCESS: render five cameras (four model cameras and one free camera) in RGB, depth and segmentation
OUTPUT: `runs/l3_2/cameras.png` and statistics per image
```python file=examples/l3_2_render.py “"”Lesson 3.2: offscreen rendering with mujoco.Renderer: RGB, depth, segmentation.
INPUT pick_place.xml at its home keyframe; needs an OpenGL backend (set MUJOCO_GL=egl or osmesa on a machine without a display) PROCESS render the front, side, top and wrist cameras plus one free camera, in RGB, metric depth and segmentation; summarize each image OUTPUT runs/l3_2/cameras.png (a figure of all images) and printed statistics
Run: MUJOCO_GL=egl python examples/l3_2_render.py “””
import sys from pathlib import Path
import matplotlib import mujoco import numpy as np
from mjcourse import model_path
matplotlib.use(“Agg”) import matplotlib.pyplot as plt # noqa: E402
OUT = Path(file).resolve().parents[1] / “runs” / “l3_2” H, W = 240, 320
def free_camera(model: mujoco.MjModel) -> mujoco.MjvCamera: cam = mujoco.MjvCamera() mujoco.mjv_defaultFreeCamera(model, cam) cam.lookat[:] = [0.45, 0.0, 0.1] # point the camera orbits around (m) cam.distance, cam.azimuth, cam.elevation = 1.3, 150.0, -30.0 # m, degrees, degrees return cam
def main() -> int: model = mujoco.MjModel.from_xml_path(str(model_path(“pick_place”))) data = mujoco.MjData(model) mujoco.mj_resetDataKeyframe(model, data, 1) mujoco.mj_forward(model, data) try: renderer = mujoco.Renderer(model, height=H, width=W) except Exception as err: # noqa: BLE001 print(f”rendering unavailable: {err}\nset MUJOCO_GL=egl or osmesa (Lesson 0.2)”) return 2
cameras = {"front": "front", "side": "side", "top": "top", "wrist": "gripper/wrist", "free": free_camera(model)}
images = {}
for label, cam in cameras.items():
renderer.disable_depth_rendering()
renderer.disable_segmentation_rendering()
renderer.update_scene(data, camera=cam)
rgb = renderer.render().copy()
renderer.enable_depth_rendering()
renderer.update_scene(data, camera=cam)
depth = renderer.render().copy()
renderer.enable_segmentation_rendering()
renderer.update_scene(data, camera=cam)
seg = renderer.render().copy() # (H, W, 2): object id, object type; background -1
images[label] = (rgb, depth, seg)
geom_ids = np.unique(seg[..., 0][seg[..., 1] == mujoco.mjtObj.mjOBJ_GEOM])
cube_pixels = {n: int(np.sum((seg[..., 1] == mujoco.mjtObj.mjOBJ_GEOM) & (seg[..., 0] == model.geom(n).id)))
for n in ("red_cube", "green_cube", "blue_cube")}
print(f"{label:>6}: rgb {rgb.shape} {rgb.dtype}; depth {depth.dtype} {depth.min():.3f} to {depth.max():.3f} m; "
f"{len(geom_ids)} geoms visible; cube pixels {cube_pixels}")
renderer.close()
OUT.mkdir(parents=True, exist_ok=True)
fig, axes = plt.subplots(3, len(images), figsize=(3.2 * len(images), 7.2))
for col, (label, (rgb, depth, seg)) in enumerate(images.items()):
axes[0, col].imshow(rgb)
axes[0, col].set_title(label)
d = axes[1, col].imshow(depth, cmap="viridis")
fig.colorbar(d, ax=axes[1, col], fraction=0.046, label="depth (m)")
is_geom = seg[..., 1] == mujoco.mjtObj.mjOBJ_GEOM
geom_id = np.ma.masked_where(~is_geom, seg[..., 0]) # background and sites masked out
cmap = matplotlib.colormaps["tab20"].with_extremes(bad="black")
axes[2, col].imshow(geom_id, cmap=cmap, vmin=0, vmax=model.ngeom, interpolation="nearest")
for row in range(3):
axes[row, col].set_xticks([])
axes[row, col].set_yticks([])
axes[0, 0].set_ylabel("RGB")
axes[1, 0].set_ylabel("depth")
axes[2, 0].set_ylabel("segmentation (geom id)")
fig.tight_layout()
fig.savefig(OUT / "cameras.png", dpi=110)
print(f"saved {OUT / 'cameras.png'}")
return 0
if name == “main”: sys.exit(main())
Output, with `MUJOCO_GL=osmesa` on the course's build machine:
```text
front: rgb (240, 320, 3) uint8; depth float32 0.729 to 2.364 m; 18 geoms visible; cube pixels {'red_cube': 148, 'green_cube': 165, 'blue_cube': 150}
side: rgb (240, 320, 3) uint8; depth float32 0.711 to 80.008 m; 21 geoms visible; cube pixels {'red_cube': 108, 'green_cube': 147, 'blue_cube': 189}
top: rgb (240, 320, 3) uint8; depth float32 0.642 to 1.400 m; 14 geoms visible; cube pixels {'red_cube': 81, 'green_cube': 2, 'blue_cube': 80}
wrist: rgb (240, 320, 3) uint8; depth float32 0.062 to 0.384 m; 13 geoms visible; cube pixels {'red_cube': 429, 'green_cube': 376, 'blue_cube': 480}
free: rgb (240, 320, 3) uint8; depth float32 0.875 to 80.008 m; 23 geoms visible; cube pixels {'red_cube': 117, 'green_cube': 147, 'blue_cube': 133}

Figure: the same state from five cameras. Rows: RGB, metric depth (colour bar in metres), segmentation by geom id with background in black. Regenerate with examples/l3_2_render.py.
Read the numbers before the pictures, because they contain three things every perception pipeline must handle.
Depth is metric, along the camera axis, and “nothing” is not infinity. The renderer converts the OpenGL depth buffer to metres (the conversion is visible in mujoco/rendering/classic/renderer.py). Pixels that see no geometry report the far clipping plane, here 80.008 m, which is model.vis.map.zfar times the model’s extent. A depth image you feed to a network or turn into a point cloud must mask those pixels, or the network learns from a wall 80 m away.
Segmentation is per object, with a type. Each pixel holds (object id, object type); for geoms the type is mjOBJ_GEOM, background is (-1, -1). Sites drawn in the scene (the placement zones here) have their own type, so filter by type before using ids.
Visibility is not presence. The green cube covers 165 pixels from the front and 2 from the top, where the arm hides it. Datasets that compute object labels from the simulator’s state, not from the image, will happily label an object the camera cannot see.
[!implementation] Images come out flipped, and the renderer fixes it OpenGL’s framebuffer has its origin at the bottom-left; image arrays have it at the top-left. The renderer flips the rows for the OpenGL backends (and not for Filament), so
render()returns images the right way up. The browser labs in this course do the same flip when they read pixels back.
All-black or all-white images. The camera is inside a geom or pointing at the sky; or lights are missing. Render with the free camera first, then with your camera, and print the camera’s position (data.cam_xpos) and axes (data.cam_xmat).
Images do not change between steps. You rendered without calling update_scene(data, ...) again. The renderer draws the scene it was last given.
ValueError: The camera "wrist" does not exist. Attached models prefix their cameras: this one is gripper/wrist.
Rendering is slow. Software rendering (OSMesa) is much slower than EGL on a GPU, and resolution costs quadratically. Measure frames per second at your dataset’s resolution before planning a large render (the Challenge in Lesson 0.2).
Render the wrist camera every 0.1 s while the gantry gripper (from Lesson 2.3) descends onto the cube, and plot the number of cube pixels and the minimum depth over time. At what height does the cube fill a quarter of the image?
Make a 10-second MP4 of the cart-pole balance from the free camera, at 30 frames per second, with simulation and video time kept in sync (render one frame per $1/30$ s of simulated time, not per step). Use any video writer you like. State in one line how you verified the frame rate.
Rendering cost and rendering realism decide what vision-based policy research can be done in MuJoCo. The classic renderer is fast and plain; the Filament backend adds physically based materials and lighting at higher cost. Neither simulates camera noise or exposure (Lesson 0.3). Policies that must transfer visually are trained with visual randomization (Level 17) or with images from other renderers, and evaluated on images they did not train on. Report the renderer, resolution and any randomization with every vision result.
{"id": "3.2-check", "title": "Knowledge check", "questions": [
{"kind": "mcq", "q": "A depth image of a tabletop scene has a maximum of 80.008 m although the farthest wall is 3 m away. What is going on?",
"options": ["A rendering bug", "Pixels that see no geometry report the far clipping plane", "Depth is in centimetres", "The camera's fovy is wrong"],
"answer": 1,
"explain": "<p>The renderer converts the depth buffer to metres; empty pixels sit at the far plane, <code>zfar * extent</code>. Mask them.</p>"},
{"kind": "mcq", "q": "Which call makes the passive viewer show the state your loop just computed?",
"options": ["<code>viewer.render()</code>", "<code>viewer.sync()</code>", "<code>mujoco.mj_forward</code>", "Nothing; it updates on its own"],
"answer": 1,
"explain": "<p><code>sync()</code> exchanges state between your <code>mjData</code> and the viewer, including mouse perturbations.</p>"},
{"kind": "open", "q": "You will build a dataset of 100 000 RGB-D frames at 640 by 480 on a CPU-only server. What do you measure before starting, and what could make you change the plan?",
"reference": "<p>Measure frames per second for RGB and depth separately at 640 x 480 with the backend you will use (OSMesa on CPU), multiply out to total hours, and compare with the physics cost per frame. If rendering dominates and the total is too long, options are: render on a GPU node with EGL, lower the resolution, render fewer cameras, or render only the frames you need (for example at 10 Hz, not every simulation step). Also check memory and disk: 100 000 frames of uint8 RGB plus float32 depth at that size is roughly 92 GB + 123 GB uncompressed.</p>"}
]}
Lesson 3.3 edits models in code with MjSpec: building scenes procedurally, attaching sub-models, and recompiling without losing the simulation state.