Why physics plus AI, and the map
The data gap in robotics, the inputs, architecture and loss taxonomy, and the physics-guided route with three worked papers.
A five-lecture PhD course from the University of Trento and TUM, rebuilt as one page you can learn from: the map of where a physics prior can enter a model, every method with its equations, four labs that run in the page, the course's own evidence read against its tables, and a timestamp for every topic in the 8 h 38 min of video.
Ten points where the brief, the lectures or the common reading of this field need tightening. Each is settled here once and used in its corrected form below. None of them changes the course's message; several sharpen it.
The course ran in Trento from 6 to 10 July 2026 as part of Neu4mes, a 1.2 million euro project funded by the Italian Ministry of University and Research under the FIS 2023 call, with Rosati Papini as principal investigator. Its aim is a method family the lecturers call model-structured neural networks, and its outputs include the open-source framework nnodely. Piccinini, now a postdoc at TUM's Autonomous Vehicle Systems lab, prepared most of the slides and gave most of the lectures; Rosati Papini interjects throughout with practitioner's notes that are some of the most useful minutes in the recording. L1 20:11
The arc is deliberate. Day one asks why data alone will not carry physical systems and builds a taxonomy of where a prior can enter; days two and three walk the taxonomy route by route; day four shows the tooling and real racing data; day five trains a controller without new data, looks at foundation models, and lists open problems.
The data gap in robotics, the inputs, architecture and loss taxonomy, and the physics-guided route with three worked papers.
PINNs from Raissi et al., data plus physics regularisation on a quadrotor, monotonicity losses, then DeLaN, LNN and HNN.
Neural ODEs and variational integrators, residual and sub-system learning, equation learning, Koopman, then the derivation of neural compositional modules.
The framework's pipeline, the 2019 longitudinal model line by line, spectral window design, and coupled longitudinal-lateral learning on a small-scale racer.
A model-structured feedforward for a full-scale race car, closed-loop training through a frozen simulator, a PINN in nnodely, StyleVLA, and open problems.
Durations from the playlist "Physics-Informed Machine Learning for Modeling, Planning, Control and Estimation of Physical Systems" (UniTrento Ingegneria Industriale, uploaded 16 Jul 2026). Every ▶ pill on this page opens the lecture at that second.
The course opens with Ken Goldberg's question from his ICRA 2026 plenary: is robotics about to have its ChatGPT moment? Deep learning with big data solved most of vision, and large language models with internet-scale text are close to solving language. Language, Piccinini argues, is a small one-dimensional world of token strings. A robot hand has 22 degrees of freedom and a humanoid around 50, and a vision-language-action model such as π0 needs synchronised images, text and motor commands for every training example. Reading the corpus behind Qwen-2.5 would take a person about 1.2 billion hours; the matching corpus for physical action does not exist on the internet, because video carries no joint torques. L1 6:21
Four gaps follow: no internet-scale open data, a real world too complex to cover in a training set, the need to generalise to unseen scenarios from little data, and safety and interpretability for public acceptance. The community fills them four ways. Data flywheels deploy fleets to harvest data (Tesla); teleoperation pays humans to drive the robot or rescue it (remote operators behind robotaxis, humanoid start-ups); simulation trades data cost for a sim-to-real gap; and physics priors reuse what engineering already knows. The course takes the fourth route and frames it as a dial between two poles. L1 13:45
A multilayer perceptron with one hidden sigmoid layer approximates any continuous function (Cybenko, 1989), so in theory structure is unnecessary. In practice every successful architecture encodes its domain: convolutions encode translation structure in images, LSTM gates keep long-range dependencies in sequences, attention encodes relations between tokens. For physical systems the domain prior is physics: equations of motion (Newton-Euler, Lagrange, Hamilton), conservation of energy and momentum, symmetries and invariances, kinematic and dynamic constraints (holonomic or not), the connectivity of bodies and joints, and problem-specific models of friction or actuators. L1 50:38
One precision for readers coming from vision: a convolution is translation-equivariant; detecting an object anywhere in the frame is the invariance obtained after pooling. The distinction matters for physics too, because a rotation of a robot's base should rotate its predicted forces (equivariance), not leave them unchanged.
Rosati Papini's favourite demonstration is the TensorFlow Playground spiral. With the raw coordinates x1, x2 as inputs, a network with several hidden layers and thousands of weights plateaus in a local minimum. Add sin(x1) and sin(x2) as inputs, because a spiral has a trigonometric parametric form, and one hidden layer separates the classes within a few epochs. The right features cut data, parameters and training time at once; that is the whole course in miniature. L1 41:15
Lecture 2 adds a side debate worth knowing: why deep beats wide when one hidden layer suffices in theory. A student cites the exponential growth of expressivity with depth; Rosati Papini adds that sparse, structured connectivity (convolutions, attention, and later the plain "+" that sums independent forces) helps fitting by removing connections that would learn spurious correlations. He points to a published study of sparse connectivity without naming it on the recording, so treat the second claim as a heuristic. L2 19:08
Any learned model, whether a neural network, a Gaussian process or a foundation model, has three ingredients: inputs and a dataset, an internal architecture, and a loss that trains it. A physics prior can enter at each, and the course names the routes after the survey of Faroughi et al. (arXiv 2211.07377, published 2024): physics-guided for inputs and data, physics-encoded for the architecture, physics-informed for the loss. Rosati Papini adds two more places: the training procedure and the validation. Click a block. L1 55:51
Physics shapes what the model sees: input features, the representation of the data, or which training data is selected or generated. The main model can stay generic.
Section 04 · L1 1:10:33
Known physics is forced into the core architecture: physics layers, physics connections and constraints on parameters. The loss can remain a plain data loss.
Sections 06 to 08 · L2 1:03:41
Governing equations enter as soft constraints, typically squared residuals, added to the data loss: ℒ = ℒdata + ℒphysics. Two uses: learn the solution of an equation that is hard to solve, or correct an imperfect model with data and generalise where data is missing.
Section 05 · L2 0:10
Physical knowledge can schedule the optimisation itself instead of the model or the objective.
Judge the trained model with the tools of physical modelling as well as a validation RMSE.
The routes compose. A network can take physics-guided inputs, have a physics-encoded body and train on a physics-informed loss, and the lecturers' closing message is that almost every published method uses one route while the open potential is in combining them. L1 1:08:29
One difference from the source survey: Faroughi et al. treat neural operators as a fourth family next to guided, informed and encoded networks. The course folds operators into the encoded route, which is defensible when the physics enters through the operator's lifting functions.
Nine short descriptions, several of them borderline on purpose. Pick a route; the explanation appears with the answer.
0 of 9 answered
The smallest chapter, and by the lecturers' account the least explored: physics decides what the model receives, and the model itself can be any general-purpose network. Four classes, of which the course works through the first two. L1 1:11:37
Compute a physical quantity first.
Raw signals pass through a physics model or state estimator (kinematics, an EKF or UKF), and its estimate joins the inputs of a neural network. Learning happens in a physically consistent space.
Choose what to feed, and what to train on.
Derived features such as v², sign(v), energies or forces; training data generated by validated simulators or post-processed until it is physically consistent.
Respect the structure of the space.
Project or augment data so it satisfies known symmetries (a car is left-right symmetric), use SO(3) and SE(3) representations and equivariant models.
Learn the bandwidth that is physics.
Transform signals with Fourier or DCT and keep the band a mechanical system actually occupies (often below 10 to 15 Hz), so the model cannot fit vibration and sensor noise.
Sideslip estimation (Gräber et al., IEEE T-IV 2019). The sideslip angle β between a car's heading and its velocity is hard to measure without expensive GPS. A physical model computes its rate β̇ from lateral acceleration, speed and yaw rate; β̇ enters a GRU alongside the raw signals, and the GRU outputs β. The hybrid beat both the same GRU without β̇ and the physical model alone. L1 1:20:03
The relation is a one-liner: for small angles, β̇ ≈ ay/v − r, with r the yaw rate. Integrating it drifts, which is why it works better as a feature than as an estimator.
Feedforward control of a linear motor (Bolderman, Lazar and Butler, IEEE CCTA 2021). A network maps the reference trajectory to the motor force. Friction depends on velocity and its sign, so ẏ and sign(ẏ) join the inputs; the same network with these two features outperforms the baseline. It is the Playground lesson on real hardware. L1 1:24:15
Physically consistent training data (PARC, Xu et al., SIGGRAPH 2025). A diffusion model generates parkour motions for simulated characters in Isaac Gym; a physics-based tracking controller replays them and outputs what is physically feasible; the corrected motions return to the dataset that retrains the generator. The loop improves on the generator alone. L1 1:27:25
The discussion at the end of Lecture 1 places a current topic in this route: reconstructing actions or forces from internet video, often by training inverse-dynamics models in simulation and applying them to real footage. The lecturers call physics-guided learning the most emergent of the three routes for exactly this reason. L1 1:33:49
Cost and failure mode. This route is the cheapest to try, since it changes no training code, and the easiest to debug. Its ceiling is the frozen preprocessing: an estimator that is wrong in a regime feeds the network a confident wrong feature, and no gradient will ever correct it. When the feature is only roughly right, letting the preprocessing train turns the method into a physics-encoded one, with the extra risk that it drifts away from its physical meaning.
The route most people mean by the phrase. Governing equations, boundary conditions or other physical laws are written as residuals that should be zero and added to the loss as squared penalties, which keeps the objective smooth for gradient descent. Two motivations run through the chapter: learn the solution of an equation that is expensive or impossible to solve directly, and combine an imperfect physical model with data so the network generalises where data is missing. L2 1:13
Raissi, Perdikaris and Karniadakis (J. Comput. Phys. 2019) start from a general nonlinear PDE for an unknown u(t, x) with parameters λ inside a nonlinear operator 𝒩. Two questions follow: given λ, can a network learn u (the forward problem)? Given data, which λ fits it (discovery, the inverse problem)? For the first, call the left side f and train a network for u with a loss in two parts. L2 5:25
Three details carry the method. The collocation points cost nothing to add, since the residual needs no measurement, so one can sample hundreds of thousands of them. The residual contains derivatives of the network with respect to its own inputs, ∂u/∂t and ∂²u/∂x², computed by automatic differentiation; Rosati Papini singles this out as the genuinely new step for the neural-network community, using the network's gradient as an output that enters the loss. And the network itself is a generic multilayer perceptron: no physics in the architecture. L2 27:36
The lecture is candid about the limits. Nothing guarantees convergence to the true solution; if the PDE is well posed and the network deep enough it works in practice. The weights between terms are a manual trade-off between honouring boundaries and honouring the equation. Inference is a forward pass instead of hours of numerical solving, but the trained network solves one problem: change the initial or boundary conditions and it must be retrained. L2 12:49 L2 33:59
Two well-documented failure modes are worth knowing before trying this at scale: imbalanced gradients between the loss terms (Wang, Teng and Perdikaris, SIAM J. Sci. Comput. 2021) and training failures on PDEs with strong convection or high-frequency solutions (Krishnapriyan et al., NeurIPS 2021). Learning a map from boundary conditions to solutions instead of one solution is the job of neural operators such as DeepONet (Lu et al., Nature Machine Intelligence 2021).
Nonlinear Schrödinger equation (Raissi et al.). A complex solution h = u + iv on x ∈ [−5, 5], t ∈ [0, π/2], known only at t = 0 as h(0, x) = 2 sech(x), with periodic boundaries in value and in ∂h/∂x. A five-layer network trained on the initial condition, the periodic boundaries and the residual reproduces the numerical solution without a single interior measurement. L2 22:20
Pneumatic soft finger (Beaber, Liu and Sun, RA-L 2024). A network maps a point (x, z) and the air pressure p to the deformations ux, uz; the loss holds the initial condition and the two Navier-Cauchy continuum equations with body forces from gravity. Benchmarked against ANSYS finite elements, the trained network predicts the finger's shape in real time across pressures. The paper calls itself physics-guided, which is the taxonomy confusion of correction 2 in one title. L2 35:01
An ODE is a PDE in one variable, so the recipe applies to any robot. Bianchi et al. (Drones 2024) learn quadrotor dynamics with a data loss on flight logs plus a physics loss on the rigid-body equations at collocation times, with γ setting how much the physics is trusted; it beat an extended Kalman filter at estimating the parameters. The lecturers then make the point they consider most useful for real systems: real data usually covers a small region of the operating domain, while an imperfect model is available everywhere. Apply the data loss where data exists and the physics loss where it does not, and the network fits the measurements while extrapolating physically. ℒphysics needs no ground truth. L2 49:49
What is computed. The system is a damped oscillator x″ + 2ζωx′ + ω²x = 0 with ω = 2 rad/s, ζ = 0.08, x(0) = 1, observed 16 times with noise inside the green window. The model is a sum of 48 Gaussian bumps; its weights enter linearly, so the data loss and the ODE residual at 120 collocation points are both quadratic and the optimum is solved exactly with a Cholesky factorisation at every slider move. A real PINN uses a multilayer perceptron trained by gradient descent: same loss, plus optimisation error. Try three things: γ at its minimum (pure data) and watch the fit collapse to zero outside the window; any γ above about 10−3; then the model's ω at 1.8 with γ = 100, where trusting a wrong model bends even the fit inside the data.
Gu, Primatesta and Rizzo (RAS 2024) learn a quadrotor's inverse dynamics, states to PWM motor commands, and add a loss that encodes conservation of momentum indirectly: the angular acceleration about each axis should move with the PWM signal that produces it, which they verify from data with clear correlations for roll, pitch and yaw. The local-monotonicity terms are soft constraints next to the data loss. L2 55:08
The assumption fails under strong aerodynamic effects, where command and acceleration decouple. Their answer is a cyclical annealing schedule that periodically drives λLM to zero after about 40 epochs, letting the network learn the aerodynamics the prior contradicts. Rosati Papini files this under physics in the training procedure: the prior is used, then deliberately switched off. L2 58:21
Cost and failure mode. No change to the architecture and no extra data, which is why this is the popular route. The price is tuning: each residual has its own units and scale, γ spans orders of magnitude (the lab's slider covers nine), and a wrong model trusted too much biases the fit everywhere. The constraint is soft, so nothing is guaranteed at deployment.
The architecture route is the largest, and the course splits it into six classes. The first two learn dynamics through the mathematics of motion itself: learn the scalar functions behind the equations of motion, or learn the vector field and integrate it with structure. L2 1:04:45
Newton's formulation needs every force and moment on a free-body diagram, which becomes a nightmare for a humanoid. Lagrange's is energy-based: write kinetic minus potential energy in generalised coordinates q (joint angles, for instance), apply the Euler-Lagrange equation, and the familiar manipulator equation falls out, including the Coriolis and centripetal term c, which needs no separate modelling. L2 1:09:00
Classical engineering measures or estimates M and g from masses, lengths and inertias, and inertias are notoriously hard to measure. Deep learning ignores the structure and learns everything, needing far more data. Rosati Papini's estimate for a well-modelled robot is that physics explains 90 to 95% of the output; relearning that part from data is waste. L2 1:12:14
Keep the equation, learn M and g.
Lutter, Ritter and Peters (ICLR 2019). Networks output M(q) and V(q) or g(q); c comes for free by automatic differentiation of M. M is parameterised as L(q)L(q)⊤ with L lower-triangular, so symmetry is structural. Torques from the equation are matched to measured torques with a plain MSE, so the loss is ordinary and the physics sits in the architecture. On a 7-DoF arm, a DeLaN feedforward plus PD feedback tracked better and generalised better than a generic feedforward network. L2 1:14:26
Learn the whole Lagrangian.
Cranmer et al. (ICLR workshop 2020). DeLaN's form T = ½q̇⊤Mq̇ holds for rigid bodies but not for a charged particle. An LNN learns ℒ(q, q̇) with a generic network and gets accelerations from the Euler-Lagrange equation, below. On a frictionless double pendulum, a baseline network integrated for hundreds of seconds loses energy and falls out of phase; the LNN keeps it. L2 1:30:18
Learn the energy, take its symplectic gradient.
Greydanus, Dzamba and Yosinski (NeurIPS 2019). The Hamiltonian ℋ(q, p), usually total energy T + V, is learned by a network; time derivatives are its partial derivatives with a sign swap. Momenta p are the catch: easy for a mass on a spring, hard to define for complex systems. Coupled with an autoencoder, an HNN learns a pendulum's dynamics from pixels by treating latents as (q, p). L2 1:45:13
A recurring exchange in Lectures 2 and 3 is worth keeping. A network trained on one-step accelerations can either gain or lose energy when integrated. Rosati Papini's experience is that networks trained on integrated multi-step losses tend to lose it, because a model that gains energy makes the loss explode later in the rollout, which gradient descent avoids. L2 1:38:46
Neural ODEs (Chen et al., NeurIPS 2018) replace the right side of dh/dt = f(h, t, θ) with a network and backpropagate through any ODE solver by the adjoint sensitivity method, with memory linear in problem size and error control. Its discrete-time version, ht+1 = ht + f(ht, θt), is a residual network. Rosati Papini notes the practical difference: a residual network learns a fixed time step, while the continuous form can be sampled anywhere, and calls the Neural ODE the most flexible tool of the chapter because any function can be plugged in. L3 3:14 L3 21:06
Euler and Runge-Kutta steps do not preserve the geometry of mechanical systems, so even a perfect vector field drifts in energy. Variational Integrator Networks (Saemundsson et al., AISTATS 2020) integrate with symplectic variational integrators that conserve energy and momentum to third order or higher. With noisy observations and a small training set, an HNN overfit and drifted while the VIN held energy on a mass-spring system and a pendulum; with a medium set the gap closed. L3 10:35
What is computed. A frictionless pendulum, ℋ = p²/2 + 1 − cos q. Three vector fields: the exact one; a black-box stand-in, ṗ = −sin q + εp, whose small error is not the gradient of any energy (the typical error of an unconstrained network, here shrunk to one term); and an HNN stand-in, the exact symplectic gradient of a slightly wrong ℋθ = p²/2 + (1 + δ)(1 − cos q). Each is integrated with the chosen scheme and compared with a fine Runge-Kutta reference. Read it in three steps: under Euler at the default step everything gains energy, the exact field included, so the integrator alone causes drift; under leapfrog the exact and HNN fields stay bounded for any horizon while the black box still drifts, so structure must sit in the field too; and the HNN conserves its own ℋθ, not the true energy, so its phase error against the exact pendulum grows steadily.
Cranmer et al.'s comparison, as the course presents it, reduces the choice to four questions: are physics priors available, do you want a differential equation, does energy conservation matter, and is the structure of the Lagrangian known? L3 17:54
| Property | Neural net | Neural ODE | HNN | DeLaN | LNN |
|---|---|---|---|---|---|
| Models a dynamical system | ✓ | ✓ | ✓ | ✓ | ✓ |
| Learns a differential equation | · | ✓ | ✓ | ✓ | ✓ |
| Learns exact conservation laws | · | · | ✓ | ✓ | ✓ |
| Learns from arbitrary coordinates | ✓ | ✓ | · | ✓ | ✓ |
| Learns arbitrary Lagrangians | · | · | · | · | ✓ |
From Cranmer et al., "Lagrangian Neural Networks" (arXiv 2003.04630), as shown on Encoded I slide 51. The lecture adds that the table is written by LNN's authors.
Where this meets robot learning. These networks assume conservative or rigid-body structure and smooth coordinates. Manipulation is dominated by contact, friction and switching, where energy is not conserved and generalised coordinates of the scene are unknown, and a VLA's action space is usually an end-effector command behind a low-level controller that already hides the arm's rigid-body dynamics. The HNN-on-pixels result is a small latent world model with a conservation prior; the open question for world models is which part of a scene has dynamics clean enough for such a prior. The robot's own body and base are the obvious candidates.
Networks stay generic black boxes but become modules inside a physical model: they correct it, replace a part that is hard to model, or preprocess sensors for it. Historically this was among the first ways of merging data and physics, and it remains the most common in robotics. L3 23:13
Learn what the model gets wrong.
Kabzan, Hewing, Liniger and Zeilinger (RA-L 2019) add Gaussian-process residuals to an eight-state single-track vehicle model, xk+1 = f(xk, uk) + Bd(d(zk) + wk), and learn them online from a dictionary of driving data inside an MPC. The lecture reports a 10% lap-time improvement over the physics-only controller. L3 27:30
Replace the part that is expensive to identify.
Wegrzynowski et al. (IROS 2024) keep a four-state vehicle model but let a network output the four tyre forces, which otherwise need a tyre test rig. The model runs inside an unscented Kalman filter whose noise covariances are trained end to end with it; it beat GRU and LSTM estimators and transferred zero-shot to unseen road conditions. L3 31:47
Let a network turn pixels into physics.
A convolutional network converts camera images into quantities a physical model consumes. When that network is frozen it is the physics-guided route; when it trains with the model it is encoded. The course mentions this class without a worked example.
Rosati Papini adds a caution against using LSTMs and GRUs as general dynamics models. They were designed for text, where an opened bracket must close hundreds of tokens later, so their gates can hold memory indefinitely. In a physical system the effect of a pedal press fades, and a gated memory is the wrong inductive bias; Piccinini extends the point to transformers and diffusion models applied to vehicle dynamics, which in their experience overfit. L3 36:02
This is the fading-memory property of correction 3 seen from the other side. A recurrent network with weight decay and early stopping can still fit fading dynamics; the argument is about what the architecture makes easy, which matters most when data is short.
Equation-learner networks replace activation functions with a dictionary of candidates (sin, cos, products, identity) and learn sparse weights that switch candidates on or off, so the trained network reads as a formula. Rosati Papini describes the training trick from the original equation learner: after some hundred epochs, inspect the weights, delete unused units and continue, so the network concentrates on the active terms. SINDy is the regression version of the same idea. Physics enters through the dictionary: if a manipulator's dynamics are trigonometric, trigonometric functions go in. L3 40:19
Neural networks map fixed-size vectors to vectors; neural operators map functions to functions and come with their own approximation theorems. The course names Koopman networks (the most popular, and in the lecturers' words very hyped at conferences), Kolmogorov-Arnold networks, DeepONets and graph neural operators. A Koopman network learns lifting functions that embed a nonlinear system in a space where it evolves linearly, plus a decoder back; once linear, the whole toolbox of linear control and linear MPC applies. Physics enters through the lifting functions, and the lecturers judge physics-informed Koopman work to be at an early stage. L3 45:44
A caveat the course leaves implicit: exact finite-dimensional linear lifts exist only for special systems, so in practice the lift is approximate and its linear prediction degrades with horizon. For control-affine systems the lifted model is usually bilinear in the input rather than linear.
| Encoded class | The network learns | Physics fixes | Pays off when | Watch for |
|---|---|---|---|---|
| Lagrangian, Hamiltonian | M and V, or the whole ℒ or ℋ | Euler-Lagrange or Hamilton's equations | conservative, rigid-body or energy-based systems; long rollouts | needs coordinates, and momenta for HNN; dissipation and contact need extensions |
| Neural ODE, VIN | the vector field f | continuous time; symplectic stepping for VIN | irregular sampling; a differential model is wanted | no conservation by itself; adjoint cost; stiffness |
| Hybrid | a residual, a sub-system or a sensor map | the rest of the physical model | a good model with known weak spots exists | the residual can absorb modelling errors and extrapolate badly |
| Topology learning | which candidate terms and connections | the dictionary | you want an equation; low-dimensional systems | the dictionary must contain the answer; noisy derivatives |
| Neural operators | maps between function spaces, Koopman lifts | lifting functions, linear lifted dynamics | families of PDE solutions; linear control of nonlinear plants | approximate lifts; few physics-encoded variants so far |
| Model-structured | input maps, regime weights, FIR responses, a residual | layers, connections and bounds from the equations | dynamics from little data; interpretability; selective adaptation | design effort per system; derivation assumes fading memory |
My synthesis of Sections 06 to 08; the class names and examples are the course's.
The lecturers' own class, which they named in papers from 2025 and 2026 to cover approaches that fit none of the existing categories. Everything lives in the architecture, under three principles, and one practical guideline: start from an imperfect physical model of the system and neuralise its architecture. L3 51:08
Physics layers. Each branch learns one named force: aerodynamic drag from v², rear braking force from the negative part of the pedal only (the rear axle of a front-wheel-drive car never pulls), front traction from the positive part. Physics connections. Newton says forces superpose, so the branches are summed, not mixed by another layer. Rosati Papini's comment on that plus sign is one of the best moments of the course: a mixing layer would add connections and let the network learn a correlation between independent forces that does not exist; the plus sign is sparsity chosen by physics. Physics constraints. Parameters that have a physical meaning get bounds through the architecture, never through a loss penalty. L3 59:37
The general recipe composes neural compositional modules (NCMs) Ψi, each a function of a window of past inputs and of scheduling variables, with sums and products. The derivation is the heart of the course and repays slow reading. L3 1:04:58
Substitute the state equation into the output equation recursively, n times.
If the window is long enough, the free response from xk−n has died out and the forced response dominates, so yk ≈ 𝒢(uk, …, uk−n). Piccinini's image: after holding the accelerator for a while, the speed you started from no longer matters. L3 1:09:16
This is the step that needs fading memory (correction 3). Choose outputs that forget, such as accelerations, rates and forces, and integrate them outside the network.
Dynamics change with speed or gear, so add scheduling variables: yk = Ψ(uk, …, uk−n, 𝐳k). This is a full NCM's signature: an input history, a regime, one output.
The output is the convolution of the input with the discrete impulse response Γ. Truncating the infinite sum gives a finite impulse response, which is exactly a linear layer g whose weights are the impulse response coefficients. L3 1:12:29
Give each of M regimes its own FIR layer 𝐠j and an activation 𝛗j(𝐳) that says how much regime j applies. The activations must be non-negative and sum to one everywhere, so no branch is favoured by construction; triangular for continuous scheduling (speed), rectangular for discrete (gears), Gaussians are also common. ⊙ is the element-wise product over the time window. L3 1:16:46
A function 𝐟j transforms the input window before scheduling. With strong priors it is a physical map (v² for drag, v·ω for lateral acceleration); with partial priors an equation learner over a dictionary (trigonometric for a manipulator); with none, a learnable Taylor polynomial a0 + a1u + a2u², or a small general-purpose network. L3 1:30:29
For a fully nonlinear system, add a general-purpose residual network ℛ in parallel. This is the complete module.
p inputs stack into a matrix of windows. With q scheduling variables (engine speed and gear, say), regime j's activation is the element-wise product of its one-dimensional activations. Products of partitions of unity are again a partition of unity, so two triangular axes give pyramids and three give the hyper-pyramids used later for steering. L3 1:37:51
Modules combine by sums (independent forces and moments, each learned by its own Ψ), products (a lateral force times the sign of the steering angle) and cascades (a steady-state module feeding a transient one). The claim is that sums, products and cascades of NCMs are flexible enough for the dynamics of physical systems. L3 1:44:22
The order of f and g can be swapped: compress time first, then schedule. The two are not equivalent; Rosati Papini uses the swap when the regime is discrete and memory across a switch is meaningless, as with gears, where the new gear filters the engine instantly. Feeding past outputs back, through the residual for instance, turns the module autoregressive, which the lecturers use mostly for control. L3 1:49:39 L3 1:53:58
Trainable or frozen activations? Trainable centres let the network place regimes where the dynamics change, but with an imbalanced dataset they migrate to the dense region and abandon the rare one. Rosati Papini's advice: if you know the data is imbalanced, freeze the activations so every part of the range keeps a model, even one fitted on six points. L3 1:22:00
The pros the lecturers list: trained f and g weights can be read physically, task complexity splits across modules, and sample efficiency and generalisation beat black boxes. The con: design effort in the f modules, and in Rosati Papini's words there is no free lunch, because you have to know the system. L3 1:47:33
Read as system identification, the module is familiar. With M = 1 and f the identity it is an FIR model; with a static f it is a Hammerstein model with FIR dynamics; with M > 1 and partition-of-unity weights it is a local-model network in the Takagi-Sugeno and LOLIMOT tradition, that is, a linear-parameter-varying FIR model; swapping f and g gives a Wiener-type structure. A student draws the parallel to interacting multiple models on the recording and the lecturer agrees. What the course adds is the grammar: modules that compose by sums, products and cascades mirroring the equations of motion, trained jointly with residual networks by gradient descent, with tooling that handles windows, closed loops and export. That grammar, not the block, is the contribution to judge. L3 1:29:25
The scheduling variable runs from 0 to 60 m/s. Pick the non-normalised Gaussian to see why the sum-to-one rule exists.
What is computed. Top: the activation functions φj over a scheduling speed and their sum, with the weights each regime gets at the probe; triangles give at most two active local models at any speed. Bottom: the weights an FIR layer converges to when the true system is a first-order lag with time constant τ and a dead time, sampled at Δt over a window of n + 1 samples, drawn oldest to newest as on the course's slide. A peak away from the newest sample is the dead time, the reading Rosati Papini demonstrates in Lecture 4. The spectral rule is the lecture's: if the useful content ends near f, keep 3 to 5 periods of f in the window. With the defaults (3 Hz, Δt = 0.05 s) it suggests 20 to 33 samples; the course chose n = 25, 1.25 s.
nnodely (read "modely": the double n stands for m) is an MIT-licensed Python framework on top of PyTorch, written by
Rosati Papini's group to build, train and deploy model-structured and other physics-embedded networks. It installs
with pip install nnodely; the stable release on PyPI is 1.5.4, with 1.5.5 in pre-release, and the main
branch already carries a NeuralODE layer with Euler and Runge-Kutta solvers that the lecture announces for the next
release. Worked applications live in a second repository, nnodely-applications.
Rosati Papini's motivation is the audience: engineers who know their physics but find raw PyTorch an obstacle, and who are better served by blocks that already mean something physical than by asking a language model for code. The pipeline has six phases: define the neural model; build the dataset from raw CSV files, with every time window derived from the model definition; train, with layer-specific helpers such as exponential initialisation of FIR weights; validate, with RMSE and the Akaike criterion; compose validated sub-models (a dynamics model and a controller) into a larger network and train again; export to native PyTorch or ONNX for C++ deployment. L4 7:28
Click a line. The code is the repository's model_longit_vehicle_dynamics.py, quoted verbatim apart from omitted plotting and file-path lines.
# Dimensions of the layersn = 25na = 21 #Create neural model inputsvelocity = Input('vel')brake = Input('brk')gear = Input('gear')torque = Input('trq')altitude = Input('alt',dimensions=na)acc = Input('acc') # Create neural network relationsair_drag_force = Linear(b=True)(velocity.last()**2)breaking_force = -Relu(Fir(W_init = 'init_negexp', W_init_params={'size_index':0, 'first_value':0.002, 'lambda':3})(brake.sw(n)))gravity_force = Linear(W_init='init_constant', W_init_params={'value':0}, dropout=0.1, W='gravity')(altitude.last())fuzzi_gear = Fuzzify(6, range=[2,7], functions='Rectangular')(gear.last())local_model = LocalModel(input_function=lambda: Fir(W_init = 'init_negexp', W_init_params={'size_index':0, 'first_value':0.002, 'lambda':3}))engine_force = local_model(torque.sw(n), fuzzi_gear) # Create neural network outputout = Output('accelleration', air_drag_force+breaking_force+gravity_force+engine_force) vehicle.addModel('acc',[out])vehicle.addMinimize('acc_error', acc.last(), out, loss_function='rmse')vehicle.neuralizeModel(0.05) data_struct = ['vel','trq','brk','gear','alt','acc']vehicle.loadData(name='trainingset', source=data_folder, format=data_struct, skiplines=1) optimizer_params = [{'params':'gravity','weight_decay': 0.1}]training_params = {'num_of_epochs':150, 'val_batch_size':128, 'train_batch_size':128, 'lr':0.00003}vehicle.saveModel()vehicle.exportONNX()
Each block of the file maps onto a piece of the NCM formalism from Section 08. The annotation names the piece and the physics behind it.
Input('x'), Output, Parameter, Constant: named signals and
learnable or fixed scalars. Dataset columns are matched by name.x.last(): the current sample. x.sw(n): a window of n past samples;
x.sw([0, 10]): the present and 10 future samples, for feedforward control. x.tw(T):
a window in seconds.Fir: linear layer over the time dimension, an impulse response. Linear: linear layer
over a feature or space dimension.Fuzzify: triangular or rectangular activation functions by number of centres, range or explicit
centres. LocalModel(input_function, output_function): a full NCM.ParamFun: any PyTorch expression with learnable parameters (saturations, local hyperplanes).
EquationLearner: dictionary-based layers.Differentiate, Integrate: time and partial derivatives, for PINN residuals and Sobolev
training. NeuralODE on the main branch.addModel, addMinimize: register sub-models and loss terms. neuralizeModel(dt):
fix the sample time and build the graph.closedLoop, connect, addClosedLoop, addConnect: wire outputs back
to inputs, within a model or across models.loadData, filterData, resamplingData; trainModel with
models= to train only some sub-models and prediction_samples for multi-step rollouts;
trainAndAnalyze.saveModel, exportPythonModel, exportONNX.Names checked against nnodely at commit 777cf84 on 1 Oct 2026.
Racing is the lecturers' testbed: closed circuits make emergency-level driving safe, the dynamics turn strongly nonlinear near the handling limits where physics models struggle, and overtaking adds adversarial decisions. All three studies sit inside a hierarchical stack of planning, feedforward and feedback control, and state estimation. L4 13:46
The naive model is a multilayer perceptron from pedal and speed to acceleration. The structured one rewrites Newton's longitudinal balance as a sum of NCMs: front force from a window of pedal values with one FIR per gear, rear braking force, drag from vx², and a lateral-force term times the sign of the road-wheel angle, itself estimated from the steering-wheel angle by a small network. On unseen data the structured network tracks the measured acceleration more closely than the unstructured one. L4 16:52
The validation step is the part to copy. The power spectral density of the residuals is lower for the structured model across the band, and it stays above the noise floor, which was estimated from constant-speed driving; a residual below the noise floor would mean the model fits noise. The signal itself meets the noise floor at 2 to 3 Hz, so nothing above that is vehicle dynamics. That fixes the window: three to five periods of the slowest relevant content, about 1 to 1.5 s, so n = m = 25 samples at 0.05 s, 1.25 s. Sampling at 20 Hz caps the analysable band at 10 Hz. Longer windows are possible but invite the network to learn long-range correlations that are not physics. L4 35:42 L4 39:52
A small-scale Roboracer car driven fast on TUM's indoor track. Two structured networks feed each other: the lateral one predicts yaw rate Ω, the longitudinal one acceleration ax, and each consumes the other's prediction internally, from windows of speed, steering angle, motor current and road slope. The lateral network starts from an empirical fact rather than Newton: curvature against steering angle is almost a line whose slope changes with speed and acceleration. A quasi-steady-state NCM learns that surface with local hyperplanes under three-dimensional triangular activations; a cascaded transient NCM with FIR outputs, scheduled by speed and acceleration in 2 × 2 regimes, learns the dynamics. The trained FIR weights rise towards the present with no dead time beyond one sample, and their mild oscillation points to higher-order dynamics. L4 1:11:16 L4 1:20:44
| Yaw rate, validation | Large set RMSE | Medium set RMSE | Small set RMSE | Small set FVU |
|---|---|---|---|---|
| G-NN, general purpose | 0.142 | 0.172 | 0.224 | 0.155 |
| MS-NN-bench, earlier structured model | 0.154 | 0.158 | 0.181 | 0.066 |
| MS-NN, ablation 1 | 0.122 | 0.110 | 0.134 | 0.053 |
| MS-NN, ablation 2 | 0.116 | 0.116 | 0.124 | 0.041 |
| MS-NN-full | 0.109 | 0.106 | 0.137 | 0.061 |
RMSE in rad/s; FVU is the fraction of variance unexplained, 1 − R². Lowest per column highlighted. nnodely deck slide 38; seeds and intervals are not shown.
Then the adaptation test, the result I find most useful in the course. Swap tyres or add 10% mass, collect 20 s of driving, and fine-tune only the sub-module physics says has changed: the transient module for tyres. Fifty to sixty epochs suffice, fast enough to run online. L4 1:28:06
| Yaw rate RMSE, zero-shot → fine-tuned | Tyre set 1 | Tyre set 2 | +10% mass |
|---|---|---|---|
| G-NN | 0.147 → 0.107 | 0.110 → 0.070 | 0.088 → 0.046 |
| MS-NN-bench | 0.102 → 0.096 | 0.079 → 0.059 | 0.072 → 0.053 |
| MS-NN-full | 0.172 → 0.085 | 0.129 → 0.053 | 0.075 → 0.045 |
| Longitudinal acceleration RMSE, zero-shot → fine-tuned | Tyre set 1 | Tyre set 2 | +10% mass |
|---|---|---|---|
| G-NN | 0.614 → 0.380 | 0.624 → 0.368 | 0.489 → 0.329 |
| MS-NN-bench | 0.462 → 0.425 | 0.373 → 0.347 | 0.396 → 0.375 |
| MS-NN-full | 0.531 → 0.395 | 0.468 → 0.334 | 0.398 → 0.290 |
Yaw rate in rad/s, acceleration in m/s². Bold: best fine-tuned value per column; red: worst zero-shot value. nnodely deck slide 39. The slide does not say how the general network was fine-tuned on the same 20 s.
Read both tables together. The full structured model does not transfer better zero-shot; after a tyre change it transfers worst. Its advantage is that it adapts best from 20 s of data, which is what a selectively trainable structure should buy, and it is a different claim from better generalisation.
TUM's controller won the 2024 Abu Dhabi Autonomous Racing League with a heuristic feedforward: a kinematic steering angle, an understeer correction, a longitudinal-acceleration correction and an offset. MS-NN-steer replaces it. L5 2:12
The structured design starts from the handling diagram, which plots how far a car's steering departs from the kinematic ideal as lateral acceleration grows. Local linear models under triangular activations learn its shape. Activations read |ay| and the local models carry sign(ay), so one parameter set serves left and right turns; the coefficients then become functions of speed, and finally of longitudinal and vertical acceleration, with each local model designed as the local approximation of a double-track vehicle model around an operating point. A transient NCM scheduled by speed and acceleration follows, and a small network maps road-wheel to steering-wheel angle. Inputs are saturated to the training range before entering, so the controller cannot extrapolate wildly on the car. L5 4:17 L5 25:22
The windows look forward, not back: 10 samples of the planned trajectory from the high-level planner. A feedforward that sees the desired future can lead the steering actuator's delay, and the trained FIR weights show it: a decaying profile means no dead time, while a peak 0.2 s into the future reads as a 0.2 s actuator delay. L5 14:52
| Steering angle error | Medium train | Medium valid | Small train | Small valid | FVU valid, small |
|---|---|---|---|---|---|
| G-NN: concatenated windows, two layers, ELU | 0.160 | 0.229 | 0.132 | 0.327 | 0.0792 |
| MS-NN-base (Piccinini et al. 2023) | 0.079 | 0.097 | 0.069 | 0.115 | 0.0026 |
| MS-NN-steer | 0.068 | 0.091 | 0.063 | 0.097 | 0.0020 |
RMSE in degrees on recorded Yas Marina telemetry. Small set: sector 3 of one lap; medium: sectors 1 and 3; validation: a second lap. Paper table shown on nnodely deck slide 65; the same excerpt reports 10 seeds in which MS-NN-steer's error barely moves and the G-NN's varies widely, with a similar picture across learning rates from 10−3 to 10−5.
Two readings the headline hides. Most of the gain over the general network was already in the 2023 base model (0.327 to 0.115 on the small set); the 2025 extension adds 16% on top. And the baseline is a two-layer perceptron on concatenated windows, so the comparison shows what structure buys over a small unstructured network, not over a tuned sequence model or an identified physical model. The seed study is the strongest evidence in the course: insensitivity to initialisation is a property a practitioner can use.
Can a controller learn to control a plant without collecting new data? Three steps. Fit a structured simulator of the plant with supervised learning. Connect an untrained structured controller to it in closed loop, the simulator's prediction feeding back as the measurement. Train the controller with the simulator frozen. Reinforcement learning would work but needs millions of interactions and treats the plant as a black box; here the simulator is differentiable, so the controller can follow its gradient. L5 33:51
The lecturers call this differentiable closed-loop controller training and use it in their lab, including with nonlinear MPC, where they report roughly an order of magnitude better sample efficiency than reinforcement learning, while RL can still find better solutions because it explores. The same gradient allows fast online adaptation when the simulator's parameters change, and the loss can be unrolled over a long horizon rather than one step. L5 41:15 L5 49:50
A student asks the right question: what if the simulator is wrong? Then the controller learns the wrong policy. The remedies named are to keep adapting the simulator on real data and fine-tune the controller, or to estimate the gradient from real measurements without a simulator. Rosati Papini adds the structural argument: a generic network controller trained through a generic network simulator finds strange inputs in regions the simulator never saw, where it extrapolates unphysically; a structured simulator extrapolates more physically, so the exploit is smaller. He cites an ICRA paper whose simulator had a physics branch and a learned residual, where gradients were backpropagated through the physics branch only. L5 44:28
The worked example is nnodely's mass-spring-damper. The direct model, two FIR branches, is exact for a linear
system; it is trained one step ahead and then refined in closed loop over 1,500-step rollouts with a smaller
learning rate, because long rollouts can destabilise training. The controller is a PID whose three gains are the
only trainable parameters; nnodely's built-in integrate and differentiate operations form the error terms, and the
models= argument freezes the simulator. The trained gains landed close to MATLAB's PID autotuner, and
adding an overshoot penalty to the loss reshaped the response. L5 50:52
Train once with no simulator error. Then set the mass error to −40%, so the simulator believes the plant is lighter than it is, reset and train again: the gains look fine on the simulator and overshoot on the real plant.
What is computed. A real plant x″ + 0.4x′ + x = F and a simulator with the same structure and adjustable errors in mass and damping, both stepped at 0.01 s. A PID with gains K = exp(θ) acts on the simulator for 8 s against a unit step. The loss is the mean squared tracking error, plus β times the mean squared overshoot, plus an effort term (0.02 · mean F²) that keeps gains moderate. Exact gradients come from forward-mode differentiation through all 800 steps, and Adam updates θ. The real plant is never used for training; its curve shows what the gains do when deployed.
For readers from model-based reinforcement learning this is familiar: policy optimisation by backpropagation through a learned model, the family of PILCO, Dreamer's actor learning in imagination and short-horizon actor-critic methods. Rosati Papini's warning is the known failure of that family, model exploitation, and his remedy, a model that extrapolates physically, is a structural answer to a problem that the RL literature usually attacks with ensembles and uncertainty penalties.
The lecturers see few examples so far and much interest: everyone at ICRA wants physics-informed foundation models, few know how to build one. Their argument is the course's opening one: embodied systems lack data, and physics can fill part of the gap. Their example comes from Piccinini's lab. L5 1:19:33
StyleVLA (Gao, Hua, Piccinini, Schäfer, Moller, Li and Betz; arXiv 2603.09482, March 2026) asks a vision-language-action model to drive in a requested style: default, balanced, comfort, sporty or safety. It fine-tunes Qwen3-VL-4B on an instruction dataset built in CARLA, 1.2k scenarios with 76k bird's-eye and 42k first-person samples. Only the vision encoder, low-rank adapters in the language model and an MLP trajectory decoder train, which fits a single consumer GPU. The model reads the image, the style prompt, the current state with 0.5 s of history and a goal, and outputs 3 s of positions, heading, speed and acceleration.
The loss mixes token cross-entropy, a regression term on the decoded trajectory, and a physics-informed kinematic consistency term (PIKC): roll the model's own previous waypoint forward with its predicted speed, heading and acceleration, and penalise the distance to the waypoint it actually predicted next. No ground truth enters the term. L5 1:26:53
On the paper's composite driving score StyleVLA reaches 0.55 on bird's-eye and 0.51 on first-person inputs, against 0.32 and 0.35 for Gemini 3 Pro used zero-shot; proprietary and open baselines, Alpamayo-R1 among them, trail. The lecture's two ablations follow.
| Training data | ADE m | FDE m | PSR |
|---|---|---|---|
| Small, 4.5k | 2.08 | 5.43 | 20.60% |
| Medium, 20k | 1.51 | 3.92 | 27.14% |
| Large, 40k | 1.47 | 3.81 | 29.37% |
| Standard, 50k | 1.17 | 3.06 | 33.19% |
| Loss, at 50k | ADE m | FDE m | PSR |
|---|---|---|---|
| CE | 1.47 | 3.82 | 29.00% |
| CE + REG | 1.21 | 3.17 | 32.08% |
| CE + REG + PIKC | 1.17 | 3.06 | 33.19% |
Outlook slide 9. ADE and FDE: average and final displacement error. PSR: share of trajectories with ADE below 1 m. No seeds or intervals on the slide.
The lecturers' take-home message: even in a four-billion-parameter model, where many would say data solves everything, a very basic physics term still helps, and they expect such examples to multiply in the next couple of years. L5 1:30:03
What the numbers support. PIKC moves ADE by 0.04 m at the largest data size, about a third of the step the regression head gave, with no seeds to tell it from run-to-run variation. The course's own thesis is that physics pays most where data is scarce, and that is the experiment missing here: the PIKC ablation at 4.5k samples. The comparison with zero-shot proprietary models measures in-domain fine-tuning as much as anything physical. The idea itself transfers cleanly to manipulation: a self-consistency penalty between predicted poses, velocities and gripper states in an action chunk costs nothing at inference and needs no labels.
A course is not a paper and owes no ablations. Its quantitative claims still teach the reader what physics priors buy, so each is read here through one checklist: what was measured, on which protocol, with how many runs, and whether the claim on the slide follows.
| Claim | Source | What was measured | Protocol notes | Reading |
|---|---|---|---|---|
| Structured networks generalise better from small data | Mungiello et al. 2026; Piccinini et al. 2025 | Offline RMSE and FVU on a held-out lap or log of the same vehicle | Same platform, same track; small baselines; seeds reported only in the steering paper | Supported for in-distribution held-out data |
| Selective fine-tuning adapts to new tyres or mass in 20 s | Mungiello et al. 2026 | Yaw-rate and acceleration RMSE before and after fine-tuning | Zero-shot the full model is worst after tyre changes; G-NN fine-tuning protocol not stated | Supported for adaptation, not for zero-shot transfer |
| MS-NN-steer beats the A2RL-winning controller | Piccinini et al. 2025 | Error against executed steering on recorded laps | No closed-loop lap with the learned feedforward reported | Partly: offline accuracy only |
| Insensitive to initialisation and learning rate | Piccinini et al. 2025 | RMSE spread over 10 seeds; learning rates 10−3 to 10−5 | Spread shown as box plots in the paper | Supported |
| GP residuals improve lap time by 10% | Kabzan et al. 2019 | Lap time in closed loop | Stated in the lecture; I did not re-read the original paper for this page | Cited, not re-examined here |
| Differentiable controller training is about 10× more sample-efficient than RL | Lecturers' differentiable MPC work | Not shown in the course | Paper not named on the recording | Not shown |
| A physics loss helps a 4B VLA | StyleVLA, arXiv 2603.09482 | ADE, FDE, PSR at 50k training samples | 0.04 m ADE gain; one run per configuration on the slide; preprint | Weak: direction plausible, size within unknown noise |
| LNN and HNN conserve energy, VIN fixes small-data drift | Cranmer 2020; Greydanus 2019; Saemundsson 2020 | Energy over long rollouts | Pendulum, double pendulum, mass-spring | Supported on toy systems |
Three directions follow for my own work on action-conditioned world models and manipulation policies. A kinematic self-consistency loss on manipulation action chunks, the PIKC idea applied to predicted poses, velocities and gripper states, is a day of code and a few GPU-days to test on LIBERO at 5 to 10% of the demonstrations, where it should matter if it matters anywhere. Physics-guided proprioceptive features for a VLA (end-effector velocity, a wrench estimated from joint torques) are the cheapest route of all, limited mainly by how realistic simulated torques are. And residual spectra are a cheap diagnostic that world-model papers almost never report: does a latent predictor's error sit on the noise floor or above it, and at which frequencies?
The adaptation result suggests a cheap experiment. In an action-conditioned latent world model for a mobile robot, route the robot's own motion through a model-structured branch: windows of commanded linear and angular velocity in, realised body twist out, scheduled by speed and payload, with FIR responses that expose actuation lag. Keep the scene latent learned. Then measure three axes separately, in Isaac Lab and on the real robot: open-loop rollout error split into ego drift and scene error, response to counterfactual commands, and planning success with latent lookahead. The decisive test is whether fine-tuning only the ego branch on seconds of real driving closes a sim-to-real gap that would otherwise need the whole model retrained.
Cost: about one GPU-week for a 3 × 3 grid (ego branch none, unicycle plus residual, or NCM; 10, 30 or 100% of real data) at three seeds. The objection I expect is that ego motion is a small share of a world model's error. So the first step is to measure that share before building anything.
Combining the three routes instead of using one; physics priors to fill the data gap of embodied foundation models; differentiable end-to-end training for performance and fast adaptation; lifelong online adaptation that is physics-aware, changing only what physics says has changed; provable, verifiable safety, hardest for foundation models; and interpretability for public acceptance. L5 1:31:08
Tap a card for an everyday analogy.
Every topic with the second it starts. Times link straight into the video.
Faithful to the sources: every equation, every table and every number in Sections 00 to 14, each with its slide or paper reference, and the nnodely code, quoted verbatim. Redrawn, not copied: the taxonomy, the spectrum, the longitudinal network, the NCM block diagram and the controller-training loop. Schematic: the four labs. They are exact computations on toy systems chosen to show one mechanism each (a linear-in-parameters stand-in for a PINN, a one-term black-box error and a perturbed Hamiltonian for Lab 2, a first-order lag for Lab 3, a linear plant for Lab 4), not reproductions of the papers' experiments.
Checked on 1 Oct 2026 in headless Chromium at 1440×900 in light and dark mode and at 390×844: no page or console errors and no horizontal scroll; all 20 equation blocks render in Noto Sans Math; every lab control was exercised and the readouts match the notes (Lab 1 extrapolation RMSE 0.742 with data only, 0.056 at γ = 10−3, 0.010 at the default; Lab 2 under leapfrog, exact field −0.01% and black box −68%; Lab 3 triangular activations summing to exactly 1; Lab 4 gains trained on a simulator 40% too light overshooting 3.3% there and 13.7% on the real plant); the quiz, map, code notes, glossary and 144-entry lecture index all respond. Touch input was not tested on a physical phone, and the equations need a browser with MathML (Chrome 109 or later, Firefox, Safari).