Study guide · Physics-informed machine learning · 2026-10-01

Where physics enters a neural network

A five-lecture PhD course from the University of Trento and TUM, rebuilt as one page you can learn from: the map of where a physics prior can enter a model, every method with its equations, four labs that run in the page, the course's own evidence read against its tables, and a timestamp for every topic in the 8 h 38 min of video.

Provenance All shown
00

Corrections before the material

Ten points where the brief, the lectures or the common reading of this field need tightening. Each is settled here once and used in its corrected form below. None of them changes the course's message; several sharpen it.

  1. Five videos, one playlist. The teaching page links a single YouTube playlist: five lectures of 1 h 37 min to 1 h 57 min, 8 h 38 min in all, plus eight PDF decks with 295 slides. The decks do not map one to one onto days (the first architecture deck spans Lectures 2 and 3, the nnodely deck spans 4 and 5), so this page indexes by topic and links each topic to its second in the video. Lecture
  2. "Physics-informed" names one of three routes. The course title uses it as an umbrella; inside the course it means a term in the loss and nothing else. Papers do not follow this: Beaber et al. (RA-L 2024) call a network trained with a PDE residual "physics-guided". Classify a method by where the prior sits, not by its title. Lecture
  3. Dropping the initial state needs a stable system. Lecture 3 derives the input-window representation of a dynamical system and says it holds for any nonlinear system. It holds for systems with fading memory. An integrator (position from pedal input), an undamped oscillator or backlash never forgets its initial state. Every case study in the course predicts a fading-memory output (accelerations, yaw rate, steering angle), so the results stand; the general claim does not. Inference
  4. Lagrangian and Hamiltonian networks conserve a learned energy. An HNN conserves Hθ exactly only in continuous time. A Runge-Kutta or Euler step still drifts, and Hθ is not the true energy. The course's own Variational Integrator Network slides show an HNN losing and gaining energy on small training sets. Paper/Code
  5. L(q)L(q)⊤ alone is only positive semi-definite. The slide says DeLaN learns "a lower-triangular matrix". Positive definiteness needs a strictly positive diagonal, which DeLaN enforces with a non-negative activation on the diagonal plus a small positive constant. Paper/Code
  6. Zero-shot, the structured model was the worst of three on yaw rate. After a tyre change the model-structured network scored yaw-rate RMSE 0.172 and 0.129 rad/s against 0.147 and 0.110 for the general-purpose network and 0.102 and 0.079 for the earlier structured benchmark. It led only after selective fine-tuning on 20 s of data, and on longitudinal acceleration with tyre set 1 the general network still won after fine-tuning (0.380 against 0.395). "Outperforming in basically all cases" holds for the fine-tuned yaw rate. Lecture
  7. "About 50% better" is 39% in RMSE and 61% in FVU, and not the best row. With the small training set the full model's validation yaw-rate RMSE is 0.137 against 0.224 for the general network. An ablated variant of the same model scores 0.124 on that column. Lecture
  8. The physics loss inside StyleVLA buys about 3% ADE, reported without seeds. Adding the kinematic-consistency term moves ADE from 1.21 to 1.17 m and FDE from 3.17 to 3.06 m; the regression head before it moved ADE from 1.47 to 1.21 m. StyleVLA is a preprint (arXiv 2603.09482), and it is a Vision Language Action model, not "Video-Language-Action" as the slide headings say. Lecture
  9. The steering results are offline. MS-NN-steer was trained and validated on recorded A2RL telemetry; its accuracy is the error against the executed steering angle. The lecture and the paper abstract report no closed-loop lap with the learned feedforward on the car. Paper/Code
  10. "YouTube has 35k hours of videos" cannot mean all of YouTube. The figure on the introduction slide is unsourced and orders of magnitude below YouTube's size; it most likely refers to some curated subset. The argument it supports, that internet video carries no action labels, does not depend on it. Inference
01

The course at a glance

The course ran in Trento from 6 to 10 July 2026 as part of Neu4mes, a 1.2 million euro project funded by the Italian Ministry of University and Research under the FIS 2023 call, with Rosati Papini as principal investigator. Its aim is a method family the lecturers call model-structured neural networks, and its outputs include the open-source framework nnodely. Piccinini, now a postdoc at TUM's Autonomous Vehicle Systems lab, prepared most of the slides and gave most of the lectures; Rosati Papini interjects throughout with practitioner's notes that are some of the most useful minutes in the recording. L1 20:11

The arc is deliberate. Day one asks why data alone will not carry physical systems and builds a taxonomy of where a prior can enter; days two and three walk the taxonomy route by route; day four shows the tooling and real racing data; day five trains a controller without new data, looks at foundation models, and lists open problems.

Lecture 21:54:10

Losses, then Lagrangian and Hamiltonian networks

PINNs from Raissi et al., data plus physics regularisation on a quadrotor, monotonicity losses, then DeLaN, LNN and HNN.

Lecture 31:56:33

Neural ODEs, hybrids, operators, model structure

Neural ODEs and variational integrators, residual and sub-system learning, equation learning, Koopman, then the derivation of neural compositional modules.

Lecture 41:46:19

nnodely and vehicle dynamics from little data

The framework's pipeline, the 2019 longitudinal model line by line, spectral window design, and coupled longitudinal-lateral learning on a small-scale racer.

Lecture 51:40:33

Steering control, controller training, foundation models

A model-structured feedforward for a full-scale race car, closed-loop training through a frozen simulator, a PINN in nnodely, StyleVLA, and open problems.

Durations from the playlist "Physics-Informed Machine Learning for Modeling, Planning, Control and Estimation of Physical Systems" (UniTrento Ingegneria Industriale, uploaded 16 Jul 2026). Every ▶ pill on this page opens the lecture at that second.

02

Why data alone will not carry physical systems

The course opens with Ken Goldberg's question from his ICRA 2026 plenary: is robotics about to have its ChatGPT moment? Deep learning with big data solved most of vision, and large language models with internet-scale text are close to solving language. Language, Piccinini argues, is a small one-dimensional world of token strings. A robot hand has 22 degrees of freedom and a humanoid around 50, and a vision-language-action model such as π0 needs synchronised images, text and motor commands for every training example. Reading the corpus behind Qwen-2.5 would take a person about 1.2 billion hours; the matching corpus for physical action does not exist on the internet, because video carries no joint torques. L1 6:21

Four gaps follow: no internet-scale open data, a real world too complex to cover in a training set, the need to generalise to unseen scenarios from little data, and safety and interpretability for public acceptance. The community fills them four ways. Data flywheels deploy fleets to harvest data (Tesla); teleoperation pays humans to drive the robot or rescue it (remote operators behind robotaxis, humanoid start-ups); simulation trades data cost for a sim-to-real gap; and physics priors reuse what engineering already knows. The course takes the fourth route and frames it as a dial between two poles. L1 13:45

Physics model equations, identified parameters 10 to 1,000 parameters Black box deep nets: thousands and up foundation models: billions this course: hybrids enough structure to need little data, enough freedom to learn what the equations miss
Redrawn from Lecture 1 (16:56 to 19:06). Parameter counts are the lecturers' orders of magnitude.

Inductive bias is not a new idea

A multilayer perceptron with one hidden sigmoid layer approximates any continuous function (Cybenko, 1989), so in theory structure is unnecessary. In practice every successful architecture encodes its domain: convolutions encode translation structure in images, LSTM gates keep long-range dependencies in sequences, attention encodes relations between tokens. For physical systems the domain prior is physics: equations of motion (Newton-Euler, Lagrange, Hamilton), conservation of energy and momentum, symmetries and invariances, kinematic and dynamic constraints (holonomic or not), the connectivity of bodies and joints, and problem-specific models of friction or actuators. L1 50:38

One precision for readers coming from vision: a convolution is translation-equivariant; detecting an object anywhere in the frame is the invariance obtained after pooling. The distinction matters for physics too, because a rotation of a robot's base should rotate its predicted forces (equivariance), not leave them unchanged.

The thirty-second demonstration

Rosati Papini's favourite demonstration is the TensorFlow Playground spiral. With the raw coordinates x1, x2 as inputs, a network with several hidden layers and thousands of weights plateaus in a local minimum. Add sin(x1) and sin(x2) as inputs, because a spiral has a trigonometric parametric form, and one hidden layer separates the classes within a few epochs. The right features cut data, parameters and training time at once; that is the whole course in miniature. L1 41:15

Lecture 2 adds a side debate worth knowing: why deep beats wide when one hidden layer suffices in theory. A student cites the exponential growth of expressivity with depth; Rosati Papini adds that sparse, structured connectivity (convolutions, attention, and later the plain "+" that sums independent forces) helps fitting by removing connections that would learn spurious correlations. He points to a published study of sparse connectivity without naming it on the recording, so treat the second claim as a heuristic. L2 19:08

03

The map: three places a prior can enter

Any learned model, whether a neural network, a Gaussian process or a foundation model, has three ingredients: inputs and a dataset, an internal architecture, and a loss that trains it. A physics prior can enter at each, and the course names the routes after the survey of Faroughi et al. (arXiv 2211.07377, published 2024): physics-guided for inputs and data, physics-encoded for the architecture, physics-informed for the loss. Rosati Papini adds two more places: the training procedure and the validation. Click a block. L1 55:51

Interactive map · click or tab to a block
inputs + dataset physics-guided architecture physics-encoded loss ℒ physics-informed training procedure validation dashed: the two extra levels Rosati Papini adds (L1 1:02:10)

Physics-guided inputs, data and representations

Physics shapes what the model sees: input features, the representation of the data, or which training data is selected or generated. The main model can stay generic.

The test: any operation on the inputs is frozen or pre-trained. If a network in the preprocessing trains together with the main model, it has become part of the architecture, and the method is physics-encoded.
  • Physics models as inputs: a Kalman filter or kinematic model computes a physical quantity first.
  • Physics-guided features and training data: v², sign(v), energies; simulated or physically corrected data.
  • Geometric learning: symmetries, SO(3) and SE(3) representations, equivariance.
  • Frequency-domain learning: Fourier or DCT inputs that keep the physical bandwidth and drop noise.

Section 04 · L1 1:10:33

Redrawn from Introduction slides 31 to 34. Colours follow the course: green inputs, orange architecture, blue loss.

The routes compose. A network can take physics-guided inputs, have a physics-encoded body and train on a physics-informed loss, and the lecturers' closing message is that almost every published method uses one route while the open potential is in combining them. L1 1:08:29

One difference from the source survey: Faroughi et al. treat neural operators as a fourth family next to guided, informed and encoded networks. The course folds operators into the encoded route, which is defensible when the physics enters through the operator's lifting functions.

Classify it yourself

Nine short descriptions, several of them borderline on purpose. Pick a route; the explanation appears with the answer.

0 of 9 answered

04

route 1 Physics-guided inputs, data and representations

The smallest chapter, and by the lecturers' account the least explored: physics decides what the model receives, and the model itself can be any general-purpose network. Four classes, of which the course works through the first two. L1 1:11:37

1 · Physics models as inputs

Compute a physical quantity first.

Raw signals pass through a physics model or state estimator (kinematics, an EKF or UKF), and its estimate joins the inputs of a neural network. Learning happens in a physically consistent space.

2 · Features and training data

Choose what to feed, and what to train on.

Derived features such as v², sign(v), energies or forces; training data generated by validated simulators or post-processed until it is physically consistent.

3 · Geometric learning

Respect the structure of the space.

Project or augment data so it satisfies known symmetries (a car is left-right symmetric), use SO(3) and SE(3) representations and equivariant models.

4 · Frequency-domain learning

Learn the bandwidth that is physics.

Transform signals with Fourier or DCT and keep the band a mechanical system actually occupies (often below 10 to 15 Hz), so the model cannot fit vibration and sensor noise.

Three worked papers

Sideslip estimation (Gräber et al., IEEE T-IV 2019). The sideslip angle β between a car's heading and its velocity is hard to measure without expensive GPS. A physical model computes its rate β̇ from lateral acceleration, speed and yaw rate; β̇ enters a GRU alongside the raw signals, and the GRU outputs β. The hybrid beat both the same GRU without β̇ and the physical model alone. L1 1:20:03

The relation is a one-liner: for small angles, β̇ ≈ ay/v − r, with r the yaw rate. Integrating it drifts, which is why it works better as a feature than as an estimator.

Feedforward control of a linear motor (Bolderman, Lazar and Butler, IEEE CCTA 2021). A network maps the reference trajectory to the motor force. Friction depends on velocity and its sign, so ẏ and sign(ẏ) join the inputs; the same network with these two features outperforms the baseline. It is the Playground lesson on real hardware. L1 1:24:15

Physically consistent training data (PARC, Xu et al., SIGGRAPH 2025). A diffusion model generates parkour motions for simulated characters in Isaac Gym; a physics-based tracking controller replays them and outputs what is physically feasible; the corrected motions return to the dataset that retrains the generator. The loop improves on the generator alone. L1 1:27:25

The discussion at the end of Lecture 1 places a current topic in this route: reconstructing actions or forces from internet video, often by training inverse-dynamics models in simulation and applying them to real footage. The lecturers call physics-guided learning the most emergent of the three routes for exactly this reason. L1 1:33:49

Cost and failure mode. This route is the cheapest to try, since it changes no training code, and the easiest to debug. Its ceiling is the frozen preprocessing: an estimator that is wrong in a regime feeds the network a confident wrong feature, and no gradient will ever correct it. When the feature is only roughly right, letting the preprocessing train turns the method into a physics-encoded one, with the extra risk that it drifts away from its physical meaning.

05

route 3 Physics-informed losses

The route most people mean by the phrase. Governing equations, boundary conditions or other physical laws are written as residuals that should be zero and added to the loss as squared penalties, which keeps the objective smooth for gradient descent. Two motivations run through the chapter: learn the solution of an equation that is expensive or impossible to solve directly, and combine an imperfect physical model with data so the network generalises where data is missing. L2 1:13

The seminal formulation

Raissi, Perdikaris and Karniadakis (J. Comput. Phys. 2019) start from a general nonlinear PDE for an unknown u(t, x) with parameters λ inside a nonlinear operator 𝒩. Two questions follow: given λ, can a network learn u (the forward problem)? Given data, which λ fits it (discovery, the inverse problem)? For the first, call the left side f and train a network for u with a loss in two parts. L2 5:25

ut+𝒩[u;λ]=0f≔ut+𝒩[u] MSE=1Nu∑i=1Nu|u(tui,xui)−ui|2+1Nf∑i=1Nf|f(tfi,xfi)|2 First sum: the known initial and boundary values. Second sum: the PDE residual at Nf collocation points sampled inside the domain. Physics-informed loss slides 4 and 5.

Three details carry the method. The collocation points cost nothing to add, since the residual needs no measurement, so one can sample hundreds of thousands of them. The residual contains derivatives of the network with respect to its own inputs, ∂u/∂t and ∂²u/∂x², computed by automatic differentiation; Rosati Papini singles this out as the genuinely new step for the neural-network community, using the network's gradient as an output that enters the loss. And the network itself is a generic multilayer perceptron: no physics in the architecture. L2 27:36

The lecture is candid about the limits. Nothing guarantees convergence to the true solution; if the PDE is well posed and the network deep enough it works in practice. The weights between terms are a manual trade-off between honouring boundaries and honouring the equation. Inference is a forward pass instead of hours of numerical solving, but the trained network solves one problem: change the initial or boundary conditions and it must be retrained. L2 12:49 L2 33:59

Two well-documented failure modes are worth knowing before trying this at scale: imbalanced gradients between the loss terms (Wang, Teng and Perdikaris, SIAM J. Sci. Comput. 2021) and training failures on PDEs with strong convection or high-frequency solutions (Krishnapriyan et al., NeurIPS 2021). Learning a map from boundary conditions to solutions instead of one solution is the job of neural operators such as DeepONet (Lu et al., Nature Machine Intelligence 2021).

Worked examples

Nonlinear Schrödinger equation (Raissi et al.). A complex solution h = u + iv on x ∈ [−5, 5], t ∈ [0, π/2], known only at t = 0 as h(0, x) = 2 sech(x), with periodic boundaries in value and in ∂h/∂x. A five-layer network trained on the initial condition, the periodic boundaries and the residual reproduces the numerical solution without a single interior measurement. L2 22:20

Pneumatic soft finger (Beaber, Liu and Sun, RA-L 2024). A network maps a point (x, z) and the air pressure p to the deformations ux, uz; the loss holds the initial condition and the two Navier-Cauchy continuum equations with body forces from gravity. Benchmarked against ANSYS finite elements, the trained network predicts the finger's shape in real time across pressures. The paper calls itself physics-guided, which is the taxonomy confusion of correction 2 in one title. L2 35:01

Beyond PDEs: data where you have it, physics where you do not

An ODE is a PDE in one variable, so the recipe applies to any robot. Bianchi et al. (Drones 2024) learn quadrotor dynamics with a data loss on flight logs plus a physics loss on the rigid-body equations at collocation times, with γ setting how much the physics is trusted; it beat an extended Kalman filter at estimating the parameters. The lecturers then make the point they consider most useful for real systems: real data usually covers a small region of the operating domain, while an imperfect model is available everywhere. Apply the data loss where data exists and the physics loss where it does not, and the network fits the measurements while extrapolating physically. ℒphysics needs no ground truth. L2 49:49

ℒ=1Ny^∑i=1Ny^1Nt∑j=1Nt|y^i(tj)−y^i,refj|2+γ1Ny^∑i=1Ny^1Nl∑k=1Nl|l(y^i(tk))|2 Data loss on Nt logged samples, physics loss on Nl collocation times; l is the ODE residual and γ the confidence in the model. Physics-informed loss slide 18.
Lab 1

Data where you have it, physics where you do not

true solution fitted model noisy data data window collocation points
RMSE inside data window-
RMSE extrapolation-
Physics residual (RMS)-
γ-

What is computed. The system is a damped oscillator x″ + 2ζωx′ + ω²x = 0 with ω = 2 rad/s, ζ = 0.08, x(0) = 1, observed 16 times with noise inside the green window. The model is a sum of 48 Gaussian bumps; its weights enter linearly, so the data loss and the ODE residual at 120 collocation points are both quadratic and the optimum is solved exactly with a Cholesky factorisation at every slider move. A real PINN uses a multilayer perceptron trained by gradient descent: same loss, plus optimisation error. Try three things: γ at its minimum (pure data) and watch the fit collapse to zero outside the window; any γ above about 10−3; then the model's ω at 1.8 with γ = 100, where trusting a wrong model bends even the fit inside the data.

Losses that are not equations

Gu, Primatesta and Rizzo (RAS 2024) learn a quadrotor's inverse dynamics, states to PWM motor commands, and add a loss that encodes conservation of momentum indirectly: the angular acceleration about each axis should move with the PWM signal that produces it, which they verify from data with clear correlations for roll, pitch and yaw. The local-monotonicity terms are soft constraints next to the data loss. L2 55:08

ℒ=ℒMSE+∑i∈{ϕ,θ,ψ}λLMiℒLMi Data loss plus one local-monotonicity term per rotation axis. Physics-informed loss slide 26.

The assumption fails under strong aerodynamic effects, where command and acceleration decouple. Their answer is a cyclical annealing schedule that periodically drives λLM to zero after about 40 epochs, letting the network learn the aerodynamics the prior contradicts. Rosati Papini files this under physics in the training procedure: the prior is used, then deliberately switched off. L2 58:21

Cost and failure mode. No change to the architecture and no extra data, which is why this is the popular route. The price is tuning: each residual has its own units and scale, γ spans orders of magnitude (the lab's slider covers nine), and a wrong model trusted too much biases the fit everywhere. The constraint is soft, so nothing is guaranteed at deployment.

06

route 2a Lagrangian, Hamiltonian and Neural ODE networks

The architecture route is the largest, and the course splits it into six classes. The first two learn dynamics through the mathematics of motion itself: learn the scalar functions behind the equations of motion, or learn the vector field and integrate it with structure. L2 1:04:45

From Newton to Lagrange

Newton's formulation needs every force and moment on a free-body diagram, which becomes a nightmare for a humanoid. Lagrange's is energy-based: write kinetic minus potential energy in generalised coordinates q (joint angles, for instance), apply the Euler-Lagrange equation, and the familiar manipulator equation falls out, including the Coriolis and centripetal term c, which needs no separate modelling. L2 1:09:00

ℒ=T−V=12𝐪˙⊤𝐌(𝐪)𝐪˙−V(𝐪)ddt∂ℒ∂𝐪˙−∂ℒ∂𝐪=𝛕 𝐌(𝐪)𝐪¨+𝐜(𝐪,𝐪˙)+𝐠(𝐪)=𝛕𝐜=𝐌˙𝐪˙−12∂∂𝐪(𝐪˙⊤𝐌(𝐪)𝐪˙) M: symmetric positive definite inertia matrix; g = ∂V/∂q: gravity; τ: non-conservative generalised forces. Encoded I, slides 8 to 10.

Classical engineering measures or estimates M and g from masses, lengths and inertias, and inertias are notoriously hard to measure. Deep learning ignores the structure and learns everything, needing far more data. Rosati Papini's estimate for a well-modelled robot is that physics explains 90 to 95% of the output; relearning that part from data is waste. L2 1:12:14

Deep Lagrangian Networks (DeLaN)

Keep the equation, learn M and g.

Lutter, Ritter and Peters (ICLR 2019). Networks output M(q) and V(q) or g(q); c comes for free by automatic differentiation of M. M is parameterised as L(q)L(q)⊤ with L lower-triangular, so symmetry is structural. Torques from the equation are matched to measured torques with a plain MSE, so the loss is ordinary and the physics sits in the architecture. On a 7-DoF arm, a DeLaN feedforward plus PD feedback tracked better and generalised better than a generic feedforward network. L2 1:14:26

Lagrangian Neural Networks (LNN)

Learn the whole Lagrangian.

Cranmer et al. (ICLR workshop 2020). DeLaN's form T = ½q̇⊤Mq̇ holds for rigid bodies but not for a charged particle. An LNN learns ℒ(q, q̇) with a generic network and gets accelerations from the Euler-Lagrange equation, below. On a frictionless double pendulum, a baseline network integrated for hundreds of seconds loses energy and falls out of phase; the LNN keeps it. L2 1:30:18

Hamiltonian Neural Networks (HNN)

Learn the energy, take its symplectic gradient.

Greydanus, Dzamba and Yosinski (NeurIPS 2019). The Hamiltonian ℋ(q, p), usually total energy T + V, is learned by a network; time derivatives are its partial derivatives with a sign swap. Momenta p are the catch: easy for a mass on a spring, hard to define for complex systems. Coupled with an autoencoder, an HNN learns a pendulum's dynamics from pixels by treating latents as (q, p). L2 1:45:13

𝐪¨=(∇𝐪˙∇𝐪˙⊤ℒ)−1[∇𝐪ℒ−(∇𝐪∇𝐪˙⊤ℒ)𝐪˙]loss‖𝐪¨−𝐪¨^‖2 𝐪˙=∂ℋ∂𝐩,𝐩˙=−∂ℋ∂𝐪ℒHNN=‖∂ℋθ∂𝐩−𝐪˙‖2+‖∂ℋθ∂𝐪+𝐩˙‖2 Top: LNN accelerations by the chain rule on the Euler-Lagrange equation. Bottom: Hamilton's equations in the unforced case and the HNN loss, with q̇ and ṗ from finite differences of the data. Encoded I, slides 22 and 31.

A recurring exchange in Lectures 2 and 3 is worth keeping. A network trained on one-step accelerations can either gain or lose energy when integrated. Rosati Papini's experience is that networks trained on integrated multi-step losses tend to lose it, because a model that gains energy makes the loss explode later in the rollout, which gradient descent avoids. L2 1:38:46

Neural ODEs and variational integrators

Neural ODEs (Chen et al., NeurIPS 2018) replace the right side of dh/dt = f(h, t, θ) with a network and backpropagate through any ODE solver by the adjoint sensitivity method, with memory linear in problem size and error control. Its discrete-time version, ht+1 = ht + f(ht, θt), is a residual network. Rosati Papini notes the practical difference: a residual network learns a fixed time step, while the continuous form can be sampled anywhere, and calls the Neural ODE the most flexible tool of the chapter because any function can be plugged in. L3 3:14 L3 21:06

Euler and Runge-Kutta steps do not preserve the geometry of mechanical systems, so even a perfect vector field drifts in energy. Variational Integrator Networks (Saemundsson et al., AISTATS 2020) integrate with symplectic variational integrators that conserve energy and momentum to third order or higher. With noisy observations and a small training set, an HNN overfit and drifted while the VIN held energy on a mass-spring system and a pendulum; with a medium set the gap closed. L3 10:35

Lab 2

Where energy drift comes from: the field or the integrator

exact field black-box field HNN field, true energy HNN field, its own ℋθ
Exact field, energy error-
Black box, energy error-
HNN, true energy error-
HNN, ℋθ error-
HNN phase error at end-

What is computed. A frictionless pendulum, ℋ = p²/2 + 1 − cos q. Three vector fields: the exact one; a black-box stand-in, ṗ = −sin q + εp, whose small error is not the gradient of any energy (the typical error of an unconstrained network, here shrunk to one term); and an HNN stand-in, the exact symplectic gradient of a slightly wrong ℋθ = p²/2 + (1 + δ)(1 − cos q). Each is integrated with the chosen scheme and compared with a fine Runge-Kutta reference. Read it in three steps: under Euler at the default step everything gains energy, the exact field included, so the integrator alone causes drift; under leapfrog the exact and HNN fields stay bounded for any horizon while the black box still drifts, so structure must sit in the field too; and the HNN conserves its own ℋθ, not the true energy, so its phase error against the exact pendulum grows steadily.

Choosing among them

Cranmer et al.'s comparison, as the course presents it, reduces the choice to four questions: are physics priors available, do you want a differential equation, does energy conservation matter, and is the structure of the Lagrangian known? L3 17:54

PropertyNeural netNeural ODEHNNDeLaNLNN
Models a dynamical system✓✓✓✓✓
Learns a differential equation·✓✓✓✓
Learns exact conservation laws··✓✓✓
Learns from arbitrary coordinates✓✓·✓✓
Learns arbitrary Lagrangians····✓

From Cranmer et al., "Lagrangian Neural Networks" (arXiv 2003.04630), as shown on Encoded I slide 51. The lecture adds that the table is written by LNN's authors.

Where this meets robot learning. These networks assume conservative or rigid-body structure and smooth coordinates. Manipulation is dominated by contact, friction and switching, where energy is not conserved and generalised coordinates of the scene are unknown, and a VLA's action space is usually an end-effector command behind a low-level controller that already hides the arm's rigid-body dynamics. The HNN-on-pixels result is a small latent world model with a conservation prior; the open question for world models is which part of a scene has dynamics clean enough for such a prior. The robot's own body and base are the obvious candidates.

07

route 2b Hybrids, discovered equations and operators

Hybrid physics-neural models

Networks stay generic black boxes but become modules inside a physical model: they correct it, replace a part that is hard to model, or preprocess sensors for it. Historically this was among the first ways of merging data and physics, and it remains the most common in robotics. L3 23:13

Residual learning

Learn what the model gets wrong.

Kabzan, Hewing, Liniger and Zeilinger (RA-L 2019) add Gaussian-process residuals to an eight-state single-track vehicle model, xk+1 = f(xk, uk) + Bd(d(zk) + wk), and learn them online from a dictionary of driving data inside an MPC. The lecture reports a 10% lap-time improvement over the physics-only controller. L3 27:30

Sub-system learning

Replace the part that is expensive to identify.

Wegrzynowski et al. (IROS 2024) keep a four-state vehicle model but let a network output the four tyre forces, which otherwise need a tyre test rig. The model runs inside an unscented Kalman filter whose noise covariances are trained end to end with it; it beat GRU and LSTM estimators and transferred zero-shot to unseen road conditions. L3 31:47

Sensor processing

Let a network turn pixels into physics.

A convolutional network converts camera images into quantities a physical model consumes. When that network is frozen it is the physics-guided route; when it trains with the model it is encoded. The course mentions this class without a worked example.

Rosati Papini adds a caution against using LSTMs and GRUs as general dynamics models. They were designed for text, where an opened bracket must close hundreds of tokens later, so their gates can hold memory indefinitely. In a physical system the effect of a pedal press fades, and a gated memory is the wrong inductive bias; Piccinini extends the point to transformers and diffusion models applied to vehicle dynamics, which in their experience overfit. L3 36:02

This is the fading-memory property of correction 3 seen from the other side. A recurrent network with weight decay and early stopping can still fit fading dynamics; the argument is about what the architecture makes easy, which matters most when data is short.

Topology learning: let the network choose the equation

Equation-learner networks replace activation functions with a dictionary of candidates (sin, cos, products, identity) and learn sparse weights that switch candidates on or off, so the trained network reads as a formula. Rosati Papini describes the training trick from the original equation learner: after some hundred epochs, inspect the weights, delete unused units and continue, so the network concentrates on the active terms. SINDy is the regression version of the same idea. Physics enters through the dictionary: if a manipulator's dynamics are trigonometric, trigonometric functions go in. L3 40:19

𝐗˙≈𝚯(𝐗)𝚵,𝚵 sparse𝛗(𝐱k+1)≈𝐊𝛗(𝐱k) Left: SINDy (Brunton, Proctor and Kutz, PNAS 2016), a library Θ of candidate functions and sparse coefficients Ξ. Right: a Koopman lift φ in which the dynamics become linear with matrix K.

Neural operators

Neural networks map fixed-size vectors to vectors; neural operators map functions to functions and come with their own approximation theorems. The course names Koopman networks (the most popular, and in the lecturers' words very hyped at conferences), Kolmogorov-Arnold networks, DeepONets and graph neural operators. A Koopman network learns lifting functions that embed a nonlinear system in a space where it evolves linearly, plus a decoder back; once linear, the whole toolbox of linear control and linear MPC applies. Physics enters through the lifting functions, and the lecturers judge physics-informed Koopman work to be at an early stage. L3 45:44

A caveat the course leaves implicit: exact finite-dimensional linear lifts exist only for special systems, so in practice the lift is approximate and its linear prediction degrades with horizon. For control-affine systems the lifted model is usually bilinear in the input rather than linear.

Encoded classThe network learnsPhysics fixesPays off whenWatch for
Lagrangian, HamiltonianM and V, or the whole ℒ or ℋEuler-Lagrange or Hamilton's equationsconservative, rigid-body or energy-based systems; long rolloutsneeds coordinates, and momenta for HNN; dissipation and contact need extensions
Neural ODE, VINthe vector field fcontinuous time; symplectic stepping for VINirregular sampling; a differential model is wantedno conservation by itself; adjoint cost; stiffness
Hybrida residual, a sub-system or a sensor mapthe rest of the physical modela good model with known weak spots existsthe residual can absorb modelling errors and extrapolate badly
Topology learningwhich candidate terms and connectionsthe dictionaryyou want an equation; low-dimensional systemsthe dictionary must contain the answer; noisy derivatives
Neural operatorsmaps between function spaces, Koopman liftslifting functions, linear lifted dynamicsfamilies of PDE solutions; linear control of nonlinear plantsapproximate lifts; few physics-encoded variants so far
Model-structuredinput maps, regime weights, FIR responses, a residuallayers, connections and bounds from the equationsdynamics from little data; interpretability; selective adaptationdesign effort per system; derivation assumes fading memory

My synthesis of Sections 06 to 08; the class names and examples are the course's.

08

route 2c Model-structured neural networks

The lecturers' own class, which they named in papers from 2025 and 2026 to cover approaches that fit none of the existing categories. Everything lives in the architecture, under three principles, and one practical guideline: start from an imperfect physical model of the system and neuralise its architecture. L3 51:08

vx,k [pk−n, …, pk] ( · )² linear: kd, bias Fdrag min(p, 0) max(p, 0) FIR layer FIR per gear Fx,rear (brakes only) Fx,front (traction) + ax,k physics layers physics connection: forces superpose learned weights
Redrawn from Encoded III slides 7 to 9 and Lecture 4: a front-wheel-drive car, braking on the rear axle. 1/m is absorbed into the learned weights.

Physics layers. Each branch learns one named force: aerodynamic drag from v², rear braking force from the negative part of the pedal only (the rear axle of a front-wheel-drive car never pulls), front traction from the positive part. Physics connections. Newton says forces superpose, so the branches are summed, not mixed by another layer. Rosati Papini's comment on that plus sign is one of the best moments of the course: a mixing layer would add connections and let the network learn a correlation between independent forces that does not exist; the plus sign is sparsity chosen by physics. Physics constraints. Parameters that have a physical meaning get bounds through the architecture, never through a loss penalty. L3 59:37

θ^=Φ_+(Φ‾−Φ_)σ(z) An unbounded learnable z mapped into physical bounds, for example a payload mass that must stay between the empty and the loaded truck. Encoded III slide 9, after Chrosniak et al. 2024.

Neural compositional modules, derived in ten steps

The general recipe composes neural compositional modules (NCMs) Ψi, each a function of a window of past inputs and of scheduling variables, with sums and products. The derivation is the heart of the course and repays slow reading. L3 1:04:58

1 · Start from any nonlinear state-space model

Substitute the state equation into the output equation recursively, n times.

xk+1=f(xk,uk),yk=h(xk,uk)⇒yk=h(f(xk−1,uk−1),uk)=⋯=𝒢(uk,…,uk−n,xk−n)

2 · Forget the initial state

If the window is long enough, the free response from xk−n has died out and the forced response dominates, so yk ≈ 𝒢(uk, …, uk−n). Piccinini's image: after holding the accelerator for a while, the speed you started from no longer matters. L3 1:09:16

This is the step that needs fading memory (correction 3). Choose outputs that forget, such as accelerations, rates and forces, and integrate them outside the network.

3 · Let the map depend on the operating regime

Dynamics change with speed or gear, so add scheduling variables: yk = Ψ(uk, …, uk−n, 𝐳k). This is a full NCM's signature: an input history, a regime, one output.

4 · If the system is linear, Ψ is one fully connected layer

The output is the convolution of the input with the discrete impulse response Γ. Truncating the infinite sum gives a finite impulse response, which is exactly a linear layer g whose weights are the impulse response coefficients. L3 1:12:29

yk=∑i=0∞Γiuk−i≈∑i=0nΓiuk−i=g(𝐮k),𝐮k=[uk−n,…,uk]

5 · If it is locally linear, blend FIR layers by regime

Give each of M regimes its own FIR layer 𝐠j and an activation 𝛗j(𝐳) that says how much regime j applies. The activations must be non-negative and sum to one everywhere, so no branch is favoured by construction; triangular for continuous scheduling (speed), rectangular for discrete (gears), Gaussians are also common. ⊙ is the element-wise product over the time window. L3 1:16:46

yk=∑j=1M[𝐮k⊙𝛗j(𝐳k)]·𝐠j,∑j=1Mϕj(z)=1,ϕj(z)≥0

6 · Add input nonlinearities where physics knows them

A function 𝐟j transforms the input window before scheduling. With strong priors it is a physical map (v² for drag, v·ω for lateral acceleration); with partial priors an equation learner over a dictionary (trigonometric for a manipulator); with none, a learnable Taylor polynomial a0 + a1u + a2u², or a small general-purpose network. L3 1:30:29

7 · Catch everything else with a residual

For a fully nonlinear system, add a general-purpose residual network ℛ in parallel. This is the complete module.

yk=∑j=1M[𝐟j(𝐮k)⊙𝛗j(𝐳k)]·𝐠j+ℛ(𝐮k) The neural compositional module Ψ. Encoded III slide 35.
NCM Ψ 𝐳k 𝐮k = [uk−n…uk] φ1(𝐳) … φM(𝐳) 𝐟1(𝐮) 𝐟M(𝐮) ⋮ ⊙ ⊙ 𝐠1 (FIR) 𝐠M (FIR) ℛ (residual) + yk input maps: physics, dictionary, Taylor or MLP time window → scalar
Redrawn from Encoded III slides 28 and 35. Dashed: regime weights from the scheduling variables. Each FIR layer compresses the time window to one number.

8 · Many inputs, many scheduling variables

p inputs stack into a matrix of windows. With q scheduling variables (engine speed and gear, say), regime j's activation is the element-wise product of its one-dimensional activations. Products of partitions of unity are again a partition of unity, so two triangular axes give pyramids and three give the hyper-pyramids used later for steering. L3 1:37:51

𝛗j(𝐳k)=𝛗j(𝐳1,k)⊙⋯⊙𝛗j(𝐳q,k)yk=ℱ(Ψ1,…,ΨQ),ℱ∈𝒞({+,×})

9 · Compose modules the way the equations compose

Modules combine by sums (independent forces and moments, each learned by its own Ψ), products (a lateral force times the sign of the steering angle) and cascades (a steady-state module feeding a transient one). The claim is that sums, products and cascades of NCMs are flexible enough for the dynamics of physical systems. L3 1:44:22

10 · Know the variants

The order of f and g can be swapped: compress time first, then schedule. The two are not equivalent; Rosati Papini uses the swap when the regime is discrete and memory across a switch is meaningless, as with gears, where the new gear filters the engine instantly. Feeding past outputs back, through the residual for instance, turns the module autoregressive, which the lecturers use mostly for control. L3 1:49:39 L3 1:53:58

Design decisions the derivation leaves to you

Trainable or frozen activations? Trainable centres let the network place regimes where the dynamics change, but with an imbalanced dataset they migrate to the dense region and abandon the rare one. Rosati Papini's advice: if you know the data is imbalanced, freeze the activations so every part of the range keeps a model, even one fitted on six points. L3 1:22:00

The pros the lecturers list: trained f and g weights can be read physically, task complexity splits across modules, and sample efficiency and generalisation beat black boxes. The con: design effort in the f modules, and in Rosati Papini's words there is no free lunch, because you have to know the system. L3 1:47:33

Read as system identification, the module is familiar. With M = 1 and f the identity it is an FIR model; with a static f it is a Hammerstein model with FIR dynamics; with M > 1 and partition-of-unity weights it is a local-model network in the Takagi-Sugeno and LOLIMOT tradition, that is, a linear-parameter-varying FIR model; swapping f and g gives a Wiener-type structure. A student draws the parallel to interacting multiple models on the recording and the lecturer agrees. What the course adds is the grammar: modules that compose by sums, products and cascades mirroring the equations of motion, trained jointly with residual networks by gradient descent, with tooling that handles windows, closed loops and export. That grammar, not the block, is the contribution to judge. L3 1:29:25

Lab 3

Inside a compositional module: regimes and impulse responses

Sum of activations, min to max-
Regimes active at probe-
Weights at probe-

The scheduling variable runs from 0 to 60 m/s. Pick the non-normalised Gaussian to see why the sum-to-one rule exists.

Window span-
Impulse response captured-
Peak weight at, read as dead time-
Nyquist limit-
Spectral rule suggests n-

What is computed. Top: the activation functions φj over a scheduling speed and their sum, with the weights each regime gets at the probe; triangles give at most two active local models at any speed. Bottom: the weights an FIR layer converges to when the true system is a first-order lag with time constant τ and a dead time, sampled at Δt over a window of n + 1 samples, drawn oldest to newest as on the course's slide. A peak away from the newest sample is the dead time, the reading Rosati Papini demonstrates in Lecture 4. The spectral rule is the lecture's: if the useful content ends near f, keep 3 to 5 periods of f in the window. With the defaults (3 Hz, Δt = 0.05 s) it suggests 20 to 33 samples; the course chose n = 25, 1.25 s.

09

nnodely: the method as a library

nnodely (read "modely": the double n stands for m) is an MIT-licensed Python framework on top of PyTorch, written by Rosati Papini's group to build, train and deploy model-structured and other physics-embedded networks. It installs with pip install nnodely; the stable release on PyPI is 1.5.4, with 1.5.5 in pre-release, and the main branch already carries a NeuralODE layer with Euler and Runge-Kutta solvers that the lecture announces for the next release. Worked applications live in a second repository, nnodely-applications.

Rosati Papini's motivation is the audience: engineers who know their physics but find raw PyTorch an obstacle, and who are better served by blocks that already mean something physical than by asking a language model for code. The pipeline has six phases: define the neural model; build the dataset from raw CSV files, with every time window derived from the model definition; train, with layer-specific helpers such as exponential initialisation of FIR weights; validate, with RMSE and the Akaike criterion; compose validated sub-models (a dynamics model and a controller) into a larger network and train again; export to native PyTorch or ONNX for C++ deployment. L4 7:28

The 2019 longitudinal model, line by line

Click a line. The code is the repository's model_longit_vehicle_dynamics.py, quoted verbatim apart from omitted plotting and file-path lines.

# Dimensions of the layersn  = 25na = 21 #Create neural model inputsvelocity = Input('vel')brake = Input('brk')gear = Input('gear')torque = Input('trq')altitude = Input('alt',dimensions=na)acc = Input('acc') # Create neural network relationsair_drag_force = Linear(b=True)(velocity.last()**2)breaking_force = -Relu(Fir(W_init = 'init_negexp', W_init_params={'size_index':0, 'first_value':0.002, 'lambda':3})(brake.sw(n)))gravity_force = Linear(W_init='init_constant', W_init_params={'value':0}, dropout=0.1, W='gravity')(altitude.last())fuzzi_gear = Fuzzify(6, range=[2,7], functions='Rectangular')(gear.last())local_model = LocalModel(input_function=lambda: Fir(W_init = 'init_negexp', W_init_params={'size_index':0, 'first_value':0.002, 'lambda':3}))engine_force = local_model(torque.sw(n), fuzzi_gear) # Create neural network outputout = Output('accelleration', air_drag_force+breaking_force+gravity_force+engine_force) vehicle.addModel('acc',[out])vehicle.addMinimize('acc_error', acc.last(), out, loss_function='rmse')vehicle.neuralizeModel(0.05) data_struct = ['vel','trq','brk','gear','alt','acc']vehicle.loadData(name='trainingset', source=data_folder, format=data_struct, skiplines=1) optimizer_params = [{'params':'gravity','weight_decay': 0.1}]training_params = {'num_of_epochs':150, 'val_batch_size':128, 'train_batch_size':128, 'lr':0.00003}vehicle.saveModel()vehicle.exportONNX()

Click a line

Each block of the file maps onto a piece of the NCM formalism from Section 08. The annotation names the piece and the physics behind it.

The vocabulary you need

  • Input('x'), Output, Parameter, Constant: named signals and learnable or fixed scalars. Dataset columns are matched by name.
  • x.last(): the current sample. x.sw(n): a window of n past samples; x.sw([0, 10]): the present and 10 future samples, for feedforward control. x.tw(T): a window in seconds.
  • Fir: linear layer over the time dimension, an impulse response. Linear: linear layer over a feature or space dimension.
  • Fuzzify: triangular or rectangular activation functions by number of centres, range or explicit centres. LocalModel(input_function, output_function): a full NCM.
  • ParamFun: any PyTorch expression with learnable parameters (saturations, local hyperplanes). EquationLearner: dictionary-based layers.
  • Differentiate, Integrate: time and partial derivatives, for PINN residuals and Sobolev training. NeuralODE on the main branch.
  • addModel, addMinimize: register sub-models and loss terms. neuralizeModel(dt): fix the sample time and build the graph.
  • closedLoop, connect, addClosedLoop, addConnect: wire outputs back to inputs, within a model or across models.
  • loadData, filterData, resamplingData; trainModel with models= to train only some sub-models and prediction_samples for multi-step rollouts; trainAndAnalyze.
  • saveModel, exportPythonModel, exportONNX.

Names checked against nnodely at commit 777cf84 on 1 Oct 2026.

10

Three case studies from autonomous racing

Racing is the lecturers' testbed: closed circuits make emergency-level driving safe, the dynamics turn strongly nonlinear near the handling limits where physics models struggle, and overtaking adds adversarial decisions. All three studies sit inside a hierarchical stack of planning, feedforward and feedback control, and state estimation. L4 13:46

1 · Longitudinal dynamics from a pedal and a speed (Da Lio, Bortoluzzi and Rosati Papini, VSD 2019)

The naive model is a multilayer perceptron from pedal and speed to acceleration. The structured one rewrites Newton's longitudinal balance as a sum of NCMs: front force from a window of pedal values with one FIR per gear, rear braking force, drag from vx², and a lateral-force term times the sign of the road-wheel angle, itself estimated from the steering-wheel angle by a small network. On unseen data the structured network tracks the measured acceleration more closely than the unstructured one. L4 16:52

ax=1m∑iFi=1m[Fxf(p,g)+Fxr(p)−kdvx2−Fyfsin(δ)] Each term becomes one NCM, Ψ1 to Ψ5, joined by the sums and the one product the equation contains. nnodely deck slide 24.

The validation step is the part to copy. The power spectral density of the residuals is lower for the structured model across the band, and it stays above the noise floor, which was estimated from constant-speed driving; a residual below the noise floor would mean the model fits noise. The signal itself meets the noise floor at 2 to 3 Hz, so nothing above that is vehicle dynamics. That fixes the window: three to five periods of the slowest relevant content, about 1 to 1.5 s, so n = m = 25 samples at 0.05 s, 1.25 s. Sampling at 20 Hz caps the analysable band at 10 Hz. Longer windows are possible but invite the network to learn long-range correlations that are not physics. L4 35:42 L4 39:52

2 · Coupled longitudinal and lateral dynamics near the limit (Mungiello et al., IEEE OJ-ITS 2026)

A small-scale Roboracer car driven fast on TUM's indoor track. Two structured networks feed each other: the lateral one predicts yaw rate Ω, the longitudinal one acceleration ax, and each consumes the other's prediction internally, from windows of speed, steering angle, motor current and road slope. The lateral network starts from an empirical fact rather than Newton: curvature against steering angle is almost a line whose slope changes with speed and acceleration. A quasi-steady-state NCM learns that surface with local hyperplanes under three-dimensional triangular activations; a cascaded transient NCM with FIR outputs, scheduled by speed and acceleration in 2 × 2 regimes, learns the dynamics. The trained FIR weights rise towards the present with no dead time beyond one sample, and their mild oscillation points to higher-order dynamics. L4 1:11:16 L4 1:20:44

Yaw rate, validationLarge set RMSEMedium set RMSESmall set RMSESmall set FVU
G-NN, general purpose0.1420.1720.2240.155
MS-NN-bench, earlier structured model0.1540.1580.1810.066
MS-NN, ablation 10.1220.1100.1340.053
MS-NN, ablation 20.1160.1160.1240.041
MS-NN-full0.1090.1060.1370.061

RMSE in rad/s; FVU is the fraction of variance unexplained, 1 − R². Lowest per column highlighted. nnodely deck slide 38; seeds and intervals are not shown.

Then the adaptation test, the result I find most useful in the course. Swap tyres or add 10% mass, collect 20 s of driving, and fine-tune only the sub-module physics says has changed: the transient module for tyres. Fifty to sixty epochs suffice, fast enough to run online. L4 1:28:06

Yaw rate RMSE, zero-shot → fine-tunedTyre set 1Tyre set 2+10% mass
G-NN0.147 → 0.1070.110 → 0.0700.088 → 0.046
MS-NN-bench0.102 → 0.0960.079 → 0.0590.072 → 0.053
MS-NN-full0.172 → 0.0850.129 → 0.0530.075 → 0.045
Longitudinal acceleration RMSE, zero-shot → fine-tunedTyre set 1Tyre set 2+10% mass
G-NN0.614 → 0.3800.624 → 0.3680.489 → 0.329
MS-NN-bench0.462 → 0.4250.373 → 0.3470.396 → 0.375
MS-NN-full0.531 → 0.3950.468 → 0.3340.398 → 0.290

Yaw rate in rad/s, acceleration in m/s². Bold: best fine-tuned value per column; red: worst zero-shot value. nnodely deck slide 39. The slide does not say how the general network was fine-tuned on the same 20 s.

Read both tables together. The full structured model does not transfer better zero-shot; after a tyre change it transfers worst. Its advantage is that it adapts best from 20 s of data, which is what a selectively trainable structure should buy, and it is a different claim from better generalisation.

3 · A steering feedforward for a full-scale race car (Piccinini et al., IEEE ITSC 2025)

TUM's controller won the 2024 Abu Dhabi Autonomous Racing League with a heuristic feedforward: a kinematic steering angle, an understeer correction, a longitudinal-acceleration correction and an offset. MS-NN-steer replaces it. L5 2:12

δk=δkin(ayk,vxk)+δus(ayk)+δax(ayk,axk)+δoff The heuristic baseline that won A2RL 2024. nnodely deck slide 62.

The structured design starts from the handling diagram, which plots how far a car's steering departs from the kinematic ideal as lateral acceleration grows. Local linear models under triangular activations learn its shape. Activations read |ay| and the local models carry sign(ay), so one parameter set serves left and right turns; the coefficients then become functions of speed, and finally of longitudinal and vertical acceleration, with each local model designed as the local approximation of a double-track vehicle model around an operating point. A transient NCM scheduled by speed and acceleration follows, and a small network maps road-wheel to steering-wheel angle. Inputs are saturated to the training range before entering, so the controller cannot extrapolate wildly on the car. L5 4:17 L5 25:22

The windows look forward, not back: 10 samples of the planned trajectory from the high-level planner. A feedforward that sees the desired future can lead the steering actuator's delay, and the trained FIR weights show it: a decaying profile means no dead time, while a peak 0.2 s into the future reads as a 0.2 s actuator delay. L5 14:52

Steering angle errorMedium trainMedium validSmall trainSmall validFVU valid, small
G-NN: concatenated windows, two layers, ELU0.1600.2290.1320.3270.0792
MS-NN-base (Piccinini et al. 2023)0.0790.0970.0690.1150.0026
MS-NN-steer0.0680.0910.0630.0970.0020

RMSE in degrees on recorded Yas Marina telemetry. Small set: sector 3 of one lap; medium: sectors 1 and 3; validation: a second lap. Paper table shown on nnodely deck slide 65; the same excerpt reports 10 seeds in which MS-NN-steer's error barely moves and the G-NN's varies widely, with a similar picture across learning rates from 10−3 to 10−5.

Two readings the headline hides. Most of the gain over the general network was already in the 2023 base model (0.327 to 0.115 on the small set); the 2025 extension adds 16% on top. And the baseline is a two-layer perceptron on concatenated windows, so the comparison shows what structure buys over a small unstructured network, not over a tuned sequence model or an identified physical model. The seed study is the strongest evidence in the course: insensitivity to initialisation is a property a practitioner can use.

11

Training a controller through a frozen simulator

Can a controller learn to control a plant without collecting new data? Three steps. Fit a structured simulator of the plant with supervised learning. Connect an untrained structured controller to it in closed loop, the simulator's prediction feeding back as the measurement. Train the controller with the simulator frozen. Reinforcement learning would work but needs millions of interactions and treats the plant as a black box; here the simulator is differentiable, so the controller can follow its gradient. L5 33:51

desired [yk … yk+q] MS-NN controller trainable θ uk MS-NN simulator frozen, pre-trained ŷk prediction fed back as the measurement ∇ℒ through the frozen simulator, ℒ = ‖y − ŷ‖²
Redrawn from nnodely deck slide 73. The loss needs only feasible desired trajectories, no measured data.
ℒ=‖𝐲−𝐲^‖2,θk+1=θk−η∇ℒk Unsupervised: the target is the planner's trajectory. Gradient descent or Adam on the controller only.

The lecturers call this differentiable closed-loop controller training and use it in their lab, including with nonlinear MPC, where they report roughly an order of magnitude better sample efficiency than reinforcement learning, while RL can still find better solutions because it explores. The same gradient allows fast online adaptation when the simulator's parameters change, and the loss can be unrolled over a long horizon rather than one step. L5 41:15 L5 49:50

A student asks the right question: what if the simulator is wrong? Then the controller learns the wrong policy. The remedies named are to keep adapting the simulator on real data and fine-tune the controller, or to estimate the gradient from real measurements without a simulator. Rosati Papini adds the structural argument: a generic network controller trained through a generic network simulator finds strange inputs in regions the simulator never saw, where it extrapolates unphysically; a structured simulator extrapolates more physically, so the exploit is smaller. He cites an ICRA paper whose simulator had a physics branch and a learned residual, where gradients were backpropagated through the physics branch only. L5 44:28

The worked example is nnodely's mass-spring-damper. The direct model, two FIR branches, is exact for a linear system; it is trained one step ahead and then refined in closed loop over 1,500-step rollouts with a smaller learning rate, because long rollouts can destabilise training. The controller is a PID whose three gains are the only trainable parameters; nnodely's built-in integrate and differentiate operations form the error terms, and the models= argument freezes the simulator. The trained gains landed close to MATLAB's PID autotuner, and adding an overshoot penalty to the loss reshaped the response. L5 50:52

Lab 4

Tune a PID by backpropagation through a frozen simulator

reference response on the simulator same gains on the real plant training loss
Iteration0
Kp, Ki, Kd-
Overshoot, sim / real-
2% settling, sim / real-
Loss-

Train once with no simulator error. Then set the mass error to −40%, so the simulator believes the plant is lighter than it is, reset and train again: the gains look fine on the simulator and overshoot on the real plant.

What is computed. A real plant x″ + 0.4x′ + x = F and a simulator with the same structure and adjustable errors in mass and damping, both stepped at 0.01 s. A PID with gains K = exp(θ) acts on the simulator for 8 s against a unit step. The loss is the mean squared tracking error, plus β times the mean squared overshoot, plus an effort term (0.02 · mean F²) that keeps gains moderate. Exact gradients come from forward-mode differentiation through all 800 steps, and Adam updates θ. The real plant is never used for training; its curve shows what the gains do when deployed.

For readers from model-based reinforcement learning this is familiar: policy optimisation by backpropagation through a learned model, the family of PILCO, Dreamer's actor learning in imagination and short-horizon actor-critic methods. Rosati Papini's warning is the known failure of that family, model exploitation, and his remedy, a model that extrapolates physically, is a structural answer to a problem that the RL literature usually attacks with ensembles and uncertainty penalties.

12

Physics priors inside a foundation model

The lecturers see few examples so far and much interest: everyone at ICRA wants physics-informed foundation models, few know how to build one. Their argument is the course's opening one: embodied systems lack data, and physics can fill part of the gap. Their example comes from Piccinini's lab. L5 1:19:33

StyleVLA (Gao, Hua, Piccinini, Schäfer, Moller, Li and Betz; arXiv 2603.09482, March 2026) asks a vision-language-action model to drive in a requested style: default, balanced, comfort, sporty or safety. It fine-tunes Qwen3-VL-4B on an instruction dataset built in CARLA, 1.2k scenarios with 76k bird's-eye and 42k first-person samples. Only the vision encoder, low-rank adapters in the language model and an MLP trajectory decoder train, which fits a single consumer GPU. The model reads the image, the style prompt, the current state with 0.5 s of history and a goal, and outputs 3 s of positions, heading, speed and acceleration.

The loss mixes token cross-entropy, a regression term on the decoded trajectory, and a physics-informed kinematic consistency term (PIKC): roll the model's own previous waypoint forward with its predicted speed, heading and acceleration, and penalise the distance to the waypoint it actually predicted next. No ground truth enters the term. L5 1:26:53

x^t+1=xt+vtcosθtΔt+0.5atcosθt(Δt)2,y^t+1=yt+vtsinθtΔt+0.5atsinθt(Δt)2 ℒpikc=1T−1∑t=0T−1(‖xt+1−x^t+1‖2+‖yt+1−y^t+1‖2) Outlook slide 7. All quantities on the right of the first line are the model's own predictions.

On the paper's composite driving score StyleVLA reaches 0.55 on bird's-eye and 0.51 on first-person inputs, against 0.32 and 0.35 for Gemini 3 Pro used zero-shot; proprietary and open baselines, Alpamayo-R1 among them, trail. The lecture's two ablations follow.

Training dataADE mFDE mPSR
Small, 4.5k2.085.4320.60%
Medium, 20k1.513.9227.14%
Large, 40k1.473.8129.37%
Standard, 50k1.173.0633.19%
Loss, at 50kADE mFDE mPSR
CE1.473.8229.00%
CE + REG1.213.1732.08%
CE + REG + PIKC1.173.0633.19%

Outlook slide 9. ADE and FDE: average and final displacement error. PSR: share of trajectories with ADE below 1 m. No seeds or intervals on the slide.

The lecturers' take-home message: even in a four-billion-parameter model, where many would say data solves everything, a very basic physics term still helps, and they expect such examples to multiply in the next couple of years. L5 1:30:03

What the numbers support. PIKC moves ADE by 0.04 m at the largest data size, about a third of the step the regression head gave, with no seeds to tell it from run-to-run variation. The course's own thesis is that physics pays most where data is scarce, and that is the experiment missing here: the PIKC ablation at 4.5k samples. The comparison with zero-shot proprietary models measures in-domain fine-tuning as much as anything physical. The idea itself transfers cleanly to manipulation: a self-consistency penalty between predicted poses, velocities and gripper states in an action chunk costs nothing at inference and needs no labels.

13

The course's evidence, read against its own tables

A course is not a paper and owes no ablations. Its quantitative claims still teach the reader what physics priors buy, so each is read here through one checklist: what was measured, on which protocol, with how many runs, and whether the claim on the slide follows.

ClaimSourceWhat was measuredProtocol notesReading
Structured networks generalise better from small dataMungiello et al. 2026; Piccinini et al. 2025Offline RMSE and FVU on a held-out lap or log of the same vehicleSame platform, same track; small baselines; seeds reported only in the steering paperSupported for in-distribution held-out data
Selective fine-tuning adapts to new tyres or mass in 20 sMungiello et al. 2026Yaw-rate and acceleration RMSE before and after fine-tuningZero-shot the full model is worst after tyre changes; G-NN fine-tuning protocol not statedSupported for adaptation, not for zero-shot transfer
MS-NN-steer beats the A2RL-winning controllerPiccinini et al. 2025Error against executed steering on recorded lapsNo closed-loop lap with the learned feedforward reportedPartly: offline accuracy only
Insensitive to initialisation and learning ratePiccinini et al. 2025RMSE spread over 10 seeds; learning rates 10−3 to 10−5Spread shown as box plots in the paperSupported
GP residuals improve lap time by 10%Kabzan et al. 2019Lap time in closed loopStated in the lecture; I did not re-read the original paper for this pageCited, not re-examined here
Differentiable controller training is about 10× more sample-efficient than RLLecturers' differentiable MPC workNot shown in the coursePaper not named on the recordingNot shown
A physics loss helps a 4B VLAStyleVLA, arXiv 2603.09482ADE, FDE, PSR at 50k training samples0.04 m ADE gain; one run per configuration on the slide; preprintWeak: direction plausible, size within unknown noise
LNN and HNN conserve energy, VIN fixes small-data driftCranmer 2020; Greydanus 2019; Saemundsson 2020Energy over long rolloutsPendulum, double pendulum, mass-springSupported on toy systems

What I would ask for next

  • A closed-loop lap with MS-NN-steer against the heuristic feedforward: lap time, lateral tracking error and feedback effort. Offline error against executed steering includes the feedback controller's corrections.
  • Two stronger baselines on the same windows: a recurrent or temporal-convolution network with weight decay and early stopping, and an identified single-track model with fitted tyre curves.
  • Seeds and intervals for the coupled-dynamics tables, which the steering paper shows the lab can produce.
  • One test outside the training envelope on a named axis: higher speed, another track or another surface. A second lap of the same circuit is held-out data, not a distribution shift.
  • The StyleVLA loss ablation at 4.5k samples with three seeds, where the course's thesis predicts the physics term should matter most.
  • A study that combines all three routes, which the lecturers name as the main open direction and the course does not yet contain.
14

What to take from it

A working recipe

  1. Try physics-guided features first. v², sign(v), estimated forces, a filtered rate: no training changes, a strong baseline for everything after.
  2. With an imperfect model and patchy data, add a physics loss where data is missing, and tune γ on the extrapolation error, never the training error.
  3. With little data and a known equation, structure the architecture after it: one branch per force, summed; bounded physical parameters; fading-memory outputs; windows sized from the residual spectrum.
  4. Validate like a physicist: residual spectra against the noise floor, trained FIR weights read as impulse responses, information criteria next to RMSE.
  5. Adapt selectively: when conditions change, fine-tune the module the physics says changed.
  6. Train controllers through structured simulators, and treat every simulator error as a policy error.

Where it meets world models and robot learning

Three directions follow for my own work on action-conditioned world models and manipulation policies. A kinematic self-consistency loss on manipulation action chunks, the PIKC idea applied to predicted poses, velocities and gripper states, is a day of code and a few GPU-days to test on LIBERO at 5 to 10% of the demonstrations, where it should matter if it matters anywhere. Physics-guided proprioceptive features for a VLA (end-effector velocity, a wrench estimated from joint torques) are the cheapest route of all, limited mainly by how realistic simulated torques are. And residual spectra are a cheap diagnostic that world-model papers almost never report: does a latent predictor's error sit on the noise floor or above it, and at which frequencies?

The idea this course raised: model-structured ego-dynamics inside a latent world model

The adaptation result suggests a cheap experiment. In an action-conditioned latent world model for a mobile robot, route the robot's own motion through a model-structured branch: windows of commanded linear and angular velocity in, realised body twist out, scheduled by speed and payload, with FIR responses that expose actuation lag. Keep the scene latent learned. Then measure three axes separately, in Isaac Lab and on the real robot: open-loop rollout error split into ego drift and scene error, response to counterfactual commands, and planning success with latent lookahead. The decisive test is whether fine-tuning only the ego branch on seconds of real driving closes a sim-to-real gap that would otherwise need the whole model retrained.

Cost: about one GPU-week for a 3 × 3 grid (ego branch none, unicycle plus residual, or NCM; 10, 30 or 100% of real data) at three seeds. The objection I expect is that ego motion is a small share of a world model's error. So the first step is to measure that share before building anything.

The lecturers' open problems

Combining the three routes instead of using one; physics priors to fill the data gap of embodied foundation models; differentiable end-to-end training for performance and fast adaptation; lifelong online adaptation that is physics-aware, changing only what physics says has changed; provable, verifiable safety, hardest for foundation models; and interpretability for public acceptance. L5 1:31:08

15

Glossary

Tap a card for an everyday analogy.

16

Lecture index

Every topic with the second it starts. Times link straight into the video.

17

What is faithful and what is schematic

Faithful to the sources: every equation, every table and every number in Sections 00 to 14, each with its slide or paper reference, and the nnodely code, quoted verbatim. Redrawn, not copied: the taxonomy, the spectrum, the longitudinal network, the NCM block diagram and the controller-training loop. Schematic: the four labs. They are exact computations on toy systems chosen to show one mechanism each (a linear-in-parameters stand-in for a PINN, a one-term black-box error and a perturbed Hamiltonian for Lab 2, a first-order lag for Lab 3, a linear plant for Lab 4), not reproductions of the papers' experiments.

Checked on 1 Oct 2026 in headless Chromium at 1440×900 in light and dark mode and at 390×844: no page or console errors and no horizontal scroll; all 20 equation blocks render in Noto Sans Math; every lab control was exercised and the readouts match the notes (Lab 1 extrapolation RMSE 0.742 with data only, 0.056 at γ = 10−3, 0.010 at the default; Lab 2 under leapfrog, exact field −0.01% and black box −68%; Lab 3 triangular activations summing to exactly 1; Lab 4 gains trained on a simulator 40% too light overshooting 3.3% there and 13.7% on the real plant); the quiz, map, code notes, glossary and 144-entry lecture index all respond. Touch input was not tested on a physical phone, and the equations need a browser with MathML (Chrome 109 or later, Firefox, Safari).