import%20marimo%0A%0A__generated_with%20%3D%20%220.24.0%22%0Aapp%20%3D%20marimo.App(width%3D%22medium%22)%0A%0A%0A%40app.cell%0Adef%20_()%3A%0A%20%20%20%20import%20marimo%20as%20mo%0A%20%20%20%20import%20numpy%20as%20np%0A%20%20%20%20import%20pandas%20as%20pd%0A%20%20%20%20import%20plotly.graph_objects%20as%20go%0A%20%20%20%20from%20plotly.subplots%20import%20make_subplots%0A%20%20%20%20import%20torch%0A%0A%20%20%20%20return%20go%2C%20make_subplots%2C%20mo%2C%20np%2C%20pd%2C%20torch%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%20Note%2008%3A%20Matrix%20Calculus%20and%20Vector%20Derivatives%0A%0A%20%20%20%20%26larr%3B%20Previous%20Note%3A%20%5B07%20Spectral%20Decomposition%5D(07_spectral_decomposition.py)%20%7C%20Next%20Note%3A%20%5B09%20Condition%20Number%5D(09_condition_number.py)%20%26rarr%3B%0A%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%20%5Ba%5D%20Why%20do%20you%20need%20to%20know%20these%20concepts%3F%0A%0A%20%20%20%20Matrix%20calculus%20is%20the%20foundational%20mathematical%20language%20of%20modern%20machine%20learning%2C%20deep%20learning%2C%20and%20numerical%20optimization.%20When%20training%20neural%20networks%20with%20millions%20or%20billions%20of%20parameters%2C%20computing%20scalar%20derivatives%20with%20respect%20to%20each%20individual%20scalar%20weight%20parameter%20is%20algebraically%20intractable%20and%20conceptually%20unwieldy.%20Matrix%20calculus%20enables%20us%20to%20express%20gradients%2C%20Jacobians%2C%20and%20Hessians%20in%20compact%2C%20parallelizable%20matrix%20and%20vector%20forms.%0A%0A%20%20%20%20Key%20motivations%20across%20data%20science%2C%20artificial%20intelligence%2C%20and%20statistics%3A%0A%20%20%20%201.%20**Reverse-Mode%20Automatic%20Differentiation%20(Backpropagation)**%3A%20Deep%20learning%20frameworks%20(PyTorch%2C%20JAX%2C%20TensorFlow)%20rely%20on%20reverse-mode%20automatic%20differentiation.%20Backpropagation%20computes%20vector-Jacobian%20products%20(VJPs)%20rather%20than%20instantiating%20full%20Jacobian%20matrices%20in%20memory%2C%20allowing%20gradient%20computation%20across%20billions%20of%20parameters%20in%20linear%20time.%0A%20%20%20%202.%20**Loss%20Surface%20Geometry%20and%20Gradient%20Descent**%3A%20The%20gradient%20vector%20%24%5Cnabla_%7B%5Cmathbf%7Bw%7D%7D%20%5Cmathcal%7BL%7D%24%20specifies%20the%20exact%20direction%20of%20steepest%20ascent%20on%20high-dimensional%20loss%20landscapes%2C%20leading%20directly%20to%20the%20canonical%20optimization%20update%20%24%5Cmathbf%7Bw%7D_%7Bt%2B1%7D%20%3D%20%5Cmathbf%7Bw%7D_t%20-%20%5Ceta%20%5Cnabla%20%5Cmathcal%7BL%7D(%5Cmathbf%7Bw%7D_t)%24.%20Crucially%2C%20the%20gradient%20is%20always%20orthogonal%20to%20the%20loss%20contour%20level%20sets.%0A%20%20%20%203.%20**Curvature%20and%20Second-Order%20Optimization%20(Hessians)**%3A%20The%20Hessian%20matrix%20%24%5Cmathbf%7BH%7D%20%3D%20%5Cnabla%5E2%20%5Cmathcal%7BL%7D%24%20governs%20optimization%20dynamics%2C%20the%20maximum%20stable%20learning%20rate%20(%24%5Ceta%20%3C%202%20%2F%20%5Clambda_%7B%5Cmax%7D%24)%2C%20Newton-Raphson%20acceleration%2C%20and%20the%20distinction%20between%20sharp%20minima%20(poor%20generalization)%20and%20flat%20minima%20(robust%20generalization).%0A%20%20%20%204.%20**Analytical%20Derivations%20of%20Closed-Form%20Estimators**%3A%20Setting%20vector%20and%20matrix%20derivatives%20to%20zero%20yields%20fundamental%20estimators%2C%20including%20Ordinary%20Least%20Squares%20(the%20Normal%20Equations%20%24%5Cmathbf%7Bw%7D%5E*%20%3D%20(%5Cmathbf%7BX%7D%5ET%20%5Cmathbf%7BX%7D)%5E%7B-1%7D%20%5Cmathbf%7BX%7D%5ET%20%5Cmathbf%7By%7D%24)%2C%20Ridge%20Regression%2C%20Fisher%20Linear%20Discriminant%2C%20Kalman%20filtering%2C%20and%20Gaussian%20Maximum%20Likelihood%20Estimators.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%20%5Bb%5D%20Concept%20explanation%20with%20their%20role%20in%20ML%2FAI%2FStats%3F%0A%0A%20%20%20%20%23%23%23%20Layout%20Conventions%3A%20Numerator%20vs%20Denominator%20Layout%0A%0A%20%20%20%20Before%20evaluating%20vector%20and%20matrix%20derivatives%2C%20one%20must%20adopt%20a%20layout%20convention.%20In%20machine%20learning%20and%20statistics%2C%20the%20standard%20convention%20is%20the%20**numerator%20layout**%20(or%20standard%20vector%20convention)%3A%0A%20%20%20%20*%20The%20gradient%20of%20a%20scalar%20function%20%24f(%5Cmathbf%7Bx%7D)%24%20with%20respect%20to%20a%20column%20vector%20%24%5Cmathbf%7Bx%7D%20%5Cin%20%5Cmathbb%7BR%7D%5En%24%20is%20defined%20as%20a%20column%20vector%20matching%20the%20dimension%20of%20%24%5Cmathbf%7Bx%7D%24%3A%20%24%5Cnabla_%7B%5Cmathbf%7Bx%7D%7D%20f%20%5Cin%20%5Cmathbb%7BR%7D%5E%7Bn%20%5Ctimes%201%7D%24.%0A%20%20%20%20*%20The%20Jacobian%20of%20a%20vector-valued%20function%20%24%5Cmathbf%7Bf%7D%3A%20%5Cmathbb%7BR%7D%5En%20%5Cto%20%5Cmathbb%7BR%7D%5Em%24%20is%20an%20%24m%20%5Ctimes%20n%24%20matrix%20whose%20%24i%24-th%20row%20is%20the%20transpose%20of%20the%20gradient%20of%20%24f_i%24.%0A%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%23%20Gradient%20of%20a%20Scalar-Valued%20Function%0A%0A%20%20%20%20Let%20%24f%3A%20%5Cmathbb%7BR%7D%5En%20%5Cto%20%5Cmathbb%7BR%7D%24%20be%20a%20differentiable%20scalar%20function%20of%20an%20%24n%24-dimensional%20vector%20%24%5Cmathbf%7Bx%7D%20%3D%20%5Bx_1%2C%20x_2%2C%20%5Cdots%2C%20x_n%5D%5ET%24.%20The%20gradient%20%24%5Cnabla_%7B%5Cmathbf%7Bx%7D%7D%20f(%5Cmathbf%7Bx%7D)%24%20is%20the%20vector%20of%20all%20first-order%20partial%20derivatives%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cnabla_%7B%5Cmathbf%7Bx%7D%7D%20f(%5Cmathbf%7Bx%7D)%20%3D%20%5Cbegin%7Bbmatrix%7D%0A%20%20%20%20%5Cfrac%7B%5Cpartial%20f%7D%7B%5Cpartial%20x_1%7D%20%5C%5C%0A%20%20%20%20%5Cfrac%7B%5Cpartial%20f%7D%7B%5Cpartial%20x_2%7D%20%5C%5C%0A%20%20%20%20%5Cvdots%20%5C%5C%0A%20%20%20%20%5Cfrac%7B%5Cpartial%20f%7D%7B%5Cpartial%20x_n%7D%0A%20%20%20%20%5Cend%7Bbmatrix%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20%23%23%23%23%20Geometric%20Properties%20of%20the%20Gradient%0A%20%20%20%20*%20**Direction%20of%20Steepest%20Ascent**%3A%20The%20directional%20derivative%20in%20unit%20direction%20%24%5Cmathbf%7Bu%7D%24%20is%20%24D_%7B%5Cmathbf%7Bu%7D%7D%20f%20%3D%20%5Cnabla%20f(%5Cmathbf%7Bx%7D)%5ET%20%5Cmathbf%7Bu%7D%20%3D%20%5C%7C%5Cnabla%20f%5C%7C_2%20%5Ccos(%5Ctheta)%24.%20This%20reaches%20its%20global%20maximum%20when%20%24%5Ctheta%20%3D%200%24%20(%24%5Cmathbf%7Bu%7D%24%20aligns%20with%20%24%5Cnabla%20f%24).%0A%20%20%20%20*%20**Orthogonality%20to%20Level%20Sets**%3A%20Along%20any%20level%20contour%20curve%20%24f(%5Cmathbf%7Bx%7D)%20%3D%20c%24%2C%20the%20directional%20rate%20of%20change%20is%20zero.%20Therefore%2C%20%24%5Cnabla%20f(%5Cmathbf%7Bx%7D)%24%20is%20strictly%20perpendicular%20(orthogonal)%20to%20the%20tangent%20hyperplane%20of%20the%20level%20curve%20at%20%24%5Cmathbf%7Bx%7D%24.%0A%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%23%20Jacobian%20of%20a%20Vector-Valued%20Function%0A%0A%20%20%20%20Let%20%24%5Cmathbf%7Bf%7D%3A%20%5Cmathbb%7BR%7D%5En%20%5Cto%20%5Cmathbb%7BR%7D%5Em%24%20be%20a%20vector-valued%20mapping%20%24%5Cmathbf%7Bf%7D(%5Cmathbf%7Bx%7D)%20%3D%20%5Bf_1(%5Cmathbf%7Bx%7D)%2C%20f_2(%5Cmathbf%7Bx%7D)%2C%20%5Cdots%2C%20f_m(%5Cmathbf%7Bx%7D)%5D%5ET%24.%20The%20**Jacobian%20matrix**%20%24%5Cmathbf%7BJ%7D%20%5Cin%20%5Cmathbb%7BR%7D%5E%7Bm%20%5Ctimes%20n%7D%24%20gathers%20all%20%24m%20%5Ctimes%20n%24%20first-order%20partial%20derivatives%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cmathbf%7BJ%7D%20%3D%20%5Cfrac%7B%5Cpartial%20%5Cmathbf%7Bf%7D%7D%7B%5Cpartial%20%5Cmathbf%7Bx%7D%7D%20%3D%20%5Cbegin%7Bbmatrix%7D%0A%20%20%20%20%5Cfrac%7B%5Cpartial%20f_1%7D%7B%5Cpartial%20x_1%7D%20%26%20%5Cfrac%7B%5Cpartial%20f_1%7D%7B%5Cpartial%20x_2%7D%20%26%20%5Cdots%20%26%20%5Cfrac%7B%5Cpartial%20f_1%7D%7B%5Cpartial%20x_n%7D%20%5C%5C%0A%20%20%20%20%5Cfrac%7B%5Cpartial%20f_2%7D%7B%5Cpartial%20x_1%7D%20%26%20%5Cfrac%7B%5Cpartial%20f_2%7D%7B%5Cpartial%20x_2%7D%20%26%20%5Cdots%20%26%20%5Cfrac%7B%5Cpartial%20f_2%7D%7B%5Cpartial%20x_n%7D%20%5C%5C%0A%20%20%20%20%5Cvdots%20%26%20%5Cvdots%20%26%20%5Cddots%20%26%20%5Cvdots%20%5C%5C%0A%20%20%20%20%5Cfrac%7B%5Cpartial%20f_m%7D%7B%5Cpartial%20x_1%7D%20%26%20%5Cfrac%7B%5Cpartial%20f_m%7D%7B%5Cpartial%20x_2%7D%20%26%20%5Cdots%20%26%20%5Cfrac%7B%5Cpartial%20f_m%7D%7B%5Cpartial%20x_n%7D%0A%20%20%20%20%5Cend%7Bbmatrix%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20The%20Jacobian%20acts%20as%20the%20optimal%20local%20linear%20map%20approximating%20%24%5Cmathbf%7Bf%7D%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cmathbf%7Bf%7D(%5Cmathbf%7Bx%7D%20%2B%20%5CDelta%5Cmathbf%7Bx%7D)%20%5Capprox%20%5Cmathbf%7Bf%7D(%5Cmathbf%7Bx%7D)%20%2B%20%5Cmathbf%7BJ%7D%20%5CDelta%5Cmathbf%7Bx%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20%23%23%23%23%20Vector-Jacobian%20Products%20(VJP)%20in%20Backpropagation%0A%20%20%20%20In%20reverse-mode%20automatic%20differentiation%20(backpropagation)%2C%20a%20scalar%20loss%20%24%5Cmathcal%7BL%7D%24%20depends%20on%20downstream%20activations%20%24%5Cmathbf%7By%7D%20%3D%20%5Cmathbf%7Bf%7D(%5Cmathbf%7Bx%7D)%24.%20To%20compute%20%24%5Cnabla_%7B%5Cmathbf%7Bx%7D%7D%20%5Cmathcal%7BL%7D%24%2C%20the%20chain%20rule%20states%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20(%5Cnabla_%7B%5Cmathbf%7Bx%7D%7D%20%5Cmathcal%7BL%7D)%5ET%20%3D%20(%5Cnabla_%7B%5Cmathbf%7By%7D%7D%20%5Cmathcal%7BL%7D)%5ET%20%5Cmathbf%7BJ%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20Rather%20than%20materializing%20the%20entire%20%24m%20%5Ctimes%20n%24%20Jacobian%20matrix%20%24%5Cmathbf%7BJ%7D%24%2C%20backpropagation%20computes%20the%20product%20directly%3A%20a%20cotangent%20vector%20%24%5Cmathbf%7Bv%7D%5ET%20%3D%20(%5Cnabla_%7B%5Cmathbf%7By%7D%7D%20%5Cmathcal%7BL%7D)%5ET%24%20contracted%20with%20%24%5Cmathbf%7BJ%7D%24.%20This%20Vector-Jacobian%20Product%20(VJP)%20requires%20%24%5Cmathcal%7BO%7D(n%20%2B%20m)%24%20memory%20instead%20of%20%24%5Cmathcal%7BO%7D(nm)%24.%0A%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%23%20Hessian%20Matrix%3A%20Second-Order%20Curvature%0A%0A%20%20%20%20For%20a%20twice-differentiable%20scalar%20function%20%24f%3A%20%5Cmathbb%7BR%7D%5En%20%5Cto%20%5Cmathbb%7BR%7D%24%2C%20the%20**Hessian%20matrix**%20%24%5Cmathbf%7BH%7D%20%3D%20%5Cnabla%5E2%20f(%5Cmathbf%7Bx%7D)%20%5Cin%20%5Cmathbb%7BR%7D%5E%7Bn%20%5Ctimes%20n%7D%24%20contains%20all%20second-order%20partial%20derivatives%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cmathbf%7BH%7D_%7Bij%7D%20%3D%20%5Cfrac%7B%5Cpartial%5E2%20f%7D%7B%5Cpartial%20x_i%20%5Cpartial%20x_j%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20By%20Clairaut's%20Theorem%20(Schwarz's%20theorem)%2C%20if%20all%20second%20partial%20derivatives%20are%20continuous%2C%20mixed%20partial%20derivatives%20commute%20(%24%5Cfrac%7B%5Cpartial%5E2%20f%7D%7B%5Cpartial%20x_i%20%5Cpartial%20x_j%7D%20%3D%20%5Cfrac%7B%5Cpartial%5E2%20f%7D%7B%5Cpartial%20x_j%20%5Cpartial%20x_i%7D%24)%2C%20making%20the%20Hessian%20unconditionally%20symmetric%20(%24%5Cmathbf%7BH%7D%20%3D%20%5Cmathbf%7BH%7D%5ET%24).%0A%0A%20%20%20%20*%20If%20%24%5Cmathbf%7BH%7D%20%5Csucc%200%24%20(positive%20definite%2C%20all%20%24%5Clambda_i%20%3E%200%24)%2C%20the%20function%20is%20strictly%20convex%20locally%2C%20and%20a%20critical%20point%20(%24%5Cnabla%20f%20%3D%20%5Cmathbf%7B0%7D%24)%20is%20a%20unique%20local%20minimum.%0A%20%20%20%20*%20If%20%24%5Cmathbf%7BH%7D%24%20has%20both%20positive%20and%20negative%20eigenvalues%2C%20the%20critical%20point%20is%20a%20**saddle%20point**.%0A%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%23%20The%20Top%20Fundamental%20Matrix%20Calculus%20Identities%0A%0A%20%20%20%20Mastering%20these%20identities%20enables%20direct%20derivation%20of%20optimization%20equations%20without%20component-wise%20expansions%3A%0A%0A%20%20%20%20%23%23%23%23%20Rule%201%3A%20Gradient%20of%20a%20Linear%20Form%0A%20%20%20%20For%20constant%20vector%20%24%5Cmathbf%7Ba%7D%20%5Cin%20%5Cmathbb%7BR%7D%5En%24%20and%20variable%20vector%20%24%5Cmathbf%7Bx%7D%20%5Cin%20%5Cmathbb%7BR%7D%5En%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cnabla_%7B%5Cmathbf%7Bx%7D%7D%20(%5Cmathbf%7Ba%7D%5ET%20%5Cmathbf%7Bx%7D)%20%3D%20%5Cmathbf%7Ba%7D%2C%20%5Cquad%20%5Cnabla_%7B%5Cmathbf%7Bx%7D%7D%20(%5Cmathbf%7Bx%7D%5ET%20%5Cmathbf%7Ba%7D)%20%3D%20%5Cmathbf%7Ba%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20Derivation%3A%20%24%5Cmathbf%7Ba%7D%5ET%20%5Cmathbf%7Bx%7D%20%3D%20%5Csum_%7Bi%3D1%7D%5En%20a_i%20x_i%24.%20Taking%20%24%5Cfrac%7B%5Cpartial%7D%7B%5Cpartial%20x_k%7D%20%5Csum_%7Bi%3D1%7D%5En%20a_i%20x_i%20%3D%20a_k%24.%20Stacking%20all%20%24k%24%20yields%20%24%5Cmathbf%7Ba%7D%24.%0A%0A%20%20%20%20%23%23%23%23%20Rule%202%3A%20Jacobian%20of%20a%20Linear%20Mapping%0A%20%20%20%20For%20constant%20matrix%20%24%5Cmathbf%7BA%7D%20%5Cin%20%5Cmathbb%7BR%7D%5E%7Bm%20%5Ctimes%20n%7D%24%20and%20variable%20vector%20%24%5Cmathbf%7Bx%7D%20%5Cin%20%5Cmathbb%7BR%7D%5En%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cfrac%7B%5Cpartial%20(%5Cmathbf%7BA%7D%5Cmathbf%7Bx%7D)%7D%7B%5Cpartial%20%5Cmathbf%7Bx%7D%7D%20%3D%20%5Cmathbf%7BA%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20Derivation%3A%20The%20%24i%24-th%20component%20of%20%24%5Cmathbf%7BA%7D%5Cmathbf%7Bx%7D%24%20is%20%24(%5Cmathbf%7BA%7D%5Cmathbf%7Bx%7D)_i%20%3D%20%5Csum_%7Bj%3D1%7D%5En%20A_%7Bij%7D%20x_j%24.%20Thus%20%24%5Cfrac%7B%5Cpartial%20(%5Cmathbf%7BA%7D%5Cmathbf%7Bx%7D)_i%7D%7B%5Cpartial%20x_k%7D%20%3D%20A_%7Bik%7D%24%2C%20which%20reconstructs%20matrix%20%24%5Cmathbf%7BA%7D%24.%0A%0A%20%20%20%20%23%23%23%23%20Rule%203%3A%20Gradient%20and%20Hessian%20of%20a%20Quadratic%20Form%0A%20%20%20%20For%20square%20matrix%20%24%5Cmathbf%7BA%7D%20%5Cin%20%5Cmathbb%7BR%7D%5E%7Bn%20%5Ctimes%20n%7D%24%20and%20variable%20vector%20%24%5Cmathbf%7Bx%7D%20%5Cin%20%5Cmathbb%7BR%7D%5En%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cnabla_%7B%5Cmathbf%7Bx%7D%7D%20(%5Cmathbf%7Bx%7D%5ET%20%5Cmathbf%7BA%7D%20%5Cmathbf%7Bx%7D)%20%3D%20(%5Cmathbf%7BA%7D%20%2B%20%5Cmathbf%7BA%7D%5ET)%5Cmathbf%7Bx%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20When%20%24%5Cmathbf%7BA%7D%24%20is%20symmetric%20(%24%5Cmathbf%7BA%7D%20%3D%20%5Cmathbf%7BA%7D%5ET%24)%2C%20this%20simplifies%20to%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cnabla_%7B%5Cmathbf%7Bx%7D%7D%20(%5Cmathbf%7Bx%7D%5ET%20%5Cmathbf%7BA%7D%20%5Cmathbf%7Bx%7D)%20%3D%202%5Cmathbf%7BA%7D%5Cmathbf%7Bx%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20The%20Hessian%20of%20a%20quadratic%20form%20is%20a%20constant%20matrix%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cnabla_%7B%5Cmathbf%7Bx%7D%7D%5E2%20(%5Cmathbf%7Bx%7D%5ET%20%5Cmathbf%7BA%7D%20%5Cmathbf%7Bx%7D)%20%3D%20%5Cmathbf%7BA%7D%20%2B%20%5Cmathbf%7BA%7D%5ET%20%3D%202%5Cmathbf%7BA%7D%20%5Cquad%20(%5Ctext%7Bwhen%20%7D%20%5Cmathbf%7BA%7D%20%3D%20%5Cmathbf%7BA%7D%5ET)%0A%20%20%20%20%24%24%0A%0A%20%20%20%20%23%23%23%23%20Rule%204%3A%20Gradients%20of%20a%20Bilinear%20Form%0A%20%20%20%20For%20matrix%20%24%5Cmathbf%7BA%7D%20%5Cin%20%5Cmathbb%7BR%7D%5E%7Bm%20%5Ctimes%20n%7D%24%20and%20independent%20vectors%20%24%5Cmathbf%7Bx%7D%20%5Cin%20%5Cmathbb%7BR%7D%5Em%24%2C%20%24%5Cmathbf%7By%7D%20%5Cin%20%5Cmathbb%7BR%7D%5En%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cnabla_%7B%5Cmathbf%7Bx%7D%7D%20(%5Cmathbf%7Bx%7D%5ET%20%5Cmathbf%7BA%7D%20%5Cmathbf%7By%7D)%20%3D%20%5Cmathbf%7BA%7D%5Cmathbf%7By%7D%2C%20%5Cquad%20%5Cnabla_%7B%5Cmathbf%7By%7D%7D%20(%5Cmathbf%7Bx%7D%5ET%20%5Cmathbf%7BA%7D%20%5Cmathbf%7By%7D)%20%3D%20%5Cmathbf%7BA%7D%5ET%20%5Cmathbf%7Bx%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20%23%23%23%23%20Rule%205%3A%20Matrix%20Derivative%20of%20an%20Inner%20Product%20%2F%20Trace%20Form%0A%20%20%20%20For%20constant%20vectors%20%24%5Cmathbf%7Ba%7D%20%5Cin%20%5Cmathbb%7BR%7D%5Em%24%2C%20%24%5Cmathbf%7Bb%7D%20%5Cin%20%5Cmathbb%7BR%7D%5En%24%20and%20variable%20matrix%20%24%5Cmathbf%7BX%7D%20%5Cin%20%5Cmathbb%7BR%7D%5E%7Bm%20%5Ctimes%20n%7D%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cnabla_%7B%5Cmathbf%7BX%7D%7D%20(%5Cmathbf%7Ba%7D%5ET%20%5Cmathbf%7BX%7D%20%5Cmathbf%7Bb%7D)%20%3D%20%5Cmathbf%7Ba%7D%5Cmathbf%7Bb%7D%5ET%0A%20%20%20%20%24%24%0A%0A%20%20%20%20This%20outer-product%20rule%20is%20the%20fundamental%20equation%20for%20updating%20weights%20in%20dense%20layers%20during%20backpropagation%3A%20%24%5Cfrac%7B%5Cpartial%20%5Cmathcal%7BL%7D%7D%7B%5Cpartial%20%5Cmathbf%7BW%7D%7D%20%3D%20%5Cboldsymbol%7B%5Cdelta%7D%20%5Cmathbf%7Bh%7D%5ET%24%2C%20where%20%24%5Cboldsymbol%7B%5Cdelta%7D%24%20is%20the%20backpropagated%20error%20vector%20and%20%24%5Cmathbf%7Bh%7D%24%20is%20the%20layer%20input%20vector.%0A%0A%20%20%20%20%23%23%23%23%20Application%3A%20Deriving%20the%20Ordinary%20Least%20Squares%20(OLS)%20Normal%20Equations%0A%20%20%20%20Consider%20the%20sum%20of%20squared%20errors%20loss%20for%20linear%20regression%20with%20feature%20matrix%20%24%5Cmathbf%7BX%7D%20%5Cin%20%5Cmathbb%7BR%7D%5E%7BN%20%5Ctimes%20d%7D%24%20and%20target%20vector%20%24%5Cmathbf%7By%7D%20%5Cin%20%5Cmathbb%7BR%7D%5EN%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cmathcal%7BL%7D(%5Cmathbf%7Bw%7D)%20%3D%20%5C%7C%5Cmathbf%7BX%7D%5Cmathbf%7Bw%7D%20-%20%5Cmathbf%7By%7D%5C%7C_2%5E2%20%3D%20(%5Cmathbf%7BX%7D%5Cmathbf%7Bw%7D%20-%20%5Cmathbf%7By%7D)%5ET%20(%5Cmathbf%7BX%7D%5Cmathbf%7Bw%7D%20-%20%5Cmathbf%7By%7D)%0A%20%20%20%20%24%24%0A%0A%20%20%20%20Expanding%20the%20product%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cmathcal%7BL%7D(%5Cmathbf%7Bw%7D)%20%3D%20%5Cmathbf%7Bw%7D%5ET%20%5Cmathbf%7BX%7D%5ET%20%5Cmathbf%7BX%7D%20%5Cmathbf%7Bw%7D%20-%202%5Cmathbf%7By%7D%5ET%20%5Cmathbf%7BX%7D%5Cmathbf%7Bw%7D%20%2B%20%5Cmathbf%7By%7D%5ET%20%5Cmathbf%7By%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20Applying%20Rule%203%20to%20the%20quadratic%20term%20and%20Rule%201%20to%20the%20linear%20term%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cnabla_%7B%5Cmathbf%7Bw%7D%7D%20%5Cmathcal%7BL%7D(%5Cmathbf%7Bw%7D)%20%3D%202%5Cmathbf%7BX%7D%5ET%20%5Cmathbf%7BX%7D%5Cmathbf%7Bw%7D%20-%202%5Cmathbf%7BX%7D%5ET%20%5Cmathbf%7By%7D%20%3D%202%5Cmathbf%7BX%7D%5ET%20(%5Cmathbf%7BX%7D%5Cmathbf%7Bw%7D%20-%20%5Cmathbf%7By%7D)%0A%20%20%20%20%24%24%0A%0A%20%20%20%20Setting%20the%20gradient%20to%20zero%20yields%20the%20celebrated%20**Normal%20Equations**%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cmathbf%7BX%7D%5ET%20%5Cmathbf%7BX%7D%20%5Cmathbf%7Bw%7D%5E*%20%3D%20%5Cmathbf%7BX%7D%5ET%20%5Cmathbf%7By%7D%20%5Cimplies%20%5Cmathbf%7Bw%7D%5E*%20%3D%20(%5Cmathbf%7BX%7D%5ET%20%5Cmathbf%7BX%7D)%5E%7B-1%7D%20%5Cmathbf%7BX%7D%5ET%20%5Cmathbf%7By%7D%0A%20%20%20%20%24%24%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%20%5Bc%5D%20Interactive%20Visualizations%3A%20Loss%20Surface%20Geometry%20and%20Gradient%20Orthogonality%0A%0A%20%20%20%20The%20interactive%20subplots%20below%20demonstrate%20the%20fundamental%20geometry%20of%20matrix%20calculus%3A%0A%20%20%20%20*%20**Left%20Panel**%3A%202D%20level%20contours%20of%20the%20quadratic%20loss%20%24f(%5Cmathbf%7Bx%7D)%20%3D%20%5Cfrac%7B1%7D%7B2%7D%20%5Cmathbf%7Bx%7D%5ET%20%5Cmathbf%7BA%7D%20%5Cmathbf%7Bx%7D%20-%20%5Cmathbf%7Bb%7D%5ET%20%5Cmathbf%7Bx%7D%24.%20At%20the%20probe%20point%20%24%5Cmathbf%7Bx%7D_0%24%2C%20the%20gradient%20%24%5Cnabla%20f%24%20is%20strictly%20orthogonal%20to%20the%20level%20contour%20tangent%20line.%20The%20red%20trajectory%20shows%20Gradient%20Descent%20steps%20converging%20to%20the%20analytical%20optimum%20%24%5Cmathbf%7Bx%7D%5E*%20%3D%20%5Cmathbf%7BA%7D%5E%7B-1%7D%5Cmathbf%7Bb%7D%24.%0A%20%20%20%20*%20**Right%20Panel**%3A%203D%20surface%20representation%20showing%20the%20loss%20bowl%2C%20the%20local%20tangent%20plane%20at%20%24%5Cmathbf%7Bx%7D_0%24%2C%20and%20the%203D%20descent%20trajectory%20down%20the%20surface.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(go%2C%20make_subplots%2C%20np)%3A%0A%20%20%20%20%23%20Construct%20symmetric%20positive-definite%20matrix%20A%20and%20vector%20b%0A%20%20%20%20a_mat%20%3D%20np.array(%5B%5B3.0%2C%201.0%5D%2C%20%5B1.0%2C%202.0%5D%5D)%0A%20%20%20%20b_vec%20%3D%20np.array(%5B2.0%2C%201.0%5D)%0A%20%20%20%20x_star%20%3D%20np.linalg.solve(a_mat%2C%20b_vec)%0A%0A%20%20%20%20def%20loss_func(x1%2C%20x2)%3A%0A%20%20%20%20%20%20%20%20return%200.5%20*%20(a_mat%5B0%2C%200%5D%20*%20x1**2%20%2B%202%20*%20a_mat%5B0%2C%201%5D%20*%20x1%20*%20x2%20%2B%20a_mat%5B1%2C%201%5D%20*%20x2**2)%20-%20(%0A%20%20%20%20%20%20%20%20%20%20%20%20b_vec%5B0%5D%20*%20x1%20%2B%20b_vec%5B1%5D%20*%20x2%0A%20%20%20%20%20%20%20%20)%0A%0A%20%20%20%20def%20grad_loss(x)%3A%0A%20%20%20%20%20%20%20%20return%20a_mat%20%40%20x%20-%20b_vec%0A%0A%20%20%20%20%23%20Grid%20for%20contour%20and%203D%20surface%0A%20%20%20%20grid_x%20%3D%20np.linspace(-1.5%2C%202.5%2C%2060)%0A%20%20%20%20grid_y%20%3D%20np.linspace(-1.5%2C%202.5%2C%2060)%0A%20%20%20%20grid_x_mesh%2C%20grid_y_mesh%20%3D%20np.meshgrid(grid_x%2C%20grid_y)%0A%20%20%20%20z_loss%20%3D%20loss_func(grid_x_mesh%2C%20grid_y_mesh)%0A%0A%20%20%20%20%23%20Gradient%20Descent%20Trajectory%0A%20%20%20%20x_init%20%3D%20np.array(%5B-1.2%2C%202.0%5D)%0A%20%20%20%20learning_rate%20%3D%200.25%0A%20%20%20%20x_current%20%3D%20x_init.copy()%0A%20%20%20%20trajectory%20%3D%20%5Bx_current.copy()%5D%0A%20%20%20%20for%20_%20in%20range(8)%3A%0A%20%20%20%20%20%20%20%20grad_val%20%3D%20grad_loss(x_current)%0A%20%20%20%20%20%20%20%20x_current%20%3D%20x_current%20-%20learning_rate%20*%20grad_val%0A%20%20%20%20%20%20%20%20trajectory.append(x_current.copy())%0A%20%20%20%20trajectory%20%3D%20np.array(trajectory)%0A%0A%20%20%20%20%23%20Orthogonal%20probe%20at%20starting%20point%0A%20%20%20%20probe_point%20%3D%20trajectory%5B0%5D%0A%20%20%20%20grad_probe%20%3D%20grad_loss(probe_point)%0A%20%20%20%20grad_norm%20%3D%20np.linalg.norm(grad_probe)%0A%20%20%20%20unit_grad%20%3D%20grad_probe%20%2F%20grad_norm%0A%20%20%20%20unit_tangent%20%3D%20np.array(%5B-unit_grad%5B1%5D%2C%20unit_grad%5B0%5D%5D)%0A%0A%20%20%20%20tangent_segment%20%3D%20np.vstack(%5Bprobe_point%20-%200.75%20*%20unit_tangent%2C%20probe_point%20%2B%200.75%20*%20unit_tangent%5D)%0A%20%20%20%20grad_segment%20%3D%20np.vstack(%5Bprobe_point%2C%20probe_point%20%2B%200.65%20*%20unit_grad%5D)%0A%20%20%20%20neg_grad_segment%20%3D%20np.vstack(%5Bprobe_point%2C%20probe_point%20-%200.65%20*%20unit_grad%5D)%0A%0A%20%20%20%20%23%203D%20Tangent%20Plane%20patch%20around%20probe_point%0A%20%20%20%20patch_x%20%3D%20np.linspace(probe_point%5B0%5D%20-%200.6%2C%20probe_point%5B0%5D%20%2B%200.6%2C%2015)%0A%20%20%20%20patch_y%20%3D%20np.linspace(probe_point%5B1%5D%20-%200.6%2C%20probe_point%5B1%5D%20%2B%200.6%2C%2015)%0A%20%20%20%20patch_x_mesh%2C%20patch_y_mesh%20%3D%20np.meshgrid(patch_x%2C%20patch_y)%0A%20%20%20%20z_probe%20%3D%20loss_func(probe_point%5B0%5D%2C%20probe_point%5B1%5D)%0A%20%20%20%20z_tangent%20%3D%20z_probe%20%2B%20grad_probe%5B0%5D%20*%20(patch_x_mesh%20-%20probe_point%5B0%5D)%20%2B%20grad_probe%5B1%5D%20*%20(%0A%20%20%20%20%20%20%20%20patch_y_mesh%20-%20probe_point%5B1%5D%0A%20%20%20%20)%0A%0A%20%20%20%20%23%203D%20trajectory%20z%20values%0A%20%20%20%20trajectory_z%20%3D%20%5Bloss_func(pt%5B0%5D%2C%20pt%5B1%5D)%20for%20pt%20in%20trajectory%5D%0A%0A%20%20%20%20fig%20%3D%20make_subplots(%0A%20%20%20%20%20%20%20%20rows%3D1%2C%0A%20%20%20%20%20%20%20%20cols%3D2%2C%0A%20%20%20%20%20%20%20%20specs%3D%5B%5B%7B%22type%22%3A%20%22xy%22%7D%2C%20%7B%22type%22%3A%20%22scene%22%7D%5D%5D%2C%0A%20%20%20%20%20%20%20%20subplot_titles%3D%5B%0A%20%20%20%20%20%20%20%20%20%20%20%20%222D%20Level%20Contours%20%26%20Gradient%20Orthogonality%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%223D%20Loss%20Surface%20Bowl%20%26%20Tangent%20Plane%22%2C%0A%20%20%20%20%20%20%20%20%5D%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%202D%20Panel%3A%20Contours%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Contour(%0A%20%20%20%20%20%20%20%20%20%20%20%20z%3Dz_loss%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dgrid_x%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dgrid_y%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20contours%3Ddict(showlines%3DTrue%2C%20start%3D-2.0%2C%20end%3D14.0%2C%20size%3D1.0)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20colorscale%3D%22Viridis%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20opacity%3D0.75%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20showscale%3DFalse%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Loss%20Contours%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%202D%20Panel%3A%20Contour%20Tangent%20Line%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dtangent_segment%5B%3A%2C%200%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dtangent_segment%5B%3A%2C%201%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%237c3aed%22%2C%20width%3D3%2C%20dash%3D%22dash%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Contour%20Tangent%20(f%3Dc)%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%202D%20Panel%3A%20Gradient%20(Ascent)%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dgrad_segment%5B%3A%2C%200%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dgrad_segment%5B%3A%2C%201%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%2Bmarkers%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%23ea580c%22%2C%20width%3D3.5)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20marker%3Ddict(size%3D7%2C%20color%3D%22%23ea580c%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22%E2%88%87f%20(Steepest%20Ascent)%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%202D%20Panel%3A%20Negative%20Gradient%20(Descent)%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dneg_grad_segment%5B%3A%2C%200%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dneg_grad_segment%5B%3A%2C%201%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%2Bmarkers%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%230284c7%22%2C%20width%3D3.5)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20marker%3Ddict(size%3D7%2C%20color%3D%22%230284c7%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22-%E2%88%87f%20(Steepest%20Descent)%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%202D%20Panel%3A%20GD%20Trajectory%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dtrajectory%5B%3A%2C%200%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dtrajectory%5B%3A%2C%201%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%2Bmarkers%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%23dc2626%22%2C%20width%3D2.5)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20marker%3Ddict(size%3D6%2C%20color%3D%22%23dc2626%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22GD%20Path%20(%CE%B7%3D0.25)%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%202D%20Panel%3A%20Minimum%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3D%5Bx_star%5B0%5D%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3D%5Bx_star%5B1%5D%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22markers%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20marker%3Ddict(size%3D13%2C%20color%3D%22%2316a34a%22%2C%20symbol%3D%22star%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Analytical%20Minimum%20x*%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%203D%20Panel%3A%20Loss%20Surface%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Surface(%0A%20%20%20%20%20%20%20%20%20%20%20%20z%3Dz_loss%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dgrid_x%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dgrid_y%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20colorscale%3D%22Viridis%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20opacity%3D0.82%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20showscale%3DFalse%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Loss%20Surface%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%203D%20Panel%3A%20Local%20Tangent%20Plane%20Patch%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Surface(%0A%20%20%20%20%20%20%20%20%20%20%20%20z%3Dz_tangent%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dpatch_x%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dpatch_y%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20colorscale%3D%5B%5B0%2C%20%22%23ea580c%22%5D%2C%20%5B1%2C%20%22%23ea580c%22%5D%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20opacity%3D0.6%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20showscale%3DFalse%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Tangent%20Plane%20at%20x%E2%82%80%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%203D%20Panel%3A%20GD%20Trajectory%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter3d(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dtrajectory%5B%3A%2C%200%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dtrajectory%5B%3A%2C%201%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20z%3Dtrajectory_z%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%2Bmarkers%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%23dc2626%22%2C%20width%3D5)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20marker%3Ddict(size%3D5%2C%20color%3D%22%23dc2626%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%223D%20Descent%20Path%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%203D%20Panel%3A%20Minimum%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter3d(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3D%5Bx_star%5B0%5D%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3D%5Bx_star%5B1%5D%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20z%3D%5Bloss_func(x_star%5B0%5D%2C%20x_star%5B1%5D)%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22markers%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20marker%3Ddict(size%3D8%2C%20color%3D%22%2316a34a%22%2C%20symbol%3D%22diamond%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%223D%20Minimum%20x*%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20fig.update_layout(%0A%20%20%20%20%20%20%20%20template%3D%22plotly_white%22%2C%0A%20%20%20%20%20%20%20%20height%3D540%2C%0A%20%20%20%20%20%20%20%20margin%3Ddict(l%3D30%2C%20r%3D30%2C%20t%3D50%2C%20b%3D30)%2C%0A%20%20%20%20%20%20%20%20xaxis%3Ddict(%0A%20%20%20%20%20%20%20%20%20%20%20%20title%3D%22x%E2%82%81%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20range%3D%5B-1.5%2C%202.5%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20zeroline%3DTrue%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20zerolinecolor%3D%22%23cbd5e1%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20gridcolor%3D%22%23f1f5f9%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20yaxis%3Ddict(%0A%20%20%20%20%20%20%20%20%20%20%20%20title%3D%22x%E2%82%82%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20range%3D%5B-1.5%2C%202.5%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20zeroline%3DTrue%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20zerolinecolor%3D%22%23cbd5e1%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20gridcolor%3D%22%23f1f5f9%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20scene%3Ddict(%0A%20%20%20%20%20%20%20%20%20%20%20%20xaxis%3Ddict(title%3D%22x%E2%82%81%22%2C%20gridcolor%3D%22%23f1f5f9%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20yaxis%3Ddict(title%3D%22x%E2%82%82%22%2C%20gridcolor%3D%22%23f1f5f9%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20zaxis%3Ddict(title%3D%22f(x)%22%2C%20gridcolor%3D%22%23f1f5f9%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20camera%3Ddict(eye%3Ddict(x%3D-1.5%2C%20y%3D-1.5%2C%20z%3D1.2))%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20legend%3Ddict(orientation%3D%22h%22%2C%20yanchor%3D%22bottom%22%2C%20y%3D-0.18%2C%20xanchor%3D%22center%22%2C%20x%3D0.5)%2C%0A%20%20%20%20)%0A%20%20%20%20return%20(fig%2C)%0A%0A%0A%40app.cell%0Adef%20_(fig%2C%20mo)%3A%0A%20%20%20%20mo.ui.plotly(fig)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%20%5Bd%5D%20Code%20Examples%0A%0A%20%20%20%20%23%23%23%20Example%201%3A%20Numerical%20and%20PyTorch%20Autograd%20Verification%20of%20the%205%20Core%20Rules%0A%0A%20%20%20%20In%20this%20example%2C%20we%20verify%20each%20of%20the%20five%20foundational%20matrix%20calculus%20rules%20by%20evaluating%3A%0A%20%20%20%201.%20The%20exact%20analytical%20expression%20derived%20via%20matrix%20calculus.%0A%20%20%20%202.%20The%20automatic%20differentiation%20gradient%20computed%20by%20PyTorch%20autograd%20(%60torch.autograd.functional.jacobian%60%20and%20%60.backward()%60).%0A%20%20%20%203.%20The%20maximum%20elementwise%20absolute%20difference%20%24%5C%7C%5Cnabla_%7B%5Ctext%7Banalytical%7D%7D%20-%20%5Cnabla_%7B%5Ctext%7Bautograd%7D%7D%5C%7C_%5Cinfty%24.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(torch)%3A%0A%20%20%20%20torch.manual_seed(47)%0A%0A%20%20%20%20%23%20Rule%201%3A%20Linear%20Form%20%E2%88%87_x%20(a%E1%B5%80%20x)%20%3D%20a%0A%20%20%20%20a_vec%20%3D%20torch.randn(3%2C%201)%0A%20%20%20%20x_vec1%20%3D%20torch.randn(3%2C%201%2C%20requires_grad%3DTrue)%0A%20%20%20%20f_linear%20%3D%20a_vec.T%20%40%20x_vec1%0A%20%20%20%20f_linear.backward()%0A%20%20%20%20grad_linear_autograd%20%3D%20x_vec1.grad%0A%20%20%20%20grad_linear_analytic%20%3D%20a_vec%0A%20%20%20%20err_rule1%20%3D%20float((grad_linear_autograd%20-%20grad_linear_analytic).abs().max().item())%0A%0A%20%20%20%20%23%20Rule%202%3A%20Jacobian%20of%20Linear%20Mapping%20J_x%20(A%20x)%20%3D%20A%0A%20%20%20%20a_matrix2%20%3D%20torch.randn(4%2C%203)%0A%20%20%20%20x_vec2%20%3D%20torch.randn(3%2C%201%2C%20requires_grad%3DTrue)%0A%0A%20%20%20%20def%20linear_mapping(x)%3A%0A%20%20%20%20%20%20%20%20return%20a_matrix2%20%40%20x%0A%0A%20%20%20%20jacobian_autograd%20%3D%20torch.autograd.functional.jacobian(linear_mapping%2C%20x_vec2).reshape(4%2C%203)%0A%20%20%20%20jacobian_analytic%20%3D%20a_matrix2%0A%20%20%20%20err_rule2%20%3D%20float((jacobian_autograd%20-%20jacobian_analytic).abs().max().item())%0A%0A%20%20%20%20%23%20Rule%203%3A%20Quadratic%20Form%20%E2%88%87_x%20(x%E1%B5%80%20A%20x)%20%3D%20(A%20%2B%20A%E1%B5%80)%20x%0A%20%20%20%20a_matrix3%20%3D%20torch.randn(3%2C%203)%0A%20%20%20%20x_vec3%20%3D%20torch.randn(3%2C%201%2C%20requires_grad%3DTrue)%0A%20%20%20%20f_quad%20%3D%20x_vec3.T%20%40%20a_matrix3%20%40%20x_vec3%0A%20%20%20%20f_quad.backward()%0A%20%20%20%20grad_quad_autograd%20%3D%20x_vec3.grad%0A%20%20%20%20grad_quad_analytic%20%3D%20(a_matrix3%20%2B%20a_matrix3.T)%20%40%20x_vec3%0A%20%20%20%20err_rule3%20%3D%20float((grad_quad_autograd%20-%20grad_quad_analytic).abs().max().item())%0A%0A%20%20%20%20%23%20Rule%204%3A%20Bilinear%20Form%20%E2%88%87_x%20(x%E1%B5%80%20A%20y)%20%3D%20A%20y%2C%20%E2%88%87_y%20(x%E1%B5%80%20A%20y)%20%3D%20A%E1%B5%80%20x%0A%20%20%20%20a_matrix4%20%3D%20torch.randn(3%2C%203)%0A%20%20%20%20x_vec4%20%3D%20torch.randn(3%2C%201%2C%20requires_grad%3DTrue)%0A%20%20%20%20y_vec4%20%3D%20torch.randn(3%2C%201%2C%20requires_grad%3DTrue)%0A%20%20%20%20f_bilinear%20%3D%20x_vec4.T%20%40%20a_matrix4%20%40%20y_vec4%0A%20%20%20%20f_bilinear.backward()%0A%20%20%20%20err_rule4_x%20%3D%20float((x_vec4.grad%20-%20a_matrix4%20%40%20y_vec4).abs().max().item())%0A%20%20%20%20err_rule4_y%20%3D%20float((y_vec4.grad%20-%20a_matrix4.T%20%40%20x_vec4).abs().max().item())%0A%20%20%20%20err_rule4%20%3D%20max(err_rule4_x%2C%20err_rule4_y)%0A%0A%20%20%20%20%23%20Rule%205%3A%20Matrix%20Derivative%20%E2%88%87_X%20(a%E1%B5%80%20X%20b)%20%3D%20a%20b%E1%B5%80%0A%20%20%20%20a_vec5%20%3D%20torch.randn(4%2C%201)%0A%20%20%20%20b_vec5%20%3D%20torch.randn(3%2C%201)%0A%20%20%20%20x_mat5%20%3D%20torch.randn(4%2C%203%2C%20requires_grad%3DTrue)%0A%20%20%20%20f_mat%20%3D%20a_vec5.T%20%40%20x_mat5%20%40%20b_vec5%0A%20%20%20%20f_mat.backward()%0A%20%20%20%20grad_mat_autograd%20%3D%20x_mat5.grad%0A%20%20%20%20grad_mat_analytic%20%3D%20a_vec5%20%40%20b_vec5.T%0A%20%20%20%20err_rule5%20%3D%20float((grad_mat_autograd%20-%20grad_mat_analytic).abs().max().item())%0A%0A%20%20%20%20rules_summary%20%3D%20%7B%0A%20%20%20%20%20%20%20%20%22Rule%22%3A%20%5B%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Rule%201%3A%20Linear%20Form%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Rule%202%3A%20Linear%20Mapping%20(Jacobian)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Rule%203%3A%20Quadratic%20Form%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Rule%204%3A%20Bilinear%20Form%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Rule%205%3A%20Weight%20Matrix%20Derivative%22%2C%0A%20%20%20%20%20%20%20%20%5D%2C%0A%20%20%20%20%20%20%20%20%22Mathematical%20Expression%22%3A%20%5B%0A%20%20%20%20%20%20%20%20%20%20%20%20%22f(x)%20%3D%20a%E1%B5%80%20x%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22f(x)%20%3D%20A%20x%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22f(x)%20%3D%20x%E1%B5%80%20A%20x%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22f(x%2C%20y)%20%3D%20x%E1%B5%80%20A%20y%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22f(X)%20%3D%20a%E1%B5%80%20X%20b%22%2C%0A%20%20%20%20%20%20%20%20%5D%2C%0A%20%20%20%20%20%20%20%20%22Analytical%20Gradient%20Formula%22%3A%20%5B%0A%20%20%20%20%20%20%20%20%20%20%20%20%22%E2%88%87_x%20f%20%3D%20a%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22J_x(f)%20%3D%20A%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22%E2%88%87_x%20f%20%3D%20(A%20%2B%20A%E1%B5%80)%20x%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22%E2%88%87_x%20f%20%3D%20A%20y%2C%20%E2%88%87_y%20f%20%3D%20A%E1%B5%80%20x%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22%E2%88%87_X%20f%20%3D%20a%20b%E1%B5%80%22%2C%0A%20%20%20%20%20%20%20%20%5D%2C%0A%20%20%20%20%20%20%20%20%22Max%20Autograd%20Discrepancy%22%3A%20%5B%0A%20%20%20%20%20%20%20%20%20%20%20%20f%22%7Berr_rule1%3A.2e%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20f%22%7Berr_rule2%3A.2e%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20f%22%7Berr_rule3%3A.2e%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20f%22%7Berr_rule4%3A.2e%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20f%22%7Berr_rule5%3A.2e%7D%22%2C%0A%20%20%20%20%20%20%20%20%5D%2C%0A%20%20%20%20%20%20%20%20%22Verification%20Status%22%3A%20%5B%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Passed%20(%3C%201e-7)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Passed%20(%3C%201e-7)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Passed%20(%3C%201e-7)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Passed%20(%3C%201e-7)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Passed%20(%3C%201e-7)%22%2C%0A%20%20%20%20%20%20%20%20%5D%2C%0A%20%20%20%20%7D%0A%20%20%20%20return%20(rules_summary%2C)%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo%2C%20pd%2C%20rules_summary)%3A%0A%20%20%20%20df_rules%20%3D%20pd.DataFrame(rules_summary)%0A%20%20%20%20mo.ui.table(df_rules)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%23%20Example%202%3A%20Training%20Linear%20Regression%20with%20Analytical%20Matrix%20Calculus%20vs%20PyTorch%20Autograd%0A%0A%20%20%20%20In%20this%20example%2C%20we%20generate%20a%20synthetic%20regression%20dataset%20with%20%24N%20%3D%2080%24%20samples%20and%20%24d%20%3D%203%24%20features%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cmathbf%7By%7D%20%3D%20%5Cmathbf%7BX%7D%20%5Cmathbf%7Bw%7D_%7B%5Ctext%7Btrue%7D%7D%20%2B%20%5Cboldsymbol%7B%5Cepsilon%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20We%20simultaneously%20optimize%20the%20parameters%20across%2010%20gradient%20descent%20iterations%20using%3A%0A%20%20%20%201.%20**Analytical%20Matrix%20Calculus%20Gradient**%3A%20%24%5Cnabla_%7B%5Cmathbf%7Bw%7D%7D%20%5Cmathcal%7BL%7D%20%3D%20%5Cfrac%7B2%7D%7BN%7D%20%5Cmathbf%7BX%7D%5ET%20(%5Cmathbf%7BX%7D%5Cmathbf%7Bw%7D%20-%20%5Cmathbf%7By%7D)%24%0A%20%20%20%202.%20**PyTorch%20Automatic%20Differentiation**%3A%20%60loss.backward()%60%0A%20%20%20%203.%20**Normal%20Equations%20Benchmark**%3A%20%24%5Cmathbf%7Bw%7D%5E*%20%3D%20(%5Cmathbf%7BX%7D%5ET%20%5Cmathbf%7BX%7D)%5E%7B-1%7D%20%5Cmathbf%7BX%7D%5ET%20%5Cmathbf%7By%7D%24%0A%0A%20%20%20%20The%20table%20below%20records%20the%20step-by-step%20equivalence%20between%20the%20analytical%20and%20automatic%20gradients%2C%20as%20well%20as%20the%20monotonic%20convergence%20toward%20the%20closed-form%20OLS%20solution.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(np%2C%20torch)%3A%0A%20%20%20%20rng%20%3D%20np.random.default_rng(42)%0A%20%20%20%20n_samples%2C%20n_features%20%3D%2080%2C%203%0A%0A%20%20%20%20%23%20Synthetic%20design%20matrix%20and%20true%20weights%0A%20%20%20%20x_data%20%3D%20rng.standard_normal((n_samples%2C%20n_features))%0A%20%20%20%20true_weights%20%3D%20np.array(%5B1.8%2C%20-2.2%2C%200.75%5D)%0A%20%20%20%20noise%20%3D%20rng.normal(0%2C%200.05%2C%20n_samples)%0A%20%20%20%20y_data%20%3D%20x_data%20%40%20true_weights%20%2B%20noise%0A%0A%20%20%20%20%23%20Closed-form%20Ordinary%20Least%20Squares%20solution%20via%20Normal%20Equations%0A%20%20%20%20w_ols%20%3D%20np.linalg.solve(x_data.T%20%40%20x_data%2C%20x_data.T%20%40%20y_data)%0A%0A%20%20%20%20%23%20Initialize%20analytical%20weights%0A%20%20%20%20w_analytic%20%3D%20np.zeros(n_features)%0A%20%20%20%20step_size%20%3D%200.08%0A%0A%20%20%20%20%23%20Initialize%20PyTorch%20tensors%20with%20float64%20precision%0A%20%20%20%20x_tensor%20%3D%20torch.tensor(x_data%2C%20dtype%3Dtorch.float64)%0A%20%20%20%20y_tensor%20%3D%20torch.tensor(y_data%2C%20dtype%3Dtorch.float64)%0A%20%20%20%20w_tensor%20%3D%20torch.zeros(n_features%2C%20dtype%3Dtorch.float64%2C%20requires_grad%3DTrue)%0A%0A%20%20%20%20optimization_records%20%3D%20%5B%5D%0A%0A%20%20%20%20for%20step%20in%20range(1%2C%2011)%3A%0A%20%20%20%20%20%20%20%20%23%201.%20Analytical%20gradient%20update%0A%20%20%20%20%20%20%20%20residuals_analytic%20%3D%20x_data%20%40%20w_analytic%20-%20y_data%0A%20%20%20%20%20%20%20%20mse_loss_analytic%20%3D%20float(np.mean(residuals_analytic**2))%0A%20%20%20%20%20%20%20%20grad_analytic%20%3D%20(2.0%20%2F%20n_samples)%20*%20(x_data.T%20%40%20residuals_analytic)%0A%20%20%20%20%20%20%20%20w_analytic%20%3D%20w_analytic%20-%20step_size%20*%20grad_analytic%0A%0A%20%20%20%20%20%20%20%20%23%202.%20PyTorch%20autograd%20update%0A%20%20%20%20%20%20%20%20residuals_torch%20%3D%20x_tensor%20%40%20w_tensor%20-%20y_tensor%0A%20%20%20%20%20%20%20%20mse_loss_torch%20%3D%20(residuals_torch**2).mean()%0A%20%20%20%20%20%20%20%20mse_loss_torch.backward()%0A%20%20%20%20%20%20%20%20with%20torch.no_grad()%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20w_tensor%20-%3D%20step_size%20*%20w_tensor.grad%0A%20%20%20%20%20%20%20%20%20%20%20%20w_tensor.grad.zero_()%0A%0A%20%20%20%20%20%20%20%20%23%20Compare%20analytical%20weights%20vs%20PyTorch%20autograd%20weights%0A%20%20%20%20%20%20%20%20w_torch_np%20%3D%20w_tensor.detach().numpy()%0A%20%20%20%20%20%20%20%20weight_discrepancy%20%3D%20float(np.max(np.abs(w_analytic%20-%20w_torch_np)))%0A%20%20%20%20%20%20%20%20dist_to_ols%20%3D%20float(np.linalg.norm(w_analytic%20-%20w_ols))%0A%0A%20%20%20%20%20%20%20%20optimization_records.append(%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Step%22%3A%20step%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22MSE%20Loss%22%3A%20f%22%7Bmse_loss_analytic%3A.4f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Analytic%20w%22%3A%20f%22%5B%7Bw_analytic%5B0%5D%3A.3f%7D%2C%20%7Bw_analytic%5B1%5D%3A.3f%7D%2C%20%7Bw_analytic%5B2%5D%3A.3f%7D%5D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Autograd%20w%22%3A%20f%22%5B%7Bw_torch_np%5B0%5D%3A.3f%7D%2C%20%7Bw_torch_np%5B1%5D%3A.3f%7D%2C%20%7Bw_torch_np%5B2%5D%3A.3f%7D%5D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Max%20Discrepancy%22%3A%20f%22%7Bweight_discrepancy%3A.2e%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Distance%20to%20OLS%20w*%22%3A%20f%22%7Bdist_to_ols%3A.4f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7D%0A%20%20%20%20%20%20%20%20)%0A%20%20%20%20return%20(optimization_records%2C)%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo%2C%20optimization_records%2C%20pd)%3A%0A%20%20%20%20df_optimization%20%3D%20pd.DataFrame(optimization_records)%0A%20%20%20%20mo.ui.table(df_optimization)%0A%20%20%20%20return%0A%0A%0Aif%20__name__%20%3D%3D%20%22__main__%22%3A%0A%20%20%20%20app.run()%0A
0e61f896744beb56cdef82c0ebbafdd1