import%20marimo%0A%0A__generated_with%20%3D%20%220.24.0%22%0Aapp%20%3D%20marimo.App(width%3D%22medium%22)%0A%0A%0A%40app.cell%0Adef%20_()%3A%0A%20%20%20%20import%20marimo%20as%20mo%0A%20%20%20%20import%20numpy%20as%20np%0A%20%20%20%20import%20pandas%20as%20pd%0A%20%20%20%20import%20plotly.graph_objects%20as%20go%0A%20%20%20%20from%20plotly.subplots%20import%20make_subplots%0A%0A%20%20%20%20return%20go%2C%20make_subplots%2C%20mo%2C%20np%2C%20pd%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%20Note%2013%3A%20Unbiased%20vs%20Consistent%20Estimators%20and%20the%20Bias-Variance-Consistency%20Trade-Off%0A%0A%20%20%20%20%26larr%3B%20Previous%20Note%3A%20%5B12%20Multivariate%20Normal%20Distribution%5D(12_multivariate_normal_distribution.py)%20%7C%20Next%20Note%3A%20%5B14%20Distribution%20of%20Minimum%5D(14_dist_of_minimum.py)%20%26rarr%3B%0A%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%20%5Ba%5D%20Why%20do%20you%20need%20to%20know%20these%20concepts%3F%0A%0A%20%20%20%20A%20fundamental%20goal%20in%20statistical%20inference%20and%20machine%20learning%20is%20estimating%20unknown%20population%20parameters%20%24%5Ctheta%24%20(such%20as%20weights%2C%20variance%2C%20or%20probabilities)%20from%20a%20finite%20dataset%20of%20observations.%20But%20how%20do%20we%20define%20what%20makes%20an%20estimator%20%24%5Chat%7B%5Ctheta%7D%24%20mathematically%20%22good%22%3F%0A%0A%20%20%20%20Two%20foundational%20properties%20govern%20estimator%20quality%3A%0A%20%20%20%20*%20**Unbiasedness**%20is%20a%20**finite-sample%20property**%3A%20Across%20repeated%20realizations%20of%20a%20fixed%20sample%20size%20%24n%24%2C%20the%20expected%20value%20of%20the%20estimator%20exactly%20matches%20the%20true%20parameter%20(%24%5Cmathbb%7BE%7D%5B%5Chat%7B%5Ctheta%7D_n%5D%20%3D%20%5Ctheta%24).%20It%20does%20not%20systematically%20overshoot%20or%20undershoot.%0A%20%20%20%20*%20**Consistency**%20is%20an%20**asymptotic%20property**%3A%20As%20the%20sample%20size%20grows%20toward%20infinity%20(%24n%20%5Cto%20%5Cinfty%24)%2C%20the%20probability%20that%20the%20estimate%20differs%20from%20the%20true%20parameter%20by%20any%20margin%20%24%5Cepsilon%20%3E%200%24%20collapses%20to%20zero%20(%24%5Chat%7B%5Ctheta%7D_n%20%5Cxrightarrow%7BP%7D%20%5Ctheta%24).%0A%0A%20%20%20%20Why%20this%20distinction%20is%20vital%20in%20machine%20learning%20and%20data%20science%3A%0A%20%20%20%201.%20**The%20Modern%20ML%20Reality%3A%20Biased%20but%20Consistent**%3A%20In%20contemporary%20machine%20learning%2C%20almost%20all%20effective%20estimators%20are%20deliberately%20biased%20to%20reduce%20variance.%20Examples%20include%20Ridge%20regression%20(%24%5Chat%7B%5Cboldsymbol%7B%5Cbeta%7D%7D_%7B%5Ctext%7Bridge%7D%7D%24)%2C%20weight%20decay%2C%20Lasso%2C%20and%20Maximum%20Likelihood%20variance%20(%24%5Chat%7B%5Csigma%7D%5E2_%7B%5Ctext%7BMLE%7D%7D%24).%20While%20biased%20for%20small%20%24n%24%2C%20they%20are%20consistent%20as%20%24n%20%5Cto%20%5Cinfty%24%20and%20dramatically%20outperform%20unbiased%20estimators%20in%20Mean%20Squared%20Error%20(MSE).%0A%20%20%20%202.%20**The%20Danger%20of%20Inconsistent%20Unbiased%20Estimators**%3A%20An%20estimator%20can%20be%20perfectly%20unbiased%20for%20every%20sample%20size%20%24n%24%2C%20yet%20completely%20useless%20because%20its%20variance%20never%20shrinks%20(e.g.%2C%20using%20only%20the%20first%20sample%20observation%20%24X_1%24%20to%20estimate%20the%20population%20mean).%20No%20matter%20how%20much%20data%20is%20collected%2C%20it%20never%20gets%20closer%20to%20the%20truth.%0A%20%20%20%203.%20**Mean%20Squared%20Error%20Decomposition**%3A%20%24%5Ctext%7BMSE%7D(%5Chat%7B%5Ctheta%7D)%20%3D%20%5Ctext%7BBias%7D%5E2(%5Chat%7B%5Ctheta%7D)%20%2B%20%5Ctext%7BVar%7D(%5Chat%7B%5Ctheta%7D)%24.%20A%20sufficient%20condition%20for%20consistency%20is%20that%20both%20the%20squared%20bias%20and%20the%20variance%20shrink%20to%20zero%20as%20%24n%20%5Cto%20%5Cinfty%24.%0A%20%20%20%204.%20**Bessel's%20Correction**%3A%20The%20Maximum%20Likelihood%20estimator%20for%20variance%20divides%20by%20%24n%24%2C%20introducing%20a%20small%20bias%20of%20%24-%5Csigma%5E2%2Fn%24%20that%20disappears%20asymptotically.%20In%20contrast%2C%20dividing%20by%20%24n-1%24%20achieves%20exact%20finite-sample%20unbiasedness.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%20%5Bb%5D%20Concept%20explanation%20with%20their%20role%20in%20ML%2FAI%2FStats%3F%0A%0A%20%20%20%20%23%23%23%20Formal%20Definition%3A%20Unbiased%20Estimator%0A%0A%20%20%20%20Let%20%24X_1%2C%20%5Cdots%2C%20X_n%24%20be%20an%20i.i.d.%20sample%20from%20a%20distribution%20parameterized%20by%20%24%5Ctheta%24.%20An%20estimator%20%24%5Chat%7B%5Ctheta%7D_n%20%3D%20g(X_1%2C%20%5Cdots%2C%20X_n)%24%20is%20**unbiased**%20if%20its%20mathematical%20expectation%20equals%20%24%5Ctheta%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cmathbb%7BE%7D%5B%5Chat%7B%5Ctheta%7D_n%5D%20%3D%20%5Ctheta%20%5Cquad%20%5Ctext%7Bfor%20all%20%7D%20n%20%5Cgeq%201%0A%20%20%20%20%24%24%0A%0A%20%20%20%20The%20**bias**%20of%20an%20estimator%20is%20defined%20as%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Ctext%7BBias%7D(%5Chat%7B%5Ctheta%7D_n)%20%3D%20%5Cmathbb%7BE%7D%5B%5Chat%7B%5Ctheta%7D_n%5D%20-%20%5Ctheta%0A%20%20%20%20%24%24%0A%0A%20%20%20%20An%20estimator%20is%20unbiased%20if%20and%20only%20if%20%24%5Ctext%7BBias%7D(%5Chat%7B%5Ctheta%7D_n)%20%3D%200%24.%0A%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%23%20Formal%20Definition%3A%20Consistent%20Estimator%0A%0A%20%20%20%20An%20estimator%20%24%5Chat%7B%5Ctheta%7D_n%24%20is%20**consistent**%20for%20%24%5Ctheta%24%20if%20it%20converges%20in%20probability%20to%20%24%5Ctheta%24%20as%20sample%20size%20%24n%20%5Cto%20%5Cinfty%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Clim_%7Bn%20%5Cto%20%5Cinfty%7D%20P(%7C%5Chat%7B%5Ctheta%7D_n%20-%20%5Ctheta%7C%20%3C%20%5Cepsilon)%20%3D%201%20%5Cquad%20%5Ctext%7Bfor%20every%20%7D%20%5Cepsilon%20%3E%200%0A%20%20%20%20%24%24%0A%0A%20%20%20%20This%20is%20written%20compactly%20as%20%24%5Chat%7B%5Ctheta%7D_n%20%5Cxrightarrow%7BP%7D%20%5Ctheta%24.%0A%0A%20%20%20%20%23%23%23%23%20Mean%20Squared%20Error%20and%20Consistency%20Criterion%0A%20%20%20%20The%20Mean%20Squared%20Error%20(MSE)%20of%20an%20estimator%20decomposes%20into%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Ctext%7BMSE%7D(%5Chat%7B%5Ctheta%7D_n)%20%3D%20%5Cmathbb%7BE%7D%5B(%5Chat%7B%5Ctheta%7D_n%20-%20%5Ctheta)%5E2%5D%20%3D%20%5Ctext%7BBias%7D%5E2(%5Chat%7B%5Ctheta%7D_n)%20%2B%20%5Ctext%7BVar%7D(%5Chat%7B%5Ctheta%7D_n)%0A%20%20%20%20%24%24%0A%0A%20%20%20%20By%20Chebyshev's%20inequality%20(Note%2010)%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20P(%7C%5Chat%7B%5Ctheta%7D_n%20-%20%5Ctheta%7C%20%5Cgeq%20%5Cepsilon)%20%5Cleq%20%5Cfrac%7B%5Cmathbb%7BE%7D%5B(%5Chat%7B%5Ctheta%7D_n%20-%20%5Ctheta)%5E2%5D%7D%7B%5Cepsilon%5E2%7D%20%3D%20%5Cfrac%7B%5Ctext%7BMSE%7D(%5Chat%7B%5Ctheta%7D_n)%7D%7B%5Cepsilon%5E2%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20Therefore%2C%20if%20%24%5Clim_%7Bn%20%5Cto%20%5Cinfty%7D%20%5Ctext%7BMSE%7D(%5Chat%7B%5Ctheta%7D_n)%20%3D%200%24%2C%20the%20estimator%20is%20guaranteed%20to%20be%20consistent.%20This%20holds%20whenever%20both%20the%20bias%20and%20variance%20vanish%20asymptotically%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Clim_%7Bn%20%5Cto%20%5Cinfty%7D%20%5Ctext%7BBias%7D(%5Chat%7B%5Ctheta%7D_n)%20%3D%200%20%5Cquad%20%5Ctext%7Band%7D%20%5Cquad%20%5Clim_%7Bn%20%5Cto%20%5Cinfty%7D%20%5Ctext%7BVar%7D(%5Chat%7B%5Ctheta%7D_n)%20%3D%200%20%5Cimplies%20%5Chat%7B%5Ctheta%7D_n%20%5Cxrightarrow%7BP%7D%20%5Ctheta%0A%20%20%20%20%24%24%0A%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%23%20The%20Four%20Canonical%20Estimator%20Archetypes%0A%0A%20%20%20%20To%20build%20intuition%2C%20consider%20estimating%20the%20population%20mean%20%24%5Cmu%24%20from%20an%20i.i.d.%20sample%20%24X_1%2C%20%5Cdots%2C%20X_n%20%5Csim%20%5Cmathcal%7BN%7D(%5Cmu%2C%20%5Csigma%5E2)%24%3A%0A%0A%20%20%20%20%23%23%23%23%201.%20Unbiased%20and%20Consistent%3A%20The%20Sample%20Mean%0A%20%20%20%20%24%24%0A%20%20%20%20T_1%20%3D%20%5Cbar%7BX%7D_n%20%3D%20%5Cfrac%7B1%7D%7Bn%7D%20%5Csum_%7Bi%3D1%7D%5En%20X_i%0A%20%20%20%20%24%24%0A%20%20%20%20*%20Expectation%3A%20%24%5Cmathbb%7BE%7D%5BT_1%5D%20%3D%20%5Cmu%20%5Cimplies%20%5Ctext%7BBias%7D%20%3D%200%24%20(Unbiased).%0A%20%20%20%20*%20Variance%3A%20%24%5Ctext%7BVar%7D(T_1)%20%3D%20%5Cfrac%7B%5Csigma%5E2%7D%7Bn%7D%20%5Cto%200%24%20as%20%24n%20%5Cto%20%5Cinfty%24%20(Consistent).%0A%0A%20%20%20%20%23%23%23%23%202.%20Biased%20but%20Consistent%3A%20Shrinkage%20Estimator%0A%20%20%20%20%24%24%0A%20%20%20%20T_2%20%3D%20%5Cfrac%7Bn%7D%7Bn%20%2B%201%7D%20%5Cbar%7BX%7D_n%0A%20%20%20%20%24%24%0A%20%20%20%20*%20Expectation%3A%20%24%5Cmathbb%7BE%7D%5BT_2%5D%20%3D%20%5Cfrac%7Bn%7D%7Bn%20%2B%201%7D%20%5Cmu%20%5Cneq%20%5Cmu%20%5Cimplies%20%5Ctext%7BBias%7D%20%3D%20-%5Cfrac%7B%5Cmu%7D%7Bn%20%2B%201%7D%24%20(Biased%20for%20every%20finite%20%24n%24).%0A%20%20%20%20*%20Asymptotics%3A%20%24%5Clim_%7Bn%20%5Cto%20%5Cinfty%7D%20%5Ctext%7BBias%7D(T_2)%20%3D%200%24%20and%20%24%5Clim_%7Bn%20%5Cto%20%5Cinfty%7D%20%5Ctext%7BVar%7D(T_2)%20%3D%20%5Clim_%7Bn%20%5Cto%20%5Cinfty%7D%20%5Cleft(%5Cfrac%7Bn%7D%7Bn%2B1%7D%5Cright)%5E2%20%5Cfrac%7B%5Csigma%5E2%7D%7Bn%7D%20%3D%200%24%20(Consistent).%0A%0A%20%20%20%20%23%23%23%23%203.%20Unbiased%20but%20Inconsistent%3A%20Single-Sample%20Estimator%0A%20%20%20%20%24%24%0A%20%20%20%20T_3%20%3D%20X_1%0A%20%20%20%20%24%24%0A%20%20%20%20*%20Expectation%3A%20%24%5Cmathbb%7BE%7D%5BT_3%5D%20%3D%20%5Cmathbb%7BE%7D%5BX_1%5D%20%3D%20%5Cmu%20%5Cimplies%20%5Ctext%7BBias%7D%20%3D%200%24%20(Unbiased).%0A%20%20%20%20*%20Variance%3A%20%24%5Ctext%7BVar%7D(T_3)%20%3D%20%5Csigma%5E2%20%5Cneq%200%24%20(Does%20not%20shrink%20with%20%24n%24).%20Because%20variance%20remains%20constant%2C%20%24T_3%24%20never%20converges%20to%20%24%5Cmu%24%20(Inconsistent).%0A%0A%20%20%20%20%23%23%23%23%204.%20Biased%20and%20Inconsistent%3A%20Shifted%20Single-Sample%20Estimator%0A%20%20%20%20%24%24%0A%20%20%20%20T_4%20%3D%20X_1%20%2B%201.0%0A%20%20%20%20%24%24%0A%20%20%20%20*%20Expectation%3A%20%24%5Cmathbb%7BE%7D%5BT_4%5D%20%3D%20%5Cmu%20%2B%201.0%20%5Cimplies%20%5Ctext%7BBias%7D%20%3D%201.0%20%5Cneq%200%24%20(Biased).%0A%20%20%20%20*%20Variance%3A%20%24%5Ctext%7BVar%7D(T_4)%20%3D%20%5Csigma%5E2%20%5Cneq%200%24%20(Inconsistent).%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%20%5Bc%5D%20Interactive%20Visualizations%3A%20Finite-Sample%20Sampling%20Distributions%20vs%20Asymptotic%20Trajectories%0A%0A%20%20%20%20The%20interactive%20subplots%20below%20display%20the%20dual%20perspectives%20of%20estimation%20theory%3A%0A%20%20%20%20*%20**Left%20Panel**%3A%20Sampling%20distributions%20at%20a%20fixed%20small%20sample%20size%20(%24n%20%3D%205%24)%20across%20%241%2C000%24%20Monte%20Carlo%20trials.%20Notice%20that%20%24T_1%24%20(Sample%20Mean)%20and%20%24T_3%24%20(Single%20Sample%20%24X_1%24)%20are%20both%20centered%20exactly%20at%20the%20true%20mean%20%24%5Cmu%20%3D%205.0%24%20(Unbiased)%2C%20but%20%24T_3%24%20has%20wide%20spread.%20In%20contrast%2C%20%24T_2%24%20(Shrinkage)%20is%20shifted%20to%20the%20left%20at%20%24%5Cfrac%7B5%7D%7B6%7D%5Cmu%20%5Capprox%204.17%24%20(Biased).%0A%20%20%20%20*%20**Right%20Panel**%3A%20Real-time%20convergence%20trajectories%20as%20sample%20size%20%24n%24%20increases%20from%20%241%24%20to%20%24300%24.%20Notice%20how%20both%20%24T_1(n)%24%20and%20%24T_2(n)%24%20rapidly%20converge%20to%20the%20true%20parameter%20line%20%24%5Cmu%20%3D%205.0%24%20(Consistent)%2C%20while%20%24T_3%24%20continues%20to%20scatter%20with%20constant%20dispersion%20(Inconsistent).%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(go%2C%20make_subplots%2C%20np)%3A%0A%20%20%20%20rng_sim%20%3D%20np.random.default_rng(42)%0A%20%20%20%20true_param_mu%20%3D%205.0%0A%20%20%20%20true_param_sigma%20%3D%201.2%0A%20%20%20%20n_monte_carlo%20%3D%201000%0A%0A%20%20%20%20%23%201.%20Left%20Panel%3A%20Fixed%20small%20sample%20size%20n%20%3D%205%0A%20%20%20%20n_fixed%20%3D%205%0A%20%20%20%20batch_samples_n5%20%3D%20rng_sim.normal(true_param_mu%2C%20true_param_sigma%2C%20size%3D(n_monte_carlo%2C%20n_fixed))%0A%20%20%20%20t1_fixed%20%3D%20np.mean(batch_samples_n5%2C%20axis%3D1)%0A%20%20%20%20t2_fixed%20%3D%20(n_fixed%20%2F%20(n_fixed%20%2B%201.0))%20*%20t1_fixed%0A%20%20%20%20t3_fixed%20%3D%20batch_samples_n5%5B%3A%2C%200%5D%0A%0A%20%20%20%20%23%202.%20Right%20Panel%3A%20Dynamic%20sample%20size%20expansion%20n%20%3D%201%20to%20300%0A%20%20%20%20n_grid_sizes%20%3D%20np.arange(1%2C%20301)%0A%20%20%20%20stream_observations%20%3D%20rng_sim.normal(true_param_mu%2C%20true_param_sigma%2C%20size%3D300)%0A%20%20%20%20t1_trajectory%20%3D%20np.cumsum(stream_observations)%20%2F%20n_grid_sizes%0A%20%20%20%20t2_trajectory%20%3D%20(n_grid_sizes%20%2F%20(n_grid_sizes%20%2B%201.0))%20*%20t1_trajectory%0A%20%20%20%20t3_independent_draws%20%3D%20rng_sim.normal(true_param_mu%2C%20true_param_sigma%2C%20size%3D300)%0A%0A%20%20%20%20fig%20%3D%20make_subplots(%0A%20%20%20%20%20%20%20%20rows%3D1%2C%0A%20%20%20%20%20%20%20%20cols%3D2%2C%0A%20%20%20%20%20%20%20%20subplot_titles%3D%5B%0A%20%20%20%20%20%20%20%20%20%20%20%20f%22Sampling%20Distributions%20at%20Fixed%20n%3D%7Bn_fixed%7D%20(Unbiasedness%20Check)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Convergence%20Trajectories%20as%20n%20%E2%86%92%20300%20(Consistency%20Check)%22%2C%0A%20%20%20%20%20%20%20%20%5D%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Left%3A%20T1%20(Sample%20Mean)%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Histogram(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dt1_fixed%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20opacity%3D0.65%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20marker_color%3D%22%232563eb%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22T%E2%82%81%3A%20Sample%20Mean%20(Unbiased)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20nbinsx%3D35%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Left%3A%20T2%20(Shrinkage)%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Histogram(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dt2_fixed%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20opacity%3D0.65%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20marker_color%3D%22%23ea580c%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22T%E2%82%82%3A%20Shrinkage%20Mean%20(Biased)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20nbinsx%3D35%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Left%3A%20T3%20(Single%20observation%20X1)%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Histogram(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dt3_fixed%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20opacity%3D0.40%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20marker_color%3D%22%239333ea%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22T%E2%82%83%3A%20Single%20Observation%20X%E2%82%81%20(Unbiased)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20nbinsx%3D35%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Left%3A%20True%20mean%20line%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3D%5Btrue_param_mu%2C%20true_param_mu%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3D%5B0%2C%20160%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%23dc2626%22%2C%20width%3D2.5%2C%20dash%3D%22dash%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22True%20Parameter%20%CE%BC%20%3D%205.0%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Right%3A%20T1%20Trajectory%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dn_grid_sizes%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dt1_trajectory%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%232563eb%22%2C%20width%3D2.5)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22T%E2%82%81(n)%20Path%20(Consistent)%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Right%3A%20T2%20Trajectory%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dn_grid_sizes%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dt2_trajectory%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%23ea580c%22%2C%20width%3D2.5)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22T%E2%82%82(n)%20Path%20(Consistent)%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Right%3A%20T3%20Trajectory%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dn_grid_sizes%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dt3_independent_draws%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22markers%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20marker%3Ddict(size%3D4%2C%20color%3D%22%239333ea%22%2C%20opacity%3D0.45)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22T%E2%82%83(n)%20Draws%20(Inconsistent)%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Right%3A%20True%20mean%20line%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3D%5B1%2C%20300%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3D%5Btrue_param_mu%2C%20true_param_mu%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%23dc2626%22%2C%20width%3D2.5%2C%20dash%3D%22dash%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20showlegend%3DFalse%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20fig.update_layout(%0A%20%20%20%20%20%20%20%20template%3D%22plotly_white%22%2C%0A%20%20%20%20%20%20%20%20barmode%3D%22overlay%22%2C%0A%20%20%20%20%20%20%20%20height%3D520%2C%0A%20%20%20%20%20%20%20%20margin%3Ddict(l%3D40%2C%20r%3D40%2C%20t%3D60%2C%20b%3D40)%2C%0A%20%20%20%20%20%20%20%20xaxis%3Ddict(title%3D%22Estimate%20Value%22%2C%20range%3D%5B1.0%2C%209.0%5D%2C%20gridcolor%3D%22%23f1f5f9%22)%2C%0A%20%20%20%20%20%20%20%20yaxis%3Ddict(title%3D%22Frequency%22%2C%20gridcolor%3D%22%23f1f5f9%22)%2C%0A%20%20%20%20%20%20%20%20xaxis2%3Ddict(title%3D%22Sample%20Size%20(n)%22%2C%20range%3D%5B1%2C%20300%5D%2C%20gridcolor%3D%22%23f1f5f9%22)%2C%0A%20%20%20%20%20%20%20%20yaxis2%3Ddict(title%3D%22Estimate%20Value%22%2C%20range%3D%5B1.0%2C%209.0%5D%2C%20gridcolor%3D%22%23f1f5f9%22)%2C%0A%20%20%20%20%20%20%20%20legend%3Ddict(orientation%3D%22h%22%2C%20yanchor%3D%22bottom%22%2C%20y%3D-0.26%2C%20xanchor%3D%22center%22%2C%20x%3D0.5)%2C%0A%20%20%20%20)%0A%20%20%20%20return%20(fig%2C)%0A%0A%0A%40app.cell%0Adef%20_(fig%2C%20mo)%3A%0A%20%20%20%20mo.ui.plotly(fig)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%20%5Bd%5D%20Code%20Examples%0A%0A%20%20%20%20%23%23%23%20Example%201%3A%20Monte%20Carlo%20Verification%20of%20Bias%2C%20Variance%2C%20and%20MSE%20Across%20the%204%20Archetypes%0A%0A%20%20%20%20In%20this%20example%2C%20we%20evaluate%20all%20four%20archetype%20estimators%20across%20sample%20sizes%20%24n%20%5Cin%20%5C%7B5%2C%2020%2C%20100%2C%20500%2C%202000%5C%7D%24%20using%20%24M%20%3D%202%2C000%24%20Monte%20Carlo%20simulations%20per%20sample%20size.%0A%0A%20%20%20%20For%20each%20estimator%2C%20we%20calculate%3A%0A%20%20%20%201.%20**Empirical%20Bias**%3A%20%24%5Chat%7B%5Cmathbb%7BE%7D%7D%5B%5Chat%7B%5Ctheta%7D%5D%20-%20%5Ctheta%24%0A%20%20%20%202.%20**Empirical%20Variance**%3A%20%24%5Cwidehat%7B%5Ctext%7BVar%7D%7D(%5Chat%7B%5Ctheta%7D)%24%0A%20%20%20%203.%20**Mean%20Squared%20Error**%3A%20%24%5Cwidehat%7B%5Ctext%7BMSE%7D%7D(%5Chat%7B%5Ctheta%7D)%20%3D%20%5Ctext%7BBias%7D%5E2%20%2B%20%5Ctext%7BVar%7D%24%0A%20%20%20%204.%20**Consistency%20Verification**%3A%20Whether%20MSE%20strictly%20converges%20to%20%240%24%20as%20%24n%20%5Cto%20%5Cinfty%24.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(np)%3A%0A%20%20%20%20rng_ex1%20%3D%20np.random.default_rng(101)%0A%20%20%20%20true_mu_val%20%3D%205.0%0A%20%20%20%20true_sigma_val%20%3D%201.5%0A%20%20%20%20num_trials%20%3D%202000%0A%20%20%20%20sample_sizes_grid%20%3D%20%5B5%2C%2020%2C%20100%2C%20500%2C%202000%5D%0A%0A%20%20%20%20archetype_results_list%20%3D%20%5B%5D%0A%0A%20%20%20%20for%20sample_n%20in%20sample_sizes_grid%3A%0A%20%20%20%20%20%20%20%20batch_draws%20%3D%20rng_ex1.normal(true_mu_val%2C%20true_sigma_val%2C%20size%3D(num_trials%2C%20sample_n))%0A%0A%20%20%20%20%20%20%20%20t1_vals%20%3D%20np.mean(batch_draws%2C%20axis%3D1)%0A%20%20%20%20%20%20%20%20t2_vals%20%3D%20(sample_n%20%2F%20(sample_n%20%2B%201.0))%20*%20t1_vals%0A%20%20%20%20%20%20%20%20t3_vals%20%3D%20batch_draws%5B%3A%2C%200%5D%0A%20%20%20%20%20%20%20%20t4_vals%20%3D%20batch_draws%5B%3A%2C%200%5D%20%2B%201.0%0A%0A%20%20%20%20%20%20%20%20estimator_configs%20%3D%20%5B%0A%20%20%20%20%20%20%20%20%20%20%20%20(%22T%E2%82%81%3A%20Sample%20Mean%20X%CC%84%22%2C%20t1_vals%2C%20%22Unbiased%22%2C%20%22Consistent%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20(%22T%E2%82%82%3A%20Shrinkage%20Mean%20%5Bn%2F(n%2B1)%5DX%CC%84%22%2C%20t2_vals%2C%20%22Biased%22%2C%20%22Consistent%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20(%22T%E2%82%83%3A%20First%20Observation%20X%E2%82%81%22%2C%20t3_vals%2C%20%22Unbiased%22%2C%20%22Inconsistent%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20(%22T%E2%82%84%3A%20Shifted%20Observation%20X%E2%82%81%20%2B%201%22%2C%20t4_vals%2C%20%22Biased%22%2C%20%22Inconsistent%22)%2C%0A%20%20%20%20%20%20%20%20%5D%0A%0A%20%20%20%20%20%20%20%20for%20label%2C%20values%2C%20unb_class%2C%20cons_class%20in%20estimator_configs%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20emp_bias%20%3D%20float(np.mean(values)%20-%20true_mu_val)%0A%20%20%20%20%20%20%20%20%20%20%20%20emp_var%20%3D%20float(np.var(values))%0A%20%20%20%20%20%20%20%20%20%20%20%20emp_mse%20%3D%20float(np.mean((values%20-%20true_mu_val)%20**%202))%0A%0A%20%20%20%20%20%20%20%20%20%20%20%20archetype_results_list.append(%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Estimator%22%3A%20label%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Sample%20Size%20(n)%22%3A%20sample_n%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Empirical%20Bias%22%3A%20f%22%7Bemp_bias%3A%2B.4f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Empirical%20Variance%22%3A%20f%22%7Bemp_var%3A.4f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Empirical%20MSE%22%3A%20f%22%7Bemp_mse%3A.4f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Theoretical%20Class%22%3A%20f%22%7Bunb_class%7D%2C%20%7Bcons_class%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22MSE%20Shrinks%20to%200%3F%22%3A%20%22Yes%22%20if%20cons_class%20%3D%3D%20%22Consistent%22%20else%20%22No%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%7D%0A%20%20%20%20%20%20%20%20%20%20%20%20)%0A%20%20%20%20return%20(archetype_results_list%2C)%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(archetype_results_list%2C%20mo%2C%20pd)%3A%0A%20%20%20%20df_archetypes%20%3D%20pd.DataFrame(archetype_results_list)%0A%20%20%20%20mo.ui.table(df_archetypes)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%23%20Example%202%3A%20Bessel's%20Correction%20and%20Covariance%20Estimation%20in%20Machine%20Learning%0A%0A%20%20%20%20When%20estimating%20population%20variance%20%24%5Csigma%5E2%24%20from%20sample%20%24X_1%2C%20%5Cdots%2C%20X_n%24%2C%20two%20estimators%20compete%3A%0A%20%20%20%201.%20**Maximum%20Likelihood%20Estimator%20(%24%5Chat%7B%5Csigma%7D%5E2_%7B%5Ctext%7BMLE%7D%7D%24)**%3A%20Divides%20by%20%24n%24.%20Biased%20for%20finite%20%24n%24%20with%20expected%20value%20%24%5Cfrac%7Bn-1%7D%7Bn%7D%5Csigma%5E2%24%2C%20but%20asymptotically%20consistent.%0A%20%20%20%202.%20**Unbiased%20Sample%20Variance%20(%24S%5E2%24)**%3A%20Divides%20by%20%24n%20-%201%24%20(Bessel's%20correction).%20Exactly%20unbiased%20for%20all%20%24n%20%5Cgeq%202%24.%0A%0A%20%20%20%20Below%2C%20we%20simulate%20true%20variance%20%24%5Csigma%5E2%20%3D%204.0%24%20across%20small%20and%20large%20sample%20sizes%20%24n%20%5Cin%20%5C%7B2%2C%205%2C%2010%2C%2050%2C%20200%5C%7D%24%20over%20%24M%20%3D%205%2C000%24%20trials%2C%20measuring%20the%20shrinkage%20of%20MLE%20bias%20toward%20zero.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(np)%3A%0A%20%20%20%20rng_bessel%20%3D%20np.random.default_rng(202)%0A%20%20%20%20true_variance%20%3D%204.0%0A%20%20%20%20true_std%20%3D%20np.sqrt(true_variance)%0A%20%20%20%20bessel_sim_trials%20%3D%205000%0A%20%20%20%20bessel_sample_sizes%20%3D%20%5B2%2C%205%2C%2010%2C%2050%2C%20200%5D%0A%0A%20%20%20%20bessel_comparison_records%20%3D%20%5B%5D%0A%0A%20%20%20%20for%20n_count%20in%20bessel_sample_sizes%3A%0A%20%20%20%20%20%20%20%20trials_matrix%20%3D%20rng_bessel.normal(0.0%2C%20true_std%2C%20size%3D(bessel_sim_trials%2C%20n_count))%0A%20%20%20%20%20%20%20%20trials_means%20%3D%20np.mean(trials_matrix%2C%20axis%3D1%2C%20keepdims%3DTrue)%0A%20%20%20%20%20%20%20%20residuals_sq%20%3D%20np.sum((trials_matrix%20-%20trials_means)%20**%202%2C%20axis%3D1)%0A%0A%20%20%20%20%20%20%20%20%23%20MLE%20variance%20(divides%20by%20n)%0A%20%20%20%20%20%20%20%20var_mle_trials%20%3D%20residuals_sq%20%2F%20n_count%0A%20%20%20%20%20%20%20%20%23%20Unbiased%20variance%20(divides%20by%20n%20-%201)%0A%20%20%20%20%20%20%20%20var_unbiased_trials%20%3D%20residuals_sq%20%2F%20(n_count%20-%201)%0A%0A%20%20%20%20%20%20%20%20mean_mle%20%3D%20float(np.mean(var_mle_trials))%0A%20%20%20%20%20%20%20%20bias_mle%20%3D%20mean_mle%20-%20true_variance%0A%20%20%20%20%20%20%20%20mean_unbiased%20%3D%20float(np.mean(var_unbiased_trials))%0A%20%20%20%20%20%20%20%20bias_unbiased%20%3D%20mean_unbiased%20-%20true_variance%0A%0A%20%20%20%20%20%20%20%20pct_bias_mle%20%3D%20(bias_mle%20%2F%20true_variance)%20*%20100.0%0A%0A%20%20%20%20%20%20%20%20bessel_comparison_records.append(%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Sample%20Size%20(n)%22%3A%20n_count%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22True%20%CF%83%C2%B2%22%3A%20f%22%7Btrue_variance%3A.2f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22MLE%20E%5B%CF%83%CC%82%C2%B2%5D%20(Divides%20by%20n)%22%3A%20f%22%7Bmean_mle%3A.4f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22MLE%20Bias%22%3A%20f%22%7Bbias_mle%3A%2B.4f%7D%20(%7Bpct_bias_mle%3A%2B.1f%7D%25)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Unbiased%20E%5BS%C2%B2%5D%20(Divides%20by%20n-1)%22%3A%20f%22%7Bmean_unbiased%3A.4f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Unbiased%20Bias%22%3A%20f%22%7Bbias_unbiased%3A%2B.4f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Consistency%20Observed%22%3A%20%22Bias%20%E2%86%92%200%20as%20n%20increases%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7D%0A%20%20%20%20%20%20%20%20)%0A%20%20%20%20return%20(bessel_comparison_records%2C)%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(bessel_comparison_records%2C%20mo%2C%20pd)%3A%0A%20%20%20%20df_bessel%20%3D%20pd.DataFrame(bessel_comparison_records)%0A%20%20%20%20mo.ui.table(df_bessel)%0A%20%20%20%20return%0A%0A%0Aif%20__name__%20%3D%3D%20%22__main__%22%3A%0A%20%20%20%20app.run()%0A
cc6bd382292eab01448c688ce3463f2f