import%20marimo%0A%0A__generated_with%20%3D%20%220.24.0%22%0Aapp%20%3D%20marimo.App()%0A%0A%0A%40app.cell%0Adef%20_()%3A%0A%20%20%20%20import%20time%0A%20%20%20%20import%20marimo%20as%20mo%0A%20%20%20%20import%20numpy%20as%20np%0A%20%20%20%20import%20pandas%20as%20pd%0A%20%20%20%20import%20plotly.graph_objects%20as%20go%0A%20%20%20%20from%20plotly.subplots%20import%20make_subplots%0A%20%20%20%20from%20sklearn.datasets%20import%20load_breast_cancer%0A%20%20%20%20from%20sklearn.model_selection%20import%20train_test_split%0A%20%20%20%20from%20sklearn.tree%20import%20DecisionTreeClassifier%0A%0A%20%20%20%20return%20(%0A%20%20%20%20%20%20%20%20DecisionTreeClassifier%2C%0A%20%20%20%20%20%20%20%20go%2C%0A%20%20%20%20%20%20%20%20load_breast_cancer%2C%0A%20%20%20%20%20%20%20%20make_subplots%2C%0A%20%20%20%20%20%20%20%20mo%2C%0A%20%20%20%20%20%20%20%20np%2C%0A%20%20%20%20%20%20%20%20pd%2C%0A%20%20%20%20%20%20%20%20time%2C%0A%20%20%20%20%20%20%20%20train_test_split%2C%0A%20%20%20%20)%0A%0A%0A%40app.cell%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%5B%E2%86%90%2034%20Mahalanobis%20Distance%5D(34_mahalanobis_distance.py)%20%7C%20%5BIndex%5D(..%2Findex.html)%20%7C%20%5B36%20Agglomerative%20Clustering%20%E2%86%92%5D(36_agglomerative_clustering.py)%0A%0A%20%20%20%20%23%20Gini%20Impurity%20vs%20Entropy%3A%20Decision%20Tree%20Splitting%20Criteria%20and%20Information%20Gain%0A%0A%20%20%20%20%23%23%20%5Ba%5D%20Why%20do%20you%20need%20to%20know%20these%20concepts%3F%0A%0A%20%20%20%20Decision%20trees%2C%20Random%20Forests%2C%20and%20Gradient%20Boosted%20Trees%20(such%20as%20XGBoost%2C%20LightGBM%2C%20and%20CatBoost)%20construct%20predictive%20models%20by%20recursively%20partitioning%20feature%20space%20into%20rectangular%20regions.%20At%20every%20internal%20node%2C%20the%20learning%20algorithm%20solves%20a%20local%20optimization%20problem%3A%20it%20searches%20over%20all%20available%20features%20and%20candidate%20thresholds%20to%20find%20the%20partition%20that%20maximizes%20the%20reduction%20in%20node%20impurity.%0A%0A%20%20%20%20%23%23%23%23%20The%20Two%20Dominant%20Purity%20Metrics%0A%20%20%20%20In%20classification%20trees%2C%20two%20impurity%20criteria%20dominate%20both%20theoretical%20literature%20and%20production%20libraries%3A%0A%20%20%20%201.%20**Gini%20Impurity**%3A%20Pioneered%20by%20Breiman%20et%20al.%20in%20the%20CART%20(Classification%20and%20Regression%20Trees)%20framework%20and%20the%20default%20criterion%20in%20Scikit-Learn.%0A%20%20%20%202.%20**Shannon%20Entropy%20(Information%20Gain)**%3A%20Introduced%20by%20Quinlan%20in%20ID3%20and%20C4.5%2C%20rooted%20in%20Claude%20Shannon's%20mathematical%20theory%20of%20communication.%0A%0A%20%20%20%20%23%23%23%23%20The%20Computational%20and%20Geometric%20Trade-Off%0A%20%20%20%20While%20both%20metrics%20measure%20the%20dispersion%20of%20class%20labels%20within%20a%20partition%2C%20they%20present%20distinct%20computational%20characteristics%3A%0A%20%20%20%20-%20**Gini%20Impurity**%20relies%20purely%20on%20arithmetic%20sums%20of%20squares%20(%241%20-%20%5Csum%20p_k%5E2%24).%20Modern%20CPUs%20and%20SIMD%20registers%20execute%20these%20vector%20dot-products%20in%20single-cycle%20operations%2C%20avoiding%20expensive%20transcendental%20function%20calls.%0A%20%20%20%20-%20**Entropy**%20evaluates%20logarithmic%20functions%20(%24-%5Csum%20p_k%20%5Clog_2%20p_k%24).%20Evaluating%20logarithms%20requires%20polynomial%20series%20expansions%20or%20hardware%20lookups%2C%20which%20incurs%20measurable%20computational%20overhead%20when%20evaluating%20millions%20of%20split%20candidates%20across%20large%20datasets.%0A%0A%20%20%20%20%23%23%23%23%20Structural%20Differences%20in%20Learned%20Trees%0A%20%20%20%20Because%20Entropy%20scales%20to%20a%20maximum%20of%201.0%20(for%20binary%20tasks)%20while%20Gini%20reaches%200.5%2C%20Entropy%20has%20a%20steeper%20curvature%20away%20from%20the%20center.%20Consequently%2C%20Entropy%20penalizes%20mixed%20impurity%20more%20aggressively%2C%20frequently%20yielding%20slightly%20more%20balanced%20trees.%20However%2C%20as%20proven%20by%20theoretical%20analysis%2C%20Gini%20is%20a%20first-order%20Taylor%20series%20approximation%20of%20Shannon%20Entropy.%20In%20practice%2C%20the%20two%20criteria%20agree%20on%20the%20optimal%20split%20point%20over%2098%25%20of%20the%20time%2C%20leading%20to%20virtually%20identical%20classification%20accuracies.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20%5Bb%5D%20Mathematical%20Foundations%20and%20Splitting%20Mechanics%0A%0A%20%20%20%20%23%23%23%201.%20Mathematical%20Definitions%20for%20a%20Node%0A%0A%20%20%20%20Let%20a%20node%20%24m%24%20contain%20%24N_m%24%20training%20observations%20belonging%20to%20%24K%24%20distinct%20classes.%20Let%20%24N_%7Bmk%7D%24%20denote%20the%20number%20of%20observations%20belonging%20to%20class%20%24k%20%5Cin%20%5C%7B1%2C%20%5Cdots%2C%20K%5C%7D%24.%20The%20empirical%20class%20probability%20distribution%20is%3A%0A%0A%20%20%20%20%24%24p_k%20%3D%20%5Cfrac%7BN_%7Bmk%7D%7D%7BN_m%7D%20%3D%20%5Cfrac%7B1%7D%7BN_m%7D%20%5Csum_%7Bi%20%5Cin%20%5Cmathcal%7BR%7D_m%7D%20%5Cmathbb%7BI%7D(y_i%20%3D%20k)%24%24%0A%0A%20%20%20%20%23%23%23%23%20Gini%20Impurity%0A%20%20%20%20The%20Gini%20impurity%20measures%20the%20expected%20error%20rate%20if%20an%20element%20from%20the%20node%20were%20randomly%20classified%20according%20to%20the%20label%20distribution%20of%20that%20node%3A%0A%0A%20%20%20%20%24%24I_G(m)%20%3D%201%20-%20%5Csum_%7Bk%3D1%7D%5EK%20p_k%5E2%20%3D%20%5Csum_%7Bk%3D1%7D%5EK%20p_k%20(1%20-%20p_k)%20%3D%20%5Csum_%7Bj%20%5Cneq%20k%7D%20p_j%20p_k%24%24%0A%0A%20%20%20%20For%20binary%20classification%20(%24K%3D2%24)%20with%20%24p_1%20%3D%20p%24%20and%20%24p_2%20%3D%201%20-%20p%24%3A%0A%0A%20%20%20%20%24%24I_G(p)%20%3D%201%20-%20%5Cleft(p%5E2%20%2B%20(1%20-%20p)%5E2%5Cright)%20%3D%202p(1%20-%20p)%24%24%0A%0A%20%20%20%20The%20Gini%20impurity%20ranges%20from%20%240%24%20(pure%20node%2C%20%24p%20%5Cin%20%5C%7B0%2C%201%5C%7D%24)%20to%20a%20maximum%20of%20%241%20-%20%5Cfrac%7B1%7D%7BK%7D%24%20(for%20binary%2C%20%24%5Cmax%20I_G%20%3D%200.5%24%20at%20%24p%20%3D%200.5%24).%0A%0A%20%20%20%20%23%23%23%23%20Shannon%20Entropy%0A%20%20%20%20Rooted%20in%20information%20theory%2C%20Shannon%20entropy%20quantifies%20the%20expected%20information%20content%20(in%20bits)%20required%20to%20identify%20the%20class%20of%20an%20observation%20drawn%20from%20node%20%24m%24%3A%0A%0A%20%20%20%20%24%24H(m)%20%3D%20-%5Csum_%7Bk%3D1%7D%5EK%20p_k%20%5Clog_2(p_k)%24%24%0A%0A%20%20%20%20with%20the%20limit%20convention%20%240%20%5Clog_2(0)%20%5Cequiv%200%24.%20For%20binary%20classification%3A%0A%0A%20%20%20%20%24%24H(p)%20%3D%20-p%20%5Clog_2(p)%20-%20(1%20-%20p)%20%5Clog_2(1%20-%20p)%24%24%0A%0A%20%20%20%20Shannon%20entropy%20ranges%20from%20%240%24%20(pure%20node)%20to%20a%20maximum%20of%20%24%5Clog_2(K)%24%20(for%20binary%2C%20%24%5Cmax%20H%20%3D%201.0%24%20bit%20at%20%24p%20%3D%200.5%24).%0A%0A%20%20%20%20%23%23%23%23%20Misclassification%20Error%0A%20%20%20%20A%20third%20intuitive%20metric%20is%20the%20misclassification%20error%20rate%3A%0A%0A%20%20%20%20%24%24I_E(m)%20%3D%201%20-%20%5Cmax_%7Bk%20%5Cin%20%5C%7B1%2C%20%5Cdots%2C%20K%5C%7D%7D%20p_k%24%24%0A%0A%20%20%20%20Although%20intuitive%2C%20misclassification%20error%20is%20rarely%20used%20as%20a%20splitting%20criterion.%20Because%20it%20is%20piecewise%20linear%20and%20not%20strictly%20concave%2C%20many%20candidate%20splits%20that%20increase%20child%20purity%20produce%20zero%20reduction%20in%20misclassification%20error%2C%20stalling%20tree%20growth.%0A%0A%20%20%20%20%23%23%23%202.%20The%20Taylor%20Series%20Connection%0A%0A%20%20%20%20Gini%20impurity%20is%20mathematically%20linked%20to%20natural%20entropy%20%24H_e(m)%20%3D%20-%5Csum_%7Bk%3D1%7D%5EK%20p_k%20%5Cln(p_k)%24%20through%20a%20first-order%20Taylor%20expansion%20of%20%24%5Cln(x)%24%20around%20%24x%20%3D%201%24.%20The%20expansion%20of%20%24%5Cln(x)%24%20is%3A%0A%0A%20%20%20%20%24%24%5Cln(x)%20%3D%20(x%20-%201)%20-%20%5Cfrac%7B(x%20-%201)%5E2%7D%7B2%7D%20%2B%20%5Cmathcal%7BO%7D((x%20-%201)%5E3)%24%24%0A%0A%20%20%20%20Setting%20%24x%20%3D%20p_k%24%20and%20neglecting%20higher-order%20terms%3A%0A%0A%20%20%20%20%24%24%5Cln(p_k)%20%5Capprox%20p_k%20-%201%20%3D%20-(1%20-%20p_k)%24%24%0A%0A%20%20%20%20Substituting%20this%20linear%20approximation%20into%20the%20natural%20entropy%20definition%20yields%3A%0A%0A%20%20%20%20%24%24H_e(m)%20%3D%20-%5Csum_%7Bk%3D1%7D%5EK%20p_k%20%5Cln(p_k)%20%5Capprox%20-%5Csum_%7Bk%3D1%7D%5EK%20p_k%20%5B-(1%20-%20p_k)%5D%20%3D%20%5Csum_%7Bk%3D1%7D%5EK%20p_k%20(1%20-%20p_k)%20%3D%20I_G(m)%24%24%0A%0A%20%20%20%20When%20scaled%20by%20a%20factor%20of%202%2C%20the%20binary%20Gini%20curve%20%242%20I_G(p)%20%3D%204p(1%20-%20p)%24%20closely%20tracks%20the%20binary%20Shannon%20entropy%20curve%20%24H(p)%24%2C%20explaining%20why%20both%20criteria%20almost%20always%20select%20identical%20splits.%0A%0A%20%20%20%20%23%23%23%203.%20Impurity%20Reduction%20and%20Information%20Gain%0A%0A%20%20%20%20Given%20a%20continuous%20or%20categorical%20split%20%24s%24%20that%20partitions%20node%20%24m%24%20into%20left%20child%20%24L%24%20and%20right%20child%20%24R%24%20with%20sample%20sizes%20%24N_L%24%20and%20%24N_R%24%20(where%20%24N_m%20%3D%20N_L%20%2B%20N_R%24)%3A%0A%0A%20%20%20%20The%20impurity%20reduction%20(or%20Information%20Gain%20when%20using%20entropy)%20is%3A%0A%0A%20%20%20%20%24%24%5CDelta%20I(m%2C%20s)%20%3D%20I(m)%20-%20%5Cleft(%20%5Cfrac%7BN_L%7D%7BN_m%7D%20I(L)%20%2B%20%5Cfrac%7BN_R%7D%7BN_m%7D%20I(R)%20%5Cright)%24%24%0A%0A%20%20%20%20The%20tree%20search%20algorithm%20selects%20the%20feature%20%24j%5E*%24%20and%20threshold%20%24t%5E*%24%20that%20maximize%20this%20gain%3A%0A%0A%20%20%20%20%24%24(j%5E*%2C%20t%5E*)%20%3D%20%5Carg%5Cmax_%7Bj%2C%20t%7D%20%5CDelta%20I(m%2C%20s(j%2C%20t))%24%24%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(np)%3A%0A%20%20%20%20%23%20Probability%20grid%20for%20binary%20classification%0A%20%20%20%20prob_grid%20%3D%20np.linspace(0.0%2C%201.0%2C%20501)%0A%0A%20%20%20%20%23%201.%20Gini%20Impurity%3A%202%20*%20p%20*%20(1%20-%20p)%0A%20%20%20%20gini_curve%20%3D%202%20*%20prob_grid%20*%20(1.0%20-%20prob_grid)%0A%0A%20%20%20%20%23%202.%20Shannon%20Entropy%3A%20-p*log2(p)%20-%20(1-p)*log2(1-p)%0A%20%20%20%20def%20binary_entropy(p_arr)%3A%0A%20%20%20%20%20%20%20%20res%20%3D%20np.zeros_like(p_arr)%0A%20%20%20%20%20%20%20%20%23%20Avoid%20log(0)%0A%20%20%20%20%20%20%20%20valid_idx%20%3D%20(p_arr%20%3E%200.0)%20%26%20(p_arr%20%3C%201.0)%0A%20%20%20%20%20%20%20%20p_val%20%3D%20p_arr%5Bvalid_idx%5D%0A%20%20%20%20%20%20%20%20res%5Bvalid_idx%5D%20%3D%20-(p_val%20*%20np.log2(p_val)%20%2B%20(1.0%20-%20p_val)%20*%20np.log2(1.0%20-%20p_val))%0A%20%20%20%20%20%20%20%20return%20res%0A%0A%20%20%20%20entropy_curve%20%3D%20binary_entropy(prob_grid)%0A%0A%20%20%20%20%23%203.%20Scaled%20Gini%3A%202%20*%20Gini%20(max%20at%201.0%20to%20overlay%20directly%20onto%20Entropy)%0A%20%20%20%20scaled_gini_curve%20%3D%202.0%20*%20gini_curve%0A%0A%20%20%20%20%23%204.%20Misclassification%20Error%3A%201%20-%20max(p%2C%201-p)%0A%20%20%20%20misclass_curve%20%3D%201.0%20-%20np.maximum(prob_grid%2C%201.0%20-%20prob_grid)%0A%0A%20%20%20%20%23%20Synthetic%20continuous%20feature%20for%20split%20optimization%20demonstration%0A%20%20%20%20np.random.seed(42)%0A%20%20%20%20n_samples%20%3D%20120%0A%20%20%20%20x_feature%20%3D%20np.sort(np.random.uniform(0.0%2C%2010.0%2C%20size%3Dn_samples))%0A%20%20%20%20%23%20True%20boundary%20around%20x%20%3D%204.8%20with%20small%20noise%0A%20%20%20%20y_labels%20%3D%20np.where(x_feature%20%2B%20np.random.normal(0%2C%200.9%2C%20size%3Dn_samples)%20%3E%204.8%2C%201%2C%200)%0A%20%20%20%20return%20(%0A%20%20%20%20%20%20%20%20binary_entropy%2C%0A%20%20%20%20%20%20%20%20entropy_curve%2C%0A%20%20%20%20%20%20%20%20gini_curve%2C%0A%20%20%20%20%20%20%20%20misclass_curve%2C%0A%20%20%20%20%20%20%20%20prob_grid%2C%0A%20%20%20%20%20%20%20%20scaled_gini_curve%2C%0A%20%20%20%20%20%20%20%20x_feature%2C%0A%20%20%20%20%20%20%20%20y_labels%2C%0A%20%20%20%20)%0A%0A%0A%40app.cell%0Adef%20_(%0A%20%20%20%20binary_entropy%2C%0A%20%20%20%20entropy_curve%2C%0A%20%20%20%20gini_curve%2C%0A%20%20%20%20go%2C%0A%20%20%20%20make_subplots%2C%0A%20%20%20%20misclass_curve%2C%0A%20%20%20%20mo%2C%0A%20%20%20%20np%2C%0A%20%20%20%20prob_grid%2C%0A%20%20%20%20scaled_gini_curve%2C%0A%20%20%20%20x_feature%2C%0A%20%20%20%20y_labels%2C%0A)%3A%0A%20%20%20%20%23%20Calculate%20candidate%20split%20curves%20along%20continuous%20feature%20x%0A%20%20%20%20thresholds%20%3D%20(x_feature%5B%3A-1%5D%20%2B%20x_feature%5B1%3A%5D)%20%2F%202.0%0A%20%20%20%20gini_gains%20%3D%20%5B%5D%0A%20%20%20%20entropy_gains%20%3D%20%5B%5D%0A%0A%20%20%20%20%23%20Parent%20node%20impurities%0A%20%20%20%20p_parent%20%3D%20np.mean(y_labels)%0A%20%20%20%20parent_gini%20%3D%202%20*%20p_parent%20*%20(1%20-%20p_parent)%0A%20%20%20%20parent_entropy%20%3D%20binary_entropy(np.array(%5Bp_parent%5D))%5B0%5D%0A%0A%20%20%20%20n_total%20%3D%20len(y_labels)%0A%20%20%20%20for%20thresh%20in%20thresholds%3A%0A%20%20%20%20%20%20%20%20left_mask%20%3D%20x_feature%20%3C%3D%20thresh%0A%20%20%20%20%20%20%20%20right_mask%20%3D%20~left_mask%0A%0A%20%20%20%20%20%20%20%20n_left%20%3D%20np.sum(left_mask)%0A%20%20%20%20%20%20%20%20n_right%20%3D%20np.sum(right_mask)%0A%0A%20%20%20%20%20%20%20%20if%20n_left%20%3D%3D%200%20or%20n_right%20%3D%3D%200%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20gini_gains.append(0.0)%0A%20%20%20%20%20%20%20%20%20%20%20%20entropy_gains.append(0.0)%0A%20%20%20%20%20%20%20%20%20%20%20%20continue%0A%0A%20%20%20%20%20%20%20%20p_left%20%3D%20np.mean(y_labels%5Bleft_mask%5D)%0A%20%20%20%20%20%20%20%20p_right%20%3D%20np.mean(y_labels%5Bright_mask%5D)%0A%0A%20%20%20%20%20%20%20%20g_left%20%3D%202%20*%20p_left%20*%20(1%20-%20p_left)%0A%20%20%20%20%20%20%20%20g_right%20%3D%202%20*%20p_right%20*%20(1%20-%20p_right)%0A%20%20%20%20%20%20%20%20g_gain%20%3D%20parent_gini%20-%20((n_left%20%2F%20n_total)%20*%20g_left%20%2B%20(n_right%20%2F%20n_total)%20*%20g_right)%0A%20%20%20%20%20%20%20%20gini_gains.append(g_gain)%0A%0A%20%20%20%20%20%20%20%20e_left%20%3D%20binary_entropy(np.array(%5Bp_left%5D))%5B0%5D%0A%20%20%20%20%20%20%20%20e_right%20%3D%20binary_entropy(np.array(%5Bp_right%5D))%5B0%5D%0A%20%20%20%20%20%20%20%20e_gain%20%3D%20parent_entropy%20-%20((n_left%20%2F%20n_total)%20*%20e_left%20%2B%20(n_right%20%2F%20n_total)%20*%20e_right)%0A%20%20%20%20%20%20%20%20entropy_gains.append(e_gain)%0A%0A%20%20%20%20gini_gains%20%3D%20np.array(gini_gains)%0A%20%20%20%20entropy_gains%20%3D%20np.array(entropy_gains)%0A%0A%20%20%20%20fig%20%3D%20make_subplots(%0A%20%20%20%20%20%20%20%20rows%3D1%2C%0A%20%20%20%20%20%20%20%20cols%3D2%2C%0A%20%20%20%20%20%20%20%20subplot_titles%3D%5B%0A%20%20%20%20%20%20%20%20%20%20%20%20%22%3Cb%3EImpurity%20Curves%20as%20a%20Function%20of%20Class%20Proportion%20p%3C%2Fb%3E%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22%3Cb%3EImpurity%20Reduction%20Gain%20Across%20Candidate%20Thresholds%3C%2Fb%3E%22%2C%0A%20%20%20%20%20%20%20%20%5D%2C%0A%20%20%20%20%20%20%20%20horizontal_spacing%3D0.12%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Panel%201%3A%20Theoretical%20Curves%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dprob_grid%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dentropy_curve%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%23DC2626%22%2C%20width%3D2.5)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Shannon%20Entropy%20H(p)%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dprob_grid%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dscaled_gini_curve%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%232563EB%22%2C%20width%3D2.5%2C%20dash%3D%22dash%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Scaled%20Gini%202%20*%20I_G(p)%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dprob_grid%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dgini_curve%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%230D9488%22%2C%20width%3D2)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Standard%20Gini%20I_G(p)%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dprob_grid%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dmisclass_curve%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%239CA3AF%22%2C%20width%3D1.5%2C%20dash%3D%22dot%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Misclassification%20Error%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Panel%202%3A%20Split%20Optimization%20Gain%20Curves%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dthresholds%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dentropy_gains%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%23DC2626%22%2C%20width%3D2.5)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Entropy%20Information%20Gain%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Scale%20Gini%20gain%20to%20overlay%20on%20same%20plot%20for%20alignment%20inspection%0A%20%20%20%20scaled_gini_gain%20%3D%20gini_gains%20*%20(np.max(entropy_gains)%20%2F%20(np.max(gini_gains)%20%2B%201e-9))%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dthresholds%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dscaled_gini_gain%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%232563EB%22%2C%20width%3D2.5%2C%20dash%3D%22dash%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Gini%20Gain%20(Aligned%20Scale)%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20best_thresh_gini%20%3D%20thresholds%5Bnp.argmax(gini_gains)%5D%0A%20%20%20%20best_thresh_entropy%20%3D%20thresholds%5Bnp.argmax(entropy_gains)%5D%0A%0A%20%20%20%20fig.add_vline(%0A%20%20%20%20%20%20%20%20x%3Dbest_thresh_gini%2C%0A%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%232563EB%22%2C%20width%3D1.5%2C%20dash%3D%22dot%22)%2C%0A%20%20%20%20%20%20%20%20annotation_text%3Df%22Best%20Gini%3A%20%7Bbest_thresh_gini%3A.2f%7D%22%2C%0A%20%20%20%20%20%20%20%20annotation_position%3D%22top%20left%22%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20fig.add_vline(%0A%20%20%20%20%20%20%20%20x%3Dbest_thresh_entropy%2C%0A%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%23DC2626%22%2C%20width%3D1.5%2C%20dash%3D%22dash%22)%2C%0A%20%20%20%20%20%20%20%20annotation_text%3Df%22Best%20Entropy%3A%20%7Bbest_thresh_entropy%3A.2f%7D%22%2C%0A%20%20%20%20%20%20%20%20annotation_position%3D%22top%20right%22%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20fig.update_xaxes(title_text%3D%22Probability%20of%20Positive%20Class%20p%22%2C%20row%3D1%2C%20col%3D1)%0A%20%20%20%20fig.update_yaxes(title_text%3D%22Impurity%20Value%22%2C%20row%3D1%2C%20col%3D1)%0A%20%20%20%20fig.update_xaxes(title_text%3D%22Candidate%20Feature%20Threshold%20x%22%2C%20row%3D1%2C%20col%3D2)%0A%20%20%20%20fig.update_yaxes(title_text%3D%22Impurity%20Reduction%22%2C%20row%3D1%2C%20col%3D2)%0A%0A%20%20%20%20fig.update_layout(%0A%20%20%20%20%20%20%20%20template%3D%22plotly_white%22%2C%0A%20%20%20%20%20%20%20%20height%3D500%2C%0A%20%20%20%20%20%20%20%20margin%3Ddict(l%3D50%2C%20r%3D40%2C%20t%3D70%2C%20b%3D50)%2C%0A%20%20%20%20%20%20%20%20legend%3Ddict(orientation%3D%22h%22%2C%20yanchor%3D%22bottom%22%2C%20y%3D-0.28%2C%20xanchor%3D%22center%22%2C%20x%3D0.5)%2C%0A%20%20%20%20)%0A%0A%20%20%20%20viz%20%3D%20mo.ui.plotly(fig)%0A%20%20%20%20return%20best_thresh_entropy%2C%20best_thresh_gini%0A%0A%0A%40app.cell%0Adef%20_()%3A%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(%0A%20%20%20%20DecisionTreeClassifier%2C%0A%20%20%20%20best_thresh_entropy%2C%0A%20%20%20%20best_thresh_gini%2C%0A%20%20%20%20load_breast_cancer%2C%0A%20%20%20%20mo%2C%0A%20%20%20%20np%2C%0A%20%20%20%20pd%2C%0A%20%20%20%20time%2C%0A%20%20%20%20train_test_split%2C%0A)%3A%0A%20%20%20%20%23%20Example%201%3A%20Pure%20Vectorized%20NumPy%20Calculation%20of%20Impurities%20for%20Multi-Class%20Scenarios%0A%20%20%20%20def%20compute_all_impurities(counts)%3A%0A%20%20%20%20%20%20%20%20total%20%3D%20np.sum(counts)%0A%20%20%20%20%20%20%20%20if%20total%20%3D%3D%200%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20return%200.0%2C%200.0%2C%200.0%0A%20%20%20%20%20%20%20%20p%20%3D%20counts%20%2F%20total%0A%20%20%20%20%20%20%20%20gini%20%3D%201.0%20-%20np.sum(p**2)%0A%20%20%20%20%20%20%20%20valid_p%20%3D%20p%5Bp%20%3E%200.0%5D%0A%20%20%20%20%20%20%20%20entropy%20%3D%20-np.sum(valid_p%20*%20np.log2(valid_p))%0A%20%20%20%20%20%20%20%20misclass%20%3D%201.0%20-%20np.max(p)%0A%20%20%20%20%20%20%20%20return%20gini%2C%20entropy%2C%20misclass%0A%0A%20%20%20%20scenario_counts%20%3D%20%5B%0A%20%20%20%20%20%20%20%20%5B50%2C%2050%5D%2C%0A%20%20%20%20%20%20%20%20%5B70%2C%2030%5D%2C%0A%20%20%20%20%20%20%20%20%5B90%2C%2010%5D%2C%0A%20%20%20%20%20%20%20%20%5B99%2C%201%5D%2C%0A%20%20%20%20%20%20%20%20%5B100%2C%200%5D%2C%0A%20%20%20%20%20%20%20%20%5B33%2C%2033%2C%2034%5D%2C%0A%20%20%20%20%20%20%20%20%5B70%2C%2020%2C%2010%5D%2C%0A%20%20%20%20%20%20%20%20%5B10%2C%2010%2C%2010%2C%2010%5D%2C%0A%20%20%20%20%5D%0A%0A%20%20%20%20scenario_results%20%3D%20%5B%5D%0A%20%20%20%20for%20counts%20in%20scenario_counts%3A%0A%20%20%20%20%20%20%20%20g%2C%20e%2C%20m%20%3D%20compute_all_impurities(np.array(counts))%0A%20%20%20%20%20%20%20%20scenario_results.append(%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Class_Distribution%22%3A%20str(counts)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Num_Classes%22%3A%20len(counts)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Gini_Impurity%22%3A%20round(g%2C%204)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Scaled_Gini_2x%22%3A%20round(2%20*%20g%2C%204)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Shannon_Entropy_bits%22%3A%20round(e%2C%204)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Misclassification_Rate%22%3A%20round(m%2C%204)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7D%0A%20%20%20%20%20%20%20%20)%0A%0A%20%20%20%20df_impurities%20%3D%20pd.DataFrame(scenario_results)%0A%0A%20%20%20%20%23%20Example%202%3A%20Threshold%20Optimization%20Comparison%0A%20%20%20%20df_splits%20%3D%20pd.DataFrame(%0A%20%20%20%20%20%20%20%20%5B%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Splitting_Criterion%22%3A%20%22Gini%20Impurity%20(CART)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Optimal_Threshold%22%3A%20round(best_thresh_gini%2C%204)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Formula%22%3A%20%221%20-%20sum(p_k%5E2)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Computational_Complexity%22%3A%20%22O(K)%20arithmetic%20additions%2Fmultiplications%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Splitting_Criterion%22%3A%20%22Shannon%20Entropy%20(C4.5%20%2F%20ID3)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Optimal_Threshold%22%3A%20round(best_thresh_entropy%2C%204)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Formula%22%3A%20%22-sum(p_k%20*%20log2(p_k))%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Computational_Complexity%22%3A%20%22O(K)%20transcendental%20logarithm%20evaluations%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7D%2C%0A%20%20%20%20%20%20%20%20%5D%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Example%203%3A%20Empirical%20Tree%20Benchmark%20on%20Real-World%20Dataset%20(Breast%20Cancer%20Wisconsin)%0A%20%20%20%20cancer%20%3D%20load_breast_cancer()%0A%20%20%20%20X_train%2C%20X_test%2C%20y_train%2C%20y_test%20%3D%20train_test_split(%0A%20%20%20%20%20%20%20%20cancer.data%2C%20cancer.target%2C%20test_size%3D0.3%2C%20random_state%3D42%2C%20stratify%3Dcancer.target%0A%20%20%20%20)%0A%0A%20%20%20%20n_iterations%20%3D%20150%0A%0A%20%20%20%20%23%20Benchmark%20Gini%20Tree%0A%20%20%20%20t0%20%3D%20time.perf_counter()%0A%20%20%20%20for%20_%20in%20range(n_iterations)%3A%0A%20%20%20%20%20%20%20%20tree_gini%20%3D%20DecisionTreeClassifier(criterion%3D%22gini%22%2C%20random_state%3D42)%0A%20%20%20%20%20%20%20%20tree_gini.fit(X_train%2C%20y_train)%0A%20%20%20%20time_gini_ms%20%3D%20(time.perf_counter()%20-%20t0)%20*%201000%20%2F%20n_iterations%0A%20%20%20%20gini_acc%20%3D%20tree_gini.score(X_test%2C%20y_test)%0A%20%20%20%20gini_depth%20%3D%20tree_gini.get_depth()%0A%20%20%20%20gini_leaves%20%3D%20tree_gini.get_n_leaves()%0A%0A%20%20%20%20%23%20Benchmark%20Entropy%20Tree%0A%20%20%20%20t0%20%3D%20time.perf_counter()%0A%20%20%20%20for%20_%20in%20range(n_iterations)%3A%0A%20%20%20%20%20%20%20%20tree_entropy%20%3D%20DecisionTreeClassifier(criterion%3D%22entropy%22%2C%20random_state%3D42)%0A%20%20%20%20%20%20%20%20tree_entropy.fit(X_train%2C%20y_train)%0A%20%20%20%20time_entropy_ms%20%3D%20(time.perf_counter()%20-%20t0)%20*%201000%20%2F%20n_iterations%0A%20%20%20%20entropy_acc%20%3D%20tree_entropy.score(X_test%2C%20y_test)%0A%20%20%20%20entropy_depth%20%3D%20tree_entropy.get_depth()%0A%20%20%20%20entropy_leaves%20%3D%20tree_entropy.get_n_leaves()%0A%0A%20%20%20%20df_benchmark%20%3D%20pd.DataFrame(%0A%20%20%20%20%20%20%20%20%5B%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Criterion%22%3A%20%22Gini%20Impurity%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Fit_Time_per_Tree_ms%22%3A%20round(time_gini_ms%2C%203)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Tree_Max_Depth%22%3A%20gini_depth%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Number_of_Leaves%22%3A%20gini_leaves%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Test_Accuracy%22%3A%20f%22%7Bgini_acc%20*%20100%3A.2f%7D%25%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Criterion%22%3A%20%22Shannon%20Entropy%20(Log%20Loss)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Fit_Time_per_Tree_ms%22%3A%20round(time_entropy_ms%2C%203)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Tree_Max_Depth%22%3A%20entropy_depth%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Number_of_Leaves%22%3A%20entropy_leaves%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Test_Accuracy%22%3A%20f%22%7Bentropy_acc%20*%20100%3A.2f%7D%25%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7D%2C%0A%20%20%20%20%20%20%20%20%5D%0A%20%20%20%20)%0A%0A%20%20%20%20table_imp%20%3D%20mo.ui.table(df_impurities)%0A%20%20%20%20table_spl%20%3D%20mo.ui.table(df_splits)%0A%20%20%20%20table_bnk%20%3D%20mo.ui.table(df_benchmark)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_()%3A%0A%20%20%20%20return%0A%0A%0Aif%20__name__%20%3D%3D%20%22__main__%22%3A%0A%20%20%20%20app.run()%0A
df9a6ad046cca729207317925367512f