import%20marimo%0A%0A__generated_with%20%3D%20%220.24.0%22%0Aapp%20%3D%20marimo.App()%0A%0A%0A%40app.cell%0Adef%20_()%3A%0A%20%20%20%20import%20marimo%20as%20mo%0A%20%20%20%20import%20numpy%20as%20np%0A%20%20%20%20import%20pandas%20as%20pd%0A%20%20%20%20import%20plotly.graph_objects%20as%20go%0A%20%20%20%20from%20plotly.subplots%20import%20make_subplots%0A%20%20%20%20from%20scipy%20import%20stats%0A%0A%20%20%20%20return%20go%2C%20make_subplots%2C%20mo%2C%20np%2C%20pd%2C%20stats%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%20Note%2019%3A%20Kendall's%20Tau-b%2C%20Concordance%2C%20and%20Robust%20Rank%20Correlation%0A%0A%20%20%20%20%26larr%3B%20Previous%20Note%3A%20%5B18%20Cramer%20V%5D(18_cramer_v.py)%20%7C%20Next%20Note%3A%20%5B20%20Spurious%20Correlation%5D(20_spurious_correlation.py)%20%26rarr%3B%0A%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%20%5Ba%5D%20Why%20do%20you%20need%20to%20know%20these%20concepts%3F%0A%0A%20%20%20%20In%20data%20science%2C%20experimental%20statistics%2C%20and%20modern%20machine%20learning%20evaluation%2C%20we%20frequently%20measure%20association%20between%20variables%20that%20are%20ordinal%20(e.g.%20Likert%20satisfaction%20scales%2C%20benchmark%20rankings%2C%20preference%20annotations%2C%20customer%20review%20stars)%20or%20continuous%20variables%20corrupted%20by%20severe%20non-Gaussian%20noise%20and%20extreme%20outliers.%0A%0A%20%20%20%20While%20Pearson's%20correlation%20coefficient%20%24r%24%20measures%20linear%20association%20between%20continuous%20variables%2C%20it%20suffers%20from%20two%20major%20limitations%3A%0A%20%20%20%201.%20It%20assumes%20linearity%20and%20bivariate%20normality%2C%20breaking%20down%20when%20the%20underlying%20relationship%20is%20non-linear%20but%20strictly%20monotonic.%0A%20%20%20%202.%20It%20possesses%20an%20unbounded%20influence%20function%3A%20a%20single%20severe%20outlier%20can%20shift%20Pearson's%20%24r%24%20from%20%24%2B0.95%24%20to%20%24-0.50%24.%0A%0A%20%20%20%20Spearman's%20rank%20correlation%20%24%5Crho%24%20mitigates%20non-linearity%20by%20converting%20raw%20values%20into%20ranks%2C%20but%20it%20does%20not%20account%20directly%20for%20tied%20observations%20and%20its%20sampling%20distribution%20converges%20slowly%20to%20normality.%0A%0A%20%20%20%20**Kendall's%20Tau-b%20(%24%5Ctau_b%24)**%20solves%20these%20challenges%20through%20pairwise%20concordance%3A%0A%20%20%20%201.%20**Probabilistic%20Interpretation**%3A%20Kendall's%20%24%5Ctau%24%20is%20defined%20directly%20as%20the%20difference%20between%20the%20probability%20of%20concordance%20and%20discordance%3A%20%24P(%5Ctext%7Bconcordance%7D)%20-%20P(%5Ctext%7Bdiscordance%7D)%24.%20A%20value%20of%20%24%2B0.60%24%20means%20that%20for%20any%20randomly%20selected%20pair%20of%20observations%2C%20the%20probability%20that%20they%20agree%20in%20ordering%20is%2060%20percentage%20points%20higher%20than%20the%20probability%20that%20they%20disagree.%0A%20%20%20%202.%20**Tie-Corrected%20Formulation%20(%24%5Ctau_b%24)**%3A%20Unlike%20Kendall's%20raw%20%24%5Ctau_a%24%20(which%20underestimates%20association%20when%20ties%20exist)%2C%20%24%5Ctau_b%24%20normalizes%20by%20the%20geometric%20mean%20of%20untied%20pairs%20in%20each%20variable%2C%20ensuring%20the%20coefficient%20reaches%20the%20canonical%20%24%5B-1%2C%201%5D%24%20bounds%20even%20on%20heavily%20discretized%20ordinal%20scales.%0A%20%20%20%203.%20**Superior%20Asymptotic%20and%20U-Statistic%20Properties**%3A%20Because%20Kendall's%20%24%5Ctau%24%20is%20an%20unbiased%20U-statistic%2C%20its%20variance%20under%20the%20null%20hypothesis%20of%20independence%20depends%20purely%20on%20the%20sample%20size%20%24n%24%20and%20tie%20counts%2C%20converging%20to%20a%20standard%20Gaussian%20distribution%20much%20faster%20than%20Spearman's%20%24%5Crho%24.%20This%20yields%20reliable%20hypothesis%20tests%20and%20confidence%20intervals%20even%20in%20small%20samples%20(%24n%20%3C%2030%24).%0A%20%20%20%204.%20**LLM%20Evaluation%2C%20Preference%20Alignment%2C%20and%20Learning%20to%20Rank%20(LTR)**%3A%20In%20RLHF%20(Reinforcement%20Learning%20from%20Human%20Feedback)%2C%20search%20engine%20ranking%2C%20and%20LLM-as-a-judge%20benchmarking%2C%20Kendall's%20%24%5Ctau_b%24%20is%20the%20standard%20metric%20for%20measuring%20inter-annotator%20agreement%20and%20ranking%20alignment%20between%20human%20annotators%20and%20language%20model%20judges.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%20%5Bb%5D%20Concept%20explanation%20with%20their%20role%20in%20ML%2FAI%2FStats%3F%0A%0A%20%20%20%20%23%23%23%201.%20Concordant%2C%20Discordant%2C%20and%20Tied%20Pairs%0A%0A%20%20%20%20Consider%20a%20bivariate%20sample%20of%20%24n%24%20observations%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5C%7B(x_1%2C%20y_1)%2C%20(x_2%2C%20y_2)%2C%20%5Cdots%2C%20(x_n%2C%20y_n)%5C%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20The%20total%20number%20of%20distinct%20pairs%20%24(i%2C%20j)%24%20with%20%241%20%5Cleq%20i%20%3C%20j%20%5Cleq%20n%24%20is%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20N_%7B%5Ctext%7Bpairs%7D%7D%20%3D%20%5Cbinom%7Bn%7D%7B2%7D%20%3D%20%5Cfrac%7Bn(n%20-%201)%7D%7B2%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20For%20any%20pair%20of%20observations%20%24(x_i%2C%20y_i)%24%20and%20%24(x_j%2C%20y_j)%24%2C%20we%20evaluate%20the%20sign%20of%20their%20coordinate%20differences%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5CDelta_%7Bij%7D%20%3D%20(x_i%20-%20x_j)(y_i%20-%20y_j)%0A%20%20%20%20%24%24%0A%0A%20%20%20%20Every%20pair%20falls%20into%20exactly%20one%20of%20four%20mutual%20categories%3A%0A%0A%20%20%20%20%23%23%23%23%20Concordant%20Pairs%20(%24C%24)%0A%20%20%20%20Both%20variables%20change%20in%20the%20same%20direction%3A%0A%20%20%20%20%24%24%0A%20%20%20%20%5CDelta_%7Bij%7D%20%3E%200%20%5Ciff%20(x_i%20%3E%20x_j%20%5Ctext%7B%20and%20%7D%20y_i%20%3E%20y_j)%20%5Ctext%7B%20or%20%7D%20(x_i%20%3C%20x_j%20%5Ctext%7B%20and%20%7D%20y_i%20%3C%20y_j)%0A%20%20%20%20%24%24%0A%0A%20%20%20%20%23%23%23%23%20Discordant%20Pairs%20(%24D%24)%0A%20%20%20%20The%20variables%20change%20in%20opposite%20directions%3A%0A%20%20%20%20%24%24%0A%20%20%20%20%5CDelta_%7Bij%7D%20%3C%200%20%5Ciff%20(x_i%20%3E%20x_j%20%5Ctext%7B%20and%20%7D%20y_i%20%3C%20y_j)%20%5Ctext%7B%20or%20%7D%20(x_i%20%3C%20x_j%20%5Ctext%7B%20and%20%7D%20y_i%20%3E%20y_j)%0A%20%20%20%20%24%24%0A%0A%20%20%20%20%23%23%23%23%20Tied%20Pairs%20on%20%24X%24%20Only%20(%24T_X%24)%0A%20%20%20%20Observations%20share%20the%20same%20%24x%24%20coordinate%20but%20differ%20in%20%24y%24%3A%0A%20%20%20%20%24%24%0A%20%20%20%20x_i%20%3D%20x_j%20%5Cquad%20%5Ctext%7Band%7D%20%5Cquad%20y_i%20%5Cneq%20y_j%0A%20%20%20%20%24%24%0A%0A%20%20%20%20%23%23%23%23%20Tied%20Pairs%20on%20%24Y%24%20Only%20(%24T_Y%24)%0A%20%20%20%20Observations%20share%20the%20same%20%24y%24%20coordinate%20but%20differ%20in%20%24x%24%3A%0A%20%20%20%20%24%24%0A%20%20%20%20x_i%20%5Cneq%20x_j%20%5Cquad%20%5Ctext%7Band%7D%20%5Cquad%20y_i%20%3D%20y_j%0A%20%20%20%20%24%24%0A%0A%20%20%20%20%23%23%23%23%20Tied%20Pairs%20on%20Both%20(%24T_%7BXY%7D%24)%0A%20%20%20%20Observations%20are%20identical%20in%20both%20dimensions%3A%0A%20%20%20%20%24%24%0A%20%20%20%20x_i%20%3D%20x_j%20%5Cquad%20%5Ctext%7Band%7D%20%5Cquad%20y_i%20%3D%20y_j%0A%20%20%20%20%24%24%0A%0A%20%20%20%20The%20sum%20of%20all%20classifications%20satisfies%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20C%20%2B%20D%20%2B%20T_X%20%2B%20T_Y%20%2B%20T_%7BXY%7D%20%3D%20%5Cbinom%7Bn%7D%7B2%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%23%202.%20Kendall's%20Tau-a%20vs.%20Tau-b%20vs.%20Tau-c%0A%0A%20%20%20%20%23%23%23%23%20Kendall's%20Tau-a%20(%24%5Ctau_a%24)%0A%20%20%20%20When%20there%20are%20no%20ties%20in%20the%20data%2C%20Kendall's%20tau%20is%20simply%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Ctau_a%20%3D%20%5Cfrac%7BC%20-%20D%7D%7B%5Cbinom%7Bn%7D%7B2%7D%7D%20%3D%20%5Cfrac%7BC%20-%20D%7D%7B%5Cfrac%7B1%7D%7B2%7D%20n%20(n%20-%201)%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20If%20ties%20are%20present%2C%20%24%5Ctau_a%24%20cannot%20reach%20%24%2B1%24%20or%20%24-1%24%20because%20ties%20reduce%20the%20maximum%20possible%20value%20of%20%24C%24%20or%20%24D%24.%0A%0A%20%20%20%20%23%23%23%23%20Kendall's%20Tau-b%20(%24%5Ctau_b%24)%0A%20%20%20%20To%20account%20for%20tied%20observations%2C%20Maurice%20Kendall%20defined%20%24%5Ctau_b%24%20by%20adjusting%20the%20denominator%20with%20the%20geometric%20mean%20of%20pairs%20that%20are%20untied%20in%20%24X%24%20and%20untied%20in%20%24Y%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Ctau_b%20%3D%20%5Cfrac%7BC%20-%20D%7D%7B%5Csqrt%7B(C%20%2B%20D%20%2B%20T_X)(C%20%2B%20D%20%2B%20T_Y)%7D%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20Equivalently%2C%20letting%20%24t_k%24%20denote%20the%20size%20of%20the%20%24k%24-th%20group%20of%20tied%20values%20in%20%24X%24%2C%20and%20%24u_m%24%20denote%20the%20size%20of%20the%20%24m%24-th%20group%20of%20tied%20values%20in%20%24Y%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20n_0%20%3D%20%5Cfrac%7Bn(n%20-%201)%7D%7B2%7D%2C%20%5Cquad%20n_1%20%3D%20%5Csum_%7Bk%7D%20%5Cfrac%7Bt_k(t_k%20-%201)%7D%7B2%7D%2C%20%5Cquad%20n_2%20%3D%20%5Csum_%7Bm%7D%20%5Cfrac%7Bu_m(u_m%20-%201)%7D%7B2%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Ctau_b%20%3D%20%5Cfrac%7BC%20-%20D%7D%7B%5Csqrt%7B(n_0%20-%20n_1)(n_0%20-%20n_2)%7D%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20Notice%20that%20%24n_0%20-%20n_1%20%3D%20C%20%2B%20D%20%2B%20T_X%24%20and%20%24n_0%20-%20n_2%20%3D%20C%20%2B%20D%20%2B%20T_Y%24.%20When%20no%20ties%20exist%2C%20%24n_1%20%3D%20n_2%20%3D%200%24%20and%20%24%5Ctau_b%24%20reduces%20identically%20to%20%24%5Ctau_a%24.%0A%0A%20%20%20%20%23%23%23%23%20Stuart's%20Tau-c%20(%24%5Ctau_c%24)%0A%20%20%20%20When%20analyzing%20rectangular%20contingency%20tables%20where%20the%20number%20of%20rows%20%24r%24%20and%20columns%20%24c%24%20differ%2C%20%24%5Ctau_b%24%20cannot%20reach%20%24%5Cpm%201%24.%20Stuart's%20%24%5Ctau_c%24%20adjusts%20for%20table%20dimensions%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Ctau_c%20%3D%20%5Cfrac%7B2%20m%20(C%20-%20D)%7D%7Bn%5E2%20(m%20-%201)%7D%2C%20%5Cquad%20%5Ctext%7Bwhere%20%7D%20m%20%3D%20%5Cmin(r%2C%20c)%0A%20%20%20%20%24%24%0A%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%23%203.%20Hypothesis%20Testing%20and%20Asymptotic%20Variance%0A%0A%20%20%20%20Under%20the%20null%20hypothesis%20%24H_0%24%20that%20%24X%24%20and%20%24Y%24%20are%20independent%20(no%20monotonic%20association)%2C%20the%20expectation%20is%20%24%5Cmathbb%7BE%7D%5B%5Ctau_b%5D%20%3D%200%24.%0A%0A%20%20%20%20For%20samples%20without%20ties%2C%20the%20exact%20null%20variance%20is%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Csigma_0%5E2%20%3D%20%5Coperatorname%7BVar%7D(%5Ctau)%20%3D%20%5Cfrac%7B2(2n%20%2B%205)%7D%7B9n(n%20-%201)%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20When%20ties%20are%20present%2C%20the%20variance%20formula%20adjusts%20for%20the%20tie%20multiplicities%20%24t_k%24%20and%20%24u_m%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20v_0%20%3D%20%5Cfrac%7Bn(n%20-%201)(2n%20%2B%205)%20-%20%5Csum%20t_k(t_k%20-%201)(2t_k%20%2B%205)%20-%20%5Csum%20u_m(u_m%20-%201)(2u_m%20%2B%205)%7D%7B18%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20%24%24%0A%20%20%20%20z%20%3D%20%5Cfrac%7BC%20-%20D%7D%7B%5Csqrt%7Bv_0%7D%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20The%20two-tailed%20%24p%24-value%20is%20then%20computed%20from%20the%20standard%20normal%20cumulative%20distribution%20function%20%24%5CPhi%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20p%20%3D%202%20%5Ccdot%20%5Cleft(1%20-%20%5CPhi(%7Cz%7C)%5Cright)%0A%20%20%20%20%24%24%0A%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%23%204.%20Mathematical%20Comparison%3A%20Pearson%20vs.%20Spearman%20vs.%20Kendall%0A%0A%20%20%20%20%7C%20Metric%20%7C%20Formula%20Basis%20%7C%20Assumptions%20%7C%20Outlier%20Sensitivity%20%7C%20Tie%20Handling%20%7C%20Primary%20Domain%20%7C%0A%20%20%20%20%7C%20%3A---%20%7C%20%3A---%20%7C%20%3A---%20%7C%20%3A---%20%7C%20%3A---%20%7C%20%3A---%20%7C%0A%20%20%20%20%7C%20**Pearson%20%24r%24**%20%7C%20Covariance%20%2F%20%24(%5Csigma_x%20%5Csigma_y)%24%20%7C%20Linear%2C%20Bivariate%20Normal%2C%20Continuous%20%7C%20Extremely%20High%20(unbounded%20influence)%20%7C%20Trivial%20(continuous)%20%7C%20Continuous%20linear%20models%2C%20OLS%20regression%20%7C%0A%20%20%20%20%7C%20**Spearman%20%24%5Crho%24**%20%7C%20Pearson%20%24r%24%20on%20Ranks%20%7C%20Monotonic%20relationship%20%7C%20Moderate%20(bounded%20by%20rank%20extremes)%20%7C%20Average%20ranks%20with%20correction%20%7C%20Continuous%20monotonic%20data%2C%20non-normal%20distributions%20%7C%0A%20%20%20%20%7C%20**Kendall%20%24%5Ctau_b%24**%20%7C%20Pairwise%20Concordance%20%2F%20Denominator%20%7C%20Monotonic%2C%20Ordinal%20scale%20%7C%20Minimal%20(pairwise%20swap%20count)%20%7C%20Explicit%20tie%20adjustment%20%24%5Ctau_b%24%20%7C%20Ordinal%20survey%20data%2C%20LLM%20alignment%2C%20small%20samples%20%7C%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(pd)%3A%0A%20%20%20%20%23%20Raw%20Survey%20Observations%3A%20Job%20Satisfaction%20(X)%20vs%20Work-Life%20Balance%20(Y)%0A%20%20%20%20_observations%20%3D%20%5B%0A%20%20%20%20%20%20%20%20(5%2C%204)%2C%0A%20%20%20%20%20%20%20%20(3%2C%203)%2C%0A%20%20%20%20%20%20%20%20(4%2C%204)%2C%0A%20%20%20%20%20%20%20%20(2%2C%202)%2C%0A%20%20%20%20%20%20%20%20(1%2C%201)%2C%0A%20%20%20%20%20%20%20%20(4%2C%203)%2C%0A%20%20%20%20%20%20%20%20(5%2C%205)%2C%0A%20%20%20%20%20%20%20%20(3%2C%202)%2C%0A%20%20%20%20%20%20%20%20(2%2C%203)%2C%0A%20%20%20%20%20%20%20%20(4%2C%204)%2C%0A%20%20%20%20%20%20%20%20(5%2C%205)%2C%0A%20%20%20%20%20%20%20%20(3%2C%203)%2C%0A%20%20%20%20%20%20%20%20(1%2C%202)%2C%0A%20%20%20%20%20%20%20%20(2%2C%201)%2C%0A%20%20%20%20%20%20%20%20(4%2C%205)%2C%0A%20%20%20%20%20%20%20%20(3%2C%202)%2C%0A%20%20%20%20%20%20%20%20(5%2C%204)%2C%0A%20%20%20%20%20%20%20%20(1%2C%202)%2C%0A%20%20%20%20%20%20%20%20(2%2C%203)%2C%0A%20%20%20%20%20%20%20%20(3%2C%203)%2C%0A%20%20%20%20%5D%0A%0A%20%20%20%20df_sample%20%3D%20pd.DataFrame(_observations%2C%20columns%3D%5B%22Job%20Satisfaction%22%2C%20%22Work-Life%20Balance%22%5D)%0A%20%20%20%20return%20(df_sample%2C)%0A%0A%0A%40app.cell%0Adef%20_(df_sample%2C%20go%2C%20make_subplots%2C%20np%2C%20stats)%3A%0A%20%20%20%20%23%20Interactive%20Visualizations%20Cell%0A%20%20%20%20%23%20Subplot%201%3A%20Jittered%20scatter%20plot%20of%20the%20survey%20observations%20(Job%20Satisfaction%20vs%20Work-Life%20Balance)%0A%20%20%20%20%23%20Subplot%202%3A%20Robustness%20breakdown%20curve%3A%20Pearson%20r%20vs%20Spearman%20rho%20vs%20Kendall%20tau-b%20under%20an%20injected%20extreme%20outlier%0A%0A%20%20%20%20_fig%20%3D%20make_subplots(%0A%20%20%20%20%20%20%20%20rows%3D1%2C%0A%20%20%20%20%20%20%20%20cols%3D2%2C%0A%20%20%20%20%20%20%20%20subplot_titles%3D(%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Ordinal%20Survey%20Observations%20(with%20Jitter%20%26%20Concordance)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Outlier%20Breakdown%3A%20Pearson%20vs.%20Spearman%20vs.%20Kendall%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20horizontal_spacing%3D0.14%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Subplot%201%3A%20Jittered%20Scatter%0A%20%20%20%20np.random.seed(42)%0A%20%20%20%20_x_vals%20%3D%20df_sample%5B%22Job%20Satisfaction%22%5D.values%0A%20%20%20%20_y_vals%20%3D%20df_sample%5B%22Work-Life%20Balance%22%5D.values%0A%20%20%20%20_jitter_x%20%3D%20_x_vals%20%2B%20np.random.uniform(-0.12%2C%200.12%2C%20size%3Dlen(_x_vals))%0A%20%20%20%20_jitter_y%20%3D%20_y_vals%20%2B%20np.random.uniform(-0.12%2C%200.12%2C%20size%3Dlen(_y_vals))%0A%0A%20%20%20%20_fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3D_jitter_x%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3D_jitter_y%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22markers%2Btext%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20marker%3Ddict(%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20size%3D12%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20color%3D_x_vals%20%2B%20_y_vals%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20colorscale%3D%22Purples%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20showscale%3DFalse%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(width%3D1.5%2C%20color%3D%22%234a154b%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20text%3D%5Bf%22(%7Bx%7D%2C%20%7By%7D)%22%20for%20x%2C%20y%20in%20zip(_x_vals%2C%20_y_vals)%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20textposition%3D%22top%20center%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20textfont%3Ddict(size%3D8%2C%20color%3D%22%23555555%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Survey%20Pairs%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20hovertemplate%3D%22Job%20Satisfaction%3A%20%25%7Bx%3A.0f%7D%3Cbr%3EWork-Life%20Balance%3A%20%25%7By%3A.0f%7D%3Cextra%3E%3C%2Fextra%3E%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Trend%20line%20for%20subplot%201%0A%20%20%20%20_slope%2C%20_intercept%2C%20_%2C%20_%2C%20_%20%3D%20stats.linregress(_x_vals%2C%20_y_vals)%0A%20%20%20%20_line_x%20%3D%20np.array(%5B1%2C%205%5D)%0A%20%20%20%20_line_y%20%3D%20_slope%20*%20_line_x%20%2B%20_intercept%0A%20%20%20%20_fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3D_line_x%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3D_line_y%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%236366f1%22%2C%20width%3D2.5%2C%20dash%3D%22dash%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Linear%20Fit%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Subplot%202%3A%20Outlier%20Stress%20Test%0A%20%20%20%20%23%20Simulate%20a%20clean%20bivariate%20normal%20sample%20(n%3D25)%20with%20true%20correlation%20~%200.85%0A%20%20%20%20%23%20Then%20displace%20one%20single%20data%20point%20(point%200)%20along%20y%20from%20%2B3%20to%20-50%0A%20%20%20%20np.random.seed(101)%0A%20%20%20%20_base_x%20%3D%20np.linspace(1%2C%2010%2C%2025)%0A%20%20%20%20_base_y%20%3D%20_base_x%20%2B%20np.random.normal(0%2C%201.0%2C%2025)%0A%0A%20%20%20%20_outlier_displacements%20%3D%20np.linspace(10%2C%20-80%2C%2040)%0A%20%20%20%20_pearson_vals%20%3D%20%5B%5D%0A%20%20%20%20_spearman_vals%20%3D%20%5B%5D%0A%20%20%20%20_kendall_vals%20%3D%20%5B%5D%0A%0A%20%20%20%20for%20_disp%20in%20_outlier_displacements%3A%0A%20%20%20%20%20%20%20%20_curr_x%20%3D%20_base_x.copy()%0A%20%20%20%20%20%20%20%20_curr_y%20%3D%20_base_y.copy()%0A%20%20%20%20%20%20%20%20_curr_y%5B0%5D%20%3D%20_disp%20%20%23%20inject%20extreme%20outlier%0A%20%20%20%20%20%20%20%20_p_val%2C%20_%20%3D%20stats.pearsonr(_curr_x%2C%20_curr_y)%0A%20%20%20%20%20%20%20%20_s_val%2C%20_%20%3D%20stats.spearmanr(_curr_x%2C%20_curr_y)%0A%20%20%20%20%20%20%20%20_k_val%2C%20_%20%3D%20stats.kendalltau(_curr_x%2C%20_curr_y)%0A%20%20%20%20%20%20%20%20_pearson_vals.append(_p_val)%0A%20%20%20%20%20%20%20%20_spearman_vals.append(_s_val)%0A%20%20%20%20%20%20%20%20_kendall_vals.append(_k_val)%0A%0A%20%20%20%20_fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3D_outlier_displacements%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3D_pearson_vals%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%23ef4444%22%2C%20width%3D3)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Pearson%20r%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20_fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3D_outlier_displacements%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3D_spearman_vals%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%23f59e0b%22%2C%20width%3D2.5%2C%20dash%3D%22dot%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Spearman%20rho%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20_fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3D_outlier_displacements%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3D_kendall_vals%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%2310b981%22%2C%20width%3D3)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Kendall%20tau-b%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20_fig.update_layout(%0A%20%20%20%20%20%20%20%20template%3D%22plotly_white%22%2C%0A%20%20%20%20%20%20%20%20height%3D500%2C%0A%20%20%20%20%20%20%20%20title%3Ddict(%0A%20%20%20%20%20%20%20%20%20%20%20%20text%3D%22Kendall's%20Tau-b%3A%20Ordinal%20Agreement%20%26%20Outlier%20Resistance%20Dynamics%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3D0.5%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20xanchor%3D%22center%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20font%3Ddict(size%3D16%2C%20family%3D%22Inter%2C%20system-ui%2C%20sans-serif%22)%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20legend%3Ddict(orientation%3D%22h%22%2C%20yanchor%3D%22bottom%22%2C%20y%3D-0.22%2C%20xanchor%3D%22center%22%2C%20x%3D0.5)%2C%0A%20%20%20%20%20%20%20%20margin%3Ddict(l%3D50%2C%20r%3D50%2C%20t%3D80%2C%20b%3D80)%2C%0A%20%20%20%20)%0A%0A%20%20%20%20_fig.update_xaxes(%0A%20%20%20%20%20%20%20%20title_text%3D%22Job%20Satisfaction%20(1-5)%22%2C%0A%20%20%20%20%20%20%20%20tickvals%3D%5B1%2C%202%2C%203%2C%204%2C%205%5D%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%20%20%20%20_fig.update_yaxes(%0A%20%20%20%20%20%20%20%20title_text%3D%22Work-Life%20Balance%20(1-5)%22%2C%0A%20%20%20%20%20%20%20%20tickvals%3D%5B1%2C%202%2C%203%2C%204%2C%205%5D%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20_fig.update_xaxes(%0A%20%20%20%20%20%20%20%20title_text%3D%22Outlier%20Y-Coordinate%20(Perturbation)%22%2C%0A%20%20%20%20%20%20%20%20autorange%3D%22reversed%22%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%20%20%20%20_fig.update_yaxes(%0A%20%20%20%20%20%20%20%20title_text%3D%22Correlation%20Coefficient%20Value%22%2C%0A%20%20%20%20%20%20%20%20range%3D%5B-0.8%2C%201.05%5D%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%20%5Bd%5D%20Code%20Examples%0A%0A%20%20%20%20Below%20we%20implement%20two%20rigorous%2C%20production-grade%20demonstrations%3A%0A%20%20%20%201.%20**Step-by-step%20calculation%20from%20scratch**%3A%20An%20exhaustive%20counting%20of%20concordant%20pairs%20(%24C%24)%2C%20discordant%20pairs%20(%24D%24)%2C%20ties%20on%20%24X%24%20(%24T_X%24)%2C%20ties%20on%20%24Y%24%20(%24T_Y%24)%2C%20and%20joint%20ties%20(%24T_%7BXY%7D%24)%2C%20manually%20computing%20%24%5Ctau_a%24%2C%20%24%5Ctau_b%24%2C%20asymptotic%20standard%20error%2C%20%24z%24-score%2C%20and%20%24p%24-value%2C%20directly%20verified%20against%20%60scipy.stats.kendalltau%60.%0A%20%20%20%202.%20**LLM-as-a-Judge%20Alignment%20Matrix**%3A%205%20evaluators%20(Human%20Expert%2C%20GPT-4o%2C%20Claude-3.5-Sonnet%2C%20Gemini-1.5-Pro%2C%20and%20Llama-3-70B)%20evaluate%20and%20rank%2010%20complex%20model%20responses.%20We%20construct%20the%20pairwise%20Kendall's%20%24%5Ctau_b%24%20correlation%20matrix%20and%20evaluate%20which%20automated%20evaluator%20exhibits%20the%20highest%20ranking%20fidelity%20with%20human%20ground%20truth.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(df_sample%2C%20np%2C%20pd%2C%20stats)%3A%0A%20%20%20%20%23%20Example%201%3A%20Pure%20Python%20manual%20computation%20of%20C%2C%20D%2C%20T_X%2C%20T_Y%2C%20T_XY%2C%20and%20Tau-b%0A%20%20%20%20_x%20%3D%20df_sample%5B%22Job%20Satisfaction%22%5D.to_numpy()%0A%20%20%20%20_y%20%3D%20df_sample%5B%22Work-Life%20Balance%22%5D.to_numpy()%0A%20%20%20%20_n%20%3D%20len(_x)%0A%0A%20%20%20%20_c%20%3D%200%0A%20%20%20%20_d%20%3D%200%0A%20%20%20%20_tx%20%3D%200%0A%20%20%20%20_ty%20%3D%200%0A%20%20%20%20_txy%20%3D%200%0A%0A%20%20%20%20for%20_i%20in%20range(_n)%3A%0A%20%20%20%20%20%20%20%20for%20_j%20in%20range(_i%20%2B%201%2C%20_n)%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20_dx%20%3D%20_x%5B_i%5D%20-%20_x%5B_j%5D%0A%20%20%20%20%20%20%20%20%20%20%20%20_dy%20%3D%20_y%5B_i%5D%20-%20_y%5B_j%5D%0A%20%20%20%20%20%20%20%20%20%20%20%20_prod%20%3D%20_dx%20*%20_dy%0A%0A%20%20%20%20%20%20%20%20%20%20%20%20if%20_prod%20%3E%200%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20_c%20%2B%3D%201%0A%20%20%20%20%20%20%20%20%20%20%20%20elif%20_prod%20%3C%200%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20_d%20%2B%3D%201%0A%20%20%20%20%20%20%20%20%20%20%20%20elif%20_dx%20%3D%3D%200%20and%20_dy%20%3D%3D%200%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20_txy%20%2B%3D%201%0A%20%20%20%20%20%20%20%20%20%20%20%20elif%20_dx%20%3D%3D%200%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20_tx%20%2B%3D%201%0A%20%20%20%20%20%20%20%20%20%20%20%20elif%20_dy%20%3D%3D%200%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20_ty%20%2B%3D%201%0A%0A%20%20%20%20_total_pairs%20%3D%20_n%20*%20(_n%20-%201)%20%2F%2F%202%0A%0A%20%20%20%20%23%20Tau-a%0A%20%20%20%20_tau_a%20%3D%20(_c%20-%20_d)%20%2F%20_total_pairs%0A%0A%20%20%20%20%23%20Tau-b%0A%20%20%20%20_denom_x%20%3D%20_c%20%2B%20_d%20%2B%20_tx%0A%20%20%20%20_denom_y%20%3D%20_c%20%2B%20_d%20%2B%20_ty%0A%20%20%20%20_tau_b_scratch%20%3D%20(_c%20-%20_d)%20%2F%20np.sqrt(_denom_x%20*%20_denom_y)%0A%0A%20%20%20%20%23%20Scipy%20reference%20verification%0A%20%20%20%20_scipy_tau%2C%20_scipy_p%20%3D%20stats.kendalltau(_x%2C%20_y)%0A%0A%20%20%20%20%23%20Hypothesis%20test%20z-statistic%20using%20standard%20tie%20correction%0A%20%20%20%20%23%20Tie%20counts%0A%20%20%20%20_%2C%20_counts_x%20%3D%20np.unique(_x%2C%20return_counts%3DTrue)%0A%20%20%20%20_%2C%20_counts_y%20%3D%20np.unique(_y%2C%20return_counts%3DTrue)%0A%0A%20%20%20%20_t_term%20%3D%20np.sum(_counts_x%20*%20(_counts_x%20-%201)%20*%20(2%20*%20_counts_x%20%2B%205))%0A%20%20%20%20_u_term%20%3D%20np.sum(_counts_y%20*%20(_counts_y%20-%201)%20*%20(2%20*%20_counts_y%20%2B%205))%0A%20%20%20%20_v0%20%3D%20(_n%20*%20(_n%20-%201)%20*%20(2%20*%20_n%20%2B%205)%20-%20_t_term%20-%20_u_term)%20%2F%2018.0%0A%20%20%20%20_z_stat%20%3D%20(_c%20-%20_d)%20%2F%20np.sqrt(_v0)%0A%20%20%20%20_p_val_scratch%20%3D%202.0%20*%20(1.0%20-%20stats.norm.cdf(np.abs(_z_stat)))%0A%0A%20%20%20%20_summary_table%20%3D%20pd.DataFrame(%0A%20%20%20%20%20%20%20%20%5B%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%22Metric%22%3A%20%22Sample%20Size%20(n)%22%2C%20%22Computed%20Value%22%3A%20str(_n)%2C%20%22Formula%20%2F%20Description%22%3A%20%22Number%20of%20bivariate%20observations%22%7D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%22Metric%22%3A%20%22Total%20Unique%20Pairs%22%2C%20%22Computed%20Value%22%3A%20str(_total_pairs)%2C%20%22Formula%20%2F%20Description%22%3A%20%22n(n%20-%201)%20%2F%202%22%7D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%22Metric%22%3A%20%22Concordant%20Pairs%20(C)%22%2C%20%22Computed%20Value%22%3A%20str(_c)%2C%20%22Formula%20%2F%20Description%22%3A%20%22(x_i%20-%20x_j)(y_i%20-%20y_j)%20%3E%200%22%7D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%22Metric%22%3A%20%22Discordant%20Pairs%20(D)%22%2C%20%22Computed%20Value%22%3A%20str(_d)%2C%20%22Formula%20%2F%20Description%22%3A%20%22(x_i%20-%20x_j)(y_i%20-%20y_j)%20%3C%200%22%7D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%22Metric%22%3A%20%22Tied%20on%20X%20only%20(T_X)%22%2C%20%22Computed%20Value%22%3A%20str(_tx)%2C%20%22Formula%20%2F%20Description%22%3A%20%22x_i%20%3D%3D%20x_j%20and%20y_i%20!%3D%20y_j%22%7D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%22Metric%22%3A%20%22Tied%20on%20Y%20only%20(T_Y)%22%2C%20%22Computed%20Value%22%3A%20str(_ty)%2C%20%22Formula%20%2F%20Description%22%3A%20%22x_i%20!%3D%20x_j%20and%20y_i%20%3D%3D%20y_j%22%7D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%22Metric%22%3A%20%22Tied%20on%20Both%20(T_XY)%22%2C%20%22Computed%20Value%22%3A%20str(_txy)%2C%20%22Formula%20%2F%20Description%22%3A%20%22x_i%20%3D%3D%20x_j%20and%20y_i%20%3D%3D%20y_j%22%7D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%22Metric%22%3A%20%22Kendall's%20Tau-a%22%2C%20%22Computed%20Value%22%3A%20f%22%7B_tau_a%3A.4f%7D%22%2C%20%22Formula%20%2F%20Description%22%3A%20%22(C%20-%20D)%20%2F%20N_pairs%20(no%20tie%20adjustment)%22%7D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%22Metric%22%3A%20%22Kendall's%20Tau-b%20(Scratch)%22%2C%20%22Computed%20Value%22%3A%20f%22%7B_tau_b_scratch%3A.4f%7D%22%2C%20%22Formula%20%2F%20Description%22%3A%20%22(C%20-%20D)%20%2F%20sqrt((C%20%2B%20D%20%2B%20T_X)(C%20%2B%20D%20%2B%20T_Y))%22%7D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%22Metric%22%3A%20%22Kendall's%20Tau-b%20(SciPy)%22%2C%20%22Computed%20Value%22%3A%20f%22%7B_scipy_tau%3A.4f%7D%22%2C%20%22Formula%20%2F%20Description%22%3A%20%22scipy.stats.kendalltau(x%2C%20y)%22%7D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%22Metric%22%3A%20%22Asymptotic%20z-Score%22%2C%20%22Computed%20Value%22%3A%20f%22%7B_z_stat%3A.4f%7D%22%2C%20%22Formula%20%2F%20Description%22%3A%20%22(C%20-%20D)%20%2F%20sqrt(Var_0)%22%7D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%22Metric%22%3A%20%22Two-tailed%20p-value%22%2C%20%22Computed%20Value%22%3A%20f%22%7B_p_val_scratch%3A.6f%7D%22%2C%20%22Formula%20%2F%20Description%22%3A%20%222%20*%20(1%20-%20Phi(%7Cz%7C))%22%7D%2C%0A%20%20%20%20%20%20%20%20%5D%0A%20%20%20%20)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(np%2C%20pd%2C%20stats)%3A%0A%20%20%20%20%23%20Example%202%3A%20LLM%20Evaluation%20Agreement%20Matrix%20(Evaluating%205%20Judges%20Across%2010%20Model%20Outputs)%0A%20%20%20%20%23%20Ranks%20assigned%20to%2010%20generated%20responses%20(1%20%3D%20Best%20response%2C%2010%20%3D%20Worst%20response)%0A%20%20%20%20_judges_data%20%3D%20%7B%0A%20%20%20%20%20%20%20%20%22Human%20Expert%22%3A%20%5B1%2C%202%2C%203%2C%204%2C%205%2C%206%2C%207%2C%208%2C%209%2C%2010%5D%2C%0A%20%20%20%20%20%20%20%20%22GPT-4o%22%3A%20%5B1%2C%202%2C%204%2C%203%2C%205%2C%207%2C%206%2C%208%2C%2010%2C%209%5D%2C%0A%20%20%20%20%20%20%20%20%22Claude-3.5-Sonnet%22%3A%20%5B1%2C%202%2C%203%2C%205%2C%204%2C%206%2C%208%2C%207%2C%209%2C%2010%5D%2C%0A%20%20%20%20%20%20%20%20%22Gemini-1.5-Pro%22%3A%20%5B2%2C%201%2C%204%2C%203%2C%206%2C%205%2C%207%2C%209%2C%208%2C%2010%5D%2C%0A%20%20%20%20%20%20%20%20%22Llama-3-70B%22%3A%20%5B3%2C%201%2C%205%2C%202%2C%207%2C%204%2C%208%2C%206%2C%2010%2C%209%5D%2C%0A%20%20%20%20%7D%0A%20%20%20%20_df_judges%20%3D%20pd.DataFrame(_judges_data%2C%20index%3D%5Bf%22Prompt%20%7Bi%2B1%7D%22%20for%20i%20in%20range(10)%5D)%0A%0A%20%20%20%20_judge_names%20%3D%20list(_judges_data.keys())%0A%20%20%20%20_n_judges%20%3D%20len(_judge_names)%0A%20%20%20%20_tau_matrix%20%3D%20np.zeros((_n_judges%2C%20_n_judges))%0A%20%20%20%20_pval_matrix%20%3D%20np.zeros((_n_judges%2C%20_n_judges))%0A%0A%20%20%20%20for%20_i%20in%20range(_n_judges)%3A%0A%20%20%20%20%20%20%20%20for%20_j%20in%20range(_n_judges)%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20_t%2C%20_p%20%3D%20stats.kendalltau(_df_judges%5B_judge_names%5B_i%5D%5D%2C%20_df_judges%5B_judge_names%5B_j%5D%5D)%0A%20%20%20%20%20%20%20%20%20%20%20%20_tau_matrix%5B_i%2C%20_j%5D%20%3D%20_t%0A%20%20%20%20%20%20%20%20%20%20%20%20_pval_matrix%5B_i%2C%20_j%5D%20%3D%20_p%0A%0A%20%20%20%20_df_tau_matrix%20%3D%20pd.DataFrame(_tau_matrix%2C%20index%3D_judge_names%2C%20columns%3D_judge_names).round(4)%0A%0A%20%20%20%20%23%20Format%20human%20alignment%20leaderboard%0A%20%20%20%20_alignment_scores%20%3D%20%5B%5D%0A%20%20%20%20for%20_judge%20in%20_judge_names%5B1%3A%5D%3A%0A%20%20%20%20%20%20%20%20_t%20%3D%20_df_tau_matrix.loc%5B%22Human%20Expert%22%2C%20_judge%5D%0A%20%20%20%20%20%20%20%20_alignment_scores.append(%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Evaluator%20Model%22%3A%20_judge%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Kendall%20Tau-b%20with%20Human%22%3A%20f%22%7B_t%3A.4f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Agreement%20Status%22%3A%20%22High%20Alignment%20(%3E%200.80)%22%20if%20_t%20%3E%3D%200.80%20else%20%22Moderate%20Alignment%22%2C%0A%20%20%20%20%20%20%20%20%7D)%0A%0A%20%20%20%20_df_leaderboard%20%3D%20pd.DataFrame(_alignment_scores).sort_values(%22Kendall%20Tau-b%20with%20Human%22%2C%20ascending%3DFalse)%0A%20%20%20%20return%0A%0A%0Aif%20__name__%20%3D%3D%20%22__main__%22%3A%0A%20%20%20%20app.run()%0A
46bc259155a5b213f0da0124ee2b9d79