import%20marimo%0A%0A__generated_with%20%3D%20%220.24.0%22%0Aapp%20%3D%20marimo.App()%0A%0A%0A%40app.cell%0Adef%20_()%3A%0A%20%20%20%20import%20marimo%20as%20mo%0A%20%20%20%20import%20numpy%20as%20np%0A%20%20%20%20import%20pandas%20as%20pd%0A%20%20%20%20import%20plotly.graph_objects%20as%20go%0A%20%20%20%20from%20plotly.subplots%20import%20make_subplots%0A%20%20%20%20from%20scipy.optimize%20import%20minimize_scalar%0A%0A%20%20%20%20return%20go%2C%20make_subplots%2C%20minimize_scalar%2C%20mo%2C%20np%2C%20pd%0A%0A%0A%40app.cell%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%5B%E2%86%90%2047%20GELU%20Activation%5D(47_gelu.py)%20%7C%20%5BIndex%5D(..%2Findex.html)%20%7C%20%5B49%20Focal%20Loss%20Balanced%20%E2%86%92%5D(49_focal_loss_balanced.py)%0A%0A%20%20%20%20%23%20Temperature-Scaled%20Softmax%3A%20Boltzmann%20Distributions%2C%20Model%20Calibration%2C%20and%20LLM%20Decoding%20Dynamics%0A%0A%20%20%20%20%23%23%20%5Ba%5D%20Why%20do%20you%20need%20to%20know%20these%20concepts%3F%0A%0A%20%20%20%20Every%20generative%20sequence%20model%E2%80%94from%20Large%20Language%20Models%20(such%20as%20GPT-4%2C%20Claude%2C%20Gemini%2C%20and%20LLaMA)%20predicting%20next-token%20probability%20distributions%20over%20vocabulary%20sizes%20exceeding%20%24100%2C000%24%20to%20reinforcement%20learning%20agents%20selecting%20discrete%20actions%E2%80%94outputs%20an%20unnormalized%20real-valued%20logits%20vector%20%24z%20%5Cin%20%5Cmathbb%7BR%7D%5EV%24.%0A%0A%20%20%20%20%23%23%23%23%20The%20Role%20of%20Temperature%20in%20Generative%20Decoding%0A%20%20%20%20In%20standard%20inference%2C%20applying%20the%20unmodified%20Softmax%20function%20(%24T%20%3D%201.0%24)%20converts%20raw%20logits%20directly%20into%20probabilities.%20However%2C%20direct%20sampling%20under%20%24T%20%3D%201.0%24%20often%20yields%20either%20repetitive%20sequences%20or%20wild%20hallucinations%20depending%20on%20the%20sharpness%20of%20the%20model's%20training%20distribution.%0A%0A%20%20%20%20Introducing%20a%20positive%20scalar%20parameter%20%24T%20%3E%200%24%20(the%20**temperature**)%20provides%20a%20continuous%20control%20mechanism%20over%20the%20entropy%20of%20the%20predicted%20distribution%3A%0A%20%20%20%20-%20**Low%20Temperature%20(%24T%20%5Cto%200%24)**%3A%20The%20distribution%20sharpens%20into%20a%20near-degenerate%20spike%20on%20the%20single%20highest%20logit.%20As%20%24T%20%5Cto%200%5E%2B%24%2C%20sampling%20converges%20to%20greedy%20deterministic%20argmax%20decoding%20(%24%5Clim_%7BT%20%5Cto%200%5E%2B%7D%20p_%7B%5Ctext%7Bmax%7D%7D%20%3D%201%24).%20Ideal%20for%20factual%20question-answering%2C%20code%20synthesis%2C%20and%20structured%20JSON%20generation.%0A%20%20%20%20-%20**Standard%20Temperature%20(%24T%20%3D%201.0%24)**%3A%20Preserves%20the%20native%20calibrated%20log-odds%20output%20by%20the%20model.%0A%20%20%20%20-%20**High%20Temperature%20(%24T%20%3E%201.0%24)**%3A%20Compresses%20differences%20between%20logits%2C%20flattening%20the%20distribution%20toward%20a%20uniform%20distribution.%20Higher%20temperatures%20increase%20Shannon%20entropy%2C%20introducing%20diversity%20and%20unexpected%20vocabulary%20selections%20for%20creative%20writing%20and%20brainstorming.%0A%0A%20%20%20%20%23%23%23%23%20Neural%20Network%20Probability%20Calibration%20(Guo%20et%20al.%2C%202017)%0A%20%20%20%20Modern%20deep%20neural%20networks%20with%20batch%20normalization%2C%20residual%20connections%2C%20and%20massive%20parameter%20counts%20are%20notoriously%20**overconfident**.%20A%20vision%20or%20classification%20model%20might%20output%20a%20predicted%20probability%20of%20%2498%5C%25%24%20on%20examples%20where%20its%20true%20empirical%20accuracy%20is%20only%20%2475%5C%25%24.%0A%0A%20%20%20%20Post-hoc%20**Temperature%20Scaling**%20provides%20a%20simple%20method%20for%20probability%20calibration%3A%20by%20optimizing%20a%20single%20scalar%20parameter%20%24T%20%3E%200%24%20to%20minimize%20validation%20negative%20log-likelihood%20(NLL)%2C%20the%20network's%20confidence%20is%20aligned%20with%20empirical%20frequency%20without%20altering%20classification%20accuracy%20or%20top-1%20predictions.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20%5Bb%5D%20Mathematical%20Foundations%20and%20Thermodynamic%20Mechanics%0A%0A%20%20%20%20%23%23%23%201.%20The%20Temperature-Scaled%20Softmax%20Function%0A%0A%20%20%20%20Let%20%24z%20%3D%20(z_1%2C%20z_2%2C%20%5Cdots%2C%20z_n)%5E%5Ctop%20%5Cin%20%5Cmathbb%7BR%7D%5En%24%20denote%20an%20unnormalized%20vector%20of%20logits.%20For%20any%20temperature%20%24T%20%3E%200%24%2C%20the%20temperature-scaled%20softmax%20operator%20%24%5Csigma_T%3A%20%5Cmathbb%7BR%7D%5En%20%5Cto%20%5CDelta%5E%7Bn-1%7D%24%20is%20defined%20as%3A%0A%0A%20%20%20%20%24%24p_i(T)%20%3D%20%5Csigma_T(z)_i%20%3D%20%5Cfrac%7B%5Cexp%5Cleft(%20%5Cfrac%7Bz_i%7D%7BT%7D%20%5Cright)%7D%7B%5Csum_%7Bj%3D1%7D%5En%20%5Cexp%5Cleft(%20%5Cfrac%7Bz_j%7D%7BT%7D%20%5Cright)%7D%2C%20%5Cquad%20%5Cforall%20i%20%5Cin%20%5C%7B1%2C%20%5Cdots%2C%20n%5C%7D%24%24%0A%0A%20%20%20%20For%20numerical%20stability%2C%20subtracting%20the%20maximum%20logit%20prevents%20floating-point%20overflow%3A%0A%0A%20%20%20%20%24%24p_i(T)%20%3D%20%5Cfrac%7B%5Cexp%5Cleft(%20%5Cfrac%7Bz_i%20-%20%5Cmax_k%20z_k%7D%7BT%7D%20%5Cright)%7D%7B%5Csum_%7Bj%3D1%7D%5En%20%5Cexp%5Cleft(%20%5Cfrac%7Bz_j%20-%20%5Cmax_k%20z_k%7D%7BT%7D%20%5Cright)%7D%24%24%0A%0A%20%20%20%20%23%23%23%202.%20Connection%20to%20Statistical%20Physics%20(Boltzmann%20Distribution)%0A%0A%20%20%20%20The%20temperature-scaled%20softmax%20is%20equivalent%20to%20the%20**Boltzmann%20(Gibbs)%20distribution**%20from%20thermodynamics.%20In%20statistical%20mechanics%2C%20the%20probability%20of%20a%20physical%20system%20occupying%20an%20energy%20state%20%24E_i%24%20at%20thermodynamic%20temperature%20%24%5Ctau%20%3D%20k_B%20T%24%20is%3A%0A%0A%20%20%20%20%24%24P(E_i)%20%3D%20%5Cfrac%7Be%5E%7B-E_i%20%2F%20(k_B%20T)%7D%7D%7B%5Cmathcal%7BZ%7D%7D%2C%20%5Cquad%20%5Ctext%7Bwhere%20%7D%20%5Cmathcal%7BZ%7D%20%3D%20%5Csum_%7Bj%3D1%7D%5En%20e%5E%7B-E_j%20%2F%20(k_B%20T)%7D%20%5Ctext%7B%20is%20the%20partition%20function%7D%24%24%0A%0A%20%20%20%20Setting%20energy%20to%20negative%20log-odds%20(%24E_i%20%3D%20-z_i%24)%20establishes%20direct%20mathematical%20equivalence%3A%20logits%20represent%20negative%20energy%20states%2C%20and%20the%20denominator%20represents%20the%20thermodynamic%20partition%20function.%0A%0A%20%20%20%20%23%23%23%203.%20Limiting%20Properties%0A%0A%20%20%20%20%23%23%23%23%20Zero%20Temperature%20Limit%20(%24T%20%5Cto%200%5E%2B%24%2C%20Argmax%20%2F%20Greedy)%0A%20%20%20%20Let%20%24%5Cmathcal%7BM%7D%20%3D%20%5Carg%5Cmax_%7Bk%7D%20z_k%24%20denote%20the%20set%20of%20maximal%20logit%20indices.%20Factoring%20out%20the%20maximum%20logit%20%24z_%7B%5Cmax%7D%24%3A%0A%0A%20%20%20%20%24%24%5Clim_%7BT%20%5Cto%200%5E%2B%7D%20p_i(T)%20%3D%20%5Cbegin%7Bcases%7D%20%5Cfrac%7B1%7D%7B%7C%5Cmathcal%7BM%7D%7C%7D%20%26%20%5Ctext%7Bif%20%7D%20i%20%5Cin%20%5Cmathcal%7BM%7D%20%5C%5C%200%20%26%20%5Ctext%7Bif%20%7D%20z_i%20%3C%20z_%7B%5Cmax%7D%20%5Cend%7Bcases%7D%24%24%0A%0A%20%20%20%20When%20the%20maximal%20logit%20is%20unique%20(%24%7C%5Cmathcal%7BM%7D%7C%20%3D%201%24)%2C%20the%20distribution%20collapses%20to%20a%20Kronecker%20delta%20centered%20on%20the%20argmax.%0A%0A%20%20%20%20%23%23%23%23%20Infinite%20Temperature%20Limit%20(%24T%20%5Cto%20%5Cinfty%24%2C%20Uniform%20Distribution)%0A%20%20%20%20As%20%24T%20%5Cto%20%5Cinfty%24%2C%20the%20exponent%20approaches%20zero%3A%20%24%5Clim_%7BT%20%5Cto%20%5Cinfty%7D%20%5Cfrac%7Bz_i%7D%7BT%7D%20%3D%200%24.%20Therefore%3A%0A%0A%20%20%20%20%24%24%5Clim_%7BT%20%5Cto%20%5Cinfty%7D%20p_i(T)%20%3D%20%5Cfrac%7Be%5E0%7D%7B%5Csum_%7Bj%3D1%7D%5En%20e%5E0%7D%20%3D%20%5Cfrac%7B1%7D%7Bn%7D%24%24%0A%0A%20%20%20%20%23%23%23%204.%20Entropy%20Dynamics%20and%20Partial%20Derivatives%0A%0A%20%20%20%20The%20Shannon%20entropy%20of%20the%20scaled%20distribution%20in%20bits%20is%3A%0A%0A%20%20%20%20%24%24H(p(T))%20%3D%20-%5Csum_%7Bi%3D1%7D%5En%20p_i(T)%20%5Clog_2%20p_i(T)%24%24%0A%0A%20%20%20%20For%20any%20non-degenerate%20logit%20vector%2C%20%24H(p(T))%24%20is%20a%20**strictly%20monotonically%20increasing%20function**%20of%20temperature%20%24T%24%3A%0A%0A%20%20%20%20%24%24%5Clim_%7BT%20%5Cto%200%5E%2B%7D%20H(p(T))%20%3D%200%20%5Ctext%7B%20bits%7D%2C%20%5Cquad%20%5Clim_%7BT%20%5Cto%20%5Cinfty%7D%20H(p(T))%20%3D%20%5Clog_2(n)%20%5Ctext%7B%20bits%7D%24%24%0A%0A%20%20%20%20Taking%20the%20partial%20derivative%20of%20probability%20%24p_i%24%20with%20respect%20to%20temperature%3A%0A%0A%20%20%20%20%24%24%5Cfrac%7B%5Cpartial%20p_i%7D%7B%5Cpartial%20T%7D%20%3D%20-%5Cfrac%7Bp_i%7D%7BT%5E2%7D%20%5Cleft(%20z_i%20-%20%5Csum_%7Bj%3D1%7D%5En%20p_j%20z_j%20%5Cright)%20%3D%20-%5Cfrac%7Bp_i%7D%7BT%5E2%7D%20%5Cleft(%20z_i%20-%20%5Cmathbb%7BE%7D_p%5Bz%5D%20%5Cright)%24%24%0A%0A%20%20%20%20-%20If%20token%20%24i%24%20has%20an%20above-average%20logit%20(%24z_i%20%3E%20%5Cmathbb%7BE%7D_p%5Bz%5D%24)%2C%20its%20probability%20**decreases**%20as%20%24T%24%20increases%20(%24%5Cfrac%7B%5Cpartial%20p_i%7D%7B%5Cpartial%20T%7D%20%3C%200%24).%0A%20%20%20%20-%20If%20token%20%24i%24%20has%20a%20below-average%20logit%20(%24z_i%20%3C%20%5Cmathbb%7BE%7D_p%5Bz%5D%24)%2C%20its%20probability%20**increases**%20as%20%24T%24%20increases%20(%24%5Cfrac%7B%5Cpartial%20p_i%7D%7B%5Cpartial%20T%7D%20%3E%200%24).%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(np)%3A%0A%20%20%20%20%23%20Candidate%20tokens%20simulating%20next-token%20completion%20for%20prompt%3A%20%22The%20capital%20of%20France%20is%22%0A%20%20%20%20token_vocab%20%3D%20%5B%22Paris%22%2C%20%22Lyon%22%2C%20%22Marseille%22%2C%20%22Europe%22%2C%20%22London%22%5D%0A%20%20%20%20%23%20Logits%20reflecting%20high%20model%20confidence%20in%20%22Paris%22%0A%20%20%20%20logits_arr%20%3D%20np.array(%5B4.8%2C%202.1%2C%201.2%2C%20-0.4%2C%20-1.8%5D)%0A%0A%20%20%20%20%23%20Temperature%20grid%20for%20entropy%20trajectory%0A%20%20%20%20temp_grid%20%3D%20np.linspace(0.08%2C%204.0%2C%20150)%0A%0A%20%20%20%20def%20stable_temp_softmax(logits%2C%20temp)%3A%0A%20%20%20%20%20%20%20%20scaled%20%3D%20(logits%20-%20np.max(logits))%20%2F%20temp%0A%20%20%20%20%20%20%20%20exp_z%20%3D%20np.exp(scaled)%0A%20%20%20%20%20%20%20%20return%20exp_z%20%2F%20np.sum(exp_z)%0A%0A%20%20%20%20%23%20Compute%20entropy%20curve%20across%20temperature%20range%0A%20%20%20%20entropy_trajectory%20%3D%20%5B%5D%0A%20%20%20%20prob_trajectories%20%3D%20%7Btok%3A%20%5B%5D%20for%20tok%20in%20token_vocab%7D%0A%0A%20%20%20%20for%20_t%20in%20temp_grid%3A%0A%20%20%20%20%20%20%20%20_p%20%3D%20stable_temp_softmax(logits_arr%2C%20_t)%0A%20%20%20%20%20%20%20%20%23%20Shannon%20entropy%20in%20bits%0A%20%20%20%20%20%20%20%20_ent%20%3D%20-np.sum(_p%20*%20np.log2(np.maximum(_p%2C%201e-15)))%0A%20%20%20%20%20%20%20%20entropy_trajectory.append(_ent)%0A%20%20%20%20%20%20%20%20for%20_idx%2C%20_tok%20in%20enumerate(token_vocab)%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20prob_trajectories%5B_tok%5D.append(_p%5B_idx%5D)%0A%0A%20%20%20%20entropy_trajectory%20%3D%20np.array(entropy_trajectory)%0A%20%20%20%20max_theoretical_entropy%20%3D%20float(np.log2(len(token_vocab)))%0A%0A%20%20%20%20%23%20Selected%20discrete%20temperatures%20for%20comparison%0A%20%20%20%20demo_temps%20%3D%20%5B0.25%2C%200.7%2C%201.0%2C%202.5%5D%0A%20%20%20%20demo_distributions%20%3D%20%7B_t%3A%20stable_temp_softmax(logits_arr%2C%20_t)%20for%20_t%20in%20demo_temps%7D%0A%20%20%20%20return%20(%0A%20%20%20%20%20%20%20%20demo_distributions%2C%0A%20%20%20%20%20%20%20%20demo_temps%2C%0A%20%20%20%20%20%20%20%20entropy_trajectory%2C%0A%20%20%20%20%20%20%20%20logits_arr%2C%0A%20%20%20%20%20%20%20%20max_theoretical_entropy%2C%0A%20%20%20%20%20%20%20%20stable_temp_softmax%2C%0A%20%20%20%20%20%20%20%20temp_grid%2C%0A%20%20%20%20%20%20%20%20token_vocab%2C%0A%20%20%20%20)%0A%0A%0A%40app.cell%0Adef%20_(%0A%20%20%20%20demo_distributions%2C%0A%20%20%20%20demo_temps%2C%0A%20%20%20%20entropy_trajectory%2C%0A%20%20%20%20go%2C%0A%20%20%20%20make_subplots%2C%0A%20%20%20%20max_theoretical_entropy%2C%0A%20%20%20%20mo%2C%0A%20%20%20%20temp_grid%2C%0A%20%20%20%20token_vocab%2C%0A)%3A%0A%20%20%20%20fig%20%3D%20make_subplots(%0A%20%20%20%20%20%20%20%20rows%3D1%2C%0A%20%20%20%20%20%20%20%20cols%3D2%2C%0A%20%20%20%20%20%20%20%20subplot_titles%3D%5B%0A%20%20%20%20%20%20%20%20%20%20%20%20%22%3Cb%3EToken%20Probability%20Distributions%20Across%20Temperature%20Levels%3C%2Fb%3E%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22%3Cb%3EShannon%20Entropy%20vs%20Temperature%20(Creativity%20Curve)%3C%2Fb%3E%22%2C%0A%20%20%20%20%20%20%20%20%5D%2C%0A%20%20%20%20%20%20%20%20horizontal_spacing%3D0.14%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Panel%201%3A%20Bar%20chart%20comparing%20temperatures%0A%20%20%20%20temp_colors%20%3D%20%7B0.25%3A%20%22%231D4ED8%22%2C%200.7%3A%20%22%232563EB%22%2C%201.0%3A%20%22%238B5CF6%22%2C%202.5%3A%20%22%23EC4899%22%7D%0A%0A%20%20%20%20for%20_t%20in%20demo_temps%3A%0A%20%20%20%20%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20%20%20%20%20go.Bar(%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20x%3Dtoken_vocab%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20y%3Ddemo_distributions%5B_t%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20name%3Df%22T%20%3D%20%7B_t%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20marker_color%3Dtemp_colors%5B_t%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20%20%20%20%20)%0A%0A%20%20%20%20%23%20Panel%202%3A%20Entropy%20Trajectory%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dtemp_grid%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dentropy_trajectory%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%232563EB%22%2C%20width%3D2.5)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Shannon%20Entropy%20H(T)%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Add%20theoretical%20maximum%20entropy%20line%0A%20%20%20%20fig.add_hline(%0A%20%20%20%20%20%20%20%20y%3Dmax_theoretical_entropy%2C%0A%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%23DC2626%22%2C%20width%3D1.5%2C%20dash%3D%22dash%22)%2C%0A%20%20%20%20%20%20%20%20annotation_text%3Df%22Max%20Entropy%20log2(5)%20%3D%20%7Bmax_theoretical_entropy%3A.2f%7D%20bits%22%2C%0A%20%20%20%20%20%20%20%20annotation_position%3D%22bottom%20right%22%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20fig.update_xaxes(title_text%3D%22Candidate%20Next%20Token%22%2C%20row%3D1%2C%20col%3D1)%0A%20%20%20%20fig.update_yaxes(title_text%3D%22Probability%20P(Token)%22%2C%20range%3D%5B0%2C%201.05%5D%2C%20row%3D1%2C%20col%3D1)%0A%20%20%20%20fig.update_xaxes(title_text%3D%22Temperature%20Parameter%20T%22%2C%20row%3D1%2C%20col%3D2)%0A%20%20%20%20fig.update_yaxes(title_text%3D%22Shannon%20Entropy%20(Bits)%22%2C%20range%3D%5B0%2C%202.5%5D%2C%20row%3D1%2C%20col%3D2)%0A%0A%20%20%20%20fig.update_layout(%0A%20%20%20%20%20%20%20%20template%3D%22plotly_white%22%2C%0A%20%20%20%20%20%20%20%20height%3D500%2C%0A%20%20%20%20%20%20%20%20margin%3Ddict(l%3D40%2C%20r%3D40%2C%20t%3D70%2C%20b%3D50)%2C%0A%20%20%20%20%20%20%20%20legend%3Ddict(orientation%3D%22h%22%2C%20yanchor%3D%22bottom%22%2C%20y%3D-0.28%2C%20xanchor%3D%22center%22%2C%20x%3D0.5)%2C%0A%20%20%20%20%20%20%20%20barmode%3D%22group%22%2C%0A%20%20%20%20)%0A%0A%20%20%20%20viz%20%3D%20mo.ui.plotly(fig)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_()%3A%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(%0A%20%20%20%20demo_distributions%2C%0A%20%20%20%20demo_temps%2C%0A%20%20%20%20logits_arr%2C%0A%20%20%20%20minimize_scalar%2C%0A%20%20%20%20mo%2C%0A%20%20%20%20np%2C%0A%20%20%20%20pd%2C%0A%20%20%20%20stable_temp_softmax%2C%0A%20%20%20%20token_vocab%2C%0A)%3A%0A%20%20%20%20%23%20Example%201%3A%20Pure%20NumPy%20Temperature-Scaled%20Softmax%20and%20Entropy%20Table%0A%20%20%20%20entropy_records%20%3D%20%5B%5D%0A%20%20%20%20for%20_t%20in%20demo_temps%3A%0A%20%20%20%20%20%20%20%20_probs%20%3D%20demo_distributions%5B_t%5D%0A%20%20%20%20%20%20%20%20_ent%20%3D%20-np.sum(_probs%20*%20np.log2(np.maximum(_probs%2C%201e-15)))%0A%20%20%20%20%20%20%20%20entropy_records.append(%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Temperature_T%22%3A%20_t%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Top_Token_Prob%20(Paris)%22%3A%20f%22%7B_probs%5B0%5D%20*%20100%3A.2f%7D%25%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Second_Token_Prob%20(Lyon)%22%3A%20f%22%7B_probs%5B1%5D%20*%20100%3A.2f%7D%25%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Lowest_Token_Prob%20(London)%22%3A%20f%22%7B_probs%5B4%5D%20*%20100%3A.2f%7D%25%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Shannon_Entropy_Bits%22%3A%20round(_ent%2C%203)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Decoding_Profile%22%3A%20(%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Deterministic%20%2F%20Greedy%22%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20if%20_t%20%3C%200.5%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20else%20(%22Balanced%20%2F%20Recommended%22%20if%20_t%20%3C%3D%201.0%20else%20%22High%20Entropy%20%2F%20Creative%22)%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7D%0A%20%20%20%20%20%20%20%20)%0A%0A%20%20%20%20df_entropy%20%3D%20pd.DataFrame(entropy_records)%0A%0A%20%20%20%20%23%20Example%202%3A%20Multinomial%20Sampling%20Simulation%20(10%2C000%20Generation%20Trials)%0A%20%20%20%20np.random.seed(42)%0A%20%20%20%20_n_trials%20%3D%2010000%0A%0A%20%20%20%20sim_records%20%3D%20%5B%5D%0A%20%20%20%20for%20_sim_t%20in%20%5B0.3%2C%201.0%2C%202.0%5D%3A%0A%20%20%20%20%20%20%20%20_probs%20%3D%20stable_temp_softmax(logits_arr%2C%20_sim_t)%0A%20%20%20%20%20%20%20%20_sampled_indices%20%3D%20np.random.choice(len(token_vocab)%2C%20size%3D_n_trials%2C%20p%3D_probs)%0A%20%20%20%20%20%20%20%20_counts%20%3D%20np.bincount(_sampled_indices%2C%20minlength%3Dlen(token_vocab))%0A%0A%20%20%20%20%20%20%20%20sim_records.append(%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Temperature%22%3A%20f%22T%20%3D%20%7B_sim_t%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Paris_Count%22%3A%20f%22%7B_counts%5B0%5D%7D%20(%7B_counts%5B0%5D%20%2F%20_n_trials%20*%20100%3A.1f%7D%25)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Lyon_Count%22%3A%20f%22%7B_counts%5B1%5D%7D%20(%7B_counts%5B1%5D%20%2F%20_n_trials%20*%20100%3A.1f%7D%25)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Marseille_Count%22%3A%20f%22%7B_counts%5B2%5D%7D%20(%7B_counts%5B2%5D%20%2F%20_n_trials%20*%20100%3A.1f%7D%25)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Europe_Count%22%3A%20f%22%7B_counts%5B3%5D%7D%20(%7B_counts%5B3%5D%20%2F%20_n_trials%20*%20100%3A.1f%7D%25)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22London_Count%22%3A%20f%22%7B_counts%5B4%5D%7D%20(%7B_counts%5B4%5D%20%2F%20_n_trials%20*%20100%3A.1f%7D%25)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7D%0A%20%20%20%20%20%20%20%20)%0A%0A%20%20%20%20df_sim%20%3D%20pd.DataFrame(sim_records)%0A%0A%20%20%20%20%23%20Example%203%3A%20Platt%20Temperature%20Scaling%20for%20Overconfident%20Neural%20Networks%0A%20%20%20%20%23%20Simulate%20an%20overconfident%20validation%20set%20of%20200%20samples%0A%20%20%20%20np.random.seed(42)%0A%20%20%20%20_n_val%20%3D%20200%0A%20%20%20%20%23%20Overconfident%20raw%20logits%3A%20large%20magnitude%20causing%20uncalibrated%2099%25%20confidence%0A%20%20%20%20raw_val_logits%20%3D%20np.random.normal(0%2C%201%2C%20(_n_val%2C%202))%0A%20%20%20%20raw_val_logits%5B%3A%2C%201%5D%20%2B%3D%20np.random.choice(%5B-3.5%2C%203.5%5D%2C%20size%3D_n_val)%20%20%23%20Extreme%20margins%0A%20%20%20%20val_labels%20%3D%20(raw_val_logits%5B%3A%2C%201%5D%20%3E%200).astype(int)%0A%20%20%20%20%23%20Introduce%20label%20noise%3A%2015%25%20random%20flips%20so%20true%20accuracy%20is%20~85%25%2C%20while%20confidence%20is%20~98%25%0A%20%20%20%20noise_mask%20%3D%20np.random.uniform(0%2C%201%2C%20_n_val)%20%3C%200.15%0A%20%20%20%20val_labels%5Bnoise_mask%5D%20%3D%201%20-%20val_labels%5Bnoise_mask%5D%0A%0A%20%20%20%20%23%20Uncalibrated%20baseline%20(T%20%3D%201.0)%0A%20%20%20%20uncal_probs%20%3D%20stable_temp_softmax(raw_val_logits.T%2C%201.0).T%0A%20%20%20%20uncal_conf%20%3D%20np.max(uncal_probs%2C%20axis%3D1)%0A%20%20%20%20uncal_acc%20%3D%20np.mean(np.argmax(uncal_probs%2C%20axis%3D1)%20%3D%3D%20val_labels)%0A%0A%20%20%20%20%23%20Loss%20function%20for%20temperature%20calibration%3A%20Negative%20Log-Likelihood%0A%20%20%20%20def%20nll_cost(temp)%3A%0A%20%20%20%20%20%20%20%20probs_t%20%3D%20stable_temp_softmax(raw_val_logits.T%2C%20temp).T%0A%20%20%20%20%20%20%20%20correct_probs%20%3D%20probs_t%5Bnp.arange(_n_val)%2C%20val_labels%5D%0A%20%20%20%20%20%20%20%20return%20-np.mean(np.log(np.maximum(correct_probs%2C%201e-15)))%0A%0A%20%20%20%20opt_res%20%3D%20minimize_scalar(nll_cost%2C%20bounds%3D(0.1%2C%2010.0)%2C%20method%3D%22bounded%22)%0A%20%20%20%20best_temp%20%3D%20float(opt_res.x)%0A%0A%20%20%20%20%23%20Calibrated%20probabilities%20under%20optimal%20T*%0A%20%20%20%20cal_probs%20%3D%20stable_temp_softmax(raw_val_logits.T%2C%20best_temp).T%0A%20%20%20%20cal_conf%20%3D%20np.max(cal_probs%2C%20axis%3D1)%0A%20%20%20%20cal_acc%20%3D%20np.mean(np.argmax(cal_probs%2C%20axis%3D1)%20%3D%3D%20val_labels)%0A%0A%20%20%20%20df_calibration%20%3D%20pd.DataFrame(%0A%20%20%20%20%20%20%20%20%5B%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Model_State%22%3A%20%22Uncalibrated%20(Default%20T%20%3D%201.0)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Optimal_Temperature_T%22%3A%20%221.00%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Mean_Model_Confidence%22%3A%20f%22%7Bnp.mean(uncal_conf)%20*%20100%3A.2f%7D%25%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Actual_Validation_Accuracy%22%3A%20f%22%7Buncal_acc%20*%20100%3A.2f%7D%25%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Confidence_Calibration_Gap%22%3A%20f%22%7Babs(np.mean(uncal_conf)%20-%20uncal_acc)%20*%20100%3A.2f%7D%25%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Negative_Log_Likelihood%22%3A%20f%22%7Bnll_cost(1.0)%3A.4f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Model_State%22%3A%20%22Temperature%20Scaled%20(Calibrated%20T*)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Optimal_Temperature_T%22%3A%20f%22%7Bbest_temp%3A.2f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Mean_Model_Confidence%22%3A%20f%22%7Bnp.mean(cal_conf)%20*%20100%3A.2f%7D%25%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Actual_Validation_Accuracy%22%3A%20f%22%7Bcal_acc%20*%20100%3A.2f%7D%25%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Confidence_Calibration_Gap%22%3A%20f%22%7Babs(np.mean(cal_conf)%20-%20cal_acc)%20*%20100%3A.2f%7D%25%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Negative_Log_Likelihood%22%3A%20f%22%7Bnll_cost(best_temp)%3A.4f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7D%2C%0A%20%20%20%20%20%20%20%20%5D%0A%20%20%20%20)%0A%0A%20%20%20%20table_entropy%20%3D%20mo.ui.table(df_entropy)%0A%20%20%20%20table_sim%20%3D%20mo.ui.table(df_sim)%0A%20%20%20%20table_cal%20%3D%20mo.ui.table(df_calibration)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_()%3A%0A%20%20%20%20return%0A%0A%0Aif%20__name__%20%3D%3D%20%22__main__%22%3A%0A%20%20%20%20app.run()%0A
66deb1675338c5d6a2b22f371bc11b6c