import%20marimo%0A%0A__generated_with%20%3D%20%220.24.0%22%0Aapp%20%3D%20marimo.App()%0A%0A%0A%40app.cell%0Adef%20_()%3A%0A%20%20%20%20import%20marimo%20as%20mo%0A%20%20%20%20import%20numpy%20as%20np%0A%20%20%20%20import%20pandas%20as%20pd%0A%20%20%20%20import%20plotly.graph_objects%20as%20go%0A%20%20%20%20from%20plotly.subplots%20import%20make_subplots%0A%20%20%20%20import%20torch%0A%20%20%20%20import%20torch.nn%20as%20nn%0A%0A%20%20%20%20return%20go%2C%20make_subplots%2C%20mo%2C%20nn%2C%20np%2C%20pd%2C%20torch%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%5B%E2%86%90%2054%20Decoding%20Strategies%5D(54_decoding_strategies.py)%20%7C%20%5BIndex%5D(..%2Findex.html)%20%7C%20%5B56%20Reparameterization%20Trick%20%E2%86%92%5D(56_reparameterization_trick.py)%0A%0A%20%20%20%20%23%2055.%20Perplexity%20in%20Language%20Modeling%3A%20Information%20Entropy%2C%20Branching%20Factor%2C%20and%20Sequence%20Likelihood%0A%0A%20%20%20%20%23%23%23%20Executive%20Summary%0A%0A%20%20%20%20In%20generative%20language%20modeling%2C%20**Perplexity%20(PPL)**%20serves%20as%20the%20canonical%20intrinsic%20benchmark%20for%20evaluating%20how%20effectively%20a%20probability%20model%20%24P_%5Ctheta%24%20captures%20the%20syntax%2C%20semantics%2C%20and%20factual%20distribution%20of%20a%20text%20corpus.%20Mathematically%20formulated%20as%20the%20exponentiation%20of%20the%20cross-entropy%20loss%2C%20perplexity%20quantifies%20the%20model's%20average%20uncertainty%20per%20token.%0A%0A%20%20%20%20Intuitively%2C%20a%20perplexity%20of%20%24K%24%20signifies%20that%20the%20model%20is%20as%20uncertain%20at%20each%20prediction%20step%20as%20if%20it%20were%20choosing%20uniformly%20at%20random%20from%20%24K%24%20equally%20likely%20alternative%20tokens%E2%80%94often%20called%20the%20**effective%20branching%20factor**.%20A%20lower%20perplexity%20indicates%20higher%20predictive%20certainty%2C%20tighter%20probability%20concentration%20around%20ground-truth%20tokens%2C%20and%20superior%20language%20modeling%20capability.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20%5Bb%5D%20Mathematical%20Foundations%20and%20Information-Theoretic%20Derivations%0A%0A%20%20%20%20%23%23%23%201.%20The%20Cross-Entropy%20Formulation%0A%0A%20%20%20%20Let%20%24W%20%3D%20(w_1%2C%20w_2%2C%20%5Cdots%2C%20w_N)%24%20be%20a%20sequence%20of%20%24N%24%20tokens%20drawn%20from%20a%20test%20corpus.%20Under%20the%20autoregressive%20probability%20chain%20rule%2C%20the%20joint%20probability%20assigned%20to%20the%20sequence%20by%20model%20%24P_%5Ctheta%24%20is%3A%0A%0A%20%20%20%20%24%24P_%5Ctheta(W)%20%3D%20P_%5Ctheta(w_1%2C%20w_2%2C%20%5Cdots%2C%20w_N)%20%3D%20%5Cprod_%7Bi%3D1%7D%5EN%20P_%5Ctheta(w_i%20%5Cmid%20w_1%2C%20%5Cdots%2C%20w_%7Bi-1%7D)%24%24%0A%0A%20%20%20%20The%20empirical%20cross-entropy%20loss%20%24%5Cmathcal%7BH%7D(W%3B%20P_%5Ctheta)%24%20per%20token%20(in%20nats)%20is%20the%20negative%20log-likelihood%20averaged%20over%20sequence%20length%20%24N%24%3A%0A%0A%20%20%20%20%24%24%5Cmathcal%7BH%7D(W%3B%20P_%5Ctheta)%20%3D%20-%5Cfrac%7B1%7D%7BN%7D%20%5Cln%20P_%5Ctheta(W)%20%3D%20-%5Cfrac%7B1%7D%7BN%7D%20%5Csum_%7Bi%3D1%7D%5EN%20%5Cln%20P_%5Ctheta(w_i%20%5Cmid%20w_%7B%3Ci%7D)%24%24%0A%0A%20%20%20%20%23%23%23%202.%20Formal%20Definition%20of%20Perplexity%0A%0A%20%20%20%20Perplexity%20is%20defined%20as%20the%20exponentiated%20cross-entropy%20loss%3A%0A%0A%20%20%20%20%24%24%5Coperatorname%7BPPL%7D(W)%20%3D%20%5Cexp%5Cleft(%20%5Cmathcal%7BH%7D(W%3B%20P_%5Ctheta)%20%5Cright)%20%3D%20%5Cexp%5Cleft(%20-%5Cfrac%7B1%7D%7BN%7D%20%5Csum_%7Bi%3D1%7D%5EN%20%5Cln%20P_%5Ctheta(w_i%20%5Cmid%20w_%7B%3Ci%7D)%20%5Cright)%24%24%0A%0A%20%20%20%20Using%20the%20properties%20of%20logarithms%20and%20products%2C%20perplexity%20can%20be%20expressed%20equivalently%20as%20the%20inverse%20geometric%20mean%20of%20the%20token%20probabilities%3A%0A%0A%20%20%20%20%24%24%5Coperatorname%7BPPL%7D(W)%20%3D%20%5Cleft(%20%5Cprod_%7Bi%3D1%7D%5EN%20P_%5Ctheta(w_i%20%5Cmid%20w_%7B%3Ci%7D)%20%5Cright)%5E%7B-%5Cfrac%7B1%7D%7BN%7D%7D%20%3D%20%5Csqrt%5BN%5D%7B%5Cfrac%7B1%7D%7B%5Cprod_%7Bi%3D1%7D%5EN%20P_%5Ctheta(w_i%20%5Cmid%20w_%7B%3Ci%7D)%7D%7D%24%24%0A%0A%20%20%20%20If%20base-2%20logarithms%20are%20used%20(measuring%20information%20in%20bits)%3A%0A%0A%20%20%20%20%24%24%5Cmathcal%7BH%7D_2(W)%20%3D%20-%5Cfrac%7B1%7D%7BN%7D%20%5Csum_%7Bi%3D1%7D%5EN%20%5Clog_2%20P_%5Ctheta(w_i%20%5Cmid%20w_%7B%3Ci%7D)%2C%20%5Cqquad%20%5Coperatorname%7BPPL%7D(W)%20%3D%202%5E%7B%5Cmathcal%7BH%7D_2(W)%7D%24%24%0A%0A%20%20%20%20%23%23%23%203.%20The%20Branching%20Factor%20Interpretation%0A%0A%20%20%20%20To%20build%20rigorous%20intuition%20for%20perplexity%2C%20consider%20two%20boundary%20cases%20on%20a%20vocabulary%20%24%5Cmathcal%7BV%7D%24%20of%20size%20%24V%24%3A%0A%0A%20%20%20%201.%20**Perfect%20Model%20(Zero%20Uncertainty)**%3A%0A%20%20%20%20If%20the%20model%20predicts%20the%20correct%20token%20with%20%24100%5C%25%24%20confidence%20at%20every%20step%20(%24P_%5Ctheta(w_i%20%5Cmid%20w_%7B%3Ci%7D)%20%3D%201.0%24%20for%20all%20%24i%24)%3A%0A%0A%20%20%20%20%24%24%5Cmathcal%7BH%7D(W)%20%3D%200%20%5Cimplies%20%5Coperatorname%7BPPL%7D(W)%20%3D%20%5Cexp(0)%20%3D%201.0%24%24%0A%0A%20%20%20%20A%20perplexity%20of%20%241.0%24%20is%20the%20theoretical%20minimum%2C%20signifying%20zero%20surprise.%0A%0A%20%20%20%202.%20**Uniform%20Random%20Baseline%20(Maximum%20Entropy)**%3A%0A%20%20%20%20If%20the%20model%20has%20learned%20nothing%20and%20assigns%20equal%20probability%20%241%2FV%24%20to%20every%20token%20in%20the%20vocabulary%3A%0A%0A%20%20%20%20%24%24%5Cmathcal%7BH%7D(W)%20%3D%20-%5Cfrac%7B1%7D%7BN%7D%20%5Csum_%7Bi%3D1%7D%5EN%20%5Cln%5Cleft(%5Cfrac%7B1%7D%7BV%7D%5Cright)%20%3D%20%5Cln%20V%20%5Cimplies%20%5Coperatorname%7BPPL%7D(W)%20%3D%20%5Cexp(%5Cln%20V)%20%3D%20V%24%24%0A%0A%20%20%20%20The%20perplexity%20equals%20the%20entire%20vocabulary%20size.%0A%0A%20%20%20%203.%20**General%20Meaning%20of%20%24%5Coperatorname%7BPPL%7D%20%3D%20K%24**%3A%0A%20%20%20%20%20%20%20If%20a%20model%20achieves%20%24%5Coperatorname%7BPPL%7D%20%3D%2012.4%24%20on%20a%20test%20set%2C%20it%20means%20that%20predicting%20each%20next%20token%20is%2C%20on%20average%2C%20as%20difficult%20for%20the%20model%20as%20choosing%20between%20%2412.4%24%20equally%20likely%20candidate%20words.%0A%0A%20%20%20%20%23%23%23%204.%20Cross-Tokenizer%20Comparison%3A%20Bits%20Per%20Byte%20(BPB)%0A%0A%20%20%20%20A%20critical%20error%20in%20NLP%20evaluation%20is%20comparing%20raw%20perplexity%20across%20models%20that%20use%20different%20tokenizers%20(e.g.%2C%20LLaMA-3%20with%20%24128%5Ctext%7Bk%7D%24%20BPE%20tokens%20vs%20GPT-2%20with%20%2450%5Ctext%7Bk%7D%24%20BPE%20tokens).%20A%20model%20with%20larger%20tokens%20will%20have%20fewer%20tokens%20per%20sentence%2C%20inflating%20per-token%20cross-entropy%20while%20reducing%20sequence%20length%20%24N%24.%0A%0A%20%20%20%20To%20ensure%20a%20mathematically%20fair%20comparison%20invariant%20to%20tokenization%20granularity%2C%20evaluations%20are%20normalized%20by%20total%20UTF-8%20byte%20count%20%24B%24%2C%20computing%20**Bits%20Per%20Byte%20(BPB)**%3A%0A%0A%20%20%20%20%24%24%5Ctext%7BBPB%7D%20%3D%20%5Cfrac%7B%5Csum_%7Bi%3D1%7D%5EN%20-%5Clog_2%20P_%5Ctheta(w_i%20%5Cmid%20w_%7B%3Ci%7D)%7D%7B%5Ctext%7BTotal%20Bytes%20%7D%20B%7D%20%3D%20%5Cfrac%7BN%20%5Ccdot%20%5Cmathcal%7BH%7D_2(W)%7D%7BB%7D%24%24%0A%0A%20%20%20%20Per-byte%20perplexity%20is%20then%3A%0A%0A%20%20%20%20%24%24%5Coperatorname%7BPPL%7D_%7B%5Ctext%7Bbyte%7D%7D%20%3D%202%5E%7B%5Ctext%7BBPB%7D%7D%24%24%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(go%2C%20make_subplots%2C%20mo%2C%20np)%3A%0A%20%20%20%20%23%20Simulated%20sentence%20demonstrating%20in-context%20perplexity%20resolution%0A%20%20%20%20%23%20As%20context%20grows%20from%20token%201%20to%20token%2010%2C%20conditional%20probability%20increases%0A%20%20%20%20test_sentence_tokens%20%3D%20%5B%0A%20%20%20%20%20%20%20%20%22Artificial%22%2C%0A%20%20%20%20%20%20%20%20%22intelligence%22%2C%0A%20%20%20%20%20%20%20%20%22models%22%2C%0A%20%20%20%20%20%20%20%20%22predict%22%2C%0A%20%20%20%20%20%20%20%20%22next%22%2C%0A%20%20%20%20%20%20%20%20%22tokens%22%2C%0A%20%20%20%20%20%20%20%20%22by%22%2C%0A%20%20%20%20%20%20%20%20%22minimizing%22%2C%0A%20%20%20%20%20%20%20%20%22cross%22%2C%0A%20%20%20%20%20%20%20%20%22entropy%22%2C%0A%20%20%20%20%20%20%20%20%22loss%22%2C%0A%20%20%20%20%5D%0A%20%20%20%20n_tokens%20%3D%20len(test_sentence_tokens)%0A%0A%20%20%20%20%23%20Simulated%20realistic%20probabilities%20reflecting%20linguistic%20constraint%20buildup%0A%20%20%20%20%23%20Token%201%20(%22Artificial%22)%3A%20ambiguous%20start%20-%3E%20p%20%3D%200.04%20(PPL%20%3D%2025.0)%0A%20%20%20%20%23%20Token%202%20(%22intelligence%22)%3A%20heavily%20conditioned%20by%20%22Artificial%22%20-%3E%20p%20%3D%200.72%20(PPL%20%3D%201.38)%0A%20%20%20%20%23%20Token%203%20(%22models%22)%3A%20moderately%20likely%20-%3E%20p%20%3D%200.28%0A%20%20%20%20%23%20Token%204%20(%22predict%22)%3A%20verb%20in%20AI%20context%20-%3E%20p%20%3D%200.45%0A%20%20%20%20%23%20Token%205%20(%22next%22)%3A%20idiomatic%20-%3E%20p%20%3D%200.65%0A%20%20%20%20%23%20Token%206%20(%22tokens%22)%3A%20highly%20constrained%20-%3E%20p%20%3D%200.82%0A%20%20%20%20%23%20Token%207%20(%22by%22)%3A%20preposition%20-%3E%20p%20%3D%200.55%0A%20%20%20%20%23%20Token%208%20(%22minimizing%22)%3A%20technical%20-%3E%20p%20%3D%200.38%0A%20%20%20%20%23%20Token%209%20(%22cross%22)%3A%20technical%20-%3E%20p%20%3D%200.75%0A%20%20%20%20%23%20Token%2010%20(%22entropy%22)%3A%20bound%20bigram%20-%3E%20p%20%3D%200.94%0A%20%20%20%20%23%20Token%2011%20(%22loss%22)%3A%20bound%20trigram%20-%3E%20p%20%3D%200.98%0A%20%20%20%20token_probs%20%3D%20np.array(%5B0.04%2C%200.72%2C%200.28%2C%200.45%2C%200.65%2C%200.82%2C%200.55%2C%200.38%2C%200.75%2C%200.94%2C%200.98%5D)%0A%20%20%20%20step_nll%20%3D%20-np.log(token_probs)%0A%0A%20%20%20%20%23%20Cumulative%20running%20perplexity%3A%20exp(1%2Ft%20*%20sum_%7Bi%3D1%7D%5Et%20-ln%20p_i)%0A%20%20%20%20cumulative_nll%20%3D%20np.cumsum(step_nll)%20%2F%20np.arange(1%2C%20n_tokens%20%2B%201)%0A%20%20%20%20cumulative_ppl%20%3D%20np.exp(cumulative_nll)%0A%20%20%20%20instantaneous_ppl%20%3D%201.0%20%2F%20token_probs%0A%0A%20%20%20%20%23%20Panel%202%20data%3A%20Theoretical%20Branching%20Factor%20vs%20Cross-Entropy%0A%20%20%20%20entropy_grid%20%3D%20np.linspace(0.0%2C%206.0%2C%20200)%0A%20%20%20%20ppl_nats%20%3D%20np.exp(entropy_grid)%0A%20%20%20%20ppl_bits%20%3D%202.0**entropy_grid%0A%0A%20%20%20%20fig%20%3D%20make_subplots(%0A%20%20%20%20%20%20%20%20rows%3D1%2C%0A%20%20%20%20%20%20%20%20cols%3D2%2C%0A%20%20%20%20%20%20%20%20subplot_titles%3D%5B%0A%20%20%20%20%20%20%20%20%20%20%20%20%22%3Cb%3EToken-Level%20Surprise%20and%20Cumulative%20Perplexity%20Trajectory%3C%2Fb%3E%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22%3Cb%3EBranching%20Factor%20Growth%3A%20PPL%20%3D%20exp(H)%20and%202%5EH%3C%2Fb%3E%22%2C%0A%20%20%20%20%20%20%20%20%5D%2C%0A%20%20%20%20%20%20%20%20horizontal_spacing%3D0.14%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Panel%201%3A%20Instantaneous%20vs%20Cumulative%20Perplexity%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Bar(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dtest_sentence_tokens%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dinstantaneous_ppl%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Instantaneous%20Token%20Surprise%20(1%2Fp_i)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20marker_color%3D%22%2393C5FD%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20opacity%3D0.7%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dtest_sentence_tokens%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dcumulative_ppl%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%2Bmarkers%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%231D4ED8%22%2C%20width%3D3.0)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20marker%3Ddict(size%3D8)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22Cumulative%20Sequence%20PPL%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D1%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Panel%202%3A%20Exponential%20Branching%20Curve%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dentropy_grid%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dppl_nats%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%23DC2626%22%2C%20width%3D2.5)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22PPL%20%3D%20exp(H_nats)%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%20%20%20%20fig.add_trace(%0A%20%20%20%20%20%20%20%20go.Scatter(%0A%20%20%20%20%20%20%20%20%20%20%20%20x%3Dentropy_grid%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20y%3Dppl_bits%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20mode%3D%22lines%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20line%3Ddict(color%3D%22%230D9488%22%2C%20width%3D2.5%2C%20dash%3D%22dash%22)%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20name%3D%22PPL%20%3D%202%5E(H_bits)%22%2C%0A%20%20%20%20%20%20%20%20)%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Annotate%20typical%20LLM%20benchmark%20regime%0A%20%20%20%20fig.add_vrect(%0A%20%20%20%20%20%20%20%20x0%3D2.0%2C%0A%20%20%20%20%20%20%20%20x1%3D3.0%2C%0A%20%20%20%20%20%20%20%20fillcolor%3D%22%23FEF3C7%22%2C%0A%20%20%20%20%20%20%20%20opacity%3D0.5%2C%0A%20%20%20%20%20%20%20%20layer%3D%22below%22%2C%0A%20%20%20%20%20%20%20%20line_width%3D0%2C%0A%20%20%20%20%20%20%20%20annotation_text%3D%22Typical%20LLM%20Regime%20(PPL%207-20)%22%2C%0A%20%20%20%20%20%20%20%20annotation_position%3D%22top%20left%22%2C%0A%20%20%20%20%20%20%20%20row%3D1%2C%0A%20%20%20%20%20%20%20%20col%3D2%2C%0A%20%20%20%20)%0A%0A%20%20%20%20fig.update_xaxes(title_text%3D%22Sequence%20Tokens%20(In-Context%20Horizon)%22%2C%20tickangle%3D-35%2C%20row%3D1%2C%20col%3D1)%0A%20%20%20%20fig.update_yaxes(title_text%3D%22Perplexity%20%2F%20Effective%20Branching%22%2C%20range%3D%5B0%2C%2030%5D%2C%20row%3D1%2C%20col%3D1)%0A%20%20%20%20fig.update_xaxes(title_text%3D%22Cross-Entropy%20Loss%20H%22%2C%20row%3D1%2C%20col%3D2)%0A%20%20%20%20fig.update_yaxes(title_text%3D%22Perplexity%20(Effective%20Choices)%22%2C%20range%3D%5B1%2C%20150%5D%2C%20row%3D1%2C%20col%3D2)%0A%0A%20%20%20%20fig.update_layout(%0A%20%20%20%20%20%20%20%20template%3D%22plotly_white%22%2C%0A%20%20%20%20%20%20%20%20height%3D520%2C%0A%20%20%20%20%20%20%20%20margin%3Ddict(l%3D40%2C%20r%3D40%2C%20t%3D70%2C%20b%3D60)%2C%0A%20%20%20%20%20%20%20%20legend%3Ddict(orientation%3D%22h%22%2C%20yanchor%3D%22bottom%22%2C%20y%3D-0.32%2C%20xanchor%3D%22center%22%2C%20x%3D0.5)%2C%0A%20%20%20%20)%0A%0A%20%20%20%20viz%20%3D%20mo.ui.plotly(fig)%0A%20%20%20%20return%20cumulative_ppl%2C%20test_sentence_tokens%2C%20token_probs%0A%0A%0A%40app.cell%0Adef%20_()%3A%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(%0A%20%20%20%20cumulative_ppl%2C%0A%20%20%20%20mo%2C%0A%20%20%20%20nn%2C%0A%20%20%20%20np%2C%0A%20%20%20%20pd%2C%0A%20%20%20%20test_sentence_tokens%2C%0A%20%20%20%20token_probs%2C%0A%20%20%20%20torch%2C%0A)%3A%0A%20%20%20%20%23%20Vectorized%20NumPy%20Perplexity%20Implementation%0A%20%20%20%20def%20calculate_perplexity_numpy(probabilities)%3A%0A%20%20%20%20%20%20%20%20%22%22%22Calculates%20exact%20perplexity%20from%20an%20array%20of%20conditional%20token%20probabilities.%0A%0A%20%20%20%20%20%20%20%20PPL%20%3D%20exp(-1%2FN%20*%20sum(ln%20p_i))%0A%20%20%20%20%20%20%20%20%22%22%22%0A%20%20%20%20%20%20%20%20probs%20%3D%20np.asarray(probabilities)%0A%20%20%20%20%20%20%20%20nll%20%3D%20-np.log(np.maximum(probs%2C%201e-15))%0A%20%20%20%20%20%20%20%20mean_nll%20%3D%20np.mean(nll)%0A%20%20%20%20%20%20%20%20ppl%20%3D%20np.exp(mean_nll)%0A%20%20%20%20%20%20%20%20return%20ppl%2C%20mean_nll%0A%0A%20%20%20%20%23%20PyTorch%20CrossEntropyLoss%20equivalence%20verification%0A%20%20%20%20torch.manual_seed(42)%0A%20%20%20%20_vocab_size%20%3D%2050%0A%20%20%20%20_seq_len%20%3D%208%0A%20%20%20%20%23%20Random%20unnormalized%20logits%3A%20(seq_len%2C%20vocab_size)%0A%20%20%20%20_dummy_logits%20%3D%20torch.randn(_seq_len%2C%20_vocab_size)%0A%20%20%20%20%23%20Random%20target%20token%20indices%0A%20%20%20%20_dummy_targets%20%3D%20torch.randint(0%2C%20_vocab_size%2C%20(_seq_len%2C))%0A%0A%20%20%20%20%23%201.%20PyTorch%20Standard%20CrossEntropyLoss%0A%20%20%20%20criterion%20%3D%20nn.CrossEntropyLoss()%0A%20%20%20%20ce_loss%20%3D%20criterion(_dummy_logits%2C%20_dummy_targets).item()%0A%20%20%20%20ppl_pytorch%20%3D%20np.exp(ce_loss)%0A%0A%20%20%20%20%23%202.%20NumPy%20from%20Softmax%20Probabilities%0A%20%20%20%20_probs_all%20%3D%20torch.softmax(_dummy_logits%2C%20dim%3D-1).numpy()%0A%20%20%20%20_target_probs%20%3D%20_probs_all%5Bnp.arange(_seq_len)%2C%20_dummy_targets.numpy()%5D%0A%20%20%20%20ppl_numpy%2C%20mean_nll_np%20%3D%20calculate_perplexity_numpy(_target_probs)%0A%0A%20%20%20%20diff_verif%20%3D%20abs(ppl_pytorch%20-%20ppl_numpy)%0A%0A%20%20%20%20df_equivalence%20%3D%20pd.DataFrame(%0A%20%20%20%20%20%20%20%20%5B%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Calculation_Method%22%3A%20%22PyTorch%20nn.CrossEntropyLoss%20Exponentiation%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Cross_Entropy_Loss%22%3A%20f%22%7Bce_loss%3A.6f%7D%20nats%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Computed_Perplexity%22%3A%20f%22%7Bppl_pytorch%3A.6f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Verification_Status%22%3A%20%22Reference%20Implementation%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Calculation_Method%22%3A%20%22Vectorized%20NumPy%20Inverse%20Geometric%20Mean%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Cross_Entropy_Loss%22%3A%20f%22%7Bmean_nll_np%3A.6f%7D%20nats%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Computed_Perplexity%22%3A%20f%22%7Bppl_numpy%3A.6f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Verification_Status%22%3A%20f%22Bitwise%20Match%20(Diff%3A%20%7Bdiff_verif%3A.2e%7D)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7D%2C%0A%20%20%20%20%20%20%20%20%5D%0A%20%20%20%20)%0A%0A%20%20%20%20%23%20Example%202%3A%20In-Context%20Token%20Breakdown%20Table%0A%20%20%20%20token_records%20%3D%20%5B%5D%0A%20%20%20%20for%20_idx%2C%20_tok%20in%20enumerate(test_sentence_tokens)%3A%0A%20%20%20%20%20%20%20%20_p%20%3D%20token_probs%5B_idx%5D%0A%20%20%20%20%20%20%20%20token_records.append(%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Position%22%3A%20_idx%20%2B%201%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Token%22%3A%20_tok%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Conditional_Probability%22%3A%20f%22%7B_p%20*%20100%3A.1f%7D%25%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Surprise_NLL_nats%22%3A%20f%22%7B-np.log(_p)%3A.3f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Instantaneous_Branching%22%3A%20f%22%7B1.0%20%2F%20_p%3A.2f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Cumulative_Sequence_PPL%22%3A%20f%22%7Bcumulative_ppl%5B_idx%5D%3A.2f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7D%0A%20%20%20%20%20%20%20%20)%0A%0A%20%20%20%20df_tokens%20%3D%20pd.DataFrame(token_records)%0A%0A%20%20%20%20%23%20Example%203%3A%20Cross-Tokenizer%20Bits%20Per%20Byte%20(BPB)%20Normalization%20Simulation%0A%20%20%20%20%23%20Target%20phrase%3A%20%22The%20quick%20brown%20fox%20jumps%20over%20the%20lazy%20dog%22%20(43%20UTF-8%20bytes)%0A%20%20%20%20target_text%20%3D%20%22The%20quick%20brown%20fox%20jumps%20over%20the%20lazy%20dog%22%0A%20%20%20%20byte_count%20%3D%20len(target_text.encode(%22utf-8%22))%0A%0A%20%20%20%20tokenizer_simulations%20%3D%20%5B%0A%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Tokenizer_Level%22%3A%20%22Character-Level%20Tokenizer%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Token_Count_N%22%3A%2043%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Per_Token_Loss_bits%22%3A%201.45%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Total_Bits%22%3A%2043%20*%201.45%2C%0A%20%20%20%20%20%20%20%20%7D%2C%0A%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Tokenizer_Level%22%3A%20%22Subword%20BPE%20(32k%20Vocab)%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Token_Count_N%22%3A%209%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Per_Token_Loss_bits%22%3A%206.80%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Total_Bits%22%3A%209%20*%206.80%2C%0A%20%20%20%20%20%20%20%20%7D%2C%0A%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Tokenizer_Level%22%3A%20%22Word-Level%20Tokenizer%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Token_Count_N%22%3A%209%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Per_Token_Loss_bits%22%3A%207.10%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%22Total_Bits%22%3A%209%20*%207.10%2C%0A%20%20%20%20%20%20%20%20%7D%2C%0A%20%20%20%20%5D%0A%0A%20%20%20%20bpb_records%20%3D%20%5B%5D%0A%20%20%20%20for%20entry%20in%20tokenizer_simulations%3A%0A%20%20%20%20%20%20%20%20tot_bits%20%3D%20entry%5B%22Total_Bits%22%5D%0A%20%20%20%20%20%20%20%20bpb%20%3D%20tot_bits%20%2F%20byte_count%0A%20%20%20%20%20%20%20%20token_ppl%20%3D%202.0%20**%20entry%5B%22Per_Token_Loss_bits%22%5D%0A%20%20%20%20%20%20%20%20byte_ppl%20%3D%202.0**bpb%0A%20%20%20%20%20%20%20%20bpb_records.append(%0A%20%20%20%20%20%20%20%20%20%20%20%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Tokenizer_Type%22%3A%20entry%5B%22Tokenizer_Level%22%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Tokens_per_Sentence%22%3A%20entry%5B%22Token_Count_N%22%5D%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Raw_Token_Perplexity%22%3A%20f%22%7Btoken_ppl%3A.2f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Total_Information_Bits%22%3A%20f%22%7Btot_bits%3A.1f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Bits_Per_Byte%20(BPB)%22%3A%20f%22%7Bbpb%3A.3f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Normalized_Byte_Perplexity%22%3A%20f%22%7Bbyte_ppl%3A.3f%7D%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%22Evaluation_Fairness%22%3A%20%22Comparable%20Across%20Tokenizers%22%20if%20%22BPB%22%20in%20%22Bits_Per_Byte%22%20else%20%22%22%2C%0A%20%20%20%20%20%20%20%20%20%20%20%20%7D%0A%20%20%20%20%20%20%20%20)%0A%0A%20%20%20%20df_bpb%20%3D%20pd.DataFrame(bpb_records)%0A%0A%20%20%20%20table_equiv%20%3D%20mo.ui.table(df_equivalence)%0A%20%20%20%20table_tokens%20%3D%20mo.ui.table(df_tokens)%0A%20%20%20%20table_bpb%20%3D%20mo.ui.table(df_bpb)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_()%3A%0A%20%20%20%20return%0A%0A%0Aif%20__name__%20%3D%3D%20%22__main__%22%3A%0A%20%20%20%20app.run()%0A
466fe14696558316157774a79f598cb9