<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Attention on Bogdan Buduroiu</title><link>https://buduroiu.com/topics/attention/</link><description>Recent content in Attention on Bogdan Buduroiu</description><generator>Hugo</generator><language>en</language><copyright>All text licensed is licensed under CC BY-NC-SA 4.0 License.</copyright><lastBuildDate>Tue, 25 Aug 2026 17:28:59 +0300</lastBuildDate><atom:link href="https://buduroiu.com/topics/attention/index.xml" rel="self" type="application/rss+xml"/><item><title>22,580: GPT-2 to Kimi K3, explained</title><link>https://buduroiu.com/links/from-gpt2-kimi-k3/</link><pubDate>Tue, 25 Aug 2026 17:28:59 +0300</pubDate><guid>https://buduroiu.com/links/from-gpt2-kimi-k3/</guid><description>&lt;p&gt;MLA, KDA, MoE, NoPE all attack either bytes transferred or bytes resident. MLA shrinks KV per token, KDA replaces a &lt;span class="katex"&gt;&lt;span class="katex-mathml"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;O&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;O(n)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;&lt;span class="katex-html" aria-hidden="true"&gt;&lt;span class="base"&gt;&lt;span class="strut" style="height:1em;vertical-align:-0.25em;"&gt;&lt;/span&gt;&lt;span class="mord mathnormal" style="margin-right:0.02778em;"&gt;O&lt;/span&gt;&lt;span class="mopen"&gt;(&lt;/span&gt;&lt;span class="mord mathnormal"&gt;n&lt;/span&gt;&lt;span class="mclose"&gt;)&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; cache with an &lt;span class="katex"&gt;&lt;span class="katex-mathml"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;O&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;O(1)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;&lt;span class="katex-html" aria-hidden="true"&gt;&lt;span class="base"&gt;&lt;span class="strut" style="height:1em;vertical-align:-0.25em;"&gt;&lt;/span&gt;&lt;span class="mord mathnormal" style="margin-right:0.02778em;"&gt;O&lt;/span&gt;&lt;span class="mopen"&gt;(&lt;/span&gt;&lt;span class="mord"&gt;1&lt;/span&gt;&lt;span class="mclose"&gt;)&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; state.&lt;/p&gt;</description></item><item><title>No Positional Embeddings (NoPE)</title><link>https://buduroiu.com/links/nope-no-positional-encoding/</link><pubDate>Tue, 25 Aug 2026 17:25:38 +0300</pubDate><guid>https://buduroiu.com/links/nope-no-positional-encoding/</guid><description>&lt;p&gt;Why do we care about avoiding rotary positional encoding in the first place? RoPE aliasing is one reason, but the second reason is that RoPE basically erases the compression we gained from MLAs low-rank latent, and forces materialising and caching the full rotated key.&lt;/p&gt;</description></item></channel></rss>