<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Paper Reading on Rui's Blog</title><link>https://low-hands.github.io/posts/paper/</link><description>Recent content in Paper Reading on Rui's Blog</description><generator>Hugo -- gohugo.io</generator><language>en</language><lastBuildDate>Sat, 14 Mar 2026 00:16:00 -0800</lastBuildDate><atom:link href="https://low-hands.github.io/posts/paper/index.xml" rel="self" type="application/rss+xml"/><item><title>Diffusion LLM paper reading</title><link>https://low-hands.github.io/posts/paper/dllm/</link><pubDate>Sat, 14 Mar 2026 00:16:00 -0800</pubDate><guid>https://low-hands.github.io/posts/paper/dllm/</guid><description>&lt;h1 id="1-residual-context-diffusion-language-models"&gt;1 Residual Context Diffusion Language Models&lt;/h1&gt;
&lt;h2 id="method"&gt;Method&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Traditional dLLMs discard the [MASK] at unpredicted positions in each round and reset them as new blank [MASK]s. To avoid wasting computational power, this method retains these discarded semantics and transforms them into vectors for the next round of prediction.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="algorithm"&gt;Algorithm&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Generate the residual vector by performing a weighted sum of the model&amp;rsquo;s predicted probability and the word embeddings:&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;$\Delta_i^{(t_k)} = \sum_{j=1}^V p_{i,j}^{(t_k)} E_{j,:}$&lt;/p&gt;</description></item><item><title>Yan:Foundational Interactive Video Generation</title><link>https://low-hands.github.io/posts/paper/yan/</link><pubDate>Mon, 09 Feb 2026 00:16:00 -0800</pubDate><guid>https://low-hands.github.io/posts/paper/yan/</guid><description>&lt;h1 id="abstract"&gt;Abstract&lt;/h1&gt;
&lt;h1 id="introduction"&gt;Introduction&lt;/h1&gt;
&lt;p&gt;The first paragraph primarily discusses the applications and deficiencies of Interactive Generative Video (IGV). AIGC is evolving from generating text and images to video synthesis, and has now progressed to IGV. IGV requires dynamic reactions to user input, with applications ranging from virtual simulation to embodied intelligence. However, the deficiencies of current methods lie in the lack of high visual fidelity, sustained temporal coherence, and rich interactivity. Generated content also remains static after creation, failing to adapt in real-time. (Not high-definition, not coherent, not interactive).&lt;/p&gt;</description></item></channel></rss>