<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Paper on moyutianzun 的博客</title><link>https://moyutianzun.com/tags/paper/</link><description>Recent content in Paper on moyutianzun 的博客</description><generator>Hugo</generator><language>zh-cn</language><copyright>moyutianzun</copyright><lastBuildDate>Tue, 02 Dec 2025 09:28:32 +0800</lastBuildDate><atom:link href="https://moyutianzun.com/tags/paper/index.xml" rel="self" type="application/rss+xml"/><item><title>context engine note</title><link>https://moyutianzun.com/blog/context-engine-note/</link><pubDate>Tue, 02 Dec 2025 09:28:32 +0800</pubDate><guid>https://moyutianzun.com/blog/context-engine-note/</guid><description>&lt;h1 style="" id="context-manager"&gt;Context Manager&lt;/h1&gt;&lt;h2 style="" id="%E5%A4%9A%E6%A8%A1%E6%80%81%E6%95%B0%E6%8D%AE%E7%9A%84%E5%A4%84%E7%90%86"&gt;多模态数据的处理&lt;/h2&gt;&lt;p style=""&gt;现在大模型系统的对话窗口和处理数据会产生大量的上下文，后续的qa往往会用到其中一部分上下文，所以建立有效的context engine是十分必要的。&lt;/p&gt;&lt;p style=""&gt;&lt;/p&gt;&lt;h3 style="" id="nlp"&gt;NLP&lt;/h3&gt;&lt;p style=""&gt;&lt;strong&gt;用时问戳标记上下文&lt;/strong&gt; &lt;u&gt;一种常见的设计是在每条信息上附加时间戳，以保留其生成的顺序&lt;/u&gt;。这种方法由于简单且维护成本低，在聊天机器人和用户活动监控中广受欢迎。然而，该方法存在若干局限性。尽管时间戳能够保持时间顺序，但它们&lt;u&gt;不提供语义结构&lt;/u&gt;，使得捕捉长程依赖关系或高效检索相关信息变得困难。随着交互数据的累积，序列呈线性增长，导致在存储和推理方面均面临可扩展性问题&lt;/p&gt;</description></item><item><title>【LLM 必读综述】Speed Always Wins：LLM高效架构调查</title><link>https://moyutianzun.com/blog/llm-bi-du-zong-shu-speed-always-wins-llmgao-xiao-jia-gou-diao-cha/</link><pubDate>Wed, 27 Aug 2025 08:06:02 +0800</pubDate><guid>https://moyutianzun.com/blog/llm-bi-du-zong-shu-speed-always-wins-llmgao-xiao-jia-gou-diao-cha/</guid><description>&lt;p style=""&gt;题目：Speed Always Wins: A Survey on Efficient Architectures for Large Language Models&lt;/p&gt;&lt;p style=""&gt;作者：孙伟高 上海人工智能实验室&lt;/p&gt;&lt;p style=""&gt;github：https://github.com/weigao266/Awesome-Efficient-Arch&lt;/p&gt;&lt;p style=""&gt;&lt;/p&gt;&lt;h1 style="" id="%E6%96%87%E7%AB%A0%E7%9A%84%E4%B8%BB%E8%A6%81%E7%9B%AE%E7%9A%84"&gt;&lt;span style="font-size: 36px"&gt;&lt;strong&gt;文章的主要目的&lt;/strong&gt;&lt;/span&gt;&lt;/h1&gt;&lt;p style="text-indent: 2em"&gt;通过分析449篇大模型的论文，汇总得到了以下能让大模型速度提升的方向：&lt;/p&gt;</description></item><item><title>如何最快速的找到最核心的几篇文章</title><link>https://moyutianzun.com/blog/arxiv/</link><pubDate>Sun, 24 Aug 2025 16:12:48 +0800</pubDate><guid>https://moyutianzun.com/blog/arxiv/</guid><description>&lt;p style=""&gt;做这期blog的动机很简单，分享一下自己如何快速的上手某个领域的论文。最核心的三个步骤我觉得分别是：&lt;/p&gt;&lt;ul&gt;&lt;li&gt;&lt;p style=""&gt;确定自己&lt;strong&gt;研究领域的key words&lt;/strong&gt;，这里是要从上到下的，如LLM -&amp;gt; 微调 和 CoT&lt;/p&gt;&lt;/li&gt;&lt;li&gt;&lt;p style=""&gt;确定自己需要&lt;strong&gt;研究的题目&lt;/strong&gt;，如 如何确定CoT的某个环节有益于微调中得到高的得分&lt;/p&gt;&lt;/li&gt;&lt;li&gt;&lt;p style=""&gt;根据“and”的检索思想，从小到大的进行检索&lt;/p&gt;&lt;ul&gt;&lt;li&gt;&lt;p style=""&gt;先检索key words，确定论文的范围&lt;/p&gt;&lt;/li&gt;&lt;li&gt;&lt;p style=""&gt;在页面内只看研究的问题是否符合自己的要求&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;&lt;/li&gt;&lt;/ul&gt;&lt;p style=""&gt;&lt;/p&gt;&lt;h1 style="" id="%E5%BF%AB%E9%80%9F%E7%A1%AE%E5%AE%9A%E8%87%AA%E5%B7%B1%E9%9C%80%E8%A6%81%E7%9A%84%E8%AE%BA%E6%96%87"&gt;快速确定自己需要的论文&lt;/h1&gt;&lt;p style="text-indent: 2em"&gt;想要快速了解某个领域的论文，一个是问，问从业者，这是最快的，有人脉，直接问最核心的开发者，受益无穷，少走很多弯路。&lt;/p&gt;</description></item><item><title>MLsys24 分类汇总</title><link>https://moyutianzun.com/blog/mlsys24-fen-lei-hui-zong/</link><pubDate>Thu, 14 Aug 2025 15:51:24 +0800</pubDate><guid>https://moyutianzun.com/blog/mlsys24-fen-lei-hui-zong/</guid><description>&lt;h1 style="" id="llm-%E6%8E%A8%E7%90%86%E4%B8%8E%E6%9C%8D%E5%8A%A1%E4%BC%98%E5%8C%96-(llm-inference-and-serving-optimization)"&gt;&lt;strong&gt;LLM 推理与服务优化 (LLM Inference and Serving Optimization)&lt;/strong&gt;&lt;/h1&gt;&lt;p style=""&gt;&lt;/p&gt;&lt;h2 style="" id="kv-%E7%BC%93%E5%AD%98%E7%AE%A1%E7%90%86%E5%92%8C%E4%BC%98%E5%8C%96-(kv-cache-management-and-optimization)"&gt;&lt;strong&gt;KV 缓存管理和优化 (KV Cache Management and Optimization)&lt;/strong&gt;&lt;/h2&gt;&lt;p style=""&gt;&lt;/p&gt;&lt;p style=""&gt;这些论文聚焦于 KV 缓存的减少、量化或重用，以提升生成推理效率和降低内存消耗。&lt;/p&gt;&lt;ul&gt;&lt;li&gt;&lt;p style=""&gt;&lt;a href="http://www.moyutianzun.cn/2025/07/30/keyformer%E9%98%85%E8%AF%BB%E7%AC%94%E8%AE%B0/"&gt;Keyformer: KV Cache reduction through key tokens selection for Efficient Generative Inference&lt;/a&gt;&lt;/p&gt;&lt;ul&gt;&lt;li&gt;&lt;p style=""&gt; (Keyformer：通过高效生成推理的关键令牌选择来减少 KV 缓存) - Muhammad Adnan 等&lt;/p&gt;</description></item></channel></rss>