<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="zh-Hans-CN">
	<id>https://ac-wiki.cn/index.php?action=history&amp;feed=atom&amp;title=Llama.cpp</id>
	<title>Llama.cpp - 版本历史</title>
	<link rel="self" type="application/atom+xml" href="https://ac-wiki.cn/index.php?action=history&amp;feed=atom&amp;title=Llama.cpp"/>
	<link rel="alternate" type="text/html" href="https://ac-wiki.cn/index.php?title=Llama.cpp&amp;action=history"/>
	<updated>2026-08-24T05:22:39Z</updated>
	<subtitle>本wiki上该页面的版本历史</subtitle>
	<generator>MediaWiki 1.45.3</generator>
	<entry>
		<id>https://ac-wiki.cn/index.php?title=Llama.cpp&amp;diff=2653&amp;oldid=prev</id>
		<title>2026年6月14日 (日) 15:26 天明</title>
		<link rel="alternate" type="text/html" href="https://ac-wiki.cn/index.php?title=Llama.cpp&amp;diff=2653&amp;oldid=prev"/>
		<updated>2026-06-14T15:26:00Z</updated>

		<summary type="html">&lt;p&gt;&lt;/p&gt;
&lt;table style=&quot;background-color: #fff; color: #202122;&quot; data-mw=&quot;interface&quot;&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;col class=&quot;diff-marker&quot; /&gt;
				&lt;col class=&quot;diff-content&quot; /&gt;
				&lt;tr class=&quot;diff-title&quot; lang=&quot;zh-Hans-CN&quot;&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;←上一版本&lt;/td&gt;
				&lt;td colspan=&quot;2&quot; style=&quot;background-color: #fff; color: #202122; text-align: center;&quot;&gt;2026年6月14日 (日) 23:26的版本&lt;/td&gt;
				&lt;/tr&gt;&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot; id=&quot;mw-diff-left-l1&quot;&gt;第1行：&lt;/td&gt;
&lt;td colspan=&quot;2&quot; class=&quot;diff-lineno&quot;&gt;第1行：&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;−&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #ffe49c; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;llama.cpp &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;是一个开源的 &lt;/del&gt;C/C++ &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;项目，旨在在本地硬件上高效运行大语言模型，最初针对 &lt;/del&gt;Meta 的 LLaMA &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;模型设计。主要用途是让用户无需依赖云服务即可在 &lt;/del&gt;CPU 或普通 GPU 上推理运行模型。核心特点包括使用 GGUF &lt;del style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;格式量化模型以降低内存占用，支持多平台（Windows、Linux、macOS、Android）及多种加速后端（如 BLAS、CUDA、Vulkan）。项目于 &lt;/del&gt;2023 年 3 月由开发者 Georgi Gerganov 首次发布，后续社区贡献使其支持更多模型架构，成为本地运行开源大模型的重要工具之一。&lt;/div&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;https://llama.app/&lt;/ins&gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-side-deleted&quot;&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt; &lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-side-deleted&quot;&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;https://github.com/ggml-org/&lt;/ins&gt;llama.cpp&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-side-deleted&quot;&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt; &lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-side-deleted&quot;&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;llama.cpp 是一个[[开源]]的 [[C 语言|&lt;/ins&gt;C&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;]]&lt;/ins&gt;/&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;[[&lt;/ins&gt;C++&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;]] 项目，旨在在本地硬件上高效运行[[大语言模型]]，最初针对 [[&lt;/ins&gt;Meta&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;]] &lt;/ins&gt;的 &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;[[LLaMA|&lt;/ins&gt;LLaMA &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;模型]]设计。主要用途是让用户无需依赖云服务即可在 [[&lt;/ins&gt;CPU&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;]] &lt;/ins&gt;或普通 &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;[[&lt;/ins&gt;GPU&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;]] &lt;/ins&gt;上推理运行模型。核心特点包括使用 &lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;[[&lt;/ins&gt;GGUF&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;|GGUF 格式]]量化模型以降低内存占用，支持多平台（[[Windows]]、[[Linux]]、[[macOS]]、[[Android]]）及多种加速后端（如 BLAS、[[CUDA]]、[[Vulkan]]）。项目于 &lt;/ins&gt;2023 年 3 月由开发者 Georgi Gerganov 首次发布，后续社区贡献使其支持更多模型架构，成为本地运行开源大模型的重要工具之一。&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-side-deleted&quot;&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt; &lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td colspan=&quot;2&quot; class=&quot;diff-side-deleted&quot;&gt;&lt;/td&gt;&lt;td class=&quot;diff-marker&quot; data-marker=&quot;+&quot;&gt;&lt;/td&gt;&lt;td style=&quot;color: #202122; font-size: 88%; border-style: solid; border-width: 1px 1px 1px 4px; border-radius: 0.33em; border-color: #a3d3ff; vertical-align: top; white-space: pre-wrap;&quot;&gt;&lt;div&gt;&lt;ins style=&quot;font-weight: bold; text-decoration: none;&quot;&gt;{{DISPLAYTITLE: llama.cpp}}&lt;/ins&gt;&lt;/div&gt;&lt;/td&gt;&lt;/tr&gt;

&lt;!-- diff cache key acwiki_mysql:diff:1.41:old-2652:rev-2653:php=table --&gt;
&lt;/table&gt;</summary>
		<author><name>天明</name></author>
	</entry>
	<entry>
		<id>https://ac-wiki.cn/index.php?title=Llama.cpp&amp;diff=2652&amp;oldid=prev</id>
		<title>天明：​创建页面，内容为“llama.cpp 是一个开源的 C/C++ 项目，旨在在本地硬件上高效运行大语言模型，最初针对 Meta 的 LLaMA 模型设计。主要用途是让用户无需依赖云服务即可在 CPU 或普通 GPU 上推理运行模型。核心特点包括使用 GGUF 格式量化模型以降低内存占用，支持多平台（Windows、Linux、macOS、Android）及多种加速后端（如 BLAS、CUDA、Vulkan）。项目于 2023 年 3 月由开发者 Georgi Ger…”</title>
		<link rel="alternate" type="text/html" href="https://ac-wiki.cn/index.php?title=Llama.cpp&amp;diff=2652&amp;oldid=prev"/>
		<updated>2026-06-14T15:24:06Z</updated>

		<summary type="html">&lt;p&gt;创建页面，内容为“llama.cpp 是一个开源的 C/C++ 项目，旨在在本地硬件上高效运行大语言模型，最初针对 Meta 的 LLaMA 模型设计。主要用途是让用户无需依赖云服务即可在 CPU 或普通 GPU 上推理运行模型。核心特点包括使用 GGUF 格式量化模型以降低内存占用，支持多平台（Windows、Linux、macOS、Android）及多种加速后端（如 BLAS、CUDA、Vulkan）。项目于 2023 年 3 月由开发者 Georgi Ger…”&lt;/p&gt;
&lt;p&gt;&lt;b&gt;新页面&lt;/b&gt;&lt;/p&gt;&lt;div&gt;llama.cpp 是一个开源的 C/C++ 项目，旨在在本地硬件上高效运行大语言模型，最初针对 Meta 的 LLaMA 模型设计。主要用途是让用户无需依赖云服务即可在 CPU 或普通 GPU 上推理运行模型。核心特点包括使用 GGUF 格式量化模型以降低内存占用，支持多平台（Windows、Linux、macOS、Android）及多种加速后端（如 BLAS、CUDA、Vulkan）。项目于 2023 年 3 月由开发者 Georgi Gerganov 首次发布，后续社区贡献使其支持更多模型架构，成为本地运行开源大模型的重要工具之一。&lt;/div&gt;</summary>
		<author><name>天明</name></author>
	</entry>
</feed>