工具与项目 3.0 · 值得看 2026-06-07 · X

Farzapedia:用日记+笔记训练个人 Wikipedia,LLM 个性化的最佳实践

Farzapedia:用日记+笔记训练个人 Wikipedia,LLM 个性化的最佳实践

回到归档

Farzapedia:用日记+笔记训练个人 Wikipedia,LLM 个性化的最佳实践

来源: karpathy
抓取时间: 2026-04-22
原始语言: 英文 → 中文

English:

Farzapedia, personal wikipedia of Farza, good example following my Wiki LLM tweet.

I really like this approach to personalization in a number of ways, compared to "status quo" of an AI that allegedly gets better the more you use it or something:

1. Explicit. The memory artifact is explicit and navigable (the wiki), you can see exactly what the AI does and does not know and you can inspect and manage this artifact, even if you don't do the direct text writing (the LLM does). The knowledge of you is not implicit and unknown, it's explicit and viewable.

2. Yours. Your data is yours, on your local computer, it's not in some particular AI provider's system without the ability to extract it. You're in control of your information.

3. File over app. The memory here is a simple collection of files in universal formats (images, markdown). This means the data is interoperable: you can use a very large collection of tools/CLIs or whatever you want over this information because it's just files. The agents can apply the entire Unix toolkit over them. They can natively read and understand them. Any kind of data can be imported into files as input, and any kind of interface can be used to view them as the output. E.g. you can use Obsidian to view them or vibe code something of your own. Search "File over app" for an article on this philosophy.

4. BYOAI. You can use whatever AI you want to "plug into" this information - Claude, Codex, OpenCode, whatever. You can even think about taking an open source AI and finetuning it on your wiki - in principle, this AI could "know" you in its weights, not just attend over your data.

So this approach to personalization puts *you* in full control. The data is yours. In Universal formats. Explicit and inspectable. Use whatever AI you want over it, keep the AI companies on their toes! :)

Certainly this is not the simplest way to get an AI to know you - it does require you to manage file directories and so on, but agents also make it quite simple and they can help you a lot. I imagine a number of products might come out to make this all easier, but imo "agent proficiency" is a CORE SKILL of the 21st century. These are extremely powerful tools - they speak English and they do all the computer stuff for you. Try this opportunity to play with one.

中文:

Farzapedia,Farza 的个人 Wikipedia,是继我的 Wiki LLM 推文之后的一个很好的范例。

相比所谓「随着使用越多 AI 就越好」的现状,这种个性化方式在很多方面深得我心:

1. 显式(Explicit)。记忆制品是显式且可导航的(那个 wiki),你可以确切看到 AI 知道什么、不知道什么,你可以检查并管理这个制品,即使你并没有直接写文本(是 LLM 在写)。关于你的知识不是隐含的、不可知的,它是显式的、可查看的。

2. 属于你(Yours)。你的数据是你的,在你自己的电脑上,不在某个 AI 提供商的系统里、无法提取。你对自己的信息拥有控制权。

3. 文件优于应用(File over app)。这里的记忆是简单文件集合,以通用格式(图片、markdown)存储。这意味着数据是互操作的:你可以通过大量工具/CLI 来处理这些信息,因为它们就是普通文件。智能体可以在这些文件上运用整套 Unix 工具链。它们天然可以读取和理解这些文件。任何类型的数据都可以作为输入导入到文件中,任何类型的界面都可以作为输出来查看它们。例如,你可以用 Obsidian 来查看,或者用 vibe coding 自己搭建一个。搜索「File over app」可以找到关于这个理念的文章。

4. BYOAI(Bring Your Own AI)。你可以使用任何你想要的 AI 来「接入」这些信息——Claude、Codex、OpenCode,随便你。你甚至可以想象用开源 AI 在你的 wiki 上微调它——原则上,这个 AI 可以「了解你」,不只是对你的数据进行处理。

所以这种个性化方法让你完全掌控。数据是你的。用通用格式。显式且可检查。用任何你想要的 AI 来处理它,让 AI 公司保持警觉!:)

当然,这不是让 AI 了解你最简单的方式——确实需要你管理文件目录等,但智能体也让这一切变得相当简单,它们可以帮你很多。我想象会有很多产品出现让这一切变得更容易,但在我看来「智能体熟练度」是 21 世纪的核心技能。这些是非常强大的工具——它们说英语,替你完成所有计算机工作。找个机会试试吧。

*以上内容由 AI Field Notes 自动抓取翻译,如有侵权请联系删除。*