Andrej Karpathy:Farzapedia用LLM把个人数据变成个人维基
原文链接: https://x.com/karpathy/status/2040572272944324650
发布时间: 2026-04-05
作者: @karpathy(Andrej Karpathy,AI研究员)
原文 / Original:
Farzapedia, personal wikipedia of Farza, good example following my Wiki LLM tweet. I really like this approach to personalization in a number of ways, compared to "status quo" of an AI that allegedly gets better the more you use it or something:
1. Explicit. The memory artifact is explicit and navigable (the wiki), you can see exactly what the AI does and does not know and you can inspect and manage this artifact, even if you don't do the direct text writing (the LLM does). The knowledge of you is not implicit and unknown, it's explicit and viewable.
2. Yours. Your data is yours, on your local computer, it's not in some particular AI provider's system without the ability to extract it. You're in control of your information.
3. File over app. The memory here is a simple collection of files in universal formats (images, markdown). This means the data is interoperable: you can use a very large collection of tools/CLIs or whatever you want over this information because it's just files. The agents can apply the entire Unix toolkit over them. They can natively read and understand them. Any kind of data can be imported into files as input, and any kind of interface can be used to view them as the output. E.g. you can use Obsidian to view them or vibe code something of your own. Search "File over app" for an article on this philosophy.
4. BYOAI. You can use whatever AI you want to "plug into" this information - Claude, Codex, OpenCode, whatever. You can even think about taking an open source AI and finetuning it on your wiki - in principle, this AI could "know" you in its weights, not just attend over your data.
So this approach to personalization puts *you* in full control. The data is yours. In Universal formats. Explicit and inspectable. Use whatever AI you want over it, keep the AI companies on their toes! :) Certainly this is not the simplest way to get an AI to know you - it does require you to manage file directories and so on, but agents also make it quite simple and they can help you a lot. I imagine a number of products might come out to make this all easier, but imo "agent proficiency" is a CORE SKILL of the 21st century. These are extremely powerful tools - they speak English and they do all the computer stuff for you. Try this opportunity to play with one.
引用推文 — Farza @FarzaTV:
"This is Farzapedia. I had an LLM take 2,500 entries from my diary, Apple Notes, and some iMessage convos to create a personal Wikipedia for me. It made 400 detailed articles for my friends, my startups, research areas, and even my favorite animes and their impact on me."
译文 / Translation:
Farzapedia(Farza的个人维基),是继我的Wiki LLM推文之后的一个很好的案例。我从多个方面非常喜欢这种个性化方式,对比那些声称"用得越多AI就越懂你"的现状:
1. 显式(Explicit)。记忆artifact是显式且可导航的(就是一个wiki),你可以精确看到AI知道什么、不知道什么,可以审查和管理这个artifact,即使你并不亲自写文本(LLM在帮你写)。关于你的知识不是隐含的、不可知的,而是显式的、可查看的。
2. 属于你(Yours)。你的数据是你自己的,在你自己的电脑上,不在某个AI提供商的系统里、无法提取。你对你的信息有完全的控制。
3. 文件优先于应用(File over app)。这里的记忆是通用格式文件(图片、markdown)的简单集合。这意味着数据是互操作的:你可以用大量工具、CLI或任何你想要的东西来处理这些信息,因为它们就是普通文件。AI Agent可以对它们使用整套Unix工具链。它们能原生地读取和理解这些文件。任何类型的数据都可以导入为文件作为输入,任何类型的界面都可以作为输出查看它们。比如你可以用Obsidian查看它们,或者凭感觉自己写个代码。搜索"File over app"可以找到关于这个理念的文章。
4. 自带AI(BYOAI)。你可以使用任何你想要的AI来"接入"这些信息——Claude、Codex、OpenCode,随你选。你甚至可以想象用开源AI在wiki上微调它——原则上,这个AI可以在它的权重中"认识"你,而不仅仅是关注你的数据。
所以这种个性化方法让你完全掌控。数据是你的。是通用格式。显式且可审查。用你想要的任何AI来处理它们,让AI公司保持警觉!当然,这不是让AI了解你最简单的方式——确实需要你管理文件目录等,但Agent也让它变得相当简单,而且它们能帮你很多。我想象会有很多产品出现来让这一切更容易,但对我来说,"Agent能力"是21世纪的核心技能。这些是非常强大的工具——它们说英语,帮你做所有电脑上的事情。试试这个机会来玩一玩Agent吧。
译文 — 引用推文(Farza @FarzaTV):
"这就是Farzapedia。我让一个LLM从我的日记、Apple Notes和一些iMessage对话中提取了2500条记录,为我创建了一部个人维基。它生成了400篇详细文章,涵盖我的朋友、我的创业项目、研究领域,甚至我最爱的动漫以及它们对我的影响。"
来源: Andrej Karpathy (@karpathy)
链接: https://x.com/karpathy/status/2040572272944324650