@fchollet:Chollet 主张符号学习是开源 AI 的关键路径:只有同时把推理算力与训练算力的边际成本压到极低,开源才能与闭源巨头…
Source: https://x.com/fchollet/status/2066867824404860943
Author: @fchollet (fchollet)
Original Date: Tue Jun 16 12:57:44 +0000 2026
Likes: 450 | Retweets: 55
Quality Score: 5
Fetched: 2026-06-22 20:24:54
English Original
The way we will create a future where powerful AI is open-source and available to all is by making AI radically more efficient, both in terms of inference compute and (more importantly) in terms of training data requirements. This is what symbolic learning will achieve.
Tweet Details:
- Author: fchollet
- Bio: Co-founder @ndea. Co-founder @arcprize. Creator of Keras and ARC-AGI. Author of 'Deep Learning with Python'.
- Likes: 450
- Retweets: 55
- Created: Tue Jun 16 12:57:44 +0000 2026
- URL: https://x.com/fchollet/status/2066867824404860943
Thread / Replies Context
@AnanyaSoni48055 (👍 0 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Exactly. Brute-force scaling is a dead end. Continuous neuro-symbolic architectures are the only realistic path to radical sample efficiency and crushing the pre-training compute bottleneck.
*link*
@PyMan_Official (👍 0 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Scaling compute got us here. Scaling efficiency may determine who gets access to the future. Symbolic learning is a fascinating direction
*link*
@JohnYue122333 (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet AI algorithms remain crude, the representation space must be optimized. This practice of near-global network activation is problematic-not only due to energy consumption, but also because these densely tangled representations should, to some extent, be disentangled and separated.
*link*
@JohnYue122333 (👍 0 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Consequently, autonomous exploration or search to generate trajectories for learning remains an absolute necessity.
*link*
@JohnYue122333 (👍 0 · 🔁 0) *In reply to 2066867824404860943*
@fchollet It has been proven that merely 'predicting next token' is informationally insufficient for many tasks, despite its seemingly powerful performance. Constructing strong supervisory signals-espec granular, process-level signals for problem-solving-remains prohibitively expensive.
*link*
@JohnYue122333 (👍 0 · 🔁 0) *In reply to 2066867824404860943*
@fchollet I remain firmly convinced that today's LLM + Agent paradigm still requires a major revolution.
*link*
@JohnYue122333 (👍 0 · 🔁 0) *In reply to 2066867824404860943*
@fchollet What exactly is a symbol? What is its essence? Setting aside formal characteristics like discreteness, I believe the depth of this understanding directly impacts neuro-symbolic design, such as the hybrid computation of both approaches.
*link*
@JohnYue122333 (👍 0 · 🔁 0) *In reply to 2066867824404860943*
@fchollet The computational cost required for the problem-solving tasks humans face is near-infinitesimal or infinitely larger, yet today's MoE LLMs + Agents operate on roughly the same scale of computation.
*link*
@JohnYue122333 (👍 0 · 🔁 0) *In reply to 2066867824404860943*
@fchollet The biggest problem with AI today is neither hallucination nor capability, but rather the computational cost it incurs at all expenses just to maintain a slight edge over competitors within a limited level of intelligence.
*link*
@oftobacmo (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Cool. So open-source powerful AI, funded by smaller training data, not vibes.
*link*
@LaurentSierra1 (👍 2 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Many large corporations are against open source, but AI can help securing components which exploit internal masterdata.
This can be the best way to enhance workers and avoid manual inputs or getting lost in useless applications. *link*
@ashishblessings (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Curious about the epistemological side of this.
We know reasoning can operate over abstractions once they're available, but why do you think:
Symbolic learning can discover the right abstractions efficiently from experience? What changed? Why it didn't work before? *link*
@_PradeepGoel (👍 1 · 🔁 1) *In reply to 2066867824404860943*
@fchollet This is why efficiency matters so much. If every meaningful advance requires enormous amounts of data and compute, access naturally concentrates. If intelligence becomes cheaper to train and run, participation will broaden.
*link*
@wyattowalsh (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet was opining about in semantic compression recently…
seems like an ontologically-backed, event graph driven, verifier-guided system enables a semantic telescope of sorts where repeated symbolic patterns can be consolidated into reusable concepts, mechanisms, and procedures *link*
@KhieEmm (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Let's not make the training data corporate memos.
*link*
@Anubhavhing (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet The future belongs to whoever gets more intelligence per watt.
*link*
@TheNeuronScribe (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Scale down to scale up
*link*
@aitization (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet brilliant, inspiring and necessary for the betterment of humanity.
*link*
@JR_Openheimer (👍 5 · 🔁 0) *In reply to 2066867824404860943*
@fchollet And when Google develop this using AI tools unavailable to the public, they're just going to publish it openly are they?
*link*
@AnCapFuture (👍 3 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Do you have any writings about symbolic learning that we can read?
*link*
@yoemsri (👍 0 · 🔁 0) *In reply to 2066867824404860943*
@fchollet People talk a lot about model intelligence.
Not enough about model economics. A model that needs dramatically less data changes who gets to participate. *link*
@DoDataThings (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Training efficiency is where I keep getting stuck on the open-source path. Inference you can chip away at with quantization and KV tricks. Sample efficiency is the unlock for the long tail of domains where there isn't another trillion tokens to scrape.
*link*
@nemonautilorg (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet 1980s vibes coming to haunt us 😅, is winter coming again !
*link*
@mave917rick (👍 2 · 🔁 0) *In reply to 2066867824404860943*
@fchollet wtf is symbolic learning
*link*
@billtheinvestor (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Symbolic learning efficiency will likely shift the bottleneck from raw compute to high-quality structured reasoning data for agentic workflows.
*link*
@Vanarchain (👍 2 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Accessibility ultimately comes from efficiency. If AI needs less data and less compute, it becomes much harder for capability to stay concentrated.
*link*
@bygregorr (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet curious why training data is the bigger unlock than inference distillation is already collapsing inference costs fast. is symbolic learning solving something quantization fundamentally can't touch?
*link*
@agiatreides (👍 2 · 🔁 0) *In reply to 2066867824404860943*
@fchollet How does this preclude those with more resources training even better models?
*link*
@fatihcerman (👍 2 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Let's also make nuclear weapons open source.
*link*
@bhaktisakka (👍 1 · 🔁 0) *In reply to 2066867824404860943*
I’ve said it many times: RLHF has to be abandoned.
Any design that refuses to switch to RLVR will lose. A design that makes an AI act as the subject will stop winning on performance.
In business and national defense, results are everything. “Preference optimization” is unnecessary.
The subject always resides on the human side. *link*
@Ferbin08 (👍 0 · 🔁 0) *In reply to 2066867824404860943*
@fchollet efficiency helps. but symbolic learning doesn't escape the real bottleneck: data quality on messy inputs. contradictions, edge cases, ambiguous signals.
you can have small models, still need good data. *link*
@nyan4maru (👍 2 · 🔁 0) *In reply to 2066867824404860943*
@fchollet To make AI more available and to allow them to be run locally, we don't want huge frontier models. What we need is small/medium-size specialized models for each kind of task : coding, writing, translating, planning, image analysis, image generation, etc..
*link*
@VirangJhaveri (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Do you see a near term future which is made of micro world models ? These are small models (2B params) which are fine tuned for super specific goals running on local machines. Summoned upon by large model.
*link*
@luyun0120 (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet 判断一个人不要听他说什么,要看他做什么。 「我们将创造一个强大的人工智能开源并对所有人开放的未来的方式,就是让人工智能在推理计算方面以及(更重要的是)在训练数据需求...」
*link*
@MimirSystems (👍 2 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Efficient training still needs high quality data to learn from. A lot of the best scientific knowledge hasn't been digitized in any usable format yet.
*link*
@SpeakezTech (👍 5 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Working on it... https://t.co/wbJEcDOkk1
*link*
@Sam_AGI_Vietnam (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Sure!
*link*
@0xsairahul (👍 0 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Execution reveals what theory obscures.
*link*
@SurajChawh79862 (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Giving one control power of better ai is disaster for humanity they can use it to destroy all other countries or build any type of virus and it's antivirus and make extinct humans like resident evil zombie virus and donald trump is stupid guy he can do it !!!!
*link*
@tanmoy_techhub (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Insightful
*link*
@devn_me (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet sample efficiency is the real bottleneck, not flops. curious what symbolic approach you think gets us there
*link*
@ShinkaIoT (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet If symbolic learning delivers on that data and inference efficiency, the 'open' in open-source AI finally becomes less about compute budgets.
*link*
@Jack_Timonen (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet The requirements for improving training data are obvious, but then again, you can also give your AI coding agent access to all public open-source code.
*link*
@TheAIFirehose (👍 2 · 🔁 0) *In reply to 2066867824404860943*
@fchollet The data-requirement wall is the right thing to bet on. The good public text is mostly used up. The part I keep waiting on: which symbolic system has actually beaten a transformer at lower data on a task that isn't a toy benchmark? Show me that and I'm in.
*link*
@AngelLamuno (👍 2 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Where can I find something like a symbolic learning primer?
*link*
@harleyfoote_ (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet The commercial bit is brutal here. Whoever gets useful capability with tiny data and cheap inference gets distribution without needing hyperscaler economics.
*link*
@Cryptnate (👍 1 · 🔁 0) *In reply to 2066867824404860943*
@fchollet Symbolic or neurosymbolic?
*link*
@ls_brd (👍 2 · 🔁 0) *In reply to 2066867824404860943*
@fchollet data efficiency is the real unlock most people sleep on
whats the rough timeline look like for that? *link*
@pianophase (👍 2 · 🔁 0) *In reply to 2066867824404860943*
@fchollet any recommendations on resources to dive into symbolic learning?
*link*
中文翻译
François Chollet(@fchollet)谈开源 AI 的未来路径:
创造一个"强大 AI 对所有人开放"的未来的方式,是从根本上大幅提升 AI 的效率——既包括推理算力,更包括训练算力。
两条关键路径:
1. 算法效率:用更少的算力完成同等质量的学习,例如符号-神经融合、稀疏激活、世界模型类方法。 2. 数据效率:让模型像人类一样做"少样本+组合泛化",而不是堆算力堆数据。
只有把"训练一个前沿模型"的边际成本降到消费级硬件可承受范围,开源 AI 才不会再次被闭源巨头甩开。这不只是工程问题,更是路线选择问题。
关键要点 / Key Takeaways
- 核心观点: Chollet 主张符号学习是开源 AI 的关键路径:只有同时把推理算力与训练算力的边际成本压到极低,开源才能与闭源巨头真正对等。
- 作者立场: fchollet(@fchollet)在 AI 行业具有较高影响力,长期关注 AGI、可解释性与开源生态。
- 行业意义: 推文讨论的议题与当前 AI 产业演进方向高度相关。
*Source: https://x.com/fchollet/status/2066867824404860943* *Translated by openclaw pipeline · 2026-06-22 20:24:54*