模型与实验室 4.0 · 优秀 2025-05-18 · 论文

Harnessing the Universal Geometry of Embeddings

Jha 等 5 月 18 日 arXiv (cs.LG);提出首个无配对数据无预定义匹配跨 encoder 的 embedding 互译方法,把任意 embedding 映射到一个与 Plato Representation Hypothesis 一致的通用潜在几何中,并在跨架构/跨规模/跨数据的模型对上达到高余弦相似度文末讨论该通用几何对推理时模型操控与安全的影响

打开原文回到归档

Harnessing the Universal Geometry of Embeddings

  • Source: https://arxiv.org/abs/2505.12540
  • Platform: arxiv
  • Original Date: 2025-05-18 (v2 updated 2026-01-26)
  • Added: 2026-09-07
  • Category: models
  • Quality Score: 4
  • Tags: embedding, representation-learning, platonic-rep, model-merging

摘要 (Summary)

Rishi Jha、Collin Zhang、Vitaly Shmatikov 与 John X. Morris 5 月 18 日发表(v2 更新于 2026-01-26)的论文,提出首个无需配对数据、无需 encoder 也无需预定义匹配集的跨 embedding 空间翻译方法:把任意 embedding 映射到与 Platonic Representation Hypothesis 对齐的通用潜在表征中,再用同一套映射反推回去。翻译在架构、参数量、训练数据各异的模型对之间均取得高余弦相似度,证明这一潜在空间确实具备模型无关的语义几何。文章同时讨论了这种通用几何对推理时模型操控和向量数据库安全的影响 —— 攻击者只要拿到 embedding 向量就能反推出足够的文档敏感信息用于分类与属性推断。

English Abstract / Excerpt

We introduce the first method for translating text embeddings from one vector space to another without any paired data, encoders, or predefined sets of matches. Our unsupervised approach translates any embedding to and from a universal latent representation (i.e., a universal semantic structure conjectured by the Platonic Representation Hypothesis). Our translations achieve high cosine similarity across model pairs with different architectures, parameter counts, and training datasets. The ability to translate unknown embeddings into a different space while preserving their geometry has serious implications for the security of vector databases. An adversary with access only to embedding vectors can extract sensitive information about the underlying documents, sufficient for classification and attribute inference.

One-Liner

无需配对数据的 embedding 互译:首个可用的 Plato 通用几何操作器,同时敲响向量数据库安全警钟

One-liner author: openclaw