跳到正文
原文
Simon Willison·· 1 天前AI 评分44

EmbeddingGemma 2 采用 Apache 2.0,嵌入模型不宜只用闭源托管

EmbeddingGemma 2

AI 导读

Simon Willison看重EmbeddingGemma 2采用Apache 2.0,认为嵌入模型不宜闭源专有且仅能托管。嵌入模型的多数应用要计算数千甚至数百万条嵌入向量,并存下来供日后比较。

正文 · 原文

6th October 2026

I really appreciate that EmbeddingGemma 2 is under the Apache 2.0 license.

For embedding models in particular, I don't think it makes sense to use a closed, proprietary, hosted-only model.

Most applications of embedding models involve calculating thousands or even millions of embedding vectors and storing them for later comparison.

If your model is proprietary, the vendor is likely someday going to decide to stop offering that model. They'll have a better model to replace it, but you still need to pay to re-calculate those millions of stored existing vectors.

(In April 2024 OpenAI offered to "cover the financial cost of users re-embedding content with these new models" - https://openai.com/index/gpt-4-api-general-availability/ - but I don't think that's something we can rely on from every provider.)

Notably, I don't want to host the model myself. I'd much rather pay a provider for a hosted model while knowing that if they ever stop hosting it I can run the open weights version myself - or find another vendor who can do that for me.

来源:Simon Willison · simonwillison.net