固定尺寸向量表示Fixed-Size Vector Representation

fixed-size vector representation,这其实就是我最近在学的核心主角了,我也是今天才知道它叫这个名字,就是embedding。
我们可以简单的认为,一个固定尺寸的向量表示一个知识点,知识库里放了成百上千这样的向量。llm收到一个query,只要把query也转换成向量然后计算它在这个向量库里距离最近的几个向量分别是什么,获取他们,那就相当于是用知识库的知识来回答用户的问题了。这里涉及很多步骤,我们一个个来。
首先我们要生成一个固定尺寸向量库,这需要使用embedding模型,它是一个可以将文本编码成固定尺寸向量表示的编码器。这里涉及到两个技术点,将哪些内容输入模型(输入数据的格式);以及,选用什么模型以及是否要对模型进行微调。
我们应该选取一种将文本输入模型进行编码的形式,我们可能有一篇新闻报道,一篇论文或者一篇个人博客,将整篇整段放进模型进行编码也许不是一个好选择。用户的问题可能难以直接和代表整篇文章的向量匹配成功,而往往是文章的标题,或者某一个关键词(如”天气很好“会和”太阳好大“匹配)。所以我们需要一个chunk策略对文本切片同时选取适合任务场景的embedding模型对文本进行编码(美食网站和健身网站对于“沙拉”和“我好爱”的向量相似度可能不同)。

同时我们还应该对用户的问题进行重新表述,并且也选取合适的embedding模型(一般是和创建RAG时使用的同一个)将问题转换为向量,这样才能进行计算。(健美先生和可爱的肥宅同时提出“给我推荐一些好吃的”的时候对应的向量表示可能不那么一致)。
向量库有了,用户的问题也变成了向量,我们可以很简单的计算两个向量的相似程度,并且得到一个相似度排序,此时我们可以选取 top k个最相似的文档和用户query一起(组成一个prompt)返回给llm,也可以做进一步处理(比如重排序等后处理)。
流程稍微有点复杂,但是合理的架构节省了大量的上下文却让llm获得了强大的检索能力,并且,这就像是它理论上能够检索也就是学习任何领域任意数量的信息。llm的知识在训练完成那一刻就停止了,在有限的模型上下文中外挂的RAG库赋予了llm更多更准确有时效性的知识,并且在一定程度上减少了幻觉问题。
还有个问题,何时使用它呢?也许我们不希望我只是问一下现在几点了然后Agent就去RAG库里翻翻找找返回 top k个最佳匹配然后整理给llm得出答案。合理的方式是让Agent自己决定是否需要调用RAG工具来获取额外的特定的信息。我们将RAG检索写成一个tool_call的形式,描述它的作用以及在何时可以调用,同时给每个RAG库带上描述它含有哪些知识的标签,从而让llm自行决定调用RAG工具并且选择最合适的知识库进行检索。还有很多优化措施,比如相似度阈值兜底,如果不够就换个库检索或者不采用等等。
references
[1]Advanced RAG on Hugging Face documentation using LangChain

Fixed-size vector representation is actually the core subject I've been learning recently—it was only today that I learned its name: embedding.
We can simply think of a fixed-size vector as representing a piece of knowledge, and a knowledge base stores hundreds or thousands of such vectors. When the LLM receives a query, we just need to convert the query into a vector, compute the closest vectors in this vector library, and retrieve them. That is essentially answering the user's question using the knowledge base. There are many steps involved; let's go through them one by one.
First, we need to create a fixed-size vector library. This requires an embedding model—an encoder that can encode text into fixed-size vector representations. Two technical points are involved here: what content to feed into the model (the format of the input data), and which model to choose and whether to fine-tune it.
We should choose a way to input text into the model for encoding. We might have a news article, a paper, or a personal blog post. Putting the whole article or paragraph into the model for encoding may not be a good choice. A user's question may not easily match the vector representing the entire article; instead, it often matches the article's title or a specific keyword (e.g., "the weather is nice" will match "the sun is big"). Therefore, we need a chunking strategy to slice the text and select an embedding model suitable for the task scenario to encode the text (a food website and a fitness website may have different vector similarities for "salad" and "I love it").

At the same time, we should also reformulate the user's question and select an appropriate embedding model (usually the same one used when creating the RAG) to convert the question into a vector so that calculations can be performed. (A bodybuilder and a cute couch potato asking "recommend me something delicious" at the same time may have vector representations that are not quite consistent).
Once we have the vector library and the user's question has been converted into a vector, we can easily compute the similarity between vectors and obtain a similarity ranking. At this point, we can select the top k most similar documents and return them together with the user's query (forming a prompt) to the LLM, or we can do further processing (such as reranking and other post-processing).
The workflow is a bit complex, but a well-designed architecture saves a lot of context while giving the LLM powerful retrieval capabilities. Moreover, it is as if the LLM can theoretically retrieve—that is, learn—any amount of information in any domain. The LLM's knowledge stops at the moment training completes; an external RAG library appended to the limited model context endows the LLM with more accurate, timely knowledge and, to some extent, reduces the hallucination problem.
There is another question: when should we use it? Perhaps we don't want the Agent to rummage through the RAG library when I just ask what time it is, retrieve the top k best matches, and organize them for the LLM to produce an answer. A reasonable approach is to let the Agent decide whether it needs to call RAG tools to obtain additional specific information. We write RAG retrieval as a tool_call, describing its purpose and when it can be invoked, and attach tags to each RAG library describing what knowledge it contains. This allows the LLM to decide on its own whether to call RAG tools and to select the most suitable knowledge base for retrieval. There are many optimization measures, such as a similarity threshold fallback—if the similarity isn't high enough, switch to another library or don't use it at all.
Fine.
references
[1]Advanced RAG on Hugging Face documentation using LangChain