mmagent(my-minimal-agent)mmagent(my-minimal-agent)

只是一个用于练习的Agent项目,我会把它接入我的一个website中作为一个ai web应用层的全栈项目,但是我可能也会把它设计成一个能在容器中运行通过API与外界通信的Agent模块,方便接入和部署。
我完全从零开始构建了这个Agent项目和我的web应用,它们可能并没有很好用,但是我可以按照自己的喜好设计它们,这也让我感到很满足。我应该只会设计一些简单的tool_call,以及试着实现完整的会话管理(其实我认为Agent循环以后会被设计的非常简单,因为模型的能力提升很快,虽然幻觉问题我想不出如何解决,目前看来上下文工程在很多场景下还是有应用的),分别是缓存层、存储层、检索层(大概是这么个意思吧)。分别用Redis, PostgreSQL和embedding向量来实现,这里可以学习一下Redis和PostgreSQL的特性,Redis存在RAM中按键检索所以快,PostgreSQL存在磁盘中所以慢但是能持久储存。以及会话是如何储存在表中的,schema是什么,每个session都有一个UUID等等。总之我会试着创造最佳的用户体验(我是讨好型人格)。
会话管理系统搭好了,采用主流的session KV + 向量库的双存储模式,PostgreSQL全量存储,Redis缓存最近10条消息缓解并发写入压力,加快会话消息读取,向量库实现跨会话消息检索(这里其实有很多可以设计的地方,可以根据设计外部知识库比如我的一些个人信息;或者工具描述向量库,从而匹配最合适的调用工具。这里就简单的将一些高价值内容存进向量库,比如用户的偏好、决策、性格等)。
就是最上面说的三层存储结构,其实要存的数据也不多,一共就两个表:sessions层在PostgreSQL中全量存储。包括uuid, user_id(后面必须要做Auth鉴权才能存进去数据,才能实现多用户多会话,每个用户管理自己的会话的功能), messages, created_at, update_at,其中messages以json形式存储数据;还有就是语义层,用来存储语义向量,同样的表只是比上面的多一个embedding向量,后续需要实现一个embedding方法来把切分后的语句选择性转换成embedding向量存储(语义切分这里也大有门道,可以针对场景选择多种方法。还有Agent的向量库检索,一般设计成tool_call,在需要的时候调用工具进行查询)。总之,Minimax和deepseek帮我做好了一些,但是我确实掌握着所有数据字段和消息储存方式,以及总共包含的4个api:
GET /api/sessions/{session_id} 用来拉取已有某个会话内容
POST /api/sessions/{session_id}/messages 用来发送消息
POST /api/session 用来创建会话
GET /api/users/{suer_id}/sessions 拉取某个用户的已有会话列表
我还可以给他们写一些状态检查或者错误码之类的,来练习一下API规范。

之后我接入一个完整的Agent对话api,后端通过http请求调用llm API,llm会自行选择是否进行工具调用(这是通过llm返回结构化信息给Runtime实现的,Runtime会读取结构体中指定的tool以及parameter),Agent在容器中执行Runtime循环(循环指的是执行工具返回执行结果给llm,llm进行下一步思考,返回消息),最终得到final answer(如何判断是否是final answer呢,有几种方法:1. 让大模型自己判断;2. 当不再发生工具调用的时候就认为是final anser;3. 超过指定步数限制,直接返回final answer)。Runtime将消息切分返回然后SSE流式渲染在前端消息窗口中,大概就这样。Agent端的tool_call,后端的RAG库,Auth鉴权用户管理,这些都可以慢慢集成。
用户鉴权
RAG库接入有两种主流方法,可以在每次回复前都进行一次检索,将检索结果拼进messages里面;也可以将RAG功能设置为一个Agent tool,让llm自行决定要不要调用RAG tool。第二种方法更合理一点,通过检测当前问题是否需要跨会话检索或者长期记忆,从而有目的的检索知识库里面的信息。主流的检索方式往往是向量召回 + 关键词/BM25 -> rerank 的检索方式,然后再将结果返回给llm。RAG的接入和检索方式都在Agent runtime设计好了,那么RAG向量库到底要怎么生成呢。
我理解的RAG库就是一堆语句信息(或者说语义)在一个高维空间的集合,每个语句都被映射成一个向量,而将它们映射为向量的这个模型是经过大量语句训练过的(就像ChatGPT一样),它很清楚不同语句之间有什么联系(比如红桃和爱丽丝,戒指和爱情),这些联系体现在高维空间中两个语句向量之间的距离(可以将其想象为一个三维坐标系中的两个物体)。所以检索的过程就像我们刚刚说的,将我们的询问转换为向量,去语义空间里计算一下距离哪些向量比较近,这些就是我们要的“记忆”,然后将它们提取出来、拼接,返回给llm。这就是RAG库的组成,碰巧我还知道这些语句和词汇都是由token组成的,如何合理选取它们并转化为RAG向量自然很大程度上影响了后面的检索结果(就像把“橘子是世界上唯一的水果”这句话转化成向量和将“橘子”、“世界”、“唯一”、“水果”分别转化成向量会导致最后的检索结果不同一眼)。

那么该如何切分段落并将其转化为RAG向量呢(还要先转换成token)。我认为切分要考虑到语义完整性、结构、以及embedding模型的能力。这里涉及两种可以挑选的策略,chunking策略(将整篇文章切分成不同段落)和tokenizer策略(将一段话拆分成最小语义单元,涉及压缩与泛化)

This is just an Agent project for practice. I will integrate it into one of my websites as a full-stack AI web application layer, but I may also design it to be an Agent module that can run in a container and communicate with the outside world via APIs, making it easy to integrate and deploy.
I built this Agent project and my web application completely from scratch. They may not be very user-friendly, but I can design them according to my own preferences, which gives me a great sense of satisfaction. I will probably design only some simple tool calls and try to implement complete session management (actually, I think the Agent loop will become very simple in the future because models are improving rapidly, though I can't figure out how to solve the hallucination problem; for now, context engineering still seems applicable in many scenarios). The components will be a cache layer, a storage layer, and a retrieval layer (approximately that concept). I'll implement them using Redis, PostgreSQL, and embedding vectors respectively. This is a good opportunity to learn about Redis and PostgreSQL features: Redis stores data in RAM and is fast for key-based retrieval, while PostgreSQL stores data on disk, so it's slower but provides persistent storage. I'll also learn how sessions are stored in tables, what the schema looks like, and that each session has a UUID. In short, I will try to create the best user experience (I have a people-pleasing personality).
The session management system is set up, adopting the mainstream dual-storage mode of session KV + vector database. PostgreSQL stores everything, Redis caches the last 10 messages to ease concurrent write pressure and speed up session message reads, and the vector database enables cross-session message retrieval (There are actually many design possibilities here. For example, an external knowledge base could be designed, such as some personal information of mine; or a tool description vector database could be used to match the most suitable tool to call. Here, we simply store some high-value content into the vector database, such as user preferences, decisions, personality, etc.).
It is the three-tier storage structure mentioned above. Actually, there isn't much data to store; there are only two tables in total: the sessions layer is fully stored in PostgreSQL, including uuid, user_id (an Auth mechanism will be needed later to store data and implement multi-user, multi-session functionality where each user manages their own sessions), messages, created_at, update_at, where messages are stored as JSON. Then there is the semantic layer, used to store semantic vectors. The table is the same, except it has an additional embedding vector. Later, an embedding method needs to be implemented to selectively convert segmented sentences into embedding vectors for storage (Semantic segmentation has many subtleties; multiple methods can be chosen based on the scenario. Also, for the Agent's vector database retrieval, it is generally designed as a tool_call, and tools are called when needed for querying). In short, Minimax and deepseek have already helped me build some parts, but I do have full control over all data fields, message storage methods, and the total of 4 APIs:
GET /api/sessions/{session_id} 用来拉取已有某个会话内容
POST /api/sessions/{session_id}/messages 用来发送消息
POST /api/session 用来创建会话
GET /api/users/{suer_id}/sessions 拉取某个用户的已有会话列表
I can also write some status checks or error codes for them to practice API specifications.

Later, I integrated a complete Agent conversation API. The backend calls the LLM API via HTTP requests, and the LLM decides whether to make tool calls (this is implemented by the LLM returning structured information to the Runtime, which reads the tool and parameters specified in the struct). The Agent executes the Runtime loop inside a container (the loop means executing a tool, returning the execution result to the LLM, which then thinks further and returns a message), ultimately obtaining a final answer (how to determine whether it's a final answer? There are several methods: 1. Let the LLM judge by itself; 2. Consider it the final answer when no more tool calls occur; 3. If the maximum step limit is exceeded, directly return the final answer). The Runtime splits the message and returns it, then renders it via SSE streaming in the frontend message window, roughly like that. The Agent-side tool_call, backend RAG library, Auth authentication and user management—all of these can be gradually integrated.
User Authentication
There are two mainstream approaches to integrating a RAG library. You can perform a retrieval before each reply and splice the retrieved results into the messages; alternatively, you can set up the RAG functionality as an Agent tool and let the LLM decide whether to call the RAG tool. The second method is more reasonable: it checks whether the current question needs cross-session retrieval or long-term memory, and then intentionally retrieves information from the knowledge base. The mainstream retrieval approach is usually vector recall + keyword/BM25 → rerank, and then the results are returned to the LLM. The RAG integration and retrieval method have already been designed in the Agent runtime, so how exactly should the RAG vector database be generated?
As I understand it, a RAG library is a collection of sentence information (or semantics) in a high-dimensional space, where each sentence is mapped to a vector. The model that maps them into vectors has been trained on a large amount of text (just like ChatGPT), and it clearly understands the relationships between different sentences (such as hearts and Alice, rings and love). These relationships are reflected in the distance between two sentence vectors in the high-dimensional space (imagine two objects in a three-dimensional coordinate system). So the retrieval process is, as we just mentioned, converting our query into a vector, calculating which vectors in the semantic space are closest to it—those are the 'memories' we want—then extracting and concatenating them and returning them to the LLM. This is what makes up a RAG library. As it happens, I also know that these sentences and words are composed of tokens, so how we reasonably select them and convert them into RAG vectors will naturally greatly affect the subsequent retrieval results (just like converting the sentence 'Oranges are the only fruit in the world' into a vector versus converting 'oranges,' 'world,' 'only,' and 'fruit' separately into vectors will lead to different retrieval results).

So how should we split paragraphs and convert them into RAG vectors (also converting them into tokens first)? I think splitting should take into account semantic completeness, structure, and the capabilities of the embedding model. There are two strategies we can choose from: the chunking strategy (splitting the entire article into different paragraphs) and the tokenizer strategy (splitting a paragraph into the smallest semantic units, involving compression and generalization).