{ Yutong Zhu }

developer • student • robot
photo

AgentAgentAgentAgentAgentAgent

webGL
webGL

Agent似乎没有那么深奥——上下文注入、会话管理、结构化工具调用、判断什么时候结束然后给用户返回消息,好像也就这样。

首先我不是要做一个多么牛逼的Agent,类似于Claude code、codex那种,虽然我觉得市面上这些牛逼的Agent其实都是coding Agent,只不过面向特定场景需求或者任务设计了特殊的架构以及工具。也许是特殊的上下文管理,或者不同的A2A通信协议,会话管理架构设计,Agent循环的逻辑设计等等。直觉上我觉得pi这个框架非常好,比较基础,而且我喜欢他的创始人以及他们公司的理念。(注入了热情的东西就会被别人喜欢)

Agent的工具系统,我们可以从工具定义,工具加载,工具执行三个方面来进行一些讨论,tool_call或者说function_call,原理都是先告诉llm在何时可以调用哪些工具,llm自行决定何时调用,调用方法就是让llm返回一段结构化数据(JSON schema),运行时读取这个结构化数据运行相应的工具(往往是一些脚本,或者联网搜索),处理工具输出并汇总整理返回给llm。其实就是给了大模型一种方法与现实的网络世界进行交互,只不过llm还不够聪明,能力还不够(context),所以才整出了这么多的subagent,MCP,上下文工程,状态管理,工具集成等等的玩意。我觉得llm智能到达一定程度,给它一个bash就完事了,它把所有事情都实现了,但是没有人会把自己交给概率。

意味不明
意味不明

话说回来,工具定义就三个变量: name description parameters,分别是工具的名字,描述和参数。这些是API协议字段,也就是 OpenAI Chat Completions API规定的,也就是说进行Completion HTTP请求时就这么传,其他的还有 model messages tool_choice以及字段里面的结构等等。对于工具的字段信息都会在创建一轮Loop的时候被加载进上下文中,它不是放在message后面,它只是作为tools字段被加载,如何处理这些信息是provider做的事情。(我记得研究说过上下文工程是很重要的,起到了引导llm的作用,一般越是在上下文的前面起到的作用越大,但是在加载工具时就不考虑这个了)

在准备好工具之后我们开始真正的创建这个agent,我们首先要注册刚刚创建的工具,这是为了让llm正确返回用于调用工具的结构体并且让这个结构体被runtime正确的execute(我要说的是对于初始化Agent时候来说的注册,实际上新建tool的时候手写好schema然后在index.js里注册)。根据我们的目的,注册总共有两个部分,将工具的名字和对应的实例放进一个Map中(这里还有execute),当llm返回一个schema的时候,根据schema中的name去map中找对应的实例然后执行;还有就是要传给llm的字段,告诉它有哪些工具怎么用,只需要传name, description, parameters:

typescript
this.toolSchemas = tools.map((t) => ({//tools是构造函数参数(可以理解为结构体中的变量),tools:Tool[] 这里tools定义为一个空数组,但是定义了它的结构,包含三个参数。
  type: "function",
  function: {
    name: t.name,
    description: t.description,
    parameters: t.parameters,
  },
}));

这里的tools在我们创建服务时使用 createTools()函数生成,这个函数在index.js里,真正的生成所有tool的JSON schema并作为一个数组传给tools,因此在创建一个新的工具时记得在这里面注册一下。

好啦好啦,llm知道你有哪些tools以及如何使用,Agent也知道怎么处理llm返回的结构体并且执行对应的tool了,这一切都是为了让这个只会说话的llm学会干点实在的活。让循环跑起来吧,不断的追加messages把llm给塞满,其实它也没多聪明是吧。

APIAPIAPIAPTAPI
APIAPIAPIAPTAPI

webGL
webGL

An Agent doesn't seem all that profound—context injection, session management, structured tool calls, and deciding when to finish and return a message to the user. That seems to be about it.

First of all, I'm not trying to build some super impressive Agent like Claude Code or Codex. Although I think all those impressive Agents on the market are actually coding Agents—they just design special architectures and tools for specific scenario requirements or tasks. Maybe special context management, or different A2A communication protocols, session management architecture design, the logic design of the Agent loop, and so on. Intuitively, I think the pi framework is very good and fairly foundational, and I like its founder and their company's philosophy. (Things infused with passion tend to be liked by others.)

The Agent's tool system can be discussed from three aspects: tool definition, tool loading, and tool execution. The principle behind tool_call, or function_call, is to first tell the LLM which tools it can call and when; the LLM decides on its own when to call them. The calling method is to have the LLM return a piece of structured data (JSON schema), and at runtime the system reads this structured data to run the corresponding tool (often some scripts, or web search), then processes the tool output and summarizes and organizes it before returning it to the LLM. In fact, this just gives the large model a way to interact with the real online world. It's only that the LLM isn't smart enough or capable enough (context), which is why so much stuff like subagents, MCP, context engineering, state management, tool integration, and so on has been created. I think that once LLM intelligence reaches a certain level, giving it a bash is all it takes—it would implement everything. But that's probably impossible; no one would entrust themselves to probability.

Here, tools is generated by the createTools() function when we create the service. This function is in index.js; it actually generates the JSON schema for all tools and passes it as an array to tools. So when creating a new tool, remember to register it here.

Unclear
Unclear

Anyway, a tool definition has just three variables: name, description, and parameters—the tool's name, description, and parameters, respectively. These are API protocol fields, as specified by the OpenAI Chat Completions API. That is, this is how you pass them when making a Completion HTTP request; other fields include model, messages, tool_choice, as well as the structures inside the fields. The tool field information is loaded into the context when a loop is created. It isn't placed after the message; it's simply loaded as the tools field. How to handle this information is the provider's job. (I remember research saying that context engineering is very important and plays a role in guiding the LLM. Generally, the earlier something appears in the context, the greater its effect, but when loading tools this isn't taken into account.)

After preparing the tools, we begin actually creating this agent. First we need to register the tools we just created. This is to let the LLM correctly return the struct used to call the tool and let the runtime correctly execute that struct. (I mean registration for initializing the Agent; in practice, when creating a new tool you hand-write the schema and then register it in index.js.) For our purposes, registration has two parts. Put the tool name and its corresponding instance into a Map (there's also execute here). When the LLM returns a schema, look up the corresponding instance in the map by the name in the schema and execute it. The other part is the fields passed to the LLM, telling it which tools exist and how to use them; you only need to pass name, description, parameters:

typescript
this.toolSchemas = tools.map((t) => ({//tools是构造函数参数(可以理解为结构体中的变量),tools:Tool[] 这里tools定义为一个空数组,但是定义了它的结构,包含三个参数。
  type: "function",
  function: {
    name: t.name,
    description: t.description,
    parameters: t.parameters,
  },
}));

Here, tools is generated by the createTools() function when we create the service. This function is in index.js; it actually generates the JSON schema for all tools and passes it as an array to tools. So when creating a new tool, remember to register it here.

Alright, alright. The LLM knows which tools you have and how to use them, and the Agent also knows how to handle the struct returned by the LLM and execute the corresponding tool. All of this is to teach this LLM that only knows how to talk to do some real work. Let the loop run: keep appending messages to stuff the LLM full. It's not actually that smart, is it?