Trust the model, tool the rest: freeing the Sardinia agent

What can my agent answer from its own knowledge vs what does it need assistance with? The answer to this question shaped the tools and knowledge base I built for this agent. Can an agent, with a good enough LLM model behind it, tell tourists general facts about Sardinian beaches, food and culture? Yes, confidently with a low chance of hallucination. Can it tell the tourist what events are on a given weekend, the opening time for a restaurant or weather conditions? It cannot. Support tools only exist for the categories the weights can’t be trusted on.

The two-state agent

Although the agent started out as a simple trip planning tool for Sardinia, the idea was to have it be a Sardinian expert, which is also capable of planning and organising a trip for the user upon request. This distinction requires some decisions to be made on how the infrastructure behind the agent is built, where the tools live, and how they are accessed by the main agent. Describing the planning tools on every answer is wasted tokens for a capability most turns never use. My solution was to set up planning as a loadable skill behind a door, which automatically loads in the extra context, updates the tool belt, and the instructions.

The how, and what it broke

The main agent needed to know when it was in planning mode, and where possible, I wanted to make the door mechanical. For any new conversation (within a new trip), the agent is told in its system prompt to call a specific tool when there is clear planning intent by the user. For old chats (or new chats in existing trips) the assumption is that they should already be in planning mode, so on the code side, if the trip object has anything in it we assume already we should be in planning mode without needing LLM judgement on that.

The new agent being allowed to answer more freely from its own knowledge saw my ‘grounding’ eval score drop from .641 (on a redesigned suite) to .347 after the first run. The agent came in with some real inventions, wrong hike distances, opening hours from memory, and answers padded with unverified restaurant names. After some considerable work on sharpening the base prompt, I was able to lift it slightly. However, the real improvement in grounding came after the web search tool was introduced, as I saw the ‘grounding’ score close at .831.

Support tools, evals first

I decided to integrate a web search, a driving time, and a weather tool. The process for building and integrating these tools with my agent isn’t too dissimilar to how I built the agent itself; start with a set of tool specific eval cases, intended to measure the effectiveness and quality of the overall agent; run the eval cases against the agent before the tool is built (expect a low result); build the tool; evaluate the results; fix and improve any caught bugs; move on. The decision to use this process meant that I was able to track step by step where the new tool was doing its job, where the limitations were, and any changes that had to be made to any system prompts.

One finding as I added more tools was that the system prompt, especially with a weaker Haiku model at the helm, needed to be very strict, without contradictions, concise, and really well structured. After a few iterations of tweaking the system prompt and re-running evals, some things I had to move into code, as some instructions simply won’t be followed consistently. The ‘model agnostic’ infrastructure I have for my agent means that I could give my agent a stronger model at any time. The reason for using a ‘weaker’ Haiku model during these early phases is I don’t want to hide any simple, fixable inefficiencies in my prompts and tools behind a stronger reasoning model which might be able to cover them up.

This block ended with a full eval re-run. The plan quality score remains lower than I expected at .591, with lots of room for improvement. The next session is likely going to be one targeting the plan directly.