Evals before optimising: the Sardinia agent
For the last few weeks, the focus of my work has been on AI agents’ evaluation systems. I am working on a project which aims to be a substitute for a traditional ‘Visitplace’ website, specifically for Sardinia. An interactive, specialist AI agent which helps tourists discover the island and plan their visits and activities.
So far, I had put together the first version of the agent, the simplest version worth measuring. My goal was to be able to grade this agent, and have a baseline score; a tangible score which I can aim to improve over time.
Why before?
Before diving deep into the agent’s system prompts, tools, RAG or any of the many fancy systems available today, my decision was to first think about how to quantify its performance, and keep the score over time. I wanted a well structured eval system in order to identify and classify any improvements, whilst also catching any potential slight regressions, or even catching bugs before they ever reach production. The eval system will be ever-evolving: as users report bugs or edge cases, and the agent becomes stronger and more complex, I will continue to expand the systems which evaluate its performance.
The system
The eval system was built to grade every one of my 33 cases across four fixed dimensions; grounding (does it state anything not backed by my resources), plan quality (are the picks and locations good), conversation (does it respond naturally, one question at a time, refuse off-topics gracefully) and consistency (does the plan live in the trip object rather than restated in chat, and do edits change only what they should). A separate unscored efficiency panel was built to measure tokens, cost, tool calls and anything which might be useful outside of the LLM behaviour.
Cases themselves come in three modes since not all behaviours need the same machinery: scripted cases are fixed user messages, cheap and easy to reproduce, right for single shot traps and clear tasks; N-1 replay cases freeze a recorded conversation as history and grade only the agent’s next reply, enabling grading on deep in conversation behaviours which you cannot otherwise consistently reproduce; and simulated user cases, putting a separate LLM actor behind a persona card, to test how an agent manages a conversation in its entirety, not just one message at a time.
Grading itself was layered so that code checks everything that code can, an LLM judge answers small per-case yes/no assertions, and per-case mark schemes carry my local knowledge. The focus on Sardinia specifically (my home island) meant that I was able to personally calibrate the LLM judges directly by grading transcripts blind (critique shadowing), landing at 88% agreement (85% on a held-out slice), with every remaining disagreement being the judge stricter than me.
It is worth mentioning that this eval is only a part of what it really takes to build a great agent; a system for collecting and categorising user feedback is also very important. A product or agent is not only as good as its benchmarks show. An agent which is optimised to be good at answering the questions to a test, may be much better on paper, but much worse in the hands of real users.
First run and results
Across 99 trials (33 cases x k=3), the picture came out the way a good baseline should; ugly in the right places. Consistency was near perfect (.992), pretty good at grounding already (.835) but left a lot of room for improvement in conversation and plan quality (.762, .585 respectively). The reliability number is the sobering one. Only 9 of the 33 cases passed all three of the trials; 18 of 33 passed at least once. So 9 cases pass or fail depending on the run, and 15 never pass at all. The same question asked three times does not get the same quality of answer. That alone justifies k=3 (running each case three times).
This first version of the evaluation system is not perfect (nor should it be at this stage), but it works. The scores came out lower in the places where I wanted headroom, as a perfect score on my first evaluation system, implies the ruler is too lenient, not that the agent is perfect. Now the focus shifts back to the agent itself, plan quality is target one.