What a knowledge graph actually buys
Knowledge graphs have been hyped to be this superpower tool for AI systems, supposedly improving the performance and efficiency of certain tasks substantially. The way the data is structured in knowledge graphs allows for quick access to the relationships between nodes, enabling a type of queries which regular databases do not intuitively allow. A small tool I built gives the LLMs the means to query a knowledge graph via Cypher, with which I ran a test to see how the same Claude Code agent performs with and without the tool, ceteris paribus.
Test
The data is from the SEC’s EDGAR. I built a knowledge graph taking data from the public filings, which outlines the relationships between entities and people (banks, accounting firms, directors, owners, etc.). My agent’s job is to answer queries about intricate multi-jump connections and relationships. I’m measuring accuracy, token cost, and time taken to answer the same set of questions given to the same agent with and without my graph query tool. I expected there to be a difference in quality of output (I did not expect the graph-less LLM agent to get every answer right) as well as in the speed of execution, as one of the benefits of a knowledge graph is how quickly you can answer these types of queries. I should mention that both agents run in the Claude Code environment, meaning they both have all the regular tools and capabilities of one of the best harnesses in the world.
Results
My eval is made up of a set of 16 questions, run 3 times each (for both agents - 96 runs in total). The questions were subdivided into ‘lookup’ questions and ‘walk’ questions, where ‘walk’ questions were the ones which required a number of hops between nodes, and where I expected the most disparity between the two agents. The results are in, and the accuracy was exactly the same: 100%, neither agent got a single question wrong, making the finding that there is no measurable difference in quality of the result with and without access to a knowledge graph in this case. The differences showed up on both speed and costs, particularly in the ‘walk’ questions; the agent with the graph tool took around half the time to answer the ‘walk’ questions, and used around 25% fewer tokens than the graph-less agent. So even though the quality was not different, there is a clear advantage in both time and cost to the agent with the knowledge graph.
A couple of considerations I must make: one is that this assumes the data is readily available to find elsewhere (in the absence of a knowledge graph). In this case the graph-less agent was able to go on the web, and find a ready-to-use, well-documented dataset on the SEC website. Other tests I ran showed that Claude Code does miss some nuances when trying to recreate a knowledge graph from scratch, and might carry those inefficiencies also when navigating a dataset to answer the queries. The second consideration is that what actually slowed the graph-less agent in my test wasn’t the number of hops - it was the fan-out at each hop: reading many filings wide to find the one connection, where the graph answers it in a single query. The quality of the queries could and should be tested with some harder questions, as my set of 16 questions didn’t include the most challenging ‘fan-out’ questions, something a capable agent may not be able to handle today. My hypothesis is that the quality of the graph-less agent will decrease as the fan-out increases. Some trickier questions are required for further clarification.