The root file is not the main event
When using a coding harness like Claude Code or Codex we often think of the root files as the main event of the entire harness. In production setups however this is rarely the case, as most of the context tends to be stripped out of the CLAUDE.md/AGENTS.md root files. I pulled apart four production harnesses this week - Remotion’s, Sentry’s, Anthropic’s own, and a 30-skill design system pack.
What these production grade harnesses share is the way in which they are structured, with each of them having a very short or no root md file, allowing the skill names and descriptions to carry much of the needed context at every turn. They all carefully make use of the skill descriptions to ensure the skill body context is loaded when needed, and for the more complex skills the reference files are loaded on demand. Most files in the four harnesses stay under 100 lines (109 out of Remotion’s 161 for example), as a way to manage the context window for performance and cost reasons. Sentry runs 40+ skills in their harness, and only 3-4 have any structure beyond the single SKILL.md file; the normal skill is ‘boring’, one file, clear description, clear instructions.
Why it works
This structure is mainly ‘context management’. As the LLM carries more tokens, its response quality tends to decrease. If the extra context carried is also not relevant to the user’s query, it becomes even worse. What Anthropic, Remotion and others do extremely well, is the structure which they use to manage and access the right amount of context, for each query, especially in the more complex skills. The skill description is the first door which allows the LLM to recognise the need to go and ‘read’ the SKILL.md to help with answering or executing the user’s query. The description is concise and situation-phrased enough for the LLM to recognise when the skill is needed, and nothing more. In the more complex skills the SKILL.md file (the first file which the LLM automatically ‘reads’ once it decides to use a skill) is itself a ‘router’: rather than bloating the context window with the entire contents of a complex skill, it directs the agent to a sub category most likely to hold the exact context the query needs. This can go on for several layers. In Remotion’s case, one of their skills points to a router pointing to another router pointing to the actual content, with each hop telling the agent to choose a path, and loading only that. Routers stay short, content mostly lives at the bottom.
It’s testable
Whether a structure works or not is testable. Run a realistic query, and watch the LLM call (or not call) the right skill and files as it works through the response. As mentioned before in some of my writing on eval systems when building an agent, I am always looking for ways to test and evaluate the work as I build. In this case, the checks are binary, the agent either went down the route you designed to get to the context it needed, or it didn’t and there is some more work to do.
The caveat
Having looked at some harnesses in great detail over the last week, a caveat comes to mind. A well structured harness can be easy to use, but it doesn’t mean it will be used. If a harness isn’t being adopted and the claim is that routing alone will fix it, there needs to be a question about whether the harness is solving the right problems to begin with. Oftentimes restructuring and optimising results in optimising the wrong things, which is a bit of a recurring theme in this industry.