I spent a good while stuffing our CLAUDE.md files with everything I could think of, and the best thing I’ve done with them since is go and delete most of it. For ages I treated that file as the place to be helpful - every convention, the big gotchas, anything I thought a future session might want to know - and what I actually built was a pretty effective way of making the model worse at its job.
The instinct behind all that bloat was fine, I just had it pointed the wrong way. We were leaning a lot harder on agentic coding across the team and I wanted those sessions to go well, so I wrote down everything I knew - how the tests were set up, what the linter would whinge about, the patterns we liked, the shape of the codebase, all the stuff that had bitten us before. Every time a session went sideways I’d go and add a paragraph explaining the thing it should have known, because if it got something wrong and I hadn’t told it, that felt like it was on me. One of my engineers at the time kept pushing back on all of it, and I kept explaining why the extra context was worth the tokens. The assumption sitting underneath that whole argument was that more detail equals better output, which I’d never actually tested, I’d just taken it as read and got on with it. It’s full of shit as it turns out. They were right and I was wrong and it took me way longer than it should have to admit that.
What it was actually doing, though, is that every session kicked off carrying way more than it needed for the job in front of it. So instead of turning up to the problem cold and having a poke around, the model turned up having already read a couple of hundred lines of my opinions about how work gets done around here, and then it went and did the thing my opinions implied. What came back was consistent, and it looked right, and a lot of the time it wasn’t the best answer available, because I’d already fenced off most of the options before anyone had even read the ticket. I’d have told you at the time this made the sessions safer, but mostly what it did was fence them in.
Worth saying that nothing measured any of this for us. There was no dashboard where problem solving quality went down and to the right. It came out of retros, which is the boring answer - just talking about what went wrong in a session, week after week, with the config sitting in a repo where changes went through small PRs the same as anything else. Have enough of those conversations in a row and the shape of the thing eventually turns up, same as any other cultural problem on a team. There’d been a bunch of small changes up and down before we got there and there’s been plenty since, but once it was properly obvious how bad it had got we did one big cut. The global file went from a couple of hundred lines to about twenty or thirty. The repo ones landed somewhere around forty or fifty each, coming from a mixed bag where some repos had a file and some had nothing at all.
The improvement was felt the day that cut merged — at least I felt it. To be fair there are a few things that I think could have stacked up on this feeling though. We had seen the models we use change over the time I had written the file, like going from Opus 4.5 to 4.7, and there is a not 0% chance that the change in model also had some impact on this. But given the change I saw on day one of cutting the file down, my gut says the smaller file did more of the lifting than the model upgrades that happened alongside the growth of the file.
Same rule at multiple layers
At the time where I was playing around with this I had the same problem at a couple of scopes - the global CLAUDE.md file (which we used in a shared repo so every engineer pulled the same thing), as well as the repo specific version of the file. I had managed to overbake the hell out of both of these. The repo specific ones were at least getting more eyes on them through regular code reviews, and by virtue of being in the repo that everyone would open when making changes. The global one being more hidden away and out of that day to day flow definitely drew less attention than something with that big of a blast radius should have.
What was left of the global one by the time I stopped hacking at it was how we think about reviews, a pointer at an optional personal preferences file so nobody had to give up their own setup to the team config, a couple of security standards, and pointers to the source control and workflow stuff. That’s it. Twenty odd lines as the standing instructions for a whole engineering org felt ridiculous the day I cut it back that far, though I got used to it pretty quickly.
The discoverability test
The rule I’d use now is pretty simple. If the model can work it out by reading the repo, don’t write it down. Not write it down shorter, just don’t write it. There’s nothing in our files describing the test framework, because the test framework is sitting right there in the config and the model can read it quicker than I can explain it. No linter docs either. What’s there instead is one line saying tests and lint pass before you open a PR, and that’s the bit it can’t work out for itself, because it’s a decision about what we expect and not something sitting in the codebase.
The stuff that gets through that test comes in about three shapes.
First one is invariants, meaning the things that compile, pass tests, and are still wrong. That’s the highest value writing you can do for a coding agent, because it’s the one category that no amount of reading the repo is ever going to turn up. On this blog’s platform, a signed cookie that gates asset delivery and must never be treated as an authorisation signal looks completely fine to a model working out intent from the code, and a model will quite happily collapse the two mechanisms into one if nothing tells it not to. Those ones need something like “if a ticket looks like it needs you to break this, stop and tell me” bolted onto them, because what actually goes wrong is the agent quietly picks a side and carries on.
Second is behavioural guardrails, and the trick with those is to state them as outcomes. Don’t refactor everything sitting next to a file you touched. Don’t go into these directories without asking. Say when you’re not sure, because a flagged uncertainty costs me nothing and a confident wrong implementation that passes its tests costs me hours. One of those came out of the team pushing back on me, and it’s in my personal instructions now - a question isn’t a criticism, I want an answer. I quite like that it outlived the team config it was written for, and it’s about the only thing to come out of this whole exercise that I’d argue is universal.
Third is mechanical gates, again written as outcomes - what has to be true before a PR opens, and leave the how to the model.
There’s a catch to all this cutting that I only half saw at the time, which is that it only works if the things you’re pointing at actually exist and can be found. Once the file stops explaining stuff and starts pointing at stuff, every pointer is a promise the repo has to keep. Our other docs were a bit sketchy on that front in places, so trimming the instructions quietly dumped a load onto documentation that wasn’t ready to carry it. Deleting the context doesn’t get rid of the need for it, you just move the need somewhere that you now have to keep up to date.
Doing the same thing on my own
I’ve since done the same trim on this blog’s platform repo, where there’s nobody around to argue with me about it. During the first proper build session I pulled the root file back to the one invariant that’s actually universal across every workspace, plus a one line index of the other twelve pointing at whichever workspace file owns each one. The full text of each one sits at its enforcement point now, so the CDK invariants load when you’re in infra/ and the auth invariants load when you’re in functions/, and a session working on the static site carries neither of them.
What the research says, for what it’s worth
I went and looked for evidence after the fact, which is the wrong way round but it’s what happened.
Anthropic’s own docs say to aim for under 200 lines per CLAUDE.md because “longer files consume more context and reduce adherence”. They’re also pretty clear that the file’s content is “delivered as a user message after the system prompt, not as part of the system prompt itself”, so it isn’t privileged instruction, it’s sitting in the window competing for attention with everything else on exactly the same terms as everything else. I knew that bit at the time, which is sort of the point. None of what I got wrong here was about how the tool works. I understood the mechanics fine and still sat there adding to the file every week, because what I had wrong was the assumption I’d layered on top of it, and that’s the kind of mistake you can keep making for ages without noticing.
Then there’s an actual study. A group at ETH Zurich ran AGENTS.md style context files across a bunch of different models and agents and found that handing one to the agent “does not generally improve task success rates, while increasing inference cost by over 20% on average”. The interesting bit is what happened when they stripped all the other documentation out of the repos first - the auto-generated context files, which had been doing nothing or slightly hurting, started helping. Which is the discoverability test coming at you from the other direction with numbers attached. A context file is only worth having to the extent it holds something the repo doesn’t already say, and if it’s just restating the discoverable stuff you’re paying for the tokens and attention and getting nothing back. Their own conclusion is more or less that these files are worth it for the non-standard stuff and not for repo overviews, which is about as close to my rule as I could have asked for.
The one bit that doesn’t line up neatly is that they found agents with a context file actually did more exploring - more searching, more reading, more tests. But the reason was the agents did pretty much whatever the file told them to, and that’s the same thing I was seeing from a different angle. My file was pointing all that effort down paths I’d already picked. Close enough that I think there’s something real in what I saw, and it’s not just me telling myself a story about it after the fact.
Still not solved
I don’t think I ever cut too far, but I also don’t think we ever got it properly right, and I’m fairly confident nothing in this space stays right for long anyway. Models change, and instructions that were doing real work six months ago end up either obvious or actively wrong, and the repo and the team keep shifting around underneath the file as well.
So what I’d actually tell you to do is the loop - keep the config in source control, change it in small PRs like anything else, and talk about what went wrong in sessions at retro until the patterns turn up. That’s how we found this one.