Essay | August 6, 2026 | 4 min read

How Do You Make Your Agent Cheaper and Better?

Efficiency stops being a vibe when you measure quality over every dollar spent.

Technical Systems
-- views

Everyone has been talking about API costs and efficiencies. The frontier model companies keep announcing that inference is getting cheaper, that their next model is more efficient, that the cost per million tokens keeps falling. Great, and useful. But nobody can actually say whether their agent is getting more efficient. There is no concrete method to measure it. And without that, every claim of efficiency is a vibe.

In a budget constrained world, the efficiency measurement starts with a very simple question: how much quality am I getting over every dollar spent on my agent?

Quality is one side of that ratio, and there are ways to measure it. Cost is the other side, and it is where most builders stop looking. So look at it. Where does the cost actually come from? More precisely, what fraction of your token cost is input, and what fraction is output?

I investigated and the answer for me was mind-blowing: 99 percent input, 1 percent output. Nearly every dollar my agent spends is on tokens I put into the model, not tokens the model gives back.

Of every 100 dollars spent on Monika's agent, 98 dollars and 93 cents goes to input and 1 dollar and 7 cents goes to output; input accounts for 99.78 percent of raw tokens.
Where my agent's dollars actually go. Input tokens dominate cost by nearly two orders of magnitude.

That number was uncomfortable at first, because it flipped where I thought the optimization opportunity was. I had been thinking about output. How the model responded, how verbose it was, whether we could get it to say the same thing in fewer words. But output is the model's choice. Input is mine. Every skill I load, every tool result I keep in context, every message from three turns ago that I decide not to trim. Those are all decisions I made, and each one shows up in my bill.

Physics has a word for this. Density is mass over volume. A dense object carries more matter in less space, and the density of a system is what tells you how efficient it is at concentrating what matters. I started thinking of my agent the same way. Its mass is quality. Its volume is tokens. Every input I add to a call adds volume, but not always mass. If the volume grows faster than the mass, my agent gets thinner, less dense, less efficient per unit of what it costs me to run.

So I built a metric that puts Quality and Cost together. I call it Harness Signal Density (HSD), measured in quality units per million tokens.

HSD = (Quality × 1,000,000) / Cost per turn in tokens

Harness Signal Density, measured in quality units per million tokens.

The formula is deliberately physics-shaped. It gives me a single scalar for my agent that goes up when quality improves, up when tokens per turn drop, and down when either moves against me. It works whether my agent is used once a day or a thousand times a day, because it measures density, not aggregate consumption. And unlike a ratio to a baseline, it is an absolute quantity, which means I can compare one experiment to another without threading everything through a starting point. What gets measured gets managed. HSD is what I chose to measure.

Cal harness evaluation showing that Cal is 3.40 percent denser than day zero, at 0.505 quality units per million tokens.
HSD in practice. The metric that lets me say whether an optimization actually made my agent denser, or just moved cost around.

Every input-token optimization I have shipped so far is one experiment against this metric. Retrieving skills instead of loading them all. Trimming stale conversation history. Caching prompts. Each one is a small hypothesis about where the mass and the volume are diverging in my agent. More will follow.

If you are building an agent harness, I would love to know: what is your input-to-output token ratio? Drop it in the replies. My hunch is that most agent harnesses are input-heavy in a way their builders have never quantified, and once you see the number, you cannot unsee where the leverage is.

I will be writing a full blog post soon on the why and the methodology behind Harness Signal Density, along with concrete strategies from my experiments and empirical results that you can apply to improve your own agent.

Want to measure your Harness Signal Density? Watch this space.

Ideas, views, and opinions are my own and do not represent those of any organization I am affiliated with.