Skip to main content

Command Palette

Search for a command to run...

Better models won’t fix a broken harness

Updated
•6 min read•View as Markdown
Better models won’t fix a broken harness

At the latest AAIF London event, five talks approached the same question from different parts of the stack: how do you actually make an agent better?

I was there to present fx, the agent harness we’re building at Vercel. But Leon Chlon’s talk was my favourite. It challenged some assumptions behind recursive self-improvement, and his view on skills was almost the opposite of mine.

The disagreement gets to a question I think matters: as models improve, what moves into the weights, and what still belongs in the system around them?

The layers of agent improvement

Max Shaposhnikov gave the evening a useful framework: model architecture and pre-training, post-training, harness engineering, and context engineering.

Each offers a different way to improve an agent.

Pre-training builds the model’s broad foundation. Post-training shapes how it applies that foundation. The harness manages its interaction with the environment: tools, execution, observations, context management, and the conditions under which it continues or stops. Context engineering determines what information reaches the model, and when.

These layers interact. A model can know how to investigate a problem without having access to the relevant logs. It can understand a codebase but lack the organisation’s current release rules. It can have the right information somewhere in its context and still fail to use it.

What Leon meant by self-improvement

Leon’s argument was more precise than “recursive self-improvement is mathematically impossible.”

He distinguished updating weights from changing a model’s behaviour through context. Training on model-generated outputs can update weights through a loss, reward, and gradient descent. With frozen weights, prompts and the KV cache can still change responses, but those changes do not add information to the weights.

His central constraint was that a model cannot reliably supply information missing from both its weights and its input. Repeatedly looping over its own outputs (or writing more files for itself to read) doesn’t necessarily resolve that gap.

I found that compelling. An agent can rearrange what it knows, explore alternatives, and correct mistakes. But if it needs evidence it doesn’t have, more generation is not a substitute for obtaining it.

Leon’s proposed starting point was therefore “I don’t know”: recognise insufficient information, abstain, and seek evidence from outside the loop.

That is a useful design principle for agents. Sometimes the productive next action is a tool call, an experiment, or a question to a human.

Improving context is still improving the system

Leon also described a more mechanistic approach to inference-time improvement.

He highlighted how changing the order of context can change an answer, even when the relevant information remains present. His proposed work uses the positional rotations in RoPE to derive a query’s sensitivity to position, then inform changes to the KV cache against an objective.

As presented, the idea was to make context interventions more predictable than simply prompting the model to try again. The talk didn’t provide enough derivation or implementation detail for me to assess its guarantees, but the direction was fascinating.

But it also reinforced why I think context engineering remains important.

Information being present is not the same as information being usable. What reaches the model, where it appears, and how it survives a long interaction can affect the outcome.

Where I disagree: skills and spiky capability

Leon’s position on skills was more sceptical. He questioned the reliability of instructions in context and suggested that widely used skills would eventually generalise into model weights.

I agree on this somewhat, common trajectories will distill into the weights and future models may perform them without needing a separate instruction file.

But models have spiky capabilities. They can succeed at a difficult task and then miss a small, consequential detail in a familiar one. Improving average capability doesn’t tell us that every procedure works reliably under every combination of constraints.

More importantly, skills don’t only teach general procedures. They also communicate local expectations.

A model may know how to write a database migration. It still needs to know how this team reviews migrations, which checks are required, and when to stop for approval.

Those expectations can change faster than model weights and likely differ between two organisations using the same model.

Why the surrounding system can matter more than training

Benchmarks give us valuable comparisons under defined conditions. Production work requires us to establish many of those conditions ourselves.

“Fix this issue” may involve discovering the right repository, finding missing evidence, interpreting an undocumented constraint, and deciding whether passing tests are sufficient.

Helin Ece Akgül’s talk on trajectory-level evaluation showed why the final outcome alone can miss important differences in how a task was completed. Taowen Liu’s talk on agent RL made the complementary point: the environment determines what an agent can practise, and the reward determines what counts as success.

For a particular deployment, I think harness and context engineering can therefore matter more than further training.

If the model already has the underlying capability, the bottleneck may be access to evidence, tool quality, context selection, or feedback. Fixing that bottleneck can be the most useful intervention.

1 Harness to rule them all?

fx supplies the agent loop, conversation and context management, tool dispatch, and runtime policy. The host supplies execution capabilities. The surrounding evaluation or RL infrastructure supplies tasks, resets, grading, rewards, and training.

That separation lets the same harness core be used for production work, evaluation, and rollout collection for training.

If we train or evaluate with one harness and deploy with another, we change what the model observes and how it acts. Reusing the core reduces that gap and makes the results easier to investigate.

It doesn’t automatically make comparisons fair. Tools, prompts, environments, budgets, and graders still need to be controlled. And fx itself is not the trainer or reward function.

Leon’s talk strengthened my view that reliable agents need external evidence and meaningful verification. Where we differ is how much of the surrounding machinery will remain necessary.

I expect models to absorb more common skills. I also expect the world to keep supplying new tools, changing rules, missing information, and unfamiliar combinations of tasks.

That gives harness and context engineering plenty of work to do and it’s why I’m helping to build fx.

X
X1Y13h ago

The point about training or evaluating with one harness and deploying with another deserves more attention. It's the agent version of train/serve skew: tool timeouts, context limits and approval rules differ between the eval run and production, so the scores describe a system nobody shipped. Reusing one harness core for both is the cleanest fix, and your point about skills carrying local expectations (how this team reviews migrations) is the part models won't absorb anytime soon.