2026 / 08 / 30 / the-skill-above-the-tool

Skill above the tool

Skill above the tool

I have caught myself doing something slightly unusual when I open a modern software project for the first time: I look for AGENTS.md.

Not because having one automatically makes a project good, but because I increasingly find it useful as a signal. A README tells me whether a project expects humans to understand it. Tests tell me whether its authors care about proving that it works. CI tells me whether they understood that repeatable processes should be automated. AGENTS.md tells me whether they have started thinking seriously about a new participant in the development process.

I think that distinction is becoming more and more important.

The skill was already moving up a level

Years ago, while I was still at university, I heard one of Prezi’s founders give a talk. One idea stayed with me: the modern successful person should not optimize for accumulating as much knowledge as possible, but for developing the methods and skills that let them reach the most useful knowledge for the problem in front of them as quickly as possible.

The web already pushed us in that direction. Search engines reduced the value of remembering where information lived. The more useful skill became knowing what to look for, how to evaluate what you found, and how to turn it into something useful.

AI feels like the same transition again, but one level higher.

Steve Jobs famously described the computer as a “bicycle for the mind”: a tool that gives the same human dramatically more leverage. AI extends that idea because it does not only accelerate access to information. It can work on the information too. It can explore a codebase, compare implementations, read documentation, form hypotheses, write code, run tests, find contradictions and revise its own work.

The bicycle has acquired pedals of its own.

Using AI is becoming a skill of its own

There is already an enormous difference between having access to modern AI tools and being effective with them, and I do not think that difference comes primarily from writing clever prompts. Prompting is part of it, but the interesting skill is much larger:

problem framing → context selection → decomposition → delegation → supervision → verification → iteration

Give the same model and repository to two engineers and you can get radically different results. One asks it to implement the Jira ticket. The other first maps the relevant part of the system, finds existing implementations, establishes invariants and pulls in references where useful. They decide what can be delegated and what requires judgment, let the agent implement, then review and test the result. Failures become context for the next iteration.

The model did not get smarter between those two sessions. The human-agent system did.

This is why describing AI-assisted development as simply “writing code faster” badly undersells what is happening. The interesting question is becoming less about how quickly I can personally produce syntax and more about how much reliable computational leverage I can direct toward a problem.

That still requires engineering skill. In many cases it requires more judgment, not less: you need to know what good looks like before you can recognize whether an agent produced it.

Why I look at AGENTS.md

This brings me back to that file. A good AGENTS.md is not interesting because agents happen to need a special Markdown file. It is interesting because writing a useful one forces a team to answer a new class of engineering questions.

What is authoritative in this repository? Which commands actually validate a change? What architectural boundaries should not be crossed? Which implementations should be treated as references? What should an agent never modify automatically? Which assumptions are important but invisible from the type system? When should it stop and ask for a decision instead of inventing one?

These are really variations of the same question: how do we turn an extremely capable but imperfect machine collaborator into a reliable participant in this particular engineering system?

A repository that has thought carefully about that feels different from one that simply gives an agent the codebase and hopes for the best. The file itself is not the maturity signal. The thinking behind it is.

Then someone started ranking it

The funny part is that I had already been thinking about all of this when an old friend of mine and former boss, Péter Varga, launched something that almost reads like an experiment designed around the same question.

It is called VibeLadder.

The format is deliberately constrained: everyone receives the same hidden application task, gets 90 minutes, can use any AI tool or model they want, and submits one self-contained result along with the transcript of how they got there. The submissions are then compared against each other.

The important part is what VibeLadder explicitly does not rank: the AI. The tool is your choice.

It ranks the human.

That made the whole thing click for me, because the actual object being measured is neither traditional coding speed nor model capability in isolation. It is much closer to the quality of the human-machine loop.

Can you turn an ambiguous goal into useful context? Can you keep an agent moving in the right direction and notice when it confidently takes a wrong turn? Can you spend your limited attention where it creates the most leverage? Can you get from intent to a working, verified result faster and more reliably than someone else with access to the same generation of tools?

Those are real engineering questions. That is why I think VibeLadder is more interesting than just another weekend coding competition: it is trying to put a number on a skill we are only beginning to recognize as a skill.

If you are experimenting seriously with AI-assisted development, it is worth climbing the ladder yourself: https://vibeladder.dev

Maybe we need new benchmarks

For a long time, we built ways of measuring software engineers around how humans produce software directly: coding exercises, algorithm problems, take-home assignments and competitive programming. Entire hiring platforms grew around making some version of that ability measurable.

Those skills are not suddenly irrelevant. But I would be surprised if the next generation of engineering benchmarks measured exactly the same thing.

If more of our work becomes directing, constraining and verifying increasingly capable agents, eventually we will want ways to measure that ability too. Maybe a 90-minute Saturday vibe-coding ladder is just a fun experiment. Or maybe experiments like this are an early version of something we will later consider completely normal.

Either way, I find that direction much more interesting than another argument about which model currently scores two points higher on a coding benchmark. The model matters. The tools matter. But there is another layer above both of them.

And that layer is becoming a skill.