Medicine distinguishes between efficacy and effectiveness. Efficacy is whether a treatment works under ideal conditions, in a trial, with a motivated patient and a careful doctor. Effectiveness is whether it works out there, in the real world, with real patients who forget to take their pills.
For AI in 2026, the second one is the interesting question: not whether it works in a demo, but whether it works out here. And I want to add two more Es: efficiency, and one that is too often skipped: ethics.
Effectiveness: yes, it works
Let me get this out of the way first: it works. I do not mean “it produces plausible-looking text”, I mean it gets real work done that I would not have gotten done otherwise.
My favourite example is asba. It reads the CD-ROM edition of Arno Schmidt’s collected works and concordance from 1998, which ships a 16-bit Windows program, a pile of Microsoft MediaView DLLs and a 275 MB concordance file. Claude reverse-engineered the file formats, and along the way found out that a second file contains the complete text of the edition, uncompressed, plus an index that maps each of the 2,252,290 concordance entries to its exact position in it. The result is a small Python tool that searches the CD directly, no emulator needed.
It took all of 20 minutes, and I wrote exactly one prompt. The transcript is public, if you want to see for yourself. Without help, this would have stayed on my list of “things I would do if I had the time” forever.
This is also why using an LLM in a harness like pi or Claude Code is not the same as “asking an AI questions”. In a chat, all you get back is text, and plausible-sounding text is what these models are built to produce — including “facts” that are simply made up. In a harness, what you get back is code, and code can be checked: it compiles or it does not, the tests pass or they do not, asba finds the word in the concordance or it does not. Those checks are only as good as whoever wrote them, though, so I read the tests myself. And when something is wrong, you can point at the failure and steer it back on course.
Claude is not conscious. But it is surprising how intelligent an LLM in a harness can appear, once it can read files, run commands and look at what happened. We used to be amazed by dogs that understand 500 words; that is going to stop being a measure of anything.
asba was small enough that I could just let it run.
For anything larger, the trick is not to.
Kristian Köhntopp describes a guide-rail method in the AGENTS.md of one of his projects: first user stories, then tickets, then code, one ticket at a time, each step committed before the next one begins.
It is basically the process a good team would follow anyway, written down so the machine has to follow it as well.
With rails like these, the output stops being a slot machine and starts being something you can review.
I tried that on a bigger project: terraform-provider-homeassistant, an OpenTofu provider that manages a Home Assistant instance as code.
The design process came from the skills at AI Hero: first a glossary, a spec and decision records, then tickets, each one a small vertical slice.
So far there are 24 decision records and 37 tickets.
The decisions stay mine. The AGENTS.md only lets an agent write a decision record after it has put the options to me and I have answered.
Once the tickets exist, several agents work at the same time.
Each one claims a ticket, works on it in its own worktree and opens a pull request.
Nothing gets merged until I have commented “LGTM”.
What it does in my actual job
asba and the Home Assistant provider are hobby projects. What about my job?
I mostly do devops. There is very little business logic in what I write; it is configuration, glue, pipelines, and the occasional operator. That is the kind of work where AI shines: well-trodden paths, lots of prior art, and a quick feedback loop that tells you whether it worked.
But I would argue that even developers who do write business logic will mostly be doing what I do: connecting one API to another API. The interesting part — deciding what should happen — was always the small part of the job. The large part was plumbing, and plumbing is what gets automated first.
None of this makes an LLM nice to use if you want to write software. If the puzzle is the point — finding the right abstraction, the moment it all clicks — the machine takes that away from you. But most of the time, I do not want to write software; I want to have written software, because what I actually want is to use it. Think of a cupboard. You can buy one: quick, cheap, and it fits your room more or less. You can build one with power tools: it fits exactly and it is yours, even if you did not cut every joint yourself. Or you can build one with hand tools, plane every board and cut every dovetail, because the building is the whole point. Off-the-shelf software is the bought cupboard, and writing code by hand is the hand tools. The LLM is the power tool in the middle: software that fits my problem, without me having to enjoy every step of making it.
The other comparison people reach for is hand-written assembler versus a compiler. It limps: a compiler is deterministic. The same source gives the same binary, and when the compiler is wrong, that is a bug somebody fixes once. The same prompt, given twice, gives you two different programs. But one part of the comparison may hold: the people who wrote assembler did not stop programming when compilers came along. They moved one level up.
A note on side effects
A harness is a program that executes arbitrary commands on your computer, chosen by a text predictor. All the checking I praised above happens after those commands have run, not before. Have a backup. Better still, let the agent that changes things run in a sandbox, and only run a planner on your own machine. Even that is not watertight. A planner also reads files and runs commands to find out what to plan, so it can run things you never meant to run, too.
Efficiency
I pay $20 a month for my Claude subscription. I am fairly sure that this does not pay for what I use.
A tool that is only cheap because somebody else is burning money to subsidise it is not efficient; it only looks efficient on my bill. Any honest assessment of “AI makes me more productive” has to include the question: at what price, once the price is real?
Ethics: not aside
When the Rust project drafted its LLM policy, the pull request declared a list of topics off-limits for its comments: the long-term social and economic impact of LLMs, their environmental impact, the copyright status of their output, and moral judgements about the people who use them. “We still consider these topics to be important, we simply do not believe this is the right place to discuss them.” That got a lot of backlash.
“Ethics aside” has become a meme, and deservedly so. It is the phrase you use when you know there is a problem and would prefer not to talk about it. So let’s not set it aside.
Copyright
Twenty years ago, the music industry sued individual people — students, single parents — for sharing a few hundred songs. Today, companies train models on everything that was ever put online, books and music included, and the answer is mostly a shrug and a licensing deal among the large players.
I do not think the conclusion is “so we should sue the AI companies as hard as we sued teenagers”. The conclusion, to me, is that the file-sharing era already showed that copyright as we have it does not fit a world where copying is free, and AI is simply the second time we are finding that out. The difference is who is doing the copying: last time it was people without lawyers, this time it is companies with lots of them.
That does not mean I want to throw copyright away. I make my living selling my intellectual output, and I am glad that the law lets me do that. Which is why it bothers me that the same law seems to bend so easily when the copier is big enough.
Who gets to own a computer
The AI build-out is buying up GPUs, memory and storage at a scale that skews the whole market for computers. The result is that it gets harder for ordinary people to own a powerful machine. For most of my life, the trend went the other way: every few years, a normal person could afford more computer than the year before. Renting compute from the few companies that can still afford the hardware is not the same thing as owning it — although, full disclosure, I work for a company that will happily rent you a machine.
Who gets asked
Data centres have to be built somewhere, and “somewhere” is always somebody’s neighbourhood. The people living there are rarely asked — the power lines, the water use and the noise get decided long before any resident gets a say.
It does not have to be like that. The data centres I work with are all in Norway, next to hydro plants. Building where clean power already is, instead of building first and fighting over the grid afterwards, is so much more sustainable — and it shows that where we put these things is a choice.
Where that leaves me
I use Claude for work, and for hobby projects like the ones above. But I am always sceptical of large companies; they are never my friend. I would like to try some of these things with local models instead, but right now I do not have access to a machine powerful enough to run them — see above, on who gets to own a computer.
So what do I tell my kids?
I do not have a good ending for this post, because I do not have a good answer to the question that actually keeps me up at night: what should I tell my kids when they ask me what they should become?
The job I have, the one that paid for the house they grew up in, is in large part the kind of job I just described as “plumbing, automated first”.
I don’t know. If you do, I would like to hear it.
Comments
With an account on the Fediverse or Mastodon, you can respond to this post. Since Mastodon is decentralized, you can use your existing account hosted by another Mastodon server or compatible platform if you don't have an account on this one. Known non-private replies are displayed below.
Learn how this is implemented here.