We are running an AI transformation on ourselves, in public

Everyone around us is talking about how much AI helps them. Almost nobody talks about how much more efficient it actually made them, and nobody at all talks about how to measure that.
We know, because we get the question constantly. Clients ask us how much more productive our engineers have become with AI. We have an answer, and it is honest, but it is a feeling based on many years of experience, not a number. And engineering is the easy case. Take finance: YouTube is full of videos about running a company's finances with AI, but nearly all of them describe narrow one-man-show setups. Does any of it work inside a company that has been operating for years, with a real finance team, real auditors, and real consequences? Nobody shows that.
I have spent 25 years automating business functions and watched technically successful projects die simply because nobody wanted to use them. That teaches you to be careful with promises. But with the release of Claude Opus 4.7, something crossed a threshold: models stopped being autocomplete with good manners and started reliably carrying multi-step work that needs context and judgment. When the tools change that much, "we are professionals, we do not chase hype" stops being a position and becomes an excuse.
At some point it simply became obvious: this future is not coming, it has arrived. And that is where we caught ourselves in the oldest pose in the trade: the shoemaker whose children go barefoot. We help financial firms adopt technology for a living, while our own company still runs mostly the way it did three years ago. Do not get us wrong, it runs well and lean, and we are quietly proud of that. But it runs without AI. Will AI make it more efficient, or just more expensive? We honestly do not know. We may well end up where we started, without AI in half of these functions, but this time with numbers, and able to look a client in the eye when they ask us that question.
So we are running the experiment on ourselves. Findev is an IT consulting firm in capital markets: 180 people, entities in several countries, several languages and currencies, counterparties across many jurisdictions, clients who trust us with sensitive data. Our example is not universal: we are an order of magnitude simpler than the firms we serve. But we are also far from a one-man show, and that middle ground is exactly where almost nobody publishes anything.
Function by function, for every recurring piece of work, we will ask the same question and publish the answer: should this stay with a human, become a deterministic script, be handed to a third-party service that has already implemented it well, or be something we build with AI ourselves?
We will publish what works, what does not, and where we simply get stuck. And what it all costs, including the token bills nobody ever shows. Costs, to be clear, not trophies: measuring effectiveness by token consumption is the new counting lines of code.
We are not doing this because we expect to get it right. We will not. Mistakes in this kind of shift are unavoidable, and so, we increasingly believe, is the shift itself. We would rather go through ours with the door open. If our wrong turns save you even one of your own, this series has already paid for itself.
We hope this made you curious enough to keep reading. We can promise the reading will be honest.