9 October 2026

The human code reviewer was never really the quality gate that people think it is" – why Spendesk's CTO pilots AI to review AI written code

Michał Szehidewicz

Michael Sols

15 min read

Spendesk's engineers already ship more AI-written code than a single reviewer can read line by line.

Alan Wright, its CTO, believes engineers should rely on automatic AI debugging under the condition that they will vouch for the changes pushed to production.


Fintech Recoded interviews demonstrate how experienced product and technology leaders make difficult decisions in an environment where resources are constantly constrained, market changes are sudden, and deadlines are non-negotiable.

Proof beats a demo

An engineering team can end up paying for an AI tool for months because someone got excited about it in week 1.

Alan Wright, Spendesk's CTO, delays that commitment until the tool survives weeks of real use, and a first demo counts for little.

What you'll learn

Why Spendesk gave AI a calculator for the sums it shows customers,

How a 12-week usage bar decides whether an AI tool becomes permanent,

Why Spendesk's CTO doesn't think code review was ever the safety net people assume.



About Alan & Spendesk

Spendesk launched its AI tool with read access only

Michał Szehidewicz: You spent your career inside companies that were already moving fast. Spotify went from about 1 million paying subscribers to over 100 million while you were there.

YOu took your first CTO role at Spendesk, aFrench company in a crowded market. Why did you say yes to it?

Alan Wright: Guilhem, Spendesk's technical co-founder, served as CTO on an interim basis until I arrived.

The company had just gone through a hard readjustment and emerged profitable, which is rare for a SaaS company at this stage.

The turnaround mattered, but it wasn't why I said yes.

Conversations with Axel, our CEO, about his ambitions for AI and innovation tipped me over the edge.

Building software in a regulated environment is hard, and I'd learned that at Spotify and again at Monzo.

A hard problem like that gives me energy.

Meeting every engineer across every squad in my first weeks gave me a second reason.

The conversations revealed a huge amount of enthusiasm at Spendesk, and the biggest unlock I found was removing the friction that was slowing the team down.

The discovery came before I applied any AI to the problem, and it let me make a real impact quickly.

You announced AI Connect at Money20/20 in June, and Pleo announced a similar product just days later.

How did the launch go for your team?

The launch went well.

We announced the beta in June and reached general availability in July.

The MCP server was the product itself, built so that customers could bring their own AI assistant rather than use ours.

Customer adoption told us plenty, but the bigger shift was in how much better we came to understand what our customers wanted from Spendesk.

We were strategic about the rollout, starting with a read-only first version.

Read-only access lets us reach customers quickly and watch how they use the connector.

We leaned on our existing API, which already used OAuth for authentication, so we didn't have to build credential handling from scratch.

Permissions and scoping came out of the box too, so we knew exactly which users could reach which data.

No user got unrestricted access to everything.

Spendesk held back the tools that write to a customer's account until customers had reason to trust the new product.

AI agents couldn't replace an engineering team

A Stack Overflow survey found 84% of developers use AI tools, but only 3% have high trust in the accuracy of the output.

Aleksei Belezeko, CTO at Vivid Money, told us his own turning point came when Opus 4.5 produced code good enough for production. Do your engineers trust AI output more than they did a year ago?

The short answer is yes, though the nuance matters more than the yes.

Step changes in the underlying models raised the ceiling, and the next generation will raise it again.

The real shift is that trust moved from the models themselves into the process around them.

We now trust AI output where tests already exist to verify it, and far less where they don't.

Giving a model the right context also raises how much we trust what it produces.

We learned that the hard way when we asked an agent to build software without giving it context on how our backend worked.

The agent had never seen that code, so it guessed, and the result wasn't good enough to ship.

We place more trust in AI now, but only where we're confident the model has the right context and we can verify what it produces.

We place our trust in the output, and a person still reviews that output every time.

Google's DORA team calls this pattern the J-curve, where AI adoption initially slows a team before it speeds it up.

Some of Spendesk's AI rollout happened before you arrived, and you're the one leading it to results now. In hindsight, what would you have done differently in that first rollout?

The product side of the rollout is hard to answer for, since much of it happened before I joined.

What I can talk about is the internal pilots we've run to accelerate our development process, and those pilots surfaced real gaps in our readiness for AI adoption.

Our initial hypothesis was that a single engineer with a set of agents could match the productivity of a full team.

We ran 2 pilots to test that, one closer to a customer-facing feature and one on a deeper, more complex backend problem.

The deeper, more complex pilot succeeded more, and it's also where we learned our biggest lessons.

A front-end engineer ran the customer-facing pilot, and the agent lacked a real understanding of our backend code.

The engineer knew what good looked like on the front end but not on the back end, so we spent a fair amount of time correcting the course.

Our second lesson was that 1 engineer with a set of agents wasn't enough on its own, at least not for what Spendesk needs.

Our third lesson came as a surprise, even in hindsight.

The engineer running the pilot ended every day exhausted, because so many agents were running in parallel that he kept switching context between them.

The surprise gave us real insight into how to scope future pilots, and into which teams and squads we should build around going forward.

Mixing skills across the team is something we'll carry into the next pilot, now that we've seen where a single engineer running many agents can break down.

Spendesk gave AI a calculator

AI Connect runs on MCP, the open standard Anthropic created before handing it to the Linux Foundation.

Being open hasn't made MCP stable, with 5 spec versions in under 2 years and 1 feature added, then removed, within 3 months.

Where did building on an open AI standard cost your team the most time?

Most of this work started when I first joined, so I'm looking back at the early stages.

MCP itself was stable enough for us to start building on.

The real challenge was that the expertise didn't exist anywhere, inside Spendesk or in the hiring market, because the standard was brand new.

Nobody had built something like this at scale before, so we started from first principles and learned through trial and error, including how to build a compliant, secure MCP server.

There was no room for mistakes, so progress was slow at first, and I had to be patient while we worked through the details.

Our first version was also slow to perform well with agents.

We'd followed the common path of wrapping an existing API one-to-one, instead of matching the tool to what an agent needed.

We refactored the MCP layer so each answer needed a small number of calls, and a lot of prompt engineering came out of that work.

Arithmetic mattered most because Spendesk is a spend platform, and CFOs and finance teams expect the numbers to add up.

An LLM can attempt arithmetic, but it doesn't always get it right. We built a dedicated maths endpoint into our MCP server, and told the model to call it instead of computing sums itself.

Customers can trust the numbers they see because the tables add up every time.

The pattern became one of our clearest lessons from building on MCP.

Every engineering team right now is full of trials, pilots, and free credits, and, eventually, someone has to decide what stays and what goes.

What does an AI tool have to show you before it becomes permanent at Spendesk?

Week 1 enthusiasm doesn't tell me much, and it never has.

I judge a tool by whether people are still using it in week 12, not by how excited they were in week 1.

On top of that, I look for whether the tool moves something we already track, rather than a flattering number invented to justify it.

Nothing about our quality or security bar can degrade, and those bars have to hold regardless of how useful the tool feels.

The last thing I look for is whether the tool gets baked into how we work, past the initial novelty and the enthusiasm of whoever found it.

Engineers pitch new tools most weeks, along with prompts and skills from a shared repository we keep tweaking, so it matters where we place our bets.

The same bar determines whether a tool earns a permanent place, rather than fading out after the first few weeks.

Three risk levels guide AI tool reviews

Spendesk holds its own payment institution license from the ACPR, so you're regulated yourselves.

Marilyn McDonald, Thredd's CTO, described a common bottleneck on this show, saying compliance gets pulled in at the end, once something is already finished.

Does an AI tool at Spendesk get a security review before your engineers can use it?

Yes, bluntly.

Spendesk Financial Services is a payment institution supervised by the ACPR, and Spendesk itself is ISO-certified, so GDPR, PSD2, DORA, and the EU AI Act all apply to us.

Every tool, AI or not, gets a security review, and that's a regulatory obligation rather than a preference.

Not every tool carries the same risk, though, so we built 3 tiers, critical, important, and flexible, as pre-approved categories for a tool or a change.

We put any tool that touches payments or payment rails in the critical tier, where it gets the highest level of control and scrutiny.

We put a change like recolouring a UI button in the flexible tier, where the team can move fast.

The tiers tell product and engineering exactly where they can move quickly and where they can't.

Bringing security in at the last minute is a recipe for failure and frustration, so we've folded security engineering into the design process from the very start.

Engineers aren't surprised by a late review because security is in the room from day one, and we know we're building compliant software from the beginning.

A JetBrains survey of over 15,000 developers found real churn over 7 months, with Copilot's share falling from 29% to 21%, Cursor's from 18% to 12%, and Codex's rising from 3% to 16%.

Dropping a tool no longer looks like a failure.

Was there an AI tool your engineering team dropped after a trial?

Yes, we've dropped a tool.

Dropping an AI tool at Spendesk rarely comes down to missing features.

The problem is usually sprawl.

Individual engineers picking their own tools works fine for 2 people, but it stops working once you have over 100 engineers inside a regulated organization.

The tools that lost out failed the scorecard on governability rather than capability.

We ran trials on two of the leading coding assistant providers, and consolidated onto one in the end, on fine margins.

Consolidating onto 1 tool comes with a real cost because people become attached to the tool they already know, and asking them to switch to a shared, golden path creates friction.

Most people here now understand the value that the golden path brings, even with that friction.

We also exclude tools on governance grounds alone, because we can't send our data to just anyone.

We watch things like data retention and how closely our data trains a vendor's models.

When we can't guarantee those things, the answer is often that the tool is great, but Spendesk still can't use it.

AI reviews the code that AI writes

A study from METR found that experienced developers were 19% slower with AI, even though they believed they were 20% faster.

Ehab Qadah from Lean summed up the gap, saying code got cheaper without the validation layer around it getting any smaller.

By how much has delivery at Spendesk sped up since your engineers started using AI?

I can't give you a clean percentage, but I can tell you what we do measure.

AI-assisted pull requests made up 16% of all our PRs in the months before I joined, and that figure is now 70%.

Tool adoption has gone from around 61% to nearly 100%, so we're using the AI tools available to us across the board.

Coding got faster, but what reaches the customer hasn't sped up to match yet.

Coding itself was never really the constraint on delivery, at least in my view.

Forrester's 2024 research, before the current wave of AI adoption, put coding at only about 24% of an engineer's working week.

The rest went to code reviews, meetings, and design workshops.

Speeding up the 24% of an engineer's week spent coding by 10 times still saves only about 21% of their overall time.

Teams hit the bottleneck on either side of the coding itself, in product discovery and planning on one side and review and deployment on the other.

Discovery, planning, review, and deployment still work roughly the way they did 2 or 3 years ago, which is exactly where we're focused right now.

Aleksei Belezeko has also told me he doesn't force a tool preference on his engineers, as long as the output matches what an AI-assisted developer delivers.

Some companies are far less relaxed about it, and a few have started building AI usage into performance reviews.

Can an engineer at Spendesk refuse to use an AI tool?

Yes, they can.

Using AI tools isn't mandated at Spendesk.

I care most about the quality bar and what an engineer gets done.

The craft of engineering itself has shifted, though, and we've had to accept that.

Some engineers were worried early on about what AI meant for their jobs, and those concerns were legitimate.

We worked through those concerns together rather than arguing them away.

The questions I hear now are sharper, like what an engineer gains from using AI, what happens when the output is wrong, and who owns the fix.

I don't expect anyone's career to be hampered for choosing not to use AI.

I would expect an engineering manager to ask why, though, if someone consistently chooses not to.

Ravneet Shah, Allica Bank's CTO, told us access to data worries her most, especially after hearing about AI agents deleting entire databases elsewhere.

Your own connector remains read-only for now, so I'd guess you're thinking about this in a similar way.

Which part of software development at Spendesk will you keep away from AI?

I agree with the premise, though AI will touch most parts of software development.

The difference lies in the level of scrutiny people apply, and the risk tiers I mentioned earlier help determine it.

The payment path, money movement, cards, ledgers, and settlement get full controls and full compliance, with no shortcuts and a person always in the loop.

Final accountability for any production change stays with a person, and we never delegate it to AI.

No agent authors code alone at Spendesk.

There's an uncomfortable truth underneath all of this that people tend to forget.

The human code reviewer was never really this quality gate that people think it was.

If code review really were that quality gate, Spendesk would never have incidents, because every bug would get caught before it reached production.

We know that isn't the case, so code review mostly functions as a confidence mechanism rather than a real safety net.

Code is also being generated at a volume that's outpacing the team's ability to review it accurately.

Augmenting the review process with AI is where we need to focus next, and it's a big part of our plan going into next year.

The tools leading this market today were barely in use in January, and the ones everyone talked about a year ago are already losing ground.

12 months is a long time in this space. Where will your engineering team be by then?

The constraint has already moved, so adding more coding tools won't make much of an impact from here.

Adoption is already high across coding, and I want to keep building on that rather than adding new tools to the pile.

I see the biggest opportunity now on the reading side of coding, meaning review, so that's where I'm putting the focus for the rest of this year and into 2027.

I expect us to build guardrails and give agents the right level of context to work with.

Building APIs and knowledge bases that supply agents with the context they need lets us hand them harder, more sophisticated problems over time.

Going into 2027, we'll keep running pilots, learning from the ones we've already run, and iterating from there.

Few playbooks exist for accelerated AI development inside a highly regulated environment, so we're learning as we go, the same way a lot of other companies are right now.

3 moves worth making

Build a deterministic tool for transaction maths

An LLM asked to total a column of transactions might return a sum that only looks correct, and that mistake might reach a customer unnoticed.

Spendesk built a maths endpoint into its MCP server and prompted the model to call that endpoint before showing any sum or average to a customer.

An engineering team building AI features can send every total a customer relies on, such as a monthly spend figure, to the same kind of endpoint, so the model never calculates that total itself.

Make engineers answerable for the AI-written code they merge

When an AI agent writes a change, the engineer who merges that change might not feel responsible for any defects hiding inside of it.

At Spendesk, Alan Wright set a rule that an engineer reviews and approves every change an AI agent writes before that change reaches production.

An engineering leader can make ownership of merged code part of each engineer's performance review, so approving AI-written code carries the same weight as writing it.

Give AI read access before write access

Letting an AI assistant read customer data usually carries less risk than letting it change that data, because a wrong answer is easier to ignore than a wrong record.

Spendesk chose the lower-risk option and launched AI Connect read-only, linking assistants such as Claude to spend data through its existing API.

A product team can add write access features such as payments and approvals in a second release once customers have seen the assistant answer correctly.

Authors

  • Michał Szehidewicz

    I'm a product delivery consultant with 12 years of experience across enterprise and scaleup businesses in financial services, SaaS, and software development. I help CTOs, CPOs, and CEOs unblock their software delivery, and I have worked in C-level management and B2B sales myself. Today, my focus is GTM engineering. I also host Fintech Recoded, an interview series with senior product and technology leaders at fintech companies. At tsh.io/fintech-recoded, you can read how these leaders make hard calls under real market constraints and uncertainty. Outside work, I'm a father of 2 with a newborn daughter and 1 unruly border collie. Everything else I love is in the mountains except the NBA.

  • Michael Sols

    Content fixer: B2B Copywriter, Marketer, Strategist. Often questioning reality — to find facts that make business decisions good, of course. Connected with technology since training his family in the basics of Windows 98. Authored brand stories out and about the software market that were published by WIRED UK, Silicon Republic, The Sun, and Vanity Fair. Also, you deserve a raise.

Imagine fintech development but 30-40% faster

free consultation

Is your core team overbooked while the backlog piles up?

Fintechs like xpate or Obligate delivered their side projects 30-40% faster with our fintech engineers equipped in an industry-designed AI framework.

How it works

Explore other interviews