Build vs Buy Is No Longer One Decision
For most of the past twenty years, build versus buy had a sensible default: buy, unless your process was unusual enough to justify the cost of building. Software vendors could spread their engineering costs across hundreds of customers, while a five-person IT team could not; buying also transferred much of the maintenance burden to someone with more people, better tooling, and a larger security budget.
AI has not reversed that logic, but it has made the old wisdom less useful. Building a working prototype is now much cheaper, while the product you buy is less fixed than it used to be because its model, hosting arrangements, sub-processors, and behaviour can all change during the life of the contract.
Microsoft 365 Copilot is a useful example. On 9 July, OpenAI announced that GPT-5.6 was now Copilot’s preferred model across Word, Excel, PowerPoint, Chat, and Cowork, even though Microsoft 365 Copilot already used model routing and, in some features, offered models from more than one provider1. Microsoft has also introduced Anthropic as a sub-processor, with tenant controls and data-processing implications that vary by region.2
The point is not that every Copilot customer experienced the same silent model swap on the same day; it is that an approval now covers a moving stack. If a risk committee approved Copilot in March, what exactly did it approve: the product, the model family, the processing arrangement, the use case, or the controls around it?
That is why the old question is becoming harder. “Build or buy?” assumes there is one system and one sourcing decision, whereas an AI-enabled service has several layers, each of which can have a different answer.
Decide what to own, layer by layer
For most firms, the useful distinction is between three layers.
The model and infrastructure layer provides the underlying AI capability and the computing power needed to run it. Few mid-sized (or even large) financial services firms should build this themselves because the major model and cloud providers operate at a scale that is difficult to reproduce locally.
The workflow and integration layer connects that capability to a real business process, deciding where information comes from, what the AI is allowed to do, which systems it can access, and when a person must intervene. This is where selective internal development can make a lot of sense.
The data and control layer covers permissions, testing, logs, human oversight, fallback procedures, and accountability. A firm can use vendors to support this layer, but it cannot outsource ownership of it.
This produces a more practical starting position: buy the model and infrastructure where appropriate, build only the workflow that makes the process distinctive, and retain ownership of the data and controls.
Building is cheaper to start
There is good evidence that AI coding tools can increase software output. Three field experiments involving 4,867 developers at Microsoft, Accenture, and a Fortune 100 company found that developers with access to an AI coding assistant completed about 26% more tasks, with less experienced developers seeing larger gains.3
That result matters, but it does not support a blanket assumption that every internal build will be 26% faster.4
METR reached a different result when it studied 16 experienced open-source developers working in repositories they knew well: with access to early-2025 AI tools, they took 19% longer to complete the selected tasks, even though they believed the tools had made them faster.
Then the result moved. When METR revisited the experiment in early 2026, the returning developers now showed a speedup of around 18%, while newly recruited developers landed near zero. METR presents the follow-up cautiously (the numbers are small, and the returning developers chose to come back), but the direction had flipped within a year5. DORA’s annual research tells the same story at organisational scale: its 2024 report found AI adoption reducing delivery throughput; its 2025 report found throughput rising as teams and tools adapted, with delivery instability persisting in both.6
The lesson is not that one study is right and another wrong. AI changes the economics of development, but the size and even the direction of the gains depend on the work, the team, and the year you measure it. Any business case built on last year’s benchmark deserves suspicion in both directions.
It is also easy to measure the wrong thing. AI reduces the time required to write a feature or assemble a prototype, but it does far less for defining requirements, connecting fragmented data, testing exceptional cases, documenting the system, monitoring it after release, and maintaining it when the original developer leaves.
The published data on AI-generated code makes the same point less politely. Veracode tested more than a hundred models on security-sensitive coding tasks and found 45% of the generated code contained vulnerabilities from the OWASP Top 10, the standard catalogue of ways software gets broken into; its spring 2026 re-test found security pass rates unchanged after two years of new model releases7. GitClear’s analysis of hundreds of millions of lines of real-world code shows code duplication reaching its highest level on record in early 2026, while refactoring, the unglamorous work of keeping a codebase healthy, continues to fall away8. Code has become cheaper to produce and, on current evidence, more expensive to look after.
For a regulated firm, those downstream activities are not finishing touches; they are much of the cost of owning software. Lower development costs are therefore a good reason to run smaller, cheaper experiments, but not a reason to treat a prototype as a production system.
Buying has become more dynamic
Buying still transfers a large amount of work to the vendor, which patches infrastructure, manages specialist engineering teams, operates security processes, and absorbs much of the cost of evaluating new models. What it no longer gives you is a static technical dependency.
Microsoft’s Anthropic arrangements illustrate the difference. Anthropic became a Microsoft sub-processor for relevant services in January 2026, with Microsoft enabling it by default for many commercial tenants while EU, EFTA, and UK tenants were initially set to opt in9. When Anthropic models are enabled, some processing falls outside Microsoft’s EU Data Boundary, so administrators need to understand both the tenant setting and the processing distinction.
This is not evidence that Microsoft bypassed its contracts; the change was documented through product terms, administrator messages, and technical guidance. It does show that a vendor’s standard change process and a regulated firm’s internal change process do not always meet in the middle.
The more extreme example came in June, when the US government imposed export controls on Anthropic’s Fable 5 and Mythos 5 models. The directive required restricting access by foreign nationals with immediate effect, and with no reliable way to verify a user’s nationality in real time, Anthropic suspended both models for all customers. Access was eventually restored in stages over the following weeks.10
Although that episode was unusual, it turned geopolitical risk into a service outage and left firms relying on those models needing a fallback, regardless of what their original vendor assessment said about uptime.
There is a concentration issue too. Three providers, Anthropic, OpenAI, and Google, account for roughly 88% of enterprise spending on AI model APIs, and many of the AI features across a typical vendor list are wrappers around the same engines11. The Financial Stability Board has identified third-party dependencies and AI service-provider concentration as vulnerabilities that could increase systemic risk12, and the point is practical as well as systemic: buying different applications does not necessarily diversify the underlying dependency if several vendors rely on the same model or cloud provider.
Buying therefore moves much of the risk out of your codebase, but it does not make the risk disappear; instead, it moves into contracts, sub-processor chains, data locations, change notices, and exit plans.
The JFSC focuses on impact, but the policies ask different questions
The JFSC’s July 2026 Guidance on the use of AI in Jersey’s financial services sector gives firms a useful way to approach this13. It creates no new requirements, but explains how existing obligations apply and makes the central principle clear: firms remain accountable for outcomes when they use AI, just as they do when they use any other technology.
Its materiality ladder distinguishes low-impact productivity tools, which need only light-touch controls, from higher-impact applications such as customer onboarding, lending decisions, or investment advice. Governance should increase where a use case affects regulated activity, customer outcomes, confidential data, or regulatory submissions.
The materiality ladder and the JFSC Outsourcing Policy answer different questions. The AI guidance calibrates governance to the impact of the use case. The Outsourcing Policy applies where a service provider performs an activity that the firm would otherwise undertake and the provider’s failure or inadequate performance would materially prevent, disrupt, or affect the firm’s continuing compliance in its regulated activity14. A bought AI service can meet that test, including where supporting IT is critical to regulatory compliance, but buying technology does not by itself make the arrangement Outsourcing.
Where the Policy applies, the governing body remains responsible and accountable. The firm must address due diligence, the contractual arrangements, ongoing monitoring, and contingency planning. Ordinarily it must notify the JFSC before appointing the provider and wait for No Objection before the activity starts, but the Policy contains specific reliefs. Qualifying Cloud, Data Centre, Cyber Security, and Digital ID Services require notification but not No Objection. Standardised public cloud-based email, such as Microsoft 365 email, is excluded altogether. Qualifying cloud sub-outsourcing does not require a separate filing at the sub-contractor layer, although the primary outsourcing arrangement remains subject to the Policy.
A foundation-model update is therefore not automatically notifiable. For existing Outsourced Activity that has already been notified, the firm must decide whether the development is a material change within the Policy. A Material Change to Outsourcing Notification will ordinarily require no further No Objection. New Outsourced Activity requires a new notification and further No Objection, subject to the Policy’s specific exceptions. Changes to providers, model supply chains, processing locations, service performance, controls, and contingency arrangements should in any event feed into the firm’s internal risk assessment and contractual oversight.
An internal build is not itself Outsourcing. It may, however, contain a separate outsourcing arrangement where a cloud platform, external API, or other provider performs an activity that satisfies the Policy test. Each external component should be assessed on its function and consequences rather than on whether the overall solution is labelled “built” or “bought”.
Both routes require ownership, but the regulatory treatment depends on what the external provider does and what happens if it fails.
What this looks like in practice
Consider a scenario most Jersey firms will recognise: using generative AI to help prepare a regulatory report.
A firm might buy an approved generative AI service rather than developing a model capable of organising source material and drafting narrative sections from scratch.
It might then build the workflow that connects the service to approved internal data and the firm’s reporting templates. That workflow could assemble the relevant source material, preserve links back to authoritative figures, flag missing commentary, and route the draft through the existing review process.
The firm must still own the controls. It decides which data can be processed, prevents the AI from inventing figures or unsupported claims, requires staff to verify the draft against approved source data, records reviewer sign-off, and maintains a manual fallback if the service is unavailable.
This is neither a pure build nor a pure buy: the model is bought, the reporting workflow is partly built, and accountability remains with the firm. For many Jersey firms, that combination will be the normal pattern.
Three questions for each use case
Once the decision is separated into layers, three questions do most of the work.
Where does our differentiation sit?
Meeting notes, generic document summaries, and drafting support are commodity capabilities, so firms should use approved features in platforms they already own or buy a suitable service. The market has largely reached the same conclusion: by 2025, 76% of enterprise AI use cases were bought rather than built, up from 53% a year earlier15. A build becomes more defensible when the value lies in a process specific to the firm: its decision logic, service model, data constraints, or way of handling exceptions.
Even then, the distinctive part is usually the workflow rather than the foundation model. MIT’s analysis of enterprise deployments points the same way: bought tools and vendor partnerships reached deployment roughly twice as often as purely internal builds, which is a caution against reading cheap prototypes as easy wins.16
There is also a category of work in Jersey where the question answers itself, because there is no buy option. Vendors build for the UK and EU mass market, and a jurisdiction of a hundred thousand people rarely earns a place on their roadmap, so the software a Jersey firm can buy tends to stop just short of the island’s specific requirements: the local regulatory return, the structure the product team has never heard of, the process shaped by Jersey law rather than FCA rules. Historically, firms absorbed these gaps with manual work: spreadsheets, rekeying between systems, and checklists that live in someone’s head.
That was a rational response when custom software cost six figures. It is much less obviously rational now. A narrow internal tool that automates one Jersey-specific process was never going to justify a vendor’s development budget, but it may now justify a few weeks of a firm’s own. This is where the falling cost of building matters most, precisely because no amount of budget could previously buy the thing at all. The same ownership tests still apply; the difference is that the alternative being displaced is not a vendor product but manual effort and its error rate.
Can we operate it after the demonstration?
Check the data, integrations, and people before comparing licence costs with development estimates. MIT’s much-quoted study of enterprise AI pilots attributed most failures to a “learning gap” (tools that don’t integrate with workflows or improve from feedback) rather than model quality17, and, when RegTech buyers are surveyed about why adoption stalls, they blame legacy systems and data, not the AI18. The practical questions are less glamorous, but more important: who will test the system and approve changes, whether the team can diagnose a bad output, whether there is enough logging to reconstruct what happened, and who will maintain it in two years.
A bought tool connected to poor data can fail as thoroughly as an internal build; it just fails with an invoice attached.
What happens when it changes or stops?
For a bought component, ask how model and sub-processor changes are communicated, what contractual rights the vendor retains, where data is processed, and how the service can be replaced.
For a built component, ask who maintains the code, which external services it still relies on, how dependencies are updated, and what happens when the person who understands it leaves.
For both, decide what the manual fallback is and which changes would trigger a fresh risk or materiality assessment.
The resulting decision rule is straightforward: buy commodity capability, build the layer where the firm’s process is genuinely different, and proceed with neither until the data, the accountable owner, and the fallback are clear.
A capability, not a procurement event
Build versus buy is no longer a decision that stays settled for the life of a system because model capability, vendor terms, regulatory guidance, and internal engineering capacity all move too quickly.
Firms do not need to re-procure every AI service every six months, but they do need a repeatable way to review material use cases, map the providers beneath them, and identify what has changed since approval.
That is the good news in the harder question: AI gives smaller firms more credible options to shape software around their own processes, while forcing them to understand which parts of those processes are worth owning.
The question is no longer simply whether to build or buy, but which layer to own and whether the firm is prepared to keep owning it.
- OpenAI, GPT-5.6 is now the preferred model in Microsoft 365 Copilot, 9 July 2026
- Microsoft, Anthropic as a sub-processor for Microsoft Online Services (checked 23 July 2026; page updated 22 July 2026). Anthropic became an active sub-processor from January 2026, default-enabled for most commercial tenants; EU, EFTA and UK tenants required explicit opt-in as Anthropic models are excluded from the EU Data Boundary. Note: Claude’s separate general availability on Azure infrastructure via Microsoft Foundry (June 2026) does not currently change this: no European data zone exists for Claude in Foundry either, with EU-native support a stated target for later in 2026. Re-verify this footnote immediately before publication; it is the fastest-moving fact in the piece.
- Cui et al., The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers: randomised trials across Microsoft, Accenture and a Fortune 100 firm, n=4,867.
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, July 2025: 16 experienced developers working in repositories they knew well. The early-2026 follow-up is METR, We are Changing our Developer Productivity Experiment Design, February 2026: returning participants estimated at ~18% speedup (wide confidence interval), new participants ~−4% (indistinguishable from zero); METR notes selection effects mean these are likely lower bounds, and characterises the evidence as weak.
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, July 2025: 16 experienced developers working in repositories they knew well. The early-2026 follow-up is METR, We are Changing our Developer Productivity Experiment Design, February 2026: returning participants estimated at ~18% speedup (wide confidence interval), new participants ~−4% (indistinguishable from zero); METR notes selection effects mean these are likely lower bounds, and characterises the evidence as weak.
- DORA, 2025 State of AI-assisted Software Development report (survey of ~5,000 technology professionals). The 2024 edition found AI adoption associated with reduced delivery throughput and stability; the 2025 edition found throughput improving while instability persisted, framing AI as an “amplifier” of existing organisational strengths and weaknesses.
- Veracode, 2025 GenAI Code Security Report: 45% of AI-generated code samples contained OWASP Top 10 vulnerabilities, across 100+ models and 80+ coding tasks. The re-test is Veracode, Spring 2026 GenAI Code Security Update, using the same testing framework; security pass rates flat at ~55% despite two years of model releases.
- GitClear, AI Copilot Code Quality: 2025 research (211 million changed lines; cloned code rose from 8.3% to 12.3% of changes while refactoring fell from roughly 25% in 2021 to under 10% in 2024) and GitClear, The Maintainability Gap: 2026 research (copy/paste up to 15.7% of changed lines in the first half of 2026, refactored “moved” code down to 3.8%; block duplication the highest on record).
- Microsoft, Anthropic as a sub-processor for Microsoft Online Services (checked 23 July 2026; page updated 22 July 2026). Anthropic became an active sub-processor from January 2026, default-enabled for most commercial tenants; EU, EFTA and UK tenants required explicit opt-in as Anthropic models are excluded from the EU Data Boundary. Note: Claude’s separate general availability on Azure infrastructure via Microsoft Foundry (June 2026) does not currently change this: no European data zone exists for Claude in Foundry either, with EU-native support a stated target for later in 2026. Re-verify this footnote immediately before publication; it is the fastest-moving fact in the piece.
- Anthropic, Statement on the directive to suspend access to Fable 5 and Mythos 5 (12 June 2026) and Redeploying Claude Fable 5 (30 June 2026). Timeline: suspension 12 June; Mythos 5 restored to approved US organisations 26 June; the Department of Commerce lifted the controls 30 June; Fable 5 restored globally from 1 July, with Mythos 5’s wider rollout following.
- Menlo Ventures, 2025: The State of Generative AI in the Enterprise, December 2025: 76% of enterprise AI use cases purchased rather than built (53% in 2024); Anthropic, OpenAI and Google together account for roughly 88% of enterprise model spend.
- Financial Stability Board, The Financial Stability Implications of Artificial Intelligence, November 2024.
- JFSC, Guidance on the use of AI in Jersey’s financial services sector (full guidance PDF), issued July 2026. “Materiality ladder” is the JFSC’s own term (Appendix 1 of the guidance).
- JFSC, Outsourcing Policy, issued 1 March 2017 and last revised 1 December 2023, effective 1 January 2024. See in particular paragraphs 2.1.1, 2.2.1, 2.2.3.12, 3.1, 3.2, 3.3, 3.5, 3.6.5, 3.6.6, and 6.4.
- Menlo Ventures, 2025: The State of Generative AI in the Enterprise, December 2025: 76% of enterprise AI use cases purchased rather than built (53% in 2024); Anthropic, OpenAI and Google together account for roughly 88% of enterprise model spend.
- MIT NANDA, The GenAI Divide: State of AI in Business 2025 (MIT Media Lab, July 2025; widely cited mirror PDF; the report has no stable official URL). Pilot failures are attributed to a “learning gap” (tools that don’t retain feedback, adapt to context, or improve over time) rather than model quality; externally sourced tools and partnerships reached deployment ~67% of the time versus ~33% for internal builds.
- MIT NANDA, The GenAI Divide: State of AI in Business 2025 (MIT Media Lab, July 2025; widely cited mirror PDF; the report has no stable official URL). Pilot failures are attributed to a “learning gap” (tools that don’t retain feedback, adapt to context, or improve over time) rather than model quality; externally sourced tools and partnerships reached deployment ~67% of the time versus ~33% for internal builds.
- FinTech Global, What are the biggest barriers to third-party RegTech adoption?, 20 July 2026, reporting the Global State of RegTech 2026 findings.
Before AI becomes a regulatory question, assess your AI exposure
Confidential. Designed for regulated firms. No obligation.
