
Your core system vendor has just added AI to the underwriting workbench. It reads submissions, drafts a summary, and flags risks outside appetite. It's already in the contract, already inside your security perimeter, and nobody has to learn a second interface.
The competing option is a separate platform that connects to that same core over an API, reads the same submissions, and writes back into the same workbench fields.
On a demo, these look like the same product. The difference shows up eighteen months later, when the model underneath one of them is retired and you find out how quickly you can move.
That's the real axis. Where the AI sits relative to your core determines whose release schedule you're on, whose contract carries your audit rights, and what you keep when you replace the core. Which tool extracts better today is the easier question, and the one that ages fastest.
Whether to buy a broad horizontal AI tool or an insurance-specific one runs on a different axis, covered in our 2026 insurance AI buying guide. That page is about vendor breadth; this one is about deployment locus.
Embedded underwriting AI ships as part of the core policy administration system or the underwriting workbench. The vendor that sells you the workbench also supplies the models, the prompts, the extraction schemas, and the release schedule. You get it through your existing contract, and you get changes to it when that vendor ships them.
Standalone underwriting AI runs as its own layer with its own contract, connected to the core over APIs. It reads from and writes to the same systems, so the underwriter still works in the workbench, but the AI's roadmap, model choices, and configuration belong to a separate vendor.
Both can look identical to the person doing the underwriting. The submission arrives, a structured summary appears, the underwriter reviews it. What differs is everything behind that screen.
Worth separating this from a question it gets confused with. Whether to buy a general-purpose AI assistant or an insurance-specific platform is about breadth — how much of the vendor's product is built for insurance. Where the AI sits relative to your core is about locus. You can have insurance-specific AI in either position, and you can have general-purpose AI in either position too.
This decision is urgent rather than academic because of a mismatch in speed. The models underneath any underwriting AI turn over far faster than core insurance systems do.
The numbers are public. Anthropic's deprecation policy commits to "at least 60 days' notice before model retirement for publicly released models," and its retirement table records Claude Opus 4.1 as retired on 5 August 2026, exactly twelve months after Anthropic announced it. Microsoft's Foundry policy sets out a full lifecycle: generally available models get 18 months, deprecation begins at 12 months when new customers lose access, and models from several providers "follow a 12-month lifecycle instead of the standard 18-month lifecycle." For Standard deployments, "Microsoft manages automatic upgrades when a model version is retired." You can pin a version instead, and the distinction matters for a validated workflow: one setting holds the version until retirement, another holds it through retirement, at which point the deployment stops serving requests.
Capability moves at a similar pace. Stanford's AI Index 2026 puts it plainly: "Evaluations intended to be challenging for years are saturated in months, compressing the window in which benchmarks remain useful for tracking progress." Frontier models gained 30 percentage points on Humanity's Last Exam in a single year.
So the underwriting AI you sign for in 2026 is running on a component with a shelf life measured in months, a core system that changes far less often. Whoever controls the release process for the AI controls how fast you can respond to that. This is the question to press hardest in a demo, and the one least likely to come up on its own.
We could not find a neutral, citable figure for how often insurance core platforms ship releases, and we'd rather say so than borrow a number from a vendor or an implementation partner. What is well sourced is the pressure the architecture creates: in Capgemini's World Property and Casualty Insurance Report 2026, 81% of insurers named legacy systems and IT architecture constraints as a barrier to scaling AI, the highest of any barrier they measured.

The instinct that embedded AI is the safer regulatory answer doesn't survive contact with the actual text. The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers says an insurer's AI program "should address the AI Systems used with respect to regulated insurance practices whether developed by the Insurer or a third-party vendor." Adopted in 24 states and the District of Columbia as of the NAIC's Spring 2026 National Meeting, it puts the obligation on you regardless of who built the model.
What does change with locus is who you can audit. The bulletin's third-party section says a program should address AI developed by a third party, which "may include, as appropriate," and "where appropriate and available," contract terms that "provide audit rights and/or entitle the Insurer to receive audit reports by qualified auditing entities" and "require the third party to cooperate with the Insurer with regard to regulatory inquiries and investigations related to the Insurer's use of the third-party's product or services." Where an examination concerns third-party models, the bulletin says an insurer "should also expect the Department to request" its "contracts with third-party AI System, model, or data vendors, including terms relating to representations, warranties, data security and privacy, data sourcing, intellectual property rights, confidentiality and disclosures, and/or cooperation with regulators."
When the AI is embedded, those terms have to reach the model's maker through your core vendor's contract. When it's standalone, you negotiate them directly. Neither is automatically better, and the question to ask in either case is the same: can you produce, on demand, a contract that gives you audit rights over the thing that made the decision?
European carriers face a sharper version. EIOPA's 2025 opinion on AI governance states that "undertakings are ultimately responsible for the AI systems that they use, regardless of whether the AI systems are developed in-house or in collaboration with third party service providers," and adds that where third-party intellectual property makes governance difficult, firms should compensate with contract clauses, external audits, and due-diligence testing. Under DORA, EU financial entities also need registers of ICT providers and contractual rights of access and audit, which have to be traceable through subcontractors.
Embedded AI ships on the core system's release train. That's genuinely an advantage when you're happy with the direction: the integration work is done, the testing is done, and the upgrade arrives as part of a release you were installing anyway.
It becomes a constraint the moment your priorities and the core vendor's diverge. If you need a new extraction behavior for a line the vendor doesn't prioritize, you wait for their roadmap. A standalone layer moves on its own schedule, which cuts both ways: you can change it whenever you like, and you also own the regression testing every time you do.
This follows directly from the clock-speed problem, and it's the dimension carriers most often skip in evaluation.
Ask any prospective vendor, embedded or standalone, three questions. Which model or models sit underneath the product today? What happens on your side when that model is retired, and how much notice do you get? Can you pin a model version for a validated workflow, and if so, for how long?
The last one matters more than it sounds. Once you've validated an underwriting workflow for a regulated use, a silent model swap underneath it is a governance event. Microsoft's own documentation draws exactly this line between deployment types that auto-upgrade and provisioned ones that don't. A vendor who can't answer clearly is telling you they haven't thought about model governance, whichever side of the line they sit on.
This is where the two options separate most sharply, and it's the dimension that's hardest to reverse.
Over two or three years of running underwriting AI, you accumulate something more valuable than the software: appetite rules encoded as checks, extraction schemas tuned to the forms your particular brokers send, mappings from carrier guidelines to specific fields, referral thresholds, and a record of every correction your underwriters made to the model's output. That configuration is the asset.
If the AI is embedded in the core and you replace the core, that asset leaves with it. You re-procure, re-configure, and re-validate on the new platform. If the AI runs alongside, replacing the core is an integration change: the connectors are rebuilt, the configuration stays.
Capgemini's data offers a related observation. Describing the roughly 10% of P&C insurers it identifies as "intelligence trailblazers," the report notes they "operate on a cloud infrastructure that lets AI run alongside legacy systems, avoiding costly core replacements." That same group shows 21% higher revenue growth and a 51% greater increase in share price than mainstream insurers, though the report measures those across 2021 to 2024 and attributes the outperformance to strategic clarity rather than to architecture. It also lists core process redesign as unfinished business even for these firms. Read it as a description of how the leaders have sequenced their work, not as proof that the sequencing caused the returns.
For underwriting summaries specifically, the evidence points to domain fit mattering more than raw model quality, and it also cautions against over-claiming.
Vals.ai's MortgageTax benchmark, reported in Stanford's AI Index, tests extraction of structured fields from real mortgage tax certificates across a dataset of 1,258 documents. The best model reached 69.4%, and the lowest of the top fifteen scored 65.9%. The Index's own reading is that "the overall accuracy level does not reach 70%, which suggests that models are not yet entirely or reliably able to extract and compute financial information from document images."
Two things follow. Document extraction in regulated finance isn't a solved problem, so treat any vendor claiming near-perfect accuracy on your forms as making a claim about their pipeline rather than about the underlying model. And because the top fifteen models cluster within a few percentage points, the differentiator is what surrounds the model: the schemas, the guideline mappings, the validation rules, and the review loop. That's a domain-knowledge problem, and it's the argument for insurance-specific tooling. Our buying guide on horizontal versus dedicated platforms works through that trade-off in full.
Neither wins outright, and the published benchmarks cut against the assumption that newer always means better.
On OmniDocBench, a CVPR 2025 document-parsing benchmark, pipeline OCR tools performed well "for commonly used data, such as academic papers and financial reports," while the authors found that "for more specialized data, such as slides and handwritten notes, general VLMs demonstrate stronger generalization." Two results from the same paper deserve more attention than they usually get. On table recognition, OCR-based models "demonstrate superior overall performance," with general vision models still lagging "behind specialized solutions." And on newspapers, the densest multi-column layouts in the set, "most VLMs fail to recognize when dealing with the Newspapers, while pipeline tools achieve significantly better performance."
Tables and dense multi-column pages are what a schedule of values and a loss run are. If your extraction problem is mostly tabular data in a predictable format, a mature OCR pipeline is a serious contender and may well be the better answer.
Where vision models pull clearly ahead is modern handwriting. On the IAM English dataset, a benchmark study against Transkribus recorded a character error rate of 9.13% for "The Text Titan I," an off-the-shelf handwriting model trained on multilingual material, against 1.71% for GPT-4o mini. On the French RIMES set the figures were 10.71% and 1.63%.
The same study is careful about how far that generalizes, and so should we be. It reports "no consistent advantage for either approach" overall, with Transkribus ahead on German, multilingual, and historical material, and it notes that accuracy declines markedly on non-English text. Performance also varies enormously inside the "vision model" category: on IAM the best scored 1.71% and the weakest 25.15%, nearly three times worse than the specialist.

The practical read is narrower than "agents beat OCR." Where your documents are tabular and your layouts are stable, pipeline OCR is cheap, fast, and often more accurate. Where they include handwriting, photographs, and layouts you don't control, the strongest vision models are substantially better. Most carriers have both problems, which is an argument for keeping the choice reversible.
One further caveat on how finished any of this is. On ReceiptBench, a 2026 benchmark built from 10,656 real receipt images, the leading model scored 0.9086 on normalizing values it had already read, but only 0.5781 on reconstructing document structure, with overall accuracy around 0.74. Turning a messy document into correctly structured records remains unsolved, which argues for an extraction layer you can keep changing as the benchmarks move.
FurtherAI runs alongside the core rather than inside it. Connectors, launched in July 2026, is a library of native integrations to the CRMs, policy administration platforms, document repositories, agency management systems, and enrichment sources already running in insurance operations. Guidewire, Duck Creek, and Majesco sit among the policy administration systems, alongside Salesforce, Microsoft Dynamics, Applied Epic, AMS360, and SharePoint. Every connector "goes through a security review before it ships," with the scope of each integration evaluated for what data it reaches, what actions it can take, and how it authenticates. Workflows read from and write to those systems directly, and admins control which connectors are active.
The practical effect is that the underwriter keeps working in the system they already use while the AI layer, its configuration, and its contract stay yours.
Upland Capital Group, an AM Best A- rated specialty P&C insurer, selected FurtherAI in October 2025 to ingest broker submissions and extract underwriting fields. VP of Digital Delivery Doug Alexander put the decision this way: "After evaluating several vendors, we chose FurtherAI for its performance, insurance expertise, and partnership approach." On the outcome side, a top-10 carrier used FurtherAI to re-map broker-submitted data into its own proprietary underwriting schema across 32 critical property fields, cutting SOV processing from one to five days down to under 10 minutes.
If you want to test the argument on your own stack, a demo is the fastest way to see what the connector layer does with your documents.
REFERENCES
Anthropic. "Model Deprecations." Anthropic Platform Documentation, accessed September 2026. platform.claude.com
Anthropic. "Claude Opus 4.1." Anthropic News, August 5, 2025. anthropic.com
Capgemini. "World Property and Casualty Insurance Report 2026." Capgemini Research Institute, May 2026. capgemini.com
Crosilla, Giorgia, Lukas Klic, and Giovanni Colavizza. "Benchmarking Large Language Models for Handwritten Text Recognition." arXiv:2503.15195, March 2025. arxiv.org
EIOPA. "Opinion on Artificial Intelligence Governance and Risk Management." European Insurance and Occupational Pensions Authority, EIOPA-BoS-25-360, August 6, 2025. eiopa.europa.eu
European Union. "Regulation (EU) 2022/2554 on Digital Operational Resilience for the Financial Sector (DORA)." Official Journal of the European Union, December 14, 2022. eur-lex.europa.eu
FurtherAI. "Complex Property SOV Intake." FurtherAI Customer Stories. furtherai.com
FurtherAI. "Introducing Connectors: One Workspace, Every Insurance System." FurtherAI, July 28, 2026. furtherai.com
Mayer Brown. "US NAIC Spring 2026 National Meeting Highlights: Innovation, Cybersecurity and Technology (H) Committee Update." Mayer Brown, April 2026. mayerbrown.com
Microsoft. "Model Deprecations and Retirements." Microsoft Foundry Documentation, accessed September 2026. learn.microsoft.com
NAIC. "Model Bulletin: Use of Artificial Intelligence Systems by Insurers." National Association of Insurance Commissioners, December 4, 2023. content.naic.org
Ouyang, Linke, et al. "OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations." CVPR 2025, arXiv:2412.07626. arxiv.org
Stanford HAI. "AI Index Report 2026, Chapter 2: Technical Performance." Stanford Institute for Human-Centered AI, 2026. hai.stanford.edu
Upland Capital Group. "Upland Capital Group Chooses FurtherAI as Strategic AI Partner to Transform Underwriting." Business Wire, October 29, 2025. businesswire.com
Wang, et al. "From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding (ReceiptBench)." arXiv:2605.22413, May 2026. arxiv.org
DISCLAIMER
This article is for general informational purposes only and does not constitute legal, regulatory, compliance, underwriting, or other professional advice. The content reflects information available as of the date of publication, and FurtherAI undertakes no obligation to update it as laws, regulations, or AI technologies evolve.
Reclaim your time for strategic work and let our AI Assistant handle the busywork. Schedule a demo to see how you can achieve more, faster.