
Scope note: this guide covers new-business submissions, from broker email to a bound file. For renewal books, where the comparison is against last year's terms rather than a fresh appetite check, see our guide to AI-powered renewal summaries. For the wider lifecycle context, start with AI for underwriting.
Delegated authority grew up. AM Best puts 2025 MGA direct premium at $108.7 billion, up from $92.3 billion in 2024, with nearly 800 MGAs now meeting the NAIC annual statement reporting threshold and more than 75% of MGA contracts granting underwriting authority, as per AM Best and Risk & Insurance. Conning, counting on a wider basis that includes Lloyd's and fronting premium, puts the 2025 total nearer $128 billion. Different bases, same direction.
That growth arrives with strings. Under the NAIC's Managing General Agents Act, an insurer must "periodically (at least semi-annually) conduct an on-site review of the underwriting and claims processing operations of the MGA," with records accessible "in a form usable by the insurer." At Lloyd's, coverholders under Lloyd's Insurance Company must be audited at least every two years — and Lloyd's own May 2026 coordinated-audit guidance describes its purpose as an approach that "reduces the regulatory burden of multiple annual audits."
Meanwhile the people who'd assemble that evidence are already underwater on clerical work. Capgemini's World Property and Casualty Insurance Report 2026 found 57% of underwriters spend most of their time on routine tasks (data gathering, document review, basic eligibility checks) and only 31% work in a workbench that surfaces AI recommendations.
So MGAs are automating, and most haven't closed the governance gap behind it: a 2026 review found 52% of UK MGAs had a formal AI governance framework in place, against 93% of Lloyd's managing agents in a Lloyd's Market Association survey, according to Insurance Business UK. That gap matters because the automation and the audit trail are the same build. Bolting evidence captured onto a pipeline that wasn't designed for it costs more than designing for it once.

Regulators are specific about the standard, and it's older than AI. The NAIC's Market Regulation Handbook directs examiners to test that "file documentation adequately supports decisions made." The Market Conduct Record Retention Model Regulation requires underwriting records — applications, rating manuals, underwriting rules, and underwriting notes — to be kept for the current policy term plus three years, and declined files, including "any documentation substantiating the decision to decline," for the current year plus three. Several states extend that to five.
Failing this is mundane and it still costs. A 2025 Pennsylvania market conduct examination cited 50 recordkeeping violations under a single statute, for failing to retain documentation supporting cancellations and nonrenewals — 23 mid-term cancellations, 19 sixty-day cancellations, and eight nonrenewals, all the same finding: no proof in the file supporting the action taken, as per Pennsylvania Insurance Department.
An audit-ready summary satisfies that standard by carrying its own evidence. Six layers, attached at the moment of processing rather than assembled later:
The exception record is the layer most pipelines skip, and it's the one auditors care about. As one reinsurance client describes its own process, the judgment step is "applying expert judgment to determine whether mismatches constitute breaches or are justifiable exceptions", as per underwriting audit case study. A summary that flags a mismatch without recording the disposition creates work rather than removing it.
For the mechanics of making each of these resolve back to a source, see our note on why AI citations matter in insurance, and on what turns a flag into a defensible finding, explainable AI for insurance audits.
In the US, the expectation is documented governance producible on demand: the NAIC's Model Bulletin on AI had been adopted in 25 jurisdictions as of the NAIC's August 6, 2026 map, New York's Circular Letter No. 7 adds model inventories and annual drift testing, and Texas's Bulletin B-0003-26 states that "if a regulated entity uses AI to make a consequential decision, TDI expects a person to review and agree with all decisions before action is taken."
In the EU, the high-risk deadline moved. Regulation (EU) 2026/1744 deferred the AI Act's Annex III obligations from August 2, 2026 to December 2, 2027, as the European Commission confirms, so guidance telling you to be compliant by August 2026 is stale. And Annex III point 5(c) covers life and health risk assessment only, which puts most MGA business under EIOPA's Opinion on AI governance (BoS-25-360) instead — a regime that asks for records sufficient "to enable their reproducibility and traceability." Not high-risk, still auditable.
Five steps. The sequence matters more than the tooling, because each step defines the evidence the next one has to carry.
1. Fix the artifact before the pipeline. Write down what a finished summary contains, field by field, for one program — not for the whole book. Which 20 to 40 fields drive the appetite decision, which guideline clause governs each, and what a reviewer is allowed to override. This document becomes the extraction schema, the rule set, and the audit checklist at the same time. Teams that skip it end up automating a summary nobody agrees on.
2. Ingest where brokers already are, and log the arrival. Submissions come as email threads with mixed attachments, not as clean portal payloads. Intake has to accept ACORD PDFs, spreadsheet SOVs, carrier-formatted loss runs, and supplementals, and it has to record what arrived, when, from whom, and in what version. That arrival log is the first link in the chain; without it, every downstream citation points at a file with no provenance.
3. Extract with field-level citation, and validate against the source. Every extracted value should carry a pointer to the page, cell, or clause it came from, plus a confidence score, and low-confidence fields should route to a human rather than pass silently. This is where most accuracy problems are actually solved — see the accuracy controls in the next section.
4. Apply your rules explicitly, and record the disposition of every exception. The eligibility check should cite the guideline clause it applied, not just return a verdict. When a value breaches a binder limit or falls outside appetite, the system flags it, a named person judges it a breach or a justified exception, and that judgment is stored with the file. Texas's June 2026 bulletin makes the human review step an explicit expectation; good practice made it one already.
5. Write back, and keep the evidence attached. Push the structured summary into the workbench or policy administration system, and into bordereaux reporting where the same fields feed carrier premium reports — our note on bordereaux management software covers that handoff. Then monitor: track accuracy by field, re-test when guidelines change, and re-validate on a set cadence. New York's circular letter asks for annual drift testing at minimum.
Pilot this on one program and one channel. Scope discipline beats breadth: it gives you a clean baseline, a real accuracy number, and an audit artifact you can show a carrier before you commit the rest of the book.

MGAs that have been through a carrier audit tend to recommend the same thing, and it isn't a brand: software where accuracy is verifiable at the field level rather than asserted at the document level. In practice that means a tool that shows you which value came from which page, scores its own confidence, routes the weak ones to a person, and lets you measure accuracy per field over time.
The reason to insist on that is technical. Retrieval alone doesn't fix grounding. Stanford's RegLab found that leading AI legal research tools, all built on retrieval, still "hallucinate between 17% and 33% of the time", as per Journal of Empirical Legal Studies, 2025. Stanford's 2026 AI Index is blunter about the failure mode: on a benchmark requiring reasoning over supplied documents, "models tend to lean on general knowledge, rather than grounding their answers in the supplied documents, even when explicitly instructed to do so." The best model tested scored 73.4%. A prompt asking for citations is not a control. A system that refuses to emit an uncited field is.
The third row is the one that separates vendors. Document-level accuracy claims hide the fields that matter: a summary can be "98% accurate" and still have the total insured value wrong. Insist on the per-field number, and on knowing what it was measured against.
Each category below is scored the same way, against the three things this article is about: whether it produces an audit trail, whether it can validate its own accuracy, and whether it holds a standard across teams.
Summary features built into the core system that already holds your policy data.
Pros: single system of record, strong transactional history, familiar to compliance. Cons: changes move at platform release cadence, unstructured documents handled poorly, extraction accuracy opaque.
Horizontal extraction platforms trained on business documents broadly, configured for insurance forms.
Pros: mature extraction, transparent confidence scores, format-agnostic, fast to prove on a sample. Cons: no underwriting logic, no guideline citation, no exception record — you build the summary layer yourself.
Scripted bots that move data between systems and populate a summary template.
Pros: predictable, cheap on stable inputs, no model risk to govern. Cons: breaks on format changes, can't read unstructured documents, produces a summary with no rationale behind it.
An internal team assembles extraction, rules, and evidence capture on foundation model APIs.
Pros: complete control, no vendor lock-in, tailored to your exact guidelines. Cons: the evaluation and evidence infrastructure is most of the work, and model upgrades force re-validation. Our build-versus-buy analysis for MGAs works through the trade.
Platforms where extraction, insurance-specific decisioning, and evidence capture are one pipeline. This is the category FurtherAI is in, so read the assessment accordingly.
Pros: insurance logic and evidence capture in one pass, deploys without replacing the workbench, accuracy reportable per field. Cons: a new system to govern and integrate, dependent on a vendor's roadmap, and underwriters still have to judge exceptions.
If you're evaluating platforms more broadly than the summary artifact, our buyer's guide to agentic AI platforms for MGAs covers the wider workflow.

For a large MGA, the answer isn't the tool with the best extraction. It's whichever platform can enforce one evidence standard while still honoring each carrier's guidelines — because at 10 or 20 programs, the failure isn't inaccuracy, it's divergence. Program A's summary proves things Program B's doesn't, and the audit finding lands on whichever team documented least.
That's a governance problem with four layers, and a platform is only worth considering if it can hold all four:
The nuance large MGAs get wrong is treating standardization as sameness. Carrier A wants occupancy classified its way; Carrier B has a different schedule. The platform has to keep the schema and the evidence rules constant while letting the mapping vary underneath. That's what the top-10 carrier in our complex property SOV case study needed: broker-submitted data re-mapped into its own proprietary underwriting schema across 32 critical property fields, with SOV processing dropping from one to five days to under 10 minutes.
The clearest result we have on standardized summary creation at scale comes from one of the largest MGAs in the US — over $1.5 billion in premium across 20+ insurance programs, serving more than a million policyholders. Average time-to-clear a submission went from about 32 minutes to about one minute, a 30x improvement, with a 200%+ efficiency gain in the first three months. In that window the workflow processed over $20 billion in total insured value and saved more than 2,000 hours of manual effort, at close to 100% accuracy, as per case study.
The standardization story shows up in how customers describe the work rather than in a metric. At McGowan Excess & Casualty, Steve Wentz put it this way: "It allowed us to map our current process and autofill our workbooks, process various carrier guidelines, and really get through a ton of information and uncover additional information that you need to underwrite" (see the announcement). Novacore, which runs more than $1.5 billion of premium across 20+ programs, framed it as an operating-model decision — CEO Aaron Miller: "FurtherAI gives us data centralization, orchestration, and a true intelligence layer across our business" (as per announcement).
And the reason to standardize is that your carriers are automating their side of the audit. One insurance company supporting over 100 MGAs cut audit time from roughly 200 hours to about 110 per MGA by automating the extraction and comparison work — the same evidence your files will be measured against (underwriting audit case study; see also underwriting audit automation for carriers). When the reviewer on the other side can examine the full population instead of a sample, documentation gaps that used to hide in the unsampled 90% stop hiding. We've written separately on sampling versus full-population review.
Report these to carriers and reinsurers with a pre-automation baseline beside each one. A number without a baseline reads as marketing.
Carriers care most about exception rate and evidence completeness. Reinsurers focus on audit hours and consistency of decisioning across the delegated book — which is the standardization question again, expressed as a number.
We build insurance-specific AI workflows where the summary and its evidence are produced in the same pass. Submissions arrive by email or API, documents are classified and extracted with field-level citation, carrier guidelines are applied with the clause cited, exceptions route to a named reviewer, and the structured output lands in the workbench or policy administration system with the trail attached.
What that has looked like: near-100% accuracy and 32-minutes-to-one-minute time-to-clear at a $1.5 billion MGA; over 95% field-level accuracy at go-live rising to 97% within six months at a top-10 global carrier; 45% less audit time at an insurer overseeing 100+ MGAs. Upland Capital Group's COO Katherine Walas has described the models as "highly accurate," adding that the partnership "has given our teams instant clarity where we used to spend hours" (as per announcement). Leavitt Group frames the design principle we'd want held to: "structured and explainable outputs that allow account managers to review results before moving work forward," in workflows "designed with clear control points" (see case study). FurtherAI is SOC 2 Type II certified.
If you want to see what an audit-ready summary looks like on one of your own programs, book a walkthrough.
REFERENCES
AM Best. "Best's Market Segment Report: Managing General Agents Adapt to Changing Demands and Added Scrutiny." June 30, 2026. ambest.com
Capgemini Research Institute. "World Property and Casualty Insurance Report 2026." May 2026. capgemini.com
Conning. "Managing General Agents: Reconfiguring the Insurance Value Chain?" July 28, 2026. conning.com
European Commission. "Regulatory Framework for AI." Last updated August 3, 2026. digital-strategy.ec.europa.eu
European Insurance and Occupational Pensions Authority. "Opinion on Artificial Intelligence Governance and Risk Management." EIOPA-BoS-25-360, August 6, 2025. eiopa.europa.eu
European Union. "Regulation (EU) 2026/1744 (Digital Omnibus on AI)." Official Journal, July 24, 2026. eur-lex.europa.eu
Insurance Business UK. "AI Adoption Outpacing Governance Across MGA Market, Warns Intersys." July 8, 2026. insurancebusinessmag.com
Lloyd's. "Coverholder and DCA Coordinated Audits Best Practice Guidance." May 2026. lloyds.com
Magesh, Varun, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning, and Daniel E. Ho. "Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools." Journal of Empirical Legal Studies, 2025. arxiv.org
National Association of Insurance Commissioners. "Implementation of NAIC Model Bulletin: Use of Artificial Intelligence Systems by Insurers." August 6, 2026. naic.org
National Association of Insurance Commissioners. "Managing General Agents Act." Model #225. naic.org
National Association of Insurance Commissioners. "Market Conduct Record Retention Model Regulation." Model #910. naic.org
National Association of Insurance Commissioners. "Market Regulation Handbook Examination Standards Summary." 2025 edition. naic.org
National Association of Insurance Commissioners. "Model Bulletin: Use of Artificial Intelligence Systems by Insurers." December 4, 2023. naic.org
New York State Department of Financial Services. "Insurance Circular Letter No. 7 (2024)." July 11, 2024. dfs.ny.gov
Pennsylvania Insurance Department. "Market Conduct Examination Report: Foremost Insurance Company." May 6, 2025. pa.gov
Risk & Insurance. "MGA Premiums Hit $108.7 Billion in 2025 as Capacity Scrutiny Tightens." July 3, 2026. riskandinsurance.com
Stanford Institute for Human-Centered Artificial Intelligence. "2026 AI Index Report, Chapter 2: Technical Performance." 2026. hai.stanford.edu
Texas Department of Insurance. "Commissioner's Bulletin B-0003-26: Use of Artificial Intelligence." June 12, 2026. tdi.texas.gov
DISCLAIMER
This article is for general informational purposes only and does not constitute legal, regulatory, compliance, underwriting, or other professional advice. The content reflects information available as of the date of publication, and FurtherAI undertakes no obligation to update it as laws, regulations, or AI technologies evolve.
Reclaim your time for strategic work and let our AI Assistant handle the busywork. Schedule a demo to see how you can achieve more, faster.