Blog/Industry
Industry

Automating Commercial Insurance Submission Packages, ACORD Forms, and Loss Runs for Insurance Brokers

August 12, 2026 17 min readDocumentIQ Team

Walk into any commercial property & casualty (P&C) brokerage on a Monday morning and the story is the same. The producer has just come off a Friday call with a manufacturing prospect. There is a submission target date on Thursday for four carrier appetite matches. The producer forwarded the whole email chain — with roughly 47 PDF attachments — to the account manager, who forwarded it to the CSR, who is now three coffees in, opening each PDF one at a time, retyping fields into the AMS (Applied Epic, Vertafore AMS360, HawkSoft, EZLynx, or whatever the shop runs), reconciling loss runs from four carriers going back five years, and trying to figure out whether the schedule of values (SOV) matches the property description on the ACORD 140. By Wednesday afternoon, if she is lucky, the submission package is clean enough to bind out to markets. On Thursday morning she starts on the next one.

This is not an edge case. This is the daily work of commercial insurance intake — at wholesale brokers, retail agencies, MGAs (Managing General Agents), MGUs (Managing General Underwriters), and the underwriting operations at every carrier that receives inbound submissions from the broker channel. And it is the single largest source of friction between "the producer just landed a lead" and "the account is on paper" in the entire commercial insurance workflow.

This guide is written for commercial insurance operations leaders, brokerage COOs, MGA underwriting managers, and heads of insurance technology at brokerages who have hit the ceiling of what a paralegal-heavy submission-intake process can carry. It covers what a real submission package contains, why traditional OCR and AMS-native intake tools break on it, what "getting it right" means when the downstream carrier will bounce a submission for a missing NAICS code or a mis-typed Total Insured Value (TIV), and how to build a production-grade extraction pipeline using DocumentIQ that ties every ACORD field, every loss run entry, and every schedule of values line item back to a single quotable account — with an audit trail that survives an E&O review. This is the workflow that modern LLM-based document extraction was built for.

What a Real Submission Package Actually Contains

For anyone outside the commercial P&C world, "submission package" sounds like it should be a single form. It is not. A single mid-market submission — for a manufacturer with three locations, a fleet of 40 trucks, 180 employees, and a workers' comp / general liability / property / auto / umbrella / cyber tower — will typically arrive as 20 to 60 individual PDFs, spread across a chain of five to twelve emails, sometimes with the same file re-attached three times because someone in the chain "clarified" one field on the ACORD.

The package almost always includes some combination of the following, and every carrier's underwriting appetite requires a slightly different subset:

  • ACORD 125 — Commercial Insurance Application. The general information form. Applicant name, mailing address, DBA, FEIN, NAICS code, SIC code, entity type (corporation, LLC, partnership, sole proprietorship), years in business, description of primary operations, subsidiaries, and prior policy history. This is the "who is this account?" foundation and every downstream form ties back to it.
  • ACORD 126 — Commercial General Liability (CGL) Section. GL exposure basis: gross sales, payroll, subcontract cost, area (sq ft), number of employees, number of admissions, products/completed operations exposure. Requested limits: per occurrence, aggregate, products aggregate, personal & advertising injury, medical payments, damage to premises rented. Additional insureds by name and endorsement form number (CG 20 10, CG 20 37, CG 20 26, etc.).
  • ACORD 140 — Property Section. Location-by-location schedule. For each location: address, occupancy, construction class (ISO class 1–6: frame, joisted masonry, non-combustible, masonry non-combustible, modified fire-resistive, fire-resistive), year built, square footage, protection class (PPC 1–10), sprinklered Y/N, alarm class, and a full schedule of values by coverage — building, business personal property (BPP), business income, extra expense, tenant improvements & betterments (TI&B).
  • Statement of Values (SOV). The property schedule in Excel or PDF format, one row per location, dozens of columns per row. For any multi-location risk larger than roughly $5M in Total Insured Value, the SOV is the underwriter's primary data artifact. Every reinsurance treaty and every catastrophe modeling run (RMS, AIR/Verisk, Karen Clark) starts from the SOV.
  • ACORD 130 — Workers' Compensation Application. Classification codes (NCCI or state-specific), payroll by class code, three-year loss experience (loss runs), experience modification factor (E-Mod / X-Mod), officer inclusion/exclusion, subcontractor exposure, and USL&H or Jones Act exposure indicators.
  • ACORD 137 — Commercial Auto (Business Auto). Vehicle schedule with VIN, year, make, model, garaging address, radius of operation, primary use (service, retail, commercial), body type, GVW, cost new, and each unit's coverage requirements (liability limit, comp/collision deductibles, hired/non-owned auto).
  • ACORD 131 — Umbrella / Excess Liability. Requested limits, underlying schedule (GL, auto, employer's liability with per-accident and disease limits), self-insured retention (SIR), and drop-down provisions.
  • ACORD 141 — Crime Section. Employee dishonesty, forgery/alteration, computer fraud, funds transfer fraud, in-transit money & securities, on-premises money & securities.
  • Cyber / Privacy Liability Supplemental Applications. Every carrier has their own — Chubb, Beazley, Coalition, At-Bay, Corvus, Travelers, CFC, Tokio Marine HCC — often 8 to 15 pages of questions on data volume, PII/PHI/PCI record counts, MFA deployment, backup posture, patching cadence, endpoint detection & response (EDR/XDR), incident response plan status, and prior claims.
  • Loss Runs — Currently Valued. One per carrier per line of coverage for the prior three to five years. The single most-scrutinized document in the entire package. Every loss with its date of loss, date reported, status (open/closed/reopened), claim number, description of loss, indemnity paid, expense paid, indemnity reserve, expense reserve, incurred (paid + reserves), recoveries, and subrogation status.
  • Currently Valued Experience Modification Worksheet. For workers' comp, the NCCI or state rating bureau's E-Mod worksheet showing the ratable losses, expected losses, D-ratio, weighting values, and calculated modifier.
  • Prior Policy Declarations. Dec pages from every expiring policy, showing bound limits, deductibles, endorsements, premium, and named insureds.
  • Financial Statements. For accounts above a size threshold (varies by carrier, often $25M in revenue or bond capacity), the last two or three years of audited financials or reviewed statements.
  • Supplemental Questionnaires. Anything the underwriter specifically asked for. Contractors get a subcontractor questionnaire and a project schedule. Habitational risks get an inspection report. Restaurants get an alcohol sales percentage schedule. Anything with product liability exposure gets a product questionnaire.
  • Broker's Cover Letter and Underwriting Notes. The broker's own summary of the account — narrative context that never lives in a fielded form, but that materially changes underwriting appetite. Loss history explanations, recent risk-improvement investments, planned expansions.

Every one of these is a PDF. Almost every one arrives by email. And every field on every one of them has to be reconciled against every other field on every other one of them before a single carrier will look at the account seriously. The name on the ACORD 125 has to match the named insured on the loss runs which has to match the entity on the FEIN records which has to match the schedule of values header. If it doesn't, the submission bounces back with a broker portal error or a friendly-but-firm email from the underwriter's assistant.

Why Manual Submission Intake Is So Painful

The pain has five layers, and every commercial brokerage COO recognizes all of them.

1. Volume compounds silently

A mid-sized retail commercial brokerage will process 300–800 new business submissions per month across producers. A larger regional brokerage will process 1,500+. A wholesale broker touching E&S markets will process 2,000–5,000, because every retail broker's submission generates a wholesale re-submission to two to four carriers apiece. An MGA writing a niche program (habitational, transportation, specialty artisan, cyber for SMB) can see 500 to 3,000 inbound submissions a month per program.

The keying tax scales linearly. Every submission is 30–90 minutes of CSR time — realistically closer to 90 for a mid-market risk with a real SOV — just to get the account rendered into the AMS in a form that a producer can actually work with. Multiply out: a shop doing 500 submissions per month is looking at 500 to 750 hours of pure data-entry labor per month before anyone underwrites, quotes, or binds anything. That is three to five full-time headcount whose entire job is retyping PDFs.

The kicker: this work does not scale linearly with revenue. A submission that turns into a quote turns into a bind pays commission. A submission that dies at the "we couldn't get it clean by the deadline" stage pays nothing. And the hit rate on inbound submissions in commercial P&C is famously low — 30% for a producer-driven retail shop with a curated pipeline, closer to 8–12% for an inbound wholesale desk. Which means 88% of the keying labor is directly wasted.

2. Every carrier's supplemental and every broker's cover letter is a different format

There is a myth that ACORD forms solve this problem. They do not, for two reasons:

  • Only the base ACORD forms are standardized. The moment coverage moves beyond the standard lines, every carrier has their own supplemental. There are, at last count, more than 40 distinct cyber supplemental applications in wide use across the US commercial market alone. There are dozens of contractor supplemental packages. Every wholesale broker's own "producer package" adds another layout on top.
  • The ACORD forms themselves have version drift. ACORD 125 has revisions across every edition year — 2013, 2014, 2016, 2017. Field labels shift. Section numbering shifts. A CSR looking at a stack of ACORDs from three different producer shops in three different edition years cannot even count on the "Description of Operations" being in the same box.

Layer on top the reality that not every producer fills every field. Half the submissions arrive with the SOV attached as a native Excel file (great — but every column header is different across brokers). A quarter arrive with the SOV as a scanned-and-photocopied handwritten sheet from a real estate portfolio manager who does not do this often. A few arrive as a PDF-of-a-scan-of-a-PDF where the ACORD is illegible in the bottom third of the page and someone has to email the producer for a redo. Rossum, ABBYY FlexiCapture, Kofax, and every template-based OCR tool that has ever tried to solve this workflow struggles here — they can be tuned to one edition of one ACORD, occasionally to a family of them, but they never generalize across the full customer mix without a full re-training cycle every time a carrier updates a supplemental (which happens at least annually per carrier). Read our OCR vs LLM extraction comparison for the underlying reason why.

3. The most valuable content on a loss run is prose

A loss run's tabular header data — claim number, date of loss, incurred, paid, reserves — is the easy part. The underwriting value of a loss run lives in the loss description narrative: two or three sentences per claim describing what actually happened. That is where an underwriter figures out whether the $85,000 general liability claim was a slip-and-fall in a lobby (low severity, well-understood exposure) or a products-liability action from a defective component that shipped to 400 customers (open-ended severity, potential class-action tail).

A useful extraction system has to pull the claim narrative out of a paragraph inside a table cell, classify the loss cause into standard Insurance Services Office (ISO) or NCCI cause-of-loss codes (fire, theft, water damage, slip-and-fall, product liability, employment practices, cyber-related, motor vehicle), identify whether the claim references a specific injured party or claimant (name redaction for privacy handling), identify the disposition (paid, litigated, subrogated, denied), and detect claims that carry litigation exposure indicators — words like "attorney of record," "demand letter," "complaint filed," "trial date," "settlement conference." Traditional OCR pipelines simply cannot do this — they can read the words but they cannot tell the difference between a routine premises liability claim narrative and a claim narrative flagging active litigation.

4. The cross-document interlocks are what an underwriter will bounce a submission over

A carrier's underwriting portal — especially at large writers like Travelers, The Hartford, Liberty Mutual, Chubb, Zurich, AIG, or CNA — will reject a submission if any of the following do not tie together across the package:

  • Named insured on the ACORD 125 matches the named insured on every ACORD supplemental and on every loss run header.
  • FEIN on the ACORD 125 matches the FEIN referenced on prior policy dec pages and on the loss run header block.
  • NAICS code and class-code payroll on the ACORD 126 (or the workers' comp ACORD 130) match the operations description on the ACORD 125 and the SIC assignment.
  • Total Insured Value on the ACORD 140 property section equals the sum of the SOV line items to the dollar — and matches the requested blanket property limit.
  • Requested limits on the umbrella application match the underlying schedule (a broker who asks for a $10M umbrella with underlying auto limits of $500K CSL will get an immediate bounce because most carriers require at least $1M underlying auto to attach an umbrella).
  • Loss run currency date. Every carrier's underwriting guideline requires loss runs "currently valued within 60 days" (some 90 days). A loss run valued as of six months ago will be rejected on receipt.
  • Loss run continuity. A three-year loss run history is standard on submission. If the account changed carriers mid-cycle and there is a two-month gap in the loss run trail, the underwriter will bounce for "prior carrier loss run required."
  • Experience modification on the WC ACORD 130 matches the current NCCI or state rating bureau publication.

Manually catching all of these interlocks means the CSR has to open every PDF and mentally reconcile every field pair — which is precisely why the average submission takes 60 to 90 minutes even with a competent operator. Miss one, and the entire submission is delayed by a day or two while the broker chases the correction. Miss two, and the carrier's underwriter deprioritizes the account behind cleaner submissions from other brokers. The producer's Thursday binder deadline slips to next Wednesday and the account starts hunting for a new broker.

5. E&O exposure is real and it lives in the intake pipeline

Every commercial broker carries an errors and omissions (E&O) policy for exactly this reason. The single largest category of E&O losses in the industry is not producer misrepresentation — it is intake-side keying errors. A CSR who transcribes a $5M building value as $500K on the SOV creates a coinsurance exposure for the insured that comes home to roost only when a large fire loss hits. A CSR who copies a wrong endorsement number into the additional-insured schedule bounces the certificate the insured presents to a general contractor, and the general contractor's back-charge lands in the broker's lap.

The industry's own claims data — ISO E&O claim frequency studies, IIABA (Big I) member surveys, IRMI's periodic E&O analyses — consistently shows that roughly 30–40% of broker E&O claims trace back to intake / policy administration errors, and the majority of those trace back to manual data-entry mistakes in the submission-to-issuance workflow. When the volume-of-keying problem is also the largest source of professional liability, the case for automating it stops being a productivity question and becomes a risk-management question.

Why the Existing Tools Do Not Fix This

Every broker of any size has tried some combination of the following, and every one of them has a specific failure mode.

AMS-native intake features. Applied Epic's IVANS Digital Transfer, Vertafore's Distribution Suite, HawkSoft's ImagePlus — all provide some intake automation, primarily for renewal downloads from carrier partners and for scanning inbound mail. None of them do meaningful extraction from a heterogeneous submission package. They index and store the PDFs but the CSR still has to open each one and re-key.

Traditional OCR / template capture. ABBYY FlexiCapture Insurance, Kofax TotalAgility for Insurance, WorkFusion, and legacy IBM Datacap installations. All can be tuned to a specific ACORD edition. All require heavy per-format template setup and re-training when forms drift. Cyber supplementals defeat them entirely because the forms change too frequently. Loss run narrative extraction is beyond their design envelope. See our DocumentIQ vs ABBYY comparison for the underlying architectural difference.

Cloud AI document services. AWS Textract Insurance, Azure AI Document Intelligence prebuilt models for W-2s and 1099s, Google Document AI. Solid at forms with known layouts. Fine for a W-9. Not designed for an ACORD 130 with 47 class codes, a 40-line SOV, and a loss run PDF from Travelers-branded ClaimCenter output that changes format twice a year. See DocumentIQ vs Textract and DocumentIQ vs Azure Document Intelligence for the trade-offs.

Point solutions for narrow slices. Products that only extract ACORDs. Products that only parse loss runs. Products that only ingest cyber applications. Every one of them solves 15% of the intake problem and forces the CSR to keep three vendors' outputs stitched together in a spreadsheet.

Outsourced BPO / KPO. Every brokerage above a certain size uses an offshore data-entry team, typically in India or the Philippines. Costs run $8–$15 per submission at scale. Cycle time is 24–72 hours per submission, which for the mid-market is exactly the window the producer does not have. Quality is inconsistent and requires QA sampling that adds another layer of cost. The ROI calculator makes the head-to-head cost model easy to work out for a given book size.

The right answer, and the one every broker who has actually done the analysis converges on, is LLM-driven extraction that reads every form contextually, understands the full package as one connected object, and produces a structured, cross-validated dataset — one that ties every ACORD field, every SOV line, and every loss run entry back to the same named-insured record, catches the interlock inconsistencies automatically, and hands the CSR an exception queue instead of an inbox of PDFs.

That is exactly what DocumentIQ is built for. See What is Intelligent Document Processing? for the framing, and What is LLM Extraction? for how the underlying approach differs from OCR.

What "Getting It Right" Actually Means

Before we sketch the extraction schema, it is worth being explicit about what the output has to do, because a lot of intake tools produce output that looks structured but is not actually useful downstream.

A production-grade extraction pipeline for a commercial P&C submission has to:

  1. Recognize the document type on receipt. Every PDF arriving in the intake pile has to be classified into one of ~20 canonical types (ACORD 125, ACORD 126, ACORD 130, ACORD 137, ACORD 140, ACORD 141, ACORD 131, SOV Excel, SOV PDF, loss run, prior dec, financial statement, cyber supplemental, contractor questionnaire, restaurant supplemental, etc.). The extraction schema depends on the type.
  2. Ground every field back to the source page and line. The output is not just { "named_insured": "Acme Manufacturing LLC" } — it is { "value": "Acme Manufacturing LLC", "source_document": "acord_125.pdf", "source_page": 1, "source_bbox": [x, y, w, h], "confidence": 0.98 }. Without this, no CSR will trust the output enough to skip verification, and no E&O carrier will accept it as a compliance artifact.
  3. Normalize into a canonical schema regardless of form edition. ACORD 125 has slightly different field labels across edition years. Every extracted field goes into the same canonical name in the output.
  4. Cross-validate across documents. The named insured on the ACORD 125 has to match the named insured on the ACORD 126, the ACORD 140, the ACORD 130, every supplemental, and every loss run. When they don't match — which they usually don't, because someone typed "LLC" once and "L.L.C." somewhere else, or the DBA is on one form and the legal name on another — the pipeline surfaces the mismatch for review instead of silently picking one.
  5. Compute TIV and reconcile against the SOV. Sum every line item on the SOV. Reconcile against the ACORD 140 total. Reconcile against the requested blanket limit. Flag any discrepancy over $50,000 (or whatever the shop's tolerance is).
  6. Parse the loss run into per-claim structured records with paid, reserves, incurred, disposition, and — critically — a classified loss cause and a litigation-exposure flag.
  7. Produce a submission-package readiness scorecard that flags every missing form, every stale loss run, every un-reconciled interlock, and every "carrier-appetite disqualifier" (e.g., the underwriter's own appetite guide says "no residential contractors" and the ACORD 125 operations description says "45% new-home framing").
  8. Emit an AMS-ready payload. The extracted, validated, and reconciled dataset renders into the exact JSON/XML shape the shop's AMS accepts on a new-business import — Applied Epic's XML feed, Vertafore's AL3/EDI, or a REST payload for a modern AMS. Or it renders as a filled-out ACORD 125 PDF ready to send to the next carrier, using the shop's own logo and cover letter.

Note the emphasis on cross-document interlocks, not just per-field extraction. This is the single biggest difference between an intake automation tool that actually saves the CSR time versus one that just moves the keying labor from "keying" to "verifying."

Field Schema — What to Actually Extract

The following schema is what a mid-market retail P&C brokerage or a wholesale broker should stand up as a starting point. Every field maps to a downstream use in the AMS, the carrier submission, or the E&O audit trail. It is intentionally exhaustive because trimming it later is easier than adding fields once the shop is committed to a workflow.

Account-level (ACORD 125 — Commercial Insurance Application)

  • named_insured — legal entity name, exactly as filed
  • dba — trade name / doing business as, if different
  • entity_type — corporation, LLC, LLP, partnership, sole proprietorship, joint venture, trust
  • fein — Federal Employer Identification Number (with format normalization)
  • naics_code — 6-digit NAICS with description
  • sic_code — 4-digit SIC (still requested by many carriers)
  • years_in_business — with founding year and years under current ownership
  • operations_description — full prose description, extracted verbatim
  • mailing_address — street, city, state, zip, county
  • primary_contact — name, title, phone, email
  • producer_name — filing broker name, license number, agency code
  • effective_date — requested policy effective date
  • prior_carrier_gl, prior_carrier_property, prior_carrier_auto, prior_carrier_wc, prior_carrier_umbrella

GL section (ACORD 126)

  • gl_exposure_basis — sales, payroll, cost, area, admissions, units
  • gl_exposure_amount — number matching the basis
  • gl_products_completed_ops_exposure — Y/N + description
  • gl_requested_limit_per_occurrence
  • gl_requested_limit_general_aggregate
  • gl_requested_limit_products_aggregate
  • gl_requested_limit_personal_advertising
  • gl_requested_limit_medical_expense
  • gl_requested_limit_damage_premises_rented
  • gl_additional_insureds — list of { name, endorsement_form, capacity }
  • gl_waiver_of_subrogation — parties named
  • gl_class_codes — list of { iso_class_code, description, exposure }

Property section (ACORD 140) — repeated per location

  • location_number
  • location_address — with geocoded lat/lon for CAT modeling
  • occupancy_description
  • iso_construction_class — 1 through 6
  • year_built
  • year_last_updated_roof, year_last_updated_electrical, year_last_updated_plumbing, year_last_updated_hvac
  • square_footage
  • number_of_stories
  • ppc — Public Protection Class, 1–10
  • sprinklered — Y/N with type (wet/dry) and coverage percentage
  • alarm_class — central station, local, none
  • building_value
  • bpp_value — business personal property
  • business_income_value — with period of restoration
  • ti_and_b_value — tenant improvements & betterments
  • coinsurance_percentage
  • deductible
  • flood_zone — FEMA zone (A, AE, V, VE, X, etc.)
  • wind_deductible_type — flat / percent
  • named_storm_deductible

SOV reconciliation

  • sov_total_building, sov_total_bpp, sov_total_bi, sov_total_tib
  • sov_total_tiv
  • sov_location_count
  • sov_acord_140_reconciled — boolean + list of discrepant fields

Workers' comp (ACORD 130)

  • wc_states — list
  • wc_class_codes — list of { ncci_code, description, payroll, number_of_employees, state }
  • wc_total_payroll
  • wc_experience_mod — X-Mod with promulgated effective date and rating bureau
  • wc_officer_inclusion — per officer: included/excluded, ownership %, elective coverage
  • wc_subcontractor_exposure — total subcontractor cost, subcontractor certificates on file Y/N
  • wc_uslh_exposure, wc_jones_act_exposure
  • wc_employers_liability_limits — bodily injury per accident, disease per employee, disease policy limit

Commercial auto (ACORD 137) — repeated per vehicle

  • vehicle_year, vehicle_make, vehicle_model, vehicle_vin
  • vehicle_body_type, vehicle_gvw, vehicle_cost_new
  • vehicle_garaging_address
  • vehicle_radius_of_operation — 0-50, 50-200, 200-500, >500 miles
  • vehicle_primary_use — service, retail, commercial, long-haul
  • vehicle_liability_limit, vehicle_comp_deductible, vehicle_collision_deductible
  • fleet_hired_auto_coverage, fleet_nonowned_auto_coverage
  • fleet_drive_other_car_coverage

Umbrella / Excess (ACORD 131)

  • umbrella_requested_limit
  • umbrella_underlying_gl, umbrella_underlying_auto, umbrella_underlying_employers_liability
  • umbrella_sir — self-insured retention
  • umbrella_follow_form — Y/N with exceptions

Cyber supplemental (fielded across all common carrier formats)

  • cyber_revenue
  • cyber_pii_records_count, cyber_phi_records_count, cyber_pci_records_count
  • cyber_mfa_deployment_percentage
  • cyber_endpoint_edr_deployment
  • cyber_backup_frequency, cyber_backup_offline_copy, cyber_backup_last_tested
  • cyber_incident_response_plan, cyber_ir_plan_last_tested
  • cyber_prior_claims_5_year — list of { date, description, indemnity, expense }

Loss run — per claim

  • claim_number
  • carrier, policy_number, policy_period_start, policy_period_end
  • line_of_coverage — GL, WC, auto, property, umbrella, crime, cyber, professional
  • date_of_loss, date_reported
  • status — open, closed, reopened, subrogated, denied
  • loss_description — extracted verbatim, plus a classified loss_cause_code
  • indemnity_paid, expense_paid, indemnity_reserve, expense_reserve, total_incurred
  • recoveries, subrogation_recovered
  • litigation_flag — boolean, derived from narrative + explicit fields
  • large_loss_flag — boolean, incurred > shop-configured threshold
  • open_reserve_flag — boolean, status=open AND indemnity_reserve+expense_reserve > 0

Package-level readiness scorecard

  • package_completeness — % of expected forms present given the requested coverage lines
  • missing_forms — list of ACORDs not present but required
  • stale_loss_runs — list of loss runs valued > 60 days ago
  • loss_run_continuity_gaps — list of coverage-period gaps in the 3-year history
  • named_insured_mismatches — list of docs where the named insured differs from ACORD 125
  • fein_mismatches
  • tiv_reconciliation_delta — $ difference between SOV total and ACORD 140 total
  • umbrella_underlying_adequacy — bool per underlying line, whether it meets attachment requirements
  • carrier_appetite_flags — list of underwriter-guideline hits (e.g., "residential contractor exposure > 20%")

That schema — with a plain-English extraction_prompt on every field, tuned once and refined over time — is enough to render a submission package into a fully quotable dataset in the AMS with zero manual keying on 70%+ of accounts on day one, climbing to 90%+ after two to three months of feedback-loop refinement.

For the underlying DocumentIQ mechanics — how each field's confidence score is computed, how few-shot examples drive accuracy on tricky carrier-specific formats, and how RAG-style search over the full package supports chat-based Q&A on the submission, see the linked glossary pages.

Building the Pipeline with DocumentIQ — End-to-End

Here is the concrete workflow. This is the pattern we have seen work at retail brokerages, wholesale brokers, and MGAs — and it maps one-to-one onto the DocumentIQ Contract Intelligence and Claims Processing function pages, both of which are close cousins of commercial P&C intake.

Step 1 — Set up the projects

Create one DocumentIQ project per document-type family. The reason to split projects rather than dump everything into one is that the extraction prompts, the annotation library, and the field schema are meaningfully different across families:

  • Project: ACORD Forms (base) — ACORD 125, 126, 130, 131, 137, 140, 141
  • Project: Statement of Values — Excel and PDF SOVs
  • Project: Loss Runs — one project per carrier family (Travelers, The Hartford, Liberty Mutual, Chubb, Zurich, AIG, CNA, and a generic catch-all), because loss-run PDF formats vary substantially per carrier
  • Project: Cyber Supplementals — one project across all cyber applications, with the schema as a superset of the common fields
  • Project: Prior Dec Pages — carrier-agnostic
  • Project: Financials — audited/reviewed financial statements

Each project has its own prompt hierarchy — a global default extraction prompt for the shop, then project-level overrides ("These are commercial P&C submission documents received via the wholesale broker channel. Named insureds may appear with LLC/L.L.C./Ltd variations that should be normalized. Dollar amounts may include or omit cents; always return two decimal places."), then per-field extraction prompts for the tricky fields.

Step 2 — Define the fields with clear extraction prompts

For every field in the schema above, write a plain-English extraction instruction. The auto-suggestion feature described in the complete guide to intelligent document processing gives a strong starting point that the intake team then refines.

For example, on gl_requested_limit_per_occurrence:

Extract the requested per-occurrence limit on the General Liability section of the ACORD 126. Look under "General Liability Limits" or "Requested Limits." Return the dollar amount as a number (e.g., 1000000 for $1M). If multiple limits are listed for different coverage parts, extract the value under "Each Occurrence" specifically. If not present, return null.

That level of specificity — and iterating on it in production — is what drives extraction accuracy from ~85% out of the box to 97%+ after tuning.

Step 3 — Annotate a canonical set of documents

For each carrier family and each ACORD edition year, upload three to five representative documents and use the PDF annotation tool to draw bounding boxes around the fields the model gets wrong or confuses. Two to three annotations per format is usually enough to correct systematic misreads.

The annotation library is the shop's own asset. As new carrier formats come through and the model bounces off them, the CSR or ops analyst adds annotations, and the pipeline learns the shop's specific book.

Step 4 — Configure the cross-document interlock rules

The interlock checks — named insured match across documents, TIV reconciliation, umbrella underlying adequacy — are configured as post-extraction rules in the pipeline. DocumentIQ's structured output makes these trivially expressible in a downstream validation service or an n8n / Zapier workflow that consumes the extraction output.

At the shop level, the typical configuration:

  • Hard block: named insured mismatches, FEIN mismatches, TIV delta > $100K, stale loss runs. These halt the pipeline and land the submission in the CSR's exception queue.
  • Soft flag: small TIV deltas, missing supplementals, umbrella underlying that barely meets attachment. These pass through but attach a note on the submission for the underwriter.
  • Auto-fix: entity type normalization ("LLC" vs "L.L.C."), dollar amount formatting, date format normalization. These are silently corrected with an audit-trail entry.

Step 5 — Configure the AMS export

DocumentIQ emits structured JSON. The final step is the transformation into the AMS's expected shape. For Applied Epic, this is a modest transform to the Applied XML new-business import feed. For Vertafore AMS360 or Sagitta, similar. For a modern AMS with a REST API, even simpler. The transform is a one-time integration job — after that, submissions land in the AMS as new-business records with every field populated, ready for the producer or account manager to work.

Step 6 — Use the chat assistant during account review

Once the extraction is complete, the whole package is chat-queryable. Underwriters and producers ask questions like:

  • "Summarize the three largest open claims across all loss runs and flag any with litigation exposure."
  • "Which locations on the SOV are in FEMA flood zones A or V?"
  • "What are the three carriers currently on the account and when do those policies expire?"
  • "Are there any subcontractor exposures called out on the WC application that aren't backed by certificates on file?"
  • "What's the total payroll on the WC application by state?"

That kind of interactive review is what separates a submission that gets underwritten seriously from one that lands at the bottom of the pile.

Metrics — What Actually Changes

Every brokerage that has stood up this pattern converges on roughly the same operational numbers. The dispersion is real but the direction is not.

  • Time per submission from intake to AMS-ready: 60 to 90 minutes manually, 8 to 15 minutes with the DocumentIQ pipeline including exception review. That is a 5–10× throughput improvement on the pure keying side.
  • Percentage of submissions that clear the CSR desk on the same business day of arrival: typically 25–40% under a manual workflow, 85–95% under the automated pipeline.
  • Cost per submission: $45–$85 fully-loaded under a US-based CSR workflow, $6–$12 including DocumentIQ platform cost and reviewer time.
  • Bind ratio lift: brokerages that can turn submissions around inside 24 hours of receipt see, empirically, a 15–25% higher bind ratio than those on a 48–72 hour cycle, because the producer hits the deadline the account is actually shopping to.
  • E&O incidence: intake-error E&O claim frequency drops by roughly the same factor as the keying labor drops, because every extracted field is grounded back to the source PDF page and bounding box, and every field has a confidence score attached.

The ROI calculator on the DocumentIQ site will let a brokerage COO plug in submission volume, CSR fully-loaded cost, and current cycle time to get a real number for their own book.

What About the Small Broker That Only Sees 40 Submissions a Month?

The math still works, but the framing is different. At small volumes, the primary win is not headcount — nobody is going to lay off their one CSR who also does customer service and claims-first-notice. The wins are:

  • The producer hits binder deadlines they were missing. One retained account at $18,000 commission covers the platform cost for a year.
  • The book-cleanliness improvement lands in the E&O renewal. E&O carriers price on incidence and on documented controls. A brokerage that can show its intake pipeline emits an auditable extraction record for every field, grounded to source, with confidence scores, gets a materially better renewal than one whose process is "we photocopy everything and file it."
  • Renewal analytics. Once six months of extraction output is in the AMS, the brokerage has a real dataset. TIV growth by client. Loss trend by industry code. Umbrella attach-point drift. That is the kind of data a growing brokerage uses to build practice groups and to renegotiate carrier appointments.

What This Looks Like for the Underwriter / Carrier

Nothing above is broker-specific. Everything applies symmetrically on the carrier / MGA underwriting-intake side. The wholesale broker sends the same 47-PDF package to the MGA underwriter, who is running the same keying-heavy triage workflow, who deprioritizes broker submissions that are messy in favor of ones that are clean, and who therefore rewards brokers who send clean packages with faster quotes and better terms.

MGAs and carrier underwriting operations use the same DocumentIQ pattern to auto-triage inbound submissions: rank on completeness score, auto-decline the ones outside appetite (flagged by the extraction pipeline against the appetite guide), and land the qualifying submissions in the underwriter's queue with every ACORD field, every SOV line, and every loss run entry already fielded. That is the workflow every carrier's "digital submission platform" is aiming at, and DocumentIQ is the tool that makes it real without a multi-year integration project.

What to Do Next

If you are running commercial P&C intake — retail brokerage, wholesale broker, MGA, MGU, or carrier underwriting operations — the practical next step is small and cheap:

  1. Pick one document class — SOVs or loss runs are usually the highest-value starting points because they consume the most CSR time per submission and drive the most downstream errors.
  2. Set up a single DocumentIQ project for that class.
  3. Feed in 30 recent submissions across your top-5 carrier formats and your top-5 producer / retail-broker formats.
  4. Define the extraction fields from the schema above.
  5. Annotate 5–10 canonical documents to teach the model your specific format mix.
  6. Compare the extracted output against your existing manually-keyed AMS records for the same 30 submissions.

You will see the accuracy floor in an afternoon, and — more importantly — you will see whether the cross-document interlock checks hold up against your real submission traffic. That is the whole value proposition. If they do, the pattern will extend cleanly to ACORDs, cyber supplementals, prior dec pages, and every adjacent document class in the intake pile.

If you want a walkthrough for your specific book mix, carrier appointments, or AMS integration, get in touch with the Algoscale team — we have rolled this pattern out for retail brokerages, wholesale brokers, and MGAs across P&C, and we can share the field schemas, annotation libraries, and interlock check rules we have found actually hold up on production submission traffic. If you want to see how the same underlying platform handles adjacent workflows in insurance operations, our P&C claims / FNOL automation guide and the Certificate of Insurance verification guide are the natural companions to this one.


Related reading:

Related DocumentIQ pages:

Related Algoscale services:

commercial insurance insurance broker ACORD forms loss run extraction submission processing professional services MGA wholesale insurance policy administration AI document extraction underwriting automation

Ready to try it yourself?

Start for Free