Mortgage AI/ML Compliance Q&A: Governance, Fair Lending, and Exam Readiness

AI/ML compliance for mortgage lenders is no longer a future-state planning exercise. Fannie Mae’s LL-2026-04 took effect August 6, 2026. Freddie Mac’s Section 1302.8 took effect March 3, 2026. The Colorado AI Act took effect February 1, 2026. The California DFPI is actively examining AI/ML programs. State AGs are bringing actions under existing UDAP authority.

This Q&A focuses on the practical questions your compliance team will face as the AI/ML governance framework comes into operational reality in 2026 — what to build, what to document, what examiners are looking at, and how to keep the program running as models evolve and rules change.

Q1: We Just Discovered LL-2026-04. What Do We Do First?

The first step is the AI/ML use case inventory. You cannot scope a governance program without knowing which systems are in scope. The inventory should identify every AI/ML system used in the mortgage lifecycle, the business function each supports, the data inputs, the model owner, the deployment date, and the underlying vendor (if third-party).

A common mistake is to start with the governance policy. A policy written before the inventory is complete is a policy that does not match the actual system footprint. The inventory drives the policy, not the other way around.

Aim to have a complete inventory within 30 days. Most lenders underestimate how many AI/ML systems they actually run — the inventory typically surfaces 30-50% more systems than the compliance team expected.

Q2: How Do We Decide Which AI/ML Systems Are “High-Impact”?

The GSE frameworks and the Colorado AI Act both use a risk-tiering approach. The relevant question for a high-impact designation is whether the system can materially affect a borrower’s loan terms, access to credit, or experience in the loan process.

High-impact systems for mortgage lenders typically include:

  • Automated underwriting systems (AUS)
  • AI-driven appraisal valuation models (AVMs with machine learning components)
  • AI-driven fraud detection models that affect application decisions
  • Pricing optimization models that influence loan pricing
  • Lead scoring models that affect credit decisions
  • Income and asset verification tools that use AI
  • Customer service chatbots that handle credit-related inquiries

Lower-risk systems include internal marketing analytics, business intelligence dashboards, and back-office automation that does not affect borrower outcomes. Document the risk tiering methodology so examiners can review the logic.

Q3: What Counts as an “AI/ML System” for Compliance Purposes?

The LL-2026-04 definition tracks the broad industry usage. An AI/ML system is any system that uses statistical learning, neural networks, natural language processing, or other machine learning techniques to produce an output from training data. Generative AI (text or image generation), predictive models, classification models, and clustering models are all in scope.

Out of scope: rule-based decision engines that do not learn from data, simple threshold-based scoring (e.g., credit score lookups without model adjustment), and traditional statistical models without a learning component. The key question is whether the system’s parameters are learned from data rather than set by human designers.

When in doubt, include the system in the inventory and document the determination. The cost of including a borderline system is much lower than the cost of an examiner discovering a missed system.

Q4: How Do We Build a Fair Lending Testing Program for AI/ML?

A defensible fair lending testing program has four components.

1. Outcomes-based testing. Compare actual loan decisions, pricing, or other outcomes across demographic segments. The relevant segments under federal law are race, national origin, sex, religion, familial status, age, and disability. Under state law, additional protected categories may apply.

2. Input-based testing. Audit model features for proxy variables that correlate with protected classes even when the protected class is not a direct input. ZIP code is a classic proxy for race. Language preference can be a proxy for national origin. Proxies are not always a violation, but they require documentation of why the proxy is appropriate and how its use is monitored.

3. Segment-level performance review. Model accuracy, false positive rates, and false negative rates should be reviewed at the segment level. A model that performs well on average but has materially different error rates across protected segments is a model that needs remediation before deployment.

4. Counterfactual testing. For loan decisions, change the protected class of an applicant and observe whether the decision changes. If the decision changes when only the protected class changes, the model is using protected class or proxy information inappropriately.

The testing should be performed before deployment, after material model changes, and at a defined cadence — at least annually for high-impact systems.

Q5: What Documentation Do Examiners Look For First?

Examiners start with the inventory and the governance policy. From there, the document request typically expands to model documentation, validation reports, fair lending testing, vendor contracts, and the most recent attestation.

The most common finding in early AI/ML examinations is a gap between what the inventory claims and what the documentation actually supports. Lenders may claim to have a fair lending testing program but produce a single slide with results, not a documented testing protocol with methodology, results, and remediation actions.

The exam-ready binder for AI/ML governance should include:

  • AI/ML use case inventory with risk tiering
  • Governance policy approved at senior committee level
  • Model documentation for each high-impact system (data sources, training methodology, performance metrics, validation results)
  • Fair lending testing reports for each high-impact system
  • Vendor contracts with AI/ML-related terms
  • Annual attestation
  • Incident log (any model failures, complaints, regulatory inquiries)

Q6: How Do We Handle AI Tools That Loan Officers Use Independently?

If loan officers use AI tools (ChatGPT, Claude, specialized mortgage AI assistants, etc.) in connection with loan files, the tools are in scope under LL-2026-04 and Section 1302.8. The lender is the deployer and is accountable for the tool’s compliance.

The minimum controls: an approved list of AI tools that may be used in connection with loans, prohibition on uploading borrower non-public personal information to tools that have not been approved, a confidentiality review of the tool’s data handling practices, and a documented training program for loan officers on the approved-use policy.

The most common exam finding: lenders have no visibility into which AI tools loan officers are using. Shadow AI use is a significant risk, and lenders are expected to address it proactively.

Q7: How Should We Structure the AI Governance Committee?

The committee structure varies by institution size. For mid-size and large lenders, the typical structure is:

AI Governance Committee: senior leadership (CRO, CIO, General Counsel, Head of Compliance, Head of Model Risk) meets quarterly or more frequently. Approves the governance policy, reviews high-impact model changes, signs off on attestations.

Model Risk Management function: dedicated staff (or a vendor) that runs the inventory, conducts validation, performs fair lending testing, maintains documentation.

Model Owners: business line leaders responsible for individual AI/ML systems. Own the system lifecycle, escalate issues to the committee, ensure documentation is current.

For smaller lenders without dedicated model risk staff, the model risk function may be outsourced or combined with compliance. The committee structure remains the same; the execution is shared.

Q8: How Do We Balance AI Innovation with AI Compliance?

The right framing is not “innovation vs. compliance” — it is “innovation with governance.” A model that cannot be explained to an examiner is a model that creates regulatory risk. A model that produces disparate outcomes is a model that creates litigation risk. Governance is what makes innovation sustainable.

The practical implementation: the governance framework should be designed to support business velocity, not slow it down. Pre-deployment validation, fair lending testing, and documentation should be efficient and well-scoped. The model risk function should be a partner to the business, not a bottleneck.

A common failure mode: over-engineering the governance process to the point where business teams route around it. The result is shadow AI use — business teams adopt new tools without governance review, and the compliance program loses visibility into the actual system footprint.

Q9: What Is the Most Common AI/ML Compliance Failure You See?

The most common failure is treating the AI/ML governance program as a documentation exercise rather than an operational one. Lenders produce a policy and a checklist, but they do not actually run the inventory, conduct the validation, or perform the fair lending testing. When the examiner asks for the supporting documentation, the program collapses.

The second most common failure is treating vendor-provided validation as a substitute for lender validation. The lender is accountable for the model’s performance in its own use context. The vendor’s validation report is an input, not an output.

The third most common failure is fair lending testing that is too narrow — testing only adverse action outcomes and missing proxy variable analysis, or testing only protected classes under federal law and missing the additional state-level protected categories.

Q10: Where Should AI/ML Compliance Be on the Q3 2026 Priority List?

For lenders that have not yet built a program, AI/ML governance should be at the top of the Q3 2026 priority list. The LL-2026-04 effective date has passed, and the first attestation cycle is approaching. The work cannot wait for Q4.

The Q3 priorities, in order:

This month: Complete the AI/ML use case inventory. Identify model owners and risk tiering.

Next 30 days: Draft the governance policy. Get committee approval.

Next 60 days: Begin fair lending testing for the highest-impact systems. Document the testing methodology.

By year-end: Complete the first round of validation. Stand up the annual attestation process. Build the exam-ready binder.

For lenders with a program already in place, the Q3 priority is to harden the documentation for examiner review and to verify the program covers the new state-level rules — particularly the Colorado AI Act if you originate or service in Colorado.

Need support on AI/ML governance, fair lending testing, or state-level compliance overlay? Synergy works with mortgage lenders on AI/ML program design, validation, multi-state compliance overlays, and exam readiness. Book a 30-minute AI/ML review.

Web Statistics