Three ways general-purpose AI gets ERISA wrong
Off-the-shelf AI is remarkably fluent in ERISA in the abstract — and remarkably unreliable the moment your specific plan is the thing that matters. The gap is predictable, and it has three shapes.
Ask a leading general-purpose model a textbook ERISA question and it will often answer well. Ask it the question your benefits team actually has — the one that turns on your plan's language, your service-crediting method, your committee's past practice — and the confident fluency becomes a liability. It will still answer confidently. It will just be answering a different plan's question.
Here are the three failure modes we see most often, and what it takes to close each one.
1. It answers the generic plan, not yours
The model has read a great deal about 401(k) plans. It has not read your adoption agreement. So when eligibility, vesting, or a distribution provision turns on a design choice your plan made — a choice the statute permits several ways — the model fills the gap with the most common answer, not the correct one. On a break-in-service or a hardship provision, "most common" and "correct for you" are frequently different plans.
What closes it: putting your documents in front of it. That does not require anyone to build you a bespoke system. It requires a workspace holding your plan document, adoption agreement, SPD, forms, and recent filings — and, critically, the discipline to point a question at those documents rather than letting the tool assume a plan term. It also requires one step almost everyone skips: verifying that the documents are actually being read. Ask which document is the plan document, what the normal retirement age is, and which section says so. If the answer names the file and the section, you are grounded. If you get a general description of what normal retirement age means, you are not, and everything downstream of that is a guess.
2. It doesn't know when to stop
The more useful failure to understand is this one. A generic model has no concept of the line between an administrative question it can handle and a judgment call that needs a lawyer. It treats "what's the deadline to deposit deferrals" and "is this arrangement a prohibited transaction" as the same kind of question — and answers both with equal confidence. The first is fine. The second is where a plan sponsor over-relying on an unsupervised tool can walk into fiduciary exposure.
The risk isn't that the tool is wrong. It's that it's wrong about a question it should never have answered.
What closes it: written criteria for what counts as a judgment call, and a tool built to apply them and say so. Not to solve the problem — a piece of software cannot resolve a question that turns on legal judgment — but to name it. The useful behavior is a tool that stops, says this one turns on judgment rather than on a plan term or a figure you can look up, says the answer is not safe to act on, and tells you to have ERISA counsel review the analysis before you commit the plan. Note what that behavior is not: it is not review, and it is not a referral. No software routes your question to a lawyer, and any tool implying that it has is telling you something untrue. Arranging the review is yours to do.
3. It's stale in a field that moves
ERISA guidance does not hold still. Between SECURE 2.0 provisions phasing in, evolving agency positions, and a steady stream of litigation reshaping fiduciary practice, an answer that was right eighteen months ago may be wrong today. A general model's knowledge has a cutoff; it doesn't quietly refresh because the law changed.
What closes it: a dated release and honest disclosure about the date. A tool should know what day its content is current to, tell you unprompted when the question is time-sensitive, and be replaced on a schedule as guidance moves. The failure to avoid is the tool that cannot tell you how old it is — because that one will eventually hand you last year's rule with this year's confidence, and nothing in the output will signal the difference.
The through-line
None of this is an argument against AI in the benefits function. It's an argument about how the tool is built, and about what the person using it has to bring. The three gaps — no grounding, no sense of its own limits, and drift — are closable, but only two of the three are closable by the tool alone. The first one needs you.
- Ground it in your documents. Insist that any assistant answer against your plans rather than a template, and verify that it is actually reading them before you trust a word of it.
- Insist it knows its limits. Require that the hard calls get named as hard calls rather than answered smoothly — and be suspicious of any tool that claims to route them somewhere.
- Keep it current. Treat the release date as part of the product, and replace the tool on a schedule rather than discovering its age from a wrong answer.
This post is general information about technology and process, not legal advice. Questions requiring legal judgment should be directed to qualified ERISA counsel of your own choosing.