Four ways general-purpose AI gets ERISA wrong
Off-the-shelf AI is remarkably fluent in ERISA in the abstract — and remarkably unreliable the moment your specific plan is the thing that matters. The gap is predictable, and it has four shapes.
Ask a leading general-purpose model a textbook ERISA question and it will often answer well. Ask it the question your benefits team actually has — the one that turns on your plan's language, your service-crediting method, your committee's past practice — and the confident fluency becomes a liability. It will still answer confidently. It will just be answering a different plan's question.
Here are the four gaps we see most often, and what it takes to close each one.
1. It does not know how to think about ERISA
This is the failure that is hardest to see, because the answer looks complete. A general model answers from broad ERISA knowledge, and broad ERISA knowledge is genuinely impressive: it can tell you what a deferral election is, what the ADP test does, and roughly how a correction works. What it does not do is traverse. ERISA is a set of interlocking provisions, and the professional habit that matters is asking what else a question drags in.
Take a missed deferral opportunity. Ask a general model how to correct one and you will likely get a serviceable answer about the qualified nonelective contribution and the missed-deferral percentage. Ask whether the correction changes the plan's testing position for that year and the question mostly does not occur to it, because you did not ask it. The correction and the test are the same fact pattern. Treating them as two conversations is how a plan ends up correcting one failure into another.
The risk here is not that general AI gets something completely wrong. It is that it is confident about an issue it should not have considered, and does not consider an issue that is essential to the question asked. A wrong answer you can catch. A confident, well-written, incomplete answer is the one that gets forwarded.
What closes it: guidance for how to think about ERISA questions rather than only what the law says. What else to consider when one ERISA issue is raised, and step-by-step instructions for working through each question raised — the way an ERISA professional approaches the work, written down and given to the model as instruction. That is not more knowledge. The model already has the knowledge. It is method, and method is the part a general model has no way to acquire from reading about the field.
2. It doesn't know your plan
The model has read a great deal about 401(k) plans. It has not read your adoption agreement. So when eligibility, vesting, or a distribution provision turns on a design choice your plan made — a choice the statute permits several ways — the model fills the gap with the most common answer, not the correct one. On a break-in-service or a hardship provision, "most common" and "correct for you" are frequently different plans.
What closes it: putting your documents in front of it. That does not require anyone to build you a bespoke system. It requires a workspace holding your plan document, adoption agreement, SPD, forms, and recent filings — and, critically, the discipline to point a question at those documents rather than letting the tool assume a plan term. It also requires one step almost everyone skips: verifying that the documents are actually being read. Ask which document is the plan document, what the normal retirement age is, and which section says so. If the answer names the file and the section, you are grounded. If you get a general description of what normal retirement age means, you are not, and everything downstream of that is a guess.
3. It's stale in a field that moves
ERISA guidance does not hold still. Between SECURE 2.0 provisions phasing in, evolving agency positions, and a steady stream of litigation reshaping fiduciary practice, an answer that was right eighteen months ago may be wrong today. A general model's knowledge has a cutoff; it doesn't quietly refresh because the law changed.
What closes it: a dated release and honest disclosure about the date. A tool should know what day its content is current to, tell you unprompted when the question is time-sensitive, and be replaced on a schedule as guidance moves. The failure to avoid is the tool that cannot tell you how old it is — because that one will eventually hand you last year's rule with this year's confidence, and nothing in the output will signal the difference.
4. It doesn't know when to stop
The more useful failure to understand is this one. A generic model has no concept of the line between an administrative question it can handle and a judgment call that needs a lawyer. It treats "what's the deadline to deposit deferrals" and "is this arrangement a prohibited transaction" as the same kind of question — and answers both with equal confidence. The first is fine. The second is where a plan sponsor over-relying on an unsupervised tool can walk into fiduciary exposure.
The risk isn't that the tool is wrong. It's that it's wrong about a question it should never have answered.
What closes it: written criteria for what counts as a judgment call, and a tool built to apply them and say so. Not to solve the problem — a piece of software cannot resolve a question that turns on legal judgment — but to name it. The useful behavior is a tool that stops, says this one turns on judgment rather than on a plan term or a figure you can look up, says the answer is not safe to act on, and tells you to have ERISA counsel review the analysis before you commit the plan. Note what that behavior is not: it is not review, and it is not a referral. No software routes your question to a lawyer, and any tool implying that it has is telling you something untrue. Arranging the review is yours to do.
The through-line
None of this is an argument against AI in the benefits function. It's an argument about how the tool is built, and about what the person using it has to bring. Of the four gaps — no method, no grounding, drift, and no sense of its own limits — three are closable by the tool. The second one is not. Nobody can supply your adoption agreement but you.
- Ground it in your documents. Insist that any assistant answer against your plans rather than a template, and verify that it is actually reading them before you trust a word of it.
- Insist it knows its limits. Require that the hard calls get named as hard calls rather than answered smoothly — and be suspicious of any tool that claims to route them somewhere.
- Keep it current. Treat the release date as part of the product, and replace the tool on a schedule rather than discovering its age from a wrong answer.
This post is general information about technology and process, not legal advice. Questions requiring legal judgment should be directed to qualified ERISA counsel of your own choosing.