For Immediate Release
A strategic examination of how language models' persuasive fluency masks unverified claims and why mathematical grounding may be the only durable solution.
Artificial intelligence systems, despite their increasing sophistication, are demonstrating a troubling tendency to confidently generate false information. This propensity for “hallucination” presenting fabricated data as fact reveals a fundamental truth problem at the heart of modern AI development. As these systems become more integrated into daily life, addressing this issue is critical to maintaining trust and ensuring responsible innovation.
That tension between what language can carry and what reality actually demands sits at the center of one of the most consequential problems in artificial intelligence today. Language models have become extraordinarily fluent. They can draft contracts, explain complex medical concepts, summarize financial reports, and write code. They can do all of this in prose that sounds authoritative, measured, and credible. The problem is not that they are wrong. The problem is that they can be wrong in language that feels right.
Researchers at GenXis Research have a name for this phenomenon: the Honesty Gap. It describes the distance between persuasive language and verified truth and it represents a fundamental challenge for anyone deploying AI in high-stakes domains.
In a research paper on the Honesty Gap, Daryl Ledyard and Philip Tyler frame the core issue with precision: "Words can escape meaning. They can rationalize, soften, blur, excuse, reframe, and drift." In human psychology, this shows up as motivated reasoning, cognitive dissonance reduction, and what researchers call ethical fading where language gradually slides away from accountability. In AI systems, the same dynamic appears as hallucination, unsupported synthesis, and citation-shaped language that lacks source custody.
The root problem, they argue, is what they call "the squishiness of words." Natural language is flexible by design. It allows approximation, metaphor, implication, emphasis, ambiguity, and context dependence. Those features make language humanly useful. They also make it, as Ledyard and Tyler put it, "a weak carrier of machine-grade certainty."
The paper proposes a formal definition: a claim is not merely a sentence it is a tuple where S is the statement, D is the domain, T is the truth condition, and E is the evidence requirement. Without those elements, language remains expressive but under-bounded. It may point toward a reality without specifying the procedure by which that reality can be checked.
This framing matters because it shifts the question from "is this sentence grammatically correct?" to "what would verified truth look like here, and how do we know when we've reached it?"
Public concern about AI has intensified because language models now operate in domains where verbal mistakes carry real consequences. Legal drafting, medical triage, education, scientific writing, financial reporting, security analysis, and software development all depend on claims that can be checked and on the ability to distinguish between statements that have been verified and statements that merely sound verified.
"The worry is not merely that systems hallucinate," Ledyard and Tyler write. "The worry is that hallucinations arrive in the same polished form as true answers. A legal citation can be fabricated in perfect legal prose. A medical explanation can sound clinically plausible while omitting a contraindication. A financial summary can appear authoritative while relying on stale facts."
In each case, the danger comes from what the researchers call "the mismatch between linguistic confidence and verified grounding."
"A legal citation can be fabricated in perfect legal prose. A medical explanation can sound clinically plausible while omitting a contraindication. A financial summary can appear authoritative while relying on stale facts."
While the AI honesty gap is a newer phenomenon, the underlying dynamic has been studied extensively in human systems particularly in education. Understanding how the gap manifests in human contexts helps illuminate why it is so difficult to close in machine contexts.
In April 2026, the U.S. Chamber of Commerce Foundation released a brief titled "The Honesty Gap: America's Academic Outcome Truth Serum," produced in partnership with the Collaborative for Student Success. The brief explains how states have lowered the bar for proficiency, creating achievement data that can paint a misleading picture one that affects students, parents, educators, and ultimately the workforce.
The Honesty Gap, as calculated in that report, measures the difference between how students perform on the National Assessment of Educational Progress (NAEP) the nationally administered gold-standard assessment and how they perform on their own state's tests. When states set lower proficiency thresholds, the reported numbers look better. The actual learning does not change.
The disparities are substantial. Consider the 2024 data:
| State | Grade | Subject | State Test Proficiency | NAEP Proficiency | Gap |
|---|---|---|---|---|---|
| Iowa | 8th | Math | 72% | 27% | 45 points |
| Virginia | 4th | Reading | 73% | 31% | 42 points |
| New York | 4th | Math | 50%+ | <40% | ~10+ points |
| Michigan | 8th | Reading | 65% | 24% | 41 points |
| Alabama | 4th | Reading | 58% | 28% | 30 points |
In Iowa, nearly three-fourths of eighth graders were considered proficient in math according to the state exam, while only a quarter met NAEP's benchmark. In Michigan, 65 percent of eighth graders were proficient in reading according to the state exam, while just 24 percent cleared the same bar on the nation's report card.
As Dale Chu wrote in a Fordham Institute commentary in February 2025: "As states continue lowering proficiency thresholds, the disconnect between what students are learning and how their progress is reported grows wider. This isn't merely a technical flaw; it's a breach of public trust."
The consequences of these gaps extend beyond policy debates. In an analysis published by the Show-Me Institute in April 2025, Cory Koedel, a tenured professor of economics and public policy at the University of Missouri-Columbia, pointed to a troubling statistic: 90 percent of parents believe their children are performing at or above grade level in reading and math, even though only about one-third of fourth- and eighth-grade students in the United States score at a proficient level on the NAEP.
"This is problematic because grades tend to carry more weight with students and parents than test scores," Koedel wrote. "Many parents assume that the grades their children receive are accurate indicators of academic progress. But this assumption is increasingly incorrect. Grades have become more and more disconnected from actual achievement."
"90 percent of parents believe their children are performing at or above grade level in reading and math, even though only about one-third of fourth- and eighth-grade students in the United States score at a proficient level on the NAEP."
Koedel's diagnosis of the human honesty gap is instructive for understanding the AI problem: "We seem to have collectively lost our appetite for bad news. Parents don't want to hear that their children are falling behind, and schools are reluctant to deliver that message. Meanwhile, states face little pushback when they lower testing standards and inflate proficiency rates."
The parallel to AI deployment is direct. Organizations do not want to hear that their AI systems are producing unreliable outputs. The people building those systems are reluctant to deliver that message. And the language of AI output polished, fluent, confident makes the problem harder to see.
Virginia offers one of the starkest examples of how the honesty gap distorts public understanding. According to data from the Thomas Jefferson Institute for Public Policy, released in February 2025, Virginia's fourth graders scored 31 percent proficient in reading and 40 percent proficient in math on the 2024 NAEP. For eighth graders, only 29 percent were proficient in both reading and math.
Parents who rely on Virginia's Standards of Learning (SOL) assessment a state-specific test with lower thresholds would see a dramatically different picture. On the 2024 SOL, 73 percent of fourth graders were proficient in reading, and 72 percent of eighth graders. In math, the SOL showed 71 percent of fourth graders proficient and 63 percent of eighth graders.
How does such a discrepancy exist? As the Thomas Jefferson Institute analysis explains, Virginia's "proficient" standards in reading on the SOL align to "below basic" on the national assessment. Virginia is one of only two states to have its "proficient" standard in reading align with "below basic" performance on the national assessment. Its math standards are only a little better aligning with "basic" on NAEP, meaning partial mastery of the skills needed for grade-level proficiency.
"You will hear that NAEP 'proficient' is too high a bar and not a good proxy for the ability to read with comprehension," said Robert Pondiscio, senior fellow at American Enterprise Institute, in the Thomas Jefferson Institute report. "A fair point as far as it goes, but I defy you to find me a single parent comfortable with her child reading at 'below basic' level."
The 2024 Nation's Report Card revealed 42 percent of Virginia fourth graders and 34 percent of eighth graders were reading below basic level on the national assessment. That means more than one-in-three Virginia students could not show even partial mastery of the reading skills necessary for grade-level proficiency.
The picture is not uniformly bleak. The Collaborative for Student Success's latest analysis notes that Massachusetts and Rhode Island closed their gaps to within 5 percentage points or less across both grades and subjects. Additionally, 14 states are holding students to an equal or higher standard than NAEP in at least one grade or subject.
As a trend, states have improved. In 2014, 23 states had "the biggest honesty gaps" in fourth-grade reading defined as 30 percentage points or larger. In 2024, only Alabama, Iowa, Nebraska, and Virginia had gaps that large in fourth-grade reading. In 2014, 14 states had "the biggest honesty gaps" in eighth-grade math. By 2024, only Iowa, Mississippi, and Virginia had gaps that large in eighth-grade math.
"To be clear, improving student outcomes takes huge commitments from states on efforts like high-quality curriculum, strong teacher development and student supports," said Jim Cowen, Executive Director of The Collaborative for Student Success. "But the truth matters. We salute the states that are embracing the issue rather than masking it or running away from it."
Virginia has committed publicly and explicitly to addressing the problem. The state redesigned its school accountability and accreditation system and committed significant funding to high-dosage tutoring and literacy initiatives.
For readers evaluating AI systems whether for internal deployment, vendor selection, or research purposes the education sector's experience offers a cautionary map. The honesty gap does not close through better intentions or more fluent language. It closes through structural changes: shared definitions, external benchmarks, verification protocols, and accountability mechanisms that reward accurate reporting over flattering reports.
The same principles apply to AI. A language model that produces polished prose is not thereby producing verified truth. Organizations that treat fluency as a proxy for reliability will find themselves making decisions based on the verbal equivalent of state test scores numbers that look good but diverge sharply from measurable reality.
Ledyard and Tyler's GenXis Research paper outlines what the antidote to verbal drift looks like: "mathematical constraint, source custody, deterministic checks, calibrated abstention, and evidence memory."
Mathematical constraint means anchoring claims in structures where the relationship between inputs and outputs can be verified formulas, algorithms, formal logic, or bounded numerical ranges rather than unbounded natural language. Source custody means tracking where information came from and requiring that origin to be verifiable before the claim is treated as reliable. Deterministic checks mean automated verification procedures that return clear yes/no answers about whether a claim holds.
Calibrated abstention is perhaps the most counterintuitive of these: a system that can reliably say "I don't know" when it lacks sufficient grounding rather than filling the gap with fluent speculation. Evidence memory means maintaining a record of what has been verified and what has not, so that claims can be traced back to their epistemic foundation.
The goal is not to reduce the utility of language. It is to create a layer of verification infrastructure that sits beneath language use something that can catch verbal drift before it compounds into misleading output.
One important clarification: the honesty gap in AI is not the same as AI dishonesty. Hallucination the generation of fluent but unverified content is a structural property of systems that optimize for linguistic coherence without optimizing for truth. It is not the same as a system deliberately lying.
A legal brief with a fabricated case citation may be the result of a model conflating sources, not a deliberate falsehood. A medical explanation that omits a contraindication may reflect gaps in the training data rather than intent to mislead. A financial summary that relies on stale facts may be drawing on outdated sources without any mechanism to flag that staleness.
The danger is that the output sounds intentional even when it is accidental. This is why Ledyard and Tyler emphasize that the problem is not less language but stronger grounding. Language remains useful it is the vehicle through which AI systems communicate. But the vehicle needs a chassis, and that chassis needs to be mathematical, deterministic, and verified.
Organizations seeking to deploy AI responsibly can draw lessons from both the AI research literature and the education sector's experience:
The education sector's experience with the honesty gap offers a clear lesson: the longer verbal drift goes unchecked, the harder it becomes to correct. State officials who lowered proficiency thresholds in one decade found it politically difficult to raise them in the next. Parents who had been told their children were proficient expected that framing to continue. Systemic incentives aligned around flattering numbers rather than accurate ones.
AI systems face similar inertial challenges. Once an organization builds workflows around AI-generated content, it becomes politically and operationally difficult to impose verification requirements that slow those workflows down. The pressure to ship, to scale, to trust the model's fluency these forces compound over time.
The antidote is early and structural. Mathematical grounding, source custody, deterministic checks, calibrated abstention, and evidence memory are not just technical features. They are the infrastructure of trust. Without them, the honesty gap widens. With them, organizations have a fighting chance of deploying AI that augments rather than undermines the reliability of their decisions.
Felix Mendelssohn understood, in the domain of music, that the most precise expressions resist translation into approximate language. The challenge of AI honesty is, at its core, the challenge of building systems that know the difference and that tell you when they don't.
###
Leadership and Authority Research
NiftyWebs