Home / Blog / Garbage In, Garbage Out
Published: September 09, 2026

Why 'Garbage In, Garbage Out' Still Applies in the Age of AI Research

Why 'Garbage In, Garbage Out' Still Applies in the Age of AI Research

Garbage in, garbage out is a maxim from 1950s computing, coined before anyone imagined a model that could write fluent, confident, footnoted prose from bad inputs. That fluency is exactly what makes the old rule more dangerous today, not less relevant. A spreadsheet built on bad data looks wrong. An AI-generated research summary built on bad or fabricated inputs reads like it was written by an expert, citations included. Epignosis Insights' review of federal AI guidance, industry codes of conduct, and disclosures from data-driven public companies shows the AI research pipeline hasn't repealed GIGO; it has simply moved the garbage further upstream, where it is harder to see and easier to trust.

The Citation Contamination Problem

The clearest evidence sits in the scientific literature itself. An audit of roughly 2.5 million PubMed-indexed papers, published as a letter to The Lancet and reported by Nature in 2026, found that fabricated references have climbed sharply since generative AI writing tools went mainstream: about one in 2,828 papers contained a fabricated citation in 2023, rising to one in 458 in 2025 and one in 277 in the first seven weeks of 2026, a twelvefold increase in roughly two years. Nature's reporting traced the sharpest jump to mid-2024, coinciding directly with the spread of AI writing tools. These are not obscure preprints; they are peer-reviewed papers that cleared review with invented sources intact, which means the contamination is already downstream in citation graphs other researchers build on.

Why Regulators Are Building Data-Quality Gates Into AI Governance

Government standard-setters have taken notice. The U.S. National Institute of Standards and Technology's AI Risk Management Framework, and its companion Generative AI Profile (NIST AI 600-1), explicitly names data quality, provenance, and information integrity as core risks organizations must map and measure before deploying generative systems, not an optional add-on to a broader compliance checklist. The framework remains voluntary, but the Federal Trade Commission, the Consumer Financial Protection Bureau, the Food and Drug Administration, and other federal regulators increasingly reference NIST's AI RMF principles in their own enforcement guidance, which means a research pipeline that skips input verification is now positioned as a governance gap, not just a quality one.

The Industry's Own Rulebook Is Racing to Catch Up

Professional research bodies are moving in the same direction from the practitioner side. The ICC/Esomar International Code of Conduct, the standard used across the global market research and insights industry, was revised specifically to address generative AI: its Article 9 makes disclosure mandatory whenever synthetic data or AI-generated content is used in published research, and its updated privacy provisions require notifying participants and authorities if AI-driven data issues occur. Esomar has since gone further, launching an AI Alliance in mid-2026 to build shared verification standards, glossary terms, and buyer guidance, because the industry's own membership recognized that self-certification without a shared quality bar was not going to hold.

What the Market Is Already Pricing In

Advisory firms are now quantifying the downstream cost of skipping verification. Gartner predicts that by 2028, half of all organizations will adopt a zero-trust posture for data governance specifically because of the growing volume of unverified AI-generated data feeding back into enterprise systems, warning that future model generations trained on unverified AI outputs risk model collapse, where a system's answers stop reflecting reality at all. That prediction lands at an awkward moment: Gartner's own 2026 CIO and Technology Executive Survey found 84 percent of respondents planned to increase GenAI funding this year, meaning investment is accelerating well ahead of the verification infrastructure Gartner says will eventually be required.

The Business Case for Curated Data

Some public companies are already building their AI strategy around this exact gap. Thomson Reuters' most recent annual filing describes a data foundation of 1.9 billion documents and 36 million editorial enhancements underpinning what the company brands fiduciary-grade AI, explicitly positioned against general-purpose tools for legal, tax, and compliance work where, as the company puts it, almost right is not good enough. The strategy shows up in the numbers: Thomson Reuters reported a 92 percent customer retention rate in 2025, evidence that professional buyers are willing to pay a premium for AI grounded in verified, curated source material rather than an open web corpus of unknown provenance.

Epignosis Insights' Verification Standard

Given fabricated citations climbing twelvefold in the scientific literature, regulators writing data-provenance requirements into AI risk frameworks, and the industry's own trade body mandating AI-use disclosure, Epignosis Insights treats source verification as a gate, not a formality. No statistic drawn from an AI-assisted research pass is published without being traced back to its original government, association, company, or news source and checked against at least one independent source that says the same thing. AI accelerates the drafting; it never gets to decide what counts as true. That distinction is the entire difference between a research brief and a plausible-sounding guess.

Frequently Asked Questions

What does 'garbage in, garbage out' mean in the context of AI research?
It means that AI models produce fluent, confident-sounding output regardless of whether their inputs or training data were accurate, so bad or fabricated source material can result in polished but false research summaries and citations.
How much has AI-related citation fabrication actually increased?
An audit of roughly 2.5 million PubMed-indexed papers found fabricated references rose from about one in 2,828 papers in 2023 to one in 277 in early 2026, a twelvefold increase in two years, as reported by Nature.
Are there official standards for verifying AI-generated research?
Yes. NIST's AI Risk Management Framework and its Generative AI Profile set out data-quality and provenance requirements in the United States, while the ICC/Esomar Code mandates disclosure of synthetic data and AI use across the global research industry.
How can organizations protect themselves from AI-driven data quality failures?
By treating verification as mandatory rather than optional: tracing every AI-assisted claim back to an original, named source and confirming it against at least one independent source before publishing or acting on it.