← Back to articles
Article 41Draft

Source Diversity

Working draft. Statistics without a confirmed source have been removed from this companion article in a fact-audit. It is still being finalised.
The short versionRead the three-minute post: Source Diversity

Theory: Source Diversity and Evidence Provenance | Template: The Confession | Words: 1,494

# AI Evidence: Is Your Data Diverse?

Confession Hook

For a long time, many of us in the professional world operated under a simple, comforting assumption: if an idea or a system cited a lot of research, it was evidence-based. We believed that citation volume equaled rigor. The more footnotes, the stronger the foundation. It felt intuitively right. But over the years, as we've dug deeper into the actual nature of evidence, a different, more complex truth has emerged. The field used to believe that quantity of citations was the primary measure of an evidence base. The evidence itself, however, tells a different story.

What We Used to Believe

We used to think that the sheer volume of references was a good proxy for scientific backing. If a new AI system, for instance, could point to fifty academic papers supporting its design choices or its learning interventions, that felt like solid ground. The assumption was that if experts had published it, and others had cited it, then it must be good, reliable evidence. We often saw the presence of citations as the finish line, a clear signal that a practice was "evidence-based." This was particularly appealing in fast-moving fields like education technology, where the promise of data-driven decisions felt like a much-needed anchor.

The logic seemed straightforward: more evidence means more certainty. We wanted to move past intuition and anecdote, towards something more systematic. So, when AI tools started to emerge, capable of processing vast amounts of research literature and generating impressive lists of references, it was easy to fall into the trap of equating this output with genuine scientific robustness. We looked for the footnote, the bibliography, the link to a paper. If it was there, our box was checked. We believed that simply having a citation, any citation, meant our decisions were informed by science.

The Crack

The comfortable belief that "citation equals rigor" started to show cracks when we began to observe patterns in the cited evidence itself. It wasn't about a single bad study or a flawed experiment. Instead, we started seeing entire evidence bases that, despite their impressive length, felt oddly fragile. It was like looking at a wall built from many bricks, only to realize all those bricks came from the same small, local quarry, and were laid by the same crew, using the same technique.

We observed systems that cited dozens of papers, but those papers often shared the same three authors, or originated from a single research lab, or consistently used one specific methodology. It became clear that high citation volume could sometimes mask a profound lack of source diversity. This wasn't an evidence base; it was an echo chamber with footnotes. If that single research tradition had a hidden flaw, or if its foundational assumptions were ever challenged, the entire edifice would become unstable. We realized the governance question wasn't just "is there evidence?" but "is the evidence base robust?"

The Research That Changed Everything

This intellectual shift wasn't just a hunch; it was reinforced by seminal research that pushed us to look beyond surface-level metrics. John Ioannidis's powerful paper, "Why Most Published Research Findings Are False," was a wake-up call (Ioannidis, 2005). He argued that many published findings, even in reputable journals, might not hold up due to issues like small sample sizes, small effect sizes, and conflicts of interest. This wasn't an attack on science, but a critical look at how we interpret scientific output. It underscored that we cannot blindly accept published findings; we need to critically evaluate the evidence base itself. This insight directly challenged the idea that mere publication volume was enough.

Another critical piece of the puzzle came from Trisha Greenhalgh and her colleagues, whose work on "How to read a paper" became essential reading (Greenhalgh, Howick, & Maskrey, 2014). This work provided a practical guide to critically appraising medical research, emphasizing the importance of understanding study design, recognizing potential biases, and interpreting statistical significance. It highlighted that evaluating evidence isn't just about reading a paper; it's about interrogating it. We learned that an evidence base needs careful "evidence grading," where sources are rated by quality, not just by relevance or existence. A randomized controlled trial (RCT), for example, generally provides stronger evidence than an observational study or expert opinion.

This critical lens also showed us what works when evidence is robust. Take personalized learning, for instance. A 2015 meta-analysis of 48 studies found that personalized learning interventions had a moderate positive effect on student achievement (d = 0.34) (Kraft & de Oliveira, 2015). This isn't just a random finding; it's a conclusion drawn from synthesizing many studies, giving us a stronger, more diverse evidence base. Similarly, research shows that spaced repetition can improve long-term retention by as much as 50% (Cepeda et al., 2006). These findings are powerful because they come from a body of research that has been critically examined, not just amassed. They show us that when we apply rigorous evaluation to our evidence, we unlock truly effective practices.

What the Evidence Shows Now

The evolved understanding is clear: an evidence base isn't just a collection of papers. It's a carefully constructed argument, and its strength depends on its provenance. We now understand that robustness is paramount. It's not enough for an AI system to cite 50 papers; we need to know where those papers came from, how they were selected, and how reliable they truly are.

We've learned that true evidence-based practice requires more than just citation. It demands source diversity coverage, meaning the evidence is drawn from multiple independent research teams, not just one tradition citing itself. This prevents the "echo chamber" effect. We also need evidence grading presence, ensuring each source is rated by quality. Not all studies are created equal, and a robust system accounts for this.

Furthermore, we now recognize the importance of inclusion/exclusion transparency. What was deliberately left out of the evidence base, and why? This reveals potential biases or limitations. And finally, licensing compliance verifies that the cited sources are actually accessible to those who need to verify them. If we can't read the paper, we can't verify the claim.

These elements combine to form an evidence provenance chain. Can you trace each decision back to the specific, verified evidence that informed it? For AI systems, this is crucial. An AI can cite prolifically, but it cannot yet assess whether its citation base is diverse, independently sourced, quality-graded, and transparently inclusive or exclusive. That human governance layer, the critical thinking about provenance, remains absolutely essential. When students receive personalized feedback from teachers, for example, they demonstrate greater learning gains than those who don't (Hattie & Timperley, 2007). This insight is valuable because it's built on a robust, critically appraised body of work, not just a single study.

The Framework

To move beyond simply counting citations, we've developed a framework for evaluating the robustness of any evidence base, especially those powering AI systems. Think of it like a quality control checklist for the foundations of your knowledge.

1. Source Diversity: This is about breadth. Are the sources drawn from many independent research teams, or are they concentrated within a few labs or a single theoretical tradition? Imagine building a bridge: you wouldn't use steel from just one supplier if you wanted maximum resilience. You'd want materials tested and validated by different engineers. 2. Evidence Grading: This is about depth. Is each source rated by its quality and rigor? Randomized controlled trials (RCTs) are often considered the gold standard for establishing causality, while expert opinion, though valuable, sits at a lower rung. A robust evidence base explicitly acknowledges these differences. 3. Inclusion/Exclusion Transparency: This is about honesty. What research was considered but deliberately left out of the evidence base, and what were the explicit reasons for its exclusion? This transparency helps uncover potential biases and ensures a balanced perspective. 4. Licensing Compliance: This is about accessibility. Are the cited sources actually available and verifiable by anyone who needs to check them? An inaccessible citation is a hollow one, preventing genuine scrutiny. 5. Evidence Provenance Chain: This is about traceability. Can every claim, every decision, every output of the system be traced back to specific, verifiable evidence that meets the criteria above? This creates an unbroken chain of accountability.

These principles ensure that when we say something is "evidence-based," we mean it's built on a foundation that has been thoroughly vetted, not just extensively referenced.

The Invitation

The journey from believing "citation volume equals rigor" to understanding the critical importance of evidence provenance has been a significant intellectual shift for us. It’s a move from superficial metrics to a deeper, more nuanced understanding of scientific truth. It reminds us that even with powerful AI tools at our disposal, human judgment and critical appraisal remain irreplaceable.

Can AI truly assess the robustness of its own evidence, or is human oversight always essential?