Information Power
Theory: Information Power | Template: The Confession | Words: 1,685
# Information Power: Quality Over Quantity in Learning
For decades, many of us in the professional world, especially those dealing with data or learning design, operated under a simple, seemingly intuitive principle: more is better. More data, more content, more participants. We believed that volume inherently brought robustness and completeness. But as we’ve dug deeper, the evidence has led us to a profound, unsettling, and ultimately liberating realization: this belief is often misguided. The true power isn't in sheer quantity, but in the focused quality of information.
What We Used to Believe
In fields like qualitative research, the gold standard for determining how much data was "enough" revolved around a concept called "saturation." The idea was straightforward: you kept gathering data – interviewing participants, for example – until no new themes or insights emerged. Once you hit that point, you'd supposedly saturated the topic, meaning you had all the information you needed. This often translated into an unspoken rule: the more interviews, the better. If you only had a few participants, you risked criticism that your findings weren't robust or comprehensive enough.
This mindset wasn't limited to research methods. In learning and development, we often designed comprehensive courses, packed with every possible detail, assuming that a broader scope meant a richer learning experience. For data science and AI, the mantra was always "feed the algorithm more data." We chased big datasets, believing that larger volumes would automatically lead to more accurate models or deeper insights. The underlying assumption was that more inputs would inevitably lead to better outputs. We defaulted to counting, rather than truly evaluating the richness of what we had.
The Turning Point
However, over time, we started noticing patterns that challenged this "more is better" assumption. We observed that adding more participants to a qualitative study didn't always yield genuinely new insights; sometimes, it just added noise or redundant information. The effort involved in recruiting, interviewing, and analyzing additional data often far outweighed the marginal new value. We found ourselves struggling to justify why a smaller, highly focused study felt so much more impactful than a larger, more diffuse one, especially when reviewers would inevitably ask, "Why only eleven participants?"
Similarly, in learning design, we saw how comprehensive courses, despite their breadth, often led to low engagement and completion rates. Imagine a Massive Open Online Course (MOOC) trying to cover an entire discipline. This suggested that simply providing a vast amount of content wasn't translating into effective learning. It felt like we were throwing a wide net, hoping to catch everything, but often coming up with very little of true value. The sheer volume became a barrier, not an asset.
The Research That Changed Everything
The real shift in our thinking came from a body of research that began to systematically question the uncritical reliance on saturation and sample size. One pivotal study by Malterud, Siersma, and Guassora (2016) introduced the "Information Power" model, fundamentally reframing how we think about sample adequacy in qualitative studies. They argued that the number of participants isn't the key; instead, it's about the power of the information each participant brings, determined by several crucial factors. This directly challenged the conventional wisdom of 'more is always better' in qualitative research (Malterud et al., 2016).
This wasn't an isolated idea. Earlier, Morse (2000) had already argued that the concept of saturation was often misused. She urged researchers to consider the study's scope, the topic's nature, data quality, and design when determining sample size, moving beyond just counting heads (Morse, 2000). This critical perspective highlighted that simply reaching a "saturation point" could be superficial if the data itself lacked depth.
Further studies reinforced this. Guest, Bunce, and Johnson (2006) conducted an experiment showing that data saturation could be reached with as few as 6 interviews in a homogenous group. This finding is powerful because it illustrates that when a sample is highly specific and relevant, you need far fewer participants to gain rich insights (Guest et al., 2006). It's not about the count, but the precision.
Hennink, Kaiser, and Marconi (2017) added an important nuance, distinguishing between 'code saturation' (no new categories emerging) and 'meaning saturation' (no new insights emerging). They argued that achieving meaning saturation is a more rigorous and valuable goal, emphasizing depth over mere breadth of codes (Hennink et al., 2017). This aligns perfectly with the information power model's focus on the quality of dialogue and the richness of the information.
Critiques like O'Reilly and Parker (2013) further exposed the pitfalls of uncritically applying saturation, calling it a potentially superficial and misleading indicator of data quality (O'Reilly & Parker, 2013). And a review by Marshall et al. (2013) found no clear relationship between sample size and the quality or richness of qualitative findings. This suggested that other factors, such as the researcher's skills and the study design, often matter more than the number of participants (Marshall et al., 2013). This meta-analysis directly undermines the assumption that larger sample sizes automatically lead to better qualitative research.
Even simulations, like van Rijnsoever (2017), showed that saturation isn't guaranteed even with large sample sizes, and its probability depends on the prevalence of themes and the quality of data collection. This challenges the idea that saturation is a reliable indicator of data quality, pushing us to consider factors beyond just the count (van Rijnsoever, 2017).
What the Evidence Shows Now
The accumulated evidence paints a clear picture: the pursuit of sheer volume, whether in research participants or learning content, often misses the point. It's not about hitting an arbitrary number of interviews or cramming every piece of information into a curriculum. Instead, it's about maximizing the "information power" of what you have. This means prioritizing depth, relevance, and specificity.
Think about it this way: five highly relevant experts, deeply interviewed, can provide far more actionable intelligence than thirty loosely connected individuals whose insights are shallow or repetitive. This principle extends beyond research. In adaptive learning, it means focusing on precisely the information a learner needs, when they need it, rather than overwhelming them with everything. For AI, it implies that a smaller, meticulously curated and labeled dataset can often yield better model performance than a massive, noisy, and poorly structured one.
We've seen that a significant portion of qualitative studies, 40% according to a 2018 study, still use saturation as the sole criterion for determining sample size (Vasileiou et al., 2018). While the median sample size in qualitative health research studies is 20 participants (Marshall et al., 2013), this number alone tells us little about the quality of the insights gained. The real lesson is that the value isn't in the count, but in the design decisions that make each data point count.
The Framework
The Information Power model provides a robust framework for making these design decisions. It guides us away from arbitrary numbers and towards a strategic assessment of our data needs. Malterud and colleagues (2016) outline five key dimensions that determine the information power of a sample:
1. Study Aim: How focused is your research question? A narrow, specific aim means you need fewer participants to achieve depth. If you're exploring a very specific phenomenon, each piece of data is highly potent. 2. Sample Specificity: How closely do your participants match the phenomenon you're studying? Highly specific participants, who are deeply immersed in the topic, provide richer, more relevant data. This is why a homogenous group can reach saturation with as few as 6 interviews (Guest et al., 2006). 3. Theoretical Backing: How much existing theory or knowledge do you have about the topic? If you're building on strong existing theory, each new data point can be interpreted more richly and contribute more significantly. You're not starting from scratch. 4. Quality of Dialogue: How deep and reflective are your interviews or data collection methods? Rich, open-ended conversations that encourage participants to elaborate and reflect will extract more valuable information per session. Think of it like mining for gold; a good technique extracts more from less ore. 5. Analysis Strategy: What kind of analysis are you planning? A detailed narrative analysis of a few cases might require fewer participants than a broad cross-case comparison looking for common patterns across many. Your analytical goals dictate your data needs.
These dimensions show us that "saturation" isn't a magic number you hit; it's a function of thoughtful design. Consider the Stanford d.school's "Design Thinking Bootleg" (Stanford d.school, 2010). Instead of an exhaustive textbook, they created a concise guide focused on core principles. Its success proves that a focused, high-quality resource can be far more effective than an overwhelming one. Similarly, IDEO U (IDEO, 2015) offers online courses that drill down into specific design thinking skills, providing in-depth instruction in targeted areas. These focused learning experiences are highly successful. Khan Academy's (Khan Academy, 2012) early success came from short, targeted video lessons on specific math concepts, building a large library of high-quality, digestible content. These examples from real-world learning and design echo the information power principle: quality and focus over sheer volume.
The Invitation
This shift in thinking – from "more is better" to "quality is power" – has profound implications across research, learning, and data science. It challenges us to be more deliberate in our design, more precise in our data collection, and more confident in the depth of our insights, even from smaller, highly focused samples. We're moving from a defensive posture when justifying sample size to an assertive argument for the power embedded in our design.
This isn't just an academic debate; it's a practical framework for building more effective adaptive learning systems, designing more impactful research, and training more accurate AI models.
_Have you experienced a similar intellectual shift in your own field, where a long-held belief about quantity gave way to a deeper understanding of quality? What specific evidence or insights changed your thinking? Is 'more data' always better, or does targeted information win?_