What if medicine in the US is just like the internet? What if it is just as difficult to separate the chaff from the wheat in medicine as it is on the web?
Both the curse and the blessing of the web is its accessibility. This means that anyone's voice can be heard. And it also means that anyone's voice can be heard. So, we are just as likely to stumble upon drivel as we are on information gold. And what takes time and skill is separating the two into neat piles, one to be ruthlessly discarded, and the other cherished for how it enriches us. To be sure without the web we might not have had access to either, and it is the egalitarian nature of the internet that gives us such a variety of sources in our information diet.
Now, let's look at medicine. Every day we hear about how much noise there is in the field, and this noise is difficult, if not impossible, to separate from the signal. Some signals are becoming much clearer, and they tell us that by being too egalitarian in medicine, we have likely been causing great harm. Take, for example, PSA and mammography screenings. The drumbeat of harm associated with these highly non-specific tests and the resultant chase after false positive results, is getting deafening, and rightfully so. Every day we hear that researchers have uncovered a breakthrough mechanism or treatment, and we hear with increasing frequency that a treatment previously thought to be sacrosanct is a bunch of rubbish. What gets lost among all this noise is the possibility of a true breakthrough in disease management or treatment or cure.
Think how hard it is to separate general valuable content from bunk on the web. Now, think of the logs of increase in the levels of difficulty of this task in medicine, where difficult concepts are further shrouded in the opaque cloth of arcane and obfuscating terminology. In fact, it is so difficult, that the class previously designated as the interpreters of this information for the lay public, physicians, are unable to keep up. There is a need for a whole new class of interpreters now -- researchers and patient advocates. And while this is good for the market and the economy, since it creates jobs that had not existed before, it begs a more critical evaluation vis a vis its impact on public's health. It also begs the question of the value of this gadgetry and information glut in medicine -- what is truly the wheat and what is the chaff? And what happens when you continuously try to drink from a fire hose? And do we turn down the stream, or is there another way?
Is it feasible to limit this stream of idea and information generation? Furthermore, is it sensible to do so? Many worry that putting limitations on this is tantamount to stifling innovation. But what is innovation? The most pertinent definition to the current discussion in the Merriam-Webster dictionary is "a new idea, method or device." Nowhere does the definition incorporate the value of this idea, method or device. Perhaps it is left to the free market to determine this value and ultimate use of such innovation. Well, in a market that claims to be free, but is filled with cynical machinations in the form of favoritism, subsidies and pricing games, is objective value really what is valued? And indeed, given the complexity of these "innovations", is it even possible for the end-user to judge their value, even if the market were free?
Yet, even despite all these challenges to establishing the value of innovation on the back end, I am not sure that centrally limiting idea generation is either feasible or right. In the case of ideas on the web, I have come to the conclusion that such microblogging platforms as Twitter can be invaluable filters of information, where my network of favorite tweeters whom I follow faithfully provides me with the wheat that has already been cleaned, yet not always overprocessed. Is this possible in medicine? I know that the FDA and CMS are supposed to provide some filtration for such medical information and interventions, but each is statutorily handcuffed and gagged not to stray beyond their legislative agendas. Therefore, a value filter should not be a body beholden to the letter of the law, or to political or financial interests. It needs to be driven by the spirit of scientific curiosity, objective evaluation and pragmatism. Most importantly, it must be open to a conversation that incorporates respectful dissent and many different perspectives.
Twitter arose out of the drive to share information, and it has shaped itself as a tool for developing value in the gargantuan and ever-growing world of yottabytes. Perhaps it is citizen bloggers and tweeters, including e-patients and clinicians and researchers and writers and others, who will ultimately solve this information glut in medicine by extracting the kernel of usefulness from this morass of vegetation. Harnessing this power systematically and accurately is the next challenge of our information age.
Because ultimately, for human cognition and health, less is more. And we are still human.
Showing posts with label p-value. Show all posts
Showing posts with label p-value. Show all posts
Tuesday, August 16, 2011
Monday, January 3, 2011
Gaol fever and intercessory prayer: Redefining the role of p-value?
Happy 2011, everyone! I hope that it is everything you want it to be. Sorry for a brief hiatus in blogging -- needed to recharge my batteries and read others' writing for a change. Well, back now. And thanks to you all for coming back too.
I want to resume our recent discussions of statistical testing in the context of biologic plausibility. We discussed the latter at length a few months ago here, and came to the conclusion that our mere impression of biologic plausibility is not a good litmus test for an association. The oft-cited discovery of H. pylori as the cause of peptic ulcer disease is a tried and true example of the knowledge we would be missing today if we used biologic plausibility as the only yardstick for measuring the prospects of research.
At the same time, we spent a fair bit of time and energy talking about p values and how they need to be used in a Bayesian manner. To review, Bayes theorem relies on pre-test probability of an association to help us understand how much stock we need to put into a finding of an association. That is, the lower the pre-test probability, the more suspicious we should be of an observed association. To put it in concrete terms, for example the finding that intercessory prayer is associated with improved health outcomes requires a much greater amount of scrutiny than one that treating a bacterial infection with an antibiotic improves survival. There is a certain mechanistic elegance to the latter that is missing in the former, unless higher powers are invoked. Here is a quote from the Cochrane meta-analysis of intercessory prayer -- I especially love the last sentence [emphasis mine]:
Consider my examples above -- those of intercessory prayer and antibiotic treatment of a bacterial infection. Let us transport ourselves to, say 18th century England, where typhus, known as "gaol fever", killed more prisoners than the executioners did. How improbable would it have seemed to the medical profession of those days that a). the disease was caused by a microorganism, and b). it could be eradicated with an antibiotic? Why, I would guess that these assertions either would appear heretical or else confirm for the religious the divine presence. Either way, the biology was lacking and the plausibility was simply not there. Yet, this does not change the reality as we understand it today. What explanations will we have 200 years from now for the occasionally observed success of intercessory prayer? And more importantly, what do we do in the meantime to tread most sensibly that purgatory between accepting absurd associations and missing the unlikely ones that are nevertheless real?
The answer may be in the p value after all. Let us model qualitatively what things might look like for intercessory prayer. Let us pretend that we have just conducted the very first randomized controlled trial of the impact of intercessory prayer on the development of post-operative infection following coronary bypass surgery among 1,200 patients. We have found that there is indeed a lowered risk of infection in the intervention group, and the difference has the p value of 0.04. Great, right? We can walk away congratulating ourselves on a positive study. Well, of course this is absurd. Even though we can come up with some remotely plausible mechanism for this potentially causal association, our pre-test probability is still minuscule. The answer at this point should obviously be what has been suggested for genome-wide interaction studies: a much lower alpha level as the significance threshold. How low? This I cannot answer yet; while the rationale is, similar to genome-wide studies, a fishing expedition without much understanding of why we should find what we should find, here we are not merely engaging in multiple hypotheses testing, the number of which could help determine the appropriate significance level. No, here we are testing a single hypothesis whose mechanism is either absent or highly biologically implausible. So, how to determine the adequate threshold for significance under these circumstances remains unclear to me at this time. I can only say that the traditional 0.05 is highly inappropriate under the circumstances early in the research efforts.
As more studies are performed, their quality and directionality of results should impact how much stock we put in the results. That is, if well done studies consistently continue to demonstrate a positive association of intercessory prayer with clinical outcomes, despite inadequate mechanistic understanding, our level of skepticism should diminish, and commensurately the acceptable alpha can creep higher. In short, the more evidence and the stronger it is, despite poor understanding of why, the more liberal we can afford to be with what we consider a significant result.
So, my point? How we interpret the significance of results needs to be fluid. A p value is not a p value is not a p value. This much embattled and misunderstood statistic may yet be the bridge between Bayesian and frequentist approaches. If we get smarter about setting its thresholds, perhaps we can keep the baby while getting rid of the rancid bath water at the same time. Of course, I am not even attempting to address all of the cognitive biases that derail us in our pursuit of scientific truths. Incorporating them into our inference testing is definitely a discussion for another day.
I want to resume our recent discussions of statistical testing in the context of biologic plausibility. We discussed the latter at length a few months ago here, and came to the conclusion that our mere impression of biologic plausibility is not a good litmus test for an association. The oft-cited discovery of H. pylori as the cause of peptic ulcer disease is a tried and true example of the knowledge we would be missing today if we used biologic plausibility as the only yardstick for measuring the prospects of research.
At the same time, we spent a fair bit of time and energy talking about p values and how they need to be used in a Bayesian manner. To review, Bayes theorem relies on pre-test probability of an association to help us understand how much stock we need to put into a finding of an association. That is, the lower the pre-test probability, the more suspicious we should be of an observed association. To put it in concrete terms, for example the finding that intercessory prayer is associated with improved health outcomes requires a much greater amount of scrutiny than one that treating a bacterial infection with an antibiotic improves survival. There is a certain mechanistic elegance to the latter that is missing in the former, unless higher powers are invoked. Here is a quote from the Cochrane meta-analysis of intercessory prayer -- I especially love the last sentence [emphasis mine]:
REVIEWER'S CONCLUSIONS: Data in this review are too inconclusive to guide those wishing to uphold or refute the effect of intercessory prayer on health care outcomes. In the light of the best available data, there are no grounds to change current practices. There are few completed trials of the value of intercessory prayer, and the evidence presented so far is interesting enough to justify further study. If prayer is seen as a human endeavour it may or may not be beneficial, and further trials could uncover this. It could be the case that any effects are due to elements beyond present scientific understanding that will, in time, be understood. If any benefit derives from God's response to prayer it may be beyond any such trials to prove or disprove.At the same time, just because we do not have a mechanistic explanation at the ready does not mean that we should discount an association. In a rather lengthy post in October I wrote about my own conflicted feelings about applying Bayesian versus frequentist (this refers to all associations standing on similar probabilistic ground prior to testing) thinking in research. Although more Bayesian in my own thinking, I recognize metacognitively that it may at times be a trap:
Yet, there is something to be said about the frequentist approach, even though it is not my way generally. The frequentist approach, which is what underlies the bulk of our traditional clinical research, does not rely on differential prior probabilities for different possible associations, but treats them all equally. Despite many disadvantages, one obvious advantage is that we do not discount potential associations that do not have biologic plausibility, given our current understanding of biology, and sometimes help us stumble on brand new hypotheses. So, clearly, there is a tension here, and I am still working on what is the better way, if any.The last sentence here implies that there is a right and a wrong way, but having spent the last several months exploring these issues, I am beginning to think that this is incorrect. In fact, all of the p value discussions are leading me to believe that both approaches are useful, and it is the nuances of when either should predominate that need to be worked out.
Consider my examples above -- those of intercessory prayer and antibiotic treatment of a bacterial infection. Let us transport ourselves to, say 18th century England, where typhus, known as "gaol fever", killed more prisoners than the executioners did. How improbable would it have seemed to the medical profession of those days that a). the disease was caused by a microorganism, and b). it could be eradicated with an antibiotic? Why, I would guess that these assertions either would appear heretical or else confirm for the religious the divine presence. Either way, the biology was lacking and the plausibility was simply not there. Yet, this does not change the reality as we understand it today. What explanations will we have 200 years from now for the occasionally observed success of intercessory prayer? And more importantly, what do we do in the meantime to tread most sensibly that purgatory between accepting absurd associations and missing the unlikely ones that are nevertheless real?
The answer may be in the p value after all. Let us model qualitatively what things might look like for intercessory prayer. Let us pretend that we have just conducted the very first randomized controlled trial of the impact of intercessory prayer on the development of post-operative infection following coronary bypass surgery among 1,200 patients. We have found that there is indeed a lowered risk of infection in the intervention group, and the difference has the p value of 0.04. Great, right? We can walk away congratulating ourselves on a positive study. Well, of course this is absurd. Even though we can come up with some remotely plausible mechanism for this potentially causal association, our pre-test probability is still minuscule. The answer at this point should obviously be what has been suggested for genome-wide interaction studies: a much lower alpha level as the significance threshold. How low? This I cannot answer yet; while the rationale is, similar to genome-wide studies, a fishing expedition without much understanding of why we should find what we should find, here we are not merely engaging in multiple hypotheses testing, the number of which could help determine the appropriate significance level. No, here we are testing a single hypothesis whose mechanism is either absent or highly biologically implausible. So, how to determine the adequate threshold for significance under these circumstances remains unclear to me at this time. I can only say that the traditional 0.05 is highly inappropriate under the circumstances early in the research efforts.
As more studies are performed, their quality and directionality of results should impact how much stock we put in the results. That is, if well done studies consistently continue to demonstrate a positive association of intercessory prayer with clinical outcomes, despite inadequate mechanistic understanding, our level of skepticism should diminish, and commensurately the acceptable alpha can creep higher. In short, the more evidence and the stronger it is, despite poor understanding of why, the more liberal we can afford to be with what we consider a significant result.
So, my point? How we interpret the significance of results needs to be fluid. A p value is not a p value is not a p value. This much embattled and misunderstood statistic may yet be the bridge between Bayesian and frequentist approaches. If we get smarter about setting its thresholds, perhaps we can keep the baby while getting rid of the rancid bath water at the same time. Of course, I am not even attempting to address all of the cognitive biases that derail us in our pursuit of scientific truths. Incorporating them into our inference testing is definitely a discussion for another day.
Thursday, December 16, 2010
P-values, Bayes and Ioannidis, oh my!
When I graduated from college in the early 1980s, much to my parents' chagrin, I was not sure what to do with my career. So, instead of following some of my more wizened classmates to Wall Street, I got a job in a very well regarded molecular endocrinology laboratory in Boston. When I look back on that time, almost 30 years ago, so much seems surreal. For example, in those days it took us well over a year to sequence a gene! How about that? Something that takes hours with today's technology took an equivalent of eternity. And what a pain it was, running those cumbersome sequencing gels, with the glass cracking in the night, negating days of work. Oh, well, times have changed and for the better, I believe. But these changes are bringing with them some odd incongruities in our prior thinking.
A few days ago, I picked up a link from Michael Pollan's twitter feed to a story that startled me: It was insinuating that all these decades of sequencing human genome to hunt for targets of disease susceptibility have come up nearly empty-handed (with a few well known exceptions). The story recounted a tale of decades of vigorous funding, meteoric career growth and very few results to date. Yet, the scientists involved are not ready to give up on human genes as the primary locus for disease susceptibility. In fact, the strong rescue bias and the fear of losing all that beautiful funding are colluding to generate some creative and complex hypotheses. I came upon one such hypothesis earlier today, following a link from Ed Yong, that British science journalist extraordinaire. But the hypothesis itself, though fascinating, is not what interested me most, no. Care to guess what did? You are correct if you said the p-value.
What grabbed me is the calculation that to identify interactions of significance, the adjusted alpha level has to be set at 10^ -12; that's 0.000000000001! Why is this, and what does it mean in the context of our recent ruminations on the topic of p-value? Well, oddly enough it ties in very nicely with Bayesian thinking and John Ioannides' incendiary assertion of lies in research. How? Follow me.
The hallmark of identifying "significant" relationships (remember that the word "significant" merely means "worthy of noting") in genetic research is searching for statistical associations between the exposure (in this case a mutation in the genetic locus) and the outcome (the specific disease in question). When we start this analysis, we have absolutely no idea which gene(s) mutation(s) is(are) likely to be at play. This means that we have no way of establishing... can you guess what? Yes, that's right, the pre-test probability. We know nothing about the pretest probability. This shotgunning screening of thousands of genes is in essence a random search for matching pairs of mutation and disease. So, why should this matter?
The reason that it matters resides in the simple, albeit overused, coin toss analogy. When you toss a coin, what are your chances of getting heads on any given toss? The odds are 1:1 every time whether you will get heads or tails, meaning that the chance of getting heads is 50% (unless the coin is rigged, in which case we are talking about a Bayesian coin, not the focus of what we are getting at here). So, if you call heads prior to any given attempt, you will be correct half the time. Now, what does this have to do with the question at hand? Well, let's go back to our definition of the p-value: a p-value represents the probability of obtaining the result (in our gene case it is the association with disease) of the magnitude observed or greater under the conditions that no real association exists. So, our customary p-value threshold of 0.05 can be verbalized as follows: "There is a 5% probability that the association of the magnitude observed or greater could happen by chance if there is in reality no association". Now, what happens if we test 20 associations? In other words, what are the chances that one of them will come up "significant" at the p=0.05 level? Yes, there is a 1 in 20 (or 5 in 100) chance that this association will hit our level of significance preset at 0.05. And it follows that the chances of getting a "significant" result grow with more testing.
This is the very reason that gene scientists have asked (and now answered) the question of what level of statistical significance (also called alpha, signified by the p-value) is acceptable when exploring hundreds of thousands of potential associations without any prior clue of what to expect. And that level is a shocking 0.000000000001!
This reinforces a couple of ideas that we have been discussing of late. First, to say simply that the p-value of 0.05 was reached is to say nothing. The p-value needs to be put into the context of a). prior probability of association and b). the number of tests of association performed. Second, as a corollary, the p-value threshold needs to be set according to these two parameters: the lower the prior probability and/or the higher the number of tests, the lower the p-value needs to be in order to be noteworthy. Third, if we pay attention only to the p-value, and particularly if we fail to understand how to set the threshold for significance, what we get is junk, and, as Ioannidis aptly points out, lies. Fourth, and final, genetic data are tougher to interpret that most of us appreciate. It is possibly even less precise than clinical data we usually discuss on this blog. And we know how uncertain that can be!
So, as sexy as the emerging field of genetics was in the 1980s, I am pretty happy that, after four years in the lab, I decided to go first to medical and then to epidemiology schools. Dealing with conditions where we at least have some notion of the pre-test probabilities makes this quantitative window through which I see healthcare just a little less opaque. And these days, I am even happier about having bucked my classmates' Wall Street trend. But that's a story for another day.
A few days ago, I picked up a link from Michael Pollan's twitter feed to a story that startled me: It was insinuating that all these decades of sequencing human genome to hunt for targets of disease susceptibility have come up nearly empty-handed (with a few well known exceptions). The story recounted a tale of decades of vigorous funding, meteoric career growth and very few results to date. Yet, the scientists involved are not ready to give up on human genes as the primary locus for disease susceptibility. In fact, the strong rescue bias and the fear of losing all that beautiful funding are colluding to generate some creative and complex hypotheses. I came upon one such hypothesis earlier today, following a link from Ed Yong, that British science journalist extraordinaire. But the hypothesis itself, though fascinating, is not what interested me most, no. Care to guess what did? You are correct if you said the p-value.
What grabbed me is the calculation that to identify interactions of significance, the adjusted alpha level has to be set at 10^ -12; that's 0.000000000001! Why is this, and what does it mean in the context of our recent ruminations on the topic of p-value? Well, oddly enough it ties in very nicely with Bayesian thinking and John Ioannides' incendiary assertion of lies in research. How? Follow me.
The hallmark of identifying "significant" relationships (remember that the word "significant" merely means "worthy of noting") in genetic research is searching for statistical associations between the exposure (in this case a mutation in the genetic locus) and the outcome (the specific disease in question). When we start this analysis, we have absolutely no idea which gene(s) mutation(s) is(are) likely to be at play. This means that we have no way of establishing... can you guess what? Yes, that's right, the pre-test probability. We know nothing about the pretest probability. This shotgunning screening of thousands of genes is in essence a random search for matching pairs of mutation and disease. So, why should this matter?
The reason that it matters resides in the simple, albeit overused, coin toss analogy. When you toss a coin, what are your chances of getting heads on any given toss? The odds are 1:1 every time whether you will get heads or tails, meaning that the chance of getting heads is 50% (unless the coin is rigged, in which case we are talking about a Bayesian coin, not the focus of what we are getting at here). So, if you call heads prior to any given attempt, you will be correct half the time. Now, what does this have to do with the question at hand? Well, let's go back to our definition of the p-value: a p-value represents the probability of obtaining the result (in our gene case it is the association with disease) of the magnitude observed or greater under the conditions that no real association exists. So, our customary p-value threshold of 0.05 can be verbalized as follows: "There is a 5% probability that the association of the magnitude observed or greater could happen by chance if there is in reality no association". Now, what happens if we test 20 associations? In other words, what are the chances that one of them will come up "significant" at the p=0.05 level? Yes, there is a 1 in 20 (or 5 in 100) chance that this association will hit our level of significance preset at 0.05. And it follows that the chances of getting a "significant" result grow with more testing.
This is the very reason that gene scientists have asked (and now answered) the question of what level of statistical significance (also called alpha, signified by the p-value) is acceptable when exploring hundreds of thousands of potential associations without any prior clue of what to expect. And that level is a shocking 0.000000000001!
This reinforces a couple of ideas that we have been discussing of late. First, to say simply that the p-value of 0.05 was reached is to say nothing. The p-value needs to be put into the context of a). prior probability of association and b). the number of tests of association performed. Second, as a corollary, the p-value threshold needs to be set according to these two parameters: the lower the prior probability and/or the higher the number of tests, the lower the p-value needs to be in order to be noteworthy. Third, if we pay attention only to the p-value, and particularly if we fail to understand how to set the threshold for significance, what we get is junk, and, as Ioannidis aptly points out, lies. Fourth, and final, genetic data are tougher to interpret that most of us appreciate. It is possibly even less precise than clinical data we usually discuss on this blog. And we know how uncertain that can be!
So, as sexy as the emerging field of genetics was in the 1980s, I am pretty happy that, after four years in the lab, I decided to go first to medical and then to epidemiology schools. Dealing with conditions where we at least have some notion of the pre-test probabilities makes this quantitative window through which I see healthcare just a little less opaque. And these days, I am even happier about having bucked my classmates' Wall Street trend. But that's a story for another day.
Monday, December 13, 2010
Can a "negative" p-value obscure a positive finding?
I am still on my p-value kick, brilliantly fueled by Dr. Steve Goodman's correspondence with me and another paper by him aptly named "A Dirty Dozen: Twelve P-Value Misconceptions". It is definitely worth a read in toto, as I will only focus on some of its more salient parts.
Perhaps the most important point that I have gleaned from my p-value quest is that the word "significance" should be taken quite literally. Here is what Merriam-Webster dictionary says about it:
It is the very meaning in point 2a that the word "significance" was meant to convey in reference to statistical testing. That is "worth noting" or "noteworthy". Nowhere do we find any references to God or truth or dogma. So, the first lesson is to drift away from taking statistical significance as the sign from God that we have discovered the absolute truth, and to focus on the fact that we need to make note of the association. The follow up to the noting action is confirmation or refutation. That is, once identified, this relationship needs to be tested again (and sometimes again and again) before we can say that there may be something to it.
As an aside, how many times have you received comments from peer reviewers saying that what you are showing has already been shown? Yet, in all of our Discussion section we are quite cautious to say that "more research is needed" to confirm what we have seen. So, it seems we -- researchers, editors and reviewers -- just need to get on the same page.
To go on, as the title of the paper states, we are initiated into the 12 common misconceptions about what p-value is not. Here is the table that enumerates all 12 (since the paper is easy to find for no fee, I am assuming that reproducing the table with attribution is not a problem):
Even though some of them seem quite similar, it is worth understanding the degrees of difference, as they provide important insights.
The one I wanted to touch upon further today is Misconception #12, as it dovetails with our prior discussion vis-a-vis environmental risks. But before we do this, it is worth defining the elusive meaning of the p-value once again: "The p-value signifies the probability of obtaining the association (or difference) of the magnitude obtained or one of greater magnitude when in reality there is no association (or difference)". So, let's apply this to an everyday example of smoking and lung cancer risk. Let's say a study shows a 2-fold increase in lung cancer among smokers compared to non-smokers, and the p-value for this association is 0.06. What this really means is that "under conditions of no true association between smoking and lung cancer, there is a 6% or less chance that a study would find a 2-fold or greater increase in cancer associated with smoking". Make sense? Yet, according to the "rules" of statistical significance, we would call this study negative. But is this a true negative? (To the reader of this blog this is obvious, but assure you that, given how cursory our reading of the literature tends to be, and how often I hear my peers discount findings with the careless "But the p-value was not significant", this is a point worth harping on).
The bottom line answer to this is found in the discussion of the bottom line misconception in the Table: "A scientific conclusion or treatment policy should be based on whether or not the p-value is significant". I would like to quote directly from Goodman's paper here, as it really drives home the idiocy of this idea:
Taken together, all these points merely confirm my prior assertion that we need to be a lot more cautious about calling results negative when deciding about potentially risky exposures than about beneficial ones. Similarly, we need to set a much higher bar for all threats to validity in studies designed to look at risky rather than beneficial outcomes (more on this in a future post). These are the principles we should be employing when evaluating environmental exposures. This becomes particularly critical in view of the startling revelations of the genome-wide association experiments findings that our genes determine a very small minority of diseases to which we are subject. This means that the ante has been upped dramatically for environmental exposures as culprits, and demands a much more serious push for the precautionary principle as the foundation for our environmental policy.
Perhaps the most important point that I have gleaned from my p-value quest is that the word "significance" should be taken quite literally. Here is what Merriam-Webster dictionary says about it:
sig·nif·i·cance
noun \sig-ˈni-fi-kən(t)s\
Definition of SIGNIFICANCE
1
a : something that is conveyed as a meaning often obscurely or indirectly
b : the quality of conveying or implying
2
a : the quality of being important : moment
b : the quality of being statistically significant
As an aside, how many times have you received comments from peer reviewers saying that what you are showing has already been shown? Yet, in all of our Discussion section we are quite cautious to say that "more research is needed" to confirm what we have seen. So, it seems we -- researchers, editors and reviewers -- just need to get on the same page.
To go on, as the title of the paper states, we are initiated into the 12 common misconceptions about what p-value is not. Here is the table that enumerates all 12 (since the paper is easy to find for no fee, I am assuming that reproducing the table with attribution is not a problem):
Even though some of them seem quite similar, it is worth understanding the degrees of difference, as they provide important insights.
The one I wanted to touch upon further today is Misconception #12, as it dovetails with our prior discussion vis-a-vis environmental risks. But before we do this, it is worth defining the elusive meaning of the p-value once again: "The p-value signifies the probability of obtaining the association (or difference) of the magnitude obtained or one of greater magnitude when in reality there is no association (or difference)". So, let's apply this to an everyday example of smoking and lung cancer risk. Let's say a study shows a 2-fold increase in lung cancer among smokers compared to non-smokers, and the p-value for this association is 0.06. What this really means is that "under conditions of no true association between smoking and lung cancer, there is a 6% or less chance that a study would find a 2-fold or greater increase in cancer associated with smoking". Make sense? Yet, according to the "rules" of statistical significance, we would call this study negative. But is this a true negative? (To the reader of this blog this is obvious, but assure you that, given how cursory our reading of the literature tends to be, and how often I hear my peers discount findings with the careless "But the p-value was not significant", this is a point worth harping on).
The bottom line answer to this is found in the discussion of the bottom line misconception in the Table: "A scientific conclusion or treatment policy should be based on whether or not the p-value is significant". I would like to quote directly from Goodman's paper here, as it really drives home the idiocy of this idea:
This misconception encompasses all of the others. It is equivalent to saying that the magnitude of effect is not relevant, that only evidence relevant to a scientific conclusion is in the experiment at hand, and that both beliefs and actions flow directly from the statistical results. The evidence from a given study needs to be combined with that from prior work to generate a conclusion. In some instances, a scientifically defensible conclusion might be that the null hypothesis is still probably true even after a significant result, and in other instances, a nonsignificant P value might still lead to a conclusion that a treatment works. This can be done formally only through Bayesian approaches. To justify actions, we must incorporate the seriousness of errors flowing from the actions together with the chance that the conclusions are wrong.When the author advocates Bayesian approaches, he is referring to the idea that a positive result in the setting of a low pre-test probability still has a very low chance of describing a truly positive association. This is better illustrated by the Bayes theorem, which allows us to quantify the result at hand ("posterior probability") to what the bulk of prior evidence and/or thought has indicated about the association ("prior probability"). This implies that the lower our prior probability, the less convinced we can be by a single positive result. As a corollary, the higher our prior probability for an association, the less credence we can put in a single negative result. So, Bayesian approach to evidence, as Goodman indicates here, can merely move us in the direction of either a greater or a lesser doubt about our results, NOT bring us to the truth or falsity.
Taken together, all these points merely confirm my prior assertion that we need to be a lot more cautious about calling results negative when deciding about potentially risky exposures than about beneficial ones. Similarly, we need to set a much higher bar for all threats to validity in studies designed to look at risky rather than beneficial outcomes (more on this in a future post). These are the principles we should be employing when evaluating environmental exposures. This becomes particularly critical in view of the startling revelations of the genome-wide association experiments findings that our genes determine a very small minority of diseases to which we are subject. This means that the ante has been upped dramatically for environmental exposures as culprits, and demands a much more serious push for the precautionary principle as the foundation for our environmental policy.
Wednesday, December 8, 2010
Getting beyond the p-value
Update 12/8/10, 9:30 AM: I just got an e-mail from Steve Goodman, MD, MHS, PhD, from Johns Hopkins about this post. Firstly, my apologies for getting his role at the Annals wrong -- he is the Senior Statistical Editor for the journal, not merely a statistical reviewer. I am happy to report that he added more fuel to the p-value fire, and you are likely to see more posts on this (you are overjoyed, right?). So, thanks to Dr. Goodman for his input this morning!
Yesterday I blogged about our preference to avoid false positive associations at the expense of failing to detect some real associations. The p value conundrum, where the arbitrary statistical significance is set at <0.05 has bothered me for a long time. I finally got curious enough to search out the origins of the p value. Believe it or not, the information was not easy to find. I have at lest 10 biostatistics or epidemiology textbooks on the shelves of my office -- not one of them went into the history of the p value threshold. But Professor Google came to my rescue, and here is what I discovered.
Using a carefully crafted search phrase, I found a discussion forum on WAME, World Association of Medical Editors, which I felt represented a credible source. Here I discovered a treasure trove of information and references to what I was looking for. Specifically, one poster referred to Steven Goodman's work, which I promptly looked up. And by the way, Steven Goodman, as it turns out isa statistical reviewer the Senior Statistical Editor for the Annals of Internal Medicine and a member of WAME. So, I went to this gem in the journal Epidemiology from May 2001, called unpretentiously "Of P-values and Bayes: A Modest Proposal". I have to say that some of the discussion was so in the weeds that even I have to go back and reread it several times to understand what the good Dr. Goodman is talking about. But here are some of the more salient and accessible points.
The author begins by stating his mixed feelings about the p-value:
The point is that we need to get beyond the p-value and develop a more sophisticated, nuanced and critical attitude toward data. Furthermore, regulatory bodies need to get a more nuanced way of communicating scientific data, particularly data evidencing harm, in order not to lose credibility with the people. Most importantly, however, we need to do a better job training researchers on the subtleties of statistical analyses, so that the p-value does not become the ultimate arbiter of the truth.
Yesterday I blogged about our preference to avoid false positive associations at the expense of failing to detect some real associations. The p value conundrum, where the arbitrary statistical significance is set at <0.05 has bothered me for a long time. I finally got curious enough to search out the origins of the p value. Believe it or not, the information was not easy to find. I have at lest 10 biostatistics or epidemiology textbooks on the shelves of my office -- not one of them went into the history of the p value threshold. But Professor Google came to my rescue, and here is what I discovered.
Using a carefully crafted search phrase, I found a discussion forum on WAME, World Association of Medical Editors, which I felt represented a credible source. Here I discovered a treasure trove of information and references to what I was looking for. Specifically, one poster referred to Steven Goodman's work, which I promptly looked up. And by the way, Steven Goodman, as it turns out is
The author begins by stating his mixed feelings about the p-value:
I am delighted to be invited to comment on the use of P-values, but at the same time, it depresses me. Why? So much brainpower, ink, and passion have been expended on this subject for so long, yet plus ca change, plus c'ést le meme chose- the more things change, the more they stay the same. The references on this topic encompass innumerable disciplines, going back almost to the moment that P-values were introduced (by R.A. Fisher in the 1920s). The introduction of hypothesis testing in 1933 precipitated more intense engagement, caused by the subsuming of Fisher's significance test into the hypothesis test machinery.1-9 The discussion has continued ever since. I have been foolish enough to think I could whistle into this hurricane and be heard. 10-12 But we (and I) still use P-values. And when a journal like Epidemiology takes a principled stand against them, 13 epidemiologists who may recognize the limitations of P-values still feel as if they are being forced to walk on one leg. 14So, here we learn that the p-value is something that has been around for 90 years and was brought into being by the father of frequentist statistics R.A. Fisher. And the users are ambivalent about it, to say the least. So, why, Goodman asks, continue to debate the value of the p-value (or its lack)? And here is the reason: publications.
Let me begin with an observation. When epidemiologists informally communicate their results (in talks, meeting presentations, or policy discussions), the balance between biology, methodology, data, and context is often appropriate. There is an emphasis on presenting a coherent epidemiologic or pathophysiologic story, with comparatively little talk of statistical rejection or other related tomfoolery. But this same sensibility is often not reflected in published papers. Here, the structure of presentation is more rigid, and statistical summaries seem to have more power. Within these confines, the narrative flow becomes secondary to the distillation of complex data, and inferences seem to flow from the data almost automatically. It is this automaticity of inference that is most distressing, and for which the elimination of P-values has been attempted as a curative.This is clearly a condemnation of the way we publish: it demands a reductionist approach to the lowest common denominator, in this case the p-value. Much like our modern medical paradigm, the p-value does not get at the real issues:
I and others have discussed the connections between statistics and scientific philosophy elsewhere, 11,12,15-22 so I will cut to the chase here. The root cause of our problem is a philosophy of scientific inference that is supported by the statistical methodology in dominant use. This philosophy might best be described as a form of naïve inductivism,23 a belief that all scientists seeing the same data should come to the same conclusions. By implication, anyone who draws a different conclusion must be doing so for nonscientific reasons. It takes as given the statistical models we impose on data, and treats the estimated parameters of such models as direct mirrors of reality rather than as highly filtered and potentially distorted views. It is a belief that scientific reasoning requires little more than statistical model fitting, or in our case, reporting odds ratios, P-values and the like, to arrive at the truth. [emphasis mine]Here is a sacred scientific cow that is getting tipped! You mean science is not absolute? Well, no, it is not, as the readers of this blog are amply aware. Science at best represents a model of our current understanding of the Universe, it builds upon itself usually in one direction, and it rarely gives an asymptotic approximation of what is really going on. Merely our current understanding of reality, given the tools we have at our disposal. Goodman continues to drive home the naïveté of our inductivist thinking in the following paragraph:
How is this philosophy manifest in research reports? One merely has to look at their organization. Traditionally, the findings of a paper are stated at the beginning of the discussion section. It is as if the finding is something derived directly from the results section. Reasoning and external facts come afterward, if at all. That is, in essence, naïve inductivism. This view of the scientific enterprise is aided and abetted by the P-value in a variety of ways, some obvious, some subtle. The obvious way is in its role in the reject/accept hypothesis test machinery. The more subtle way is in the fact that the P-value is a probability - something absolute, with nothing external needed for its interpretation.In fact the point is that the p-value is exactly NOT absolute. The p-value needs to be judged relative to some other standard of probability, for example the prior probability of an event. And yet what do we do? We worship at the altar of the p-value without giving any thought to its meaning. And this is certainly convenient for those who want to invoke evidence of absence of certain associations, such as toxic exposures and health effects, for example, when the reality simply indicates absence of evidence.
The point is that we need to get beyond the p-value and develop a more sophisticated, nuanced and critical attitude toward data. Furthermore, regulatory bodies need to get a more nuanced way of communicating scientific data, particularly data evidencing harm, in order not to lose credibility with the people. Most importantly, however, we need to do a better job training researchers on the subtleties of statistical analyses, so that the p-value does not become the ultimate arbiter of the truth.
Subscribe to:
Posts (Atom)
