Showing posts with label confounding by indication. Show all posts
Showing posts with label confounding by indication. Show all posts

Tuesday, April 24, 2012

How to avoid the "Titanic effect" in Pharma

Today I was going to tell you the tale of my son's broken wrist (he is fine now, this happened in January, but the insurance issues are fascinating), but I got distracted thinking about another fascinating subject that many do not understand well: confounding by indication. I especially started thinking about it in the context of how decisions and policies are made, and how not having the right data at the right time leads to this "Titanic effect" for a technology. What do I mean by this? Well, let me explain.

Some say the Titanic sank simply because of poor preparation -- not enough life boats, not enough training on the evacuation procedure, in other words "not enough imagination" to plan for a catastrophe. It was derailed in its course by an entirely predictable natural calamity that had not been planned for adequately, even though the risk was obvious in retrospect. Was this just on of those "unintended consequences" that could have been avoided with more clear vision? Perhaps, but the Titanic is, ahem, water under the bridge. But we can focus on some more mundane and current potential missteps and make some guesses.

Let's talk about medical technologies, and drugs in particular. Let us say that there is a new sepsis drug that has been tested among patients with sepsis but without organ failure. This drug appears to prevent organ failure in a fraction of the treated patients, and also reduces mortality by 6%. The only obstacle to widespread use of this drug is its acquisition cost, which is much higher than what the hospital's critical care pharmacist is used to paying for other drugs. Because of this high cost, the drug, despite being on the formulary, gets administered only to those patients who have developed not one, but two organ failures. The savvy pharmacist looks at the outcomes of these patients and, after comparing them to those of the patients who did not receive the drug, concludes that the new sepsis drug, instead of saving lives, actually kills. The P&T committee discusses this, dumps the drug from the formulary and other hospitals follow suit. What's wrong with this picture?

Several fallacies are at work here, including an overly broad inference of causality and bias. But the most important lesson is to do with confounding: because of its apparent expense, the drug has been niched into a population of patients who a). were not the ones that exhibited the evidence of benefit in the trials, and b). have a very high risk of mortality at baseline. So, not only is it not valid to conclude that the drug killed these patients, but it is not even valid to say that the drug does not work -- it may well work in the populations that it was shown to work in, but not in this, much more ill, population. You see the difference? It is like saying that you umbrella failed to keep you dry when you opened it only after you already got soaked.

So confounding by indication is one reason that drugs "fail" -- they are given to people who are by definition not going to do well, and the confirmation bias pushes us to say see, it's expensive and doesn't work. So how do we overcome this phenomenon and make sure that appropriate patients get access to useful technologies? I believe I have a very simple answer: don't squeeze the toothpaste out of the tube if you don't want to have to cram it back in. Huh?

In other words, do what I always advocate: be ready with the relevant data before the train leaves the station, before the cat gets out of the bag, before the horse gets out of the barn. It is very well known that cognitive biases, once established, are difficult to overcome. The pharmacist's first concern is for being able to use his very limited resources efficiently, and to guard from spending his monthly budget on a potentially useless intervention in a single patient only to be left with no resources to care for all of the other patients. Yet many manufacturers at launch send their reps to the pharmacist with two virtually unrelated stories: one about efficacy and the other about the acquisition price and its impact on his budget. When the drug is expensive, the efficacy pales in comparison to the price tag, and the pharmacist has no choice but to restrict the use of the drug, thereby consigning it to failure by confounding by indication. Sound familiar?

Is there a way to avoid this scenario? I think so. It is self-evident that you have to have good data. The surprising thing is that good data are necessary, but not sufficient: the timing of these data is critical as well. It is easier to help people form an opinion where none exists than to change one that is already there. So, to be successful, the manufacturer with a good technology must have a coherent effectiveness and cost-effectiveness proposition right out of the gate. Not only that, but it is imperative to help the clinician understand what patients might benefit from the technology (no, not all patients should be on your drug). This is the kind of a collaboration that will ultimately benefit all stake holders: 1). Appropriate patients will get the opportunity at better outcomes, 2). The pharmacist will understand up front the value proposition and the potential scope of use, and 3). The manufacturer will profit from providing a beneficial service. Isn't this the intent of all this drug development?

If all this seems all too obvious, it is because this is not rocket science. But why, then, do I see so many companies get into trouble with this very scenario? Is it just the case of "best laid plans" or is it a real blind spot that needs to be illuminated? You tell me. Given the investment that goes into drug development, I think it makes sense to approach this gap earnestly, instead of just shuffling the deck chairs on the Titanic.      

 
If you like Healthcare, etc., please consider a donation (button in the right margin) to support development of this content. But just to be clear, it is not tax-deductible, as we do not have a non-profit status. Thank you for your support!

Thursday, January 20, 2011

To guideline or not to guideline, that is the question in... pneumonia?

Addendum 1/20/11, 1:27 PM
I want to add something to this, since I have been reflecting on the data more. It turns out that about 3/4 of all patients had an organism isolated felt to be causative of their pneumonia. Among these patients, over 80% in each group received empiric treatment that covered the pathogen. This means that 4 out of 5 patients in both groups received appropriate antibiotic coverage. What the authors skimmed over briefly is to talk about de-escalation. De-escalation is the guideline recommended strategy which entails reducing the spectrum of treatment after culture results become available to only those antibiotics that cover what has grown out. So, if, say, a patient is being empirically treated for Pseudomonas aeruginosa with double coverage, and the culture grows our MRSA and no Pseudomonas, the two anti-pseudomonal drugs should be stopped immediately. The investigators state that they did apply a de-escalation protocol, and that by day 3 50% and by day 5 75% were essentially de-escalated. The fact that they state this in the Discussion section makes me think that this was inserted in response to a reviewer. It is a pity that they did not include de-escalation in their stratified analysis, as it may be at least somewhat explanatory for the findings. 

I always felt that there was something intangible and intuitive about my assessments of the critically ill for whom I cared. I could not always explain why I thought one particular patient was more ill than the next, but there was that little something that I must have noticed out of the corner of my eye, and if I tried too hard to focus on it, it would disappear like a puff of smoke. Yet, docs make these pre-conscious assessments all the time. And though these hints drive treatment choices, they are distinctly difficult to quantify scientifically.

A new paper that was just published in The Lancet Infectious Diseases online is a great illustration of what happens when our analyses fail to account for these intuitions. The phenomenon is referred to as "confounding by indication", and it is the perennial plague of observational clinical research. Just to summarize, the study was an observational study of guideline implementation for the treatment of healthcare-associated pneumonia among ICU patients. The central guideline was that for the choice of empiric antibiotics selection. The initial choice of antibiotics, even before the definitive results of cultures are available, is based on the clinician's best guess at what organism(s) may be causing the pneumonia. Among these severely ill patients, the risk of having a bug that is resistant to many antibiotics is higher than for patients who come from the community with pneumonia, and this propensity drives the recommendation for a broader antibiotic coverage for these cases. It has been shown by us and many others that missing this initial opportunity to cover the bug(s) adequately subjects patients to a doubling or even trebling of the risk of death, regardless of whether the coverage is broadened later to include the culprit organism(s).

Back to the study. The four academic medical center that participated in it enrolled 303 eligible patients, of whom 129 were treated with antibiotic combinations that comported with the guideline recommendations (guideline compliant treatment) and 174 received other combinations that did not fit the guideline recommendations (guideline non-compliant). To their surprise, the investigators discovered that 28-day survival was actually higher in the non-compliant group than in the compliant one. And even after doing a great job of adjusting for many potential factors that made the groups different, this paradoxical disparity persisted, with an overall near-doubling in the hazard of death at 28 days in the compliant as opposed to the noncompliant group. Now, this is a fine how-do-you-do! So, does this mean that the guideline is actually killing people by advocating broader coverage? Well, not so fast.

First, I have to acknowledge that I may be engaging in rescue bias right now. Having said this, taking biological plausibility into account, the findings are very likely explained by confounding by indication. Namely, the docs who choose, say, dual rather than single therapy against gram-negative bacteria may be pre-consciously incorporating some intangible patient data into their choices, data that are not well represented by either laboratory values or disease severity scoring systems. I know this is a bit "soft" and maybe even "touchy-feely", but ask any doc, and s/he will confirm this phenomenon.

On the other hand, to be fair and balanced, I do have to agree that there may be other explanations. These include the possibility that our guideline recommendations, never really prospectively validated, may be wrong. Perhaps there is something about the untoward effects of these broad spectrum regimens that is at play. Maybe it is as simple as the "no free lunch" principle, and that even in the situation of covering appropriately broadly, introducing additional drugs increases not only their benefits, but also the risks associated with them. Finally, I have to acknowledge the possibility that we just have no clue what any of this means because our understanding of how antibiotics work in the setting of these types of pneumonia is flawed.

Now, let's put all of this in the context of our multiple discussions about data and knowledge on this web site. Several factors suggest that my initial explanation is correct. The bulk of the evidence points to the fact that skimpy early coverage increases the risk of death. Also, over a century of understanding and the durability of the germ theory imply that antibiotics are important in treating serious bacterial infections. So, the pre-test probability of the validity of the finding in the paper is pretty low. This is not to say that the study should not inject caution and self-examination into how we treat severe pneumonia; it absolutely should! This is also a place where we definitely need well designed interventional studies to confirm (or debunk) what we think we know to be true. In the meantime, as we often intone on this blog, let us not throw the baby out with the bath water.

Disclosure: I have done a lot of work in this area, so I have a potential intellectual COI with the study. Also, at least some of my research has been funded by the manufacturers of some of the antibiotics included in the guidelines.

Wednesday, January 12, 2011

Reviewing medical literature, part 3: Threats to validity

You have heard this a thousand times: no study is perfect. But what does this mean? In order to be explicit about why a certain study is not perfect, we need to be able to name the flaws. And let's face it: some studies are so flawed that there is no reason to bother with them, either as a reviewer or as an end-user of the information. But again, we need to identify these nails before we can hammer them into a study's coffin. It is the authors' responsibility to include a Limitations paragraph somewhere in the Discussion section, in which they lay out all of the threats to validity and offer educated guesses as to the importance of these threats and how they may be impacting the findings. I personally will not accept a paper that does not present a coherent Limitations paragraph. However, reviewers are not always, as, shall we say, hard assed about this as I am, and that is when the reader is on her own. Let us be clear: even if the Limitations paragraph is included, the authors do not always do a complete job (and this probably includes me, as I do not always think of all the possible limitations of my work). So, as in everything, caveat emptor! Let us start to become educated consumers.

There are four major threats to validity that fit into two broad categories. They are:
A. Internal validity
  1. Bias
  2. Confounding/interaction
  3. Mismeasurement or misclassification
B. External validity
  4. Generalizability
Internal validity refers to whether the study is examining what it purports to be examining, while external validity, synonymous with generalizability, gives us an idea about how broadly the results are applicable. Let us define and delve into each threat more deeply.

Bias is defined as "any systematic error in the design, conduct or analysis of a study that results in a mistaken estimate of an exposure's effect on the risk of disease" (the reference for this is Schlesselman JJ, as cited in Gordis L, Epidemiology, 3rd edition, page 238). I think of bias as something that artificially makes the exposure and the outcome either occur together or apart more frequently than they should. For example, the INTERPHONE study has been criticized for its biased design, in that it defined exposure as at least one cellular phone call every week. Now enrolling such light users can really result in such a small exposure as not to be able to detect any increase in adverse events. This is an example of a selection bias, by far the most common form that bias takes. Another example of a frequent bias is encountered in retrospective case-control studies where people are asked to recall distant exposures. Take for example middle-aged women with breast cancer who are asked to recall their diets when they were in college. Now, ask the same of similar women without breast cancer. What you are likely to get is the effect, absent in women without cancer, of seeking an explanation for the cancer that expresses itself in a bias in what women with cancer recall eating in their youth. So, a bias in the design can make the association seem either stronger or weaker than it is in reality.

I want to skip over confounding and interaction at the moment, as these threats deserve a post of their own, which is forthcoming. Suffice it to say here that a confounder is a factor related to both, the exposure and the outcome. An interaction is also referred to as effect modification or effect heterogeneity. This means that there may be population characteristics that alter the response to the exposure of interest. Confounders and effect modifiers are probably the trickiest concepts to grasp. So, stay tuned for a discussion of those.

For now, let us move on to measurement error and misclassification. Measurement error, resulting in misclassification, can happen at any step of the way: it can be in the primary exposure, a confounder, or the outcome of interest. I run into this problem all the time in my research. Since I rely on administrative coding for a lot of the data that I use, I am virtually certain that the codes routinely misclassify some of the exposures and confounders that I deal with. Take Clostridium difficile as an example. There is an ICD-9 code to identify it in administrative databases. However, we know from multiple studies that it is not all that sensitive or all that specific; it is merely good enough, particularly for making observations over time. But even for laboratory values there is a certain potential for measurement error, though we seem to think that lab results are sacred and immune to mistakes. And need I say more about other types of medical testing? Anyhow, the possibility of error and misclassification is ubiquitous. What needs to be determined by the investigator and the reader alike is the probability of that error. If the probability is high, one needs to understand whether it is a systematic error (for example, a coder always more likely than not to include C. diff as a diagnosis) or a random one (a coder is just as likely to include as not to include a C diff diagnosis). And while a systematic error may result in either a stronger or a weaker association between the exposure and the outcome, a random, or non-differential, misclassification will virtually always reduce the strength of this association.

And finally, generalizability is a concept that helps the reader understand what population the results may be applicable to. In other words, will the data be applied strictly to the population represented in the study? If so, is it because there are biological reasons to think that the results would be different in a different population? And if so, is it simply the magnitude of the association that can be expected to be different or is it possible that even the direction could change? In other words, could something found to be beneficial in one population be either less beneficial or even more harmful in another? The last question is the reason that we perseverate on this idea of generalizability. Typically, a regulatory RCT is much less likely to give us adequate generalizability than a well designed cohort study, for example.

Well, these are the threats to validity in a nutshell. In the next post we will explore much more fully the concepts of confounding and interaction and how to deal with them either at the study design or study analysis stage.            

Tuesday, March 9, 2010

Lies, big lies and... epidemiology?

Now, as you know, I am a big fan of epidemiology. I do not believe that a randomized controlled trial is the be-all-and-end-all in evidence generation, and a well done observational study can add to our reservoir of knowledge much more efficiently. Of course, I, as many others, acknowledge certain limitations of epidemiologic design. However, many of them can be overcome with careful design and analysis.

I have to confess, though, that over the last week I have bumped into two news stories that have made me cringe. The first, reported a couple of days ago and based on a Kaiser study, showed that people
drinking at least four cups of coffee a day were 18 percent less likely to be admitted with a heart rhythm disturbance than those who drank no coffee at all.
So, great, the public may take away the message that drinking coffee prevents a-fib. Well, I have to say that the reporting of this was measured and tried to avoid this unfortunate inference of causality. Yet, it left enough room to imply that yes, perhaps there is a causal link. So, what's wrong with that?

Bear with me while I bring in the second example of a study that bugged me this week, based on the Women's Health Initiative. You may recall that the WHI is the large NIH-sponsored study that a few years ago turned hormone replacement therapy on its head. The study had a randomized component and an observational component. So, the newest analysis shows that women
who drank the equivalent of one to two drinks a day -- be it beer, wine, or liquor -- were 30% less likely than non-drinkers to become overweight or obese.
 Do you see the similarities? So why am I bothered? To me this is the classic case of a high potential for confounding by indication. What's that you say? That is a situation in which a subject that has the exposure in question (in these two cases coffee and alcohol) is inherently different from one who does not, and this difference is what determines the probability of the exposure itself. Why should this present a problem in a study where the authors carefully adjusted for confounding, which is true for both of the studies? It is a problem because the kind of confounding that this represents is impossible to tease out without real-time attention to the subject.

Here is how it would work in the case of coffee study. Say I am a person with paroxysmal (occasional) a-fib, and I have noticed that if I drink so many cups of coffee per day, I get into brief episodes of palpitations. Not enough to send me to the doctor's office or the hospital, but enough to start thinking about cutting out caffeine. So, I stop drinking caffeine, and continue with my baseline frequency of a-fib attacks. You see the problem? Is it possible (or even probable) that those people who drink four or more cups of coffee per day somehow have an inherently higher threshold for slipping into their a-fib than those who do not? And if the answer is "yes", then the four cups become a marker for someone who can tolerate them, rather than the cure for a-fib. You can construct a similar explanation with the two drinks and weight.

So, while I love epidemiology and its methods, I am wary of hanging my hat on associations that may likely be explained by confounding by indication. And although the stories were reported with many caveats, human nature may prevent us from hearing the nuance. It is clear that in both these instances the burden of proof is on the researchers to show me that I am wrong.