Showing posts with label hte. Show all posts
Showing posts with label hte. Show all posts

Thursday, May 24, 2012

Septic shock doesn't need my rescue bias

Rescue bias and auxiliary hypothesis bias are tempting distractions when it comes to dealing with data that are counter to our preconceived notions. In general, the role of our cognitive biases in all kinds of discourse, including clinical and scientific, is under appreciated. For this reason I devoted fully two chapters of my book to cognitive biases.

Which makes it even more frightening that, as I read this NEJM paper on the failed drotrecogin PROWESS-SHOCK study, I am still looking for reasons why it failed, other than that just doesn't work. After so many years of trials and tribulations with Xigris, and hopes that this lone therapy ever to have been approved for treating this deadly disease, I am having a hard time letting go.

Yet the rial was meticulous, as all trials designed and overseen by B. Taylor Thompson are -- he is just a brilliant clinical researcher that we all need to learn from. The design considered all the right issues a priori -- multiple interim looks at the data, a possible need to increase the power, heterogeneous treatment effect, and others. Although there were a few numerical imbalances between the treatment and the placebo groups -- blood cultures were positive 4% more frequently and the offending pathogen overall was 5% more likely to be identified in the Xigris than in the placebo group -- I have to accept that these are not the reasons for the observed failure to improve either the 28-day or the 90-day mortality.

I have seen and even blogged about these data before -- the press release had some of them, but, more importantly, Taylor presented them at the Society of Critical Care Medicine meeting back in January. So, this is nothing new. Yet reading all of the details in the pages of the NEJM brings back that pre-conscious cognitive wall that I have to climb over in order to get to reality. And the reality is that the drug just did not work.

But the reality is also that sepsis patients have a better chance of surviving this deadly assault than they did 10 years ago. In the original PROWESS study 28-day mortality in the placebo group was 31%, and these were all kinds of sepsis patients. In the current study, among most severe sepsis patients, those with septic shock, the 28-day mortality was on the order of 25% in both arms -- that's a staggering reduction in this most ill septic population! We really need to appreciate this. In part this trend may be due to some evolution of the disease-host interaction. But in part it has to be because of the concerted effort to understand sepsis better, to study various treatment options, and, most importantly, to implement these learnings at the bedside. In no small part we owe these advances in sepsis care to the short life of drotrecogin alpha.

The data are clearly anticlimactic, and for this reason the current paper has no sex appeal -- just look at the (lack of) press coverage about it. Yet, this is a clear example of countering the traditional publication bias: Here is a manufacturer-funded negative study published in a premier peer-reviewed journal. I realize that it was always a high-profile undertaking, but let's pause for a moment to enjoy this small victory for those who have railed against the spread of biased information in the medical literature.

I don't know if and when another therapy will emerge courageous enough to brave the rough waters of sepsis and septic shock. But one thing is clear: Xigris has served out its purpose, and no amount of rescue bias from me will save it. Rest in peace.    

If you like Healthcare, etc., please consider a donation (button in the right margin) to support development of this content. But just to be clear, it is not tax-deductible, as we do not have a non-profit status. Thank you for your support!

Thursday, February 16, 2012

Medicine: The art of applied science

I read this NPR article this morning and had to do a post in response. The gist is that the military is turning to what we might call the less conventional (for us in the West) medical modalities to deal with the injuries sustained by the current crop of vets. Instead of getting them hooked on pain meds for life (we saw plenty of this in the VAs in the '80s and '90s among Vietnam vets), they are turning to stuff like massage and acupuncture. And, predictably, it is stirring up controversy.

The story that is told is of a Sgt. Rick Remalia who fractured his back and pelvis in Afghanistan:
Remalia broke his back, hip and pelvis during a rollover caused by a pair of rocket-propelled grenades in Afghanistan. He still walks with a cane and suffers from mild traumatic brain injury. Pain is an everyday occurrence, which is where the needles come in.
And lately he has been receiving acupuncture treatments, with this result:
"I've had a lot of treatment, and this is the first treatment that I've had where I've been like, OK, wow, I've actually seen a really big difference," he says.
And incidentally, her gets these treatments from a military physician, who, herself a skeptic, admits to perceiving a personal benefit from her own exposure to it:
"I actually had a demonstration of acupuncture on me, and I'm not a spring chicken," she says, "and it didn't make me 16 again, but it certainly did make me feel better than I had, so I figured, hey ... let's give it a shot with our soldiers here."
So, all good so far, right? Well, Harriet Hall is quoted in the same article, and to her this falls right into what she likes to call "quack-ademic" medicine. She says,
"The military has led the way on trauma care and things like that, but the idea that putting needles in somebody's ear is going to substitute for things like morphine is just ridiculous," Hall says.
Now, as you know, I have had some debates with the SBM crowd in the past, and as it turns out, we agree on the science more than we disagree. However, I am thinking that this argument is not about science, but about politics.

I am well aware that a group of anecdotes does not amount to science. And I am also well aware that what we are hearing here are anecdotes. But here is the thing: when your kid tells you that she likes chocolate ice cream better than vanilla, do you ask for evidence that chocolate is better than vanilla at the population level? No, that's absurd! OK, you say, but this is a strawman: nobody is going for a claim of superiority of chocolate ice cream over vanilla. That is true, but is this about the science or about being able to make a claim? If my kid likes chocolate, why not let her have that when ice cream is on the menu? If acupuncture seems to provide some relief to Sgt. Remalia, why not let him have that relief? After all, whose opinion about what works counts in this individual example, ours or the patient's? And if the ethics of using placebo are the concern, there is nothing wrong with letting him know that in large clinical trials the evidence is equivocal, which means that it may work for some and not for others. In fact, this might be a good disclaimer to make before commencing any treatment, one with the right to claims and one without.

Another argument is that there is no way that insurance (or our taxes) should pay for this unproven treatment. Still about science? Do any of you want to stand up and tell Sgt. Remalia, who fought for our freedom, that we will not pay for the only thing that seems to help him, that is pretty cheap and safe and that has very few, if any, long-term adverse effects, in stark contrast to pain killers? Yes, I understand that this is not science, but is there no room for humanism in the practice of medicine? After all we have throaty debates as to whether or not it is ethical to deny a $100,000 payment for a treatment that, on average, prolongs life by 2 weeks. Surely, denying Sgt. Remalia access to this relief would diminish our humanity. And what about the costs of treating addiction to pain killers?

So, here are my points:
1. I completely agree that that acupuncture "works" for Sgt. Remalia, does not mean that "acupuncture works" in the scientific sense. It may or may not work; furthermore, our current models of the universe do not allow us to have an adequate mechanistic explanation. But that is not the point -- it works for this young man whose life will never be the same because he signed up to defend his country. To this extent his "claim" has all kinds of internal validity.
2. Making claims is subject to legal and regulatory frameworks that have very little to do with science. I have done much blogging on clinical vs. statistical considerations in clinical research that feeds regulatory approvals and hence claims, and I remain of the opinion that a lot of the acceptable claims are specious. I know, I know, this is a "tu quoque" argument, but if we are talking about the goose and the gander, well...
3. Whether or not a treatment should be paid for is more prone to political than evidence-based decisions. Given that most medicines work in a minority of patients, and none comes without adverse events, the extent of which remains largely unknown because of our negligence to build real regulatory systems to quantify them, we are spending a lot of dollars on stuff that does not work at the individual level.

Medicine has to be part science and part art; in fact the art is in how and when to apply the science. That latter portion must be about humanism.
   

Thursday, February 24, 2011

New treatments: What benefits at what costs

Yesterday brought quite a bit of press coverage to a small biotechnology company in Cambridge called Vertex. All this attention was spurned by their gene therapy trial results in cystic fibrosis. The treatment, aimed at a genetic mutation present in about 4% of all CF sufferers, was able to improve the volume that a patient can force out of his lungs in 1 second by over 10%, from about 65% to 75%. Matthew Herper of Forbes on his blog, while being duly impressed by the results, also cautioned that the annual price tag for this medicine is likely to reach $250,000 per patient. So, what does all of this mean in the context of our ongoing national discussion about the value of therapies? Well, let's break things down a bit.

First, let's talk about CF. This is a genetic disorder that essentially makes mucus very sticky. Among its many effects, in its most familiar manifestation this mucus plugs up the airways making it difficult to breathe and predisposing the person to frequent and serious lung infections. When I was a resident back in the early '90s, I remember a devastating case of a young man in his late teens with CF whom we all knew so well from his frequent admissions for exacerbations. Though he was pretty high on the lung transplant list, he ended up succumbing to a devastating pneumonia in our ICU, leaving behind a devoted sister who had been fortunate enough to benefit from a transplant several years earlier. This was a typical course in those days: a brief life punctuated by frequent exacerbations, hospitalizations, antibiotics, gastrointestinal complications, and early death in the second or at best third decade of life with very little hope of procreation. Over the last 20 years things have changed dramatically in the treatment of CF: fewer exacerbations, much lengthened life expectancy and a good chance of having children. Yet we cannot attribute most of these changes to dramatic new breakthrough therapies. To be sure, while there have been tweaks to how we give antibiotics and how pancreatic enzymes are administered to replace the digestive enzymes that the pancreas in CF is unable to produce, most of the progress can be attributed to the increased attention to detail and the advent of almost ruthless care coordination at specialty centers. As a Fellow in the '90s I participated in a clinic where CF patients were transitioning from care by pediatric Pulmonologists to that by adult doctors. The CF specialist running this clinic did not only know all of his patients and their family members by names, but was available 24/7 to them and to his staff for consultation. This is the kind of dedication and vigilance necessary to improve the outcomes in CF.

Now, let's talk about the lesion addressed in the Vertex trial. The type of chronic lung disease caused by CF is called "obstructive." Simply put, it makes exhaling the air in the lungs difficult to do. On lung testing one manifestation of obstruction is the amount of air one is able to force out of his lungs in the first second of the effort, and this is called the FEV1, or forced expiratory volume in 1 second. Another important measure of the degree of obstruction is the amount of air that this volume expired in 1 second represents as a proportion of all of the air in the lungs that can be expired, known as the FVC or forced vital capacity. We say that if the FEV1/FVC ratio is under 75%, then obstruction is present. The size of FEV1 helps us understand how bad the obstruction is.

With this as a background, the primary outcome in many obstructive lung disease trials is the improvement in the FEV1. In the specific trial discussed, the average starting FEV1 in the intervention group was about 65%, which falls in the mild-to-moderate category of obstruction. What this means in terms of symptoms can vary widely. The 10% absolute improvement seen in the intervention group resulted in the average FEV1 of about 75% after treatment, definitely representing fairly mild obstruction (generally FEV1 over 80% is considered to be in the normal range). And this truly is impressive. However, equally interesting is the information that is not in the press coverage, largely based on press releases and sound bites from company executives, since the peer reviewed study is not available at this point. We are not told, but led to assume that, the control group started out on average with a similar deficit in lung function. We are informed that the treatment patients were 55% less likely than placebo patients to have an exacerbation of their disease, yet we do not know what the absolute numbers are; that is we are not told what proportion in each group had an exacerbation, how frequently or how severely. So, this 55%, in the absence of context, while an attention grabber, is not a substantive number. Herper does tell us that there was a remarkable difference in the weight gain (a desirable outcome in the CF population), on average 6.8 lb in the treatment vs. 0.9 lb in the placebo group. This is truly impressive, though it would be even more so if I knew that the trial was double blind, a piece of information I did not notice in any of the reports. Some of the reports have also alluded to symptomatic improvement in shortness of breath, though nowhere did I see this quantified.

The most important piece of data, however, is conspicuously absent from all the stories. What is the proportion of patients who responded to therapy? Why is this important? Well, we know that far from everyone responds to every treatment that they ostensibly qualify for; this is referred to as the heterogeneous treatment effect, or HTE. It is very likely that the 10% improvement in the FEV1 represents at once an inflated estimate referent to those non- or under-responders and a muted one for those patients with a terrific response. The question of a minimal clinically significant change in the FEV1 has haunted the lung trials community for a long time now. Yet, without setting some threshold for a minimum FEV1 improvement that correlates with a meaningful improvement in symptoms, one cannot quantify how well the drug works and hence articulate its value. This is crucial when trying to justify the ostensibly exorbitant price tag anticipated for this drug. How many patients will we need to treat in order to have one of them respond meaningfully with an improvement not just in a laboratory number, but also in their lives? If this targeted drug produces a desirable response even in 50% of all patients with the specific mutation it targets, then it means that we need to spend $500,000 annually to obtain a meaningful improvement in symptoms in one CF patient. But what if it only works this way in 20%? Then we will need to treat 5 patients with this drug to obtain 1 meaningful response at the price of $1.25 million annually. This becomes a bit more daunting, particularly given that the costs will have to be covered through some kind of public or pooled funds and given that this is one of many therapies in the pipeline likely to come with a similar conundrum.

I am not implying that improving a single life is not worth $1.25 million annually. In fact, it may well be a bargain. My point is that these are the serious discussions we need to have as a society, so that when the time comes to make these choices, the discussion will not be subverted by a few loud voices sensationalizing "death panel" slogans. Manufacturers need to know that they should disclose full data, not just selective tidbits that highlight benefits only, but also those difficult pieces of information that shed light on their costs. On our part, we need to understand the gargantuan effort and resources these companies expend to tame these elusive wild therapies that hold so much more promise in the abstract than they end up embodying.

We tread a fine line here. Information and how we assimilate it are the next frontier for cogent decision making. We need to get educated about this now because this train is leaving the station regardless of how we feel about it.                            

Thursday, January 13, 2011

Reviewing medical literature part 3 continued: threats to validity

As promised, today we talk about confounding and interaction.

A confounder is a factor related to both, the exposure and the outcome. Take for example the relationship between alcohol and head and neck cancer. While we know that heavy alcohol consumption is associated with a heightened risk of head and neck cancer, we also know that people who consume a lot of alcohol are also more likely to be smokers, and smoking in turn raises the risk of H&N CA. So, in this case smoking is a confounder of the relationship between alcohol consumption and the development of H&N CA. It is virtually impossible to get rid of all confounding completely in any study design, save for possibly in a well designed RCT, where randomization presumably assures equal distribution of all characteristics; and even there you need an element of luck. In observational studies our only hope to deal with confounding is through statistical manipulation we call "adjustment", as it is virtually impossible to chase it away any other way. And in the end we still sigh and admit to the possibility of residual confounding. Nevertheless, going through the exercise is still necessary in order to get closer to the true association of the main exposure and the outcome of interest.

There are multiple ways of dealing with the confounding conundrum. The techniques used are matching, stratification, regression modeling, propensity scoring and instrumental variables. By far the most commonly used method is regression modeling. This is a rather complex computation that requires much forethought (in other words, "Professional driver on a closed circuit; don't try this at home"). The frustrating part is that, just because the investigators did the regression, does not mean that they did it right. Yet word limits for journal articles often preclude authors from giving enough detail on what they did. At the very least they should tell you what kind of a regression they ran and how they chose the terms that went into it. Regression modeling relies on all kinds of assumptions about the data, and it is my personal belief, though I have no solid evidence to prove it, that these assumptions are not always met.

And here are the specific commonly encountered types of regressions and when each should be used:
1. Linear regression. This is a computation used for outcomes that are continuous variables (i.e., variables represented by a continuum of numbers, like age, for example). This technique's main assumption is that the exposure and outcome are related to each other in a linear fashion. The resulting beta coefficient is the slope of this relationship if it is graphed.
2. Logistic regression. This is done when the outcome variable is categorical (i.e., one of two or more categories, like gender, for example, or death). The result of a logistic regression is an adjusted odds ratio (OR). It is interpreted as an increase or a decrease in the odds of the outcome occurring due to the presence of the main exposure. Thus, a OR of 0.66 means that there is a 34% reduction in the odds (used interchangeably with risk, though this is not quite accurate) of the outcome due to the presence of the exposure. Conversely, a OR of 1.34 means the opposite, or a 34% increase in the odds of the outcome if the exposure is present.
3. Cox proportional hazards. This is a common type of a model developed for a time to event, also known as "survival analysis" (even if not done for survival per se as the outcome). The resulting value is a hazard ratio (HR). For example, if we are talking about a healthcare-associated infection's impact on the risk of remaining in the hospital longer, a HR of, say, 1.8 means that a HAI increases the risk of being in the hospital by 80% at any time during the hospitalization. To me this tends to be the most problematic technique in terms of assumptions, as it requires that the risk of an even stays constant throughout the time frame of the analysis, and how often does this hold true? For this reason the investigators should be explicit about whether or not they tested for the assumption of proportional hazards and whether this was met.

Let's now touch upon the other techniques that help us to unravel confounding. Matching is just that: it is a process of matching subjects with the primary exposure to those without in a cohort study or subjects with the outcome to those without in a case-control study, based on certain characteristics, such as age, gender, comorbidities, disease severity, etc.; you get the picture. By its nature, matching reduces the amount of analyzable data, and thus reduces the power of the study. So, is is most efficiently applied in a case-control setting, where it actually improves the efficiency of enrollment.

Stratification is the next technique. The word "stratum" means "layer", and stratification refers to describing what happens to the layers of the population of interest with and without the confounding characteristic. In the above example of smoking confounding the alcohol and H&N CA relationship, stratifying the analyses by smoking (comparing the H&N CA rates among drinkers and non-drinkers in the smoking group separately from the non-smoking group) can divorce the impact of the main exposure from that of the confounder on the outcome. This method has some distinct intuitive appeal, though its cognitive effectiveness and efficiency dwindle the more strata we need to examine.

Propensity scoring is gaining popularity as an adjustment method in the medical literature. A propensity score is essentially a number, usually derived from a regression analysis, giving the propensity of each subject for a particular exposure. So, in terms of smoking, we can create a propensity score based on other common characteristics that predict smoking. Interestingly, some of these characteristics will be present also in people who are not smokers, yielding a similar propensity score in the absence of this exposure. Matching smokers to non-smokers based on the propensity score and examining their respective outcomes allows us to understand the independent impact of smoking on, say, the development of coronary artery disease. As in regression modeling, the devil is in the details. Some studies have indicated that most papers that employ propensity scoring as the adjustment method do not do this correctly. So, again, questions need to be asked and details of the technique elicited. There is just no shortcut to statistics.

Finally, a couple of words about instrumental variables. This method comes to us from econometrics. An instrumental variable is one that is related to the exposure but not the outcome. One of the most famous uses of this method was published by a fellow you may have heard of, Mark McClellan, where he looked at the proximity to a cardiac intervention center as the instrumental variable in the outcomes of acute coronary events. Essentially, he argued, the randomness of whether or not you are close to a center randomizes you to the type of treatment you get. Incidentally, in this study he showed that invasive interventions were responsible for a very small fraction of the long-term outcomes of heart attacks. I have not seen this method used that much in the literature I read or review, but am intrigued by its potential.

And now, to finish out this post, let's talk about interaction. "Interaction" is a term mostly used by statisticians to describe what epidemiologists call "effect modification" or "effect heterogeneity". It is just what the name implies: there may be certain secondary exposures that either potentiate or diminish the impact of the main exposure of interest on the outcome. Take the triad of smoking, asbestos and lung cancer. We know that the risk of lung cancer among smokers who are also exposed to asbestos is far higher than among those who have not been exposed to asbestos. Thus, asbestos modifies the effect of smoking on lung cancer. So, to analyze those smokers exposed to asbestos together with those who were not will result in an inaccurate measure of the association of smoking with lung cancer. More importantly, it will fail to recognize this very important potentiator of tobacco's carcinogenic activity. To deal with this, we need to be aware of the potentially interacting exposures, and either stratify our analyses based on the effect modifier or work the interaction term (usually constructed as a product of the two exposures, in out case smoking and asbestos) into the regression modeling. In my experience as a peer reviewer, interactions are rarely explored adequately. In fact, I am not even sure that some investigators understand the importance of recognizing this phenomenon. Yet, the entire idea of heterogeneous treatment effect (HTE) and our pathetic lack of understanding of its impact on our current bleak therapeutic landscape, is the result of this very lack of awareness. The future of medicine truly hinges on understanding interaction. Literally. Seriously. OK, at least in part.

In the next installment(s) of the series we will start tackling study analyses. Thanks for sticking with me.        

Thursday, November 18, 2010

Patient empowerment: A tango worth dancing

I like to poke fun at real estate agents (please, forgive me if you are one, it is all in good fun). My experience has been that, despite what I describe as my preferences, they always end up showing me what they have, even if it does not bear the remotest resemblance to what I need. This holds true for politicians, with this cardinal rule: always answer the question you want to answer, rather than the one being asked. Well, now that I think about it, it is also true for modern medicine. Here is how.

This morning I attended the annual fundraising breakfast for an organization that started locally, but is spreading nationally and even internationally. The group is MotherWoman, and it arose from a realization that women's post-partum needs were not being met. At the regional level, women suffering post-partum depression had no resources available to them, and the local clinicians were not clued in to the problem to the point where they could offer targeted help. And although the organization has grown and matured around issues of parenting in general, they are still known for their robust community outreach to help manage PPD.

So, every year they have a fundraising breakfast, and every year I set a limit for the sum I will write in that box on the check, and every year, after hearing inspirational stories of real women and families, I go over this limit. This year was no different. The speakers were fantastic -- passionate, committed, authentic. Sharon Lerner, the journalist and author of The War on Moms, was a featured speaker. But the most touching of all was the talk by a young mother, followed by one by her husband, about their family's struggle with profound dark and shattering PPD. Aside form the fact that there was not a dry eye in the audience, the story had yet another effect, on me specifically: it clarified for me the importance of empowering mothers, fathers, families and all patients to advocate for themselves. The story was one of missed opportunities to listen to, to connect with, to help someone whose life was suddenly too heavy to carry. Instead of offering individual care and support, the woman was prescribed a one-size fit-all fix, which she was not willing to accept. Yet the healthcare system failed to offer a viable alternative. Just as in the case of the real estate agent and the politician, the property shown and the question answered were not this woman's. Instead, they were generic solutions driven by our increasingly widget-oriented approach to healthcare.

She and her family did eventually get the help they needed, not the least of which came from MotherWoman, and she is well on her way out of the jungle of PPD. And just as this family made clear, over and over again I hear people tell me that they do not blame their healthcare providers, they do not consider them evil, and they actually appreciate the efforts made by them on their behalf. But these efforts fall short when we are forced to measure individuals with a population-based yardstick, leaving both the patients and the physicians frustrated.

The way we derive evidence for our practices needs to change. There are many underutilized tools already available to help us understand inter-individual variations, and there are many more that need to be developed and used. But as in any relationship, patients need to learn to speak up and make their needs heard. We need to develop a common vocabulary so that clinicians can become aware of our needs and be sensitive to them.

As I tell my children when they bicker and come to me complaining about each other, it takes two to tango. Physicians are busier than ever, we live in the era of the incredible shrinking appointment, and the proliferation of gadgetry and other pressures on the doctors' attention and time is crushing. Just as we are not afraid to let the real estate agent know to change course, just as we demand that politicians answer our questions, so in this the patients need to learn to advocate for themselves. This is surely a culture shift. Yet this is one dance that is worth performing well.

Update 11/19/2010:
I encourage you to go here to read Sharon Lerner's remarks.          

Wednesday, October 20, 2010

What-based medicine?

Evidence-based medicine started out with Archie Cochrane in 1972 and really hit big time with David Sackett, David Eddy and others in the '90s. While initially meant to ensure equitable use of a limited resource through evidence-based decisions, EBM now drives reimbursement, quality metrics, MD grades, and opinion. As with market success of a drug thrust into a generally healthy population, where safety signals are bound to overshadow any measurable benefit (think Vioxx), there are likewise unintended consequences to such widespread adoption, or we might even say bastardization, of EBM.

Here is how I see it. EBM has given the opportunity to develop markets, vast markets of diseases that would not have existed or been given the time of day without it. "Diseases" such as pre-diabetes, pre-hypercholesterolemia, pre-hypertension, pre-osteoporosis, all have been brought to our consciousness through the engine of EBM. The "E" in EBM is what is so poorly defined, and your E may not be my E, as keenly illustrated in several recent editorials in both the popular and scientific journals. Where we have arrived is at the lowest common denominator of E: statistical significance. The argument goes that a response statistically different from that to a placebo constitutes evidence of efficacy. Indeed, this is what our regulatory agencies rely on in their deliberations of market approvals. Once on the market, the Wild West of the marketing and prescribing giant is rarely held to a higher standard than the proof obtained in the laboratory of a clinical trial. Yet, those of us familiar with the methods realize that a randomized controlled clinical trial, no matter how well executed, is but the beginning of gathering evidence of effect. Just look at a typical trial's list of inclusion and exclusion criteria. It is no wonder that the worn meme of "not my patient" is so prevalent among docs being handed this new "evidence".

Several mechanisms have been developed to overcome this limitation of pre-approval trials. One is a meta-analysis, where many similarly designed trials contribute the data to arrive at a point estimate for the effect in a much larger population. The issue with this is that it in no way helps us understand what the intervention does in the real world to a real individual patient. In fact, a meta-analysis can make us a bit more certain about the consistency of the magnitude of this effect, but will in no way diverge from what is seen in the controlled environment of a trial. And we know that in most instances therapies lose in the magnitude of their effectiveness when tested in the real world setting.

Another interesting discussion happening in the literature is about heterogeneous treatment effect (HTE), which I have discussed here. The idea is that a study reports the aggregate findings through measures of central tendency, such as the means and medians, as well as some of the scatter around them. What these numbers fail to tell us is how an individual patient can be expected to respond to a particular intervention. The inherent inter-individual variability in disease and response, as well as intra-individual variability in these factors over time pose major challenges in interpreting the single curve presented in a clinical trial report. Although methodologically rigorous subgroup analyses have been proposed as a potential solution to this conundrum, they are a thorn in the side of statisticians due to their implications for probability of errors, and, even when done well, are still not likely to get down to the level of granularity needed at the bedside. And of course, if a trial is not representative of the population that the intervention will eventually address in the real world, then all of these machinations will yield rather limited information.

So, what are we left with? We are still left with "E" that is far from perfect, yet it is this "E" that drives the business of medicine. I was reminded of this by this op-ed from Gil Welch from Dartmouth, where he mentions screening mammography as a metric that feeds into publicly reported MD grades. Clearly, the issue of screening mammography is not at all settled, yet the "evidence" is driving an important policy.

I hope you, my reader, do not think that I am an opponent of EBM. I have seen the dirty underbelly of business-based medicine, where an expedient diagnosis of a common ailment provides a steady revenue stream from subsequent treatment and complications associated with it. I have witnessed willfullness-based medicine and laziness-based medicine. I have also seen medicine practiced with compassion and utmost respect for the nexus of evidence and the individual patient. It is the dogmatic know-it-all approach to our evolving understanding of human health that gets us into trouble every time. Patients are not happy, clinicians are not happy, and even the bureaucrats are beginning to feel the sting of EBM's unbridled success. Men like Cochrane, Eddy and Sackett understood the nuances of the clinical encounter and advocated that it take place within the context of EBM, not that EBM replace these nuances. The bastardization of their intent is palpable and fraught. We will be reaping the fruit of their success for a long time to come.                    

Thursday, October 14, 2010

The reality of "science-based medicine"

Dr. Novella continues with his egregious oversimplification of the concept of science-based medicine here. I again felt compelled to respond. And given my previous difficulties with getting my responses accepted by the web site, I thought I'd post it here too. Tell me what you think.

Dear Dr. Novella,

Once again I have to agree with some of your premises, but disagree with your misguided leaps of illogic. I agree that if a modality has not been proven effective, the only way it should be left alone to be used by the public is if it has been shown to be safe. Alas, the risk-benefit equation is an individual choice, and we cannot impose our quantitative bottom line on it. Your assertion that scientific medicine is being eschewed because of acceptance of alternative modalities is as flawed as maintaining that a rain dance brings on rain. I know you said the relationship was complicated, but let's be honest: you think that CAM acceptance is killing allopathic medicine.

Now, let's get on to "proof" in science-based medicine. As you well know, while we do have evidence for efficacy and safety of some modalities, many are grandfathered without any science. Even those that are shown to have acceptable efficacy and safety profiles as mandated by the FDA, are arguably (and many do argue) not all that. There is an important concept in clinical science of heterogeneous response to treatment, HTE, which I have addressed extensively on my blog. I did not make it up, it is very real, and it is this phenomenon that makes it difficult to predict how an individual will respond to a particular intervention. This confounds much of what we think is God's own word on what is supposed to work in allopathic medicine.

Finally, do you really think that agents that are approved based on a 2-week prolongation of median survival in a desperately ill population of patients are used because of their supposed scientific merit? I would have to argue that there is a lot of subcortical emotional thinking that goes into these decisions. Can you really prove to me that a 2-week increase in median survival is not tantamount to placebo effect, aka type I error? Yet this is science-based. I think if I had a horrible disease, I might opt for acupuncture to make me feel better in the weeks I have left rather than rely on this kind of "science" to prolong my misery by 2 weeks.

Bottom line, we need to appreciate that none of the science is all that straightforward. Let us not dumb down the arguments and create false dichotomies. If we do, no one wins.

Wednesday, September 29, 2010

Disruptive innovation in healthcare: Overcoming HTE

I have been working on a talk for the American College of Chest Physicians (Chest) annual meeting. The session is on alternative research study designs, and I chose to talk about the N of 1 trials. I think that you can get an idea of why I chose this topic from reading my previous posts here, here and here. Doing my research, I came upon a term "heterogeneous treatment effect", or HTE, that is well worth exploring. I think that every clinician who has ever seen patients is familiar with this effect, but let us trace its explanation.

This excellent article in the UK Independent summarizes the premiss with some scathing comments from none other than GSK's chief geneticist, Dr. Allen Roses:
"The vast majority of drugs - more than 90 per cent - only work in 30 or 50 per cent of the people," Dr Roses said. "I wouldn't say that most drugs don't work. I would say that most drugs work in 30 to 50 per cent of people. Drugs out there on the market work, but they don't work in everybody."
There is even a table presented with response rates by therapeutic area, though the reference(s) is(are) not cited, so, please, take with a grain of salt:
Response rates
Therapeutic area: drug efficacy rate in per cent
  • Alzheimer's: 30
  • Analgesics (Cox-2): 80
  • Asthma: 60
  • Cardiac Arrythmias: 60
  • Depression (SSRI): 62
  • Diabetes: 57
  • Hepatits C (HCV): 47
  • Incontinence: 40
  • Migraine (acute): 52
  • Migraine (prophylaxis): 50
  • Oncology: 25
  • Rheumatoid arthritis: 50
  • Schizophrenia: 60
In essence, what Dr. Roses was referring to is the phenomenon of HTE, described aptly by Kent et al. as the fact that "the effectiveness and safety of a treatment varies across the patient population". The authors preface it by saying that 
Although “evidence-based medicine” has become the dominant paradigm for shaping clinical recommendations and guidelines, recent work demonstrates that many clinicians’ initial concerns about “evidence-based medicine” come from the very real incongruence between the overall effects of a treatment in a study population (the summary result of a clinical trial) and deciding what treatment is best for an individual patient given their specific condition, needs and desires (the task of the good clinician). The answer, however, is not to accept clinician or expert opinion as a replacement for scientific evidence for estimating a treatment’s efficacy and safety, but to better understand how the effectiveness and safety of a treatment varies across the patient population (referred to as heterogeneity of treatment effect [HTE]) so as to make optimal decisions for each patient.
Ah, so it is not your imagination: when someone brings an evidence-based guideline to you, and insists that unless you comply 95% of the time, you are providing less than great quality of care, and you say "this does not represent my patients", you are actually not crazy. To be sure, a good EBPG will apply to most patients encountered with the particular condition. But the devil, of course is in the details. As I have already pointed out, we impose statistical principles onto data to whip it into submission. When we do a good job, we acknowledge the limitations of what measures of central tendency provide us with. But so much of the time I see physicians relying on the p value alone to compare the effects, that I am convinced that the variation around the center is mostly lost on us. And further, how does this variance help a clinician faced with an individual patient who has at best a probability of response on some continuum of a population of probabilities? And more importantly, what will this individual patient's risk-benefit balance be for a particular therapy?

I think what I am walking away with after thinking about this issue is that it is of utmost importance to understand what kind of data have gone into a recommendation. What is the degree of HTE in the known research, and specifically, what is known about the population that your patient represents. The less HTE and the more knowledge about the specific subgroups, the more confident you can be that the therapy will work. Ultimately, however, each patient is a universe onto herself, since no two people will share same genetics, environmental exposures, chronic condition profile or other treatments, to name just a few potential characteristics that may impact response to therapy.

This is the reason that we need better trials, where people are represented more broadly, leading to an increase in external validity. To make this information useful at the bedside, we need a priori plans to analyze many different subgroups, as that will give clinicians at least some granularity so desperately needed in the office. And while pharmacogenomics may be helpful, I am sure that it will not be the panacea for reducing all of this complexity to zero.

Until technology gives us a better way (assuming that it will), where possible, a systematic approach to treatment trials should be undertaken. Later I will blog about N of 1 trials, which, though not appropriate in every situation, may be quite helpful in optimizing treatment in some chronic conditions. With the advent of health IT, these trials may become less daunting and, in aggregate, provide some very useful generalizable information on what happens in the real world. Each clinician will need to take some ownership in advancing our collective understanding of the diseases s/he treats. This may truly be the disruptive innovation we are all looking for to improve the quality of care not just to please the bureaucrats, but to promote better health and quality of life.