Showing posts with label false positive. Show all posts
Showing posts with label false positive. Show all posts

Tuesday, August 14, 2012

BTL reader question: How do you get to 2%?

I have started a FAQ page on the BTL book web site here, and I will cross-post the discussion here on the blog. This will give us an opportunity to have a more interactive discussion, if necessary, with additional comments and questions.

Here is the inaugural installment.

On August 13, 2012, this question came in via Twitter:




Well, here is the answer (and thank you for the question, Tia!)

First the problem. At the bottom of page 74 and going on to the top of page 75 I discuss the question posed in a 1978 New England Journal of Medicine paper by Casscells and colleagues to 60 physicians and physicians-in-training at Harvard Medical School. The problem went like this:
 
"If a test to detect a disease whose prevalence is 1/1000 has a false positive rate of 5 per cent, what is the chance that a person found to have a positive result actually has the disease, assuming that you know nothing about the person's symptoms or signs?"

The question clearly mimics a disease screening situation. The answer is simple yet elusive. Let us assume that 1,000 people are tested. Among them only 1 person has the actual disease. However, given that the false positive rate is 5%, we also know that out of the 1,000 people tested, 50 will have a false positive test. Assuming that the single person with the disease also has a positive test, we can expect 51 people to test positive. But since only 1 out of these 51 people with a positive test has the disease, the answer to the question above is 1/51=2%. This is a pretty shocking realization, given that a large plurality of the Harvard doctors and trainees chose 95% as their answer. 

So, be careful not to let your intuition override the data when making medical decisions!


If you like Healthcare, etc., please consider a donation (button in the right margin) to support development of this content. But just to be clear, it is not tax-deductible, as we do not have a non-profit status. Thank you for your support!

Wednesday, July 25, 2012

Medicine as the trolley problem

Are you familiar with the trolley problem? It is an ethics dilemma first formulated by the great Philippa Foot as a part of a series of such dilemmas. Her formulation goes roughly like this. Imagine there is a tram hurtling down a track. If it keeps going straight, it will hit and kill 5 people who are working on that track. The conductor is able to throw a switch and divert the train to another part of the track, where 1 single worker will be killed by the trolley. The question is what should the conductor do? Most people when asked respond that yes, he should throw the switch and sacrifice 1 life to save 5. After all, the net benefit is n=4.

There are literally thousands of alternative formulations of this problem, but one of them from the philosopher Judith Jarvis Thomson merits special consideration. The problem starts out similarly, with 5 lives on a track in potential peril. The vantage point and the solution are quite different, though. Now there is a bridge over the rail track, and a very large man is looking at the tracks from the bridge. One way to stop the train is to throw a heavy object in its path, like this large man, for example. You are on the bridge standing behind the man. Would you be justified in pushing him off the bridge in front of the tram to meet his death in order to spare the 5 workers down the tracks? Most people when faced with this formulation say an emphatic "no." This is somehow puzzling, since the net benefit is the same, n=4, as in the original Foot formulation.

Philosophy professors have puzzled over this difference for decades, and there are several potential explanations for why we respond differently to the two scenarios. One explanation has to do with the proximity of the operator (conductor in the first case and the person doing the pushing in the second) to the sacrificial lamb -- in the first case one is enough removed from the action of killing by merely redirecting the tram, whereas in the second the action is, well, more active, and the operator is actually pushing an innocent person to his death.

Though in some ways the scenarios seem to bear no practical distinction from one another, we see the morals and ethics of each differently. This difference in the view point is instructive to the field of medicine, where it has implications to how policy relates to the individual patient encounter. Here is what I mean.

Suppose you are a policy maker, and you recommend that every woman at age 40 start to receive an annual screening mammogram to reduce deaths from breast cancer. At the population level, if we screen 1,000 women for about 30 years, we will save approximately 8 of them from a breast cancer death. (Yes, it's 8, not 80, and not 800). At the same time, among these 1,000 women, there will be over 2,000 false alarms, and over 150 of these will result in an unnecessary biopsy. Some of these biopsies will incur further complications, though currently we  do not seem to have the data to quantify this risk. But what if even one of these biopsies were to lead to death of or another dire lasting complication in a woman who turned out not to have cancer? And by the way the accounting is not all that different when applying the new USPSTF mammography screening recommendations. Well, then we have the trolley problem, don't we? We are potentially sacrificing 1 individual to save 8. And who does the sacrificing is where the variations of the trolley problem come in.

Payers levy financial penalties on primary care physicians when they fail to comply with screening recommendations in their patient panels. The payer certainly sees this issue as the original formulation of the problem: Why not throw this financial switch to achieve net life savings? But for a clinician who deals with the individual patient this may be akin to pushing her over the bridge toward a potentially fatal event. Because we don't have a crystal ball, we cannot say which woman will die or incur a terrible complication. But the same population data that tell us about benefits must also give us pause when reflecting on the risks. Add the ubiquitous uncertainty (and lack of data) into this equation, and the implications are even more shocking. So, while making policy recommendations based on population data is sensible, policing uniform application of these recommendations to individual patients is fraught: of course, clinicians and patients need to be cautious about making individual decisions even when in population data benefits outweigh risks.

On the surface risk-benefit equations for many interventions may appear favorable, leading to blanket policy recommendations to employ them on everyone who qualifies. In the office, the clinician, caught in a tug of war between mountains of new literature and the ever-shrinking appointment times, is hard-pressed to take the time to consider these recommendations in the context of the individual patient. And furthermore, financial incentives from payers act as a short-hand justification, a "nudge," for doing as recommended rather than for giving it thought. So, who must look out for the patient's interest? The patient, that's who. Who understands the patient's attitude toward the risks and the benefits? The patient, that's who. Who now has to be responsible for making the ultimate informed decision about which track to stand on? The patient, that's who.

For me the trolley problem gives clarity to the reservations that I walk around with every day. I have done a lot of soul searching about why it is that, even if the benefits seem to outweigh the risks, I am still more often than not skeptical about whether a particular intervention is right for me. And since every intervention in medicine has a real risk, though mostly quite low, of going terribly awry, my skepticism is justified. This is my approach to evaluating these risks and benefits, based on my values and my understanding of the data as it is today.

What's the answer to this ethical conundrum in medicine? I cannot see that policy makers will stop throwing the switch in the near future, and so as a society we will be forced to accept the tram's collateral damage. And while this may make sense in an area such as vaccination, where thousands of lives can be saved by sacrificing a very few by throwing the switch, in most everyday less clear-cut medical decisions the answer is less clear-cut. Will doctors rebel against being forced to throw some patients on the tracks in order to save some marginally larger number of others? I don't think that they have the time or the energy or the incentive to do this, since the framing of the switch-throwing is through the rhetoric of "evidence." Right or wrong, doctors are shackled by the stigma of ignorance that comes with not following evidence-based guidelines, and this may act to perpetuate blind compliance. This leaves the patients, for some of whom the right thing will be just to get themselves off the tracks altogether, far away from the hurtling trolley until its brakes are fixed.                        

If you like Healthcare, etc., please consider a donation (button in the right margin) to support development of this content. But just to be clear, it is not tax-deductible, as we do not have a non-profit status. Thank you for your support!

Friday, June 29, 2012

Molecular diagnostics: Making the uncertainties more certain?

Scott Hensley over at the NPR's Shots blog posted a story about the recently approved molecular diagnostic test that can rapidly identify several gram-positive bacteria that cause blood stream infections.  This is indeed important, since conventional microbiologic techniques rely on bacterial growth, which can take up to 2 to 3 days. This is too long to wait to identify the bug that is the cause of a serious infection. What doctors have done to date is make the best guess based on several factors, including the type of a patient, the source of the infection and the patterns of bacterial resistance at their site, to tailor empiric antibiotic coverage. The sicker the patient, the broader the coverage, until the culture results come back, when the doctor is meant to alter this treatment accordingly, either by narrowing or broadening the spectrum. The pitfalls of this work flow are obvious -- too many points where error can enter the equation. So on the surface the new tests are a great advance. And they actually are, but they are not free of problems, and we need to be very explicit confronting them.

Each diagnostic test can be evaluated on its sensitivity (how well it identifies the problem when the problem exists), specificity (how rarely it identifies a problem when it does NOT exist), and positive (what proportion of all positive tests represents true problem) and negative (what proportion of all negative tests represents a true absence of problem). Sensitivity and specificity are intrinsic properties of the test and can be altered only by how the test is performed. Positive and negative predictive values are dependent not only on the test and how it is done, but also on the population that is getting tested.

Let's take Nanosphere's test in Scott's story. If you trawl the company's web site, you will find that the sensitivity and specificity of this technology is close to 100%, if not 100%, the "gold standard" for comparison being conventional microbiology culture. And perhaps this is really the case in these very specialized hands that were testing the diagnostic. If these characteristics remain at 100%, disregard the rest of this post, please. However, the odds that they will remain at 100% in the wild of clinical practice are slim. But I am willing to give them 99% on each of these characteristics nevertheless.

OK, so now we have a near-perfect test that is available for anyone to use. Imagine that you are an ED doc at the beginning of your shift. An ambulance pulls up and rolls a septic patient into an empty bay. The astute ED nurses rush into settle the patient, and, as a part of the protocol, take a sample of blood for determining the pathogen that is making your patient sick. You quickly start the patient on broad spectrum antibiotics and walk away to take care of the next patient that has just rolled in with a heart attack. A few hours later, the septic patient, who is still in the ED because there are no ICU beds for him yet, is pretty stable, and you get the lab result back: he has MRSA sepsis. You pat yourself on the back because one of the antibiotics that you ordered was vancomycin, which should cover this bug quite adequately. You had also put him on ceftazidime to cover any potential gram-negative critters that may be lurking within as well. Now that you have the data, though, you can stop ceftaz and just continue vanc. The patient finally gets a bed upstairs, and your shift is over and you go home withe a sense of accomplishment.

The next morning you come in refreshed with your double-venti iced macchiato in your hand, sit at the computer and check on the septic patient. You are shocked to find out that last night he decompensated, went into shock and is now requiring breathing assistance and 3 vasopressors to maintain his blood pressure. You scratch your head wondering what happened. Then you come upon this crazy blog post that tells you.

Here is what happened. What you (and these tests) did not take into account is the likelihood of MRSA being the real problem rather than just a decoy false positive. Let's run some numbers. The literature tells us that the likelihood of MRSA causing sepsis is on the order of 5%. Let's create a 2x2 square to figure out what this means for the value of a positive test, shall we?


MRSA present
MRSA absent
Total
Test +
495
95
590
Test -
5
9405
9410
Total
500
9,500
10,000

What this says is the following. We have 10,000 patients roll into our ED with sepsis (in reality there are about 1/2 million to 1 million sepsis cases in the US annually), and we test them all with this great new test that has 99% sensitivity and 99% specificity. Of these 10,000, fifty five hundred (thanks, Brad, for noticing this error!) are expected to have MRSA. Given this situation, we are likely to get 590 positive tests, of which 95, or 16%, will be false positive. Face-palm, you drop your head on the desk realizing that Mr. Sepsis from yesterday was probably one of these 16 per 100 false positives, and MRSA is probably not the cause of his infection.

You begin to wonder what if your lab really did not get the sensitivity and specificity of 99%, but more like 98%? Still pretty generous, but what if? You start writing madly on a napkin that you grabbed at Starbucks, and your jaw drops when you see your 2x2:


MRSA present
MRSA absent
Total
Test +
490
190
680
Test -
10
9310
9320
Total
500
9,500
10,000

Wow, you think, this false positive rate is now nearly 30% (190/680)! You can't believe that you could be jeopardizing your patients' lives 3 times out of 10 because you are under the mistaken impression that they have MRSA sepsis. This is unacceptable. But can you really trust yourself with these calculation? You have to do one more thing to convince yourself. What if your lab only gets 97% specificity and sensitivity? What then? You choke when you see the numbers:




MRSA present
MRSA absent
Total
Test +
485
285
770
Test -
15
9215
9230
Total
500
9,500
10,000


It's and OMG moment -- nearly 40% would be treated for MRSA when they potentially have something else.

But you, my dear reader, realize that in the real world docs are not that likely to remove gram-negative coverage if MRSA shows up as the culprit pathogen. Why should you think otherwise, when there is so much evidence that people are not that great about de-escalating antimicrobial coverage in response to culture data? But then I have to ask you what's the use of this new test if no one will act on it anyway? In other words, how is it expected to help curb the rise of resistance? In fact, given the false positive MRSA rates we see above, might there not even be a paradoxical increase in the proliferation of resistance?

The point is this: We are about to see many new molecular diagnostic technologies on the market that have really really high sensitivity and specificity. The fly in this ointment, of course is the pre-test probability of the bug causing the problem. Look how in a very low risk group (5% MRSA) even a near-perfect test's value of a positive is reduced by almost a ridiculous magnitude. Do feel free to check my math.

So you trudge into the hospital the next day for your shift and check on Mr. Sepsis one more time. Sure enough, his conventional blood culture grew out E. coli, a gram-negative bug. You notice that he is turning around, though, ceftazidime having been restarted by the astute intensivist (well, I am a bit biased here, of course). All is well in the world once again. Except you hear an ambulance pull up and the nurse talking on the phone to the EMTs -- it's another sepsis on your hands. What are you going to do now?
              

If you like Healthcare, etc., please consider a donation (button in the right margin) to support development of this content. But just to be clear, it is not tax-deductible, as we do not have a non-profit status. Thank you for your support!

Tuesday, March 27, 2012

Coronary CT angiography: More and less

I am still scratching my head over this study that was just published in the NEJM on coronary CT angiography among patients admitted with a suspected coronary syndrome. There are many potentially confusing points about it and how it was reported. The three things that I found most confusing were:
1. The formulation of the null hypothesis
2. The definition of the outcomes
3. The flow of patients through the diagnostic algorithm.
Let's see if we can clear some of that confusion.

The intent of the study was to see how adding the CCTA to the usual diagnostic testing in the ED impacted cardiac mortality or a myocardial infarction within 30 days. The null hypothesis was posed in an interesting way:
The study was powered to test the null hypothesis that the rate of major cardiac events among patients who did not have clinically significant coronary artery disease as assessed by CCTA would exceed 1%. 
This is confusing, and I had to read and reread several times to understand what it meant. I finally realized that what they were saying was that in order to disprove that CCTA is not a useful test (this is the alternative hypothesis) they would have to show that the rates of the primary outcome were not above 1%. Make sense?

The next issue I had trouble with, as I always do in cardiology studies is their choice of endpoints. The primary endpoint was 30-day cardiac death OR MI. This means that anyone who either died of a cardiac cause or had a heart attack within this time period was counted as an event, which would argue against CCTA usefulness if they reached the critical mass of over 1%. Mind you this was limited to those patients whose CCTA did not reveal significant disease. There were some secondary outcomes examined as well, all at the 30-day time point: death, MI, revascularization procedure, and resource utilization. Note these outcomes were not combined, but were examined singly, and the pool of patients for these were all those randomized.

As an aside, cardiology studies frequently use combined outcomes, such as death and MI due to sample size considerations. That is both cardiac death and MI should be rare events in the group examined. When these events are rare, in order to get at their statistical significance exceedingly large sample sizes are needed. For this reason, these trials frequently combine several events, so as to enrich the frequency and commensurately drop the needed sample size.

But here is where it gets a little confusing. Looking at the "Safety" section of the Results, the authors state that no one who received a CCTA had a cardiac death or an MI within the 30-day period. However, 1% of the patients randomized into the CCTA arm did have a MI within 30 days. How can this be? Well, we need to keep in mind that not all those randomized to CCTA (n=908) actually got a CCTA (n=767), and that the majority of those who did have one did NOT have significant coronary disease (n=640). Thus, both the numerator and the denominator for the primary and the secondary outcomes are different.

Then there is this quick sentence in the "Efficiency and Use of Resources" section:
Coronary disease was more likely to be diagnosed in patients in the CCTA group than in patients in the traditional-care group (9.0% vs. 3.5%; difference, 5.6 percentage points; 95% CI, 0 to 11.2).
It is almost an afterthought, but it is important. It is puzzling that the outcomes in the two groups are completely identical, and this is not limited to those in the primary endpoint pool, namely people without significant disease. The question arises about what this means. Since there is no difference in cardiac death or MI, this increase in the diagnostic rate may imply overdiagnosis in the group receiving CCTA. There is a slight increase in the revascularization rates in the CCTA group (3%) over the standard care group (1%). This endpoint is much more subjective that most people realize, as a lot of judgment by the cardiologist and the surgeon goes into the decision. So this endpoint does not rule out overdiagnosis as a possibility. On the other hand, the study was not powered to detect a difference in the secondary outcomes, so it may be that the diagnoses are valid, but we do not have enough events to judge.

One final point of frustration with the study reporting was hitting dead ends in the flow of patients through the respective algorithms. I was, of course, looking for the rates of false positive CCTA tests and their outcomes. Unfortunately, I kept getting stuck at the catheterization and stress test steps, not knowing who exactly went on to have this testing. Since in the CCTA out of the 767 tests 47 were indeterminate and 80 were indicative of moderate-to-severe coronary disease, the 37 catheterizations performed were likely among these patients, though we are not told for sure. Of these catheterizations, 9 were negative for a significant stenosis, but again we are not told how this reflects back to the CCTA results. In fairness it is worth noting that fewer people in the CCTA arm (18%) underwent a follow-up test, compared to the standard care arm (62%). But again, this is just one test replacing another, or even being added on top of others. In addition, more of the former group were discharged from the ED (50%) than among the latter (23%) without being admitted.

So what does it all say to me? Well, CCTA seems to aid in the diagnosis of coronary disease among certain people presenting to the ED with chest pain. Given such a low pre-test probability of significant coronary disease (17%, derived by dividing the negative CCTA [640] by the total CCTA tests [767] and subtracting that from 1), I have to wonder about the performance of the test that is quite sensitive but not that specific (high risk of a false positive in a low-risk population). And even though fewer patients needed to be admitted from the ED in the CCTA group, I wonder if this does not have more to do with the differential availability of the testing rather than with its superiority.

Given the context that I laid out above for the results presented, I would not be rushing to adopt CCTA for all patients who present to the ED with a certain type of chest pain. And after all, though a "black swan" event, catastrophes do occur in follow-up catheterization even when the patient is disease-free. This is a good reason to be careful and circumspect in our quest for primum non nocere.      

If you like Healthcare, etc., please consider a donation (button in the right margin) to support development of this content. But just to be clear, it is not tax-deductible, as we do not have non-profit status. 
Thank you for your support!

Thursday, March 22, 2012

More on "mammography saves lives" story


A story in HealthDay with the title of "Two Studies Find Routine Mammography Saves Lives" talks about studies presented at the annual meeting of the American Society for Clinical Oncology (ASCO) European Breast Cancer Conference (EBCC). The studies, both from the Netherlands, allegedly showed that mammography does indeed save lives. If true, these data would contrast with the preponderance of evidence that has stirred up such a ruckus recently about the utility of mammography, complete with references to rationing and death panels. But let's look at what's reported in today's story more closely.

Cutting through all the definitive bravado, here is a little piece of science that was reported:
Compared with the pre-screening period 1986 to 1988, deaths from breast cancer among women aged 55-79 fell by 31 percent in 2009," Jacques Fracheboud, a senior researcher at the Erasmus University Medical Center in Rotterdam, said in a meeting news release. We found there was a significant change in the annual increase in breast cancer deaths: before the screening program began, deaths were increasing by 0.3 percent a year, but afterwards there was an annual decrease of 1.7 percent," he added. "This change also coincided with a significant decrease in the rates of breast cancers that were at an advanced stage when first detected.
Note, the reference is to deaths from breast cancer without any mention of all-cause mortality. (You can read about why the latter is important here.)

The report next states that over the first 20 years of the screening program
... 13.2 million breast cancer screening examinations were performed among 2.9 million women (an average of 4.6 examinations per woman), resulting in nearly 180,000 referral recommendations, nearly 96,000 biopsies and more than 66,000 breast cancer diagnoses.
So, doing the math, I come up with about 31% false positive rate at the biopsy stage (that's 96,000 biopsies minus 66,000 positives for cancer, all divided by the 96,000 total biopsies). If we use the 180,000 "referral recommendations" as our denominator of all positive tests, and stick with the 66,000 true positive rate, then the false positives grow to (180,000-66,000)/180,000 = 0.63, or 63%. If we spread the 33,000 false positives over the 13.2 million examinations, that equates to 0.25% chance for a false positive. Yet the report goes on to say that (emphasis mine):
For a woman who was 50 in 1990 and had 10 screenings over 20 years, the cumulative risk of a false-positive result (something being detected that turned out not to be breast cancer) was 6 percent.
Six percent? This is clearly a place where my high school math teacher's mantra of "show your work" is applicable.

The next piece of information that I would like to understand better is this:
Over-diagnosis (detection of breast tumors that would never have progressed to be a problem) occurred in 2.8 percent of all breast cancers diagnosed in the total female population and 8.9 percent of screening-detected breast cancers."
How exactly was this computed? Again a case for "show-your-work."

And then there is this (emphasis mine):
Regular screening "decreases deaths by over 30 percent, [with] limited harm and reasonable costs. Additionally, cancers are detected at an earlier stage, which means not only decreased mortality but also morbidity; the patient may not have to have chemotherapy or a mastectomy," she noted.
OK, so, if I got it right, it is breast cancer mortality that is decreased by 31%, not all-cause mortality. This really should have been spelled out more clearly, not to mention that the actual, or absolute, reduction likely pales in comparison to this relative drop. And what about diagnosing earlier stage disease? Lead time bias, anyone?

The second study was a computer model, and I will not go through it at this time as I need to move on to other work. But you get the picture: the numbers given in the report are limited and at times they don't add up. Mixing up cancer mortality with all-cause mortality leads to erroneous conclusions. And finally, forgoing reporting on the absolute risk reduction in favor of the inflated relative reduction is not helpful for understanding the true risks involved.

One final thought: Yes, I do have cognitive biases, and it is difficult for me to avoid them. I happen to fall into the camp that thinks screening for sublclinical diseases, at least in our current technological setting, is disease mongering. At the same time, I would like to think that if the data really showed a significant benefit without great risks, I would give them a second look.

The bottom line is this: at least for me, the report confused the issue more than it has clarified. Perhaps the study, once published, will answer all of the questions that I have posed adequately. But at this stage, it is a shame that such strong statements as...
"These results show why mammography is such an effective screening tool," said one U.S. expert, Dr. Kristin Byrne, chief of breast imaging at Lenox Hill Hospital in New York City. She was not involved in the new research.
 and this...
"We are convinced that the benefits of the screening program outweigh all the negative effects," Fracheboud said.
 ... are not backed up by appropriate evidence.

h/t to @ElaineSchattner for the story


If you like Healthcare, etc., please consider a donation (button in the right margin) to support development of this content. But just to be clear, it is not tax-deductible, as we do not have a non-profit status. 

Thank you for your support!

Thursday, March 8, 2012

Three central questions about medical technologies

Every day my inbox gets filled with announcements for and invitations to attend all kinds of conferences. While many of them are of the traditional medical education sort, more and more I hear about meetings where new technologies and gadgetry in healthcare are the focus. And while the former feature healthcare professionals speaking medicalese from the stage to rapt audiences of other healthcare professionals in dimly lit halls, the latter capture their audience' imaginations with the promise of the future, in all its glitter and glory. And naturally, the latter are what attract techies and patients alike, steamrolling over our staid medieval medical conventions. I heard that the HIMSS conference in Las Vegas last month attracted 37,000 attendees! And the excitement was palpable even through Twitter feeds. This is clearly the preferred way to effect public engagement with healthcare.

But here is the thing: medical advances happen much more slowly than the speed of technology development. That is why, year after year, we go to our professional society meetings and have deja vus all over again. Year after year we hear the same people present the same studies, sometimes, if we are lucky, with a slightly different twist. The last group of breakthroughs I heard about at one of our critical care meetings was over a decade ago. And we have had to backpedal from that quite a bit with the removal of Xigris from the market, and with the realization that tight glucose control had to be used with extreme caution, so as not to kill more critically ill patients than it was meant to save.

This disconnect between the glacial pace of true progress in the clinical sciences and the lightening speed of technological progress raises some obvious questions about our assumptions. If we are not making dramatic breakthroughs in medicine every day, what is this breakneck pace of technology innovation delivering, save for the glitter and a seductive promise of health and wealth? Is there evidence supporting this promise? You might counter by saying that we have years, maybe even decades, of translational catching up to do, bringing all the advances from the bench to the bedside. I would have to say that the magnitude of such advances, as well as clinicians' resistance to them, may have been overstated. You might also point out that the vast advances in computing capabilities have not penetrated sufficiently into our healthcare system, and there I cannot disagree. But whether bringing these advances into clinic without careful planning will improve our health or our healthcare finances remains in question.

Moreover, I see numerous downsides to rushing ahead without thinking through what we are rushing toward. Ostensibly, the light at the end of this bright technological tunnel is better health. What concerns me is that the journey has become a sort of an end in itself: the sheer beauty of the tunnel has itself become the prize. To regain our compass, we need to ask three tough questions, each of them central to understanding the impending advent of too much out-of-context information:
1. Do we really want to walk around with sensors (pdf document, see page 16)? Personally, I find 24/7 monitoring of our vital functions to be a depressing prospect. Furthermore, I sincerely doubt that this is a better (or more cost-effective) way to achieve health than through focus on public health and socioeconomic equity.
2. What is the use of having your genome in your pocket?
How is that going to help us at the stage where all we can do is identify certain levels of risk (bracketed by broad intervals of uncertainty) in isolation from all the influences that modify that risk?
3. Do we want an epidemic of false positive findings and pseudo-disease?
We have enough trouble interpreting positive findings from medical screening tests. Do we really want the public falling prey to the anxiety and over-testing that false positives bring? As a healthcare system, can we sustain such an avalanche? As clinicians are we able to mange these screening snafus?
(I have done so many posts on this issue that you can barely navigate this site without stumbling over their debris).

All of the above is not sexy or shiny, and it brings in the ugly four-letter word "risk" to balance the discussion of the holy grail of benefit. So, how do we sell it to people so hypnotized by technology? I am asking this question quite seriously. I, and many of you out there, really would like people to become more cognizant of these nuances, but how do we accomplish this? Do we need to change our byzantine approach to medical meetings to start attracting a broader audience for our messages? If so, let's get started. Perhaps other forums that are already exceedingly successful can teach us how. I commend TEDMED (which I am attending as a Front-Line Scholar) for starting to bring this important viewpoint to their audiences. Now, how about having HIMSS, Health 2.0 and others join to clarify this other edge of the medical sword? Perhaps we can prevent the wild pendulum swings by thinking things through now. And think how much more credibility we all will have if, by thinking and debating now, we minimize the unintended consequences later?                            

Tuesday, February 28, 2012

The flu and likelihood ratios

An interesting study was just published in the Annals of Internal Medicine. It was a meta-analysis of rapid influenza diagnostic tests (RIDT) and their characteristics. Since we have been spending so much time talking about test characteristics, this study provides a nice opportunity to discuss another way of looking at the values of positive and negative tests. This is going to be a fairly short post, since I simply want to discuss these additional tools.

In the current study the investigators evaluated how well these RIDTs predicted the disease. As we have discussed in multiple places on this blog (here, here and here, to name a few), it matters whether the test is used for screening or diagnosis. In this case, the testing was done in symptomatic populations, so for diagnostic reasons. The authors report that there was quite a bit of heterogeneity in their findings, but the ultimate result is reported as a positive (34.5) and negative (0.38) likelihood ratios. What are they and how do we interpret them?

A positive likelihood ratio, or LR+ is the ratio between sensitivity and 1-specificity (LR+=[sensitivity]/[1-specificity]). Sensitivity is the proportion of patients with the disease who are identified as having the disease, or true positives (TP), and specificity is the proportion of persons without the disease who are identified as not having the disease, or true negatives (TN). The opposite of TN, 1-TN, is the false positives (FP). So, the LR+ equates to the TP/FP, or the odds that a positive indicates true disease. In the current study it is 34, meaning that the odds are 34 to 1 that a positive test indicates the presence of the disease. Another way of putting it is that of the 35 total positive test results, 34 (97.1%) represent true disease. This is essentially equivalent to the positive predictive value (PPV).

Now, let's examine the negative likelihood ratio, or LR-. This is defined as the ratio between the opposite of sensitivity (1-sensitivity) and specificity (LR-=[1-sensitivity]/specificity]). 1-sensitivity is the proportion that are false negative (FN), while the specificity is the proportion of persons without the disease who are identified as such. In the study this LR- was 0.38, meaning that the odds that a negative result truly indicates the absence of disease are about 1 to 2 (0.36:1), or not so great. In other words, out of the total of 3 negative tests, 2 are truly negative, while 1 is a false negative, giving us the negative predictive value (NPV) of about 65% (actually it is 1-0.36=0.64, or 64%).

So, there you have it. The clinical take-away, as the authors noted, is that these RIDTs are good at ruling in the flu, but not at ruling it out. In other words, the problem here is the opposite of what we discussed in all those previous test, or the rate of false negatives. And this makes sense, given that the pre-test probability is reasonably enriched in populations with symptoms, in addition to the relatively poor sensitivities of these technologies.  

  

Wednesday, February 22, 2012

Endometriosis and cancer: When a "breakthrough" may not be all that

There is an interesting new study that was just published online at the Lancet Oncology. It is a pooled analysis of a bunch of case-control studies to explore the association between endometriosis and certain types of ovarian cancer. Here is the abstract:

Background

Endometriosis is a risk factor for epithelial ovarian cancer; however, whether this risk extends to all invasive histological subtypes or borderline tumours is not clear. We undertook an international collaborative study to assess the association between endometriosis and histological subtypes of ovarian cancer.

Methods

Data from 13 ovarian cancer case—control studies, which were part of the Ovarian Cancer Association Consortium, were pooled and logistic regression analyses were undertaken to assess the association between self-reported endometriosis and risk of ovarian cancer. Analyses of invasive cases were done with respect to histological subtypes, grade, and stage, and analyses of borderline tumours by histological subtype. Age, ethnic origin, study site, parity, and duration of oral contraceptive use were included in all analytical models.

Findings

13 226 controls and 7911 women with invasive ovarian cancer were included in this analysis. 818 and 738, respectively, reported a history of endometriosis. 1907 women with borderline ovarian cancer were also included in the analysis, and 168 of these reported a history of endometriosis. Self-reported endometriosis was associated with a significantly increased risk of clear-cell (136 [20·2%] of 674 cases vs 818 [6·2%] of 13 226 controls, odds ratio 3·05, 95% CI 2·43—3·84, p<0·0001), low-grade serous (31 [9·2%] of 336 cases, 2·11, 1·39—3·20, p<0·0001), and endometrioid invasive ovarian cancers (169 [13·9%] of 1220 cases, 2·04, 1·67—2·48, p<0·0001). No association was noted between endometriosis and risk of mucinous (31 [6·0%] of 516 cases, 1·02, 0·69—1·50, p=0·93) or high-grade serous invasive ovarian cancer (261 [7·1%] of 3659 cases, 1·13, 0·97—1·32, p=0·13), or borderline tumours of either subtype (serous 103 [9·0%] of 1140 cases, 1·20, 0·95—1·52, p=0·12, and mucinous 65 [8·5%] of 767 cases, 1·12, 0·84—1·48, p=0·45).

Interpretation

Clinicians should be aware of the increased risk of specific subtypes of ovarian cancer in women with endometriosis. Future efforts should focus on understanding the mechanisms that might lead to malignant transformation of endometriosis so as to help identify subsets of women at increased risk of ovarian cancer.

Funding

Ovarian Cancer Research Fund, National Institutes of Health, California Cancer Research Program, California Department of Health Services, Lon V Smith Foundation, European Community's Seventh Framework Programme, German Federal Ministry of Education and Research of Germany, Programme of Clinical Biomedical Research, German Cancer Research Centre, Eve Appeal, Oak Foundation, UK National Institute of Health Research, National Health and Medical Research Council of Australia, US Army Medical Research and Materiel Command, Cancer Council Tasmania, Cancer Foundation of Western Australia, Mermaid 1, Danish Cancer Society, and Roswell Park Alliance Foundation.
Alas, I do not have access to the full article (paywall), so cannot go though it in a detailed way. Nevertheless, we can try to put these findings in perspective. So, briefly, the investigators put together data from many case-control studies and discovered that the risk of some, though not all, ovarian cancers was 2-3 times higher in the presence of endometriosis than in its absence, and concluded that clinicians should be aware of this increase in risk. Fair? But you know I am going to deconstruct it, right? Here we go.

By now we all understand what a case-control study is, right? It is a study where cases (those patients with the disease of interest) are compared to controls (subjects who are in all ways the same as the cases with the exception that they do not harbor the disease in question). So the study identifies subjects with an outcome, and follows them backward to the exposure that is of interest vis-a-vis this outcome. These studies are notoriously difficult to do well, particularly when it comes to the choice of a control, and very few do it well in my experience. I cannot comment on the current conglomeration of 13 of them, so will not venture a guess on whether some or how many of them may lead us astray. Though these studies are difficult, for various reasons they are the way to go when examining an uncommon outcome. So the choice of the design is legit.

No, let's examine the risk for misclassification for both the disease and the exposure. I think you will agree that a case of ovarian cancer is difficult to misclassify, so I will not pick on this as a potentially major threat to the validity. But what about endometriosis? This is a chameleonic condition that is probably way under-recognized. Its symptoms and signs are varied and, unless the studies required a look "inside," which I sincerely doubt they did -- note, the abstract states that this exposure was self-reported -- there is a very real and grave threat to validity here. Ever heard of recollection bias? If such exists, then more women with cancer are likely to report symptoms that may be indicative of endometriosis than those without cancer. So, the observed increase in the risk of ovarian CA in the presence of endometriosis may be due to just that -- a recollection bias.

These limitations notwithstanding, the authors felt that the study was a breakthrough (emphasis mine):
"This breakthrough could lead to better identification of women at increased risk of ovarian cancer and could provide a basis for increased cancer surveillance of the relevant population, allowing better individualization of prevention and early detection approaches such as risk-reduction surgery and screening,” lead author Celeste Leigh Pearce, at the University of Southern California, Los Angeles, said in a journal news release. 
What does Dr. Pearce mean by "increased cancer surveillance?" Is she talking about screening women with endometriosis because of this possibly heightened risk? And if so, is this really wise? Let's simulate some of these numbers.

Let us suppose that endometriosis does indeed increase the risk of ovarian cancer 2-3-fold. This means that the incidence now goes from 13 per 100,000 women up to 39 per 100,000. Let us now also assume that there is a test that is 99% sensitive (able to identify ovarian CA when it is present) and 99% specific (able to demonstrate that no ovarian CA is present when it is not present). Recall that at the population incidence of ovarian CA, the USPSTF does not recommend screening due to a very high risk of a false positive. The question is does this 3-fold elevation in risk change the positive predictive value of screening substantially enough for it now to be recommended? I think I know the answer, but let's go through the exercise anyway, just to be explicit.
Disease present
Disease absent
Total
Test+
39
1,000
1,038
Test-
0
98,961
98,962
Total
39
99,961
100,000

So, the corresponding positive predictive value is... drum roll, please... 3.7%. This means that out of 100 women who have tested positive, fully 96 have a false positive result and are now likely to be subjected to invasive procedures. If we imagine that a test can have near-perfect specificity of 99.99% (no test that I know of can come close to this in any consistent way), still 20% of all positive results are false positives. So, is this indeed food for screening thought? I really don't think so, particularly given that the current risk calculation is likely a gross over-estimate.

So, there you have it. I don't think I am engaging in hyperbole when I say that "breakthrough" is very likely an overstatement.                  
  

Saturday, February 18, 2012

The implications of a blood test for depression

So, a coupe of days ago we spent considerable (virtual) ink on discussing the risk of a false positive result in the setting of screening for rare events. Today something else has caught my eye: blood test for depression. The Atlantic reported this much-retweeted piece on the same day that I was droning on about lung cancer screening. So what is this about, and does it bear any similarity to what we discussed here?

Let us examine the lede of the Atlantic article:
New research shows that blood screenings can accurately spot multiple telltale biomarkers in patients with classic symptoms of depression.
The writer uses the word "screening" while talking about patients with "classic symptoms of depression." This choice of language is problematic. In clinical medicine the term "screening" generally refers to a population without any signs or symptoms of disease. Think of breast cancer and prostate cancer screening. When symptoms or signs are present, the testing becomes diagnostic, not screening, in its purpose. This difference is actually critical to appreciate in the context of what we talked about in the lung cancer screening post. The presence of signs that make a disease suspect presumably increase the pre-test probability of that disease. This means that the prevalence of this disease is higher in the population that has these particular signs/symptoms than in the overall population without them. And recall that it is this very pre-test probability that drives the predictive value of a positive test. Namely, the higher the pre-test probability of the disease, the more credence we can put in a positive test result.

OK, so let's move on to the data. I went to the primary source, but all I could get to was the abstract (paywall and all), so bear in mind that I do not have all of the data. I am reproducing the abstract here for your convenience:
Despite decades of intensive research, the development of a diagnostic test for major depressive disorder (MDD) had proven to be a formidable and elusive task, with all individual marker-based approaches yielding insufficient sensitivity and specificity for clinical use. In the present work, we examined the diagnostic performance of a multi-assay, serum-based test in two independent samples of patients with MDD. Serum levels of nine biomarkers (alpha1 antitrypsin, apolipoprotein CIII, brain-derived neurotrophic factor, cortisol, epidermal growth factor, myeloperoxidase, prolactin, resistin and soluble tumor necrosis factor alpha receptor type II) in peripheral blood were measured in two samples of MDD patients, and one of the non-depressed control subjects. Biomarkers measured were agreed upon a priori, and were selected on the basis of previous exploratory analyses in separate patient/control samples. Individual assay values were combined mathematically to yield an MDDScore. A ‘positive’ test, (consistent with the presence of MDD) was defined as an MDDScore of 50 or greater. For the Pilot Study, 36 MDD patients were recruited along with 43 non-depressed subjects. In this sample, the test demonstrated a sensitivity and specificity of 91.7% and 81.3%, respectively, in differentiating between the two groups. The Replication Study involved 34 MDD subjects, and yielded nearly identical sensitivity and specificity (91.1% and 81%, respectively). The results of the present study suggest that this test can differentiate MDD subjects from non-depressed controls with adequate sensitivity and specificity. Further research is needed to confirm the performance of the test across various age and ethnic groups, and in different clinical settings.
So, what did they really do? Well, let us go through the info applying the PICO framework. The population (P) is people with a major depressive disorder as diagnosed by clinical criteria. The intervention (I) is the new multi-assay serum test for 9 biomarkers associated with depression. The comparator (C) is the clinical diagnosis of MDD, and the outcome (O) is the concordance of the serum test and the clinical diagnosis. OK so far?

Bear in mind that there were actually two studies, and here is how they played out. For the first study the researchers recruited 36 patients with MDD (disease present) and 43 subjects without MDD (disease absent). Given the sensitivity of 91.7% and specificity of 81.3%, here are the results:

Disease present
Disease absent
Total
Test+
33
8
41
Test-
3
35
38
Total
36
43
79




Based on these numbers, the positive predictive value is 80.4% and the negative predictive value is 92.1%. What does this mean? This means that in a population with a 45.6% (36/79) prevalence of MDD, only 20% of all positive tests will be false positives, or identifying the disease when it is absent. Conversely, of all negative tests, 8% will be false negative, or missing the disease when it is present. And for the second study, where 34 MDD patients were involved, frankly not enough information is given in the abstract to say anything about it -- I do not have the denominator (the total pool of subjects including those with and without MDD), and therefore cannot say anything about the PPV or NPV.

So what does all of this mean? Well, there are 3 take-home points:
1. When a test is used to diagnose rather than to screen for a disease, you are dealing with a population that has a higher pre-test probability of the disease. So, when the pre-test probability is close to 50%, even a test with suboptimal sensitivity and specificity can be fairly accurate.
2. Your test is only as good as the "gold standard" against which it is being tested. In this case we are talking about a clinical diagnosis of a major depression. The assumption here is that this gold standard test is perfect already. In the absence of anything else to compare it to, it really is: 100% sensitive, 100% specific and quick. How can you improve upon that? And if this is the case, then why do we need a serum test that will give us false results a good part of the time? One argument for this is given here:
...one of the paper’s co-authors said at the very least establishing a physiological link to depression will hopefully get patients to look at their depression as a treatable condition rather than something that’s wrong with their minds. 
But I guess I am not sure that this is really a valid reason for developing a test. It is much like looking for biological mechanism for homosexuality for the purpose of proving that it is OK to be gay. I already know that it is OK, and find its biological origins of mere intellectual curiosity with little practical consequences. Perhaps we just need to change our minds about it, that is all.
3. Finally, would this test be used to screen people for depression? In medicine there is a temptation to go after "the answer" even when the question is rather oblique. In other words, will the testing in the wild of clinical practice really be limited to those with suspected MDD, or is it likely to metastasize into others, those with milder presentations or even those whom the clinician just finds annoying? If it is the latter (and I can almost guarantee that), then we are in deep doo-doo as far as false positive rates are concerned. If you think that we have had an epidemic of depression up until now, just you wait.

My final word for the day is "caution." I want to be very clear that asking scientific questions is never a bad idea, and that the answers do not always have to bring practical or applied value. I just want to inject some caution into the breathless discussion of screening for everything and our dogged search for "hard" evidence.