Showing posts with label screening. Show all posts
Showing posts with label screening. Show all posts

Tuesday, August 14, 2012

BTL reader question: How do you get to 2%?

I have started a FAQ page on the BTL book web site here, and I will cross-post the discussion here on the blog. This will give us an opportunity to have a more interactive discussion, if necessary, with additional comments and questions.

Here is the inaugural installment.

On August 13, 2012, this question came in via Twitter:




Well, here is the answer (and thank you for the question, Tia!)

First the problem. At the bottom of page 74 and going on to the top of page 75 I discuss the question posed in a 1978 New England Journal of Medicine paper by Casscells and colleagues to 60 physicians and physicians-in-training at Harvard Medical School. The problem went like this:
 
"If a test to detect a disease whose prevalence is 1/1000 has a false positive rate of 5 per cent, what is the chance that a person found to have a positive result actually has the disease, assuming that you know nothing about the person's symptoms or signs?"

The question clearly mimics a disease screening situation. The answer is simple yet elusive. Let us assume that 1,000 people are tested. Among them only 1 person has the actual disease. However, given that the false positive rate is 5%, we also know that out of the 1,000 people tested, 50 will have a false positive test. Assuming that the single person with the disease also has a positive test, we can expect 51 people to test positive. But since only 1 out of these 51 people with a positive test has the disease, the answer to the question above is 1/51=2%. This is a pretty shocking realization, given that a large plurality of the Harvard doctors and trainees chose 95% as their answer. 

So, be careful not to let your intuition override the data when making medical decisions!


If you like Healthcare, etc., please consider a donation (button in the right margin) to support development of this content. But just to be clear, it is not tax-deductible, as we do not have a non-profit status. Thank you for your support!

Wednesday, July 25, 2012

Medicine as the trolley problem

Are you familiar with the trolley problem? It is an ethics dilemma first formulated by the great Philippa Foot as a part of a series of such dilemmas. Her formulation goes roughly like this. Imagine there is a tram hurtling down a track. If it keeps going straight, it will hit and kill 5 people who are working on that track. The conductor is able to throw a switch and divert the train to another part of the track, where 1 single worker will be killed by the trolley. The question is what should the conductor do? Most people when asked respond that yes, he should throw the switch and sacrifice 1 life to save 5. After all, the net benefit is n=4.

There are literally thousands of alternative formulations of this problem, but one of them from the philosopher Judith Jarvis Thomson merits special consideration. The problem starts out similarly, with 5 lives on a track in potential peril. The vantage point and the solution are quite different, though. Now there is a bridge over the rail track, and a very large man is looking at the tracks from the bridge. One way to stop the train is to throw a heavy object in its path, like this large man, for example. You are on the bridge standing behind the man. Would you be justified in pushing him off the bridge in front of the tram to meet his death in order to spare the 5 workers down the tracks? Most people when faced with this formulation say an emphatic "no." This is somehow puzzling, since the net benefit is the same, n=4, as in the original Foot formulation.

Philosophy professors have puzzled over this difference for decades, and there are several potential explanations for why we respond differently to the two scenarios. One explanation has to do with the proximity of the operator (conductor in the first case and the person doing the pushing in the second) to the sacrificial lamb -- in the first case one is enough removed from the action of killing by merely redirecting the tram, whereas in the second the action is, well, more active, and the operator is actually pushing an innocent person to his death.

Though in some ways the scenarios seem to bear no practical distinction from one another, we see the morals and ethics of each differently. This difference in the view point is instructive to the field of medicine, where it has implications to how policy relates to the individual patient encounter. Here is what I mean.

Suppose you are a policy maker, and you recommend that every woman at age 40 start to receive an annual screening mammogram to reduce deaths from breast cancer. At the population level, if we screen 1,000 women for about 30 years, we will save approximately 8 of them from a breast cancer death. (Yes, it's 8, not 80, and not 800). At the same time, among these 1,000 women, there will be over 2,000 false alarms, and over 150 of these will result in an unnecessary biopsy. Some of these biopsies will incur further complications, though currently we  do not seem to have the data to quantify this risk. But what if even one of these biopsies were to lead to death of or another dire lasting complication in a woman who turned out not to have cancer? And by the way the accounting is not all that different when applying the new USPSTF mammography screening recommendations. Well, then we have the trolley problem, don't we? We are potentially sacrificing 1 individual to save 8. And who does the sacrificing is where the variations of the trolley problem come in.

Payers levy financial penalties on primary care physicians when they fail to comply with screening recommendations in their patient panels. The payer certainly sees this issue as the original formulation of the problem: Why not throw this financial switch to achieve net life savings? But for a clinician who deals with the individual patient this may be akin to pushing her over the bridge toward a potentially fatal event. Because we don't have a crystal ball, we cannot say which woman will die or incur a terrible complication. But the same population data that tell us about benefits must also give us pause when reflecting on the risks. Add the ubiquitous uncertainty (and lack of data) into this equation, and the implications are even more shocking. So, while making policy recommendations based on population data is sensible, policing uniform application of these recommendations to individual patients is fraught: of course, clinicians and patients need to be cautious about making individual decisions even when in population data benefits outweigh risks.

On the surface risk-benefit equations for many interventions may appear favorable, leading to blanket policy recommendations to employ them on everyone who qualifies. In the office, the clinician, caught in a tug of war between mountains of new literature and the ever-shrinking appointment times, is hard-pressed to take the time to consider these recommendations in the context of the individual patient. And furthermore, financial incentives from payers act as a short-hand justification, a "nudge," for doing as recommended rather than for giving it thought. So, who must look out for the patient's interest? The patient, that's who. Who understands the patient's attitude toward the risks and the benefits? The patient, that's who. Who now has to be responsible for making the ultimate informed decision about which track to stand on? The patient, that's who.

For me the trolley problem gives clarity to the reservations that I walk around with every day. I have done a lot of soul searching about why it is that, even if the benefits seem to outweigh the risks, I am still more often than not skeptical about whether a particular intervention is right for me. And since every intervention in medicine has a real risk, though mostly quite low, of going terribly awry, my skepticism is justified. This is my approach to evaluating these risks and benefits, based on my values and my understanding of the data as it is today.

What's the answer to this ethical conundrum in medicine? I cannot see that policy makers will stop throwing the switch in the near future, and so as a society we will be forced to accept the tram's collateral damage. And while this may make sense in an area such as vaccination, where thousands of lives can be saved by sacrificing a very few by throwing the switch, in most everyday less clear-cut medical decisions the answer is less clear-cut. Will doctors rebel against being forced to throw some patients on the tracks in order to save some marginally larger number of others? I don't think that they have the time or the energy or the incentive to do this, since the framing of the switch-throwing is through the rhetoric of "evidence." Right or wrong, doctors are shackled by the stigma of ignorance that comes with not following evidence-based guidelines, and this may act to perpetuate blind compliance. This leaves the patients, for some of whom the right thing will be just to get themselves off the tracks altogether, far away from the hurtling trolley until its brakes are fixed.                        

If you like Healthcare, etc., please consider a donation (button in the right margin) to support development of this content. But just to be clear, it is not tax-deductible, as we do not have a non-profit status. Thank you for your support!

Tuesday, July 10, 2012

DHHS: Does this lie make me look stupid?

Update, July 12, 4:30 PM Eastern:
Just got this extra lame reply from healthfinder:

Dear Ms. Zilberberg, Thank you for contacting healthfinder.gov.  healthfinder is a government Web site featuring prevention and wellness information and tools to help you and those you care about stay healthy. At healthfinder.gov, you will find:
 ·         interactive tools like menu planners and health calculators
·         online checkups
·         printable information that you can share with a family member or take to the doctor.
 healthfinder.gov is coordinated by the Office of Disease Prevention and Health Promotion (ODPHP), U.S. Department of Health and Human Services and the National Health Information Center (NHIC). NHIC links people to organizations that provide reliable health information. All of healthfinder.gov’s topics and tools go through subject matter expert reviews. As a result of these reviews, sentences and wording sometimes get updated and/or changed. This particular topic has already been reviewed, and the content team will be rewording the language; the word “best” will be removed from that sentence. This change will be reflected on the site in the next scheduled healthfinder.gov update. Sincerely, 
healthfinder.gov TeamNational Health Information Centerhealthfinder.gov is coordinated by the Office of Disease Prevention and Health Promotion (ODPHP), U.S. Department of Health and Human Services and the National Health Information Center (NHIC).

HOW ABOUT INCLUDING A DISCUSSION OF SAFE SEX?!!!!!!! Idiotic.




Update July 11, 10:50 AM Eastern
I have just sent the following e-mail to healthfinder.gov at the address healthfinder@nhic.org. I urge everyone who reads this to send them the same or a similar message. And if you do, please, leave a comment below to let everyone know.
Hello, 
I wanted to let you know that the information you posted on this web page on Pap testing is erroneous and misleading. Telling women that the "best" way to prevent cervical cancer is through a regular Pap test is not supported by evidence. The "best" way is to prevent HPV infection by engaging in safe sexual intercourse. As a public health communicator you are doing a tremendous disservice to the public.  
I urge you to change this message to reflect reality. 
Thank you. 
Marya Zilberberg, MD, MPH, FCCP 

There is pounding in my temples, my back muscles are in a spasm, and I might even be turning green and busting out of my clothes. What caused all this? This innocent-looking tweet from the Department of Health and Human Services:


I had to do a double take. My blood started to boil almost immediately. But I persisted, clicked on the link, and saw this:


The first sentence really says "The best way to prevent cervical cancer is to get regular Pap tests." Jaw, meet floor. What does the word "prevent" really mean? I went to The Free Dictionary for enlightenment:



Just as I had suspected: to avert, to keep from happening. And how does a Pap test keep the cancer away? It finds "abnormal cells before they turn into cancer." And where do abnormal cells come from? God, right? Well, no, they are mostly associated with an HPV infection, which comes from exposing yourself to unprotected sexual intercourse, usually with someone whose HPV status you don't know. You see where I am going with this? The message here is that there is nothing more effective at preventing cervical cancer than having a Pap test to detect early changes and lop out the misbehaving piece of your cervix. Are they serious? Is this really the "best way"? Let's examine the meaning of "best":


   
I guess beauty (and value) are in the eye of the beholder. Does subjecting yourself to a surgical procedure that may leave your cervix unable to help your uterus to maintain a pregnancy qualify as "surpassing all others in excellence" or as "most desirable"? Not in my book, not when a little advanced planning and a nickel for a condom could could keep that horse from leaving the barn in the first place. True prevention does not take place in a doctor's office, and it is a mistake to equate screening to prevention.

Come on, DHHS, who writes your stuff? Fire them! You are risking your credibility. What's next? "Bulimia is the best way to prevent obesity"?    

If you like Healthcare, etc., please consider a donation (button in the right margin) to support development of this content. But just to be clear, it is not tax-deductible, as we do not have a non-profit status. Thank you for your support!

Tuesday, June 12, 2012

Healthfinder.gov: Education or indoctrination?

Ever heard of healthfinder.gov? It's a web site from the US Department of health and Human Services
...where you will find information and tools to help you and those you care about stay healthy.
Sounds like a laudable goal, right? Great! Now, help me! Here is the "help" that I found when I went to the page called "Colorectal Cancer Screening: Questions for the doctor":

What do I ask the doctor?

It helps to have questions for the doctor written down ahead of time. Print out these questions and take them to your next appointment. You may want to ask a family member or close friend to come with you to take notes.
So far so good. But here is the list that follows:
  • What puts me at risk for colorectal cancer?
  • When do I need to start getting tested?
  • How often do I need to get tested?
  • What screening test do you recommend? Why?
  • What’s involved in screening? How do I prepare?
  • Are there any dangers or side effects involved?
  • How long will it take to get the results?
  • What can I do to reduce my risk of colorectal cancer?
Note the wording: "When do I need to start getting tested?" "How often do I need to get tested?" And these "needs" come well before the "why?" In fact, the "why" never really comes. The oblique "why" about which test is recommended is too little too late. The real "why" is why, or even whether, I need to get tested in the first place. I am happy to see a question on the dangers of screening, but again it leaves plenty of room for the clinician to minimize and patronize.

The list of questions is built upon one (erroneous) assumption: Everyone is bound to perceive the risk-benefit equation of colorectal cancer screening the same way. We know this is false, and each person needs to make an individual decision based in what we know today and according to the values he/she places on the outcomes. The way the questions are written, they simply reinforce the bullying attitude of the screening bias, making those who swim against this tide feel irrational and unreasonable. But may I point out that some of us spoke out against universal mammography screening even before it became the main-stream recommendation? So perhaps there are good reasons to be more cautious with screening for everything, even colon cancer.

Science evolves, our knowledge evolves. What we think we know today will be modified tomorrow. I take a strong exception to this dogmatic and one-sided formulation of how to have a discussion about testing whose risk and benefit profile may not (and should not) elicit the same unbridled enthusiasm from everyone. So please, healthfinder.gov, rethink your "helpful" questions so as to educate, rather than indoctrinate.

Hat tip to @DCPatient for pointing me to this page  

If you like Healthcare, etc., please consider a donation (button in the right margin) to support development of this content. But just to be clear, it is not tax-deductible, as we do not have a non-profit status. Thank you for your support!

Thursday, March 22, 2012

More on "mammography saves lives" story


A story in HealthDay with the title of "Two Studies Find Routine Mammography Saves Lives" talks about studies presented at the annual meeting of the American Society for Clinical Oncology (ASCO) European Breast Cancer Conference (EBCC). The studies, both from the Netherlands, allegedly showed that mammography does indeed save lives. If true, these data would contrast with the preponderance of evidence that has stirred up such a ruckus recently about the utility of mammography, complete with references to rationing and death panels. But let's look at what's reported in today's story more closely.

Cutting through all the definitive bravado, here is a little piece of science that was reported:
Compared with the pre-screening period 1986 to 1988, deaths from breast cancer among women aged 55-79 fell by 31 percent in 2009," Jacques Fracheboud, a senior researcher at the Erasmus University Medical Center in Rotterdam, said in a meeting news release. We found there was a significant change in the annual increase in breast cancer deaths: before the screening program began, deaths were increasing by 0.3 percent a year, but afterwards there was an annual decrease of 1.7 percent," he added. "This change also coincided with a significant decrease in the rates of breast cancers that were at an advanced stage when first detected.
Note, the reference is to deaths from breast cancer without any mention of all-cause mortality. (You can read about why the latter is important here.)

The report next states that over the first 20 years of the screening program
... 13.2 million breast cancer screening examinations were performed among 2.9 million women (an average of 4.6 examinations per woman), resulting in nearly 180,000 referral recommendations, nearly 96,000 biopsies and more than 66,000 breast cancer diagnoses.
So, doing the math, I come up with about 31% false positive rate at the biopsy stage (that's 96,000 biopsies minus 66,000 positives for cancer, all divided by the 96,000 total biopsies). If we use the 180,000 "referral recommendations" as our denominator of all positive tests, and stick with the 66,000 true positive rate, then the false positives grow to (180,000-66,000)/180,000 = 0.63, or 63%. If we spread the 33,000 false positives over the 13.2 million examinations, that equates to 0.25% chance for a false positive. Yet the report goes on to say that (emphasis mine):
For a woman who was 50 in 1990 and had 10 screenings over 20 years, the cumulative risk of a false-positive result (something being detected that turned out not to be breast cancer) was 6 percent.
Six percent? This is clearly a place where my high school math teacher's mantra of "show your work" is applicable.

The next piece of information that I would like to understand better is this:
Over-diagnosis (detection of breast tumors that would never have progressed to be a problem) occurred in 2.8 percent of all breast cancers diagnosed in the total female population and 8.9 percent of screening-detected breast cancers."
How exactly was this computed? Again a case for "show-your-work."

And then there is this (emphasis mine):
Regular screening "decreases deaths by over 30 percent, [with] limited harm and reasonable costs. Additionally, cancers are detected at an earlier stage, which means not only decreased mortality but also morbidity; the patient may not have to have chemotherapy or a mastectomy," she noted.
OK, so, if I got it right, it is breast cancer mortality that is decreased by 31%, not all-cause mortality. This really should have been spelled out more clearly, not to mention that the actual, or absolute, reduction likely pales in comparison to this relative drop. And what about diagnosing earlier stage disease? Lead time bias, anyone?

The second study was a computer model, and I will not go through it at this time as I need to move on to other work. But you get the picture: the numbers given in the report are limited and at times they don't add up. Mixing up cancer mortality with all-cause mortality leads to erroneous conclusions. And finally, forgoing reporting on the absolute risk reduction in favor of the inflated relative reduction is not helpful for understanding the true risks involved.

One final thought: Yes, I do have cognitive biases, and it is difficult for me to avoid them. I happen to fall into the camp that thinks screening for sublclinical diseases, at least in our current technological setting, is disease mongering. At the same time, I would like to think that if the data really showed a significant benefit without great risks, I would give them a second look.

The bottom line is this: at least for me, the report confused the issue more than it has clarified. Perhaps the study, once published, will answer all of the questions that I have posed adequately. But at this stage, it is a shame that such strong statements as...
"These results show why mammography is such an effective screening tool," said one U.S. expert, Dr. Kristin Byrne, chief of breast imaging at Lenox Hill Hospital in New York City. She was not involved in the new research.
 and this...
"We are convinced that the benefits of the screening program outweigh all the negative effects," Fracheboud said.
 ... are not backed up by appropriate evidence.

h/t to @ElaineSchattner for the story


If you like Healthcare, etc., please consider a donation (button in the right margin) to support development of this content. But just to be clear, it is not tax-deductible, as we do not have a non-profit status. 

Thank you for your support!

Thursday, March 15, 2012

PSA screening: Does it or doesn't it?

A study in the NEJM reports that after 11 years of follow up in a very large cohort of men randomized either to PSA screening every 4 years (~73,000 subjects) or to no screening (~89,000 subjects) there was both a reduction in death and no mortality advantage. How confusing can things get? Here is a screenshot of today's headlines about it from Google News:






























How can the same test cut prostate cancer deaths and at the same time not save lives? This is counter-intuitive. Yet I hope that a regular reader of this blog is not surprised at all.  For the rest of you, here is a clue to the answer: competing risks.

What's competing risks? It is a mental model of life and death that states that there are multiple causes competing to claim your life. If you are an obese smoker, you may die of a heart attack or diabetes complications or a cancer, or something altogether different. So, if I put you on a statin and get you to lose weight, but you continue to smoke, I may save you from dying from a heart attack, but not from cancer. One major feature of the competing risks model that confounds the public and students of epidemiology alike is that these risks can actually add up to over 100% for an individual. How is this possible? Well, the person I describe may have (and I am pulling these numbers out of thin air) a 50% risk of dying from a heart attack, 30% from lung cancer, 20% from head and neck cancer, and 30% from complications of diabetes. This adds up to 130%; how can this be? In an imaginary world of risk prediction anything is possible. The point is that he will likely die of one thing, and that is his 100% cause of death.

Before I get to translating this to the PSA data, I want to say that I find the second paragraph in the Results section quite problematic. It tells me how many of the PSA tests were positive, how many screenings on average each man underwent, what percentage of those with a positive test underwent a biopsy, and how many of those biopsies turned up cancer. What I cannot tell from this is precisely how many of the men had a false positive test and still had to undergo a biopsy -- the denominators in this paragraph shape-shift from tests to men. The best I can do is estimate: 136,689 screening tests, of which 16.6% (15,856) were positive. Dividing this by 2.27 average tests per subject yields 6,985 men with a positive PSA screen, of whom 6,963 had a biopsy-proven prostate cancer. And here is what's most unsettling: at the cut-off for PSA level of 4.0 or higher, the specificity of this test for cancer is only 60-70%. What this means is that at this cut-off value, a positive PSA would be a false positive (positive test in the absence of disease) 30-40% of the time. But if my calculations are anywhere in the ballpark of correct, the false positive rate in this trial was only 0.3%. This makes me think that either I am reading this paragraph incorrectly, or there is some mistake. I am especially concerned since the PSA cut-off used in the current study was 3.0, which would result in a rise in the sensitivity with a concurrent decrease in specificity and therefore even more false positives. So this is indeed bothersome, but I am willing to write it off to poor reporting of the data.

Let's get to mortality. The authors state that the death rates from prostate cancer were 0.39 in the screening group and 0.50 in the control group per 1,000 patient-years. Recall from the meat post that patient-years are roughly a product of the number of subjects observed by the number of years of observation. So, again, to put the numbers in perspective, the absolute risk reduction here for an individual over 10 years is from 0.5% to 0.39%, again microscopic. Nevertheless, the relative risk reduction was a significant 21%. But of course we are only talking about deaths from prostate cancer, not from all other competitors. And this is the crux of the matter: a man in the screening group was just as likely to die as a similar man in the non-screening group, only causes other than prostate cancer were more likely to claim his life.

The authors go through the motions of calculating the number needed to invite for screening (NNI) in order to avoid a single prostate cancer death, and it turns out to be 1,055. But really this number is only meaningful if we decide to get into death design in a something like "I don't want to die of this, but that other cause is OK" kind of a choice. And although I don't doubt that there may be takers for such a plan, I am pretty sure that my tax dollars should not pay for it. And thus I cast my vote for "doesn't."       

Wednesday, February 22, 2012

Endometriosis and cancer: When a "breakthrough" may not be all that

There is an interesting new study that was just published online at the Lancet Oncology. It is a pooled analysis of a bunch of case-control studies to explore the association between endometriosis and certain types of ovarian cancer. Here is the abstract:

Background

Endometriosis is a risk factor for epithelial ovarian cancer; however, whether this risk extends to all invasive histological subtypes or borderline tumours is not clear. We undertook an international collaborative study to assess the association between endometriosis and histological subtypes of ovarian cancer.

Methods

Data from 13 ovarian cancer case—control studies, which were part of the Ovarian Cancer Association Consortium, were pooled and logistic regression analyses were undertaken to assess the association between self-reported endometriosis and risk of ovarian cancer. Analyses of invasive cases were done with respect to histological subtypes, grade, and stage, and analyses of borderline tumours by histological subtype. Age, ethnic origin, study site, parity, and duration of oral contraceptive use were included in all analytical models.

Findings

13 226 controls and 7911 women with invasive ovarian cancer were included in this analysis. 818 and 738, respectively, reported a history of endometriosis. 1907 women with borderline ovarian cancer were also included in the analysis, and 168 of these reported a history of endometriosis. Self-reported endometriosis was associated with a significantly increased risk of clear-cell (136 [20·2%] of 674 cases vs 818 [6·2%] of 13 226 controls, odds ratio 3·05, 95% CI 2·43—3·84, p<0·0001), low-grade serous (31 [9·2%] of 336 cases, 2·11, 1·39—3·20, p<0·0001), and endometrioid invasive ovarian cancers (169 [13·9%] of 1220 cases, 2·04, 1·67—2·48, p<0·0001). No association was noted between endometriosis and risk of mucinous (31 [6·0%] of 516 cases, 1·02, 0·69—1·50, p=0·93) or high-grade serous invasive ovarian cancer (261 [7·1%] of 3659 cases, 1·13, 0·97—1·32, p=0·13), or borderline tumours of either subtype (serous 103 [9·0%] of 1140 cases, 1·20, 0·95—1·52, p=0·12, and mucinous 65 [8·5%] of 767 cases, 1·12, 0·84—1·48, p=0·45).

Interpretation

Clinicians should be aware of the increased risk of specific subtypes of ovarian cancer in women with endometriosis. Future efforts should focus on understanding the mechanisms that might lead to malignant transformation of endometriosis so as to help identify subsets of women at increased risk of ovarian cancer.

Funding

Ovarian Cancer Research Fund, National Institutes of Health, California Cancer Research Program, California Department of Health Services, Lon V Smith Foundation, European Community's Seventh Framework Programme, German Federal Ministry of Education and Research of Germany, Programme of Clinical Biomedical Research, German Cancer Research Centre, Eve Appeal, Oak Foundation, UK National Institute of Health Research, National Health and Medical Research Council of Australia, US Army Medical Research and Materiel Command, Cancer Council Tasmania, Cancer Foundation of Western Australia, Mermaid 1, Danish Cancer Society, and Roswell Park Alliance Foundation.
Alas, I do not have access to the full article (paywall), so cannot go though it in a detailed way. Nevertheless, we can try to put these findings in perspective. So, briefly, the investigators put together data from many case-control studies and discovered that the risk of some, though not all, ovarian cancers was 2-3 times higher in the presence of endometriosis than in its absence, and concluded that clinicians should be aware of this increase in risk. Fair? But you know I am going to deconstruct it, right? Here we go.

By now we all understand what a case-control study is, right? It is a study where cases (those patients with the disease of interest) are compared to controls (subjects who are in all ways the same as the cases with the exception that they do not harbor the disease in question). So the study identifies subjects with an outcome, and follows them backward to the exposure that is of interest vis-a-vis this outcome. These studies are notoriously difficult to do well, particularly when it comes to the choice of a control, and very few do it well in my experience. I cannot comment on the current conglomeration of 13 of them, so will not venture a guess on whether some or how many of them may lead us astray. Though these studies are difficult, for various reasons they are the way to go when examining an uncommon outcome. So the choice of the design is legit.

No, let's examine the risk for misclassification for both the disease and the exposure. I think you will agree that a case of ovarian cancer is difficult to misclassify, so I will not pick on this as a potentially major threat to the validity. But what about endometriosis? This is a chameleonic condition that is probably way under-recognized. Its symptoms and signs are varied and, unless the studies required a look "inside," which I sincerely doubt they did -- note, the abstract states that this exposure was self-reported -- there is a very real and grave threat to validity here. Ever heard of recollection bias? If such exists, then more women with cancer are likely to report symptoms that may be indicative of endometriosis than those without cancer. So, the observed increase in the risk of ovarian CA in the presence of endometriosis may be due to just that -- a recollection bias.

These limitations notwithstanding, the authors felt that the study was a breakthrough (emphasis mine):
"This breakthrough could lead to better identification of women at increased risk of ovarian cancer and could provide a basis for increased cancer surveillance of the relevant population, allowing better individualization of prevention and early detection approaches such as risk-reduction surgery and screening,” lead author Celeste Leigh Pearce, at the University of Southern California, Los Angeles, said in a journal news release. 
What does Dr. Pearce mean by "increased cancer surveillance?" Is she talking about screening women with endometriosis because of this possibly heightened risk? And if so, is this really wise? Let's simulate some of these numbers.

Let us suppose that endometriosis does indeed increase the risk of ovarian cancer 2-3-fold. This means that the incidence now goes from 13 per 100,000 women up to 39 per 100,000. Let us now also assume that there is a test that is 99% sensitive (able to identify ovarian CA when it is present) and 99% specific (able to demonstrate that no ovarian CA is present when it is not present). Recall that at the population incidence of ovarian CA, the USPSTF does not recommend screening due to a very high risk of a false positive. The question is does this 3-fold elevation in risk change the positive predictive value of screening substantially enough for it now to be recommended? I think I know the answer, but let's go through the exercise anyway, just to be explicit.
Disease present
Disease absent
Total
Test+
39
1,000
1,038
Test-
0
98,961
98,962
Total
39
99,961
100,000

So, the corresponding positive predictive value is... drum roll, please... 3.7%. This means that out of 100 women who have tested positive, fully 96 have a false positive result and are now likely to be subjected to invasive procedures. If we imagine that a test can have near-perfect specificity of 99.99% (no test that I know of can come close to this in any consistent way), still 20% of all positive results are false positives. So, is this indeed food for screening thought? I really don't think so, particularly given that the current risk calculation is likely a gross over-estimate.

So, there you have it. I don't think I am engaging in hyperbole when I say that "breakthrough" is very likely an overstatement.                  
  

Saturday, February 18, 2012

The implications of a blood test for depression

So, a coupe of days ago we spent considerable (virtual) ink on discussing the risk of a false positive result in the setting of screening for rare events. Today something else has caught my eye: blood test for depression. The Atlantic reported this much-retweeted piece on the same day that I was droning on about lung cancer screening. So what is this about, and does it bear any similarity to what we discussed here?

Let us examine the lede of the Atlantic article:
New research shows that blood screenings can accurately spot multiple telltale biomarkers in patients with classic symptoms of depression.
The writer uses the word "screening" while talking about patients with "classic symptoms of depression." This choice of language is problematic. In clinical medicine the term "screening" generally refers to a population without any signs or symptoms of disease. Think of breast cancer and prostate cancer screening. When symptoms or signs are present, the testing becomes diagnostic, not screening, in its purpose. This difference is actually critical to appreciate in the context of what we talked about in the lung cancer screening post. The presence of signs that make a disease suspect presumably increase the pre-test probability of that disease. This means that the prevalence of this disease is higher in the population that has these particular signs/symptoms than in the overall population without them. And recall that it is this very pre-test probability that drives the predictive value of a positive test. Namely, the higher the pre-test probability of the disease, the more credence we can put in a positive test result.

OK, so let's move on to the data. I went to the primary source, but all I could get to was the abstract (paywall and all), so bear in mind that I do not have all of the data. I am reproducing the abstract here for your convenience:
Despite decades of intensive research, the development of a diagnostic test for major depressive disorder (MDD) had proven to be a formidable and elusive task, with all individual marker-based approaches yielding insufficient sensitivity and specificity for clinical use. In the present work, we examined the diagnostic performance of a multi-assay, serum-based test in two independent samples of patients with MDD. Serum levels of nine biomarkers (alpha1 antitrypsin, apolipoprotein CIII, brain-derived neurotrophic factor, cortisol, epidermal growth factor, myeloperoxidase, prolactin, resistin and soluble tumor necrosis factor alpha receptor type II) in peripheral blood were measured in two samples of MDD patients, and one of the non-depressed control subjects. Biomarkers measured were agreed upon a priori, and were selected on the basis of previous exploratory analyses in separate patient/control samples. Individual assay values were combined mathematically to yield an MDDScore. A ‘positive’ test, (consistent with the presence of MDD) was defined as an MDDScore of 50 or greater. For the Pilot Study, 36 MDD patients were recruited along with 43 non-depressed subjects. In this sample, the test demonstrated a sensitivity and specificity of 91.7% and 81.3%, respectively, in differentiating between the two groups. The Replication Study involved 34 MDD subjects, and yielded nearly identical sensitivity and specificity (91.1% and 81%, respectively). The results of the present study suggest that this test can differentiate MDD subjects from non-depressed controls with adequate sensitivity and specificity. Further research is needed to confirm the performance of the test across various age and ethnic groups, and in different clinical settings.
So, what did they really do? Well, let us go through the info applying the PICO framework. The population (P) is people with a major depressive disorder as diagnosed by clinical criteria. The intervention (I) is the new multi-assay serum test for 9 biomarkers associated with depression. The comparator (C) is the clinical diagnosis of MDD, and the outcome (O) is the concordance of the serum test and the clinical diagnosis. OK so far?

Bear in mind that there were actually two studies, and here is how they played out. For the first study the researchers recruited 36 patients with MDD (disease present) and 43 subjects without MDD (disease absent). Given the sensitivity of 91.7% and specificity of 81.3%, here are the results:

Disease present
Disease absent
Total
Test+
33
8
41
Test-
3
35
38
Total
36
43
79




Based on these numbers, the positive predictive value is 80.4% and the negative predictive value is 92.1%. What does this mean? This means that in a population with a 45.6% (36/79) prevalence of MDD, only 20% of all positive tests will be false positives, or identifying the disease when it is absent. Conversely, of all negative tests, 8% will be false negative, or missing the disease when it is present. And for the second study, where 34 MDD patients were involved, frankly not enough information is given in the abstract to say anything about it -- I do not have the denominator (the total pool of subjects including those with and without MDD), and therefore cannot say anything about the PPV or NPV.

So what does all of this mean? Well, there are 3 take-home points:
1. When a test is used to diagnose rather than to screen for a disease, you are dealing with a population that has a higher pre-test probability of the disease. So, when the pre-test probability is close to 50%, even a test with suboptimal sensitivity and specificity can be fairly accurate.
2. Your test is only as good as the "gold standard" against which it is being tested. In this case we are talking about a clinical diagnosis of a major depression. The assumption here is that this gold standard test is perfect already. In the absence of anything else to compare it to, it really is: 100% sensitive, 100% specific and quick. How can you improve upon that? And if this is the case, then why do we need a serum test that will give us false results a good part of the time? One argument for this is given here:
...one of the paper’s co-authors said at the very least establishing a physiological link to depression will hopefully get patients to look at their depression as a treatable condition rather than something that’s wrong with their minds. 
But I guess I am not sure that this is really a valid reason for developing a test. It is much like looking for biological mechanism for homosexuality for the purpose of proving that it is OK to be gay. I already know that it is OK, and find its biological origins of mere intellectual curiosity with little practical consequences. Perhaps we just need to change our minds about it, that is all.
3. Finally, would this test be used to screen people for depression? In medicine there is a temptation to go after "the answer" even when the question is rather oblique. In other words, will the testing in the wild of clinical practice really be limited to those with suspected MDD, or is it likely to metastasize into others, those with milder presentations or even those whom the clinician just finds annoying? If it is the latter (and I can almost guarantee that), then we are in deep doo-doo as far as false positive rates are concerned. If you think that we have had an epidemic of depression up until now, just you wait.

My final word for the day is "caution." I want to be very clear that asking scientific questions is never a bad idea, and that the answers do not always have to bring practical or applied value. I just want to inject some caution into the breathless discussion of screening for everything and our dogged search for "hard" evidence.  
           
   

Thursday, February 16, 2012

In medicine, beware of what seems too good to be true

Update 2/17/12:
A reader brought to my attention (thanks!) a very slight inaccuracy in the first table below, which I have corrected. I did the calculations in Excel, which, as you may know, likes to round numbers. 

File this under "misleading." Here is the story:
What's the Latest Development? 
A California start up has developed a breath test that can diagnose lung cancer with a 83 percent accuracy and distinguish between different types of the disease. The procedures which currently exist to test for lung cancer, which is the leading cause of cancer deaths worldwide, result in too many false positives, meaning unnecessary biopsies and radiation imaging. The new devices works by drawing breath "through a series of filters to dry it out and remove bacteria, then [carries it] over an array of sensors."  
What's the Big Idea? 
The company is now testing a version of the machine 1,000 times more accurate than its latest model, which could increase the accuracy of diagnoses to 90 percent, the level likely needed to take the device to market. Because the machine is not specific to a particular group of chemicals, the breath tester could, in principle, test for any disease that has a metabolic breath signature, for example, tuberculosis. "A breath signature could give a snapshot of overall health," says the company's founder, Paul Rhodes. 
Am I just being a luddite by not getting, well, breathless about this? I'll just lay out my argument, and you can be the judge.

There is not doubt that lung cancer is a devastating disease, and we have not done a great job reducing its burden or the associated mortality. However, there are several issues with what is implied above, and some of the assumptions are unclear. First, what does "accuracy" mean? In the world of epidemiology it refers to how well the test identifies true positives and true negatives. If that is in fact what the story means, then 83% may not be bad; we'll regroup on that point at the end late in this post. This brings me to my second point: what is the gold standard that the test is being measured against? In other words, what is it that has the 100% accuracy in lung cancer detection? Is it a chest X-ray, a CT scan, a biopsy, what?

The SEER database, the most rigorous source of cancer statistics in the US, classifies tissue diagnosis as the highest evidence of cancer. However, in some cases a clinical diagnosis is acceptable. The inference of cancer when no tissue is examined is possible when weighing patient risk factors and the behavior of the tumor. So, you see where I am going here? The gold standard is tissue or tumor behavior in a specific patient. Is that what this technology is being measured against? We need to know. And here is another consideration. What if the tissue provides a cancer diagnosis, but the cancer is not likely to become a problem, like in the prostate cancer story, for example?

But all of these issues are but a prelude to what is the real problem with a technology like the one described: the predictive value of a positive test. The story even alludes to this, pointing the finger at other current-day technologies and their rates of false positivity, and away from itself. Yet, in fact, this is the crux of the matter for all diagnostics. Let me show you what I mean.

The incidence of lung cancer in the US is on the order of 60 cases per 100,000 population. Now, let us give this test a huge break and say that it yields (consistently) 99% sensitivity (identifies patients with cancer when cancer is really present) and 99% specificity (identifies patients without cancer when they really do not have cancer). What will this look like numerically given the incidence above if we test 100,000 people?

Cancer present
Cancer absent
Total
Test +
59
999
1,0589
Test -
1
98,941
98,9421
Total
60
99,940
100,000

If we add up all the "wrong" test results, the false negative (n=1) and the false positives (n=999), we arrive at a 1% "inaccuracy" rate, or 99% accuracy. But what is hiding behind this 99% accuracy is the fact that of all those people with a positive test only a handful, a paltry 6%, actually have cancer. And what does this mean to the other 94%? Additional testing, a lot of it invasive. And what does this testing mean for the healthcare system? You connect the dots.

Let's explore a slightly different scenario. Let us assume that there is a population of patients whose risk for developing lung cancer is 10 times higher than the population average. Let us say that their incidence is 600 cases per 100,000 population. Let us perform the same calculation assigning this same bionic accuracy to the test:

Cancer present
Cancer absent
Total
Test +
594
994
1,588
Test -
6
98,406
98,412
Total
600
99,400
100,000
The accuracy remains at 99%, but the value of the positive test rises to 37%. Still, 63% of all people testing positive for cancer will go on to unnecessary testing. And imagine the numbers when we try to screen millions of people, rather than just 100,000.

Let us do just one final calculation. Let us reflect the data back to the test in question, where the article claims that the accuracy of the next version of the technology will be 90%. Assuming a high risk population (600 cases per 100,000 population), what does a positive result mean?

Cancer present
Cancer absent
Total
Test +
540
9,940
10,480
Test -
60
89,460
89,520
Total
600
99,400
100,000
From this table, the accuracy is indeed 90%, concealing the very low value of a positive test of 5%! This means that of the people testing positive for lung cancer with this technology, 95% will be false positives! What is most startling is that to arrive at the same mediocre 37% value for a positive test that we saw above in this population, we would need a population where cancer incidence is a whopping 6,000 per 100,000, or 6%!

I do not want to belabor this issue any further. Screening for disease that is not yet a clinical problem is fraught with many problems, and manufacturers need to be aware of these logic pitfalls. What I have shown you here is that even when the "accuracy" of a test is exquisitely (almost impossibly) high, it is the pre-test probability of, or the patient's risk for the disease that is the overwhelming driver of false positives. Therefore, I give you this conclusion: beware of tests that sound too good to be true -- most of the time they are.

h/t to @gingerly_onward for the story link