Showing posts with label Statistics. Show all posts
Showing posts with label Statistics. Show all posts

Friday, June 13, 2014

Immortality

A few months ago, a paper was published in the International Journal of Epidemiology that caused a sensation in it's home country - Denmark. Using registry data, and the diagnosis of non-melanomatous skin cancer as a proxy for sun exposure, the authors found that the OR for mortality in individuals who developed skin cancer compared to those who did not was 0.53. The p-value was an unfeasibly low 2x10E-308. However, when the results were stratified by age, the OR was a much more reasonable 0.97 in the sun-exposed group. This fact was conveniently forgotten when the study was reported in the media which uncritically stated that people who get more sun will live longer - something which upset the Danish cancer societies immensely.

So why was there such a disparity between the two results and how could it be possible that individuals with cancer could live so much longer (8.5 years on average) than those who did not get cancer. This month, a mea culpa editorial was published in IJE explaining their mistake along with an article discussing the entire issue. What the authors did in this case was fail to account for the immortal time bias. In this study, individuals entered the cohort when they were 40 years old. However, most people did not develop skin cancer until they were in their 60s or 70s. As a result, for the skin cancer cohort, there were approximately 20 years where they could not have died (they had to survive until they developed cancer and were "immortal" until that time). Someone who did not develop skin cancer could have died at any time during that 20 year period. To demonstrate how this works, the authors of the follow-up study, using the same dataset, randomly allocated a "lottery prize" to a proportion of individuals with a mean age of 68 who lived in Denmark over the last 20 years. Again, in this case, a person who received the prize would have had to have lived until the time that it was awarded while those who did not get the prize could die at any time. The results of the simulation study were very similar to the skin cancer study with an OR for all-cause mortality of 0.5 for the prizewinners and a similarly outrageous p-value.

This issue was first described in the 1800s when it was noticed that generals and bishops live longer than lieutenants and curates. Again, this is because one has to survive to an older age to become a bishop or a general and not because there is something inherent in these ranks that lead to an improvement in mortality.

This example is particularly egregious and the editors of IJE have to be congratulated for the way that they dealt with it. However, there are more subtle examples, one of which, highlighted in the follow-up article, appeared in JASN in 2010. This paper found that survival after transplant failure was improved by nephrectomy. However, the mean follow-up in this paper was only 2.93 years and the mean time to nephrectomy was 1.66 years. Thus, about half of the follow-up in individuals who had a nephrectomy was immortal time - they could not die during that follow-up period because they had to survive to the time of surgery. Most of the difference could be accounted for by this bias.

Similarly, a study in JASN in 2007, found that individuals enrolled in a multidisciplinary care clinic were more likely to survive than those who were not. Again, the time of the first MDC clinic was about 1 year after enrollment in the study. Individuals who entered the MDC program had a year of immortal time compared to individuals who did not enter MDC. This was pointed out in a follow-up article in KI. The authors of the original paper reanalyzed their data to account for this bias and found that MDC was still associated with improved survival although the magnitude of the effect was substantially less.

This is a fascinating issue and probably affects more cohort studies than we think. As a reviewer, I'll certainly try to look out for it in the future.

Tuesday, November 20, 2012

Miracle Drug - Part 2

In a previous post, we discussed the dangers of subgroup analyses. Specifically, performing increasing numbers of post-hoc subgroup analyses increases the risk of a type 1 error while the inevitable reduction in numbers tested in subgroup analyses decreases power and increases the risk of a type 2 error.

This is a particular problem in nephrology where there is a lack of randomized trials and a lot of our information is gleaned from subgroup analyses of trials that were never intended to specifically test patients with CKD. A paper was recently published in CJASN which examined the issue of subgroup analysis in the nephrology literature and suggested some potential guidelines for reporting:

1. Plan subgroup analyses during the study's design period
2. Limit the analyses to variables where a plausible hypothesis exists for why there might be a difference between groups
3. In the methods section, list the pre-specified analyses and explain the rationale for each analysis
4. Do not report subgroup analyses in the abstract
5. Use statistical tests for interaction and report effect estimates (confidence intervals rather than p-values)
6. Avoid over-interpretation of subgroup analysis results, explain that this is hypothesis-generating rather than definitive and emphasize the overall study results

The idea is to prevent fishing - performing multiple post-hoc analyses until one is found that is positive. Technically, if one is going to do this, there should be a correction for multiple statistical testing but this is rarely done. The authors of the review examined published randomized trials in 4 major journals and found that about 1/3 of nephrology papers published subgroup analyses (which was actually below average). However, more than 3/4 of these had issues with the analysis including the lack of pre-specified subgroups and inappropriate statistical testing. 

The take home from this for the general audience who will not necessarily be performing these analyses is to carefully examine any trial that makes far-reaching claims based on a subgroup analysis. Remember that in the majority of cases, the trials were never designed to look at these subgroups and that these results should not be regarded the same way as the primary results of the papers. Similarly, a negative result in a subgroup analysis should be examined carefully to determine if the study had sufficient power to come to this conclusion. This is certainly not to say that they should never be done. Some extremely valuable information has come from these analyses and one should never assume that all groups will respond in the same way to the same treatments.

Wednesday, May 30, 2012

Miracle Drug?



The above figure compares long term survival in a subgroup of a trial that was published in Circulation in 1980. This was a randomized controlled trial of just over 1000 patients with known cardiovascular disease who were treated with medical therapy alone. The patients were randomized to two treatment groups and were followed for 5 years. In the primary analysis, there was no statistically significant difference in survival between the two groups. However, a subgroup analysis that compared only patients with 3-vessel disease and LV dysfunction at baseline (~200 patients in each group) found that the outcomes were significantly better in group B (p=0.025).

So what was this treatment that was so successful in reducing mortality in group B? There was no treatment. The patients were randomized to the two groups and then simply followed with usual therapy. This study was designed to show the danger of subgroup analyses and why they should be taken with a grain of salt. When the authors looked deeper into the data, it became apparent that the patients in group B were not as sick as those in group A and that the survival difference was non-significant in a multivariable analysis. However, when they further stratified the patients by only including those with no history of congestive heart failure, the difference between the groups became more significant and remained significant in the multivariable analysis (p=0.01).

We are often confronted by negative clinical trials in nephrology and other disciplines and there is a natural tendency to try and find something useful when these trials are completed. Like any multiple comparison, if you do enough subgroup analyses, you will eventually find one that is significant. Any good statistician will tell you that this needs to be accounted for in the final analysis but this is not necessarily always done. Think about this trial when you are reading about the wonderful effects of a treatment that was negative for most but efficacious for a small group of patients with very specific attributes.