Showing posts with label outcome measures. Show all posts
Showing posts with label outcome measures. Show all posts

Saturday, 30 March 2013

Simulation Based Confidence Intervals for Functions with Complicated Derivatives


Micha Mandel has a new paper in the American Statistician. This concerns the simulation based Delta method approach to obtaining asymptotic confidence intervals which has been used quite extensively for obtaining confidence intervals of transition probabilities in multi-state models. This essentially involves assuming and constructing confidence intervals for by simulating and then considering the empirical distribution of

A formal justification for the approach is laid out and some simulations of its performance compared to the standard Delta method in some important cases is given. It is stated that the utility of the simulation method is not in situations when the Delta method fails, but in situations where the calculating the derivatives needed for the Delta method is difficult. In particular, it still requires the functional g to be differentiable. This seems to down play the simulation method slightly. One advantage of the simulation method is that it is not necessary to linearize g around the MLE. To take a pathological example consider data that are . Suppose we are interested in estimating . When the coverage of the Delta method confidence interval will be anti-conservative because has a point of inflection at . The simulation based will work fine in this situation (as obviously would just constructing a confidence interval for and just cubing the upper and lower limits!).

Friday, 26 October 2012

Constrained parametric model for simultaneous inference of two cumulative incidence functions


Haiwen Shi, Yu Cheng and Jong-Hyeon Jeong have a new paper in Biometrical Journal. This paper is somewhat similar in aims to the pre-print by Hudgens, Li and Fine in that it is concerned with parametric estimation in competing risks models. In particular, the focus is on building models for the cumulative incidence functions (CIFs) but ensuring that the CIFs sum to less than 1 at the asymptote as time tends to infinity. Hudgens, Li and Fine dealt with interval censored data but without covariates. Here, the data are assumed to be observed up to right-censoring but the emphasis is on simultaneously obtaining regression models directly for each CIF in a model with two risks.

The approach taken in the current paper is to assume that the CIFs will sum to 1 at the asymptote, to model the cause 1 CIF using a modified three-parameter logistic function with covariates via an appropriate link function. The CIF for the second competing risk is assumed to also have a three-parameter logistic form, but covariates only affect this CIF through the probability of this risk ever occurring.

When a particular risk in a competing risks model is of primary interest, the Fine-Gray model is attractive because it makes interpretation of the covariate effects straightforward. The model of Shi et al seems to be for cases where both risks are considered important, but still seems to require that one risk be considered more important. The main danger of the approach seems to be that the model for the effect of covariates on the second risk may be unrealistic, but will have an effect on the estimates for the first risk. If we only care about the first risk the Fine-Gray model would be a safer bet. If we care about both risks it might be wiser to choose a model based on the cause-specific hazards, which are guaranteed to induce a model with well behaved CIFs albeit at the expense of some interpretability of the resulting CIFs.

Obtaining a model with a direct CIF effect for each cause seems an almost impossible task because, if we allow a covariate to effect the CIF in such a way that a sufficiently extreme covariate leads to a CIF arbitrarily close to 1, it must be having a knock-on effect on the other CIF. The only way around this would be to have a model that assigns maximal asymptote probabilities to the CIFs at infinity that are independent of any covariates e.g. where are increasing functions taking values in [0,1] and . The need to restrict the to be independent of covariates would make the model quite inflexible however.

Tuesday, 14 August 2012

Absolute risk regression for competing risks: interpretation, link functions, and prediction

Thomas Gerds, Thomas Scheike and Per Andersen have a new paper in Statistics in Medicine. To a certain extent this is a review paper and considers models for direct regression on the cumulative incidence function for competing risks data. Specifically models of the form where is a known link function and is the cumulative incidence function for event 1 given covariates X. The Fine-Gray model is a special case of this class of models, where a complementary log-log link is adopted. Approaches to estimation based on inverse probability of censoring weights and jackknife based pseudo-observations are considered. Model comparison based on predictive accuracy as measured through Brier score and model diagnostics based on extended models allowing time dependent covariate effects are also discussed.
The discussion gives a clear account of the various pros and cons of direct regression of the cumulative incidence functions. In particular, an obvious, although perhaps not always sufficiently emphasized issue is that if, in a model with two causes, a Fine-Gray (or other direct model) is fitted to the first cause, and another to the second cause, the resulting predictions will not necessarily have the property that an issue that is not problematic if the second cause is essentially a nuisance issue, but obviously problematic if both causes are of interest. In such cases regression of the cause-specific-hazards is preferable even if it makes interpreting the effect on the cumulative intensity functions more difficult.

Wednesday, 18 April 2012

Bayesian inference of the fully specified subdistribution model for survival data with competing risks

Miaomiao Ge and Ming-Hui Chen have a new paper in Lifetime Data Analysis. This considers methods for Bayesian inference in the Fine-Gray model for competing risks. In order to perform the Bayesian analysis it is necessary to fully specify the model for the competing risk(s) that are not of direct interest in the analysis. Ge and Chen propose a model in which for failure time ,
and a proportional hazards model affects via
.
Note that does not affect or .
The authors consider approaches to non-parametric Bayesian inference using Gamma process priors, but also use piecewise constant hazard models (with fixed cut points) where the hazard in each time period has an independent gamma prior.

Monday, 26 March 2012

A note on the decomposition of number of life years lost according to causes of death

Per Kragh Andersen has a new paper available as a Department of Biostatistics, Copenhagen Research Report. He shows that the integral of the cumulative incidence function of a particular risk has an interpretation as the expected number of life years lost due to this cause, i.e.

It is argued that this is a more appropriate quantification of the effect of a cause-of-death than using a hypothetical estimate of life expectancy without a particular cause of death, which is reliant on an (untestable) assumption of independent competing risks.

Regression models based around expected "life years lost" are proposed using the pseudo-observations method.

Thursday, 5 January 2012

Estimating treatment effects with treatment switching via semicompeting risks models: an application to a colorectal cancer study

Donglin Zeng, Qingxia Chen, Ming-Hui Chen and Joseph Ibrahim have a new paper in Biometrika. This considers the problem of comparing survival times in clinical trials in which there may be an intermediate disease progression event which may cause the treatment (if initially randomized to control) to be switched. Previous approaches to this problem have been proposed, through univariate survival methods. Here, the authors instead consider a multi-state (semi-competing risks) model. Interest lies in determining the survival distribution under a treatment given no switching.

In line with the other recent Biometrika paper, a pattern-mixture type parametrization is adopted in that a logistic regression component is defined for the probability of progression before death and separate conditional hazard function for time to death given no-progression and time to progression given progression, plus time to death from progression. Switching is assumed to be non-informative of outcome given that progression has occurred and the act of switching has a proportional effect on the hazard of death from progression.

Besides rigorous proofs of results, there doesn't seem to be any substantial conceptual advances in the paper, though it does seem to represent a better approach to the specific problem than the previous approaches.

Tuesday, 13 December 2011

A dynamic model for the risk of bladder cancer progression

Núria Porta, M. Luz Calle, Núria Malats and Guadalupe Gómez have a new paper in Statistics in Medicine. This develops a model for progression of bladder cancer with particular emphasis on predicting future risk given events up to a certain point in time.

In many ways the paper is taking a similar approach to Cortese and Andersen in explicitly modelling a time dependent covariate (here recurrence) in order to obtain predictions.
They fit a semi-parametric Cox Markov multi-state model to the data and define a prediction process


where is the time of the second event, is the type of the second event where P denotes progression and represents the history of the process up to time t. Analogously to outcome measures like the cumulative incidence functions, this predictive process is a function of the transition intensities. They also consider time dependent ROC curves to assess the improvement in classification accuracy that can be achieved by taking into account past history in addition to baseline characteristics.

Tuesday, 10 May 2011

A proportional hazards regression model for the subdistribution with right-censored and left-truncated competing risks data

Xu Zhang, Mei-Jie Zhang and Jason Fine have a new paper in Statistics in Medicine. This covers the same ground as the paper by Geskus in Biometrics, in developing an approach to fitting the Fine-Gray proportional subdistribution hazard model for competing risks data with left truncated and right censored observations by using inverse probability weights (IPW). Bizarrely, the paper makes no reference at all to the Geskus paper. Presumably this is because the paper was first submitted in 2009 before Geskus's work was published (April 2010). However, it is strange that neither the authors nor the referees became aware of the work in the interim (i.e. acceptance of the paper wasn't until March 2011).

What is interesting is the differences between the approach taken in this paper compared to Geskus. The authors work on the basis that since X = min(T,C) is only observable if X > L, where T is the time of failure, L the time of left truncation and C the time of right censoring, the IPW should be calculated conditional on L < X. Zhang et al use a stabilised weight rather than the IPW to reduce the variability in the original weight. The weights they derive seem quite different to Geskus's as they depend on an estimate of overall survival, which will have to depend on the covariates if the semi-parametric model for the subdistribution hazard is to apply.
The authors suggest using Aalen additive hazard models for the overall survival (thus allowing for time varying covariate effects that can ensure the weights are consistent with the proportional subdistribution hazard model).

Zhang et al start from the general case where the truncation and censoring distributions depend on covariates (but are independent conditional on these covariates), though they only detail non-parametric estimates of the weights. Geskus argued that even if the censoring/truncation distribution depended on covariates that didn't imply it was necessary to include these covariates in the weightings.

Given these discrepancies it would be of interest to contrast and compare the two approaches to the same problem. If both approaches are effective, Geskus's seems preferable because the weights are much easier to calculate.

Wednesday, 23 February 2011

Estimating and testing for center effects in competing risks

Sandrine Katsahian and Christian Boudreau have a new paper in Statistics in Medicine. This develops methods for including frailty terms within a Fine-Gray competing risks model in order to account for clustering, e.g. effects of different centres.

Since the Fine-Gray model is essentially just a standard Cox proportional hazards regression model with additional time dependent weights, based on the censoring distribution, for individuals who have had a competing event, methods appropriate for standard Cox frailty models can be readily adapted.

Katsahian and Boudreau closely follow the approach taken by Ripatti and Palmgren (Biometrics, 2000). They assume a Gaussian frailty. Computation of the likelihood requires integrating out the frailty terms. Here this is performed using a Laplace approximation. A difficulty with the Laplace approximation is that it still requires the modal value of the frailty distribution conditional on the data and current values of the parameters. The authors therefore take a profile likelihood approach in which they fix the frailty variance and maximize the likelihood term with respect to both the regression parameters and the frailty terms, . Having obtained, and they can then plug into the Laplace approximation to get the profile likelihood for . The procedure gives a local approximation for which can be used to suggest the updated estimate. Thus the process involves alternating between two Newton-Raphson algorithms until convergence.

Monday, 21 February 2011

Testing Markovianity in the Three-state Progressive Model via future-past Association

Mar Rodríguez-Girondo and Jacobo de Uña Álvarez have a paper currently available as a Universidade de Vigo Discussion paper in Statistics and Operations Research. They develop a test for the Markov property in a progressive three-state model subject to continuous observation up to right-censoring. The test is based on calculating Kendall's Tau at each time point t, which involves calculating the difference between the concordance and discordance probabilities for two pairs of (Z,T) where Z=sojourn in state 1, T=time to entry in state 3, given both subjects are in state 2 at time t, i.e. Z <= t < T. If the Markov property holds we expect Kendall's Tau to stay at around zero. For a non-Markov process tau would vary with time away from 0. An estimator for Tau is developed and a bootstrap resampling algorithm is proposed to estimate a p-value for tau at a fixed time point t, based on independently sampling Z and T from their empirical distributions. A trace of p-values at a grid of time points can then be produced.

Unlike the Cox-proportional hazards approach where a single statistic is produced, here the p-value varies depending on the choice of t chosen. A superior power to the Cox-PH approach was obtained in simulations but only for a good choice of t. An omnibus statistic based on some weighted integral of the absolute value (or square) of tau over the observation range would be useful.

The paper only deals with the progressive 3-state case. It's unclear how the method would be extended to incorporate complicated past history in a more general multi-state model. However, there is scope to extend it to testing whether a particular state within a model has semi-Markov dependencies. In its current form, while the test is an interesting concept, using a simple Cox-PH seems a much more attractive prospect in practice. Update: This paper is now published in Biometrical Journal.

Thursday, 6 January 2011

Analyzing Competing Risk Data Using the R timereg Package

The first paper in the special issue of Journal of Statistical Software is by Thomas Scheike and Mei-Jie Zhang and illustrates the use of the R package timereg to fit the flexible direct binomial regression based competing risks models proposed by Scheike et al, Biometrika 2008. These models, which rely on inverse probability of censoring weights allow both proportional and additive models for the cumulative incidence functions. In particular, this offers a goodness-of-fit test for the Fine-Gray model, testing for time dependence in the covariate effects on the subdistribution hazard.
timereg offers some interesting functionality, for instance confidence bands on the cumulative incidence functions can be computed (albeit via bootstrap resampling).

The package timereg has other very useful functions not directly featured in the paper. In particular, it can fit Aalen additive hazard models which have yet to be used much for modelling transition intensities for multi-state models (Shu & Klein, Biometrika 2005 being the exception).

Friday, 3 December 2010

Interpretability and importance of functionals in competing risks and multi-state models

Per Kragh Andersen and Niels Keiding have a new paper currently available as a Department of Biostatistics, Copenhagen research report. They argue that three principles should be adhered to when constructing functionals of the transition intensities in competing risks and illness-death models. The principles are:

1. Do not condition on the future.
2. Do not regard individuals at risk after they have died.
3. Stick to this world.

They identify several existing ideas that violate these principles. Unsurprising, the latent failure times model for competing risks rightly comes under fire for violating (3), i.e. to say anything about hypothetical survival distributions in the absence of the other risks requires making untestable assumptions. Semi-competing risks analysis where one seeks the survival distribution for illness in the absence of death has the same problem.

The subdistribution hazard from the Fine-Gray model violates principle 2 because it involves the form . Andersen and Keiding say this makes interpretation of regression parameters difficult because they are log(subdistribution hazard ratios). The problem seems to be that many practitioners interpret the coefficients as if they are standard hazard ratios. The authors go on to say that linking covariates directly to cumulative incidence functions is useful. The distinction between this and the Fine-Gray model is rather subtle as in the Fine-Gray model (when covariates are not time dependent): i.e. b is essentially interpreted as a parameter in a cloglog model.
The conditional probability function recently proposed by Allignol et al has similar problems with principle 2.

Principle 1 is violated in the pattern-mixture parametrisation. This is where we consider the distribution of event times conditional on the event type, e.g. the sojourn in state i given the subject moved to state j. This is used for instance in flow-graph Semi-Markov models.

A distinction that is perhaps needed that isn't really made clear in the paper is that there is a difference between violating the principles for mathematical convenience e.g. for model fitting and violating the principles in terms of the actual inferential output. Functionals to be avoided should perhaps be those where no easily interpretable transformation to a sensible measure is available. Thus a pattern-mixture parametrisation for a semi-Markov model without covariates seems unproblematic, since we can retrieve the transition intensities. However, when covariates are present the transition intensities will have complicated relationships to the covariates without an obvious interpretation.

**** UPDATE : The paper is now published in Statistics in Medicine. *****

Monday, 1 November 2010

A regression model for the conditional probability of a competing event: application to monoclonal gammopathy of unknown significance

Arthur Allignol, Aurélien Latouche, Jun Yan and Jason Fine have a new paper in Applied Statistics (JRSS C). The paper concerns competing risks data and develops methods for regression analysis of the probability of a competing event conditional on no competing event having occurred. In terms of the cumulative incidence functions, for the case of two competing events, this can be written as . In some applications this quantity may be more useful than either the cause-specific hazards or the cumulative incidence functions themselves. One approach to regression is this scenario might be to compute pseudo-observations and perform the regression using those. The authors instead propose use of temporal process regression (Fine, Yan and Kosorok 2004), allowing estimation of time dependent regression parameters, by considering the cross-sectional data at each event time.

Tuesday, 17 August 2010

A novel semiparametric regression method for interval-censored data




Seungbong Han, Adin-Cristian Andrei and Kam-Wah Tsui have a paper currently available as a University of Wisconsin Biostatistics and Medical Informatics Department working paper. This essentially extends the concept of pseudo-observations, which up till now have only concerned right-censored, to interval censored survival data. The idea remains the same except that S(t) is estimated using the NPMLE of the survival function (e.g. via Turnbull self-consistent estimator or iterative convex minorant).

The paper is a little disappointing in only providing a very brief heuristic justification for using pseudo-observations in the interval-censoring case. For right-censored data, an estimate of the baseline survival function can be obtained as well as the regression parameters. No discussion of whether this is possible for the interval censored case is given. However, the baseline estimates are likely to be highly unreliable (e.g. non-monotonic) because particular subjects may have extreme influence because they effect where the mass points of the NPMLE occur. For example, the plot above is based on 1000 subjects with survival generated from an exponential with rate 0.25, subject to independent current status observation (uniformly on (0,10)). Estimating the baseline survival from the pseudo-observations (calcuated at times 1,2,3,...,10) leads to a survivor function which increases at time 5. It seems necessary that there should be more consideration of this issue as well as the choice of how many time points to evaluate the pseudo-observations at.

The authors choose to transform the pseudo-observations before regressing on the covariates, rather than using a link function in a GLM. One problem with this approach is presumably that if the estimate of S(t) is 0 or 1, g(S(t))=-Inf or Inf.

On a practical level the authors use Icens to calculate the NPMLE. As noted previously the MLEcens package seems to perform considerably better than Icens and would presumably speed up computation of the pseudo-observations method.

A natural next step would be to consider pseudo-observations for interval censored multi-state data. The lack of non-parametric methods except in a few simple cases is an obvious bar to development in this direction.

Update: The paper has now been published in Communications in Statistics - Simulation and Computation.

Thursday, 17 June 2010

Estimating summary functionals in multistate models with an application to hospital infection data

Arthur Allignol, Martin Schumacher and Jan Beyersmann have a new paper in Computational Statistics. This considers methods for obtaining outcome measures based on estimates of the cumulative transition intensities in nonhomogeneous Markov models. Specifically they consider hospital length of stay data and consider estimating the expected excess time spent in hospital given an infection has occurred by time s, compared to if no infection has occurred by time s. For nonhomogeneous models this is clearly a function of the time s. They consider possible weighting methods to get a reasonable summary measure.

The plausibility of a Markov model is tested informally by including time since infection as a time dependent covariate in a Cox PH model. Some discussion of the extension of the methods to the non-Markov case is given.

Monday, 14 June 2010

Regression analysis of censored data using pseudo-observations

Erik Parner and Per Kragh Andersen have a new paper available as a research report at the Department of Biostatistics, University of Copenhagen. This develops STATA routines for implementing the pseudo-observations method of performing direct regression modelling of complicated outcome measures (such as cumulative incidence functions or overall survival times) for multi-state models subject to right censoring. The paper is essentially the STATA equivalent of the 2008 paper by Klein et al which developed similar routines for SAS and R.

[Update: The paper is now published in The STATA Journal]

Thursday, 18 March 2010

Nonparametric estimator for the survival function of quality adjusted lifetime (QAL) in a three-state illness–death model

Biswabrata and Dewanji have a new paper in the Journal of the Korean Statistical Society. This considers estimation of quality adjusted lifetime (QAL) in illness-death models under right censoring. In contrast to their previous work, estimation is non-parametric rather than parametric. A semi-Markov assumption, where sojourn times in each state are assumed independent, is taken.

Tuesday, 12 January 2010

Estimating disease progression using panel data

Micha Mandel has a new paper in Biostatistics. This considers estimation for panel observed data relating to MS progression. The underlying process can have backward transitions, so reaching a higher state is not in itself considered as evidence of progression. Instead, the patient is required to have stayed in the higher state for some period of time (e.g. 6 months). The quantities of interest in the study is therefore the time taken to first have stayed in state 3 for 6 months. For a continuous time, time-homogeneous Markov process expressions for the mean time and the distribution function are obtained.

As noted by Mandel, the estimates obtained are strongly dependent on the Markov assumption and time homogeneity. This is demonstrated in a simulation study. Mandel notes that more methods for semi-Markov models are required, the recent paper on phase-type semi-Markov models may be of use. Indeed, in the simulations Mandel actually uses phase-type distributions to create a semi-Markov process. However, the MS dataset may be too small to reliably estimate a semi-Markov model.

Informative observation times are a potential problem in the MS study. Mandel shows that observations that occur away from the scheduled 26-week gap time, are more likely to involve a transition, suggesting these observation times are informative. As a result, only observations within a 4 week period of the scheduled visit time are included in the analysis.

A comparison with a discrete-time Markov model approach showed that estimates based on assuming a discrete-time process gave larger estimates of the hitting times.

Monday, 21 December 2009

Estimating life expectancy of demented and institutionalized subjects from interval-censored observations of a multi-state model

Joly, Durand, Helmer and Commenges have a new paper in Statistical Modelling. This concerns the estimation of life expectancy in patients with dementia. Unlike similar analysis by Van den Hout and Matthews, here there are 5 states rather than 3 since they consider instiutionalization as an additional event. Age at death is known up to right-censoring but other transitions are interval censored. Like previous papers by Joly and Commenges, a penalized likelihood approach is taken. The penalized likelihood is approximated by cubic M-splines, with the degree of penalization chosen via an approximate cross-validation score. This allows smooth, flexible non-homogeneous intensities. A Markov assumption is assumed but semi-Markov models can also be fitted provided the model is progressive.

Life expectancies are found by integrating the estimated transition probabilities. Like Van den Hout and Matthews, a parametric bootstrap approach based on simulating from the asymptotic normal distribution of the parameters is used to obtain confidence bands. However, these will typically underestimate variability because the penalization factor is taken to be fixed.

Thursday, 5 November 2009

Analyzing longitudinal data with patients in different disease states during follow-up and death as final state

Le Cessie, de Vries, Buijs and Post have a new paper in Statistics in Medicine. This is concerned with estimating mean quality of life in breast cancer patients at different time points. Standard approaches to analyzing such longitudinal data would be generalized estimating equations (GEE). However, observations are often missing and assuming such data are missing completely at random (MCAR) is unrealistic or even missing at random. In the current study a three-state progressive illness-death model is considered where the illness state refers to presence of a relapse. Both Markov (or clock-forward) and semi-Markov (or clock-reset) models are considered. There was continuous observation of the illness-death process, whereas the quality of life was observed at a common set of time points. The authors propose to model quality of life scores conditional on the state occupied in the multi-state model. A more realistic missingness model can then be adopted by assuming MAR conditional on the occupied state. Inverse probability weighting is used to deal with the missing data. Standard error estimation is performed by bootstrapping.

While the model gives an improved picture compared to ignoring the disease state, the model still makes the assumption that quality of life is dependent on time and current disease state but not on the time since entry into the current disease state.