Thursday, 29 September 2011

A multi-state model for the analysis of changes in cognitive scores over a fixed time interval

Arnold Mitnitski, Nader Fallah, Charmaine Dean and Kenneth Rockwood have a new paper in Statistical Methods in Medical Research. The paper develops a model to describe the trajectory of cognitive function test data. The responses are test scores out of 100, but the authors choose to group responses in 12 states. Additionally subjects may die before the following assessment. Assessments occur at (roughly) equally spaced intervals so a discrete-time model is adopted. Essentially the data are then ordinal longitudinal data.

The novel aspect of the model is to assume that, conditional on survival between time j-1 and j, the state at time j follows a truncated Poisson distribution on , with the mean of the Poisson distribution taken as a linear function of the state at time j and covariates (including age and/or time since baseline measurement). A separate logistic regression model is applied to the deaths. This avoids having to pretend the data are continuous as one might if a linear mixed model were used, and also avoids there being a very large number of unknown parameters, as there would be if a general discrete-time Markov model were applied. However, the truncated Poisson distribution model makes strong assumptions about the conditional distribution of the states, which may or may not be well supported by the data. The authors note that the model fits better if 12 states are used rather than 16. Whether accommodating the proposed model should be a criterion for choosing the number of states to use in the model is questionable.

It is not clear a multi-state model is particularly appropriate for modelling a response with such a large number of responses. It might be better to follow a latent trait approach where for some Normally distributed latent variable that evolves with time deterministically with the addition of stationary Gaussian noise (not necessarily independent), where the are boundary values to be estimated. Similarly existing approaches to modelling MMSE based on a much simpler classification into cognitive-normal and cognitive-impaired with the possibility of misclassification considered (see e.g. Van den Hout and Matthews) are likely to give more meaningful results, even if data from raw MMSE scores is sacrificed.

Wednesday, 24 August 2011

Maximum likelihood analysis of semicompeting risks data with semiparametric regression models

Yi-Hau Chen has a new paper in Lifetime Data Analysis which extends his 2010 JRSS B paper from competing risks to semi-competing risks data. Essentially in both cases the main idea is to model the dependence between the competing risks by assuming their event time distributions are related via some family of copulas. Mathematically this approach is quite elegant as it allows regression models to be built on the marginal distributions of each failure time, with the inherent dependency in the censoring accounted for through the copula. From a practical perspective, particularly with semi-competing risks data and medical applications one has to question the sensibleness of the model and the objective of modelling marginal distributions.

It seems most useful to follow Xu, Kalbfleisch and Tai and view semi-competing risks as an illness-death model. After accounting for covariates, a patient's illness time and death time can be related either due to a shared frailty term, which it may be sensible to assume is determined from the outset, or through onset of illness causing death to occur sooner than it would have done. In the copula model these two distinct factors get pooled together. It is questionable how well the copula model would perform when the true process has a more event determined dependence.

More importantly the question has to be asked why you would want to try and estimate the "illness free" survival distribution? This breaks Andersen and Keiding's guideline to "Stick to this world". Illness (or relapse) is never going to be eliminated. More sensible measures like the cumulative incidence function of death (without illness having occurred) can of course be derived from Chen's copula model, although analogously to the case of semi-parametric models on cause-specific hazards, the effect of covariates on the CIF may be complicated.

Monday, 8 August 2011

Joint modelling of longitudinal outcome and interval-censored competing risk dropout in a schizophrenia clinical trial

Ralitza Gueorguieva, Robert Rosenheck and Haiqun Lin have a new paper in JRSS A. The paper concerns the joint modelling of a longitudinal outcome and an interval censored competing risks outcome that explains drop-out. As is common with these joint longitudinal and survival types of models the two processes are linked via a normally distributed vector of random effects. The novelty of the paper is in the survival part is a competing risks process and the event time is interval censored. The authors adopt a parametric model for the competing risks, using the family of distributions proposed by Sparling et al (Biostatistics, 2006). This makes inference somewhat more straightforward than it would be if a non-parametric baseline cause-specific hazards were used. As recently noted, parametric treatment of competing risks data is surprisingly rare. One problem faced by the authors is that the hazard family of Sparling, while allowing closed form expressions for interval censored univariate survival data, do not result in closed form expressions for interval censored competing risks data (except in special cases). Instead a numerical integral has to be competed. The presence of the overall random effects would mean the likelihood requires nested integration. To avoid this problem the authors adopt an approximation to the true likelihood for competing risks data. If a patient is known to have had a failure of type j in the interval [t0,t1] the authors assume that the patient is censored of all risks except risk j at time t0. It is clear that this approximation will lead to systematic bias as the time at risk from each failure type will be underestimated so the hazards will tend to be overestimated. The amount of bias will depend on the typical length of the intervals [t0,t1].

For the CATIE data example the proposed approximation is probably not an issue. The drop out (competing risks) part of the model is not the primary focus of the inference, and it is really the relative hazards of different types of drop out rather than their absolute values that is important in determining the trajectories of the longitudinal measure without drop out. For instance the estimates for simulated data of a similar type are close to unbiased.
However in extreme cases like current status competing risks data the approximation will do extremely badly.

Current status observation of a three-state counting process with application to simultaneous accurate and diluted HIV test data

Karen McKeown and Nicholas Jewell have a new paper in Canadian Journal of Statistics as part of the Kalbfleisch and Lawless special issue. The paper considers non-parametric inference for three-state progressive models subject to current status data observation. The authors note previous work by van der Laan and Jewell (Annals of Statistics, 2003) that shows that a naive estimator using only information on the first event cannot be improved upon in a fully non-parametric setting. The authors consider situations where additional assumptions are made about the waiting time (between the first and second events). In the motivating example, there is a fairly extreme case of current status data where all subjects are observed at the same time point (i.e. a cross-sectional sample). In this case, if an assumption that the distribution for the first event time is assumed to be locally linear, then only the mean waiting time is relevant. The authors then consider a couple of different scenarios with a view to choosing an optimal assumed mean waiting time that minimizes the mean squared error of some quantity of interest relating to the first event time distribution (e.g. the mean cumulative hazard in some time interval).

Friday, 29 July 2011

Parametric inference for time-to-failure in multi-state semi-Markov models: A comparison of marginal and process approaches

Yang Yang and Vijayan Nair have a new paper in Canadian Journal of Statistics as part of the Kalbfleisch and Lawless special issue. The paper considers inference for progressive semi-Markov processes for complete, right-censored and interval-censored data for parametric models with Gamma or inverse Gamma sojourn distributions. The advantage of considering Gamma or inverse Gamma distributions is that their convolutions have a closed form meaning the overall failure (absorption) time distribution is available in closed form. Inference using only the time-to-failure data is therefore relatively straightforward. The main focus of the paper is investigating the loss in efficiency of only using time-to-failure data when additional data on intermediate states (either through continuous observation or panel data) are available. The authors show that the loss in efficiency can be rather substantial, particularly if interest lies in making predictions about time to failure given an existing process history rather than just mean or median survival. It is therefore suggested that information on intermediate states should be incorporated where possible despite this presenting computationally difficulties in the case of panel data.

Wednesday, 13 July 2011

Combined survival analysis of cardiac patients by a Cox PH model and a Markov chain

Michal Shauly, Gad Rabinowitz, Harel Gilutz and Yisrael Parmet have a new paper in Lifetime Data Analysis. This considers methods for modelling the effect of a mixture of time dependent and time constant covariates on overall survival, with the complication that the time dependent covariates are only observed at a discrete set of time points. They propose to firstly fit a Cox proportional hazard model assuming all covariates are fixed at their baseline values. The main modelling approach is to assume a discrete-time homogeneous Markov model with states corresponding to the combinations of the time dependent covariates (which are categorical or need to be categorized) and death. Transitions between all covariates states are assumed to be possible between each time point. The Cox model is used to determine which of the constant covariates should be considered in the Markov model. For this the authors propose to again categorize the covariates and consider a separate Markov model for each level of the covariates. Having obtained the estimates from the Markov models, it is then possible to calculate expected survival times for patients conditional on their baseline characteristics.

In general the approach proposed is reasonably sensible. However, there is panel data available for the time dependent covariates. It therefore seems possible to fit a time continuous Markov model to the data using methods appropriate for panel data (e.g. Kalbfleisch and Lawless, 1985). This approach has the advantage that the exact time of the death events can still be used.

The authors rely on categorization throughout. While this seems necessary for the time dependent covariates, there seems scope for using multinomial logit models for other covariates. Similarly, by allowing a different mortality probability for each combination of covariates they are effectively fitting covariate models with interactions (i.e. the effect of being in covariate level 2 compared to 1, is different depending on which level(s) of the other time dependent covariate(s) a subject is in). While such interactions may be necessary, it might be better to allow simpler models where only the evolution of the time dependent covariates is kept general. This is another advantage of a continuous time model since covariate effects could remain on the hazard (transition intensity) scale as in the Cox PH model.

Finally, the authors give a partial justification of the use of a time homogeneous Markov model through the Cox PH model having an approximately constant baseline hazard. It should be noted that a time homogeneous Markov model does not imply a constant absorption hazard (unless the model begins in the quasi-stationary distribution). Conversely, while a constant hazard might suggest homogeneity more than non-homogeneity, it is nevertheless possible to construct non-homogeneous processes with constant (or near constant) marginal absorption hazards. The authors do however report a statistic which gives a better justification of homogeneity.

Friday, 10 June 2011

Comparison of prediction models for competing risks with time-dependent covariates

Giuliana Cortese, Thomas Gerds and Per Kragh Andersen have a new paper available as a University of Copenhagen Department of Biostatistics technical report. The paper concerns the development of models for prediction for competing risks in the presence of internal time dependent covariates. This is a follow-up to Cortese and Andersen's 2010 Biometrical Journal paper. Like the previous paper, the authors compare a multi-state modelling approach that explicitly models the progression of the (categorical) time dependent covariate and its effect on the cause specific hazards, and a landmarking approach that sets (arbitrarily chosen) time points and performs separate regressions to estimate the hazards for conditional on the value of the time dependent covariate at time . The authors consider two modelling approaches under landmarking, one based on Cox regression of cause-specific hazards and the other based on Fine-Gray subdistribution hazards.

They compare the predictive ability of the models to predict the outcome by landmark given data up to . This is assessed by using a time dependent Brier score (Gerds & Schumacher, 2006). Rather than use inverse probability weighting, the authors instead use a pseudo-value to estimate outcomes when a subject is lost to follow-up between and . The authors perform the comparison using a bone marrow transplant study where the competing events are relapse and death, and the internal time dependent covariate is the development of Graft versus Host Disease (GvHD). The predictive abilities are estimated via cross-validation involving randomly choosing 2/3 of patients as training data and using the remainder as test data, repeating the process 100 times. For the data considered the three methods performed equally well in terms of prediction error. As might be expected, there was significantly improved predictive ability of these models compared to one that ignored GvHD (i.e. only considered baseline time constant covariates).

As the authors note, there are advantages and disadvantages to both approaches. The multi-state modelling approach requires modelling of the covariate process (e.g. Markov or semi-Markov assumptions and proportional hazard assumptions on the effect of baseline covariates on transition rates through covariate states) and requires a categorical covariate. Landmarking can accommodate continuous covariates but relies on an arbitrary set of landmark times and requires fitting regressions at each landmark. The extra modelling required for the multi-state approach may either be a blessing, in terms of having the potential to give a greater insight into the whole process, or a curse (questions of robustness to incorrect modelling assumptions).