Tuesday, 22 September 2009

Estimating stroke-free and total life expectancy in the presence of non-ignorable missing values

Van den Hout and Matthews have a new paper in JRSS A. This deals with interval censored data from a three-state disease model where subjects may miss scheduled interviews meaning the disease status is not observed. A joint model for the disease state and an observation indicator is developed, being a continuous time generalisation of Cole et al (2005) that allows information from exact death times to be included. Conditional on the disease state and measured covariates, the observation indicator is governed by a logistic model.
In general the method seems promising for dealing with informative observation when the potential observation times are known.

Though not noted, the model can be expressed as a hidden Markov model. The authors state that the logistic model and the three-state Markov model are estimated separately. It is not made clear how this is achieved since the logistic model depends on the unobserved states of the Markov model. In some cases the missing state will in fact be known, for instance if the sequence is 1,-,1 or 2,-,2. However, for sequences like 1,-,2 or 1,-,3 it is not possible to establish the unobserved state.

The main aim of the analysis is to obtain estimates of life expectancy, disease free life expectancy and post-disease life-expectancy. These are complicated functions of the parameter vector as they involve integrals of transition probabilities. In addition to the method of Aalen et al 1997, Van den Hout and Matthews additionally propose to use a Metropolis algorithm to get confidence intervals for the life expectancies. The resulting intervals have a Bayesian interpretation, being the credible intervals from an improper uniform prior, but will not be invariant to changes in parametrisation. In practice, the intervals may give good frequentist coverage, particularly for large samples. However, Van den Hout and Matthews seem to be implying the intervals have exact coverage (apart from Monte-Carlo error through the Metropolis algorithm) which is a substantial misconception. Moreover, no mention of the procedure being Bayesian is given.

Thursday, 20 August 2009

Joint Modeling of Self-Rated Health and Changes in Physical Functioning

Hubbard, Inoue and Diehr have a new paper in press in JASA. This applies the time-transformation model proposed by Hubbard et al (Biometrics, 2008). The data are panel observed with an assessment of disability at each observation plus a self-rated measure of health. The disability measure is assumed to be a 5 state time non-homogeneous Markov model, with 4 levels of disability and death as an absorbing state. Backward transitions between disability levels are permitted. Two parametric forms for the time transformation were considered: a power transformation implying monotonicity of all intensities with time, and a two parameter transformation implying the intensities are all unimodal. The non-parametric transformation proposed in Hubbard et al (2008) are not considered here.

Disability is jointly modelled with the self-rated measure of health which is dichotomised as healthy or unhealthy. This health outcome may depend on both the current and past values of disability and other covariates. There would be obvious problems of missing data if the past history of disability is included due to the panel observation. The authors only consider models where the health outcome depends on current (+ predicted future) levels of disability but not past levels. Linear logistic models are used to relate the health outcome to the observed levels of disability and other covariates. Rudimentary goodness-of-fit for the multi-state model is carried out using the prevalence-counts method of Gentleman et al (Stats in Med, 1994), while the logistic model is assessed using the Hosmer-Lemeshow test.

Tuesday, 4 August 2009

Model diagnostics for multi-state models

Titman and Sharples have a new review paper in SMMR. This considers methods for assessing fit in parametric, panel observed multi-state models. The primary focus is on the assessment of time homogeneous Markov models, although there is also a section on hidden Markov models that occur if states are considered to be observed with classification error. Methods for fitting more complicated models such as non-homogeneous and random effects models are also reviewed. A simple graphical generalization of the prevalence counts method of Gentleman et al (Stats in Med, 1994) is also developed.

Monday, 3 August 2009

Estimating dementia-free life expectancy for Parkinsons patients using Bayesian inference and microsimulation

Van den Hout and Matthews have a new paper in Biostatistics involving a random-effects Markov model with time (age) dependent intensities. The methodology is close to that used by Pan et al and Wu et al, using a WinBUGS/OpenBUGS Bayesian approach. They use a three-state illness-death model without recovery. A more sophisticated multivariate log-normal random effect on the effects of age on the intensities is used with a Wishart prior is used on the covariance matrix, which is more appropriate than the Gamma(e,e) type priors used by Pan et al. Like their recent Applied Statistics paper, time dependencies in the intensities are accounted for by assuming an individual that is observed at times t1 and t2 has a constant matrix of intensities between those points, but different assumptions are used to calculate life expectancies. The main methodological development is obtaining life expectancy estimates through 'microsimulation.' This is deemed necessary because there are two levels of variation: variation in the posterior of the parameters and variation from the random effects distribution conditional on the parameters. 'Microsimulation' (or simulation) just approximates the integral over the random effects distribution.

Tuesday, 28 July 2009

Nonparametric inference and uniqueness for periodically observed progressive disease models

Beth Griffin and Stephen Lagakos have a new paper in Lifetime Data Analysis. They consider panel observed progressive disease model (chain-of-events) data. The NPMLE estimator under a discrete-time semi-Markov assumption was developed by Sternberg and Satten (Biometrics, 1999). For datasets where individuals are observed at different times, some discretization of the data is required. An issue with the NPMLE is that it is not guaranteed to be unique and therefore reporting a single NPMLE may be misleading. The paper develops procedures for determining which components of the NPMLE are unique based on considering various re-parameterizations of the likelihood. The method is demonstrated on three example datasets including one on bronchiolitis obliterans syndrome in post-lung transplantation patients and one on primary HIV infection. In addition, the authors also provide a more intuitive algorithm for obtaining the NPMLE than the self-consistency algorithm of Sternberg and Satten.

Wednesday, 22 July 2009

On Induced Dependent Censoring for Quality Adjusted Lifetime (QAL) Data in Simple Illness-Death Model

A new paper by Pradhan and Dewanji in Statistics and Probability Letters considers the problem of induced dependent censoring in quality adjusted lifetime data. Quality adjusted survival time and quality adjusted censoring times are correlated even if the raw survival and censoring times are independent. Kaplan-Meier based estimates of QAL using the QA survival and censoring times will therefore be biased. The paper investigates the nature of the correlation and bias for the case of a simple three-state illness-death model. Under a semi-Markov assumption, they show that QA survival and censoring are positively correlated when the healthy state has greater utility than illness, but the correlation is negative if the relative utilities are reversed.

Tuesday, 23 June 2009

About Earthquake Forecasting by Markov Renewal Processes

There is a new paper in Methodology and Computing in Applied Probability by Garavaglia and Pavani. This concerns the forecasting of earthquakes. They analyse data on severe earthquakes in Turkey during the 20th century. Two approaches were considered. Firstly a two-state semi-Markov model is proposed to model the process. The states represent the occurrence of earthquakes of magnitude 5.5 - 6.3 and greater than 6.3 respectively. However, the probability that the next earthquake is of magnitude greater than 6.3 is seen to not depend on the magnitude of the last earthquake. Hence, instead a model based only on the inter-event times is considered using an exponential-Weibull mixture distribution.