Showing posts with label Biostatistics. Show all posts
Showing posts with label Biostatistics. Show all posts

Tuesday, 5 February 2013

The gradient function as an exploratory goodness-of-fit assessment of the random-effects distribution in mixed models


Geert Verbeke and Geert Molenberghs have a new paper in Biostatistics. The paper proposes the use of the gradient function (or equivalently the directional derivatives) of the marginal likelihood with respect to the random effects distribution, as a way of assessing goodness-of-fit in a mixed model. They concentrate on cases related to standard longitudinal data analysis using linear (or generalized linear) mixed models, however the method can be extended to other mixed models, such as clustered multi-state models with multivariate (log)-normal random effects.

If we consider data from units i with observations x_i, given a mixing distribution G, we can say the marginal density is given by .

The gradient function is then taken as where N is the total number of independent clusters

The use of the gradient function stems from finite mixture models and in particular the problem of finding the non-parametric maximum likelihood estimate of the mixing distribution. At the NPMLE the gradient function has a supremum of 1. If instead we assume there is a parametric mixing distribution, under correct specification the gradient function should be close to 1 across all values of u. Verbeke and Molenberghs use this property to construct an informal graphical diagnostic of the appropriateness of the proposed random effects distribution.

An advantage of the approach is that essentially no additional calculations are required to compute the measure, above and beyond those already needed for estimation of the parametric mixture model itself. A current limitation of the approach is that there is no formal test to assess whether the observed deviation is statistically significant. It is stated that this is ongoing work. It seems reasonably straightforward to show that the gradient function will tend to a Gaussian process with mean 1 but with a quite complicated covariance structure. Obtaining some nice asymptotics for a statistic based either on the maximum deviation from 1 or some weighted integral of the distance from 1 therefore seems unlikely. However, it may be possible to obtain a simulation based p-value by simulating from the limiting Gaussian process.

Thursday, 28 April 2011

Bayesian evidence synthesis for a transmission dynamic model for HIV among men who have sex with men

Presanis, De Angelis, Goubar, Gill and Ades have a new paper in Biostatistics. The paper attempts to use several disparate sources of data on prevalence and incidence to estimate transition rates and prevalences in a multi-state model of transmission of HIV among men who have sex with men (ironically abbreviated as MSM) via Bayesian evidence synthesis. The multi-state model has 4 transient states relating to "Eligible (non MSM)", "Susceptible (MSM)", "Undiagnosed" and "Diagnosed". Subjects enter aged 15 (or later due to migration) and may exit due to death or reaching age 45.

It should be noted that the multi-state model is deterministic and not stochastic, i.e. the proportions in each state at a given time conditional on the parameters and starting conditions are the solution of an ordinary differential equation. While this is unproblematic in terms of getting estimates of the expected proportions in each state at a given time, there must be some level of under representation of the uncertainty. In particular the likelihood assumes that the number of deaths from HIV from the SOPHID dataset at time is where is the number in the diagnosed state at time and is the probability of death in one year from HIV for the year . In reality D(t) should be stochastic rather than deterministic and as such we would expect the number of deaths to have a higher variance than assumed by the binomial model.

In principle the covariance matrix of the numbers in each state for each year could be derived, i.e. given the Markov model, we have N independent individuals each of whom's state occupancy at time is multinomial given the occupancy at time . A more realistic model would then assume that conditional on the parameters, the state occupation counts have a multivariate normal distribution (e.g. analogous to the approaches taken in the estimation of aggregate Markov data) and, for instance, the number of deaths from HIV is binomial with a denominator that is itself normally distributed. Quite possibly this extra source of uncertainty is negligible compared to the vast existing uncertainties but it nevertheless ought to be explored.

Monday, 21 March 2011

Joint model with latent state for longitudinal and multistate data

Dantan, Joly, Dartigues and Jacqmin-Gadda have a new paper in Biostatistics. There is a wide literature on joint models of survival with longitudinal data with some extensions for joint competing-risks survival and longitudinal data (e.g. Elashoff et al). Dantan et al develop a joint model for multi-state survival data and longitudinal data. The standard approach to these joint models is to have a random effect that it shared across the model for the survival data (e.g. appearing as a frailty) and the longitudinal model (e.g. random slopes and intercepts in a generalized linear mixed model). Rather than follow this convention, Dantan et al allow a slightly more direct correspondence between the survival and longitudinal parts. Specifically, they have a progressive 4-state multi-state model relating to healthy, pre-diagnosis, illness and death states. The pre-diagnosis state is unobservable and entry into this state corresponds to the time a which the slope of decline in the longitudinal biomarker changes. The model is an extension on the random change-point model proposed by Jacqmin-Gadda et al (Biometrics, 2006), as here it is not necessary to make unrealistic assumptions about death being non-informative censoring for illness.

For the PAQUID dataset on cognitive decline, the baseline transition intensities are of Weibull form for healthy to pre-diagnosis and pre-diagnosis to illness (the latter being Weibull w.r.t time since pre-diagnosis). The hazard of death is assumed to depend only on age (via piecewise constant intensities) and the value of the longitudinal marker, Y(t), but not explicitly on the current disease state. This assumption seems to be primarily for computational reasons.

A limitation of the approach taken in the paper is that dementia is only diagnosed at clinic visits, so in effect is interval censored between the current and last clinic visit. The authors just assume entry into the dementia state occurred at the midpoint between clinic visits. Through simulation in the supplementary materials the authors show this doesn't cause serious bias.
However, a further issue is that the dependency of the dementia age on the observed value of the marker at a clinic visit would presumably mean the assumption of independence between dementia age and observation errors conditional on the random effects (slopes, intercepts and change-time) would be inappropriate. To what extent is the apparent increase in the decline of cognition before diagnosis due to such an artifact?

Tuesday, 29 June 2010

A multistate model for events defined by prolonged observation

Vern Farewell and Li Su have a new paper in Biostatistics. This models remission in psoriatic arthritis, based on panel data assessing joints at clinic visits. Existing models use a two-state model. However, a spell in remission should last a discernible amount of time. Rather than specify an artificial length of time (e.g. 6 months) e.g. a guarantee time in a state, Farewell and Su adopt a model with two states referring to remission: they refer to these as "early stage remission" and "established remission". It is assumed an individual must progress through both early and established remission before returning to active disease. A subject is observed to be in early stage remission if they have no active joints at a visit not preceded by at least 2 other zero count visits. They are in established remission if there is a zero count and at least 2 previous zero counts. State misclassification is allowed in the model through misclassification of the active disease count, i.e. patients may have 0 active joints without being in remission. It is assumed that misclassification to the early stage remission is possible but not to the established remission stage. In the example the misclassification probability is also allowed to depend on whether the previous observed count was zero or not.

The basic problem with the method is that having the states defined by the pattern of previous observations means that it is essentially impossible for the observed data (in terms of the three-state model) to come from the claimed Markov model: In the observed data, an established remission stage must be preceded by two early stage remission observations. Yet the actual Markov model allows the passage time from active disease to established remission to be arbitrarily close to zero (e.g. just the sum of two independent exponential - or perhaps piecewise exponential - distributions). As a result it is not clear how to interpret the resulting transition intensity estimates since the estimated process will not reproduce the original data.

Misclassification is effectively dealt with twice in the model. Firstly in an ad hoc way through rules on what early and established remission are. Then by allowing these observed states to have classification error over some true states. But in fitting a hidden Markov model the misclassification is assumed independent conditional on the underlying state. There is then an inherent contradiction because on the one hand the model says P(Observed zero | Active disease) >0, but at the same time P(Observed zero and two previous zero | Active disease) =0.

An approach using a guarantee time (e.g. Kang and Lagakos (2007)) or perhaps an Erlang distribution through latent states would be far more satisfactory even if it might require "special software". Potentially the guarantee time could be dependent on covariates.

Tuesday, 12 January 2010

Estimating disease progression using panel data

Micha Mandel has a new paper in Biostatistics. This considers estimation for panel observed data relating to MS progression. The underlying process can have backward transitions, so reaching a higher state is not in itself considered as evidence of progression. Instead, the patient is required to have stayed in the higher state for some period of time (e.g. 6 months). The quantities of interest in the study is therefore the time taken to first have stayed in state 3 for 6 months. For a continuous time, time-homogeneous Markov process expressions for the mean time and the distribution function are obtained.

As noted by Mandel, the estimates obtained are strongly dependent on the Markov assumption and time homogeneity. This is demonstrated in a simulation study. Mandel notes that more methods for semi-Markov models are required, the recent paper on phase-type semi-Markov models may be of use. Indeed, in the simulations Mandel actually uses phase-type distributions to create a semi-Markov process. However, the MS dataset may be too small to reliably estimate a semi-Markov model.

Informative observation times are a potential problem in the MS study. Mandel shows that observations that occur away from the scheduled 26-week gap time, are more likely to involve a transition, suggesting these observation times are informative. As a result, only observations within a 4 week period of the scheduled visit time are included in the analysis.

A comparison with a discrete-time Markov model approach showed that estimates based on assuming a discrete-time process gave larger estimates of the hitting times.

Monday, 3 August 2009

Estimating dementia-free life expectancy for Parkinsons patients using Bayesian inference and microsimulation

Van den Hout and Matthews have a new paper in Biostatistics involving a random-effects Markov model with time (age) dependent intensities. The methodology is close to that used by Pan et al and Wu et al, using a WinBUGS/OpenBUGS Bayesian approach. They use a three-state illness-death model without recovery. A more sophisticated multivariate log-normal random effect on the effects of age on the intensities is used with a Wishart prior is used on the covariance matrix, which is more appropriate than the Gamma(e,e) type priors used by Pan et al. Like their recent Applied Statistics paper, time dependencies in the intensities are accounted for by assuming an individual that is observed at times t1 and t2 has a constant matrix of intensities between those points, but different assumptions are used to calculate life expectancies. The main methodological development is obtaining life expectancy estimates through 'microsimulation.' This is deemed necessary because there are two levels of variation: variation in the posterior of the parameters and variation from the random effects distribution conditional on the parameters. 'Microsimulation' (or simulation) just approximates the integral over the random effects distribution.