Showing posts with label random effects. Show all posts
Showing posts with label random effects. Show all posts

Tuesday, 31 July 2012

A multistate modelling approach for pancreatic cancer development in genetically high risk families

Kolamunnage-Dona, Vitone, Greenhalf, Henderson and Williamson have a new paper in Applied Statistics. This uses a competing risks model with shared frailties to model data on the progression of pancreatic cancer in the presence of clustering and informative censoring. Clustering is present due to data being available on patients from the same family groups. A patient having a resection causes censoring of the main event, time to pancreatic cancer, but is likely to be informative. This is dealt with by allowing the cause-specific hazards to depend on the same frailty term. The methodology in the paper is very similar to that of Huang and Wolfe (Biometrics, 2002), the only extension being that the current formulation allows for the possibility of time dependent covariates. It isn't clear what if any complication this adds to the original procedure in Huang and Wolfe. An MCEM algorithm is used where the E-step is approximated by using Metropolis-Hastings in order to calculate the required expected quantities. It's not clear what the authors mean when they say the Metropolis-Hastings step use "a vague prior for the frailty variance." Hopefully they mean an improper uniform prior as otherwise they would be pointlessly adding Bayesian features to an otherwise frequentist estimating procedure. On a related computational point, since the random effect for each cluster is one dimensional, I suspect using a Laplace approximation to compute the required integrals at each step would perform quite well and be a lot faster than using Metropolis-Hastings

In the analysis there does seem evidence for a shared frailty within clusters, but it appears that the parameter which links the frailty in the time to pancreatic cancer to the time to resection intensity is hard to identify having a very wide 95% confidence interval encompassing strong negative dependence through independence to strong positive dependence. The typical cluster size in the data is quite small (e.g. median is 3) and this is probably insufficient as you would ideally need some subjects to fail and some to be informatively censored in each cluster to assess their association. The point estimate is negative implying a counter intuitive negative association between resection and pancreatic cancer. The authors suggest as a model extension to allow a specific bivariate frailty linking competing risks (presumably within individuals?) - which is unlikely to be helpful.

Sunday, 29 July 2012

Mixture distributions in multi-state modelling: Some considerations in a study of psoriatic arthritis

Aidan O'Keeffe, Brian Tom and Vern Farewell have a new paper in Statistics in Medicine. This considers random effects models for clustered multi-state models, specifically considering the psoriatic arthritis example considered in their previous paper. The particular emphasis in the current paper is comparing models using a continuous gamma frailty term with an extended model that additionally allows a "stayer" component. The latter can be thought of as a joining of the continuous random effects model (e.g. Cook et al 2004) and the "mover-stayer" model (e.g. Cook et al 2002). In a discussion on random effects, particularly where there is a question of finite mass points, it is strange there is no mention of the non-parametric mixing distribution (Laird, JASA 1978). While this hasn't really been considered for continuous time processes (the nearest example is Frydman's model) it wouldn't be particularly hard to implement at least the (not entirely reliable) EM algorithm type approach used in discrete-time by Maruotti and Rocci. The apparent presence or absence of a "stayer" component is likely to be heavily dependent on the parametric assumptions made about the rest of the mixing distribution. A very small random effect is indistinguishable from a zero random effect. The authors do emphasize the need to consider several possible mover-stayer models and consider the biological plausibility of them. It is also worth mentioning that all these issues will also hinge on the appropriateness of other assumes e.g. conditionally time homogeneous Markov processes.

Wednesday, 20 June 2012

Investigating hospital heterogeneity with a multi-state frailty model

Benoit Liquet, Jean Francois Timsit and Virginie Rondeau have a new paper in BMC Medical Research Methodology. They consider data on the evolution of patient's status for patients in intensive care units. The data originate from 16 distinct ICUs and there is therefore inherent clustering in the data. A progressive four-state model, with admission as state 0, ventilator-associated pneumonia infection (VAP) as state 1, death as state 2 and discharge as state 3 is considered. To account for the clustering, shared Gamma frailty terms are included in the intensities for particular transitions. To maintain simplicity of the method the authors consider cases where either each transition intensity has a ICU related frailty that is assumed independent of frailties for other intensities, or where there are individual level frailties that act in common across intensities. In each, there is only one level of clustering (i.e. either independent frailties on each intensity common to each centre or a subject specific intensity affecting multiple intensities). In the first case, a non-homogeneous Markov or semi-Markov allows separate models to be fitted for each transition using methods applicable to univariate survival analysis. In the second case, multiple intensities can be fitted in a single survival model by specifying separate "strata" for each intensity. In either case the presence of only one level of clustering makes estimation, at least if a Gamma frailty is assumed, relatively straightforward. The authors use the R package frailtypack which Rondeau maintains. This allows fully parametric models (via either Weibull or piecewise constant intensities) or semi-parametric models using spline intensities fitted using penalized likelihood. The emphasis on simple(r) approaches, for which existing software exist, is a good one. But some mention of the existing capacity of the more general survival package to fit Cox models with gamma frailties by using frailty() in the formula.

The wider issue of what to do in the case of nested frailties, i.e. where there are multiple layers of clustering, is an interesting one, but is not discussed which is surprising given that frailtypack has some capability for fitting such models. It is also questionable whether the assumption of independent frailties across intensities is a realistic one. To some extent this doesn't matter because if the data are Markov or semi-Markov conditional on the frailties, then even if the frailties are dependent assuming independence should produce unbiased estimates of the marginal frailty distribution (provided it is Gamma). Similarly estimates of the individual frailties should be reasonably unbiased (empirical Bayes estimates aren't fully unbiased anyway) but will be inefficient if frailty terms across intensities are actually correlated. It would be relatively straightforward to look at the correlation in the empirical Bayes estimates of the frailties associated with the intensities under the independence model as a semi-informal diagnostic. The difficult part would be fitting the multivariate frailty model if the independence model appeared inadequate.

Tuesday, 27 March 2012

Modeling Left-truncated and right-censored survival data with longitudinal covariates

Yu-Ru Su and Jane-Ling Wang have a new paper in the Annals of Statistics. This considers the problem of modelling survival data in the presence of intermittently observed time varying covariates when the survival times are both left truncated as well as right-censored. They consider a joint model which involves assuming there exists a random effect which influences both the longitudinal covariate values (which are assumed to be a function of the random effects plus Gaussian error) and the survival hazard. Considerably work has been done in this area in cases where the survival times are merely right-censored (e.g. Song, Davidian and Tsiatis, Biometrics 2002). The authors show that the addition of left-truncation complicates inference quite considerably; firstly because the parameters affecting the longitudinal component may not be identifiable and secondly because the score equations for the regression and baseline hazard parameters become much more complicated than in the right-censoring case. To alleviate this problem, the authors propose to use a modified likelihood rather than either the full or conditional likelihood. The full likelihood can be expressed in terms of an integral over the conditional distribution of the random effect, given the event time occurred after the truncation time. The proposed modification is to instead integrate over the unconditional random effect distribution. Heuristically this is justified by noting that
and
where is the random effect. The authors also show inference based on this modified likelihood gives consistent and asymptotically efficient estimators of the regression parameters and the baseline survival hazard.

An EM algorithm to obtain the MMLE is outlined, in which the E-step involves a multi-dimensional integral which the authors evaluate through Monte Carlo approximation. The implementation of the EM algorithm is simplified if the random effect is assumed to have a multivariate Normal distribution.

Friday, 24 February 2012

Estimating survival of dental fillings on the basis of interval-censored data and multi-state models

Pierre Joly, Thomas Gerds, Vibeke Qvist, Daniel Commenges and Niels Keiding have a new paper in Statistics in Medicine. This considers the estimation of survival times of dental fillings from interval censored data. A particular feature of the data is that there is inherent clustering in the form of multiple fillings from the same child.

Bizarrely the authors claim "we are not aware of any paper combining multi-state models for clustered data with interval censoring" implying they (and presumably the referees as well) are unaware of both Cook, Yi, Lee and Gladman (Biometrics, 2004) and
Sutradhar and Cook (JRSS C, 2008), the latter having "clustered", "multistate" and "interval-censored" all in the title!

A progressive four-state model is assumed for each filling, with the states consisting of Treatment, Filling failure, endodontic complication and exfoliation (an absorbing state).
For exfoliation, age of child is taken as the time scale meaning that the time of treatment is taken as a left-truncation time. For transitions from treatment to the other two states, time since treatment is taken as the time scale. Weibull transition intensities are assumed. However, monitoring ended once any filling event (filling failure or endodontic complication) had occurred. Because the exact time of an endodontic complication is known if it occurred during the monitoring period, the transition intensity from endodontic complication to exfoliation is not relevant to the likelihood. The authors also argue that it is necessary to assume that the intensity from filling failure to exfoliation and from treatment to exfoliation is the same, due to never being able to observe a filling failure to exfoliation transition. Strictly speaking, under the assumption of non-informative observation times, there should be some information in the data to estimate something about the separate intensities based on the proportion of cases where a filling was observed compared to the proportion where exfoliation occurred without observing a filling. Indeed in the research report by Frydman et al (2008), using a subset of the data in the current paper and a three-state verison of the model, a discrete-time NPMLE for the intensities was developed.

Random effects are incorporated into the model in a hierarchical way, with a dentist level random effect that affects the intensity to filling failure or endodontic complication and correlated child level random effects determining the correlation to time to failures (thus affecting the transitions to filling failure and endodontic complication) and time to exfoliation (thus affecting only the exfoliation transition intensity). The random effects are taken to be multivariate Normal with a log-additive effect on the intensities. Calculating the likelihood requires numerical integration, which here is achieved via Gauss-Hermite quadrature. 30 quadrature points were used - this seems a rather small number for a multi-dimensional integral. Cook et al (2004) avoided attempting to get a strict approximation to the multivariate Normal by formally restricting the random effects to have a discrete distribution. Sutradhar and Cook (2008) used an MCEM in order to apply a continuous random effects distribution. That approach is likely to be more computationally intensive that Gauss-Hermite quadrature (on 30pts) but more accurate. The recent suggestion by Putter and van Houwelingen to use a simple two-component mixture frailty could also be adaptable for this situation.

Two filling types, amalgam and glass ionomer are compared in terms of probability of surviving without complication, with amalgam performing somewhat better.

Monday, 6 February 2012

A mixed non-homogeneous hidden Markov model for categorical data, with application to alcohol consumption

Antonello Maruotti and Roberto Rocci have a new paper in Statistics in Medicine. This develops a hidden Markov model for modelling longitudinal data on alcohol consumption in discrete time. The observed data are taken to consist of a three-level ordinal variable denoting whether no drinking (0 drinks), light drinking (1–c drinks), and intense drinking (c+ drinks) occurred in that period of time. The model considered is both time non-homogeneous and mixed, in the sense that there is additional patient level heterogeneity after accounting for covariates. Rather than specifying a continuous distribution for the random effects, the authors adopt the non-parametric mixing distribution approach. Computationally, a finite mixture random effect is much simpler than having a continuous random effect if the random effect is multi-dimensional. However, computation of the full non-parametric maximum likelihood estimate of the mixing distribution is not in itself straightforward. The authors adopt the approach of Aitkin (Statistics and Computing, 1996) which is essentially to work up from a small number of components, performing an EM-algorithm at the fixed level of mixture components. EM based approaches to obtaining the NPMLE of a mixing distribution are known to perform badly and approaches using directional derivatives are preferred (see for instance Wang 2007, JRSS B). The best model, with m components, is assumed to have been reached once taking m+1 components does not produce a better model in terms of AIC or BIC. The main issue with this approach is that the EM algorithm is typically very sensitive to the initial parameter values chosen and prone to fail to find a global maximum. A further danger with these models is to ascribe too great a physical significance to the mixture components estimated.

To reach the final model, choices have to be made regarding: the categorization for the responses (observed level of drinking), the latent Markov states for the HMM (e.g. 2, 3 or 4 latent states), the number of mixture components for the random effect (how many "archetypes" of longitudinal behavior) and the degree of time non-homogeneity of transition probabilities. As a result, while the model is likely to explain the observed data reasonably well, a leap of faith is required to believe the model is an accurate representation of the process of binge drinking/alcoholism.

Monday, 28 November 2011

Frailties in multi-state models: Are they identifiable? Do we need them?

Hein Putter and Hans van Houwelingen have a new paper in Statistical Methods in Medical Research. This reviews the use of frailties within multi-state models, concentrating on the case of models observed up to right-censoring and primarily looking at models where the semi-Markov (Markov renewal) assumption is made. The paper clarifies some potential confusion over when frailty models are identifiable. For instance, for simple survival data with no covariates, if a non-parametric form is taken for the hazard then no frailty distribution is identifiable. A similar situation is true for competing risks data. Things begin to improve once we can observe multiple events for each patient (e.g. in illness-death models), although clearly if the baseline hazard is non-parametric, some parametric assumptions will be required for the frailty distribution. When covariates are present in the model, the frailty term has a dual role of modelling dependence between transition times but also soaking up lack of fit in the covariate (e.g. proportional hazards) model.

The authors consider two main approaches to fitting a frailty; assuming a shared Gamma frailty which acts as a multiplier on all the intensities of the multi-state model and a two-point mixture frailty, where there are two mixture components with a corresponding set of multiplier terms for the intensities for each component. Both approaches admit a closed form expression for the marginal transition intensities and so are reasonably straightforward to fit, but the mixture approach has the advantage of permitting a greater range of dependencies, e.g. it can allow negative as well as positive correlations between sojourn times.

In the data example the authors consider the predictive power (via K-fold cross-validation) of a series of models on a forward (progressive) multi-state model. In particular they compare models which allow sojourn times in later states to depend on the time to first event, with frailty models and find the frailty model performs a little better. Essentially, these two models do the same thing and as the authors note, a more flexible model for the effect of time of first event on the second sojourn time may well give a better fit than the frailty model.

Use of frailties in models with truly recurrent events seems relatively uncontroversial. However for models where few intermediate events are possible their use, rather than some non-homogeneous model allowing dependence both on time since entry to the state and time since initiation, the choice is largely related to either what is more interpretable in terms of the application at hand or possibly what is more computationally convenient.

Monday, 14 November 2011

Bayesian inference for an illness-death model for stroke with cognition as a latent time-dependent risk factor

Van den Hout, Fox and Klein Entink have a new paper in Statistical Methods in Medical Research. This looks at joint modelling of cognition data measured through the Mini-Mental State Examination (MMSE) and stroke/survival data modelled as a three-state illness-death model. The MMSE produces item response type data and the longitudinal aspects are modelled through a latent growth model (with random slope and intercept). The three-state multi-state model is Markov conditional on the latent cognitive function. Time non-homogeneity is accounted for in a somewhat crude manner by pretending that age and cognitive function vary with time in a step-wise fashion.

The IRT model for MMSE allows the full information from the item response data to be used (e.g. rather than summing responses to individual questions to get a score). Similarly, compared to treating MMSE as an external time dependent covariate, the current model allows prediction of the joint trajectory of stroke and cognition. However, the usefulness of these predictions are constrained by how realistic the model actually is. The authors make the bold claim that the fact that a stroke will cause a decrease in cognition (which is not explicitly accounted for in their model) does not invalidate their model. It is difficult to see how this can be the case. The model constrains the decline in cognitive function to be linear (with patient specific random slope and intercept). If it actually falls through a one-off jump, then the model will still try to fit a straight line to a patient's cognition scores. Thus the cognition before the drop will tend to be down-weighted. It is therefore quite feasible that the result that lower cognition causes strokes is mostly spurious. One possible way of accommodating the likely effect of a stroke on cognition would be to allow stroke status to be a covariate in the linear growth model e.g. for multi-state process and latent trait , we would take


where the may be correlated random effects.

Thursday, 6 October 2011

A case-study in the clinical epidemiology of psoriatic arthritis: multistate models and causal arguments

Aidan O'Keeffe, Brian Tom and Vern Farewell have a new paper in Applied Statistics. They model data on the status of 28 individual hand joints in psoriatic arthitis patients. The damage to joints for a particular patient is expected to be correlated. Particular questions of interest are whether the damage process has 'symmetry', meaning if a specific joint in one hand is damaged does this increase the hazard of the corresponding joint in the other hand becoming damaged, and whether activity, meaning whether a joint is inflamed, causes the joint damage.

Two approaches to analysing the data are taken. In the first each pair of joints (from the left and right hands) is modelled by a 4 state model where the initial state is no damage to either joint, the terminal state is damage to both joints and progession to the terminal state may be through either a state representing damage only to the left joint or through a state representing damage only to the right joint. The correlation between joints within a patient is incorporated by allowing a common Gamma frailty which affects all transition intensities for all 14 pairs of joints. Symmetry can then be assessed by comparing the hazard of damage to the one joint with or without damage having occurred to the other joint. Under this approach, activity is incorporated only as a time dependent covariate. There are two limitations to this. Firstly, panel observation means that some assumption about the status of activity between clinic visits has to be made (e.g. it keeps the status observed at the last clinic visit until the current visit) and, from a causal inference perspective, activity is an internal time dependent covariate so treating it as fixed makes it harder to infer causality.

The second approach seeks to address these problems by jointly modelling activity and joint damage as linked stochastic processes. In this model each of the 28 joints is represented by a three-state model where the first two states represent presence and absence of activity for an undamaged joint, whilst state three represents joint damage. For this model, rather than explicitly fitting a random effects model, a GEE type approach is taken where the parameter point estimates are based on maximizing the likelihood assuming independence of the 28 joint processes, but standard errors are calculated using a sandwich estimator based on patient level clustering. This approach is valid provided the point estimates are consistent, which will be the case if the marginal processes are Markov. For instance if the transition intensities of type {ij} are linked via a random effect u such that
but crucially if other transition intensities also have associated random effects, these must be independent of .

The basic model considered is (conditionally) Markov. The authors attempt to relax this assumption by allowing the transition intensities to joint damage from the non-active state to additionally depend on a time dependent covariate indicating whether activity has ever been observed. Ideally the intensity should depend on whether the patient has ever been active rather than whether they have been observed to be active. This could be incorporated very straightforwardly by modelling the process as a 4 state model where state 1 represents never active, no damage, state 2 represents active, no damage, state 3 represents not currently active (but have been in the past) and state 4 represents damage. Clearly, we cannot always determine whether the current occupied state is 1 or 3 so the observed data would come from a simple hidden Markov model.

On a practical level, the authors fit the likelihood for the gamma random effects model by using the integrate function in R, for each patient's likelihood contribution for each likelihood evaluation in the (derivative-free) optimization. Presumably this is very slow. Using a simpler more explicit quadrature formulation would likely improve speed (at a very modest cost to accuracy) because the contributions for all patients for each quadrature point could be calculated at the same time and the overall likelihood contributions could then be calculated in the same operation. Alternatively, the likelihood for a single shared Gamma frailty is in fact available in closed form. This follows from arguments extending the results in the tracking model of Satten (Biometrics, 1999). Essentially, we can write every term in the conditional likelihood as some weighted sum of exponentials:

the product of these terms is then still in this form:

Calculating the marginal likelihood then just involves computing a series of Laplace transforms.
Provided the models are progressive, keeping track of the separate terms is feasible, although the presence of time dependent covariates makes this approach less attractive.

Tuesday, 4 October 2011

A Bayesian Simulation Approach to Inference on a Multi-State Latent Facotr Intensity Model

Chew Lian Chua, G.C. Lim and Penelope Smith have a new paper in the Australian and New Zealand Journal of Statistics. This models the trajectory of credit ratings of companies. While there are 26 possible rating classes, these are grouped to the 4 broad classes: A,B,C,D, Movement between any two classes is deemed possible resulting in 12 possible transition types. The authors develop methods for fitting what is termed a multi-state latent factor intensity model. This is essentially a multi-state model with a shared random effect that affects all the transition intensities, but the random effect is allowed to evolve in time. For instance, the authors assume the random effect is an AR(1) process.

There seems to be some confusion in the formulation of the model as to whether the process is in continuous or discrete time: The process is expressed in terms of intensities which are "informally" defined as
where is a counting process describing the number of s to k transitions that have occurred to time t and is the ith observation time. But if the time between observations were variable, the AR(1) formulation (which takes no account of this) would not seem sensible.

Nevertheless the idea of having a random effect that can vary temporally within a multi-state model (proposed by Koopman et al) is interesting, though obviously presents various computational challenges which is the main focus of Chua, Lim and Smith's paper.

Monday, 8 August 2011

Joint modelling of longitudinal outcome and interval-censored competing risk dropout in a schizophrenia clinical trial

Ralitza Gueorguieva, Robert Rosenheck and Haiqun Lin have a new paper in JRSS A. The paper concerns the joint modelling of a longitudinal outcome and an interval censored competing risks outcome that explains drop-out. As is common with these joint longitudinal and survival types of models the two processes are linked via a normally distributed vector of random effects. The novelty of the paper is in the survival part is a competing risks process and the event time is interval censored. The authors adopt a parametric model for the competing risks, using the family of distributions proposed by Sparling et al (Biostatistics, 2006). This makes inference somewhat more straightforward than it would be if a non-parametric baseline cause-specific hazards were used. As recently noted, parametric treatment of competing risks data is surprisingly rare. One problem faced by the authors is that the hazard family of Sparling, while allowing closed form expressions for interval censored univariate survival data, do not result in closed form expressions for interval censored competing risks data (except in special cases). Instead a numerical integral has to be competed. The presence of the overall random effects would mean the likelihood requires nested integration. To avoid this problem the authors adopt an approximation to the true likelihood for competing risks data. If a patient is known to have had a failure of type j in the interval [t0,t1] the authors assume that the patient is censored of all risks except risk j at time t0. It is clear that this approximation will lead to systematic bias as the time at risk from each failure type will be underestimated so the hazards will tend to be overestimated. The amount of bias will depend on the typical length of the intervals [t0,t1].

For the CATIE data example the proposed approximation is probably not an issue. The drop out (competing risks) part of the model is not the primary focus of the inference, and it is really the relative hazards of different types of drop out rather than their absolute values that is important in determining the trajectories of the longitudinal measure without drop out. For instance the estimates for simulated data of a similar type are close to unbiased.
However in extreme cases like current status competing risks data the approximation will do extremely badly.

Monday, 21 March 2011

Joint model with latent state for longitudinal and multistate data

Dantan, Joly, Dartigues and Jacqmin-Gadda have a new paper in Biostatistics. There is a wide literature on joint models of survival with longitudinal data with some extensions for joint competing-risks survival and longitudinal data (e.g. Elashoff et al). Dantan et al develop a joint model for multi-state survival data and longitudinal data. The standard approach to these joint models is to have a random effect that it shared across the model for the survival data (e.g. appearing as a frailty) and the longitudinal model (e.g. random slopes and intercepts in a generalized linear mixed model). Rather than follow this convention, Dantan et al allow a slightly more direct correspondence between the survival and longitudinal parts. Specifically, they have a progressive 4-state multi-state model relating to healthy, pre-diagnosis, illness and death states. The pre-diagnosis state is unobservable and entry into this state corresponds to the time a which the slope of decline in the longitudinal biomarker changes. The model is an extension on the random change-point model proposed by Jacqmin-Gadda et al (Biometrics, 2006), as here it is not necessary to make unrealistic assumptions about death being non-informative censoring for illness.

For the PAQUID dataset on cognitive decline, the baseline transition intensities are of Weibull form for healthy to pre-diagnosis and pre-diagnosis to illness (the latter being Weibull w.r.t time since pre-diagnosis). The hazard of death is assumed to depend only on age (via piecewise constant intensities) and the value of the longitudinal marker, Y(t), but not explicitly on the current disease state. This assumption seems to be primarily for computational reasons.

A limitation of the approach taken in the paper is that dementia is only diagnosed at clinic visits, so in effect is interval censored between the current and last clinic visit. The authors just assume entry into the dementia state occurred at the midpoint between clinic visits. Through simulation in the supplementary materials the authors show this doesn't cause serious bias.
However, a further issue is that the dependency of the dementia age on the observed value of the marker at a clinic visit would presumably mean the assumption of independence between dementia age and observation errors conditional on the random effects (slopes, intercepts and change-time) would be inappropriate. To what extent is the apparent increase in the decline of cognition before diagnosis due to such an artifact?

Wednesday, 23 February 2011

Estimating and testing for center effects in competing risks

Sandrine Katsahian and Christian Boudreau have a new paper in Statistics in Medicine. This develops methods for including frailty terms within a Fine-Gray competing risks model in order to account for clustering, e.g. effects of different centres.

Since the Fine-Gray model is essentially just a standard Cox proportional hazards regression model with additional time dependent weights, based on the censoring distribution, for individuals who have had a competing event, methods appropriate for standard Cox frailty models can be readily adapted.

Katsahian and Boudreau closely follow the approach taken by Ripatti and Palmgren (Biometrics, 2000). They assume a Gaussian frailty. Computation of the likelihood requires integrating out the frailty terms. Here this is performed using a Laplace approximation. A difficulty with the Laplace approximation is that it still requires the modal value of the frailty distribution conditional on the data and current values of the parameters. The authors therefore take a profile likelihood approach in which they fix the frailty variance and maximize the likelihood term with respect to both the regression parameters and the frailty terms, . Having obtained, and they can then plug into the Laplace approximation to get the profile likelihood for . The procedure gives a local approximation for which can be used to suggest the updated estimate. Thus the process involves alternating between two Newton-Raphson algorithms until convergence.

Sunday, 21 November 2010

The analysis of tumorigenicity data using a frailty effect

Yang-Jin Kim, Chung Mo Nam, Youn Nam Kim, Eun Hee Choi and Jinheum Kim have a new paper in the Journal of the Korean Statistical Society. This concerns the analysis of tumorigenicity data where a three-state model with states representing tumour-free, with tumour and death is used. The main characteristic of the data is that disease status (presence of tumour) can only be determined at either death or sacrifice - meaning we have data with possible observations: sacrifice without tumour, sacrifice with tumour, death without tumour and death with tumour.
Note that while the time of the tumour is always interval censored, if the mouse dies before sacrifice this death time is taken to be known exactly.

Kim et al extend the model proposed by Lindsey and Ryan (1993, Applied Statistics) by allowing a shared-frailty term that affects both tumour onset rate and the hazard of death given the presence of a tumour. Like Lindsey and Ryan, piecewise constant intensities are used to model the baseline intensities, with a Cox proportional hazards model for the effect of covariates. The same baseline hazard is used for pre- and post-tumour hazard of death, with the presence of tumour changing the effect of covariates. For some reason the introduction talks about n^{1/3} convergence rates of the non-parametric survival distribution for current status data. This doesn't seem relevant here given that the piecewise constant intensities model is parametric.

An EM algorithm is used to maximize the likelihood. Gauss-Hermite quadrature is required to perform the M-step due to the presence of the frailty terms. The authors appear to be basing standard error estimates on the complete data (rather than observed data) likelihood.

The new model is applied to the same data as used in Lindsey and Ryan. In addition, in a simulation it is shown that there is some reduction in bias compared to Lindsey and Ryan's method is the proposed model is correctly specified.

Thursday, 27 May 2010

Incorporating frailty in a multi-state model: application to disease natural history modelling of adenoma-carcinoma in the large bowel

Yen, Chen, Duffy and Chen have a new paper in SMMR. This considers methods for incorporating frailty into a progressive multi-state model under interval censoring (the data resemble current status data being a case-cohort design with known sampling probabilities). They note that the accommodation of frailty is made easier by considering a model where only one transition intensity is subject to heterogeneity. The main deficiency of the paper is the lack of reference to the fairly wide existing literature on frailty and random effects models for panel observed multi-state models. In particular, the tracking model of Satten is directly applicable to their data. In the tracking model, a common frailty affects all transition intensities and results in a likelihood that doesn't require numerical integration. Some comment on the extendability of their model to the case of multiple observations per patient would also have been useful.

Wednesday, 27 January 2010

A nonstationary Markov transition model for computing the relative risk of dementia before death

Yu et al have a new paper in Statistics in Medicine. This is concerned with estimating the risk of dementia before death. A 5 state multi-state model is used, with three transient states representing levels of cognitive impairment, plus two absorbing states dementia and death. Unlike other recent dementia studies using multi-state models, improvements, as well as deteriation, in cognitive ability are assumed possible. A discrete-time approach is taken. This has some advantages in terms of the flexibility in modelling possible in terms of incorporating non-homogeneity over time and a frailty term - the Markov assumption. A drawback of using a discrete-time approach in the current study was that the study did not have equally spaced observation times, but it was necessary to assume this was the case for the discrete-time model. This is likely to cause some bias.

The main theoretical development in the paper is expressions for the mean and variance of the time spent in each transient state before absorption in the non-homogeneous case.

Tuesday, 24 November 2009

Statistical Analysis of Illness-Death Processes and Semicompeting Risks Data

Xu, Kalbfleisch and Tai have a new paper in Biometrics. They make a compelling case against a latent failure times approach to the analysis of semi-competing risks data, advocating to instead take a classical cause-specific hazards approach. Moreover, they note that semi-competing risks is essentially just an illness-death model. They consider models where the intensities have a shared Gamma frailty and present methods for (Non-parametric) maximum likelihood estimation. In addition, covariates can be included acting proportionally on the conditional hazards (an alternative approach with proportionality on the marginal hazards is also outlined).

Monday, 3 August 2009

Estimating dementia-free life expectancy for Parkinsons patients using Bayesian inference and microsimulation

Van den Hout and Matthews have a new paper in Biostatistics involving a random-effects Markov model with time (age) dependent intensities. The methodology is close to that used by Pan et al and Wu et al, using a WinBUGS/OpenBUGS Bayesian approach. They use a three-state illness-death model without recovery. A more sophisticated multivariate log-normal random effect on the effects of age on the intensities is used with a Wishart prior is used on the covariance matrix, which is more appropriate than the Gamma(e,e) type priors used by Pan et al. Like their recent Applied Statistics paper, time dependencies in the intensities are accounted for by assuming an individual that is observed at times t1 and t2 has a constant matrix of intensities between those points, but different assumptions are used to calculate life expectancies. The main methodological development is obtaining life expectancy estimates through 'microsimulation.' This is deemed necessary because there are two levels of variation: variation in the posterior of the parameters and variation from the random effects distribution conditional on the parameters. 'Microsimulation' (or simulation) just approximates the integral over the random effects distribution.

Monday, 10 November 2008

Analysis of interval-censored data from clustered multistate processes: application to joint damage in psoriatic arthritis

Rinku Sutradhar and Richard J. Cook have a new paper in JRSS C. This builds on previous work on random effects multi-state models by Cook, Yi, Lee and Gladman (Biometrics, 2004). Here rather than using a discrete random effects distribution and assuming time homogeneity, they use an MCEM algorithm to allow multivariate normal random effects. In addition, piecewise constant transition intensities allow the assumption of time homogeneity to be relaxed.