Showing posts with label hidden Markov model. Show all posts
Showing posts with label hidden Markov model. Show all posts
Friday, 26 October 2012
Estimating parametric semi-Markov models from panel data using phase-type approximations
Andrew Titman has a new paper in Statistics and Computing. This extends previous work on fitting semi-Markov models to panel data using phase-type distributions. Here, rather than assume a model in which each state is assumed to have a Coxian phase-type sojourn distribution, the model assumes a standard parametric sojourn distribution (e.g. Weibull or Gamma). The computational tractability of phase-type distributions is exploited by approximating the parametric semi-Markov model by a model in which each state has a 5-phase phase-type distribution. In order to achieve this, a family of approximations to the parametric distribution, with scale parameter 1, is developed by solving a relatively large one-off optimization assuming the optimal phase-type parameters for given shape parameters evolve as B-spline functions. The resulting phase-type approximations can then be scaled to give an approximation for any combination of scale and shape parameter and then embedded into the overall semi-Markov process. The resulting approximate likelihood appears to be very close to the exact likelihood, both in terms of shape and magnitude.
In general a 2-phase Coxian phase-type model is likely to give similar results to a Weibull or Gamma model. The only advantages of the Weibull or Gamma model are that the parameters are identifiable under the null (Markov) model so standard likelihood ratio tests can be used to compare models (unlike for the 2-phase model). Also the Weibull or Gamma model requires fewer parameters so may be useful for smaller sample sizes.
Friday, 7 September 2012
Effect of vitamin A deficiency on respiratory infection: Causal inference for a discretely observed continuous time non-stationary Markov process
Mingyuan Zhang and Dylan Small have a paper to appear in The Canadian Journal of Statistics, currently available here. The paper uses a multi-state model approach to obtain estimates of the causal effects of vitamin D deficiency on respiratory infection.
The observed data consist of observations of respiratory infection status, vitamin D deficiency status and whether the child is stunted at time t. Each of these is a binary variable, leading to 8 possible observation patterns.
The data are assumed to be generated from an underlying non-homogeneous Markov chain on 32 states, consisting of a latent 4-level definition of vitamin D deficiency, the stunting variable, the observed respiratory infection status and additionally a counter-factual respiratory infection status defined as the status at time t hat would have occurred had the child maintained the lowest level of vitamin D deficiency from time 0 to time t.
Let
The overall model is a pretty innovative use of a hidden Markov model structure to obtain those causal estimates. The true process is assumed to occur in continuous time. However, it is desired that the underlying transition intensities are not time constant. As a result, the authors choose to approximate the process by one in discrete time (with some similarities to the approach of Bacchetti et al 2010).
In practical terms, the weakness of the model seems to be the assumption that the relative effect of a current vitamin deficiency compared to a perfect record of vitamin D levels, is both constant in time and does not depend on the past history of vitamin D deficiency. The latter assumption is essentially the Markov assumption and would be quite difficult to relax. The observed infection status is effectively a misclassified version of the counter-factual infection status. As a result the former assumption could be relaxed by letting the
The definition of the process as initially being continuous is a little artificial and ill-specified in places. For instance, the transition intensities are defined in terms of a logit transformation from the outset. Also, it is stated that the underlying 32 state process (including both of
Monday, 6 February 2012
A mixed non-homogeneous hidden Markov model for categorical data, with application to alcohol consumption
Antonello Maruotti and Roberto Rocci have a new paper in Statistics in Medicine. This develops a hidden Markov model for modelling longitudinal data on alcohol consumption in discrete time. The observed data are taken to consist of a three-level ordinal variable denoting whether no drinking (0 drinks), light drinking (1–c drinks), and intense drinking (c+ drinks) occurred in that period of time. The model considered is both time non-homogeneous and mixed, in the sense that there is additional patient level heterogeneity after accounting for covariates. Rather than specifying a continuous distribution for the random effects, the authors adopt the non-parametric mixing distribution approach. Computationally, a finite mixture random effect is much simpler than having a continuous random effect if the random effect is multi-dimensional. However, computation of the full non-parametric maximum likelihood estimate of the mixing distribution is not in itself straightforward. The authors adopt the approach of Aitkin (Statistics and Computing, 1996) which is essentially to work up from a small number of components, performing an EM-algorithm at the fixed level of mixture components. EM based approaches to obtaining the NPMLE of a mixing distribution are known to perform badly and approaches using directional derivatives are preferred (see for instance Wang 2007, JRSS B). The best model, with m components, is assumed to have been reached once taking m+1 components does not produce a better model in terms of AIC or BIC. The main issue with this approach is that the EM algorithm is typically very sensitive to the initial parameter values chosen and prone to fail to find a global maximum. A further danger with these models is to ascribe too great a physical significance to the mixture components estimated.
To reach the final model, choices have to be made regarding: the categorization for the responses (observed level of drinking), the latent Markov states for the HMM (e.g. 2, 3 or 4 latent states), the number of mixture components for the random effect (how many "archetypes" of longitudinal behavior) and the degree of time non-homogeneity of transition probabilities. As a result, while the model is likely to explain the observed data reasonably well, a leap of faith is required to believe the model is an accurate representation of the process of binge drinking/alcoholism.
To reach the final model, choices have to be made regarding: the categorization for the responses (observed level of drinking), the latent Markov states for the HMM (e.g. 2, 3 or 4 latent states), the number of mixture components for the random effect (how many "archetypes" of longitudinal behavior) and the degree of time non-homogeneity of transition probabilities. As a result, while the model is likely to explain the observed data reasonably well, a leap of faith is required to believe the model is an accurate representation of the process of binge drinking/alcoholism.
Tuesday, 5 April 2011
Progression of liver cirrhosis to HCC: an application of hidden Markov model.
Nicola Bartolomeo, Paolo Trerotoli and Gabriella Serio have a new paper in BMC Medical Research Methodology. This applies a three state hidden Markov model to data on the progression of liver cirrhosis to Hepatocellular carcinoma. A time homogeneous continuous time progressive model is fitted with death as the absorbing state. Covariate effects are included via a proportional intensities model.
Schoenfeld residuals, which are appropriate for right censored data, are applied here as a test of proportionality. It isn't made clear precisely how this is done here. If time of death is known exactly then a Schoenfeld type residual could be defined for the times of death replacing the standard formulation

with

where
If the times of death are interval censored then this approach is inappropriate.
On a more trivial level the matrix of misclassification probabilities is missing a 1 in the third row corresponding to the absorbing state.
Schoenfeld residuals, which are appropriate for right censored data, are applied here as a test of proportionality. It isn't made clear precisely how this is done here. If time of death is known exactly then a Schoenfeld type residual could be defined for the times of death replacing the standard formulation
with
where
On a more trivial level the matrix of misclassification probabilities is missing a 1 in the third row corresponding to the absorbing state.
Monday, 21 February 2011
A Hidden Markov Model for Informative Dropout in Longitudinal Response Data with Crisis States
Spagnoli, Henderson, Boys and Houwing-Duistermaat have a new paper in Statistics and Probability Letters. The paper is concerned with the modelling of discrete-time continuous response longitudinal data in the presence of informative dropout. The process of dropout is modelled by a three-state (potentially non-homogeneous) discrete time Markov model. The three states correspond to stable, crisis and dropout. The crisis state is characterized by a higher probability of
dropout and a shift in the mean response. The shift is random but is fixed for each individual so that repeated visits to the crisis state have a cumulative effect on the mean. The model is motivated by studies in which dropout is more likely to occur after the treatment has been ineffective for a period of time. A linear-mixed model, with random slope and intercept, is taken for the responses, with the response at time m being shifted by a random quantity, d, times by the number of time periods spent in the crisis state.
The authors put the model within the standard Rubin framework of MCAR/MAR/MNAR, making a distinction between observable and latent filtration. They are careful to formulate the model so that the latent mechanism for dropout only depends on the past and not the future.
The model is applied to schizophrenia data and the Leiden 85+ data. In the former case, interest is in the mean response curves in the hypothetical situation of no dropout. However for the Leiden 85+ data dropout is death and thus such curves would have little meaning. For both cases the crisis-state model represents a substantial improvement in likelihood compared to a two-state model.
dropout and a shift in the mean response. The shift is random but is fixed for each individual so that repeated visits to the crisis state have a cumulative effect on the mean. The model is motivated by studies in which dropout is more likely to occur after the treatment has been ineffective for a period of time. A linear-mixed model, with random slope and intercept, is taken for the responses, with the response at time m being shifted by a random quantity, d, times by the number of time periods spent in the crisis state.
The authors put the model within the standard Rubin framework of MCAR/MAR/MNAR, making a distinction between observable and latent filtration. They are careful to formulate the model so that the latent mechanism for dropout only depends on the past and not the future.
The model is applied to schizophrenia data and the Leiden 85+ data. In the former case, interest is in the mean response curves in the hypothetical situation of no dropout. However for the Leiden 85+ data dropout is death and thus such curves would have little meaning. For both cases the crisis-state model represents a substantial improvement in likelihood compared to a two-state model.
Friday, 7 January 2011
Multi-State Models for Panel Data: The msm Package for R
The final paper in the special issue in Journal of Statistical Software is by Chris Jackson and is about his package msm. msm has been around for many (over 8) years and has steadily built up functionality over time. As noted by Hein Putter in his introduction, msm is different from the other packages featured in the issue in that it deals with panel observed/interval censored data and concentrates on parametric models. In particular, hidden Markov models with both discrete and continuous responses may be fitted.
The paper covers much old ground, the early part repeating similar themes from Jackson & Sharples 2002 (Statistics in Medicine) and Jackson et al 2003 (JRSS D), the section on model diagnostics closely follows Titman and Sharples 2010 (SMMR). One error in msm is in the implementation of prevalence counts/plots. It is made clear by Gentleman et al (1994, Stat Med) and Titman and Sharples that the denominator for the counts at time t is the number "under observation". In particular, subjects who reach the absorbing state should stay under observation only until the time at which they would have been censored. msm assumes subjects who reach the absorbing state remain under observation indefinitely which leads to overestimation of the empirical prevalence in the absorbing state(s). This can be seen clearly in the bottom-right panel of figure 4 of the paper that is suggesting spurious lack of fit. Patients either need to be removed from observation at a known administrative censoring time, or else the censoring distribution needs to be empirically estimated to allow time-dependent weighting of dead patients.
In Section 6 Jackson discusses model extensions which are generally not available in the package, but seems to suggest that to maintain generality these extensions, such as a wider range of time inhomogeneous models, random effects models and semi-Markov models, will not be incorporated into msm.
In general msm is a very good package and lots of effort (e.g. C programming, use of first derivatives, direct coding of transition probabilities for more basic model structures) has been put in to provide a fast performance. However, some improvement in computation speed is no doubt possible since msm uses the BFGS algorithm in optim to fit models rather than applying the well-known Fisher scoring algorithm.
The paper covers much old ground, the early part repeating similar themes from Jackson & Sharples 2002 (Statistics in Medicine) and Jackson et al 2003 (JRSS D), the section on model diagnostics closely follows Titman and Sharples 2010 (SMMR). One error in msm is in the implementation of prevalence counts/plots. It is made clear by Gentleman et al (1994, Stat Med) and Titman and Sharples that the denominator for the counts at time t is the number "under observation". In particular, subjects who reach the absorbing state should stay under observation only until the time at which they would have been censored. msm assumes subjects who reach the absorbing state remain under observation indefinitely which leads to overestimation of the empirical prevalence in the absorbing state(s). This can be seen clearly in the bottom-right panel of figure 4 of the paper that is suggesting spurious lack of fit. Patients either need to be removed from observation at a known administrative censoring time, or else the censoring distribution needs to be empirically estimated to allow time-dependent weighting of dead patients.
In Section 6 Jackson discusses model extensions which are generally not available in the package, but seems to suggest that to maintain generality these extensions, such as a wider range of time inhomogeneous models, random effects models and semi-Markov models, will not be incorporated into msm.
In general msm is a very good package and lots of effort (e.g. C programming, use of first derivatives, direct coding of transition probabilities for more basic model structures) has been put in to provide a fast performance. However, some improvement in computation speed is no doubt possible since msm uses the BFGS algorithm in optim to fit models rather than applying the well-known Fisher scoring algorithm.
Tuesday, 29 June 2010
Hidden Markov models with arbitrary state dwell-time distributions
Langrock and Zucchini have a new paper in Computational Statistics and Data Analysis. This develops models methodology for fitting discrete-time hidden semi-Markov models. They demonstrate that hidden semi-Markov models can be approximated through hidden Markov models with aggregate blocks of states. Dwell-time in a particular state is made up of a series of latent states. An individual enters a state in latent state 1. After k steps in a state where k < m , a subject either leaves the state with probability c(k) or progresses to the next latent state with probability 1-c(k). c(k) therefore represents the hazard rate of the dwell time distribution. If time m in the state is reached then the subject may either stay in latent state m with some probability c(m) or else leave the state. Hence the tail of the dwell-time distribution is constrained to be geometric. By choosing m sufficiently large a good approximation to the desired distribution can be found.
A multistate model for events defined by prolonged observation
Vern Farewell and Li Su have a new paper in Biostatistics. This models remission in psoriatic arthritis, based on panel data assessing joints at clinic visits. Existing models use a two-state model. However, a spell in remission should last a discernible amount of time. Rather than specify an artificial length of time (e.g. 6 months) e.g. a guarantee time in a state, Farewell and Su adopt a model with two states referring to remission: they refer to these as "early stage remission" and "established remission". It is assumed an individual must progress through both early and established remission before returning to active disease. A subject is observed to be in early stage remission if they have no active joints at a visit not preceded by at least 2 other zero count visits. They are in established remission if there is a zero count and at least 2 previous zero counts. State misclassification is allowed in the model through misclassification of the active disease count, i.e. patients may have 0 active joints without being in remission. It is assumed that misclassification to the early stage remission is possible but not to the established remission stage. In the example the misclassification probability is also allowed to depend on whether the previous observed count was zero or not.
The basic problem with the method is that having the states defined by the pattern of previous observations means that it is essentially impossible for the observed data (in terms of the three-state model) to come from the claimed Markov model: In the observed data, an established remission stage must be preceded by two early stage remission observations. Yet the actual Markov model allows the passage time from active disease to established remission to be arbitrarily close to zero (e.g. just the sum of two independent exponential - or perhaps piecewise exponential - distributions). As a result it is not clear how to interpret the resulting transition intensity estimates since the estimated process will not reproduce the original data.
Misclassification is effectively dealt with twice in the model. Firstly in an ad hoc way through rules on what early and established remission are. Then by allowing these observed states to have classification error over some true states. But in fitting a hidden Markov model the misclassification is assumed independent conditional on the underlying state. There is then an inherent contradiction because on the one hand the model says P(Observed zero | Active disease) >0, but at the same time P(Observed zero and two previous zero | Active disease) =0.
An approach using a guarantee time (e.g. Kang and Lagakos (2007)) or perhaps an Erlang distribution through latent states would be far more satisfactory even if it might require "special software". Potentially the guarantee time could be dependent on covariates.
The basic problem with the method is that having the states defined by the pattern of previous observations means that it is essentially impossible for the observed data (in terms of the three-state model) to come from the claimed Markov model: In the observed data, an established remission stage must be preceded by two early stage remission observations. Yet the actual Markov model allows the passage time from active disease to established remission to be arbitrarily close to zero (e.g. just the sum of two independent exponential - or perhaps piecewise exponential - distributions). As a result it is not clear how to interpret the resulting transition intensity estimates since the estimated process will not reproduce the original data.
Misclassification is effectively dealt with twice in the model. Firstly in an ad hoc way through rules on what early and established remission are. Then by allowing these observed states to have classification error over some true states. But in fitting a hidden Markov model the misclassification is assumed independent conditional on the underlying state. There is then an inherent contradiction because on the one hand the model says P(Observed zero | Active disease) >0, but at the same time P(Observed zero and two previous zero | Active disease) =0.
An approach using a guarantee time (e.g. Kang and Lagakos (2007)) or perhaps an Erlang distribution through latent states would be far more satisfactory even if it might require "special software". Potentially the guarantee time could be dependent on covariates.
Monday, 14 June 2010
An application of hidden Markov models to French variant Creutzfeldt-Jakob disease epidemic
Chadeau-Hyam et al have a new paper in Applied Statistics (JRSS C). This is concerned with modelling vCJD in France. A 5 state multi-state model is assumed, with states representing susceptible to infection, asymptomatic infection, clinical vCJD, death from vCJD and death from causes other than vCJD. The data available are extremely sparse since no reliable test is available to distinguish susceptible from asymptomatic. Indeed the only data actually observed are the yearly transitions from infected to clinical vCJD and clinical vCJD to death. As a result, pseudo-observed quantities, estimated in previous studies or from general population data are used to get quantities such as the numbers susceptible. Various approximations in terms of the number and type of transitions possible by an individual in one year are also made. Some simulations are performed which suggest the results are reasonably robust to these approximations.
The most interesting methodological aspect of the paper is the use of an (approximately) Erlang distribution for the incubation time (rather than an Exponential). This is achieved by assuming that the incubation state is made up of 11 latent phases.
The most interesting methodological aspect of the paper is the use of an (approximately) Erlang distribution for the incubation time (rather than an Exponential). This is achieved by assuming that the incubation state is made up of 11 latent phases.
Wednesday, 5 May 2010
Multistate Markov models for disease progression in the presence of informative examination times
Sweeting, Farewell and De Angelis have a new paper in Statistics in Medicine. This deals with the problem of a panel observed disease process where the examination times are generated by a non-ignorable mechanism. This is in general a very difficult problem. The authors consider a special case where, while the disease process is only observed at informative examination times, an auxiliary variable is able to be observed at a full set of planned (or ignorable) times. They consider a model where the disease process is missing at random (MAR) conditional on the values of the auxiliary variable. The likelihood then becomes of the form of a (partially) hidden Markov model. An alternative missing not at random model (MNAR) with missingness dependent on the current state is also considered. However the comparison between the MAR and MNAR models is somewhat unfair. In the simulations and the application, the disease process is only observed at a minority of planned examination times. The auxiliary variable has a fairly strong correlation with the disease process, meaning significant information about the process can be obtained through using the auxiliary variable in a hidden Markov model. Indeed, even without the suspicion of informative missingness the HMM would be worth using (if the strong assumptions about the relationship with the auxiliary variable could be reliably assumed) for improved efficiency. However, the MNAR model does not use the auxiliary variable at all. A better model to compare would be a hybrid of the two where a partially HMM is used in conjunction with the logistic model for the MNAR (although this wouldn't necessarily be consistent under AD-MAR unless the logistic model was correctly specified). This does not seem like a difficult extension.
A further shortcoming of the paper is the lack of any suggestions for model checking. Measurements from the auxiliary variable are assumed to be Normally distributed and independent conditional on the underlying disease states at the examination times. This seems fairly unrealistic and should at least be supported by the data.
A further shortcoming of the paper is the lack of any suggestions for model checking. Measurements from the auxiliary variable are assumed to be Normally distributed and independent conditional on the underlying disease states at the examination times. This seems fairly unrealistic and should at least be supported by the data.
Monday, 22 February 2010
Non-Markov Multistate Modeling Using Time-Varying Covariates
Bacchetti et al have a new paper in The International Journal of Biostatistics. This considers modelling progression of liver fibrosis due to hepatitis C following liver transplant using a 5-state progressive multi-state model. The data are panel observed at irregular time points, so it would be most natural to model the data in continuous time. In addition the observed states are subject to classification error. To avoid having to making either Markov or time homogeneity assumptions, the authors adopt a discrete time assumption: assuming 4 time periods per year. They then model the transition probabilities as linear on the log-odds scale, depending on covariates such as medical center, donor age, year of transplant as well as log time since entry into the current state (relaxing the Markov assumption). Potentially, time since transplant could also be included in the model (for non-homogeneity).
The main challenge in fitting the models is to enumerate all possible complete "paths" of true states at the discrete time points that could result in the observed data. Obviously an iterative algorithm is needed for this. The computational complexity will depend on the complexity of the state transition matrix, the misclassification probability matrix and the degree of discretization.
The main drawback of approximating a continuous time process by a discrete-time process is the restriction that only 1 transition may occur between time points. While this can be acceptable for progressive models such as the one considered here, it may be more problematic when backward transitions are allowable.
The authors have developed an R package called mspath, designed to complement the existing package for continuous-time Markov and hidden Markov model msm, to allow non-Markov models through the discrete-time approximation.
The main challenge in fitting the models is to enumerate all possible complete "paths" of true states at the discrete time points that could result in the observed data. Obviously an iterative algorithm is needed for this. The computational complexity will depend on the complexity of the state transition matrix, the misclassification probability matrix and the degree of discretization.
The main drawback of approximating a continuous time process by a discrete-time process is the restriction that only 1 transition may occur between time points. While this can be acceptable for progressive models such as the one considered here, it may be more problematic when backward transitions are allowable.
The authors have developed an R package called mspath, designed to complement the existing package for continuous-time Markov and hidden Markov model msm, to allow non-Markov models through the discrete-time approximation.
Monday, 23 November 2009
Semi-Markov models with phase-type sojourn distributions
Titman and Sharples have a new paper in Biometrics. This concerns the fitting of semi-Markov models to panel observed data. They propose to fit models where the sojourn time in each state has a phase-type sojourn distribution - i.e. corresponds to the time to absorption of some time homogeneous Markov model. The advantage of this specification is that, unlike general semi-Markov models, the likelihood remains analytically tractable, falling within a hidden Markov model framework. This also makes the extension to models where the observations are subject to misclassification error straightforward, at least theoretically. A two-phase Coxian phase-type distribution is proposed for the sojourn time, allowing increasing, decreasing or constant hazards with respect to time since entry into the state.
While the phase-type framework makes computation of the likelihood more straightforward, model fitting is still potentially problematic due to possible problems of parameter estimability. Also since certain parameters of the phase-type model are unidentifiable under a Markov model meaning an (approximate) modified likelihood ratio test is required to test the Markov assumption.
While the phase-type framework makes computation of the likelihood more straightforward, model fitting is still potentially problematic due to possible problems of parameter estimability. Also since certain parameters of the phase-type model are unidentifiable under a Markov model meaning an (approximate) modified likelihood ratio test is required to test the Markov assumption.
Wednesday, 11 November 2009
Computation of the asymptotic null distribution of goodness-of-fit tests for multi-state models
Andrew Titman has a new paper in Lifetime Data Analysis. This is essentially a continuation of previous papers by Aguirre-Hernandez and Farewell and by Titman and Sharples on Pearson-type goodness-of-fit tests for Markov and hidden Markov models on panel observed data. A practical problem with the tests is that the null distribution depends on the true parameter value and the observation scheme and that a chi-squared approximation can perform inadequately. A parametric bootstrap could be used to find the upper 95% point of the distribution. However, for many models the re-fitting required may take an unacceptable amount of time. Titman shows that, conditional on a fixed observation scheme, the asymptotic distribution can be expressed as a weighted sum of independent
random variables, where the weights depend on the true parameter values. A simulation study shows that computing the weights based on the maximum likelihood estimate of the parameter values, gives tests of close to the appropriate size for realistic sample sizes. The method can be applied to both Markov and misclassification-type hidden Markov models, but only when all transitions are interval-censored.
Tuesday, 4 August 2009
Model diagnostics for multi-state models
Titman and Sharples have a new review paper in SMMR. This considers methods for assessing fit in parametric, panel observed multi-state models. The primary focus is on the assessment of time homogeneous Markov models, although there is also a section on hidden Markov models that occur if states are considered to be observed with classification error. Methods for fitting more complicated models such as non-homogeneous and random effects models are also reviewed. A simple graphical generalization of the prevalence counts method of Gentleman et al (Stats in Med, 1994) is also developed.
Monday, 20 April 2009
Parameter estimation in a model for misclassified Markov data - a Bayesian approach.
Rosychuk and Islam have a paper in Computational Statistics and Data Analysis. This concerns parameter estimation in a two-state recurrent misclassification type hidden Markov model, where the Markov process is assumed to be continuous time and in equilibrium and is observed at discrete, equally spaced time points. A Bayesian approach to estimation is considered via Gibbs sampling. To avoid identifiability issues, the misclassification probabilities are constrained to be below 0.5. An additional issue is the choice of starting values of the transition probabilities for the latent Markov process. Values based on simple correction formulae previously developed by Rosychuk and Thompson appear to perform better than values based on taking naive estimates of the transition probabilities of the observed process.
Tuesday, 24 March 2009
Estimating life expectancy in health and ill health by using a hidden Markov model
Van den Hout, Jagger and Matthews have a paper to appear in JRSS C. The paper applies the misclassification hidden Markov model, developed by Satten and Longini and Jackson and Sharples, to modelling of data on cognitive impairment in the elderly and its effect on mortality. Patients with a cognition score (MMSE) below 22 were considered impaired. However, cognitive decline is considered to be progressive so backwards transitions in the dataset are explained through misclassification.
The main aim of the paper is to estimate life expectancies in the non-impaired and impaired states. As mortality will be highly dependent on age, non-homogeneous transition intensities are required. Rather than employ the standard approach of piecewise constant intensities, the authors instead include age as a log-linear time dependent covariate and assume that an individual observed at ages t and u, for t < u, has constant intensity Q(t) for the interval (t,u). This will clearly result in some degree of bias, particularly if observation times are widely spaced. Life expectancy is then calculated by assuming intensities are constant in 1 year intervals. As this is different from how the data were estimated, the bias may be further compounded.
Rudimentary goodness-of-fit is carried out by comparing estimated survival curves from the HMM with a Cox-regression performed directly on the survival data. It is worth noting that this approach could be problematic in certain circumstances because the HMM is not nested within the Cox-regression model, so there might be discrepancies between the curves even if the HMM is correctly specified.
The main aim of the paper is to estimate life expectancies in the non-impaired and impaired states. As mortality will be highly dependent on age, non-homogeneous transition intensities are required. Rather than employ the standard approach of piecewise constant intensities, the authors instead include age as a log-linear time dependent covariate and assume that an individual observed at ages t and u, for t < u, has constant intensity Q(t) for the interval (t,u). This will clearly result in some degree of bias, particularly if observation times are widely spaced. Life expectancy is then calculated by assuming intensities are constant in 1 year intervals. As this is different from how the data were estimated, the bias may be further compounded.
Rudimentary goodness-of-fit is carried out by comparing estimated survival curves from the HMM with a Cox-regression performed directly on the survival data. It is worth noting that this approach could be problematic in certain circumstances because the HMM is not nested within the Cox-regression model, so there might be discrepancies between the curves even if the HMM is correctly specified.
Subscribe to:
Posts (Atom)