Showing posts with label Bayesian. Show all posts
Showing posts with label Bayesian. Show all posts
Wednesday, 11 July 2012
Bayesian analysis of a disability model for lung cancer survival
Armero, Cabras, Castellanos, Perra, Quirós, Oruezábal and Sánchez-Rubio have a new paper in Statistical Methods in Medical Research. This develops a Bayesian three-state illness-death type model to the progression and survival of lung cancer patients. The data considered are assumed to be complete up to right censoring (in reality there may be some interval censoring but the authors argue the patients can be considered `quasi-continuously' followed. A Weibull semi-Markov model is assumed for the transition intensities and covariates are accommodated via an accelerated failure time model (it's worth noting that for a Weibull distribution the proportional hazard and accelerated failure time models are equivalent up to reparameterization). A feature of the dataset used is the rather small sample size (35 patients) which is perhaps the strongest reason for taking a parametric and Bayesian approach in this case.
Wednesday, 18 April 2012
Bayesian inference of the fully specified subdistribution model for survival data with competing risks
Miaomiao Ge and Ming-Hui Chen have a new paper in Lifetime Data Analysis. This considers methods for Bayesian inference in the Fine-Gray model for competing risks. In order to perform the Bayesian analysis it is necessary to fully specify the model for the competing risk(s) that are not of direct interest in the analysis. Ge and Chen propose a model in which for failure time
,
and a proportional hazards model affects
via
.
Note that
does not affect
or
.
The authors consider approaches to non-parametric Bayesian inference using Gamma process priors, but also use piecewise constant hazard models (with fixed cut points) where the hazard in each time period has an independent gamma prior.
Note that
The authors consider approaches to non-parametric Bayesian inference using Gamma process priors, but also use piecewise constant hazard models (with fixed cut points) where the hazard in each time period has an independent gamma prior.
Friday, 9 March 2012
Estimating Discrete Markov Models From Various Incomplete Data Schemes
Alberto Pasanisi, Shuai Fu and Nicolas Bousquet have a new paper in Computational Statistics & Data Analysis. This considers approaches to Bayesian inference for time-homogeneous discrete-time Markov models under incomplete observation. Firstly, they consider where there are missing observations in a sequence of states (considering different missingness assumptions). Secondly, they consider the case of aggregate data where all that is known is the number of subjects in each state at each time. A Bayesian approach is adopted throughout, which the authors claim is the most convenient in this situation.
The part involving incomplete data follows similar ground to Deltour et al (Biometrics, 1999). The problem only becomes non-trivial if the missingness mechanism is non-ignorable.
The treatment of aggregate data is incorrect as the authors state the likelihood as being the product of independent multinomial random variables with probabilities corresponding to the probability of being in a state at time t given the initial state distribution at time 0. As a result they claim the likelihood is proportional to the case of current status data where each subject or unit is only observed once. The reason that likelihood-based inference for aggregate data is so difficult is that we observe all units multiple times but don't know number or nature of the transitions that occurred. Hence, the full likelihood would require summing over all possible transitions consistent with the aggregate counts. Kalbfleisch and Lawless (Canadian Journal of Statistics, 1984) derived the mean and covariance of the aggregated counts across times to establish a least-squares estimation procedure. Pasanisi et al's procedure is only relevant when the data consist of a series of independent cross-sectional surveys at different time points all assumed to come from different units. An MCMC or simulation based approach would be necessary to compute the exact likelihood or posterior distribution in the true aggregate data case, which the authors did not pursue. However, the gain in efficiency compared to the least-squares approach is probably not worth the trouble except for very small counts. Crowder and Stephens (2011) pursued an approach based on matching the coefficients of the probability generating function of the aggregate counts.
The part involving incomplete data follows similar ground to Deltour et al (Biometrics, 1999). The problem only becomes non-trivial if the missingness mechanism is non-ignorable.
The treatment of aggregate data is incorrect as the authors state the likelihood as being the product of independent multinomial random variables with probabilities corresponding to the probability of being in a state at time t given the initial state distribution at time 0. As a result they claim the likelihood is proportional to the case of current status data where each subject or unit is only observed once. The reason that likelihood-based inference for aggregate data is so difficult is that we observe all units multiple times but don't know number or nature of the transitions that occurred. Hence, the full likelihood would require summing over all possible transitions consistent with the aggregate counts. Kalbfleisch and Lawless (Canadian Journal of Statistics, 1984) derived the mean and covariance of the aggregated counts across times to establish a least-squares estimation procedure. Pasanisi et al's procedure is only relevant when the data consist of a series of independent cross-sectional surveys at different time points all assumed to come from different units. An MCMC or simulation based approach would be necessary to compute the exact likelihood or posterior distribution in the true aggregate data case, which the authors did not pursue. However, the gain in efficiency compared to the least-squares approach is probably not worth the trouble except for very small counts. Crowder and Stephens (2011) pursued an approach based on matching the coefficients of the probability generating function of the aggregate counts.
Sunday, 1 January 2012
Bayesian analysis of multistate event history data: beta-Dirichlet process prior
Yongdai Kim, Lancelot James and Rafael Weissbach have a new paper in Biometrika. This develops a conjugate prior process suitable for non-parametric and semi-parametric Bayesian modelling of right-censored multi-state Markov data. The model is parametrised in terms of the sum of the intensities out of each state

and instantaneous transition probabilities

A possible choice for a prior process is a Dirichlet distribution but this is not independent in the limit of a continuous time process. Instead the authors propose a new beta-Dirichlet process consisting of a beta distributed part which determines the increment in
(between 0 and 1) and a Dirichlet part determining the instantaneous transition probabilities for each particular transition. The authors prove this prior process is conjugate in the continuous limit.
A semi-parametric regression model is proposed, which the authors term as a semi-proportional intensities model. This consists of a proportional intensities model for the all-cause hazard of exiting state h and a multinomial type model for the instantaneous transition probabilities out of state h and bears some resemblance to the vertical modeling parametrization for competing risks regression.
In an aside the authors claim that interval censoring can easily be dealt with by treating the unknown transition time as missing data that can be accounted for in the Gibbs sampling. This only works under the assumption that only one transition can have occurred between examination times. While other authors have made this assumption (e.g. Foucher et al 2007) it is dubious to say the least and likely to result in biased estimates. Similarly, the authors claim right-censoring can be dealt with by treating a censoring event as an additional state. While this will obviously allow the observed process to be modelled, it is not clear how this approach would allow the underlying process (without censoring) is estimated?
and instantaneous transition probabilities
A possible choice for a prior process is a Dirichlet distribution but this is not independent in the limit of a continuous time process. Instead the authors propose a new beta-Dirichlet process consisting of a beta distributed part which determines the increment in
A semi-parametric regression model is proposed, which the authors term as a semi-proportional intensities model. This consists of a proportional intensities model for the all-cause hazard of exiting state h and a multinomial type model for the instantaneous transition probabilities out of state h and bears some resemblance to the vertical modeling parametrization for competing risks regression.
In an aside the authors claim that interval censoring can easily be dealt with by treating the unknown transition time as missing data that can be accounted for in the Gibbs sampling. This only works under the assumption that only one transition can have occurred between examination times. While other authors have made this assumption (e.g. Foucher et al 2007) it is dubious to say the least and likely to result in biased estimates. Similarly, the authors claim right-censoring can be dealt with by treating a censoring event as an additional state. While this will obviously allow the observed process to be modelled, it is not clear how this approach would allow the underlying process (without censoring) is estimated?
Labels:
Bayesian,
Biometrika,
Markov,
non-parametric,
right censoring,
semi-parametric
Thursday, 28 April 2011
Bayesian evidence synthesis for a transmission dynamic model for HIV among men who have sex with men
Presanis, De Angelis, Goubar, Gill and Ades have a new paper in Biostatistics. The paper attempts to use several disparate sources of data on prevalence and incidence to estimate transition rates and prevalences in a multi-state model of transmission of HIV among men who have sex with men (ironically abbreviated as MSM) via Bayesian evidence synthesis. The multi-state model has 4 transient states relating to "Eligible (non MSM)", "Susceptible (MSM)", "Undiagnosed" and "Diagnosed". Subjects enter aged 15 (or later due to migration) and may exit due to death or reaching age 45.
It should be noted that the multi-state model is deterministic and not stochastic, i.e. the proportions in each state at a given time conditional on the parameters and starting conditions are the solution of an ordinary differential equation. While this is unproblematic in terms of getting estimates of the expected proportions in each state at a given time, there must be some level of under representation of the uncertainty. In particular the likelihood assumes that the number of deaths from HIV from the SOPHID dataset at time
is
where
is the number in the diagnosed state at time
and
is the probability of death in one year from HIV for the year
. In reality D(t) should be stochastic rather than deterministic and as such we would expect the number of deaths to have a higher variance than assumed by the binomial model.
In principle the covariance matrix of the numbers in each state for each year could be derived, i.e. given the Markov model, we have N independent individuals each of whom's state occupancy at time
is multinomial given the occupancy at time
. A more realistic model would then assume that conditional on the parameters, the state occupation counts have a multivariate normal distribution (e.g. analogous to the approaches taken in the estimation of aggregate Markov data) and, for instance, the number of deaths from HIV is binomial with a denominator that is itself normally distributed. Quite possibly this extra source of uncertainty is negligible compared to the vast existing uncertainties but it nevertheless ought to be explored.
It should be noted that the multi-state model is deterministic and not stochastic, i.e. the proportions in each state at a given time conditional on the parameters and starting conditions are the solution of an ordinary differential equation. While this is unproblematic in terms of getting estimates of the expected proportions in each state at a given time, there must be some level of under representation of the uncertainty. In particular the likelihood assumes that the number of deaths from HIV from the SOPHID dataset at time
In principle the covariance matrix of the numbers in each state for each year could be derived, i.e. given the Markov model, we have N independent individuals each of whom's state occupancy at time
Friday, 17 December 2010
Estimating the distribution of the window period for recent HIV infections: A comparison of statistical methods
Michael Sweeting, Daniela De Angelis, John Parry and Barbara Suligo have a new paper in Statistics in Medicine. This considers estimating the time to a biomarker crossing a threshold from seroconversion in doubly-censored data (denoted T). Such methods have been used in the past to model time to seroconversion from infection. Sweeting et al state in the introduction that a key motivation is to allow the prevalence of recent infection to be estimated. This is defined as

where S(t) is the time from seroconversion to biomarker crossing. Here we see there is a clear assumption that T is independent of calendar time d. This is consistent with the framework first proposed by De Gruttola and Lagakos (Biometrics, 1989) where the time to first event is assumed to be independent of the time between first and second events, the data are essentially a three-state progressive semi-Markov model where interest lies in estimating the sojourn distribution in state 2. Alternatively, the methods of Frydman (JRSSB, 1992) are based on a Markov assumption. Here the marginal distribution of the sojourn in state 2 could be estimated by integrating over the estimated distribution of times to entry into state 2. Sweeting et al however treat it as standard bivariate survival data (where X denotes the time to serocoversion and Z denotes time to biomarker crossing) using MLEcens for estimation. It's not clear whether this is sensible if the aim is to estimate P(d) as defined above. Moreover, Betensky and Finklestein (Statistics in Medicine, 1999) suggested an alternative algorithm would be required for doubly-censored data because of the inherent ordering (i.e. Z>X). Presumably as long as the intervals for seroconversion and crossing are disjoint the NPMLE for S is guaranteed to have support only at non-negative times.
Sweeting et al make a great deal about the fact that since the NPMLE of (X,Z) is non-unique, e.g. the NPMLE can only ascribe mass to regions of the (X,Z) plane, and that T = Z - X, there will be huge uncertainty over the distribution of S between the extremes of assuming mass only in the lower left corner or only in the upper right corner of the support rectangles. They provide a plot to show this in their data example. The original approach to doubly-censored data of De Gruttola and Lagakos assumes a set of discrete time points of mass for S (the generalization to higher order models was given by Sternberg and Satten). This largely avoids any problems of non-uniqueness (though some problems remain see e.g. Griffin and Lagakos 2010).
For the non-parametric approach, Sweeting et al seem extremely reluctant to make any assumptions whatsoever. In contrast, they go completely to town on assumptions once they get onto their preferred Bayesian method. It should firstly be mentioned that the outcome time is time to crossing of a biomarker and there is considerable auxiliary data available in the form of intermediate measurements. Thus, under any kind of monotonicity assumptions, we can see that extra modelling of the growth process of the biomarker has merit. Sweeting et al use a parametric, mixed effects growth model, to model the growth of the biomarker. The growth is measured in terms of time since seroconversion, which is unknown. They consider two methods: a naive method that assumes seroconversion occurs at the midpoint of the time interval and a uniform prior method that assumes a priori that the seroconversion time is distributed uniformly within the times at which it is interval censored (i.e. last time before seroconversion and the first biomarker measurement time). The first method is essentially like imputing the interval midpoint for X. The second method is like assuming the marginal distribution for X is uniform.
Overall, the paper comes across as a straw man argument against non-parametric methods appended onto a Bayesian analysis that makes multiple (unverified) assumptions.
where S(t) is the time from seroconversion to biomarker crossing. Here we see there is a clear assumption that T is independent of calendar time d. This is consistent with the framework first proposed by De Gruttola and Lagakos (Biometrics, 1989) where the time to first event is assumed to be independent of the time between first and second events, the data are essentially a three-state progressive semi-Markov model where interest lies in estimating the sojourn distribution in state 2. Alternatively, the methods of Frydman (JRSSB, 1992) are based on a Markov assumption. Here the marginal distribution of the sojourn in state 2 could be estimated by integrating over the estimated distribution of times to entry into state 2. Sweeting et al however treat it as standard bivariate survival data (where X denotes the time to serocoversion and Z denotes time to biomarker crossing) using MLEcens for estimation. It's not clear whether this is sensible if the aim is to estimate P(d) as defined above. Moreover, Betensky and Finklestein (Statistics in Medicine, 1999) suggested an alternative algorithm would be required for doubly-censored data because of the inherent ordering (i.e. Z>X). Presumably as long as the intervals for seroconversion and crossing are disjoint the NPMLE for S is guaranteed to have support only at non-negative times.
Sweeting et al make a great deal about the fact that since the NPMLE of (X,Z) is non-unique, e.g. the NPMLE can only ascribe mass to regions of the (X,Z) plane, and that T = Z - X, there will be huge uncertainty over the distribution of S between the extremes of assuming mass only in the lower left corner or only in the upper right corner of the support rectangles. They provide a plot to show this in their data example. The original approach to doubly-censored data of De Gruttola and Lagakos assumes a set of discrete time points of mass for S (the generalization to higher order models was given by Sternberg and Satten). This largely avoids any problems of non-uniqueness (though some problems remain see e.g. Griffin and Lagakos 2010).
For the non-parametric approach, Sweeting et al seem extremely reluctant to make any assumptions whatsoever. In contrast, they go completely to town on assumptions once they get onto their preferred Bayesian method. It should firstly be mentioned that the outcome time is time to crossing of a biomarker and there is considerable auxiliary data available in the form of intermediate measurements. Thus, under any kind of monotonicity assumptions, we can see that extra modelling of the growth process of the biomarker has merit. Sweeting et al use a parametric, mixed effects growth model, to model the growth of the biomarker. The growth is measured in terms of time since seroconversion, which is unknown. They consider two methods: a naive method that assumes seroconversion occurs at the midpoint of the time interval and a uniform prior method that assumes a priori that the seroconversion time is distributed uniformly within the times at which it is interval censored (i.e. last time before seroconversion and the first biomarker measurement time). The first method is essentially like imputing the interval midpoint for X. The second method is like assuming the marginal distribution for X is uniform.
Overall, the paper comes across as a straw man argument against non-parametric methods appended onto a Bayesian analysis that makes multiple (unverified) assumptions.
Tuesday, 26 October 2010
Parameterization of treatment effects for meta-analysis in multi-state Markov models
Malcolm Price, Nicky Welton and Tony Ades have a new paper in Statistics in Medicine. This considers how to parameterize Markov models of clinical trials, so that meta-analysis can be performed. They consider data that is panel observed (at regular time intervals) but aggregated into the number of transitions of each type among all patients between two consecutive observations. As a result, it is necessary to assume a time homogeneous Markov model, since there is no information to consider time inhomogeneity or patient inhomogeneity. These types of cost-effectiveness trials are usually analyzed using a discrete time framework (Markov transition model), with the assumption that only one transition is allowed between observations. Price et al advocate using a continuous time method. The main advantage of this is that fewer parameters are required e.g. rare jumps to distant states can be explained by the presence of multiple jumps rather than having to include the transition probability as an extra parameter in the model.
Price et al consider asthma trial data where they are interested in combining data from 5 distinct two-arm RCTs, there is some overlap between the treatments tested so evidence networks comparing treatments can be constructed.
The main weakness of the approach is the reliance on the DIC. Different models for the set of non-zero transition intensities are considered using the DIC, with the models allowing different intensities for each trial treatment arm. While later in the paper there are benefits of adopting the Bayesian paradigm, here it doesn't seem useful. For these fixed effects models and given the flat priors chosen, DIC should (asymptotically) be the same as AIC. However, AIC is dubious here also because of the aspect of testing on the boundary of the parameter space and is likely to put too much favour on simpler models in this context. Likelihood ratio tests based on mixtures of Chi-squared distributions (Self and Liang, 1987) could be applied as in Gentleman et al 1994. However, judging from the example data given for treatments A and D, the real issue is lack of data, e.g. there are very few transitions to and from state X.
Price et al consider asthma trial data where they are interested in combining data from 5 distinct two-arm RCTs, there is some overlap between the treatments tested so evidence networks comparing treatments can be constructed.
The main weakness of the approach is the reliance on the DIC. Different models for the set of non-zero transition intensities are considered using the DIC, with the models allowing different intensities for each trial treatment arm. While later in the paper there are benefits of adopting the Bayesian paradigm, here it doesn't seem useful. For these fixed effects models and given the flat priors chosen, DIC should (asymptotically) be the same as AIC. However, AIC is dubious here also because of the aspect of testing on the boundary of the parameter space and is likely to put too much favour on simpler models in this context. Likelihood ratio tests based on mixtures of Chi-squared distributions (Self and Liang, 1987) could be applied as in Gentleman et al 1994. However, judging from the example data given for treatments A and D, the real issue is lack of data, e.g. there are very few transitions to and from state X.
Monday, 14 June 2010
An application of hidden Markov models to French variant Creutzfeldt-Jakob disease epidemic
Chadeau-Hyam et al have a new paper in Applied Statistics (JRSS C). This is concerned with modelling vCJD in France. A 5 state multi-state model is assumed, with states representing susceptible to infection, asymptomatic infection, clinical vCJD, death from vCJD and death from causes other than vCJD. The data available are extremely sparse since no reliable test is available to distinguish susceptible from asymptomatic. Indeed the only data actually observed are the yearly transitions from infected to clinical vCJD and clinical vCJD to death. As a result, pseudo-observed quantities, estimated in previous studies or from general population data are used to get quantities such as the numbers susceptible. Various approximations in terms of the number and type of transitions possible by an individual in one year are also made. Some simulations are performed which suggest the results are reasonably robust to these approximations.
The most interesting methodological aspect of the paper is the use of an (approximately) Erlang distribution for the incubation time (rather than an Exponential). This is achieved by assuming that the incubation state is made up of 11 latent phases.
The most interesting methodological aspect of the paper is the use of an (approximately) Erlang distribution for the incubation time (rather than an Exponential). This is achieved by assuming that the incubation state is made up of 11 latent phases.
Tuesday, 22 September 2009
Estimating stroke-free and total life expectancy in the presence of non-ignorable missing values
Van den Hout and Matthews have a new paper in JRSS A. This deals with interval censored data from a three-state disease model where subjects may miss scheduled interviews meaning the disease status is not observed. A joint model for the disease state and an observation indicator is developed, being a continuous time generalisation of Cole et al (2005) that allows information from exact death times to be included. Conditional on the disease state and measured covariates, the observation indicator is governed by a logistic model.
In general the method seems promising for dealing with informative observation when the potential observation times are known.
Though not noted, the model can be expressed as a hidden Markov model. The authors state that the logistic model and the three-state Markov model are estimated separately. It is not made clear how this is achieved since the logistic model depends on the unobserved states of the Markov model. In some cases the missing state will in fact be known, for instance if the sequence is 1,-,1 or 2,-,2. However, for sequences like 1,-,2 or 1,-,3 it is not possible to establish the unobserved state.
The main aim of the analysis is to obtain estimates of life expectancy, disease free life expectancy and post-disease life-expectancy. These are complicated functions of the parameter vector as they involve integrals of transition probabilities. In addition to the method of Aalen et al 1997, Van den Hout and Matthews additionally propose to use a Metropolis algorithm to get confidence intervals for the life expectancies. The resulting intervals have a Bayesian interpretation, being the credible intervals from an improper uniform prior, but will not be invariant to changes in parametrisation. In practice, the intervals may give good frequentist coverage, particularly for large samples. However, Van den Hout and Matthews seem to be implying the intervals have exact coverage (apart from Monte-Carlo error through the Metropolis algorithm) which is a substantial misconception. Moreover, no mention of the procedure being Bayesian is given.
In general the method seems promising for dealing with informative observation when the potential observation times are known.
Though not noted, the model can be expressed as a hidden Markov model. The authors state that the logistic model and the three-state Markov model are estimated separately. It is not made clear how this is achieved since the logistic model depends on the unobserved states of the Markov model. In some cases the missing state will in fact be known, for instance if the sequence is 1,-,1 or 2,-,2. However, for sequences like 1,-,2 or 1,-,3 it is not possible to establish the unobserved state.
The main aim of the analysis is to obtain estimates of life expectancy, disease free life expectancy and post-disease life-expectancy. These are complicated functions of the parameter vector as they involve integrals of transition probabilities. In addition to the method of Aalen et al 1997, Van den Hout and Matthews additionally propose to use a Metropolis algorithm to get confidence intervals for the life expectancies. The resulting intervals have a Bayesian interpretation, being the credible intervals from an improper uniform prior, but will not be invariant to changes in parametrisation. In practice, the intervals may give good frequentist coverage, particularly for large samples. However, Van den Hout and Matthews seem to be implying the intervals have exact coverage (apart from Monte-Carlo error through the Metropolis algorithm) which is a substantial misconception. Moreover, no mention of the procedure being Bayesian is given.
Labels:
Bayesian,
interval censoring,
JRSS A,
missing data,
outcome measures,
parametric
Monday, 3 August 2009
Estimating dementia-free life expectancy for Parkinsons patients using Bayesian inference and microsimulation
Van den Hout and Matthews have a new paper in Biostatistics involving a random-effects Markov model with time (age) dependent intensities. The methodology is close to that used by Pan et al and Wu et al, using a WinBUGS/OpenBUGS Bayesian approach. They use a three-state illness-death model without recovery. A more sophisticated multivariate log-normal random effect on the effects of age on the intensities is used with a Wishart prior is used on the covariance matrix, which is more appropriate than the Gamma(e,e) type priors used by Pan et al. Like their recent Applied Statistics paper, time dependencies in the intensities are accounted for by assuming an individual that is observed at times t1 and t2 has a constant matrix of intensities between those points, but different assumptions are used to calculate life expectancies. The main methodological development is obtaining life expectancy estimates through 'microsimulation.' This is deemed necessary because there are two levels of variation: variation in the posterior of the parameters and variation from the random effects distribution conditional on the parameters. 'Microsimulation' (or simulation) just approximates the integral over the random effects distribution.
Monday, 20 April 2009
Parameter estimation in a model for misclassified Markov data - a Bayesian approach.
Rosychuk and Islam have a paper in Computational Statistics and Data Analysis. This concerns parameter estimation in a two-state recurrent misclassification type hidden Markov model, where the Markov process is assumed to be continuous time and in equilibrium and is observed at discrete, equally spaced time points. A Bayesian approach to estimation is considered via Gibbs sampling. To avoid identifiability issues, the misclassification probabilities are constrained to be below 0.5. An additional issue is the choice of starting values of the transition probabilities for the latent Markov process. Values based on simple correction formulae previously developed by Rosychuk and Thompson appear to perform better than values based on taking naive estimates of the transition probabilities of the observed process.
Subscribe to:
Posts (Atom)