Skip to content

Implement Finite Mixture Models for Multinomial Variables Using e.g., flexmix #233

Description

@matteodelucchi

Currently, abn (v3.x) uses two approaches for handling multinomial variables in the frequentist paradigm (method = "mle"):

  1. nnet for basic MLE estimation (softmax link, neural network with no hidden layers)
  2. mclogit for multinomial variables with random effects to control for clustering​

Both approaches assume a single, homogeneous population in which all observations follow the same multinomial response structure with respect to the parent nodes in the DAG. However, this may not adequately capture latent heterogeneity where different subpopulations exhibit fundamentally different relationships.

Proposed Enhancement: Finite Mixture Models

The flexmix package provides a general framework for finite mixture modeling using the EM algorithm. It can fit mixtures of multinomial logit models via FLXMRmultinom, enabling the identification of latent classes with distinct parameter estimates.​

Potential Advantages

  • Latent Heterogeneity Detection: Identify unobserved subgroups within data where the relationship between parent nodes and multinomial responses differs substantially
    • Example: In clinical data, different patient subgroups may have distinct disease progression patterns not explained by observed covariates alone
  • Improved Model Fit: When population heterogeneity exists but cannot be captured by observed variables, mixture models can provide a better fit and predictive accuracy​
  • Enhanced Interpretability:
    • Mixture component membership probabilities reveal which observations belong to which latent class
    • Component-specific parameters show how DAG relationships vary across subpopulations
    • Could lead to new biological or epidemiological insights about distinct mechanistic pathways?
  • Flexibility in Modeling Complex Phenomena:
    • Handle overdispersion beyond what random effects can capture
    • Model structural changes in relationships (e.g., different causal mechanisms in subgroups)
    • Natural extension to the existing ABN framework without changing the core structure

Implementation Considerations

Maybe something along these lines:

# Conceptual example
buildScoreCache/fitAbn(..., 
       method = "mle",
       mixture = TRUE,          # Enable mixture modeling
       n_components = 2:4,      # Test 2-4 latent classes
       multinomial_engine = "flexmix")  # Use flexmix for multinomial nodes

Key Design Questions:

  1. Scope: Should mixtures apply to:
    • Only multinomial variables? (initial implementation)
    • All node types in the network? (future extension)
    • Selected nodes specified by the user?
  2. Component Number Selection:
    • Implement BIC/AIC criteria for selecting optimal K
    • Allow user-specified K or data-driven selection
    • Multiple random initializations for EM algorithm stability​
  3. Integration with Existing Framework:
    • How to represent mixture components in DAG visualization?
    • How to report component-specific parameters vs. pooled effects?
    • Compatibility with buildScoreCache() for structure learning
  4. Computational Complexity:
    • EM algorithm + structure learning = substantial computational burden?
    • May need parallel processing enhancements?
    • Consider limiting to structure learning within components vs. across all data

Challenges to Address

  1. Identifiability: Mixture models can have identification issues; need to implement constraints or post-processing to ensure interpretability​
  2. Interpretability Trade-off: Mixture parameters are more complex than standard regression; documentation must guide users carefully
  3. Computational Cost: Structure learning is already intensive; mixture modeling adds another layer.
    • May need approximations or heuristics.
    • Consider limiting to post-hoc analysis after the structure is learned.
  4. Integration with Random Effects: How do mixture models interact with existing mclogit random effects?
    • Mixture-of-mixed-effects models? (very complex)
    • Choose one or the other per node/per model?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestquestionFurther information is requested

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions