一覧 Case Studies Clarifying complex factors using exploratory and confirmatory factor analysis

Clarifying complex factors using exploratory and confirmatory factor analysis

In this case study, I describe how Dr.DataScience used its expertise in factor analysis to help a client who wished to reveal latent structures and common factors within a complex dataset they held. In research and business that handle a wide range of variables, deeply understanding the meaning of each variable and their mutual relationships is essential to obtaining substantive insights.

By thoroughly evaluating the characteristics of the data and deriving a statistically valid and practical factor structure, Dr.DataScience helped the client make data-driven decisions with greater confidence. In accordance with our confidentiality agreement, no specific survey item names or figures are disclosed; however, the factor-analysis methods used, the process, and the type of findings obtained are the same as in the actual analysis.

Background and objective

The client had collected diverse information (a set of variables) about a certain subject. They held the hypothesis that these variables might be underpinned by several common latent “factors.” They had a strong wish to understand the data more simply and essentially by revealing this latent structure. In particular, they aimed to solve challenges such as handling too many variables efficiently and organizing the complex relationships among variables.

Data and variables

This analysis used anonymized, large-scale survey data. The subjects were data collected for a particular purpose. Specifically, they included a wide range of measured items concerning the subjects’ attributes, behaviors, and attitudes.

  1. Variables to be analyzed
    • Multiple measured variables, set by the client, that were expected to be explained by latent factors. These mainly consisted of items measured on Likert scales rated in steps from “strongly agree” to “strongly disagree.”
  2. Other data
    • In addition to the Likert-scale data above, there were also attribute data such as sex, age group, region, and occupation, which were considered for possible use in interpretation and segmentation after the factor analysis.

Analytical methods

  1. Data cleaning and preprocessing
    • I identified and excluded from the analysis nominal-scale variables unsuited to factor analysis, as well as variables judged inappropriate for the particular purpose (for example, items with an extremely large number of “don’t know” responses).
    • For some variables, I made corrections such as deleting particular values in order to treat them as ordinal scales.
    • For variables with a large influence, such as “age,” I devised ways to use them in the analysis by aggregating them into appropriate categories (for example, as age groups) to prevent bias in the analysis.
  2. Variable selection
    • To enhance the accuracy and validity of the factor analysis, I carefully selected the variables to be used through the following multi-stage process.
    • Ceiling/floor effect analysis: I checked whether a variable’s values were excessively skewed toward a particular extreme, and judged which items should be excluded.
    • I–T correlation (item–total correlation) analysis: I checked how strongly each item moved together with the overall tendency, and selected items with high correlation coefficients.
    • Good–Poor analysis: I adopted items showing a statistically significant difference between the high-scoring and low-scoring groups.
    • Inter-item correlation analysis: where pairs of items had very high mutual correlation, I excluded one of them to avoid redundancy.
    • KMO (Kaiser–Meyer–Olkin) test: I evaluated whether the data were suitable for factor analysis (whether there was sufficient commonality among variables). I confirmed that the overall and individual KMO values met the criteria, ensuring the validity of conducting factor analysis.
  3. Determining the number of factors
    • To determine the optimal number of factors (the number of latent factors), I comprehensively evaluated multiple statistical indices.
    • Proportion of variance explained: I selected a number of factors whose cumulative proportion provided sufficient explanatory power.
    • VSS (Very Simple Structure): I evaluated the simple-structure fit of the factor solution; larger values suggested a number of factors with higher fit.
    • MAP (Minimum Average Partial): I evaluated how much of the variance of the inter-item correlation matrix was explained by the factors, and considered the number of factors that minimized the value.
    • I rigorously evaluated the optimal number of factors using a wide range of indices, including the chi-squared value, RMSR (Root Mean Square of Residuals), Fit, RMSEA (Root Mean Square Error of Approximation), BIC/SABIC, SRMR (Standardized Root Mean Square Residual), eCRMS, Complex, and parallel analysis.
  4. Exploratory factor analysis
    • For the selected variables, I performed exploratory factor analysis using maximum-likelihood estimation (Promax rotation).
    • This analysis confirmed that the variables were classified into five factors.
    • I associated variables whose factor loadings met a particular criterion with each factor. Variables below the criterion were judged to be unrelated to any factor and were excluded.
    • I also evaluated communality, an index of association with the factors as a whole; although some variables fell below the criterion value, the discrepancy was small and judged not to be a problem.
    • I also checked complexity (how strongly a variable relates to multiple factors), which indicated simplicity of structure.
    • I checked the inter-factor correlations and confirmed there was no excessive overlap among factors, while noting that the correlation between some factors tended to be somewhat high.
    • In the goodness-of-fit evaluation, indices such as RMSR and Fit indicated that a good model had been built.
  5. Confirmatory factor analysis
    • To verify the validity of the factor structure obtained from the exploratory factor analysis, I performed confirmatory factor analysis.
    • From the factor loadings, I judged that the association between each factor and each variable was strong.
    • I checked the values of the unique variances (the portion of each variable not explained by the factors); while noting that some variables did not meet the criterion, the small inter-error covariances led me to judge that the errors were independent, supporting the reliability of the factor loadings.
    • The reliability coefficient Cronbach’s α reached a practically acceptable criterion for all factors, confirming high factor reliability. The possibility of redundancy was also extremely low.
    • The factor covariances tended to be somewhat high in places, but not to the extent of strongly recommending merging factors.
    • In the goodness-of-fit evaluation, good results were obtained on multiple indices such as CFI, TLI, GFI, AGFI, SRMR, and RMSEA, and the model was judged to be highly reliable and valid.

Overview of the main results and clinical considerations

Through a multifaceted and rigorous factor analysis, it became clear that five statistically distinct latent factors existed within the client’s diverse data. These factors were strongly associated with particular groups of variables, and I succeeded in organizing the data’s complexity into a more understandable form with substantive meaning.

This finding provides strong grounds for the client to focus on the most influential factors among an enormous amount of data and make more effective and efficient decisions in future research and development or in considering clinical interventions. For example, where a particular factor is suggested to be important, concentrating resources on the group of variables related to that factor makes it possible to maximize outcomes.

Dr.DataScience’s contribution

This case demonstrated how deeply and practically Kenichiro Suzuki, the statistical-analysis specialist of Dr.DataScience, can contribute to revealing the latent structure within a client’s complex data and deriving practical insights.

For the enormous set of variables the client held, I first performed thorough data cleaning and variable selection to build a dataset optimal for factor analysis. On that basis, by following a rigorous, multi-index process for determining the number of factors, I identified the optimal number of latent factors free from arbitrariness.

Furthermore, by using a two-stage approach of exploratory and confirmatory factor analysis, I ensured the statistical validity and reliability of the derived factor structure at the highest level. In particular, as with the RMST model application example, I am proud that Dr.DataScience’s ability to deeply understand the characteristics of the data and to select and apply the optimal analytical method accordingly was fully demonstrated here as well.

The resulting factor structure gave the client a clear guide for grasping the overall picture of the data and formulating a concrete action plan. Dr.DataScience draws out the true insights hidden behind complex data and powerfully advances the client’s data-driven decision-making.

© Dr.データサイエンス. All Rights Reserved.