AVE 0.48. HTMT 0.87. Scalar invariance not supported.
Three red numbers in the results panel, after months of refining items and collecting data. The obvious question is whether the whole thing has to start over.
Falling short of a threshold is not the same as failing. It is also not something you can carry regardless of severity. There is a line between a shortfall that sends you back and a shortfall you can state as a limitation and move past.
This article covers how to locate the source of the shortfall, when each of those two paths applies, and how to write a limitations section that reviewers accept.
Locate the Level Before Judging the Number
The same number calls for entirely different responses depending on where it originates. The starting point is not the failing index but the layer the failure sits in.
An item-level problem shows up as one item with a low loading, or one item that keeps appearing across several diagnostics. Usually the wording is ambiguous or the item asks two things at once. Revising or replacing that item resolves it.
A factor-level problem is different in kind. AVE that stays under 0.5 or HTMT above 0.85 points at the conceptual boundary of the factor rather than at any single item. Two factors you defined as theoretically distinct are being read as one thing by respondents. Swapping items will not fix that; the conceptual design has to be revisited.
A group-level problem is invariance holding for some factors but not others, which is where the partial invariance handling from the previous article applies.
Is AVE Below 0.5 Automatically a Failure?
No. AVE ≥ 0.5 is the Fornell & Larcker (1981) criterion, and a relaxation traced to the same paper is widely used: if AVE falls below 0.5 but composite reliability exceeds 0.6, convergent validity can still be treated as adequate.
One caution about citing it. That relaxation is conventionally attributed to Fornell & Larcker (1981), but there is disagreement about where in the original it appears. If you are submitting to a journal with strict review, check the original (Journal of Marketing Research, 18(1), 39-50) before citing it.
| Index | Threshold | Source |
|---|---|---|
| AVE | ≥ 0.5 | Fornell & Larcker (1981) |
| CR | ≥ 0.7 | Hair et al. |
| HTMT | < 0.85 | Henseler et al. (2015) |
| CFI | ≥ 0.90 | Hu & Bentler (1999) |
| RMSEA | ≤ 0.08 | Browne & Cudeck (1993) |
Reading one number in isolation is the mistake. AVE of 0.48 with CR of 0.82 and items that hang together theoretically is thin grounds for discarding a factor. AVE of 0.49 with CR of 0.61 and uniformly weak loadings looks similar on the first number and is a much worse position.
Go Back, or Carry It Forward?
There are two paths: return to an earlier stage, revise, and rerun what follows; or state the limitation explicitly and proceed with the current results.
Researchers often read the second path as a compromise on quality. It is not. Recognizing a limitation and constraining your conclusions to what the data support is a defensible methodological choice. The riskier move is trimming items until the thresholds clear and losing the construct the scale was built to measure.
The weight of the decision falls on how severe and how widespread the shortfall is. A shortfall confined to one factor and close to the threshold leaves room to carry it as a limitation. Shortfalls across several factors, or well below the cutoff, mean the interpretation is being pushed past what the data hold, and going back is the better call. Remaining time and the feasibility of recollecting data are real inputs too.
Writing Limitations Reviewers Accept
What separates a strong limitations section from a weak one is not whether the fact was stated but whether its effect on the conclusions was addressed.
Reporting the number concretely comes first. "Some indices fell slightly short of the recommended values" reads as concealment. "AVE for Factor 3 was 0.47, below the recommended 0.5" is better. The reviewer is already looking at your table.
Bounding the interpretation comes next. Say what you cannot claim because of the shortfall, before anyone asks. If convergent validity is weak for one factor, state that path coefficients involving that factor should be read with caution, and name which part of the conclusion is affected.
Connecting to future work closes it. Describing how the items could be improved, or which sample would be needed for revalidation, turns the limitation into the next step of a program rather than an admission.
Reviewers do not expect flawless results. They distinguish researchers who know where their results fall short from those who do not.
What modidoc Does at the Strategic Judgment Stage
modidoc's strategic judgment stage collects the below-threshold items surfaced during rigorous validation and group fairness verification, classifies each as an item-, factor-, or group-level issue, and lays the revise-and-return path against the proceed-with-limitations path side by side. It also indicates which path it recommends based on the proportion of failing items and their distance from the threshold. That recommendation line is an operational rule set by modidoc rather than a published standard, and the final call stays with the researcher. This stage is implemented internally as the C7 strategic judgment engine.
Frequently Asked Questions
Can I publish with AVE below 0.5?
Yes. AVE ≥ 0.5 is a recommended criterion from Fornell & Larcker (1981), not an absolute condition. When CR is a stable 0.7 or above and the items are theoretically coherent, stating the shortfall as a limitation and proceeding is common. Shortfalls across multiple factors, or values dropping below 0.4, are grounds for revisiting the factor structure instead.
What should I do when CFA fit is poor?
Not reach for modification indices first. Check whether specific items carry low loadings and whether the problems cluster in one factor, and establish the level of the issue before adjusting anything. A model that only clears the thresholds after several modifications is a model built to fit the data, and the question of how many modifications you applied is one reviewers ask.
How long should a limitations section be?
Structure matters more than length. Three or four sentences carrying the number, the interpretive bound it imposes, and a direction for follow-up work is enough. A long list of limitations that never addresses their effect reads as weak regardless of length.
The next article looks at checking whether structural equation modeling is feasible before the main data collection. It covers model identification, sample size and power, and screening mediation and moderation at the planning stage.
Previous: Measurement Invariance Before Group Comparison