Reliability is whether your instrument measures consistently, and validity is whether it measures what it claims to measure. For a questionnaire-based PhD, examiners usually expect Cronbach’s alpha or composite reliability of at least 0.70, average variance extracted (AVE) of at least 0.50, and evidence of discriminant validity such as HTMT values below 0.85 or 0.90, all calculated from your own data.
This guide explains each measure, the thresholds and where they come from, what to do when a value falls short, and how to report the results in a thesis. It is written for survey-based studies in management, commerce, education, psychology and social sciences, with a short note on qualitative and experimental work. Where the instrument section fits in the methodology chapter is covered in our methodology guide.
Last reviewed September 2026.
What is the difference between reliability and validity?
A bathroom scale that always shows 2 kg too much is reliable (it gives the same wrong answer every time) but not valid. A scale that gives a different reading each time you step on it cannot be valid, because it is not even consistent. That is the relationship: reliability is necessary for validity, but not enough on its own.
In a questionnaire study, reliability asks whether the items of each scale measure consistently. Validity asks whether the scale measures the construct you named, such as “job satisfaction”, rather than something else, such as general mood.
| Concept | Question it answers | Usual evidence in a thesis |
|---|---|---|
| Internal consistency reliability | Do the items of one scale hang together? | Cronbach’s alpha, McDonald’s omega, composite reliability |
| Test-retest reliability | Does the instrument give similar scores when repeated? | Correlation or intraclass correlation (ICC) between two administrations |
| Inter-rater reliability | Do two coders or raters agree? | Cohen’s kappa, ICC, percentage agreement |
| Content validity | Do the items cover the construct, and only the construct? | Expert review; content validity index (CVI) or ratio (CVR) |
| Face validity | Do respondents understand the items as intended? | Pilot feedback, cognitive interviews |
| Convergent validity | Do items meant to measure one construct share enough variance? | Factor loadings, average variance extracted (AVE) |
| Discriminant validity | Is each construct distinct from the others? | HTMT ratio, Fornell–Larcker criterion, cross-loadings |
| Criterion validity | Does the score relate to an outcome it should predict or match? | Correlation with an established measure or later outcome |
Most survey theses need internal consistency, content validity, convergent validity and discriminant validity. Test-retest and criterion validity are needed mainly when you are developing a new instrument or using one in a very different setting.
Cronbach’s alpha or composite reliability: which should you report?
Cronbach’s alpha (Cronbach, 1951) is the measure every examiner knows. It is easy to calculate in SPSS (Analyze, Scale, Reliability Analysis). Its weakness is an assumption that every item relates equally strongly to the construct. When items differ, as they usually do, alpha tends to underestimate reliability. It also rises simply because a scale has more items.
Composite reliability (CR) uses the actual loading of each item, so it does not need that assumption. If you are doing SEM, report CR alongside alpha. SmartPLS also reports rho_A, which usually falls between the two.
McDonald’s omega is the equivalent of CR for studies that do not use SEM. JASP, jamovi and the R package psych all calculate it. Many methodologists now prefer it to alpha, but most Indian examiners will still expect to see alpha too. Report both.
What if alpha is below 0.70?
- Check reverse-coded items first. A negatively worded item that was not recoded before analysis is the single most common reason for a low alpha. SPSS shows it as a negative corrected item-total correlation.
- Look at “alpha if item deleted”. If removing one item raises alpha noticeably, read that item again. Was it understood? Was it translated well?
- Consider the scale length. Two- and three-item scales often have alphas around 0.60 even when they work well. For two-item scales, the Spearman-Brown coefficient is a better measure.
- Do not delete items just to reach 0.70. Each deletion narrows what the scale measures. Deleting two items from a four-item scale is hard to defend at a viva.
What thresholds do examiners expect?
The numbers below are what most reviewers and examiners in management and social science expect. They are conventions from the methodological literature, not laws. Cite where each comes from, so an examiner can see you did not invent it.
| Measure | Commonly cited threshold | Source usually cited | Notes |
|---|---|---|---|
| Cronbach’s alpha | ≥ 0.70 (≥ 0.60 sometimes accepted in exploratory work) | Nunnally (1978); Hair et al. | Rises with the number of items; above about 0.95 suggests redundant items |
| Composite reliability (CR) | ≥ 0.70; values above 0.95 are a warning | Hair et al. (2019) | Preferred to alpha in SEM because it weights items by their loadings |
| Indicator loading | ≥ 0.708 (0.40–0.70: consider deleting only if it improves CR or AVE) | Hair et al. (2019) | A loading of 0.708 means the construct explains about half the item’s variance |
| Average variance extracted (AVE) | ≥ 0.50 | Fornell and Larcker (1981) | Below 0.50, more error than construct variance on average |
| HTMT | < 0.85 (conservative) or < 0.90 for conceptually close constructs | Henseler, Ringle and Sarstedt (2015) | Now preferred to the Fornell–Larcker criterion |
| Fornell–Larcker criterion | √AVE of each construct greater than its correlations with other constructs | Fornell and Larcker (1981) | Still reported, but shown to miss many discriminant validity problems |
| CFA fit (CB-SEM) | CFI and TLI ≥ 0.95, RMSEA ≤ 0.06, SRMR ≤ 0.08 | Hu and Bentler (1999) | Many theses use CFI ≥ 0.90 as “acceptable”; say which rule you follow |
Two cautions. First, thresholds differ slightly between textbooks. That is fine; choose one source, cite it, and apply it consistently. Second, meeting every threshold does not prove the measurement is good. A scale can pass every test and still measure the wrong thing if the items were poorly chosen in the first place. That is why content validity comes first.
How do you establish convergent and discriminant validity?
Convergent validity
Convergent validity asks whether the items of one construct share a large part of their variance. You check it through loadings and AVE. An AVE of 0.50 means that, on average, the construct explains half the variance in its items.
If AVE falls short, look at the lowest-loading items. An item loading below 0.40 is usually removed. Items between 0.40 and 0.70 are removed only if doing so raises AVE or CR above the threshold, and only if the construct’s meaning survives. Record each deletion and the reason.
Discriminant validity
Discriminant validity asks whether constructs that are meant to be different are, in your data, actually different. Many theses report only the Fornell–Larcker criterion: the square root of each construct’s AVE should exceed its correlations with every other construct. Henseler, Ringle and Sarstedt (2015) showed by simulation that this test often fails to detect real problems, and proposed the heterotrait-monotrait ratio (HTMT). SmartPLS reports HTMT by default; for AMOS, you can calculate it from the item correlations or use a published plug-in or spreadsheet.
If HTMT is above 0.90 for two constructs, your respondents may not distinguish them. Common causes are overlapping items, or two constructs that are genuinely very close (for example, “satisfaction” and “loyalty” in some service studies). Options include removing overlapping items, merging the constructs with a theoretical justification, or reporting the limitation honestly.
Where CFA fits
If you use covariance-based SEM, a confirmatory factor analysis of the measurement model comes before the structural model. Report the fit indices, and if you add error covariances to improve fit, justify each one theoretically. Adding them one after another until fit improves is a practice examiners increasingly question. Our AMOS vs SmartPLS guide explains how the two approaches assess measurement differently.
How do you show content validity?
Content validity is established before data collection, by showing that the items cover the construct properly. For a scale adapted from the literature, the usual steps are:
- Define each construct clearly, citing the source definition.
- Give the draft questionnaire to a panel of experts, commonly five to ten: subject faculty, and one or two practitioners from the population you are studying.
- Ask them to rate each item’s relevance, often on a four-point scale, and to comment on clarity.
- Calculate an item-level content validity index (the proportion of experts rating the item 3 or 4) if your field expects numbers. A common guide, from Polit and Beck’s work, is an I-CVI of at least 0.78 with six or more experts. Lawshe’s content validity ratio is an alternative, with critical values that depend on panel size.
- Revise, and record what changed.
The pilot study then checks face validity and gives you a first reliability estimate. Designing items and running the pilot are covered in the questionnaire design guide.
Qualitative and experimental studies
Qualitative research uses different criteria. Lincoln and Guba’s trustworthiness framework (credibility, transferability, dependability, confirmability) is the one most examiners recognise. Describe what you actually did: member checking, an audit trail, peer debriefing, reflexive notes. In experimental and engineering work, reliability means repeatability and calibration; report replicates, standard deviations and instrument calibration rather than alpha.
What should you report in the thesis?
Report reliability and validity from your main sample, in the chapter where you present the measurement results. A statement such as “the scale has been validated by Smith (2010), α = 0.89” is not enough; that was Smith’s sample, in another country, perhaps in another language.
- A table with, for each construct: number of items retained, alpha, CR (and rho_A if using SmartPLS), AVE.
- A table of loadings for every item, with any deleted items listed and the reason for deletion.
- A discriminant validity table: HTMT values, and Fornell–Larcker if your examiners expect it.
- For CB-SEM, the fit indices of the measurement model before and after any changes, with each change justified.
- The source of each scale and the reliability reported by its original authors, for comparison.
- How content validity was established: number and type of experts, what they changed.
- The pilot study: sample size, reliability from the pilot, and changes made as a result.
Put the reliability and validity results before any hypothesis testing. Examiners read in that order: if the measurement is weak, the tests that follow mean little. Once the measurement is sound, the statistical test guide helps you choose what comes next.
Sources
- Henseler, J., Ringle, C. M. and Sarstedt, M. (2015). A new criterion for assessing discriminant validity in variance-based structural equation modeling. Journal of the Academy of Marketing Science 43(1), 115–135
- Fornell, C. and Larcker, D. F. (1981). Evaluating structural equation models with unobservable variables and measurement error. Journal of Marketing Research 18(1), 39–50
- Hair, J. F., Risher, J. J., Sarstedt, M. and Ringle, C. M. (2019). When to use and how to report the results of PLS-SEM. European Business Review 31(1), 2–24
- Hu, L. and Bentler, P. M. (1999). Cutoff criteria for fit indexes in covariance structure analysis. Structural Equation Modeling 6(1), 1–55
- Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika 16(3), 297–334
FAQ
Questions scholars ask
Is Cronbach’s alpha of 0.65 acceptable?
It is below the usual 0.70 convention. Some authors accept 0.60 for exploratory research or very short scales. Check reverse-coded items first, report the value honestly, and discuss it as a limitation if you keep the scale.
Do I need to test validity if I used a standard, published scale?
Yes. Validity belongs to how a scale performs in a particular population, not to the scale itself. A scale validated in the US may behave differently among Indian respondents, especially after translation.
Should I report reliability from the pilot or the main study?
Both. The pilot value shows the instrument was ready for use; the main-study value is the one your conclusions rest on.
My AVE is 0.46 but CR is 0.82. Is that a problem?
Many papers cite Fornell and Larcker (1981) as allowing convergent validity when AVE is below 0.50 but CR is above 0.60, though methodologists disagree on whether the original paper supports that reading. It can be argued if you state it openly, but expect a question about it.
Can I delete items to improve AVE?
You can delete weak items, but each deletion must be justified and reported, and the construct must still be adequately covered. Many examiners question scales reduced to two items.
