No statistical test can fix a dataset with duplicate entries, impossible values, or inconsistent coding. Clean the data first. It's faster than debugging strange results later.

Check for duplicates

Duplicate participant IDs or repeated rows silently inflate your sample size. Sort and scan before anything else.

Look for impossible values

An age of 200, a blood pressure of 0: run frequencies and descriptives on every variable before you analyse anything, and query anything outside a plausible range.

Standardise coding

"M", "Male", and "1" in the same column will silently break a categorical analysis. Recode consistently and document your coding scheme.

Handle missing data deliberately

Decide, and state, how you'll handle missing values before you run a single test, not after you notice your sample size has quietly shrunk.