Pandas for Beginners: Stop Using dropna() on Every Missing Value

When starting out with data analysis in Python, missing data (NaN or None) is usually the first roadblock you hit. The temptation is to wipe out any row with empty cells using .dropna().
In a small practice project, that might seem fine. In a real-world dataset or an interview assessment, indiscriminately deleting rows will distort your statistical results and destroy critical data.
Here is the quick breakdown of when to drop vs. when to fill missing values using Pandas.
- When to Drop Rows (dropna): Only drop rows when the missing column is the unique identifier (like an ID) or the target variable you are trying to predict.
import pandas as pd
# Drops rows ONLY if the 'user_id' column is missing
df = df.dropna(subset=['user_id'])
2. When to Impute Values (fillna): If a numerical column (such as salary or age) has scattered missing values, filling them with the median is often safer than wiping out entire rows:
# Calculate the median of the column
median_salary = df['salary'].median()
# Fill the missing values in place
df['salary'] = df['salary'].fillna(median_salary)
3. Handling Text Data: For categorical or text columns, replace missing entries with an explicit "Unknown" category so downstream filters or models still capture the row:
df['department'] = df['department'].fillna("Unknown")
What to Build Next
Don't just read code—run it. Open a free notebook in Google Colab, import a public dataset from Kaggle, and inspect your missing values using df.isnull().sum().
