Skip to main content

Command Palette

Search for a command to run...

Pandas for Beginners: Stop Using dropna() on Every Missing Value

Updated
•2 min read•View as Markdown
Pandas for Beginners: Stop Using dropna() on Every Missing Value
W
Engineering student exploring Data, AI, Python, SQL, and Information Systems. I share practical tutorials, projects, lessons learned, and resources to help students and beginners build real technical skills.

When starting out with data analysis in Python, missing data (NaN or None) is usually the first roadblock you hit. The temptation is to wipe out any row with empty cells using .dropna().

In a small practice project, that might seem fine. In a real-world dataset or an interview assessment, indiscriminately deleting rows will distort your statistical results and destroy critical data.

Here is the quick breakdown of when to drop vs. when to fill missing values using Pandas.

  1. When to Drop Rows (dropna): Only drop rows when the missing column is the unique identifier (like an ID) or the target variable you are trying to predict.
import pandas as pd

# Drops rows ONLY if the 'user_id' column is missing
df = df.dropna(subset=['user_id'])

2. When to Impute Values (fillna): If a numerical column (such as salary or age) has scattered missing values, filling them with the median is often safer than wiping out entire rows:

# Calculate the median of the column
median_salary = df['salary'].median()

# Fill the missing values in place
df['salary'] = df['salary'].fillna(median_salary)

3. Handling Text Data: For categorical or text columns, replace missing entries with an explicit "Unknown" category so downstream filters or models still capture the row:

df['department'] = df['department'].fillna("Unknown")

What to Build Next

Don't just read code—run it. Open a free notebook in Google Colab, import a public dataset from Kaggle, and inspect your missing values using df.isnull().sum().