How to Deal With Missing Data: Databases, Zeroes and Imputation

Even in this age of ubiquitous computing where all kinds of data constantly flow around all of us through every conceivable electronic device, knowing everything about everyone all the time is just not possible. Some say that marketers collect more data in one hour than they did in a year in the ’70s. But linking all those data points to a known individual (or even an anonymous match key) is always a challenge due to privacy issues, data ownership or lack of a common key by which data are combined. Statisticians always want more variables for better predictability, but, like in the olden days, modeling still is about “making the best of what we know.”
Then, what to do with the “unknowns”? Do we just dismiss them and move on? Properly treating missing data may boost targeting efficiency as not all missing data are created equal, and missing data often contain interesting stories behind them. For example, certain variables may be missing only for very rich people and very poor people, as their residency may not be as exposed as others. That in itself is a story. Some data may be missing in certain geographic regions or for certain age groups. “Not” having access to broadband may mean something interesting, too.

Filling in the Blanks
Like other targeting challenges, missing-data management starts with proper database design. Even at the data collection stage, reasons why certain data points are missing should not be ignored. If you are dealing with numeric data, such as dollars, frequency counts, dates, etc., why are they missing? Is it because they are really unknown and incalculable (no transaction to deal with), or a simple issue of mismatches among different data marts and sources? Database managers may not always know the actual reasons why they are missing, but they should never blindly fill the missing values with “0”s. Zeros must be reserved for known and verified zeros.

Stephen H. Yu is a world-class database marketer and Associate Principal, Analytics & Insights Practice Lead for eClerx. Stephen has a proven track record in comprehensive strategic planning and tactical execution, effectively bridging the gap between the marketing and technology world with a balanced view obtained from over 28 years of experience in best practices of database marketing. Prior to eClerx, he served as VP, Data Strategy & Analytics at Infogroup, and previously he was the founding CTO of I-Behavior Inc. “As a long-time data player with plenty of battle experiences, I would like to share my thoughts and knowledge that I obtained from being a bridge person between the marketing world and the technology world. In the end, data and analytics are just tools for decision-makers; let’s think about what we should be (or shouldn’t be) doing with them first. And the tools must be wielded properly to meet the goals, so let me share some useful tricks in database design, data refinement process and analytics.” Reach him at

Related Content