Anecdotes and data
Anecdotes by themselves might make for bad decisions, but using them as a hypothesis engine for data analysis can give you a massive superpower
We data people love making fun of anecdotes, and people who rely on them.
“The plural of anecdote is not data”, we say.
We have even come up with a portmanteau “anecdata”, to make fun of people who quote anecdotes in their work in the hope that it substitutes for actual data.
I’ve mostly worked in “data science” in the “big data” era (2011 onwards), and if there is one thing that big data gave us, it is the belief that bigger data means bigger insight.
However, as I have been reflecting upon my ~15-year career in data science recently (no no - I’m not retiring or even pivoting away; I just like to think about the past a lot), I’m coming to believe that anecdotes have been unfairly maligned.
Yes, anecdotes can lead to flawed conclusions. They can lead to all kinds of cognitive biases. They can result in making decisions that are bad for the majority (because you chose a minority anecdote). On their own, they may not be the best way get insights from data. However, the moment you combine anecdotes with robust data analyses, you have yourself a superpower.
Essentially, anecdotes are excellent hypothesis generators!
The Scientific Process
Think about the process you follow when you try to get insight from data. If you are doing it right, it needs to follow something that approximates the “scientific process”. You start with a hypothesis. And then figure out how you can test it using data. Then you gather the said data, clean it, process it, test the hypothesis, and iterate (the great thing with data analysis is that each iteration of analysis gives you a bunch of hypotheses that keep the process going). And then somewhere you have insight.
In fact, this was the entirety of the content of a workshop I had conducted at The Fifth Elephant back in 2012.
All this is all well and good but if you have analysed a lot of datasets (like I have), you will realise that sometimes the sticking point is the initial hypothesis. Without an initial hypothesis, you really don’t know where to begin.
How many times have you sat down with a dataset to “analyse” and then not known where to begin? How many times have you found yourself stuck with a bunch of datasets, not knowing what to do? Even when you look at data cleaning - it’s clearly a context-sensitive process, and without a hypothesis or question to answer, you don’t know how exactly to clean the data.
And this is where anecdotes come in - they are an excellent starting point for your analyses.
Businesspeople propose anecdotes. Data people dispose them.
You would have noticed this if you have worked closely with business teams. Most of the times, they questions they pop at you is driven by some anecdote. Some customer would have gotten angry about something, and the business decides that deserves a data investigation. A CXO would have made some observation “on the assembly line” on a rare visit there, and that warrants another data investigation.
The weekly or monthly business reviews might be a sea of data, but you will notice that the real questions for you will come while discussing some tiny part of the business, rather than from the aggregate data.
All anecdotes. If you act upon these anecdotes, you risk making faulty decisions. However, the moment you treat them as hypotheses to kick off your scientific data analysis processes, you know which of them to take action on, and how.
It doesn’t always require business people to generate hypotheses using anecdotes - if as a data team you are solely relying on that, you may not have sufficient hypotheses. What a good data analytics team needs to do is to also generate its own hypotheses, and looking for anecdotes is a great way to get there.
During downtimes while at Delhivery, I would just pull up a random set of packages and look at all their scans to track their journey towards being delivered. This process led to a whole bunch of interesting hypotheses, some of which led to insights that were impactful at the business level.
At other times, while reviewing models, rather than simply look at aggregate statistics, I would get the team to pull up a small number of examples where the model was failing. Yes, you can argue they were anecdotes, but deeply analysing these few cases that went wrong gave far more insight than simply torturing the aggregate statistics.
Hypotheses in the times of AI
Recently, while building Babbage Insight, one of the challenges in building the Proactive Insight Engine we were building was - where do the hypotheses come from? Building an investigation engine is not that hard. The real challenge with proactive insights is - how good are the hypotheses that you are investigating? In the case of Babbage, we had built out our engine to hypothesise based on aggregate statistics. Now I’m starting to wonder if we missed a trick (or few) by relying on statistics and not getting into the anecdotal territory.
On a broader note, one of my discomforts with analytics copilots is that they inevitably solve for the ability to pull data and do analysis, and the quality of insight is ultimately dependent upon the hypotheses that are being fed into them. And if companies aren’t structuring these copilots in a way that people can bombard them with hypotheses (and anecdotes), then they are not making very good use of these copilots.
As a corollary - one skill that is essential for any “analysis copilot” system is the ability to say “there is nothing spectacular in your hypothesis”. Or course, the challenge with building this in is that the initial product sale might become harder!


