Sent to 18,897 subscribers
Huy Nguyen · August 18, 2021
Inner Join is Holistics's weekly business intelligence newsletter. This week: abductive thinking, correlation vs causation, and what in the world is going on with data catalogs?
Induction, Deduction ... Abduction?
One of the things we do a lot in data is some form of inductive thinking or deductive thinking.
A quick refresher, for those of you who've forgotten what the two terms mean:
- Inductive thinking — You look at a whole mess of data, and synthesise some meaning from it. e.g. "From the geolocations of pings from our app, it seems like we're mostly used by golfers in Hawaii." Wait, what? "Yeah — the majority of the pings seem to be from golf courses, and many of them are from the fairways."
- Deductive thinking — You look at a pattern of data, and work through some set of logically true statements in order to come to a conclusion. "Servers can't send pings when they're down, and we got a ping every minute over the last 24 hours. Therefore our servers weren't down over the last 24 hours. The problem must be something else."
But there's also abductive reasoning, and abductive reasoning is what interests us today.
Abductive reasoning is 'making a best guess inference'. For instance, "our app is mostly used on golf courses, and it seems likely that they're golfers. I think we should try and sell them golf equipment — I bet we'd get good conversions on those ads."
The thinking here isn't really induction — you're not synthesising an explanation from a mass of data, but instead applying a set of logical reasoning steps to suggest a course of action. But it's also not deduction, because your logical reasoning isn't watertight:
- Our app is used by golfers (you think — this is the result from an inductive thinking process earlier!).
- Golfers like to buy golfing equipment.
- Therefore golfers will buy golfing equipment from your app, an app that they mostly look at while outdoors, in the middle of a fairway.
With abduction you're making a best guess. It's not entirely logical, and it's not a watertight case, but it seems worth testing!
In order words, abduction is the kind of thinking that entrepreneurs do.
This explains why we're talking about abduction today. A reductive view of data analysis is that you just deal with induction or deduction — that is, either you synthesise an explanation for data you've observed, or you reason using strong logic. But occasionally you encounter people (mostly from the business end of things?) who spitball — that is, who come up with some weird-ass idea to try based off some data you give them — and they go out and do it, and then they recruit you for some data collection on the tail end.
They're not being illogical. They're not being 'unrigorous' — or at least they're not being unrigorous in a bad way. They're doing abductive reasoning.
Don't discount them! You might be tempted to object, because data folk are rigorous and intelligent and analytical and used to precision. But hopefully abductive thinking gives you another bucket to use when evaluating their actions. And there's good reason to let them do what they do — they're simply using data the way an improv comedian might use a prompt — as a springboard into something possibly good.
Here's the Wikipedia page if you want to read more.
I wrote a LinkedIn post recently where I said:
A lot of job roles are created by mixing elements of two distinct roles together:
- A Product Manager is someone who has more business context than a Software Engineer, and knows more about technical/engineering than a Business Person.
- A Data Scientist is someone who has more business context than a Statistician, and knows more about statistics than a Data Analyst.
So recently in Analytics land, they just invented another role — the Analytics Engineer.
- 😅 An Analytics Engineer is something who has more business context than a Data Engineer, and knows more about software best practices than a Data Analyst.
…
I wonder what else can be created (that is useful, of course)?
This is something that we've talked about before — but I'm looking at the data engineer end of things, and wondering if there are other combinations that we've not thought of.
Something to chew on.
Insights From Elsewhere
When Correlation is Better Than Causation — Brittany Davis over at the Narrator.ai blog on something closely related to today's topic — how 'that's just correlation' isn't necessarily a bad thing, and why it shouldn't be used to shut down data discussions.
What in the World is Going On With Data Catalogs? — Gordon Wong and Barr Moses on how data catalogs are at the stage of their lives where they're trying to 'be everything to everyone'. My biases are clear here: Wong is always worth reading. Every time I talk to him he is 100% focused on delivering business value, is pragmatic in all the right ways, and so when he says that data catalog offerings currently lack a user story, I pay attention.
He's probably right.
Why Observability Requires a Distributed Column Store — This one is for the more data engineering oriented readers out there. Alex Vondrak from Honeycomb.io writes about why they decided to build their own distributed columnar store. The goal: high query performance for the observability work that they do.
The piece goes into a nice level of detail on why and how they came to that conclusion.
That's it for this week! If you enjoyed this newsletter, I'd be very appreciative if you forwarded it to a friend. And if you have any feedback for me, hit the reply button and shoot me an email — I'm always happy to hear from readers.
As always, I wish you a good week ahead,
Warmly,
Huy,
Co-founder, Holistics.
PS: If you've not seen it already, we've got a guidebook to bring you up to speed on the ins-and-outs of a contemporary analytics stack — get your copy here.
Get Inner Join in your inbox
A business intelligence newsletter for data practitioners — one considered read, straight to your inbox.