Sent to 19,158 subscribers
Huy Nguyen · September 15, 2021
Inner Join is Holistics's weekly business intelligence newsletter. This week: the tension between ad-hoc queries and 'the work', why data scientists shouldn't need to know Kubernetes and thinking with external representations.
The Tension Between Ad-Hoc Queries and The Work
A couple of weeks ago I linked to Benn Stencil's Data's Big Whiff, where I quoted:
The other half of our jobs is doing analysis directly. This work is mostly commonly referred to as ad hoc analysis, though some people call it advanced analytics, or decision science, or just “answering questions.” This is, presumably, what want to do rather than build the dashboards we complain about; we build self-serve tools, we say, so that we can focus on this type of work. Looker sells this promise directly: “Looker helps to streamline processes to save valuable time, freeing up data scientists to focus on the more rewarding aspects of their job.”
We prefer this work in part because it's less tedious than adding the 1,000th filter to a dashboard, and in part because this is the work that actually matters. **Ad hoc analysis is meant to support ad hoc decisions. These decisions are, almost by definition, the most important decisions companies make—they're the ones you only get to make once.* Jeff Bezos' famous one-way doors are the stuff of ad hoc analysis, not a BI report or self-serve dashboard.*
Stencil then goes on to say that operational dashboards aren't the important thing in data; ad-hoc analysis is what is truly important — and yet we don't really have good ways to capture such ad-hoc analysis.
I both agree and disagree with his argument (Must we really have tools for that? What if you put it in a memo, and it gets transferred to the heads of all the leaders? Who says organizational memory has to live inside a BI system?), but that's besides the point. I want to talk about something related instead.
In the day-to-day work of a BI team, there's this interesting tension that arises between ad-hoc questions, as well as 'The Work' — i.e. the daily, operational grind of creating dashboards and maintaining dashboards and making sure your data team is happy and investigating data errors and complaining about BigQuery costs and the rest of that jazz.
Sure, Stencil has a point when he argues that the operational grind isn't as important as the 'bet the company' moments that come from ad-hoc requests, but in certain types of companies (e.g. low margins, or high operational requirements due to regulation), this grind of more mundane dashboard-based metrics may mean the difference between profitability and death. (See Amazon's adapted DMAIC process as one example).
Iteration on the day-to-day metrics really, really matter in this scenarios. And of course, in such companies, and at certain times, the data team can be busier with more 'mundane' dashboard-work that is still rather important to the company.
I'll give you an example.
Let's say that you have a product team focused on improving the on-boarding experience of your app. As part of this, the data team is involved with the product team, and assists product with coming up with better metrics to track. Some of this demands new instrumentation (which means bringing engineering in — plus the data team needs to know what events mean what, because they want to document this in a metadata layer); other parts of this is brainstorming with product people to come up with a shared understanding of important metrics.
Of course, if you use Holistics then your product people can self-serve, assuming the necessary metrics are all available to them in the dataset interface. But the coordination work still needs to happen — product, data, and engineering need to be on the same page. This is the sort of 'day-to-day grind' that is important but that isn't as important as an ad-hoc question (one that may, for instance, change the entire direction of the company).
During times where the 'mundane' workload is high, I've found that data teams in general get more irritable with serving ad-hoc requests. Sure, 'self-service' is part of the answer, but some amount of coordination work still needs to happen when you're iterating on operational metrics. That takes up the data team's time.
So one thing that I've been mulling over is: how do you balance between ad-hoc metrics and more 'day-to-day' operational BI stuff?
I don't yet have the answers (and I'll tell you if we do!), but I really appreciate Stencil's essay because it points to the importance of ad-hoc metrics in the grand scheme of things.
It's easy to get lost in the day-to-day grind of business intelligence. Appreciating the importance of ad-hoc analysis is something to keep in mind.
Insights From Elsewhere
Why Data Scientists Shouldn't Need to Know Kubernetes — Chip Huyen, with a really thoughtful, really well-argued piece on why data scientists shouldn't need to be full-stack:
This post is to argue that while it's good for data scientists to own the entire stack, they can do so without having to know K8s if they leverage a good infrastructure abstraction tool that allows them to focus on actual data science instead of getting YAML files to work.
The post starts with a hypothesis that the expectation for full-stack data scientists comes from the fact that their development and production environments are vastly different. It continues to discuss two steps of the solutions to bridge the gap between these two environments: the first step is containerization and the second step is infrastructure abstraction.
The essay then goes on to say that infrastructure abstraction is relatively new, and then walks you through the various workflow orchestration tools that currently exist on the market.
Thinking with External Representations — this is a more cognitive science thing than I normally link to, but I thought it was pretty good.
Why do people create extra representations to help them make sense of situations, diagrams, illustrations, instructions and problems? The obvious explanation—external representations save internal memory and computation—is only part of the story. I discuss seven ways external representations enhance cognitive power.
Worth thinking about when you write a memo or a Jupyter notebook or a dashboard for others to use.
Building the Analytics Team at Wish — An oldie but goodie from back in 2018: Samson Hu has a wonderful series of posts on building the analytics team at Wish — a low margin business if there ever was one — that's well worth a read.
Enjoy.
Want to write for us? We're looking to hire a few part-time Analytics Content Writers here at Holistics. Apply here, or reach out and let me know. :-)
That's it for this week! If you enjoyed this newsletter, I'd be very appreciative if you forwarded it to a friend. And if you have any feedback for me, hit the reply button and shoot me an email — I'm always happy to hear from readers.
As always, I wish you a good week ahead,
Warmly,
Huy,
Co-founder, Holistics.
PS: If you've not seen it already, we've got a guidebook to bring you up to speed on the ins-and-outs of a contemporary analytics stack — get your copy here.
Get Inner Join in your inbox
A business intelligence newsletter for data practitioners — one considered read, straight to your inbox.