How Data Leaders Evaluate BI Tools: BI Requirements You Might Miss (A Report)

We analyzed 5,490 signals from BI buyers and customers: what decides the deal (pricing, AI trust, embedding, RLS, semantic-layer lock-in) and what breaks after you sign.

October 01, 2026 · 17 min read · Huy Nguyen
How Data Leaders Evaluate BI Tools: BI Requirements You Might Miss (A Report)

Choosing a BI tool commits a data team for two years or more, so you don't want to make that decision gets made on the strength of a 45-minute demo. The demo covers what the vendor chooses to show, so the problems teams run into later usually sit in the parts the demo skipped.

We analyzed 5,490 real signals from BI buyers and customers: 2,090 tagged signals and 894 direct questions from evaluation calls, plus 2,506 support tickets, all from 2025–2026. Three patterns hold across all of it. Buyers open with pricing and AI, but the questions that decide the purchase cluster in areas where they rarely get straight answers: embedding, row-level security, semantic-layer portability, and the true cost of querying a warehouse live. The share of evaluations that raised AI rose from about 38% in 2025 to about 65% in 2026, and the deciding question is now "how do I know the AI isn't confidently wrong?" The problems that show up in support after purchase line up closely with the questions buyers skipped during evaluation.

This post combines three views of the same buyers: what they ask when evaluating a BI tool, how they are adopting AI inside analytics, and what breaks after they buy. All of it comes from our own record of sales and support conversations at Holistics, which makes it a record of what people did and asked rather than a survey of what they say they would do.

What this measures: A "signal" is a distinct data point extracted from a real customer conversation: a tagged market signal from a sales call, a direct buyer question, or a support ticket. An "AI signal" is a distinct instance of a buyer raising AI during an evaluation, whether as a requirement, an objection, a question, or a workaround they described.

Methodology: The corpus is 5,490 signals: 2,090 tagged market signals and 894 direct buyer questions extracted from recorded sales and onboarding calls (2025–2026) across ~120 accounts, plus 2,506 support interactions (1,662 Zendesk tickets, 2025–2026; 844 Slack Connect threads, 2025–2026), read alongside 280 deal and customer profiles (2025–2026). Signals were tagged by intent and clustered; support tickets were theme-tallied over the full title set with ~110 read in depth; AI signals were counted per distinct source. Counts are approximate and themes overlap (a signal can touch two themes). All companies and individuals are anonymized to industry descriptors; competitor, warehouse, and LLM-vendor names are retained as third parties. Quotes are paraphrased.

Conflict of interest and sample bias. Holistics sells a BI tool built on a semantic layer, with AI that answers from that layer, and every signal in this post comes from our own sales and support conversations. That creates a conflict of interest, because one of the central findings (that buyers trust AI more when it answers from a governed semantic layer) matches what we sell. The sample carries several biases that follow from where it comes from:

  • Who the buyers are. Most of our customers are computer software companies, so questions about embedding, multi-tenancy, and per-tenant security show up far more often here than they would for retail, healthcare, or financial-services teams.
  • Why they were talking to us. These are teams that chose to evaluate Holistics, which means they were already more interested than average in semantic layers, BI-as-code, and embedded analytics, and our own team led many of the calls where those topics came up.
  • What the tickets describe. The support tickets are about our product, so the ranking of what breaks after purchase reflects Holistics customers' experience rather than BI tools in general.
  • How much data there is. The window is two years (2025–2026) and about 120 accounts, so small shifts in a few deals can move the percentages noticeably.

For those reasons, the findings are for reference purposes only. They are most useful as a list of questions to check in your own evaluation, and they should be read as one vendor's view of its own pipeline rather than a measure of the BI market.

If you lead a data team and want to know what other teams ask before they buy, and what they wish they had asked, the rest of this post walks through it.

What buyers ask about, ranked

Money and AI take up the most airtime, but the questions that decide the purchase come from governance and architecture further down the list.

Bar chart of the 14 themes BI buyers ask about, ranked by approximate question count out of 894, grouped into money (~230), trust (~220) and lock-in (~210)

Rank What buyers ask about Approx. questions Where it comes up
1 Pricing, licensing & packaging ~160 Every deal, every stage
2 AI & natural-language querying ~140 The most discussed product topic
3 Visualization & dashboard customization ~90 Demo stage
4 Embedding & multi-tenancy ~90 SaaS vendors reselling analytics
5 Row-level security, permissions & governance ~80 Technical gate
6 Semantic layer & modeling architecture ~80 Where deals are won or lost
7 Caching, performance & warehouse query cost ~70 Sharpest technical objection
8 Data warehouse connectivity & sources ~60 Qualification
9 Data modeling mechanics & metric authoring ~60 Onboarding
10 MCP / AI-agent / analytics-as-code ~50 Technical buyers
11 Security, compliance & data residency ~50 Legal review
12 Migration & switching cost ~40 Incumbent replacement
13 Onboarding, support & time-to-value ~40 Late-stage
14 Admin & usage monitoring ~30 Post-sale

One question can touch more than one theme, so the counts in this table add up to more than the 894 questions in the dataset.

The seven questions behind most evaluations

1. Pricing (~160 questions). Buyers spend more energy working out the pricing model than negotiating the number: per seat, per consumption, or flat, and whether an inactive user still costs money. They ask whether embedded pricing is separate from internal-user pricing, what the security tier adds and whether it raises the price of internal users too, and whether licenses can be added or dropped mid-contract with a cap on annual increases.

2. AI & natural-language querying (~140 questions). These questions come in two halves: "can a non-technical user just ask?" and "can I trust and govern the answer?" The trust section below covers both in detail.

3. Embedding & multi-tenancy (~90 questions). Buyers want to know whether embedding works by iframe or API, and whether external users are provisioned first or passed in the embed token at load time. They ask whether a change to a dashboard shared by 100 tenants means one edit or 100, how embedded users are counted for billing, and whether the embed and its AI chat can be white-labeled and scoped per customer. (For how the main options handle this, see our best embedded analytics tools.)

4. Row-level security & governance (~80 questions). The questions here are about whether the user identity is passed from the buyer's app and persists for the whole session, whether PII columns can be masked by user attribute with one global definition, and whether business-user work can go through review before it becomes official.

5. Semantic layer & modeling architecture (~80 questions). Buyers ask where the semantic layer should live (in the BI tool or the warehouse), how the vendor's modeling language differs from dbt, Cube, or LookML and why it is proprietary, and how hard it would be to leave in three years if their logic lives in that language.

6. Caching, performance & warehouse cost (~70 questions). The questions cover live query versus cache and whether cache duration can be set per report, whether every filter change and first external viewer fires a new billable warehouse query, what caching adds to the warehouse bill, and how the tool holds up at 300–1,000 concurrent users.

7. Analytics-as-code & MCP (~50 questions). Technical buyers ask whether there is an MCP server so they can drive the tool from Claude or a terminal, and whether everything is Git-backed so content can move between dev and prod by copying it.

The money questions: roughly a quarter of everything asked

Pricing was the most common topic among buyer questions, at around 160 of the 894 questions, and the warehouse-cost questions (~70) add to it. Together, money accounts for close to a quarter of every question data teams ask.

What buyers pressed on was the model, more than the number:

  • "Is this per seat, per query, or flat, and does a user who logs in twice a year cost the same as a daily builder?"
  • "If I embed analytics for my customers, is that priced like my internal users? What happens at 5,000 external accounts?"
  • "Does every filter change and every first external view fire a new billable warehouse query?"

Why it's hard: a tool that looks cheap on today's headcount can cost much more once view-only users, embedded tenants, and warehouse queries are counted. Pricing that passes per-query warehouse cost through can still work out cheaper for many teams, but only after they have modeled their own query volume, which a demo never does.

How to get it answered: the most useful request is a written quote for two scenarios, one where most users are view-only and one at your realistic embedded scale, before any feature discussion. A second check is whether cache duration can be set per report, so a dashboard that changes monthly stops re-querying the warehouse every hour. A vendor that can price both scenarios during the call has usually done it for other customers.

Some cost questions went unanswered even on the call, and vendor pages rarely address them: what caching adds to warehouse egress and ingress for a comparable client in actual dollars, and whether builder refreshes (in addition to filter changes and first external views) trigger billable queries.

The trust questions: about a quarter, with AI growing fast

This group is AI (~140 questions) plus row-level security and governance (~80). It is where demos look most impressive and where buyers most often stall, and its AI half nearly doubled its share of evaluations between 2025 and 2026.

How fast AI became an evaluation criterion

The shift shows up even inside a two-year window. In 2025, AI came up in about 38% of evaluations (roughly 12 of 32 archived deals), and in 2026 in about 65% (roughly 92 of 142). AI came up in about 140 of the 894 buyer questions (roughly 16%), second only to pricing. One buyer said AI now "takes up half a demo." Since every tool now has AI, its presence tells a buyer very little, and the real difference between tools is whether the answers can be trusted. (For a side-by-side look at how tools compare on this, see our best AI analytics tools.)

Bar chart showing AI raised in about 38% of BI evaluations in 2025 (12 of 32 deals) and about 65% in 2026 (92 of 142 deals)

What data teams want from AI

The asks fall into four groups, ordered by how many deals raised them:

  1. Natural-language self-serve for non-technical users (~50 deals). A director or frontline lead asks a question in plain language and gets a correct answer without an analyst in the loop.
  2. Bring-your-own-agent and MCP workflows (~40 deals). Technical buyers who already work inside an AI coding agent often raise this by name. One said, "Several of our people work in Claude all day, so an MCP server is a must-have."
  3. Building charts and dashboards by prompt (~30 deals).
  4. AI that maintains and authors the models and dashboards (~24 deals).

What stops them

Buyers in very different industries described the same fear, an AI that is confidently wrong:

  • "AI is very good at high-confidence nonsense, and that's exactly what we need to avoid."
  • "An inaccurate AI answer is worse than no AI answer."

The specific failures they named are the ones teams run into in practice. Hallucinated answers came up in about 24 deals, and inconsistency (two different answers to the same question) came up repeatedly. In about 16 deals the concern was that the AI doesn't know the business's own definitions; one team watched an AI compute average daily rate with cancelled orders wrongly counted in the denominator. In about 20 deals buyers tested whether the AI respects row-level security and embed-token scope, because a tool that ignores them can show one customer another customer's data, and several said they explicitly did not want an open, client-facing chatbot. About 12 deals reacted strongly and positively to being able to inspect the AI-generated query as plain SQL.

Chart of what data teams want from AI, what stops them, and what they now require, by approximate number of deals

What earns trust

The teams that got comfortable adopting AI had one thing in common, and it showed up in 50 to 60 deals: they trusted AI far more when it answered from a governed semantic layer than when it wrote SQL from scratch against raw tables. (This is the finding most affected by the conflict of interest described above, since these buyers came to a semantic-layer vendor and heard our team explain the approach.) In that setup the AI expresses intent, and the semantic layer writes the actual query against metrics the team has already defined and verified. One buyer described it this way: "It's almost as if the AI is a human user building a report, rather than looking at the database and making assumptions."

This approach has a real cost, and buyers raised it too. The semantic layer itself becomes work, and teams whose metrics change often worried that "we could spend all our time maintaining the layer instead of analyzing." A governed layer earns trust, but a team should budget for the time it takes to maintain one.

How to test it in the demo: three checks cover most of the risk. The same question asked twice should return the same answer. The tool should show the query it generated as plain SQL you can read. A restricted user asking a broad question should only see the rows that user is allowed to see.

Two more findings belong in the requirements doc. Bring-your-own-model is now a standard ask, appearing in about 48–56 deals compared with about 30–36 that raised AI cost. Regulated and embedded teams want to use their own Claude or Gemini key so their data doesn't route through a model they didn't choose, and "it's OpenAI-only?" has become an objection on its own. Flat or included AI pricing also beats metered pricing, and opaque per-credit AI pricing cost at least one vendor a deal outright. Both questions are worth asking before the demo moves on.

Teams also aren't waiting for a tool to arrive. In about 24 deals, teams described improvising before they bought anything: pasting their schema and CSV exports into ChatGPT or Claude, standing up their own MCP servers to query the warehouse in plain language, and then checking every answer by hand because they don't fully trust it yet. That gap between wanting self-serve AI and double-checking a chatbot is the clearest sign of how much demand there is.

The lock-in questions: the quarter that predicts regret in year two

This group is the semantic layer (~80 questions), embedding and multi-tenancy (~90), and migration (~40). It gets the least time in a demo and has the most influence on whether a team is still happy with the tool in year two.

The questions that decided deals:

  • "Where does the modeling logic live: in the tool, or in my warehouse where I control it?"
  • "If a hundred of my customers share one dashboard and I change it, do I edit it once or a hundred times?"
  • "Is the user's identity passed from my app, and does it hold for a whole session, so a customer only ever sees their own rows?"

Buyers asked one question almost word for word whenever a proprietary modeling language came up, and it was the most emotionally loaded question in the dataset:

"If we adopt your proprietary language and later want to leave, how hard is that?"

Several related questions went unanswered on the call and on vendor pages. Buyers asked whether they could expose their semantic layer through MCP to other tools and agents or whether it stays locked inside the BI tool, and how reliably an AI coding agent can write a proprietary query language it has barely seen in training. They also raised modeling cases that generic documentation skips: defining a metric differently per line of business while still rolling up to one governed global definition, and onboarding new tenants when SOC 2 requires a separate warehouse dataset per tenant, without rebuilding models each time.

Why it's hard: every answer in this section describes a switching cost the buyer is signing up for, and vendors have an incentive to keep switching costs vague. Answers like "it's flexible" or "most customers stay" usually mean the question hasn't been answered.

How to get it answered: the exit question is worth asking early, and how directly the vendor answers matters more than the claim itself. Two follow-ups help: whether your model definitions live in a Git repo you own, and whether the vendor can show a real embed with row-level security enforced, live on their own data, during the call.

What breaks after you sign: the 2,506-ticket view

The evaluation questions show what buyers asked. This section covers what customers filed for help with afterward: 2,506 support interactions (Zendesk tickets and Slack threads) from 2025–2026, all from Holistics customers. About 40% of the subject lines carry a failure word such as "error," "broken," or "not working," so a large share of what teams file is break/fix, and the ranking of what breaks is useful on its own.

By share of support volume, the top categories are data modeling and semantic-layer logic (~20–24%), visualization and dashboard behavior (~18–20%), account, billing and compliance (~15%), scheduling and delivery (~9–11%), and permissions and access control (~9–10%). Further down are BI-as-code publishing and Git (~7–8%), query performance and warehouse cost (~5–7%), connectivity (~5%), embedding (~4–5%), and AI features (~2–3%).

The most useful finding is the overlap. The tickets teams file after buying are things they could have tested during evaluation:

  • Warehouse cost arrives as a bill rather than an error. Teams find out months after signing that the generated SQL scans more than they expected (fan-out joins, no aggregate awareness) when the warehouse invoice climbs. This belongs in the POC.
  • Row-level security passes the demo and then breaks on the shape of a real dataset. In one case a misconfigured RLS dataset left dashboards showing every row to everyone, a live data exposure that a demo would never surface.
  • Connectivity comes down to yes-or-no facts that a checklist misses: a specific auth method, identifier case-sensitivity, private-key passphrase support. These are worth confirming before signing.
  • Ownership when someone leaves is the most frequently repeated request in the support data: what happens to dashboards and schedules when their owner leaves the company. Teams rarely evaluate it, and most of them run into it.
  • Export fidelity and AI data policy both tend to be discovered after go-live: whether exported numbers match what is on screen, and where AI prompts go and how long they are retained.

The mix of tickets in 2025–2026 also says something about where BI work has moved. Data modeling is the top technical driver, and BI-as-code (Git, publish, deploy) plus AI and MCP together account for roughly a tenth of all tickets, including questions about AI access control and data retention that a BI support queue would rarely have seen a few years ago. Teams evaluating a tool with an as-code semantic layer and built-in AI should expect their own support questions to look like this mix.

Six questions to bring to your next vendor call

The whole dataset reduces to a short list of questions that predict satisfaction better than a feature comparison:

  1. What's my bill when most users are view-only, and what happens at embedded scale?
  2. Can I set cache duration per report, and does every filter fire a billable warehouse query?
  3. Does the AI give the same answer twice, show its SQL, and respect a restricted user's permissions?
  4. Can I bring my own model and key, and how is AI priced?
  5. If I build in your semantic layer and later leave, what does that migration cost me?
  6. Can you show a live embed with row-level security enforced on your own data?

The six questions to bring to a BI vendor call, grouped into money, trust and lock-in

Any buyer can ask these without special access. What they require is staying on an uncomfortable topic for a few more minutes when the vendor would rather move on to the next chart. Teams that ask them tend to choose better, and teams that skip them usually find the gaps after signing, once the tool is in daily use.

We pulled these findings from our own sales and support records because analyst surveys and listicles don't capture what people ask when real budgets and real data are involved. The full report has all fourteen question themes with examples, the AI adoption data in more depth, and the full support analysis.


Findings drawn from Holistics' first-party record of real BI buyers and customers: 2,090 tagged signals and 894 direct questions from evaluation calls (2025–2026), across ~120 accounts with recorded calls, plus 2,506 support interactions (2025–2026), read alongside 280 deal and customer profiles. Most of our customers are computer software companies. Companies and individuals are anonymized; figures are aggregate, and shares are approximate because a single signal can touch more than one theme. Because the data comes from Holistics' own pipeline, the findings are for reference purposes only.