Collect a four-week baseline before changing architecture. Pull refresh history for certified datasets and count rows where Status = Failed and the app is on an executive dashboard path.
Metrics to log weekly
- Failed refreshes touching exec workspaces (absolute count + % of scheduled runs).
- Analyst hours recovering packs (timecoded: re-run refresh, export CSV, rebuild pivot) after GatewayNotReachable or credential errors.
- Meetings that started with a stale Snapshot or reverted to Excel — note date and deck.
- Mean time to restore after credential rotation or gateway node loss (minutes from first Failed to first Completed).
What you can claim — and what you cannot
Reliability ROI is avoided cost (rework hours × fully loaded rate) plus regained decision time. Do not invent revenue uplift or blanket productivity percentages. Contractual uptime commitments belong in the commercial schedule — not in marketing copy.
Spend order that matches the maths
First kill single points of failure: one personal gateway, one shared mailbox password in Manage gateways. Second: query folding and overtime refreshes (Native Query missing). Third: incremental refresh / capacity. Capacity SKUs without hygiene rarely move the needle in your baseline spreadsheet.