Let’s talk money. Not the fun kind—the “How much is all this AI actually costing us?” kind.
Because as AI moves into everyday workflows, companies are starting to reckon with what it all costs:
"How do we keep costs from creeping up as AI adoption grows?"
"What does a realistic AI budget look like for next quarter, or next year?"
Databricks, a company that knows a thing or two about running AI at scale, recently shared what it’s learned from managing AI spend across its own organization—and there are some good lessons for the rest of us, regardless of our tech stack.
This week, we're breaking down what they learned, plus the tech updates on our radar, and where we hope to see some of you in person in September!
Below you'll find:
Tips to Keep AI Spend Under Control
Tech Insights: Unity AI Gateway, Omnigent, Genie Ontology
Events: AI Agents and the End of Static BI (A8 live stream), Databricks Meetup (Chicago), dbt Summit (Vegas)
What Databricks’ AI Spend Can Teach the Rest of Us
Databricks’ latest webinar Govern AI Spend at Scaleused its own internal AI use cases to share lessons learned on controlling AI spend.
Yes, their budgets may be bigger than most, but the lessons are practical, regardless of your token spend, model, or AI stack:
#1 - Not every task deserves your most powerful model.
This was the most concrete lesson from Databricks’ own experience.
They benchmarked coding agents against real tasks and found that dialing down reasoning effort could significantly reduce costs without materially affecting success rates. The same idea applies to model selection: simple tasks can go to smaller, cheaper models, while more complex work gets the additional horsepower.
#2 - Make AI spend visible to the people creating it.
Databricks initially encouraged employees to experiment with coding agents, but as adoption grew, so did costs.
One of the simplest things they did was make AI consumption visible. They gave users and teams visibility into their own spend and some ownership over managing it. Once engineers could see what their choices were costing, they had an incentive to use AI more thoughtfully.
#3 - Optimize the task, not just the token.
An agent that normally spent around $50 on a task ran up $1000s after unnecessarily reading through a huge set of log files while debugging.
When AI costs climb, look at agent behavior. By analyzing agent traces, Databricks found waste like repeatedly sending the same repository context without caching it. Repeated context, unnecessary reasoning, retries, tool calls, and inefficient agent behavior can all drive up the cost of completing a task.
#4 - Put guardrails in place to avoid unexpected spend.
They experienced how quickly autonomous workflows can create unexpected spend.
Rather than waiting for that spend to show up on a monthly bill, set budgets and thresholds at the user, task, agent, or workflow level, and create a path for additional spend when there’s a legitimate reason for it.
Tech Updates
Our consultants connect biweekly to talk through the tech updates worth watching—here's what came out of this week's session.
Unity AI Gateway is Generally Available
Databricks is turning AI Gateway into the control tower for enterprise AI usage and spend
AI Gateway gives admins centralized visibility and cost controls across users, models, agents, and providers, including budgets and alerts. But one of the more interesting features is the ability to analyze historical usage and estimate where intelligent model routing could have reduced costs.
That gives teams a way to prove the potential ROI before changing how production traffic is routed. And as AI usage spreads across models, agents, coding tools, and MCP servers, Gateway increasingly becomes the governance layer sitting in the middle of all of it.
Omnigent, Databricks' Ultra Harness, is in Hypergrowth Mode
And you don't need Databricks to use it
Omnigent is Databricks' harness-of-harnesses, sitting on top of 13+ coding harnesses like Claude Code and Codex, has moved from v0.2 at Summit to v0.10 already. You can run it completely standalone with your existing Claude or Codex subscription, or connect it to Unity AI Gateway to add Databricks context, governance, and centralized control.
The interesting part is how many different problems it can solve:
→ At the simplest level, it gives you one front door to multiple AI coding tools — intelligent routing instead of constantly switching between them, plus the ability to share context with other users.
→ Connecting it to a Databricks workspace and AI Gateway takes that further, adding shared workspace context and making the orchestration even more powerful.
→ At the enterprise level, Omnigent starts to look less like a developer productivity tool and more like a governed way for data leaders to make multiple models, harnesses, and AI workflows available across teams and vendors.
Two Genie Ontology features worth knowing: OntaRank, which prioritizes answers based on business context and semantics rather than just query matching, and Pages, which lets teams manually curate trusted business context for Genie to draw on.
AI is moving from simply querying your data to actually understanding how your business thinks about it.
Events
Looking for a good excuse to step away from your day-to-day and learn something new? Here are a few events that we think are worth your time.
The Future of Trusted Analytics: AI Agents and the End of Static BI
We had the best time hosting author and data leader Matt Housley and discussing how to adapt your BI solutions in the era of AI Agents. Watch the recording